Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
# Partial Morpheme Health Check

## Goal

Extend `GrammarHealthChecker` on pull request #475 to identify every partially
analyzed HermitCrab morpheme. The finding should tell grammar authors to finish
the incomplete analysis and explain that partial morphemes can broaden search
and prevent safe final-template pruning.

This is a production-readiness diagnostic. It does not change grammar loading,
parsing, synthesis, or the conservative final-template correctness guard.

## Diagnostic contract

Add the stable code `hc-partial-morpheme` with warning severity. Emit one
finding for each distinct `Morpheme` whose `IsPartial` property is true.

The check covers:

- lexical entries in every stratum;
- ordinary morphemic morphological rules in every stratum; and
- morphemic rules referenced by affix-template slots.

The same rule object can be referenced more than once, so enumeration must use
reference identity and report it once. Each finding's first and only subject is
the partial `Morpheme`, allowing a host to navigate to the original object.

The message identifies whether the subject is a lexical entry or morphological
rule, names it using the best available identifier, and recommends supplying
its missing category or template/slot analysis. It also states that leaving the
morpheme partial can broaden analysis and disable safe final-template pruning.

## Placement

`GrammarHealthChecker.Check(Language)` will invoke a private partial-morpheme
check alongside the two existing checks. The implementation stays inside the
`netstandard2.0` HermitCrab library and remains diagnostic-only.

No new parser option or model field is introduced. `Morpheme.IsPartial` remains
the owner of the decision; the health checker reports that published fact and
does not re-derive partiality from POS, slots, or feature structures.

## Tests

Tests will be written and observed failing before production code changes. They
will prove that:

1. a partial lexical entry produces one actionable warning and exposes the
entry as its subject;
2. a partial ordinary affix rule produces one warning;
3. a partial template rule produces one warning even if referenced by multiple
slots or templates;
4. non-partial morphemes produce no partial-morpheme warning; and
5. the existing checks continue to compose with the new check.

The targeted HermitCrab suite and formatting check must pass before the branch
is pushed.

## Relationship to pull request #491

Pull request #491 keeps its safe default: final-template pruning remains
disabled wherever partial morphemes make the stronger conclusion unsafe. The
new health finding gives grammar authors an actionable route to remove that
performance blocker instead of weakening the correctness guard or silently
forcing the optimization.
325 changes: 325 additions & 0 deletions src/SIL.Machine.Morphology.HermitCrab/GrammarHealthChecker.cs
Original file line number Diff line number Diff line change
@@ -0,0 +1,325 @@
using System;
using System.Collections.Generic;
using System.Linq;
using SIL.Machine.Annotations;
using SIL.Machine.FeatureModel;
using SIL.Machine.Morphology.HermitCrab.MorphologicalRules;
using SIL.ObjectModel;

namespace SIL.Machine.Morphology.HermitCrab
{
/// <summary>
/// Checks a loaded <see cref="Language"/> for problems HermitCrab does not otherwise report:
/// segments used without a declaration, declared segments with duplicate phonological feature
/// bundles, and morphemes whose analysis is marked partial. These problems can silently refuse
/// words, make morpheme identification unreliable, or broaden analysis enough to disable safe
/// final-template pruning. This checker surfaces them before the grammar ships. It is diagnostic
/// only: it never changes how a <see cref="Language"/> parses.
/// </summary>
public static class GrammarHealthChecker
{
/// <summary>
/// Runs every check against <paramref name="language"/> and returns the findings, in the
/// order the checks ran. An empty list means every registered check passed, not that nothing
/// was checked -- see <see cref="GrammarHealthCodes"/> for what each finding's code means.
/// </summary>
public static IList<GrammarHealthFinding> Check(Language language)
{
if (language == null)
throw new ArgumentNullException("language");

var findings = new List<GrammarHealthFinding>();
CheckDuplicateFeatureBundles(language, findings);
CheckUndeclaredSegments(language, findings);
CheckPartialMorphemes(language, findings);
return findings;
}

private static void CheckPartialMorphemes(Language language, List<GrammarHealthFinding> findings)
{
var seen = new HashSet<Morpheme>(new ReferenceEqualityComparer<Morpheme>());

foreach (Stratum stratum in language.Strata)
{
foreach (LexEntry entry in stratum.Entries)
CheckPartialMorpheme(entry, seen, findings);

foreach (Morpheme rule in stratum.MorphologicalRules.OfType<Morpheme>())
CheckPartialMorpheme(rule, seen, findings);

foreach (AffixTemplate template in stratum.AffixTemplates)
{
foreach (MorphemicMorphologicalRule rule in template.Slots.SelectMany(slot => slot.Rules))
CheckPartialMorpheme(rule, seen, findings);
}
}
}

private static void CheckPartialMorpheme(
Morpheme morpheme,
HashSet<Morpheme> seen,
List<GrammarHealthFinding> findings
)
{
if (!morpheme.IsPartial || !seen.Add(morpheme))
return;

string kind;
string name;
var rule = morpheme as MorphemicMorphologicalRule;
if (rule != null)
{
kind = "Morphological rule";
name = FirstNonEmpty(rule.Name, rule.Id, rule.Gloss);
}
else
{
kind = "Lexical entry";
name = FirstNonEmpty(morpheme.Id, morpheme.Gloss);
}

findings.Add(
new GrammarHealthFinding(
GrammarHealthSeverity.Warning,
GrammarHealthCodes.PartialMorpheme,
string.Format(
"{0} '{1}' is partially analyzed. Supply its missing category or template/slot analysis; "
+ "leaving it partial can broaden analysis and disable safe final-template pruning.",
kind,
name
),
new object[] { morpheme }
)
);
}

private static string FirstNonEmpty(params string[] values)
{
return values.FirstOrDefault(value => !string.IsNullOrEmpty(value)) ?? "unnamed";
}

// Every table's segments must have distinct phonological feature bundles, or a segment-changing
// rule cannot tell them apart.
private static void CheckDuplicateFeatureBundles(Language language, List<GrammarHealthFinding> findings)
{
// No feature system means every bundle is the same empty struct by construction (see
// PhonologicalBundle), not a collision.
if (language.PhonologicalFeatureSystem.Count == 0)
return;

foreach (CharacterDefinitionTable table in language.CharacterDefinitionTables)
{
List<CharacterDefinition> segmentDefs = table
.Where(cd => cd.Type == HCFeatureSystem.Segment)
.OrderBy(cd => cd.Representations.First(), StringComparer.Ordinal)
.ToList();

// ValueEquals is the model's own deep, order-independent feature-value equality.
var groups = new List<List<CharacterDefinition>>();
foreach (CharacterDefinition cd in segmentDefs)
{
FeatureStruct bundle = PhonologicalBundle(cd);
List<CharacterDefinition> group = groups.FirstOrDefault(g =>
PhonologicalBundle(g[0]).ValueEquals(bundle)
);
if (group == null)
{
group = new List<CharacterDefinition>();
groups.Add(group);
}
group.Add(cd);
}

foreach (List<CharacterDefinition> group in groups)
{
if (group.Count < 2)
continue;

string names = string.Join(", ", group.Select(cd => cd.Representations.First()));
var subjects = new List<object> { table };
subjects.AddRange(group);
findings.Add(
new GrammarHealthFinding(
GrammarHealthSeverity.Warning,
GrammarHealthCodes.DuplicateFeatureBundle,
string.Format(
"Character definition table '{0}' has {1} segments with an identical "
+ "phonological feature bundle, so a segment-changing rule cannot reliably "
+ "tell them apart: {2}.",
table.Name,
group.Count,
names
),
subjects
)
);
}
}
}

// Strips Type (constant per segment) and any synthesized StrRep, neither of which the grammar author chose.
private static FeatureStruct PhonologicalBundle(CharacterDefinition cd)
{
FeatureStruct bundle = cd.FeatureStruct.Clone();
bundle.RemoveValue(HCFeatureSystem.Type);
bundle.RemoveValue(HCFeatureSystem.StrRep);
return bundle;
}

// Every segment the grammar actually uses must be declared in the table it is used against.
private static void CheckUndeclaredSegments(Language language, List<GrammarHealthFinding> findings)
{
var declaredTables = new HashSet<CharacterDefinitionTable>(language.CharacterDefinitionTables);

foreach (Stratum stratum in language.Strata)
{
foreach (LexEntry entry in stratum.Entries)
{
foreach (RootAllomorph allomorph in entry.Allomorphs)
{
CheckSegmentsDeclared(
allomorph.Segments,
string.Format(
"Lexical entry '{0}' allomorph '{1}'",
entry.Id,
allomorph.Segments.Representation
),
findings
);
}
}

// A rule reached only through an affix-template slot is not necessarily in
// stratum.MorphologicalRules; a rule reached both ways must still report once.
var seenRules = new HashSet<IMorphologicalRule>(new ReferenceEqualityComparer<IMorphologicalRule>());

foreach (IMorphologicalRule rule in stratum.MorphologicalRules)
CheckRuleSegmentsDeclared(rule, seenRules, findings);

foreach (AffixTemplate template in stratum.AffixTemplates)
{
foreach (MorphemicMorphologicalRule rule in template.Slots.SelectMany(slot => slot.Rules))
CheckRuleSegmentsDeclared(rule, seenRules, findings);
}
}

foreach (NaturalClass naturalClass in language.NaturalClasses)
{
var segmentClass = naturalClass as SegmentNaturalClass;
if (segmentClass == null)
continue;

foreach (CharacterDefinition cd in segmentClass.Segments)
{
if (cd.CharacterDefinitionTable != null && declaredTables.Contains(cd.CharacterDefinitionTable))
continue;

findings.Add(
new GrammarHealthFinding(
GrammarHealthSeverity.Error,
GrammarHealthCodes.UndeclaredSegment,
string.Format(
"Natural class '{0}' references a segment ('{1}') that does not belong to any "
+ "character definition table in this language.",
naturalClass.Name,
cd.Representations.Count > 0 ? cd.Representations.First() : cd.FeatureStruct.ToString()
),
new object[] { naturalClass, cd }
)
);
}
}
}

private static void CheckRuleSegmentsDeclared(
IMorphologicalRule rule,
HashSet<IMorphologicalRule> seen,
List<GrammarHealthFinding> findings
)
{
if (!seen.Add(rule))
return;

var affixRule = rule as AffixProcessRule;
if (affixRule != null)
CheckAffixAllomorphsDeclared(affixRule.Name, affixRule.Allomorphs, findings);

// A sibling of AffixProcessRule, not a subclass, but it owns the same allomorph type.
var realizationalRule = rule as RealizationalAffixProcessRule;
if (realizationalRule != null)
CheckAffixAllomorphsDeclared(realizationalRule.Name, realizationalRule.Allomorphs, findings);

var compoundingRule = rule as CompoundingRule;
if (compoundingRule != null)
{
foreach (CompoundingSubrule subrule in compoundingRule.Subrules)
{
foreach (InsertSegments insert in subrule.Rhs.OfType<InsertSegments>())
{
CheckSegmentsDeclared(
insert.Segments,
string.Format(
"Compounding rule '{0}' inserted segments '{1}'",
compoundingRule.Name,
insert.Segments.Representation
),
findings
);
}
}
}
}

private static void CheckAffixAllomorphsDeclared(
string ruleName,
IEnumerable<AffixProcessAllomorph> allomorphs,
List<GrammarHealthFinding> findings
)
{
foreach (AffixProcessAllomorph allomorph in allomorphs)
{
foreach (InsertSegments insert in allomorph.Rhs.OfType<InsertSegments>())
{
CheckSegmentsDeclared(
insert.Segments,
string.Format(
"Morphological rule '{0}' inserted segments '{1}'",
ruleName,
insert.Segments.Representation
),
findings
);
}
}
}

// Same GetMatchingStrReps lookup used to render a shape back to text; boundary/anchor nodes are
// structural, not graphemes.
private static void CheckSegmentsDeclared(Segments segments, string where, List<GrammarHealthFinding> findings)
{
CharacterDefinitionTable table = segments.CharacterDefinitionTable;
foreach (ShapeNode node in segments.Shape)
{
if (node.Annotation.Type() != HCFeatureSystem.Segment)
continue;
if (table.GetMatchingStrReps(node).Any())
continue;

findings.Add(
new GrammarHealthFinding(
GrammarHealthSeverity.Error,
GrammarHealthCodes.UndeclaredSegment,
string.Format(
"{0} contains a segment with feature bundle {1} that character definition table "
+ "'{2}' does not declare.",
where,
node.Annotation.FeatureStruct,
table.Name
),
new object[] { table, segments, node }
)
);
}
}
}
}
Loading
Loading