Skip to content

Add UniqueStemMSAs - #1148

Draft
jtmaxwell3 wants to merge 1 commit into
mainfrom
add-UniqueStemMSAs
Draft

jtmaxwell3 wants to merge 1 commit into
mainfrom
add-UniqueStemMSAs

Conversation

@jtmaxwell3

@jtmaxwell3 jtmaxwell3 commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Start here: UniqueStemMSAs in Src/LexText/ParserCore/HCLoader.cs, then its two call sites in LoadLexEntries.

What it does. When HCLoader translates a FieldWorks lexicon to HermitCrab, it builds one HermitCrab lexical entry per stem MSA. Users often create a new, identical MSA when adding a sense instead of reusing the existing one, so an entry with three identical "Noun" MSAs becomes three HermitCrab entries. They have the same allomorphs, category and features, and differ only in gloss. Every word built on that stem then parses three times, and the parser does three times the work for it. This PR keeps only the first stem MSA in each group that IMoStemMsa.EqualsMsa calls equal (category, inflection class, features, from-categories, production restrictions), for ordinary entries and for variants of a main entry.

The unknown you start with: what changes for users? Parses of affected stems are no longer duplicated, and the surviving parse carries the kept MSA's gloss. Equivalent MSAs in different entries are never merged.

Where to look

  1. Which duplicate survives. "First" is owning-collection order, which is unrelated to which MSA a sense uses. In the Aweti project, the entry kãᴷ lists an unreferenced MSA first; its one sense (to.dry) points at the second. After this PR its only entry carries a null gloss, and HCParser shows the placeholder gloss ksQuestions. Preferring an MSA that SenseWithMsa resolves would avoid this.
  2. Affix MSAs are unchanged. LoadMorphologicalRules still builds one rule per affix MSA. The Aweti reduplication entry ll has 9 derivational MSAs, 6 of them redundant copies of an earlier one. They still become 9 rules, so the Aweti reduplication case this change was written for is not fixed here.
  3. No test. HCLoaderTests has no duplicate-MSA case.

Downstream. PanGloss (a Rust port of HermitCrab) mirrors this loader and has a matching port waiting on this PR: sillsdev/PanGloss#3.

Verification. Not built or run. The duplicate counts come from a script over the project .fwdata files (see Evidence).

Next: decide between keep-first and prefer-referenced, and whether affix MSAs belong in this PR or a follow-up; then add an HCLoaderTests case.


Evidence: duplicate MSAs in five real projects

Counted directly from the .fwdata XML: an MSA counts as a duplicate when its serialized content, with GUIDs removed and owned objects inlined, equals an earlier MSA of the same class in the same entry. That is stricter than EqualsMsa (ordered, exact), so these counts are lower bounds.

Project Stem MSAs (duplicates) Derivational affix MSAs (duplicates) Inflectional affix MSAs (duplicates)
Aweti 804 (1, in kãᴷ) 44 (6, all in ll) 78 (0)
Sena 1,394 (1, in amyali) 11 (0) 61 (0)
Amharic 77 (0) 4 (0) 55 (0)
Indonesian 66 (0) 9 (0) 3 (0)
Mbugwe 163 (0) 20 (0) 203 (0)

In Sena's amyali the referenced MSA comes first, so keep-first happens to keep the right one there. In Aweti's kãᴷ it does not.

Original description

This makes stem MSAs unique within a lexicon entry when translating from FieldWorks to HermitCrab. MSAs are defined locally to a lexicon entry, and are supposed to be unique, but the user sometimes duplicates MSAs when defining new senses rather than reusing existing ones. Having duplicate MSAs affects the performance of Hermit Crab, since it produces duplicate rules if the senses only differ semantically, not syntactically. This not only introduces spurious ambiguity, but allows the rule to be applied multiple times (otherwise, rules can only be applied once). This is particularly bad for the Aweti grammar, which has multiple versions of the reduplication rules that only differ semantically.

How the gloss is lost

LoadLexEntry sets hcEntry.Gloss = GetGloss(msa) and Properties[MsaID] = msa.Hvo. GetGloss returns SenseWithMsa(msa)?.Gloss, which is null for an MSA no sense references. At parse time HCParser resolves the MSA from MsaID and again calls SenseWithMsa; a null sense yields ParserCoreStrings.ksQuestions as the gloss.


This change is Reviewable

@johnml1135 johnml1135 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Withdrawn; points will follow individually.

@johnml1135 johnml1135 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keeping the first of several equal MSAs changes which MSA the parser reports. Glosses are mostly fine: SenseWithMsa returns only the first sense that uses an MSA, so a project whose senses correctly share one MSA also shows one gloss per parse, and deduplicating makes duplicate data behave like clean data there.

What breaks is identity. Saved analyses and ad-hoc prohibitions point at a specific MSA. After the loader drops a duplicate, analyses using it stop matching (ParseResult.MatchesIWfiAnalysis compares MsaRA by identity) and get a parser "disapproves" from SetUnsuccessfulParseEvals, and prohibitions naming it are discarded in LoadMorphemeCoOccurrenceRules.

Suggestion: fix the data, not the loader. liblcm already merges EqualsMsa-equal MSAs when entries are merged (LexEntry.MergeObject, OverridesLing_Lex.cs ~3524, LT-13007), and MergeObject redirects every reference to the dropped MSA and combines reference collections such as slots.

  1. Now: a Tools > Utilities fixer, next to DuplicateAnalysisFixer, that does that merge within each entry. Senses, analyses and prohibitions all follow the kept MSA, and the duplicate parses go away.
  2. With the upcoming liblcm model bump: the same merge as a one-time data migration, so every project is repaired once (draft Jira issue below).

If the loader deduplication stays until projects are repaired, it would need:

  1. Dropped MSAs registered in m_morphemes under the kept entry, so prohibitions naming them still apply.
  2. Parse filing that matches saved analyses by group, not exact MsaRA, with the dropped-to-kept map passed from the loader.
  3. A predictable survivor: the MSA of the first sense, which gives the same gloss a clean project would show.
  4. HCLoaderTests for the grouping, variants and prohibitions, plus a ParseFiler test for a saved analysis on a dropped MSA.
  5. A scope decision on affix MSAs, where the Aweti reduplication duplicates live.

For the Aweti speed problem, sillsdev/machine#519 takes the worst word from out-of-memory to 9 s by pruning reduplication splits soundly, without touching MSAs.

Draft Jira issue (not filed yet)

  • Type: Task, project LT, Affects Version FW 9.3
  • Summary: Data migration: merge equivalent MSAs within each lexical entry
  • What this is: a one-time liblcm data migration that merges EqualsMsa-equal MSAs within each entry, using the existing MergeObject path.
  • Who it affects: projects with duplicate grammatical info on an entry. The parser builds one entry per duplicate, so words get duplicate parses.
  • Why it matters: it removes the duplicate parses at the source, while senses, saved analyses and ad-hoc prohibitions keep pointing at a valid MSA.
  • Links: LT-22615 (related: creating a distinct MSA on purpose), LT-13007 (precedent), this PR.

@jtmaxwell3

Copy link
Copy Markdown
Collaborator Author

@johnml1135 A possible disadvantage of changing the data is that external data (like interlinear texts) may become invalid. We should check with Jason before pursuing this approach. Merging entries is not a good analogy, since one of the entries disappears in any case. If we merge MSAs, then all of the senses are still there.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants