Add UniqueStemMSAs - #1148
Add UniqueStemMSAs#1148jtmaxwell3 wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Keeping the first of several equal MSAs changes which MSA the parser reports. Glosses are mostly fine: SenseWithMsa returns only the first sense that uses an MSA, so a project whose senses correctly share one MSA also shows one gloss per parse, and deduplicating makes duplicate data behave like clean data there.
What breaks is identity. Saved analyses and ad-hoc prohibitions point at a specific MSA. After the loader drops a duplicate, analyses using it stop matching (ParseResult.MatchesIWfiAnalysis compares MsaRA by identity) and get a parser "disapproves" from SetUnsuccessfulParseEvals, and prohibitions naming it are discarded in LoadMorphemeCoOccurrenceRules.
Suggestion: fix the data, not the loader. liblcm already merges EqualsMsa-equal MSAs when entries are merged (LexEntry.MergeObject, OverridesLing_Lex.cs ~3524, LT-13007), and MergeObject redirects every reference to the dropped MSA and combines reference collections such as slots.
- Now: a Tools > Utilities fixer, next to
DuplicateAnalysisFixer, that does that merge within each entry. Senses, analyses and prohibitions all follow the kept MSA, and the duplicate parses go away. - With the upcoming liblcm model bump: the same merge as a one-time data migration, so every project is repaired once (draft Jira issue below).
If the loader deduplication stays until projects are repaired, it would need:
- Dropped MSAs registered in
m_morphemesunder the kept entry, so prohibitions naming them still apply. - Parse filing that matches saved analyses by group, not exact
MsaRA, with the dropped-to-kept map passed from the loader. - A predictable survivor: the MSA of the first sense, which gives the same gloss a clean project would show.
HCLoaderTestsfor the grouping, variants and prohibitions, plus aParseFilertest for a saved analysis on a dropped MSA.- A scope decision on affix MSAs, where the Aweti reduplication duplicates live.
For the Aweti speed problem, sillsdev/machine#519 takes the worst word from out-of-memory to 9 s by pruning reduplication splits soundly, without touching MSAs.
Draft Jira issue (not filed yet)
- Type: Task, project LT, Affects Version FW 9.3
- Summary: Data migration: merge equivalent MSAs within each lexical entry
- What this is: a one-time liblcm data migration that merges
EqualsMsa-equal MSAs within each entry, using the existingMergeObjectpath. - Who it affects: projects with duplicate grammatical info on an entry. The parser builds one entry per duplicate, so words get duplicate parses.
- Why it matters: it removes the duplicate parses at the source, while senses, saved analyses and ad-hoc prohibitions keep pointing at a valid MSA.
- Links: LT-22615 (related: creating a distinct MSA on purpose), LT-13007 (precedent), this PR.
|
@johnml1135 A possible disadvantage of changing the data is that external data (like interlinear texts) may become invalid. We should check with Jason before pursuing this approach. Merging entries is not a good analogy, since one of the entries disappears in any case. If we merge MSAs, then all of the senses are still there. |
Start here:
UniqueStemMSAsinSrc/LexText/ParserCore/HCLoader.cs, then its two call sites inLoadLexEntries.What it does. When
HCLoadertranslates a FieldWorks lexicon to HermitCrab, it builds one HermitCrab lexical entry per stem MSA. Users often create a new, identical MSA when adding a sense instead of reusing the existing one, so an entry with three identical "Noun" MSAs becomes three HermitCrab entries. They have the same allomorphs, category and features, and differ only in gloss. Every word built on that stem then parses three times, and the parser does three times the work for it. This PR keeps only the first stem MSA in each group thatIMoStemMsa.EqualsMsacalls equal (category, inflection class, features, from-categories, production restrictions), for ordinary entries and for variants of a main entry.The unknown you start with: what changes for users? Parses of affected stems are no longer duplicated, and the surviving parse carries the kept MSA's gloss. Equivalent MSAs in different entries are never merged.
Where to look
kãᴷlists an unreferenced MSA first; its one sense (to.dry) points at the second. After this PR its only entry carries a null gloss, andHCParsershows the placeholder glossksQuestions. Preferring an MSA thatSenseWithMsaresolves would avoid this.LoadMorphologicalRulesstill builds one rule per affix MSA. The Aweti reduplication entryllhas 9 derivational MSAs, 6 of them redundant copies of an earlier one. They still become 9 rules, so the Aweti reduplication case this change was written for is not fixed here.HCLoaderTestshas no duplicate-MSA case.Downstream. PanGloss (a Rust port of HermitCrab) mirrors this loader and has a matching port waiting on this PR: sillsdev/PanGloss#3.
Verification. Not built or run. The duplicate counts come from a script over the project
.fwdatafiles (see Evidence).Next: decide between keep-first and prefer-referenced, and whether affix MSAs belong in this PR or a follow-up; then add an
HCLoaderTestscase.Evidence: duplicate MSAs in five real projects
Counted directly from the
.fwdataXML: an MSA counts as a duplicate when its serialized content, with GUIDs removed and owned objects inlined, equals an earlier MSA of the same class in the same entry. That is stricter thanEqualsMsa(ordered, exact), so these counts are lower bounds.kãᴷ)ll)amyali)In Sena's
amyalithe referenced MSA comes first, so keep-first happens to keep the right one there. In Aweti'skãᴷit does not.Original description
This makes stem MSAs unique within a lexicon entry when translating from FieldWorks to HermitCrab. MSAs are defined locally to a lexicon entry, and are supposed to be unique, but the user sometimes duplicates MSAs when defining new senses rather than reusing existing ones. Having duplicate MSAs affects the performance of Hermit Crab, since it produces duplicate rules if the senses only differ semantically, not syntactically. This not only introduces spurious ambiguity, but allows the rule to be applied multiple times (otherwise, rules can only be applied once). This is particularly bad for the Aweti grammar, which has multiple versions of the reduplication rules that only differ semantically.
How the gloss is lost
LoadLexEntrysetshcEntry.Gloss = GetGloss(msa)andProperties[MsaID] = msa.Hvo.GetGlossreturnsSenseWithMsa(msa)?.Gloss, which is null for an MSA no sense references. At parse timeHCParserresolves the MSA fromMsaIDand again callsSenseWithMsa; a null sense yieldsParserCoreStrings.ksQuestionsas the gloss.This change is