HermitCrab optimization round record: 23 attempts; a faster package that is not mergeable - #490
johnml1135 wants to merge 2 commits into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #490 +/- ##
=======================================
Coverage 74.07% 74.07%
=======================================
Files 456 456
Lines 38164 38164
Branches 5228 5228
=======================================
Hits 28270 28270
Misses 8734 8734
Partials 1160 1160 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
004e199 to
3252bbc
Compare
|
Two further optimization ideas were investigated against this ledger, on a branch based on Neither appears in the existing 22 rows. Row 5 covers deterministic synthesis fold sharing and the rejected two-pass experiment covers nondeterministic traversal; neither is the idea below. Row 6 is the synthesis-side sibling of the second idea. Idea A — a
|
| grammar | words | rule attempts off → on | Σ det arc checks off → on | suppressed | Σ nondet arc checks | det share of FST work |
|---|---|---|---|---|---|---|
| Indonesian | 121 | 15,040 → 15,040 | 34,386 → 34,386 | 0 | 990,111 | 3.4 % |
| Mbugwe | 37 | 4,351,113 → 4,351,113 | 760,833 → 760,829 | 4 | 21,334,695 | 3.4 % |
| Sena | 35 | 22,416,707 → 22,416,707 | 194,014 → 192,515 | 192 | 124,891,129 | 0.16 % |
| Amharic | 37 | 71,315 → 71,315 | 9,979 → 9,979 | 0 | 44,559,287 | 0.02 % |
| Aweti | 22 | 24,274,820 → 24,274,820 | 192,827 → 192,779 | 8 | 97,502,953 | 0.20 % |
Parse-signature multisets identical on every completed word. The reason it is a no-op is structural: the deterministic traversal only runs on the synthesis side (analysis affix matchers set AllSubmatches, analysis rewrite/metathesis set Nondeterministic), and synthesis shapes carry almost no optional annotations, so the skip-vs-consume fork it would collapse rarely occurs. Deterministic traversal is 0.02–3.4 % of all CheckInputMatch calls, and dedup removes ≤ 0.8 % of that.
Where the duplicate explosion the idea was aiming at actually lives is the nondeterministic set, which already exists: on Amharic it suppressed 30.1 M candidate instances against 15.2 M pushed (66 % duplicates), and on one heavy word 82.7 M against 38.5 M.
Two latent issues noticed while auditing the existing keys, neither reachable from HermitCrab today:
NondeterministicFstTraversalMethod's key carries the output action list but not which annotations the actions applied to. Two paths through optional annotations can share state, index, registers and action sequence while having modified different output annotations. Unreachable because no HermitCrab FST carriesIFstOperations.- Both nondeterministic keys omit
VariableBindings, whichCheckInputMatchreads. Whether two equal(State, AnnotationIndex, Registers)keys can carry different bindings on a real pattern is untested. - Both store the mutable
Register[,]directly in the key whileExecuteCommandslater mutates it. This can only cause a missed dedup, never a wrong one.
Idea B — canonicalize the order of independent rules in an Unordered analysis stratum
The analysis-side counterpart of row 6. In an Unordered stratum the cascade explores both [A,B] and [B,A]; if A and B commute, one order is redundant work. The prune skips a dispatch when the candidate is orthogonal to the trail's last rule and precedes it in source order — adjacent-transposition normalization, with non-orthogonal rules and templates acting as barriers.
Orthogonality was implemented as three cumulative levels — L0 shape (pure prefix vs pure suffix, single application, edge-blind stem pattern, no ModifyFromInput/variables), L1 + syntactic-feature commutation decided by enumerating the finite feature lattice the two rules mention in both the analysis (IsUnifiable gate then Add/Clear) and synthesis (Unify then PriorityUnion) directions, L2 + MPR/stem-name/partial/blocking. A second, independently written census over the XML agreed with the object-model implementation on every unordered stratum: 176 L0 pairs, 60 L1, 36 L2 across the five grammars.
Measured: 0.00–0.05 % of rule-unapplication attempts, parity clean on every word.
| grammar | words | attempts off → on | reduction | dispatches | skipped |
|---|---|---|---|---|---|
| Indonesian | 121 | 15,040 → 15,040 | 0.00 % | 21,030 | 0 |
| Mbugwe | 37 | 4,351,113 → 4,351,113 | 0.00 % | 1,631,322 | 858 |
| Sena | 35 | 22,416,707 → 22,404,555 | 0.05 % | 6,603,760 | 12,104 |
| Amharic | 29 | 36,774 → 36,758 | 0.04 % | 27,264 | 16 |
| Aweti | 12 | 3,014,439 → 3,014,439 | 0.00 % | 1,213,430 | 0 |
| Sena at L0 (unsound control, same 35 words) | 35 | 25,823,709 → 24,179,995 | 6.37 % | 7,442,610 | 67,090 |
The L0 control is what makes this a finding rather than a null result: widening Sena from 2 qualifying pairs to 52 moves the reduction from 0.05 % to 6.37 %, so the limit is the predicate, not the hook. But each narrowing step is load-bearing — of Sena's 52 L0 pairs, 30 fail L1 because their syntactic features genuinely do not commute, and 20 fail L2 on the partial-rule clause, which is necessary there because that stratum holds 24 affix templates so IsLastAppliedRuleFinal takes a value and both final-template gates in SynthesisAffixProcessRule can fire. 6.37 % is what an incorrect prune wins.
A trail-depth census explains why widening cannot reach further. Instrumenting every dispatch with the length of the trail it extends:
- Mbugwe, 155,672 dispatches: 68 % sit on trails 9+ rules deep, but at only 0.8 % does the trail's last rule have any orthogonal partner even at L0.
- A heavy Sena word, 76,777 dispatches: 67 % at depth 9+, 16.6 % partner-available, 2.3 % actually skipped → the 6.37 %.
So the permutation blow-up is real and deep, but the rules filling those trails are ones the criterion excludes on shape (Mbugwe rejects 57 pairs MultiInsert, 17 AllomorphOutput, 21 RuleKind). The largest excluded class is same-side pairs (Sena: 84 rejections, the C(13,2) suffix pairs), and those can never commute — stem+A+B ≠ stem+B+A is arithmetic, not conservatism.
One consequence worth separating from performance. The prune is not signature-preserving and cannot be: a parse signature carries MorphemesInApplicationOrder, so unapplying a prefix before a suffix and the reverse are two distinct parses of one word even when the rules commute. What it preserves is the morph sequence and the root; what it drops is derivational history. Those two rows are indistinguishable to a reader, so this is an ambiguity reduction as much as an optimization — and it would need the correctness treatment rather than being presented as invisible.
This is now the fifth instance of the pattern in this ledger's "lesson that generalises": apparent redundancy that evaporates once the key is complete.
On rows 10 and 11 — reachability
The first reading of our data looked like it reopened the lexical reachability gate: the heavy Sena word made 272,152 unapplication attempts and the Aweti word 505,600, each finding zero roots. It does not reopen it. Those lexical-lookup counts come from runs truncated by a step budget, so they show the search never reached lookup within budget, not that lookup would have failed. Row 11's finding — that the expensive words fail on checks running after lexical lookup succeeds — is untouched by anything measured here.
On the complexity-cap step budget
b3fd2b55 was cherry-picked to bound the pathological words. It works where the cost is many rule applications: an Aweti word that previously crashed a test host at 12.1 GB after 1,717 s completes in 50 s at 200,000 steps, and whole-corpus runs go from 22/40 words to 40/40 at 757 MB peak instead of 12 GB.
It does not bound a word whose cost is one expensive traversal. An Amharic word ran 2,280 s under a 200,000-step budget with the nondeterministic counters frozen at 217,493 traversals and 53.9 M arc checks across successive samples — stuck inside a single matcher call. The budget is checked between leaf-rule applications and cannot interrupt one. Anyone relying on it as a safety bound should know that.
Also worth flagging for consumers: the budget ships on by default (2,000,000 steps, 10 s). A 10-second default silently truncates parses.
Reproduction
Branch based on master a4b2974. Measurements used a census harness that parses each word twice in one process with MaxDegreeOfParallelism: 1 and tracing on (so the analysis memo and MergeEquivalentAnalyses are off and counts are the unmemoized search), compares complete parse-signature multisets, and reports MorphologicalRuleAnalysis trace nodes as attempts and CheckInputMatch calls as arc checks. Grammars are FieldWorks GenerateHCConfig exports of five projects; word lists are the first 40 entries except Indonesian (all 121) and Sena (entries 1001–1040, since the head of that list is punctuation). Machine-safety aborts and test-host crashes are recorded as such and excluded from ratios, never counted as agreement.
|
Two observations from the same investigation that belong in the record but are not optimizations, so they did not fit the rows above. A repeatable memory wall at 12.2–12.6 GB, independent of any capAcross the five-grammar runs, 13 words (Sena and Aweti) ended as test-host crashes rather than completions, and they did so at a strikingly consistent managed heap: 12,225 / 12,259 / 12,270 / 12,278 / 12,282 / 12,312 / 12,323 / 12,324 / 12,336 / 12,340 / 12,359 / 12,365 / 12,563 MB. The consistency holds regardless of the per-process limit in force — including runs whose limit was 20 GB, so these are not the harness stopping them. The clustering at ~12.3 GB rather than at whatever ceiling was configured is the interesting part; it is the signature one would expect from a single collection outgrowing the 2 GB per-object limit rather than from address-space or RAM exhaustion. A crash dump on one of these words would name the collection. Reported as an observation only — no diagnosis attempted. For scale, one Aweti word (line 183 of its list) reached 13.2 GB and 30.1 M nondeterministic traversals / 139.3 M arc checks in 53 minutes without leaving analysis, and never reported a lexical lookup. Under a 200,000-step budget the same word completes in 231 s at 676 MB with 505,600 rule-unapplication attempts and 0 deterministic traversals. Where analysis time goes, by traversal familyThe counters added for this work split every
This is offered as calibration for anyone sizing a synthesis-side optimization: on the two grammars with the heaviest analysis, 99.8 % or more of FST work is on the nondeterministic side. It is consistent with the ledger's Amharic breakdown (analysis cascade ~95 % of analysis, all synthesis 0.3 %), measured independently and at the arc-check level rather than by wall-time bucket. |
|
Added one record from a September round that post-dates this set: Three things it contributes to this ledger:
Also closes the two extensions before anyone builds them — environment matchers are 0.20%/1.18% of traversal instances on Amharic/Mbugwe (this is #515), and the analysis rules' 93.95% is 99.7% alternate captures because
🤖 Generated with Claude Code |
2026-09-23 round: the missing Aweti speedup was reduplication, not traversalOne new row closes the gap #511 opened. Aweti index 182 had never completed here, on any stack we tried (#491, #494, #511, FieldWorks #1148, forced final templates, FLEx's parallel/no-memo mode). A copy-agreement prune in reduplication unapplication makes it finish in 9 s, and the whole Aweti list runs 4.95× faster with 0 divergences across five grammars. Proposed as #519 (on by default) and shipped in PanGloss v0.4.0. New row
Why the ledger missed it: every prior row measured Sena, Mbugwe and Amharic, where full-copy rules are absent or rare. The 19 docs mention reduplication once (the probe-design classifier, marked "Unknown"). What else this round settled
PanGloss v0.4.0 (same prune, Rust)Release binaries, same machine, v0.3.3 vs v0.4.0: identical analyses on all 7,455 words both complete across the five grammars. Aweti: 172 -> 191 words complete within 20 s, 5.3x on the 172 shared words, median word 188 -> 26 ms, word-list index 182 60 s timeout -> 2.7 s. Sena, Amharic and Mbugwe are within noise to about 10% slower; a per-call allocation on rules that copy nothing is being removed. Open since this round
Method notes worth keeping
🤖 Generated with Claude Code |
Full docs-only record of the September 2026 HermitCrab optimization round: 22 ledger attempts (2 earlier shipped wins, 12 closed as measured-does-not-pay, 3 closed before building, 5 reopened this round), the probe design needed to rebuild the measurement instrumentation from scratch, the per-mechanism disposition (9 retained mechanisms with individual records, 4 rejected with individual records), and the branch/worktree inventory. The round's own conclusion, from the disposition doc: do not merge the branch as a package (896 lines across 28 files, not reviewable as one change, and the package speedup could not be attributed to its parts); isolate what's worth isolating separately, and refocus on PR #491 and grammar hygiene where the large returns are. Kept here as a durable "do not retry this" record, not as a merge candidate. Copied from branch perf/hc-optimization-archive onto current master, replacing the stale two-file version of this PR. Minimal fixups only: the ledger, probe-design, disposition, and counts docs were read in full to check for (a) references to branches deleted since the archive was cut and (b) private corpus data. Found and fixed: the disposition doc's several references to `perf/hc-edge-prefilter`, whose own branch and archive tag have both since been deleted; the mechanism itself is not lost, since it also lives in the whole-round archive (branch `origin/perf/hc-optimization-archive`) copied in by this same commit, so the references now name that ref instead of calling the mechanism unrecoverable. The branch-wide parity sentence in the counts doc is qualified to say what `WordAnalysisSignature` actually verifies on the FLEx exports (morph counts, since `Morpheme.Id` is empty), matching the caveat the disposition doc already recorded. The finished execution-plan checklist is dropped now that its outcome is captured in the disposition record instead. No private word forms or corpus data found; the docs already record only word-list indices, counts, and ratios, per the round's own scrubbing note. No findings, numbers, or conclusions were otherwise altered. Docs only, no code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
PR #511 proposes skipping traversal instances whose (State, AnnotationIndex) was already pushed. The motivating pathology is real and reproduces on a 2-state fsa: Advance forks per Optional annotation, so instances are exponential in optional count and independent of the state count. The key is too coarse. A 20,000-case differential fuzz finds 112 cases where Match() changes, 108 from ignored variable bindings and 4 from shortened ranges on the deterministic path, with AllMatches().First() identical in all 20,000 as the control. A narrowed gate -- deterministic method, no capture groups -- measures 0 divergences with the full bound preserved. Two censuses close the extensions: environment matchers are 0.20%/1.18% of traversal instances on Amharic/Mbugwe, and the analysis rules' apparent 93.95% collapse is 99.7% alternate captures, because AnalysisAffixProcessRule and AnalysisCompoundingRule set AllSubmatches and enumerate morph boundaries on purpose. That makes a fourth row for the apparent-vs-sound table. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
47318eb to
5bbfa01
Compare
Quick summary
Preserves the HermitCrab optimization round's experiments, measurements, and disposition so future
work can reuse the evidence instead of repeating it. The package measured faster (1.76x Amharic,
1.83x Mbugwe) but was archived: unattributable per mechanism, and its parity gate compared morph
counts only. Documentation only: no parser behavior changes, and the archived mechanisms
(including the edge prefilter) stay off master.
Where to look
docs/hermitcrab-perf-2026-09-counts.md:59: the legacyWordAnalysisSignaturecomparedmorph counts only on the FLEx exports (
Morpheme.Idis empty).docs/hermitcrab-perf-2026-09-disposition.mdpoints to branchperf/hc-optimization-archive,which still holds the edge prefilter (
git grep EdgePrefilterEnabled).docs/hermitcrab-optimization-ledger.md: 23 attempts, five rejected-optimization records.Deliberately not included
parse-identity signature first, not this PR.
Validation
git diff --check origin/master...HEAD-- clean, no output.pwsh -NoProfile -File scripts/comment-hygiene.ps1 -BaseRef origin/master-- no in-scope files(docs-only diff).
dotnet build,dotnet test,dotnet csharpier check .,./local_check.sh-- docs-onlychange, no
.cs/.csprojtouched.Reading this a year from now
This record covers the September 2026 HermitCrab optimization round: 18 documentation files, 23
ledger attempts, nine retained experimental mechanisms, and five rejected or closed experiments,
plus measurement methods and the final disposition. "Retained" does not mean merged or
production-ready -- the package itself was rejected for merge.
The package's branch-wide sample was faster on all five languages measured, "stable" on two of them
per the probe rules (1.76x Amharic, 1.83x Mbugwe), but the package figure could not be attributed to
any one mechanism -- only the deferred template clone was ever timed alone, at 1.054x. That
non-attribution, not an absence of speedup, is why the package was archived rather than merged.
Correctness limit: the branch-wide "parity" gate used
WordAnalysisSignature, which joinsMorpheme.Id-- empty for every FLEx export used, so on Sena, Mbugwe, and Amharic the signaturedegenerates to a morph count per analysis. It verified counts, not analyses. Any renewed work needs
a gloss + allomorph-index + syntactic-feature-structure signature, or better, a complete
parse-identity multiset including multiplicities.
Evidence, scope, and provenance
Durable records: ledger,
disposition,
counts, probe design,
nine mechanism records under
docs/optimizations/, and five rejection records underdocs/rejected-optimizations/(the fifth,state-position-traversal-dedup.md, is PR #511'straversal dedupe, added after the original round closed).
The unmerged package is archived on branch
perf/hc-optimization-archive. The separately isolatedperf/hc-edge-prefilterbranch andits own archive tag were deleted, but the edge-prefilter mechanism (source and tests) still lives in
the whole-round archive above -- confirmed directly with
git grep EdgePrefilterEnabledagainstboth refs.
Historical verification (prior measurement, not rerun for this description update): 603 of 604
non-explicit NUnit tests passed, one skipped for Windows symlink privileges.
Original work generated with Claude Code.
🤖 Generated with Claude Code
This change is