Skip to content

HermitCrab optimization round record: 23 attempts; a faster package that is not mergeable - #490

Open
johnml1135 wants to merge 2 commits into
masterfrom
docs/hc-optimization-ledger
Open

johnml1135 wants to merge 2 commits into
masterfrom
docs/hc-optimization-ledger

Conversation

@johnml1135

@johnml1135 johnml1135 commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Quick summary

Preserves the HermitCrab optimization round's experiments, measurements, and disposition so future
work can reuse the evidence instead of repeating it. The package measured faster (1.76x Amharic,
1.83x Mbugwe) but was archived: unattributable per mechanism, and its parity gate compared morph
counts only. Documentation only: no parser behavior changes, and the archived mechanisms
(including the edge prefilter) stay off master.

Where to look

  • Parity claims -- docs/hermitcrab-perf-2026-09-counts.md:59: the legacy WordAnalysisSignature compared
    morph counts only on the FLEx exports (Morpheme.Id is empty).
  • Archived code -- docs/hermitcrab-perf-2026-09-disposition.md points to branch perf/hc-optimization-archive,
    which still holds the edge prefilter (git grep EdgePrefilterEnabled).
  • The register -- docs/hermitcrab-optimization-ledger.md: 23 attempts, five rejected-optimization records.

Deliberately not included

  • No source changes, no re-measurement: reviving any archived mechanism needs a complete
    parse-identity signature first, not this PR.

Validation

  • git diff --check origin/master...HEAD -- clean, no output.
  • pwsh -NoProfile -File scripts/comment-hygiene.ps1 -BaseRef origin/master -- no in-scope files
    (docs-only diff).
  • Skipped: dotnet build, dotnet test, dotnet csharpier check ., ./local_check.sh -- docs-only
    change, no .cs/.csproj touched.

Reading this a year from now

This record covers the September 2026 HermitCrab optimization round: 18 documentation files, 23
ledger attempts, nine retained experimental mechanisms, and five rejected or closed experiments,
plus measurement methods and the final disposition. "Retained" does not mean merged or
production-ready -- the package itself was rejected for merge.

The package's branch-wide sample was faster on all five languages measured, "stable" on two of them
per the probe rules (1.76x Amharic, 1.83x Mbugwe), but the package figure could not be attributed to
any one mechanism -- only the deferred template clone was ever timed alone, at 1.054x. That
non-attribution, not an absence of speedup, is why the package was archived rather than merged.

Correctness limit: the branch-wide "parity" gate used WordAnalysisSignature, which joins
Morpheme.Id -- empty for every FLEx export used, so on Sena, Mbugwe, and Amharic the signature
degenerates to a morph count per analysis. It verified counts, not analyses. Any renewed work needs
a gloss + allomorph-index + syntactic-feature-structure signature, or better, a complete
parse-identity multiset including multiplicities.

Evidence, scope, and provenance

Durable records: ledger,
disposition,
counts, probe design,
nine mechanism records under docs/optimizations/, and five rejection records under
docs/rejected-optimizations/ (the fifth, state-position-traversal-dedup.md, is PR #511's
traversal dedupe, added after the original round closed).

The unmerged package is archived on branch
perf/hc-optimization-archive. The separately isolated perf/hc-edge-prefilter branch and
its own archive tag were deleted, but the edge-prefilter mechanism (source and tests) still lives in
the whole-round archive above -- confirmed directly with git grep EdgePrefilterEnabled against
both refs.

Historical verification (prior measurement, not rerun for this description update): 603 of 604
non-explicit NUnit tests passed, one skipped for Windows symlink privileges.

Original work generated with Claude Code.

🤖 Generated with Claude Code


This change is Reviewable

@codecov-commenter

codecov-commenter commented Aug 27, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 74.07%. Comparing base (7109531) to head (5bbfa01).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #490   +/-   ##
=======================================
  Coverage   74.07%   74.07%           
=======================================
  Files         456      456           
  Lines       38164    38164           
  Branches     5228     5228           
=======================================
  Hits        28270    28270           
  Misses       8734     8734           
  Partials     1160     1160           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@johnml1135
johnml1135 force-pushed the docs/hc-optimization-ledger branch from 004e199 to 3252bbc Compare September 15, 2026 13:11
@johnml1135 johnml1135 changed the title HermitCrab optimization ledger: 22 attempts, no speedup, one located target HermitCrab optimization round record: 22 attempts, 9 retained mechanisms, no net speedup as a package Sep 15, 2026
@johnml1135

Copy link
Copy Markdown
Collaborator Author

Two further optimization ideas were investigated against this ledger, on a branch based on master (a4b2974). Both are closed. Recording them here because the ledger's purpose is to stop the work being redone, and one of them looks irresistible on paper.

Neither appears in the existing 22 rows. Row 5 covers deterministic synthesis fold sharing and the rejected two-pass experiment covers nondeterministic traversal; neither is the idea below. Row 6 is the synthesis-side sibling of the second idea.

Idea A — a traversed dedup set for the deterministic FSA traversal

NondeterministicFsaTraversalMethod keeps a traversed set keyed on (State, AnnotationIndex, Registers); DeterministicFsaTraversalMethod does not. The motivating observation was that a deterministic FST can still be traversed nondeterministically when the input is nondeterministic — Annotation.Optional makes Advance fork skip-vs-consume, and the forks reconverge.

The mechanism is real and the key is complete on that path (no output operations, no priorities, and Matcher.Compile only determinizes when there are no variables, so VariableBindings is always null). On a synthetic input with 8 consecutive optional annotations it takes instance pushes from 510 to 44 with identical matches.

It does not pay on real grammars. Five FieldWorks exports, tracing on, each word parsed twice (dedup off then on), one word per process, no parse timeout:

grammar words rule attempts off → on Σ det arc checks off → on suppressed Σ nondet arc checks det share of FST work
Indonesian 121 15,040 → 15,040 34,386 → 34,386 0 990,111 3.4 %
Mbugwe 37 4,351,113 → 4,351,113 760,833 → 760,829 4 21,334,695 3.4 %
Sena 35 22,416,707 → 22,416,707 194,014 → 192,515 192 124,891,129 0.16 %
Amharic 37 71,315 → 71,315 9,979 → 9,979 0 44,559,287 0.02 %
Aweti 22 24,274,820 → 24,274,820 192,827 → 192,779 8 97,502,953 0.20 %

Parse-signature multisets identical on every completed word. The reason it is a no-op is structural: the deterministic traversal only runs on the synthesis side (analysis affix matchers set AllSubmatches, analysis rewrite/metathesis set Nondeterministic), and synthesis shapes carry almost no optional annotations, so the skip-vs-consume fork it would collapse rarely occurs. Deterministic traversal is 0.02–3.4 % of all CheckInputMatch calls, and dedup removes ≤ 0.8 % of that.

Where the duplicate explosion the idea was aiming at actually lives is the nondeterministic set, which already exists: on Amharic it suppressed 30.1 M candidate instances against 15.2 M pushed (66 % duplicates), and on one heavy word 82.7 M against 38.5 M.

Two latent issues noticed while auditing the existing keys, neither reachable from HermitCrab today:

  • NondeterministicFstTraversalMethod's key carries the output action list but not which annotations the actions applied to. Two paths through optional annotations can share state, index, registers and action sequence while having modified different output annotations. Unreachable because no HermitCrab FST carries IFstOperations.
  • Both nondeterministic keys omit VariableBindings, which CheckInputMatch reads. Whether two equal (State, AnnotationIndex, Registers) keys can carry different bindings on a real pattern is untested.
  • Both store the mutable Register[,] directly in the key while ExecuteCommands later mutates it. This can only cause a missed dedup, never a wrong one.

Idea B — canonicalize the order of independent rules in an Unordered analysis stratum

The analysis-side counterpart of row 6. In an Unordered stratum the cascade explores both [A,B] and [B,A]; if A and B commute, one order is redundant work. The prune skips a dispatch when the candidate is orthogonal to the trail's last rule and precedes it in source order — adjacent-transposition normalization, with non-orthogonal rules and templates acting as barriers.

Orthogonality was implemented as three cumulative levels — L0 shape (pure prefix vs pure suffix, single application, edge-blind stem pattern, no ModifyFromInput/variables), L1 + syntactic-feature commutation decided by enumerating the finite feature lattice the two rules mention in both the analysis (IsUnifiable gate then Add/Clear) and synthesis (Unify then PriorityUnion) directions, L2 + MPR/stem-name/partial/blocking. A second, independently written census over the XML agreed with the object-model implementation on every unordered stratum: 176 L0 pairs, 60 L1, 36 L2 across the five grammars.

Measured: 0.00–0.05 % of rule-unapplication attempts, parity clean on every word.

grammar words attempts off → on reduction dispatches skipped
Indonesian 121 15,040 → 15,040 0.00 % 21,030 0
Mbugwe 37 4,351,113 → 4,351,113 0.00 % 1,631,322 858
Sena 35 22,416,707 → 22,404,555 0.05 % 6,603,760 12,104
Amharic 29 36,774 → 36,758 0.04 % 27,264 16
Aweti 12 3,014,439 → 3,014,439 0.00 % 1,213,430 0
Sena at L0 (unsound control, same 35 words) 35 25,823,709 → 24,179,995 6.37 % 7,442,610 67,090

The L0 control is what makes this a finding rather than a null result: widening Sena from 2 qualifying pairs to 52 moves the reduction from 0.05 % to 6.37 %, so the limit is the predicate, not the hook. But each narrowing step is load-bearing — of Sena's 52 L0 pairs, 30 fail L1 because their syntactic features genuinely do not commute, and 20 fail L2 on the partial-rule clause, which is necessary there because that stratum holds 24 affix templates so IsLastAppliedRuleFinal takes a value and both final-template gates in SynthesisAffixProcessRule can fire. 6.37 % is what an incorrect prune wins.

A trail-depth census explains why widening cannot reach further. Instrumenting every dispatch with the length of the trail it extends:

  • Mbugwe, 155,672 dispatches: 68 % sit on trails 9+ rules deep, but at only 0.8 % does the trail's last rule have any orthogonal partner even at L0.
  • A heavy Sena word, 76,777 dispatches: 67 % at depth 9+, 16.6 % partner-available, 2.3 % actually skipped → the 6.37 %.

So the permutation blow-up is real and deep, but the rules filling those trails are ones the criterion excludes on shape (Mbugwe rejects 57 pairs MultiInsert, 17 AllomorphOutput, 21 RuleKind). The largest excluded class is same-side pairs (Sena: 84 rejections, the C(13,2) suffix pairs), and those can never commute — stem+A+B ≠ stem+B+A is arithmetic, not conservatism.

One consequence worth separating from performance. The prune is not signature-preserving and cannot be: a parse signature carries MorphemesInApplicationOrder, so unapplying a prefix before a suffix and the reverse are two distinct parses of one word even when the rules commute. What it preserves is the morph sequence and the root; what it drops is derivational history. Those two rows are indistinguishable to a reader, so this is an ambiguity reduction as much as an optimization — and it would need the correctness treatment rather than being presented as invisible.

This is now the fifth instance of the pattern in this ledger's "lesson that generalises": apparent redundancy that evaporates once the key is complete.

On rows 10 and 11 — reachability

The first reading of our data looked like it reopened the lexical reachability gate: the heavy Sena word made 272,152 unapplication attempts and the Aweti word 505,600, each finding zero roots. It does not reopen it. Those lexical-lookup counts come from runs truncated by a step budget, so they show the search never reached lookup within budget, not that lookup would have failed. Row 11's finding — that the expensive words fail on checks running after lexical lookup succeeds — is untouched by anything measured here.

On the complexity-cap step budget

b3fd2b55 was cherry-picked to bound the pathological words. It works where the cost is many rule applications: an Aweti word that previously crashed a test host at 12.1 GB after 1,717 s completes in 50 s at 200,000 steps, and whole-corpus runs go from 22/40 words to 40/40 at 757 MB peak instead of 12 GB.

It does not bound a word whose cost is one expensive traversal. An Amharic word ran 2,280 s under a 200,000-step budget with the nondeterministic counters frozen at 217,493 traversals and 53.9 M arc checks across successive samples — stuck inside a single matcher call. The budget is checked between leaf-rule applications and cannot interrupt one. Anyone relying on it as a safety bound should know that.

Also worth flagging for consumers: the budget ships on by default (2,000,000 steps, 10 s). A 10-second default silently truncates parses.

Reproduction

Branch based on master a4b2974. Measurements used a census harness that parses each word twice in one process with MaxDegreeOfParallelism: 1 and tracing on (so the analysis memo and MergeEquivalentAnalyses are off and counts are the unmemoized search), compares complete parse-signature multisets, and reports MorphologicalRuleAnalysis trace nodes as attempts and CheckInputMatch calls as arc checks. Grammars are FieldWorks GenerateHCConfig exports of five projects; word lists are the first 40 entries except Indonesian (all 121) and Sena (entries 1001–1040, since the head of that list is punctuation). Machine-safety aborts and test-host crashes are recorded as such and excluded from ratios, never counted as agreement.

@johnml1135

Copy link
Copy Markdown
Collaborator Author

Two observations from the same investigation that belong in the record but are not optimizations, so they did not fit the rows above.

A repeatable memory wall at 12.2–12.6 GB, independent of any cap

Across the five-grammar runs, 13 words (Sena and Aweti) ended as test-host crashes rather than completions, and they did so at a strikingly consistent managed heap: 12,225 / 12,259 / 12,270 / 12,278 / 12,282 / 12,312 / 12,323 / 12,324 / 12,336 / 12,340 / 12,359 / 12,365 / 12,563 MB. The consistency holds regardless of the per-process limit in force — including runs whose limit was 20 GB, so these are not the harness stopping them.

The clustering at ~12.3 GB rather than at whatever ceiling was configured is the interesting part; it is the signature one would expect from a single collection outgrowing the 2 GB per-object limit rather than from address-space or RAM exhaustion. A crash dump on one of these words would name the collection. Reported as an observation only — no diagnosis attempted.

For scale, one Aweti word (line 183 of its list) reached 13.2 GB and 30.1 M nondeterministic traversals / 139.3 M arc checks in 53 minutes without leaving analysis, and never reported a lexical lookup. Under a 200,000-step budget the same word completes in 231 s at 676 MB with 505,600 rule-unapplication attempts and 0 deterministic traversals.

Where analysis time goes, by traversal family

The counters added for this work split every CheckInputMatch call between the two FSA traversal families. Deterministic traversal — which is the synthesis side only, since analysis affix matchers set AllSubmatches and analysis rewrite/metathesis set Nondeterministic — accounts for:

grammar deterministic share of arc checks
Indonesian 3.36 %
Mbugwe 3.44 %
Aweti 0.20 %
Sena 0.16 %
Amharic 0.02 %

This is offered as calibration for anyone sizing a synthesis-side optimization: on the two grammars with the heaviest analysis, 99.8 % or more of FST work is on the nondeterministic side. It is consistent with the ledger's Amharic breakdown (analysis cascade ~95 % of analysis, all synthesis 0.3 %), measured independently and at the arc-check level rather than by wall-time bucket.

@johnml1135

Copy link
Copy Markdown
Collaborator Author

Added one record from a September round that post-dates this set: docs/rejected-optimizations/state-position-traversal-dedup.md, covering PR #511's (State, AnnotationIndex) traversal dedupe.

Three things it contributes to this ledger:

  1. A reproduction of the motivating pathology, with the mechanism named. It is not the state count — TraversalMethodBase.Advance forks an instance per Optional annotation, so instances are exponential in optional count and independent of |States|. Measured 2,097,150 instances popped on a 2-state FSA at 20 optional annotations.
  2. A fourth row for "the lesson that generalises." Apparent collapse 93.95%, sound collapse 0.28%. Same shape as rows 5, 6 and the fold-entry census: a key that omits what makes instances distinct shows large shareable work that is not there. Here the omitted state is variable bindings (which filter later arcs) and registers (which carry the captures AllSubmatches callers consume).
  3. A narrowed form that does pay, unlike most of this ledger: restricted to the deterministic method with no capture groups, the differential fuzz goes 112 divergences → 0 while the group-free microbenchmark stays cell-for-cell identical to the unnarrowed change.

Also closes the two extensions before anyone builds them — environment matchers are 0.20%/1.18% of traversal instances on Amharic/Mbugwe (this is #515), and the analysis rules' 93.95% is 99.7% alternate captures because AnalysisAffixProcessRule and AnalysisCompoundingRule set AllSubmatches and enumerate morph boundaries deliberately.

docs/rejected-optimizations/two-pass-nondeterministic-traversal.md now cross-references it: that experiment kept bindings in the key and rebuilt captures and was too expensive; this one drops both and is unsound. Opposite errors, same boundary.

🤖 Generated with Claude Code

@johnml1135

Copy link
Copy Markdown
Collaborator Author

2026-09-23 round: the missing Aweti speedup was reduplication, not traversal

One new row closes the gap #511 opened. Aweti index 182 had never completed here, on any stack we tried (#491, #494, #511, FieldWorks #1148, forced final templates, FLEx's parallel/no-memo mode). A copy-agreement prune in reduplication unapplication makes it finish in 9 s, and the whole Aweti list runs 4.95× faster with 0 divergences across five grammars. Proposed as #519 (on by default) and shipped in PanGloss v0.4.0.

New row

# Attempt Mechanism Result Verdict
23 Copy-agreement prune Reject a reduplication match whose copies can't unify segment by segment; keep copies with optional nodes or modifications Aweti 182: OOM → 9 s. Aweti list: 172 → 192 complete, 4.95× on the shared 172. Mbugwe ~0.2% fewer candidates; Sena/Amharic untouched Retained, on by default (#519); PanGloss v0.4.0

Why the ledger missed it: every prior row measured Sena, Mbugwe and Amharic, where full-copy rules are absent or rare. The 19 docs mention reduplication once (the probe-design classifier, marked "Unknown").

What else this round settled

  • The Aweti blow-up is 12 full-copy rules (mrule25-mrule36), not traversal. Analysis captures each copy independently and uses only the first, so every split becomes an analysis. Across the list, 99.7% of the matches that can be judged disagree.
  • Add allMatches to Traverse function to improve performance #511 as reviewed was unsound and gave no gain on our grammars. The old key produced 112/20k fuzz divergences; a narrowed form was fuzz-clean but no faster. See "Open since this round" for the redesign.
  • #1148 changes 1 entry / 1 allomorph on Aweti and Sena, and no speed. It can also drop a co-occurrence prohibition that references a removed duplicate MSA (reproduced).
  • FLEx never runs the Memoize HermitCrab's sequential analysis cascade #456 memo. FieldWorks (HCParser.cs:178) and the HC tool construct Morpher with the default parallelism, and the memo only engages at degree 1. On 21 completing Aweti words, the default mode uses 2.93× the CPU of the memo mode, with identical analyses. That makes it a separate FieldWorks change worth measuring.
  • Master is not worse than Maxwell's build. FLEx's own mode (parallel, no memo) also OOMs on index 182. His reported 83 s → 3 s and 4 → 1 min match this prune's effect; how his Add allMatches to Traverse function to improve performance #511/#1148 stack got there is still unexplained.
  • Closed, don't revisit:
  • Open, research-only:
    • Lazy composed FST for the phonology (unmeasured).
    • Two-way FSTs for productive copying: literature only, no HC mapping yet.

PanGloss v0.4.0 (same prune, Rust)

Release binaries, same machine, v0.3.3 vs v0.4.0: identical analyses on all 7,455 words both complete across the five grammars. Aweti: 172 -> 191 words complete within 20 s, 5.3x on the 172 shared words, median word 188 -> 26 ms, word-list index 182 60 s timeout -> 2.7 s. Sena, Amharic and Mbugwe are within noise to about 10% slower; a per-call allocation on rules that copy nothing is being removed.

Open since this round

Method notes worth keeping

  • Parity used gloss/allomorph/index signatures. Morpheme.Id is empty in FLEx exports.
  • The fixture edge-cases-zero-width-morpheme-identity-stability is nondeterministic on master (off/off hashes differ); exclude it from any off/on diff.
  • A seeded fuzz that compares "baseline minus classified" against "pruned" only tests the mechanism, not soundness. Soundness came from conformance, the parameterized ReduplicationRules/ModifyFromInputRules tests, three mutants, and the real-grammar sweep.

🤖 Generated with Claude Code

johnml1135 and others added 2 commits September 25, 2026 18:56
Full docs-only record of the September 2026 HermitCrab optimization round:
22 ledger attempts (2 earlier shipped wins, 12 closed as measured-does-not-pay,
3 closed before building, 5 reopened this round), the probe design needed to
rebuild the measurement instrumentation from scratch, the per-mechanism
disposition (9 retained mechanisms with individual records, 4 rejected with
individual records), and the branch/worktree inventory.

The round's own conclusion, from the disposition doc: do not merge the branch
as a package (896 lines across 28 files, not reviewable as one change, and the
package speedup could not be attributed to its parts); isolate what's worth
isolating separately, and refocus on PR #491 and grammar hygiene where the
large returns are. Kept here as a durable "do not retry this" record, not as
a merge candidate.

Copied from branch perf/hc-optimization-archive onto current master, replacing
the stale two-file version of this PR. Minimal fixups only: the ledger,
probe-design, disposition, and counts docs were read in full to check for (a)
references to branches deleted since the archive was cut and (b) private
corpus data. Found and fixed: the disposition doc's several references to
`perf/hc-edge-prefilter`, whose own branch and archive tag have both since
been deleted; the mechanism itself is not lost, since it also lives in the
whole-round archive (branch
`origin/perf/hc-optimization-archive`) copied in by this same commit, so the
references now name that ref instead of calling the mechanism unrecoverable.
The branch-wide parity sentence in the counts doc is qualified to say what
`WordAnalysisSignature` actually verifies on the FLEx exports (morph counts,
since `Morpheme.Id` is empty), matching the caveat the disposition doc already
recorded. The finished execution-plan checklist is dropped now that its
outcome is captured in the disposition record instead. No private word forms
or corpus data found; the docs already record only word-list indices, counts,
and ratios, per the round's own scrubbing note. No findings, numbers, or
conclusions were otherwise altered. Docs only, no code changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
PR #511 proposes skipping traversal instances whose (State, AnnotationIndex)
was already pushed. The motivating pathology is real and reproduces on a
2-state fsa: Advance forks per Optional annotation, so instances are
exponential in optional count and independent of the state count.

The key is too coarse. A 20,000-case differential fuzz finds 112 cases where
Match() changes, 108 from ignored variable bindings and 4 from shortened
ranges on the deterministic path, with AllMatches().First() identical in all
20,000 as the control. A narrowed gate -- deterministic method, no capture
groups -- measures 0 divergences with the full bound preserved.

Two censuses close the extensions: environment matchers are 0.20%/1.18% of
traversal instances on Amharic/Mbugwe, and the analysis rules' apparent 93.95%
collapse is 99.7% alternate captures, because AnalysisAffixProcessRule and
AnalysisCompoundingRule set AllSubmatches and enumerate morph boundaries on
purpose. That makes a fourth row for the apparent-vs-sound table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@johnml1135
johnml1135 force-pushed the docs/hc-optimization-ledger branch from 47318eb to 5bbfa01 Compare September 25, 2026 22:57
@johnml1135 johnml1135 changed the title HermitCrab optimization round record: 22 attempts, 9 retained mechanisms, no net speedup as a package HermitCrab optimization round record: 23 attempts; a faster package that is not mergeable Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants