Skip to content

perf: speed up dynamic XML output and constrained scalars - #137

Merged
nth-bailey merged 17 commits into
mainfrom
perf/xml-runtime-investigation
Oct 4, 2026
Merged

nth-bailey merged 17 commits into
mainfrom
perf/xml-runtime-investigation

Conversation

@nth-bailey

Copy link
Copy Markdown
Collaborator

Dynamic XML parsing clones rich scalar definitions for every occurrence and recompiles each scalar regex for every value. Writing also formats lists through intermediate strings and sends append-only output through a cursor. This change borrows frame-owned scalar metadata by index, reuses anchored regexes in a concurrent cache, formats lists into one string, selects prevalidation obligations before traversing values, and writes directly to the output vector.

The cache retains 16 entries with keys up to 4,096 bytes; larger patterns compile uncached. Compilation and matching occur outside its lock. Salted victim selection avoids systematic misses in a 17-pattern cycle without increasing capacity. Schema edits before sharing, validation errors, split text, nested/mixed order and nil reads retain their behavior.

Three alternating Criterion processes per revision confirm 18–20% faster plain writes and about 45% faster enum reads versus 0.34.6. Warm pattern fixtures improve 42–44× on reads and 92–104× on writes. With 17 distinct patterns, reads improve 6.6× and writes 36×; at 64 patterns, reads are approximately unchanged and writes improve 7×. These are fixture-specific results, and the cache limits entries/key retention rather than total regex memory. Plain catalog reads remain about 5% slower than 0.27.0.

The report links exact source revisions, locks, every retained sample, instruction profiles and positive/negative experiments. Generated model sources are byte-identical; four-round XML/Serde controls show no material batch slowdown, with small micro-operation differences retained in the tables.

Validation:

  • Full quality gate: strict Clippy, 321 Rust tests, 112 Python tests, 100% statement/branch coverage.
  • Refreshed CType sample: 31/31 schema expectations and 23/28 instance round trips; the same five existing failures as the recorded reference.
  • Strict documentation build, benchmark/tool formatting, dependency equality and archive reconstruction checks passed. Builds and measurements ran serially with memory caps; no OOM occurred.

Two newly noticed correctness limitations—escaped XSD enumeration attributes and writing nil mixed scalars—have minimal reproducers showing identical control/candidate errors. They remain separate follow-up work. Main is unchanged; this PR is a draft for reviewing the optimization and its evidence.

@nth-bailey
nth-bailey marked this pull request as ready for review October 4, 2026 03:58
@nth-bailey
nth-bailey merged commit 95ad852 into main Oct 4, 2026
30 checks passed
@nth-bailey
nth-bailey deleted the perf/xml-runtime-investigation branch October 4, 2026 04:14
@nth-bailey
nth-bailey restored the perf/xml-runtime-investigation branch October 4, 2026 04:15
@nth-bailey
nth-bailey deleted the perf/xml-runtime-investigation branch October 4, 2026 04:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant