Skip to content

feat: unified generation options, xs:pattern/xsi:type semantics, base-field inlining fixes, and phf tag dispatch - #56

Merged
nth-bailey merged 10 commits into
mainfrom
cli-and-pattern-validation
Sep 23, 2026
Merged

nth-bailey merged 10 commits into
mainfrom
cli-and-pattern-validation

Conversation

@nth-bailey

@nth-bailey nth-bailey commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Summary (10 commits)

CLI / schema semantics

Codegen correctness fixes

  • 93713b3 fix(rust-codegen): inline xsd:extension base fields in generated structs
  • 5e559ee fix(java-codegen): inline base fields in plain record mode
  • c16e57d fix: simpleContent text codecs, split-text accumulation, and transcoder entity refs

#49: perfect-hash element tag dispatch

  • dce03c1 feat(rust-codegen): --feature phf emits a per-struct __{Struct}ElementId enum + phf::Map dispatch at Event::Start/Event::Empty; default output stays byte-identical. Includes benches/tag_dispatch.rs (criterion, tiers 16/120/600/1500, hit+miss), scripts/perf_stat.{c,sh} hardware counters, tests (18/18 rust codegen incl. 3 new phf tests, 35/35 CLI), verify_codegen.sh phf step, and full results/tradeoff analysis in docs/benchmarks/rust-phf-dispatch.md.
    • Measured: throughput crossover ≈ 835 elements (phf wins above; miss crossover ≈ 400), match's L1I misses grow 277× from 120→1500 tags while phf stays flat; recommendation documented: auto-enable at ≥ ~900 elements.
    • Small-schema cost: +0.04 % .text, +144 B .rodata, rebuild delta within noise.
    • Compile-time/size at the 600/1500 tiers exceeded the safe RAM cap for every variant (scales with field count, not dispatch) → deferred to [Benchmark] Large-schema compile-time & binary-size comparison (600/1500 elements) deferred from #49 #55.
  • 020794f docs: link that caveat to [Benchmark] Large-schema compile-time & binary-size comparison (600/1500 elements) deferred from #49 #55; fix garbled memcap sentence in SKILL.md

Review fixes

  • d6f567d enforces full-value XSD patterns in all seven generators, resolves xsi:type values against in-scope namespace bindings, rejects invalid boolean simpleContent, refreshes Python subclass dispatch after schema caching, and adds regression coverage and documentation.

Benchmark host safety

  • scripts/memcap.sh: portable 60 %-of-available-RAM guard (systemd scope + ulimit -v fallback, nesting-aware) wired into benchmarks/run_all.sh, benchmarks/cli/benchmark.sh, scripts/perf_stat.sh so heavy builds OOM-kill in-cgroup instead of freezing 8-GiB dev boxes. CI never references those entry points.

Verification (full AGENTS checklist, all green)

  • cargo fmt --check · cargo clippy --workspace --all-targets -- -D warnings (0 warnings)
  • cargo test --workspace (all suites)
  • ./scripts/test_codegen.sh (all 7 language suites + CLI integration)
  • ./scripts/verify_codegen.sh (7 ecosystems × backends smoke, incl. new --feature phf + ELEMENT_DISPATCH assertion)
  • Python gate: maturin develop · ruff check + ruff format --check · pytest --cov=polyxml --cov-branch --cov-fail-under=100 → 100.00 %, 72 passed

Issues

Unify backend, style, and feature options across CLI and manifests, retain deprecated aliases, and validate target compatibility before generation. Add shell completion, startup benchmarks, documentation, and regression coverage.

Refs #50
…ns in all 7 codegens (issue #54)

Collapse same-restriction alternatives into one (a)|(b) group in the parser while keeping each derivation step a separate flat entry, so downstream ANDing yields W3C OR-within-restriction, AND-across-steps semantics without changing the IR shape.

Add pattern enforcement where it was missing (Rust, Go, C++), and fix anchoring/escaping in Java (search not full-match), C# (unanchored regex, backslash escapes), and TypeScript TypeBox (duplicate pattern keys -> Type.Intersect). Patterned simple types resolve through primitive_base so validators operate on primitives.

Add test_pattern_codegen.rs, which executes generated validators in all 7 languages, wired into scripts/test_codegen.sh; document the semantics in the codegen skill.
Delete the 7 hidden legacy flags (--zod, --source-gen, --record-kind,
--rkyv, --builder, --codec and the deprecation-warning machinery) plus
their polyxml.toml counterparts; unknown manifest keys now fail loudly
via deny_unknown_fields. --zero-copy stays a first-class visible option
because --feature cannot express false (owned Rust output).

Co-Authored-By: opencode@users.noreply.github.com
… (issue #53)

- ModelSchema gains is_abstract plus a variants registry (OnceLock, keyed
  by type QName local part), populated from XSD extension chains in from_ir
  and from Python subclass hierarchies in the PyO3 layer.
- from_ir flattens base-chain fields into derived runtime schemas
  (mirroring the Java codegen), so dispatched records keep inherited fields.
- Parser dispatches xsi:type at every frame site (root, nested, list,
  empty, iterparse); unknown selectors on abstract types fail loudly
  listing known derivations, plain content attributes named "type" never
  misdispatch.
- Serializer re-emits xsi:type for variant records and declares xmlns:xsi
  whenever dispatch is possible; xml_to_json keeps variant-only fields.
- Python codegen emits Meta.abstract; polyxml.serialize grows target_type
  for base-declared root round trips.
- New docs/guides/polymorphism.md documents the strategy, limitations, and
  escape hatches; core-engine and codegen skills updated.

Tests: 7 Rust (test_issue53_xsi_type.rs), 11 Python (test_xsi_type.py at
100% coverage), test_python_abstract_meta_emission, audit item 7 re-pointed.
- Add flatten_fields(s, ir): walks the base chain root-first with a cycle
  cut, then appends own fields with the most-derived declaration of a
  schema name shadowing inherited duplicates (keeps the simpleContent
  `value` field singular instead of emitting a `value_2` duplicate).
- Use it at all four field-iteration sites — compute_types_with_lifetime,
  emit_struct, the has_patterns probe, and the decode field_metas build —
  so the struct body, `<'a>` propagation, validate_patterns, and both
  codec directions agree on one list. Previously Rust was the only backend
  that dropped inherited fields: a derived InternationalAddress could not
  store or round-trip its base's street/city/country, and a derived type
  whose own fields were all numeric would have emitted Cow<'a, str> in a
  struct declared without the lifetime parameter.
- Regression tests cover two- and three-level chains, lifetime propagation
  from inherited string fields, codec decode/encode of inherited members,
  the simpleContent shadow rule, and circular bases; the codegen skill
  now records the invariant plus two related unfixed gaps found while
  verifying (Java record mode bypasses the model_fields base walk, and
  generated Rust codecs never consume FieldKind::Text).

Tests: 4 new cases in test_rust_codegen.rs (13 green), cargo test
--workspace, test_codegen.sh (TS modules enabled), verify_codegen.sh
across all 7 ecosystems, Python gate at 100% coverage; generated output
additionally compiled and round-tripped in a scratch cargo project
(inherited decode, required-field enforcement, encode parity).
- model_fields always runs the base walk now: the
  `use_records && !emit_builder && !emit_direct_codec` guard made the
  DEFAULT record output declare only a derived type's own components, so
  an InternationalAddress record could not even store its base's
  street/city/country. Builder and direct-codec modes already collected;
  only the plain record path dropped inherited data.
- The local collect walker is replaced by the shared flatten_fields
  helper (moved from rust/mod.rs to codegen/mod.rs), which also brings
  Java its first shadow rule: a derived declaration of the same schema
  name wins, so simpleContent chains emit one `value` component instead
  of a suffixed duplicate in records, classes, and codecs alike.
- POJO/class emission is unchanged: it re-slices the own-fields tail for
  declarations and keeps `extends Base`; builders keep their inherited
  boundary arithmetic.

Regression tests: plain-record component assertions (root-first order,
inherited id/unit present, exactly one value) plus a javac+java E2E that
constructs derived records positionally across complexContent and
simpleContent chains; codegen skill records the fix.

Tests: 2 new cases in test_java_codegen.rs (15 green), cargo test
--workspace (26 suites), test_codegen.sh (TS modules enabled),
verify_codegen.sh across all 7 ecosystems, Python gate at 100%
coverage; fixture records additionally compiled and executed via javac
(three-level chain + simpleContent unit inheritance).
…er entity refs

- emit FieldKind::Text codecs (emit_text_content_parse/_serialize): decode
  reads the element content through the scalar/facet path and encode writes
  it back; only scalar-backed values are emitted (struct-typed value stays a
  documented cross-backend gap pending issue #51 item 4 IR modeling)
- read_element_text (both zero_copy variants) now appends segments and
  resolves GeneralRef/CData like parser.rs::append_general_ref: quick-xml
  splits 'a &amp; b' into Text/GeneralRef/Text, so the old last-wins loop
  silently dropped every entity-bearing or CDATA text value
- transcoder: resolve general refs instead of dropping them via _ => {},
  unescape attribute values, and drop trim_text(true) — per-segment
  trimming ate spaces adjacent to refs (x &amp; y -> x&y); the End arm
  trims the assembled buffer instead
- add unused_assignments to the generated module allow-header (the text
  path assigns its slot unconditionally)
- regression tests: test_rust_simple_content_text_codec,
  test_rust_read_element_text_accumulates_refs_and_cdata,
  test_schemaless_preserves_refs_cdata_and_attr_entities; skills updated
Emit a per-struct __{Struct}ElementId enum plus a phf::Map keyed by tag
name at Event::Start/Event::Empty when the `phf` feature is requested
(#49). Default output stays byte-identical; attrs/unions/enum values
keep match dispatch. Criterion bench (benches/tag_dispatch.rs, tiers
16/120/600/1500, hit+miss) and perf counters (scripts/perf_stat.{c,sh})
show match's L1I misses grow 277x from 120->1500 tags while phf stays
flat; throughput crossover lands near 835 elements, so >=~900 is the
documented auto-enable threshold. Closes the uniquify-parity gap
discovered while testing group emission.

Also adds scripts/memcap.sh, a portable 60%-of-available-RAM guard
(systemd-run scope with ulimit fallback) wired into the benchmark entry
points so heavy builds OOM-kill inside their cgroup instead of freezing
the host. Compile-time/size comparison at the 600/1500 tiers exceeded
the cap for every variant (scales with field count, not dispatch) and is
documented as deferred to a follow-up issue.

Docs: docs/benchmarks/rust-phf-dispatch.md (throughput, counters, size,
threshold, repro), memcap notes in docs/benchmarks.md and
benchmarks/README.md, SKILL.md updated.
The size section's "deferred to a follow-up issue" caveat now points at
the concrete issue. Also fixes a garbled memcap sentence in SKILL.md
(POLYXML_MEMCAP_LEVEL is set automatically, not set by the caller) and
drops an unverifiable exit-code claim from the OOM encounter record.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment