Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
0d7a16c
perf: avoid lexical-list scans over ordinary repeated records
nth-bailey Oct 3, 2026
246fcf5
perf: select constrained fields before serializer prevalidation
nth-bailey Oct 3, 2026
a6ecc08
perf: retain scalar field indices instead of cloning schema metadata
nth-bailey Oct 3, 2026
dd3bbc7
perf: reuse anchored scalar regexes in a bounded concurrent cache
nth-bailey Oct 3, 2026
dafb7dd
perf: stream lexical-list formatting into one output string
nth-bailey Oct 3, 2026
7f50cbd
perf: append XML events directly to the output byte vector
nth-bailey Oct 3, 2026
d682520
perf: unify scalar parsing around indices into frame-owned schemas
nth-bailey Oct 3, 2026
e8bf306
test: isolate mixed parser coverage from existing schema and nil writ…
nth-bailey Oct 3, 2026
c65d47d
bench: retain repeatable runtime diagnostics and constraint workloads
nth-bailey Oct 3, 2026
5a6be5c
bench: expose scalar pattern cache working-set boundaries
nth-bailey Oct 3, 2026
8cd0950
bench: record input byte lengths for pattern working sets
nth-bailey Oct 3, 2026
ac485e0
bench: snapshot diagnostic harnesses and reject uncommitted core sources
nth-bailey Oct 3, 2026
45f5af3
bench: summarize generated controls and document preflight rebuild ha…
nth-bailey Oct 3, 2026
19e7628
style: sort diagnostic summarizer imports
nth-bailey Oct 3, 2026
c2ea06b
perf: avoid cyclic scalar-pattern cache misses with salted victim sel…
nth-bailey Oct 3, 2026
7f845ae
docs: record the cache eviction experiment and its bounded working set
nth-bailey Oct 3, 2026
1819c91
docs: publish runtime optimization comparisons and retained experimen…
nth-bailey Oct 3, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .agents/skills/polyxml-benchmark-workflow/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,3 +66,9 @@ description: Use when running, adding, or publishing PolyXML language and runtim
revision order and repeat uncertain results before calling a regression.
Check per-process medians as well as pooled samples; do not treat correlated
samples as independent process repetitions or call noise a speedup.

11. For runtime optimization experiments, consult
[polyxml-runtime-investigation](../polyxml-runtime-investigation/SKILL.md).
Its allocation and Callgrind consumers are diagnostic only. Confirm wins
with the uninstrumented Criterion runner, and anchor write-only filters as
`^serialization/` because `serialization` also matches deserialization.
24 changes: 24 additions & 0 deletions .agents/skills/polyxml-core-engine/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -250,3 +250,27 @@ records need not be walked for lexical-list validation at the outer field when
its value type cannot be a lexical list. Preserve all constraints, mutable-schema
semantics and full read/write regressions; source inspection alone does not prove
which change caused a measured slowdown. Keep generated and dynamic timings separate.

## Runtime hot paths

Active scalar parser state stores field or mixed-branch indices into the schema
owned by its stack frame. Resolve metadata at the end event rather than cloning
rich enum/pattern definitions per occurrence. Keep split text, nil and nested
mixed-content tests when changing this state.

Scalar regex validation uses the private process-wide `pattern` cache, keyed by
the exact original pattern and preserving anchored matching. Compilation and
matching run outside its lock. The cache admits 16 entries and keys up to 4,096
bytes; larger keys bypass caching. These bound entries and key retention, not
compiled-regex memory. At capacity it selects a victim using the existing
randomized hasher and an eviction nonce; FIFO would miss on every access for a
cycle of 17 patterns. Public schema metadata remains mutable before sharing,
so avoid cached flags that become stale after edits.

The serializer writes directly into an append-only byte vector and checks for
lexical-list obligations before traversing repeated values. Ordinary nested
records validate their own fields. Lexical lists append tokens to one string;
retain whitespace-token rejection, escaping and empty-list behavior.

For measured changes, use
[polyxml-runtime-investigation](../polyxml-runtime-investigation/SKILL.md).
44 changes: 44 additions & 0 deletions .agents/skills/polyxml-runtime-investigation/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
---
name: polyxml-runtime-investigation
description: Diagnose and experimentally optimize regressions in PolyXML's dynamic Rust XML runtime using allocation counts, instruction profiles, and same-host latency comparisons.
---

# Dynamic XML runtime investigation

Read [the investigation toolkit](../../../benchmarks/rust-runtime-investigation/README.md)
for exact commands and diagnostic boundaries. Keep generated XML/Serde consumers
separate from the schema-driven `PolyValue` runtime.

- Start with an explicit detached baseline and a committed candidate on a branch.
Keep core sources stable while consumers build and run. Retain small experiments
independently so a failed hypothesis can be backed out without losing evidence.
- Use the counting allocator to distinguish allocation growth from CPU work.
Requested bytes are cumulative requests, not peak or retained memory. Larger
schema types can grow a fixed frame reserve without increasing allocation count.
- Use Callgrind's `measured_region` toggle for a fixed operation count. Its simulated
instruction count identifies candidates; verify changes with uninstrumented
Criterion consumers. Extracting Valgrind locally is an option when installation
would unnecessarily change the host.
- Preflight companion tooling before timing. Even `uv run runner.py --help` can
synchronize and rebuild an editable binding. After preparing its environment,
use `uv run --no-sync` for companion runs. If a build overlaps timing, retain
and mark that attempt as contaminated, then repeat it.
- Run heavy work serially, with one Cargo worker and the repository memory cap.
Short comparisons screen hypotheses. Repeat retained improvements with longer
measurements, alternating revision order and preserving every raw sample.
- Anchor Criterion filters: `^serialization/` means writes only; `serialization`
also matches deserialization. Remove stale results only in the runner's owned
scratch directory before measurement, not from retained raw evidence.
- Check rich enums, patterns, lists and mixed branches alongside plain records.
Warm pattern-cache gains do not establish cold-start, churn or concurrent
throughput. Entry/key bounds do not constitute a compiled-regex byte budget.
- Preserve validation, errors, split Text/CData/GeneralRef handling, nil reads,
nesting and metadata mutation semantics. Borrow scalar metadata from the schema
owned by a frame instead of cloning rich scalar definitions per element.
- Include the 17-pattern cycle when changing eviction: a 16-entry FIFO has
systematic misses, while salted victim selection preserves reuse at the same
capacity. Confirm hot cases too; pressure improvements alone do not establish
universal throughput gains.
- Publish positive and negative experiments with exact source revisions, raw
evidence and remaining regressions. Run the full quality gate before pushing
the experiment branch; a branch request does not authorize a merge to main.
1 change: 1 addition & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ equivalent) and never at this root.
| :--- | :--- | :--- | :--- |
| Rust core engine | [`crates/polyxml-core/benches/`](../crates/polyxml-core/benches/) | [Criterion.rs](https://github.com/bheisler/criterion.rs) | `cargo bench --bench core_benchmarks` |
| Rust same-host XML/Serde regression | [`rust-xml-regression/`](rust-xml-regression/README.md) | Generated consumers + Criterion | See suite README |
| Rust runtime investigation | [`rust-runtime-investigation/`](rust-runtime-investigation/README.md) | Dynamic core constraints, allocations and instruction profiles | See suite README |
| Rust tag dispatch | [`crates/polyxml-core/benches/`](../crates/polyxml-core/benches/) | [Criterion.rs](https://github.com/bheisler/criterion.rs) + `perf stat` | `cargo bench --bench tag_dispatch` ([results](../docs/benchmarks/rust-phf-dispatch.md)) |
| Rust end-to-end dispatch | [`rust-phf-e2e/`](rust-phf-e2e/README.md) | Generated decoders | `./benchmarks/rust-phf-e2e/run.sh` |
| Rust dispatch compile cost | [`rust-phf-compile/`](rust-phf-compile/README.md) | Sequential capped release builds, GNU time, size | See suite README |
Expand Down
2 changes: 2 additions & 0 deletions benchmarks/rust-runtime-investigation/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
target/
__pycache__/
115 changes: 115 additions & 0 deletions benchmarks/rust-runtime-investigation/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
# Dynamic Rust XML runtime investigation

These tools diagnose the schema-driven `PolyValue` runtime. Generated Rust XML
and JSON/Serde consumers have separate suites in
[`rust-xml-regression`](../rust-xml-regression/README.md).

Start from a committed core implementation and an explicit detached baseline.
Run builds, profiles and latency measurements serially through `scripts/memcap.sh`.
Preflight companion tools before timing: even `uv run ... --help` can rebuild an
editable extension. Use `uv run --no-sync` after preparing that environment.
The runners use one Cargo worker, systemd memory limits and disabled swap. Keep
source files stable until a comparison finishes; source revisions describe HEAD,
so an uncommitted core edit makes that identification incomplete.

## Latency comparisons

For plain records, use the current core harness for both revisions:

```bash
python3 benchmarks/rust-xml-regression/run_core.py \
--baseline ../polyxml-perf-control \
--output benchmarks/rust-runtime-investigation/target/evidence/final-core \
--rounds 3 --samples 100 --warmup 3 --measurement 5
```

For scalar constraints and content-model validation, use the alternate harness:

```bash
python3 benchmarks/rust-xml-regression/run_core.py \
--baseline ../polyxml-perf-control \
--harness benchmarks/rust-runtime-investigation/constraints.rs \
--output benchmarks/rust-runtime-investigation/target/evidence/final-constraints \
--rounds 3 --samples 100 --warmup 3 --measurement 5
python3 benchmarks/rust-runtime-investigation/summarize.py \
benchmarks/rust-runtime-investigation/target/evidence/final-constraints
```

Both harnesses check decoded values and XML round trips before timing. The
constraint harness measures 1,000 repeated values: a 16-member enum, one scalar
pattern, a string restriction with a pattern and length, lexical integer lists,
and a content regex. Pattern cases measure **warm reuse**, including value
validation but excluding the first compilation. They do not characterize cold
patterns, cache churn, many-pattern schemas or multithreaded throughput.

Criterion filters are regexes: use `--filter '^serialization/'` for writes only.
The unanchored filter `serialization` also matches `deserialization`. The runner
clears its own Criterion directory before each process, retains separate results
and alternates baseline/candidate order. Short runs (`--rounds 2 --samples 50
--warmup 1 --measurement 2`) screen hypotheses; confirm useful changes with longer
runs. The summarizer takes the median of independent process mean estimates and
retains each process's confidence interval and delta. It does not manufacture a
confidence interval across process repetitions.

## Allocation diagnostics

```bash
python3 benchmarks/rust-runtime-investigation/run.py \
--repo ../polyxml-perf-control --label control \
--output benchmarks/rust-runtime-investigation/target/evidence/control
python3 benchmarks/rust-runtime-investigation/run.py \
--repo . --label candidate \
--output benchmarks/rust-runtime-investigation/target/evidence/candidate
```

The standalone release consumer enables a counting allocator only around the
measured operation. It counts allocations, reallocations and requested bytes;
requested bytes are cumulative requests, **not peak RSS or retained memory**.
Schema construction, ten warmup operations, and correctness checks are outside
the counted region. The returned value or byte buffer is destroyed in that
region. Use count 0 for the sensor and 1,000/10,000 for catalogs. Diagnostic
binaries contain instrumentation even when counting is disabled; never use
these binaries for latency claims. Public schema/scalar/value sizes are recorded
to distinguish per-frame growth from per-value heap allocation.

## Instruction profiles

With Valgrind installed, profile the same consumer and fixed number of operations:

```bash
scripts/memcap.sh valgrind --tool=callgrind --collect-atstart=no \
--toggle-collect=measured_region \
--callgrind-out-file=target/read.callgrind \
--log-file=target/read-valgrind.txt --error-exitcode=93 \
benchmarks/rust-runtime-investigation/target/candidate/target/release/xml-diagnostics \
read 1000 25 instructions
callgrind_annotate --auto=no --inclusive=no --threshold=95 \
target/read.callgrind > target/read-annotated.txt
```

`measured_region` is deliberately not inlined. Collection excludes startup and
warmup; the release consumer retains line tables for attribution. Callgrind
counts executed instructions in a simulation; those counts explain hypotheses,
but do not measure hardware cycles or prove a latency improvement. If Valgrind
is extracted into an isolated directory, set `VALGRIND_LIB` to its matching
`usr/libexec/valgrind` directory and invoke its executable by absolute path.

Retain baseline and candidate profiles, source harnesses, lockfiles, build logs,
metadata, every Criterion sample and verification logs under a dated
`docs/benchmarks/data/` directory. Link the resulting report from the benchmark
index. Include experiments that failed to improve performance, and state what
remains slower than the historical baseline.

`cache_boundaries.rs` is a second alternate harness for the cache's working-set
boundary: 1, 16, 17 and 64 distinct patterns, eight occurrences per pattern.
Reads cycle pattern keys to expose FIFO misses above capacity and compare
replacement policies. Writes visit each
field's eight occurrences consecutively, so they can still reuse each compiled
regex locally even when the full schema exceeds capacity. Run it with the same
`--harness` option and retain both cases; field order affects reuse.

For the generated sensor control, run the XML/Serde regression runner and use
`python3 benchmarks/rust-runtime-investigation/summarize.py --generated DIR` on
its output directory. This computes each process's median of seven timing
samples, then the median across processes; it retains the process deltas. The
statistic differs from the Criterion process means above.
70 changes: 70 additions & 0 deletions benchmarks/rust-runtime-investigation/cache_boundaries.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
use criterion::{
criterion_group, criterion_main, BenchmarkId, Criterion, SamplingMode, Throughput,
};
use polyxml::schema::{FieldKind, FieldSchema, ModelSchema, ScalarType, ValueType};
use polyxml::{deserialize, serialize};
use std::{hint::black_box, sync::Arc};

// Cycle through more than the admitted cache capacity to compare eviction policies.
// Schema construction and correctness checks remain outside measurement.
fn fixture(patterns: usize) -> (Arc<ModelSchema>, Vec<u8>) {
let mut builder = ModelSchema::builder("Root");
for i in 0..patterns {
builder = builder.field(FieldSchema::new(
format!("values{i}"),
format!("Value{i}").as_bytes(),
FieldKind::Element,
ValueType::List(Box::new(ValueType::Scalar(ScalarType::Pattern(
Box::new(ScalarType::String),
vec![format!("P{i}-[0-9]{{3}}")],
)))),
));
}
let schema = builder.build();
let mut xml = String::from("<Root>");
for _ in 0..8 {
for i in 0..patterns {
xml.push_str(&format!("<Value{i}>P{i}-123</Value{i}>"));
}
}
xml.push_str("</Root>");
let value = deserialize(xml.as_bytes(), Arc::clone(&schema)).unwrap();
for i in 0..patterns {
let values = value.get(&format!("values{i}")).unwrap().as_list().unwrap();
assert_eq!(values.len(), 8);
assert!(values
.iter()
.all(|value| value.as_str() == Some(format!("P{i}-123").as_str())));
}
let output = serialize("Root", &value, &schema, None).unwrap();
assert_eq!(value, deserialize(&output, Arc::clone(&schema)).unwrap());
(schema, xml.into_bytes())
}

fn benchmarks(c: &mut Criterion) {
for operation in ["cache_read", "cache_write"] {
let mut group = c.benchmark_group(operation);
group.sampling_mode(SamplingMode::Flat);
for patterns in [1, 16, 17, 64] {
let (schema, xml) = fixture(patterns);
let value = deserialize(&xml, Arc::clone(&schema)).unwrap();
group.throughput(Throughput::Bytes(xml.len() as u64));
group.bench_with_input(
BenchmarkId::new("active_patterns", patterns),
&patterns,
|b, _| {
b.iter(|| {
if operation == "cache_read" {
black_box(deserialize(black_box(&xml), Arc::clone(&schema)).unwrap());
} else {
black_box(serialize("Root", black_box(&value), &schema, None).unwrap());
}
});
},
);
}
group.finish();
}
}
criterion_group!(benches, benchmarks);
criterion_main!(benches);
98 changes: 98 additions & 0 deletions benchmarks/rust-runtime-investigation/constraints.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
use criterion::{
criterion_group, criterion_main, BenchmarkId, Criterion, SamplingMode, Throughput,
};
use polyxml::ir::RestrictionFacets;
use polyxml::schema::{FieldKind, FieldSchema, ModelSchema, ScalarType, ValueType};
use polyxml::{deserialize, serialize, PolyValue};
use std::{hint::black_box, sync::Arc};

fn fixture(kind: &str, count: usize) -> (Arc<ModelSchema>, Vec<u8>) {
let scalar = match kind {
"enum" => ScalarType::Enum((0..16).map(|i| format!("STATE{i}")).collect()),
"pattern" => ScalarType::Pattern(
Box::new(ScalarType::String),
vec!["[A-Z]{3}[0-9]{3}".into()],
),
"restricted" => ScalarType::Restricted(
Box::new(ScalarType::String),
Box::new(RestrictionFacets {
patterns: vec!["[A-Z]{3}[0-9]{3}".into()],
length: Some(6),
..Default::default()
}),
),
"list" => ScalarType::List(Box::new(ScalarType::Int)),
"content" => ScalarType::String,
_ => unreachable!(),
};
let mut builder = ModelSchema::builder("Root").field(FieldSchema::new(
"values",
b"Value",
FieldKind::Element,
ValueType::List(Box::new(ValueType::Scalar(scalar))),
));
if kind == "content" {
builder = builder
.content_pattern(polyxml::schema::compile_content_pattern("(?:Value;)*").unwrap());
}
let schema = builder.build();
let mut xml = String::from("<Root>");
for i in 0..count {
let value = match kind {
"enum" => format!("STATE{}", i % 16),
"pattern" | "restricted" => format!("ABC{:03}", i % 1000),
"list" => format!("{i} {} -1", i + 1),
_ => format!("record-{i}"),
};
xml.push_str(&format!("<Value>{value}</Value>"));
}
xml.push_str("</Root>");
let value = deserialize(xml.as_bytes(), Arc::clone(&schema)).unwrap();
let values = value.get("values").unwrap().as_list().unwrap();
assert_eq!(values.len(), count);
for (i, value) in values.iter().enumerate() {
match kind {
"enum" => assert_eq!(value.as_str(), Some(format!("STATE{}", i % 16).as_str())),
"pattern" | "restricted" => {
assert_eq!(value.as_str(), Some(format!("ABC{:03}", i % 1000).as_str()))
}
"list" => assert_eq!(
value,
&PolyValue::List(vec![
PolyValue::Int(i as i64),
PolyValue::Int(i as i64 + 1),
PolyValue::Int(-1)
])
),
_ => assert_eq!(value.as_str(), Some(format!("record-{i}").as_str())),
}
}
let output = serialize("Root", &value, &schema, None).unwrap();
assert_eq!(value, deserialize(&output, Arc::clone(&schema)).unwrap());
(schema, xml.into_bytes())
}
fn benchmarks(c: &mut Criterion) {
for operation in ["constraint_read", "constraint_write"] {
let mut group = c.benchmark_group(operation);
group.sampling_mode(SamplingMode::Flat);
for kind in ["enum", "pattern", "restricted", "list", "content"] {
let (schema, xml) = fixture(kind, 1000);
let value = deserialize(&xml, Arc::clone(&schema)).unwrap();
group.throughput(Throughput::Bytes(xml.len() as u64));
group.bench_function(BenchmarkId::new(kind, 1000), |b| {
if operation == "constraint_read" {
b.iter(|| {
black_box(deserialize(black_box(&xml), Arc::clone(&schema)).unwrap())
});
} else {
b.iter(|| {
black_box(serialize("Root", black_box(&value), &schema, None).unwrap())
});
}
});
}
group.finish();
}
}
criterion_group!(benches, benchmarks);
criterion_main!(benches);
Loading
Loading