Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
7 changes: 7 additions & 0 deletions .agents/skills/polyxml-core-engine/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -272,5 +272,12 @@ lexical-list obligations before traversing repeated values. Ordinary nested
records validate their own fields. Lexical lists append tokens to one string;
retain whitespace-token rejection, escaping and empty-list behavior.

An explicit null in the ordered mixed item stream is a present nil element,
including for nested branches. Serialize it with a local instance namespace
binding; avoid shadowing the element QName's prefix. Keep empty strings distinct
from nulls and preserve item order. Mixed branch schemas do not retain nillable
constraints, and the dynamic reader remains permissive; this round-trip behavior
does not establish nillability validation or ordinary optional-field nil output.

For measured changes, use
[polyxml-runtime-investigation](../polyxml-runtime-investigation/SKILL.md).
35 changes: 35 additions & 0 deletions .agents/skills/polyxml-runtime-investigation/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,11 @@ separate from the schema-driven `PolyValue` runtime.
also matches deserialization. Remove stale results only in the runner's owned
scratch directory before measurement, not from retained raw evidence.
- Check rich enums, patterns, lists and mixed branches alongside plain records.
The `mixed.rs` consumer checks text/scalar/enum/pattern/nested ordered items.
Text/GeneralRef/CData boundaries may coalesce when written: compare adjacent
text's concatenated value, rather than treating token segmentation as XML
semantics. Keep nil correctness tests separate from baseline timing consumers
when the baseline cannot serialize nil.
Warm pattern-cache gains do not establish cold-start, churn or concurrent
throughput. Entry/key bounds do not constitute a compiled-regex byte budget.
- Preserve validation, errors, split Text/CData/GeneralRef handling, nil reads,
Expand All @@ -42,3 +47,33 @@ separate from the schema-driven `PolyValue` runtime.
- Publish positive and negative experiments with exact source revisions, raw
evidence and remaining regressions. Run the full quality gate before pushing
the experiment branch; a branch request does not authorize a merge to main.

For mixed schemas with many possible child tags, measure the linear branch scan
with `mixed_branches.rs`, including sparse one-item controls. A temporary index
can borrow kind keys and branch references for large repeated payloads without
persisting stale mutable metadata. Preserve the original first-match behavior
for duplicate kind names; test metadata edits, tagged records, nil and unknown
kinds. Keep small tables on a linear path and measure the chosen crossover.

Check the indexing cutoff with short documents too. A 256-branch table built for
64 mixed items was 30% slower, despite a 54% gain at 1,000 items. Requiring at
least 64 items and half as many items as branches avoids that observed crossover
regression. Retain 32/64/128-item controls, and keep the cutoff heuristic separate
from semantic behavior: duplicate first-match ordering and metadata edits still
need independent correctness tests.

Also measure concentrated tag reuse and text-heavy content. Even above the
size cutoff, repeated first-branch hits were 11–20% slower with a hash table.
A bounded, evenly spaced sample can select linear lookup when the observed
branch-search work is low. Treat sampling as a performance heuristic only;
normal writer validation must still examine every item and preserve errors.

Offset samples within their bins and include a power-of-two item count: a
strictly evenly spaced sample can repeatedly hit the first tag when the choice
cycle divides its stride, overlooking a distributed payload's lookup work.

For large investigations, retain Criterion process directories in per-experiment
archives rather than thousands of loose JSON files in a PR. Verify the original
manifest, archive every file with its relative path, read back and compare each
member's checksum before removing loose copies, and retain member/top-level
manifests plus the packaging script. Extract into scratch to rerun summaries.
32 changes: 32 additions & 0 deletions benchmarks/rust-runtime-investigation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,3 +113,35 @@ For the generated sensor control, run the XML/Serde regression runner and use
its output directory. This computes each process's median of seven timing
samples, then the median across processes; it retains the process deltas. The
statistic differs from the Criterion process means above.

## Ordered mixed content

Use `--harness benchmarks/rust-runtime-investigation/mixed.rs` for 1,000
text, scalar, enum, warm-pattern or nested mixed items. Every decoded payload
is checked before timing. Adjacent text events can be coalesced on output, so
the text case compares the concatenated text value after a round trip; element
cases compare the complete ordered value. These cases contain no nil items and
can compare revisions whose nil writer is broken. Cover nil output separately
with `test_mixed_nil`; a timing harness must not turn a known correctness
failure into a performance result.

`mixed_branches.rs` varies the child-tag table across 1, 16, 64 and 256 entries
and the payload across 1, 32, 64, 128, 1,000 and 1,024 items. It checks every kind against its
expected wire name and every integer before timing, then checks the full
round trip. Use `^branch_write/.*/1000$` to screen repeated writes; retain sparse
and read controls before promoting a branch-lookup optimization. A temporary
index trades a per-container allocation for faster repeated lookup. It must
preserve first-match behavior and rebuild after mutable metadata edits.

The mixed branch consumer also includes 32, 64 and 128 items to check index
construction cost. The trigger requires at least 64 branches, at least 64 items,
and items numbering at least half the branch count. This is a measured heuristic,
not a universal break-even guarantee; tag distribution and text-only content can
affect how much lookup work is saved.

The final selector samples at most 16 items spread across the payload, requiring at least
four tagged samples and an average linear-search depth of 16 before indexing.
This avoids the observed repeated-first-tag regression and skips a table for
text-only content. It is a heuristic, so retain concentrated/distributed cases
and the actual sampled payload along with timing evidence. All items still pass
through the normal writer validation.
120 changes: 120 additions & 0 deletions benchmarks/rust-runtime-investigation/mixed.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,120 @@
use criterion::{
criterion_group, criterion_main, BenchmarkId, Criterion, SamplingMode, Throughput,
};
use polyxml::schema::ModelSchema;
use polyxml::schema_parser::XsdParser;
use polyxml::{deserialize, serialize, PolyValue};
use std::{hint::black_box, sync::Arc};

fn fixture(kind: &str, count: usize) -> (Arc<ModelSchema>, Vec<u8>) {
let xsd = r#"<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema">
<xs:simpleType name="State"><xs:restriction base="xs:string"><xs:enumeration value="Ready"/><xs:enumeration value="Done"/></xs:restriction></xs:simpleType>
<xs:simpleType name="Code"><xs:restriction base="xs:string"><xs:pattern value="[A-Z]{3}[0-9]{3}"/></xs:restriction></xs:simpleType>
<xs:element name="Root"><xs:complexType mixed="true"><xs:choice minOccurs="0" maxOccurs="unbounded">
<xs:element name="Text" type="xs:string"/>
<xs:element name="Number" type="xs:int"/>
<xs:element name="State" type="State"/>
<xs:element name="Code" type="Code"/>
<xs:element name="Child"><xs:complexType><xs:sequence><xs:element name="Value" type="xs:string"/></xs:sequence></xs:complexType></xs:element>
</xs:choice></xs:complexType></xs:element>
</xs:schema>"#;
let ir = XsdParser::new().parse_str(xsd).unwrap();
let schema = ModelSchema::from_ir(&ir, Some("Root")).unwrap();
let mut xml = String::from("<Root>");
for index in 0..count {
let part = match kind {
"text" => format!("text-{index}&amp;<![CDATA[more]]>"),
"scalar" => format!("<Text>item-{index}&amp;<![CDATA[more]]></Text>"),
"enum" => "<State>Ready</State>".into(),
"pattern" => format!("<Code>ABC{:03}</Code>", index % 1000),
"nested" => format!("<Child><Value>item-{index}</Value></Child>"),
_ => unreachable!(),
};
xml.push_str(&part);
}
xml.push_str("</Root>");
let value = deserialize(xml.as_bytes(), Arc::clone(&schema)).unwrap();
let items_name = &schema.fields[schema.mixed_content.as_ref().unwrap().items_index].name;
let items = value.get(items_name).and_then(PolyValue::as_list).unwrap();
if kind == "text" {
// Text, GeneralRef and CData become separate ordered items.
assert_eq!(items.len(), count * 3);
for (index, items) in items.chunks_exact(3).enumerate() {
for (item, expected) in
items
.iter()
.zip([format!("text-{index}"), "&".into(), "more".into()])
{
assert_eq!(item.get("kind").and_then(PolyValue::as_str), Some("#text"));
assert_eq!(
item.get("value").and_then(PolyValue::as_str),
Some(expected.as_str())
);
}
}
} else {
assert_eq!(items.len(), count);
for (index, item) in items.iter().enumerate() {
let value = item.get("value").unwrap();
match kind {
"scalar" => assert_eq!(value.as_str(), Some(format!("item-{index}&more").as_str())),
"enum" => assert_eq!(value.as_str(), Some("Ready")),
"pattern" => assert_eq!(
value.as_str(),
Some(format!("ABC{:03}", index % 1000).as_str())
),
"nested" => assert_eq!(
value.get("value").and_then(PolyValue::as_str),
Some(format!("item-{index}").as_str())
),
_ => unreachable!(),
}
}
}
let output = serialize("Root", &value, &schema, None).unwrap();
let reparsed = deserialize(&output, Arc::clone(&schema)).unwrap();
if kind == "text" {
// XML writers coalesce adjacent text segments. Compare their text value
// rather than treating the reader's segmentation as XML semantics.
let text = |value: &PolyValue| -> String {
value
.get(items_name)
.unwrap()
.as_list()
.unwrap()
.iter()
.map(|item| item.get("value").unwrap().as_str().unwrap())
.collect()
};
assert_eq!(text(&value), text(&reparsed));
} else {
assert_eq!(value, reparsed);
}
(schema, xml.into_bytes())
}

fn benchmarks(c: &mut Criterion) {
for operation in ["mixed_read", "mixed_write"] {
let mut group = c.benchmark_group(operation);
group.sampling_mode(SamplingMode::Flat);
for kind in ["text", "scalar", "enum", "pattern", "nested"] {
let (schema, xml) = fixture(kind, 1000);
let value = deserialize(&xml, Arc::clone(&schema)).unwrap();
group.throughput(Throughput::Bytes(xml.len() as u64));
group.bench_function(BenchmarkId::new(kind, 1000), |b| {
if operation == "mixed_read" {
b.iter(|| {
black_box(deserialize(black_box(&xml), Arc::clone(&schema)).unwrap())
});
} else {
b.iter(|| {
black_box(serialize("Root", black_box(&value), &schema, None).unwrap())
});
}
});
}
group.finish();
}
}
criterion_group!(benches, benchmarks);
criterion_main!(benches);
74 changes: 74 additions & 0 deletions benchmarks/rust-runtime-investigation/mixed_branches.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
use criterion::{
criterion_group, criterion_main, BenchmarkId, Criterion, SamplingMode, Throughput,
};
use polyxml::schema::ModelSchema;
use polyxml::schema_parser::XsdParser;
use polyxml::{deserialize, serialize, PolyValue};
use std::{hint::black_box, sync::Arc};

fn fixture(branches: usize, count: usize) -> (Arc<ModelSchema>, Vec<u8>) {
let declarations = (0..branches)
.map(|i| format!(r#"<xs:element name="B{i}" type="xs:int"/>"#))
.collect::<String>();
let xsd = format!(
r#"<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema"><xs:element name="Root"><xs:complexType mixed="true"><xs:choice minOccurs="0" maxOccurs="unbounded">{declarations}</xs:choice></xs:complexType></xs:element></xs:schema>"#
);
let ir = XsdParser::new().parse_str(&xsd).unwrap();
let schema = ModelSchema::from_ir(&ir, Some("Root")).unwrap();
let body = (0..count)
.map(|i| format!("<B{}>{i}</B{}>", i % branches, i % branches))
.collect::<String>();
let xml = format!("<Root>{body}</Root>").into_bytes();
let value = deserialize(&xml, Arc::clone(&schema)).unwrap();
let mixed = schema.mixed_content.as_ref().unwrap();
let field = &schema.fields[mixed.items_index].name;
let items = value.get(field).and_then(PolyValue::as_list).unwrap();
assert_eq!(items.len(), count);
for (i, item) in items.iter().enumerate() {
assert_eq!(item.get("value"), Some(&PolyValue::Int(i as i64)));
let kind = item.get("kind").and_then(PolyValue::as_str).unwrap();
let branch = mixed
.branches
.iter()
.find(|b| b.variant_name == kind)
.unwrap();
assert_eq!(branch.xml_name, format!("B{}", i % branches).as_bytes());
}
let output = serialize("Root", &value, &schema, None).unwrap();
assert_eq!(value, deserialize(&output, Arc::clone(&schema)).unwrap());
(schema, xml)
}

fn benchmarks(c: &mut Criterion) {
for operation in ["branch_read", "branch_write"] {
let mut group = c.benchmark_group(operation);
group.sampling_mode(SamplingMode::Flat);
for branches in [1, 16, 64, 256] {
for count in [1, 32, 64, 128, 1000, 1024] {
let (schema, xml) = fixture(branches, count);
let value = deserialize(&xml, Arc::clone(&schema)).unwrap();
group.throughput(Throughput::Bytes(xml.len() as u64));
group.bench_with_input(
BenchmarkId::new(format!("branches_{branches}"), count),
&count,
|b, _| {
b.iter(|| {
if operation == "branch_read" {
black_box(
deserialize(black_box(&xml), Arc::clone(&schema)).unwrap(),
);
} else {
black_box(
serialize("Root", black_box(&value), &schema, None).unwrap(),
);
}
});
},
);
}
}
group.finish();
}
}
criterion_group!(benches, benchmarks);
criterion_main!(benches);
Loading