Skip to content

[Target: Wasm & TS] WebAssembly Streaming Runtime (polyxml-wasm) for Node.js, Bun, and Browsers #47

Description

@nth-bailey

🎯 Motivation & Background

Currently, PolyXML supports TypeScript interface and schema generation (Zod, Valibot, TypeBox). However, at runtime in JavaScript environments (Node.js, Bun, Cloudflare Workers, modern browsers), developers still rely on pure JavaScript XML parsers (e.g. fast-xml-parser, xml2js, or DOMParser).

In high-throughput microservices and data-intensive browser apps (financial reconciliation, geospatial GIS, telemetry):

  1. Allocation pressure: Compare the heap and GC effects of JS parsers with Wasm byte transfer and object conversion.
  2. Throughput: Measure end-to-end parsing on representative document sizes.
  3. Engine parity: Make PolyXML's Rust transcoder available in Node, Bun, and browsers while measuring Wasm-specific costs.

🛠️ Proposed Solution

Create a new workspace crate: crates/polyxml-wasm utilizing wasm-bindgen and wasm-pack:

1. Incremental Streaming API

Accept byte input as Uint8Array / ArrayBuffer, and process a web ReadableStream incrementally for repeated records:

import { createPolyXml } from '@polyxml/wasm';

const polyxml = await createPolyXml();
const completeDocument = polyxml.xmlToJson(xmlBytes);
for await (const record of polyxml.parseStream(readableStream)) {
  consume(record);
}

2. Fast Path Serialization / Transcoding

Provide in-Wasm transcoding without allocating intermediate JS objects:

// XML <-> JSON directly inside Wasm linear memory
const jsonBytes = polyxml.xmlToJsonBytes(xmlBytes);

🔬 Benchmarking & Tradeoff Analysis

To rigorously evaluate whether the Wasm boundary crossing is worth the architectural cost, a dedicated benchmark suite must be created.

Benchmark Setup & Tooling

  • Harness: reproducible timed runs across Node.js, Bun, and headless Chromium.
  • Payload Tiers:
    • Micro (1 KB - 5 KB): Small single-record payloads (REST / Webhook messages).
    • Medium (500 KB - 2 MB): Typical batch documents (e.g., SEPA pacs.008 payments).
    • Large (50 MB - 200 MB): Heavy telemetry, NeTEx public transit dumps, or defense track logs.
  • Competitors: fast-xml-parser, xml2js, native browser DOMParser.

Metrics to Measure

  1. Throughput (MB/s and ops/sec): End-to-end time to deserialize into JS objects.
  2. Memory Footprint & GC Pauses: Heap deltas (process.memoryUsage().heapUsed) and GC observations where the runtime exposes reliable measurements; report limitations.
  3. Bundle Size Impact: report Wasm and JS wrapper bytes and gzip sizes; distinguish these from an application bundle.

Hypothesized Tradeoffs & Evaluation Criteria

  • The "Small Payload" Boundary Penalty:
    • Hypothesis: For tiny XML payloads (< 5 KB), JavaScript-to-Wasm FFI boundary crossing and string copying into Wasm linear memory may be slower than a lightweight pure JS parser.
    • Decision Rule: If pure JS is faster for micro-payloads, document a hybrid threshold recommendation or provide an inline zero-overhead fast-path.
  • The "Large Payload" Evaluation:
    • Hypothesis to test: Wasm may improve throughput or reduce JS heap pressure for large documents; copies and output object creation still have costs.
    • Decision rule: Report measured throughput, end-to-end allocations, and GC pauses. Document where Wasm helps or hurts; no fixed speedup or GC threshold is required.

✅ Acceptance Criteria

Publishing the package to npm is tracked separately in #71.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions