scan_utf8.valid, the json module's is_utf8 (decision 38), ends with utf8.zig's machine over the octets after its last whole block: at most 15, plus up to 3 of a character the last block cuts, taken again. For a buffer shorter than 16 octets, that loop is all of the work.
In that loop the machine's needed field lives in a stack slot, while low and high stay in registers. The loop loads needed for every octet and stores it after every octet from 0x80 up. The tests' ReleaseSafe build (zig build test-json -Drelease) at e631d40 shows it on every runner target:
- aarch64, baseline CPU:
ldrb w2, [sp, #0xc] for each octet, then strb w17, [sp, #0xc] after a continuation octet or strb w18, [sp, #0xc] after a first octet. The M-series host's build uses [x29, #-0x1] the same way.
- x86-64, baseline CPU and the AVX2 object:
movzbl -0x11(%rbp) for each octet, and movb %al, -0x11(%rbp) after an octet from 0x80 up.
In a run of non-ASCII octets, each octet's load of needed waits on the store the octet before it made. Nothing has measured what that costs: bench-json's "UTF-8 validation against simdutf" times whole files, which run this loop once each. The buffers it touches are the ones 9ed8911's fix of the tail's blocks touched.
Not investigated:
- why LLVM keeps
needed in memory. It is the struct's only field narrower than an octet (u2);
- whether the encoder's and decoder's scalar paths, which keep a
Utf8 in their state, reload it the same way inside a call.
Utf8 is the machine zig build test replays against decision 28's proved vectors (utf8_vectors.txt), so a change to its fields must keep that replay passing.
The check: the loop's disassembly on the three runner targets with no stack access per octet, and a timing of is_utf8 over buffers shorter than a group, ASCII and non-ASCII, as ratios within one run.
scan_utf8.valid, the json module'sis_utf8(decision 38), ends with utf8.zig's machine over the octets after its last whole block: at most 15, plus up to 3 of a character the last block cuts, taken again. For a buffer shorter than 16 octets, that loop is all of the work.In that loop the machine's
neededfield lives in a stack slot, whilelowandhighstay in registers. The loop loadsneededfor every octet and stores it after every octet from 0x80 up. The tests' ReleaseSafe build (zig build test-json -Drelease) at e631d40 shows it on every runner target:ldrb w2, [sp, #0xc]for each octet, thenstrb w17, [sp, #0xc]after a continuation octet orstrb w18, [sp, #0xc]after a first octet. The M-series host's build uses[x29, #-0x1]the same way.movzbl -0x11(%rbp)for each octet, andmovb %al, -0x11(%rbp)after an octet from 0x80 up.In a run of non-ASCII octets, each octet's load of
neededwaits on the store the octet before it made. Nothing has measured what that costs: bench-json's "UTF-8 validation against simdutf" times whole files, which run this loop once each. The buffers it touches are the ones 9ed8911's fix of the tail's blocks touched.Not investigated:
neededin memory. It is the struct's only field narrower than an octet (u2);Utf8in their state, reload it the same way inside a call.Utf8is the machinezig build testreplays against decision 28's proved vectors (utf8_vectors.txt), so a change to its fields must keep that replay passing.The check: the loop's disassembly on the three runner targets with no stack access per octet, and a timing of
is_utf8over buffers shorter than a group, ASCII and non-ASCII, as ratios within one run.