Skip to content

perf: keep shared-memory locking out of line - #72

Merged
explodingcamera merged 2 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/cold-shared-memory
Sep 26, 2026
Merged

explodingcamera merged 2 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/cold-shared-memory

Conversation

@matthargett

Copy link
Copy Markdown
Contributor

Since the threads proposal landed, with_memory! takes a shared memory's lock inline and holds the guard until the access ends, and this actually dropped tinywasm's leaderboard placement in my WasmBench app on multiple devices. With the inline lock, load and store handler contain the lock and unlock calls, even though core::hint::cold_path() marks that branch cold. To keep values alive across those calls, each handler saves extra callee-saved registers on every access, whether or not the memory is shared.

This runs the shared-memory body in a #[cold], #[inline(never)] helper that takes the lock. (There is a core::hint::cold_path() in there, but when I looked at the disassembled code for the Release build, it didn't affect codegen and Rust docs verify that its not necessarily meant to influence codegen.) The ordinary-memory path stays inline, which is seemingly more common in the wild, but if you have some specific workload you're optimizing for and need both paths to be equally fast, I can go back to the drawing board.

Everything measured with XCode 27's PMU template, the same approach I used for measuring and optimizing other projects (WAMR, wgpu-native, femtovg, BabylonNative, xmrsplayer, etc) on this set of hardware.

On aarch64 (iPhone XS A12 codegen, nightly-tail-calls):

handler before after
I32Store 188 instructions, 5 saved register pairs, 17 calls 137 instructions, 3 pairs, 8 calls
I32Load 186 instructions, 5 pairs, 16 calls 142 instructions, 3 pairs, 8 calls

Three memory.copy / data-segment bodies used ? inside the macro for a unit result. They now return the Result and apply ? outside the macro, so a body means the same thing on both paths.

iPhone 12 (A14) efficiency cores. Per-call medians of the kernel's retired-instruction and cycle counters, 5 interleaved runs of each build:

row instructions cycles
byte-hash loop over one memory −3.6% −4.2%
convolution 256×256 (scalar) −5.7% −3.3%
sieve (scalar) −4.2% −7.8%
crc32 64 KB (scalar) −4.0% −6.3%
xmrsplayer −2.7% −1.1%
graphql-validation (AS) −3.4% −2.7%
shared memory: plain load/store −4.2% −4.1%
shared memory: atomic RMW −6.3% −10.7%
ordinary memory: plain load/store −5.8% −4.4%
fib(30), no memory ops −0.1% −0.0%

**Apple Watch Series 10.**Per-call medians of the kernel's retired-instruction and cycle counters:

row instructions cycles
byte-hash loop over one memory −4.6% −8.6%
convolution 256×256 (scalar) −5.8% −4.8%
sieve (scalar) −4.2% −8.3%
crc32 64 KB (scalar) −4.0% −6.2%
xmrsplayer −2.9% −3.6%
graphql-validation (AS) −3.3% −2.8%
fib(30), no memory ops −0.2% −0.0%

WasmBench on Apple Watch Series 10, tinywasm with this change

The screenshot is one 1-second run. Full data: tinywasm on the Apple Watch Series 10.

with_memory! locked a shared memory inline and held the guard to the end
of the access, so every load and store handler contained the lock and
unlock calls and saved extra callee-saved registers on every access,
shared memory or not. Run the shared-memory body in a cold, out-of-line
helper instead; the ordinary-memory path stays inline.

The three bodies that used `?` for a unit result (the memory.copy range
checks and active data segment initialization) now return the Result and
apply `?` outside the macro, so a body means the same thing in both
paths.
Atomic operations usually target a shared memory, so for them the
out-of-line path was the common one and cost an extra call: about 1-2%
more cycles per atomic read-modify-write. They now use
with_memory!(@lock_inline ...), the previous inline form without the
cold_path() hint.

Per access, shared memory, M4 (P-cores / E-cores), against next:
atomic RMW -5.6% instructions, -3.4% / -7.8% cycles; plain load and
store -3.2% instructions, +1.9% / -6.9% cycles. Ordinary memory:
-5.8% instructions, -5.3% / -5.0% cycles.
@explodingcamera

Copy link
Copy Markdown
Owner

Great! Yea I also noticed some perf regressions when introducing shared memories that I want to revisit later, having them be cold is totally fine. I'm thinking of maybe even putting some wasm features like threads and simd behind feature flags to improve performance and code size when they aren't needed at all.

@explodingcamera
explodingcamera merged commit a361d8f into explodingcamera:next Sep 26, 2026
11 checks passed
@matthargett

Copy link
Copy Markdown
Contributor Author

Great! Yea I also noticed some perf regressions when introducing shared memories that I want to revisit later, having them be cold is totally fine. I'm thinking of maybe even putting some wasm features like threads and simd behind feature flags to improve performance and code size when they aren't needed at all.

I'll keep that compile-time configurability in mind for the optimized SIMD stuff I'm working on now. For xmrsplayer, SIMD dispatch being hyper efficient was one of the things I worked on a lot in my WAMR fork, and in other runtimes I've worked on in the last ~10 years.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants