perf: keep shared-memory locking out of line - #72
explodingcamera merged 2 commits into
Conversation
with_memory! locked a shared memory inline and held the guard to the end of the access, so every load and store handler contained the lock and unlock calls and saved extra callee-saved registers on every access, shared memory or not. Run the shared-memory body in a cold, out-of-line helper instead; the ordinary-memory path stays inline. The three bodies that used `?` for a unit result (the memory.copy range checks and active data segment initialization) now return the Result and apply `?` outside the macro, so a body means the same thing in both paths.
Atomic operations usually target a shared memory, so for them the out-of-line path was the common one and cost an extra call: about 1-2% more cycles per atomic read-modify-write. They now use with_memory!(@lock_inline ...), the previous inline form without the cold_path() hint. Per access, shared memory, M4 (P-cores / E-cores), against next: atomic RMW -5.6% instructions, -3.4% / -7.8% cycles; plain load and store -3.2% instructions, +1.9% / -6.9% cycles. Ordinary memory: -5.8% instructions, -5.3% / -5.0% cycles.
|
Great! Yea I also noticed some perf regressions when introducing shared memories that I want to revisit later, having them be cold is totally fine. I'm thinking of maybe even putting some wasm features like threads and simd behind feature flags to improve performance and code size when they aren't needed at all. |
I'll keep that compile-time configurability in mind for the optimized SIMD stuff I'm working on now. For xmrsplayer, SIMD dispatch being hyper efficient was one of the things I worked on a lot in my WAMR fork, and in other runtimes I've worked on in the last ~10 years. |
Since the threads proposal landed,
with_memory!takes a shared memory's lock inline and holds the guard until the access ends, and this actually dropped tinywasm's leaderboard placement in my WasmBench app on multiple devices. With the inline lock, load and store handler contain thelockandunlockcalls, even thoughcore::hint::cold_path()marks that branch cold. To keep values alive across those calls, each handler saves extra callee-saved registers on every access, whether or not the memory is shared.This runs the shared-memory body in a
#[cold],#[inline(never)]helper that takes the lock. (There is acore::hint::cold_path()in there, but when I looked at the disassembled code for the Release build, it didn't affect codegen and Rust docs verify that its not necessarily meant to influence codegen.) The ordinary-memory path stays inline, which is seemingly more common in the wild, but if you have some specific workload you're optimizing for and need both paths to be equally fast, I can go back to the drawing board.Everything measured with XCode 27's PMU template, the same approach I used for measuring and optimizing other projects (WAMR, wgpu-native, femtovg, BabylonNative, xmrsplayer, etc) on this set of hardware.
On aarch64 (iPhone XS A12 codegen,
nightly-tail-calls):I32StoreI32LoadThree memory.copy / data-segment bodies used
?inside the macro for a unit result. They now return theResultand apply?outside the macro, so a body means the same thing on both paths.iPhone 12 (A14) efficiency cores. Per-call medians of the kernel's retired-instruction and cycle counters, 5 interleaved runs of each build:
**Apple Watch Series 10.**Per-call medians of the kernel's retired-instruction and cycle counters:
The screenshot is one 1-second run. Full data: tinywasm on the Apple Watch Series 10.