[feat] 31-perf: the perf probe, core stutter fixes, LOD, quality in the headset (K4) - #249
Merged
Merged
Conversation
- scripts/perf-games.cjs: per game, a fresh page installs the game's modules from their real zips, loads its real .tpscene through the Games-tab path (importSessionZip + requestLoadSession) from the scenes repo at a git ref (PERF_SCENES_REF, default preview-1-17) or a directory, presses Play and the game's own HUD Play/Start, then measures 10 s at CPU throttle x4: draw calls + triangles per display frame (every render() pass summed), geometries/textures, a texture-MB estimate, lights and shadow-casting lights/meshes, meshes/instanced/unculled, p50/p95/p99 frame ms and the heap delta. Markdown + JSON into after-31/31-perf/. - --vr walks a fake XR session's left stick forward the whole window (fakeXR.cjs); --profile adds CDP CPU self-time and allocation sampling (collected objects included) per function, top 30 each. - Baseline on 1.17 (Radeon 890M): Waves p50 700 ms, Dungeon Realms ~10 MB/s heap churn, Football 336 calls + 7 lights, Jam Room 910 calls. No behaviour change (a script), so no counterfactual; the table is the artefact. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Measured with scripts/perf-games.cjs --profile (desktop, CPU x4, Radeon 890M), 1.17 ->
this commit: Waves p50 699.9 -> 16.7 ms, Dungeon Realms p99 116.7 -> 16.8 ms and heap
churn +97 MB/10 s -> none, Football p50 49.9 -> 33.4 ms.
- ModuleContent.svelte: rows (each carrying its THREE group) lived in a deep $state and
`use:rowMenu={row}` made svelte deep_read the action parameter - the WHOLE scene graph
through `parent`, every typed-array index, on every refresh (1 s + every pokeScene),
mounted in every mode incl. VR. 88% of Waves' CPU, ~200 MB/s garbage; Dungeon Realms'
1 Hz hitch. $state.raw, and the 1 s catch-up skips a hidden list.
- dungeonPlay.dungeonData: a whole-scene getObjectByName per frame from the walker,
flier and VR walk in EVERY game (a miss walks everything). Hit cached + re-validated by
parent chain, miss remembered 250 ms; walkable() stops allocating arrays per call.
- flowRuntime: resolveInputs / evalNode input() scanned every edge and, per edge, every
node (O(consumers x edges x nodes) per tick). Indexed per graph array (nodeById,
edgesInto), rebuilt only when a graph changes.
- moveSmoothing: a watching peer searched the tree per eased body and pokeScene'd (every
$objectsGroup subscriber) EVERY frame while a remote physics stream eased. The ease keeps
its object; the UI hears the motion at 10 Hz and once on landing.
- VRObjectsPanel: always mounted, and rebuilt its row list on every scene poke while
closed. Only while open.
- perf-games.cjs: --profile names the nearest app frame calling the heaviest functions;
--vr measures with post/AO off (the headset's direct render), documented as one eye.
Suite perf-stutter (new), one counterfactual guard per fix - with the five fixes reverted:
deep-reads 0 -> 9 getter reads; edge reads 401 -> 20451; dungeon traversals 1/1 ->
100/100; eased frames <=1 search + <=1 poke -> 59 + 30; closed-panel rebuilds 0 -> 8.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- src/lib/lod.js: dense meshes (>= 3000 tris) in objectsGroup and under the module world
root get meshoptimizer-simplified levels (ratios 0.5 / 0.25 / 0.1), built ONCE per asset
in the existing decimation worker (keyed by a content signature, so fifty placements of
one pack piece or ten parses of one GLB share one set). THE TREE IS NEVER TOUCHED: the
swap happens in scene.onBeforeRender / onAfterRender, inside renderer.render() only, so
every serializer, picking, BVH, physics hull and mesh tool sees the source geometry.
Distances are in world RADII of the mesh with a hysteresis band; a `lodBias` store pulls
them in (the quality governor writes it in P3). A geometry swap or an in-place edit
stands the entry down at once and rebuilds after it has held still 1.5 s. Skinned
meshes, morph targets and (in auto mode) instanced meshes are left alone;
`userData.lod = false` opts a mesh out. LOCAL, never replicated.
- src/lib/lodCore.js: the pure rule (pickLevel with hysteresis, levelsToBuild,
normalizeLodOptions, lodBiasFor, geometrySignature) - vitest.
- decimateCore: an optional per-job `errorCap` (absent = the import gate's 0.05,
byte-identical); a far LOD level asks for the count, not a surface error.
- api.lod(object, {ratios, distances, minTriangles}) -> {meshes, ready, remove}, journalled.
- Settings > "Simplify distant models" (lodEnabled, LOCAL, default on); App boots it and
the debug hook gains `lod` (the three tails 219/219/219).
Suite lod (new, 20 checks): levels built (9024 -> 4511 -> 2256 -> 902), the frame's
triangles from 400 m drop 18110 -> 1866 and a middle distance draws a middle level, the
mesh holds its source geometry after a render and toJSON writes only it, a geometry swap
draws whole at once, twins share one cache entry, the bias pulls the edge in, api.lod on
scene-root module content + teardown on deactivate.
COUNTERFACTUAL in-suite: with LOD switched off the far render draws the full 18110.
vitest lodCore (14): hysteresis both ways, radii not metres, bias, floors, signature.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ty (P3, K4)
Until now the adaptive quality governor never ran in VR: its frames came from
sceneBudget's window rAF, which an immersive session suspends, so a Quest sat at full
quality however badly it judged.
- FRAMES: the XR session's own requestAnimationFrame (a second callback beside three's)
feeds every frame interval, deciding every 250 ms.
- THRESHOLDS from the headset's refresh rate (xrThresholds: over = 1.3 budgets, under =
1.12 budgets; 72 Hz -> 18.1 / 15.6 ms), re-read on frameratechange, through a new
governor.setThresholds(). The static VR pair (13.9 / 11.1) could never recover at 72 Hz:
a healthy frame reads 13.9.
- ENTRY FLOOR: auto mode enters a session at level 1 (shadows off, the Quest budget) and
holds it as a recovery floor (governor.setFloor) for the session; exit gives the level
back. The opt-out holds: with auto quality off nothing changes a level.
- RESOLUTION: an XR framebuffer's size is fixed at entry and three already runs maximum
foveation, so the scale the session needed is applied to the NEXT entry
(xrScaleAfter -> setFramebufferScaleFactor at sessionend).
- LOD: every published level sets lod.lodBias = 0.85^level (floor 0.35).
- api.quality = {level (getter, 0 best), max, labels, vr, onChange(fn) -> off}; the
listener fires on level CHANGES only and is released with the module.
- A headset is always "heavy" (governed); the sampler's last reading is a desktop one.
MERGE NOTE (31-game-shell edits the same file): their applyGameQuality pin must win over
the entry floor - `if (gameForcedLevel !== null) return` at the top of startXRQuality's
floor block. Hunks were kept apart from theirs (decideNow / after autoQuality.subscribe).
perf-governor section 6b (8 checks): a session starts at level 1 with shadows off even on
a light scene, judges by 72 Hz, missed XR frames take the next step and pull the LOD bias
in, on-time frames recover to the floor and never past it, exit restores level 0 and the
desktop thresholds, the next-entry scale is kept. COUNTERFACTUAL in-suite: with auto quality
off a session changes no level. vitest: the 72 Hz counterfactual (static pair stuck at
level 1, xrThresholds recovers), missed-frame step + recovery, scale rule.
lod suite 8.3/8.4: api.quality's shape and onChange hearing 2 then 0.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…et, the probe (P4)
- api.quality / api.lod with a worked example (cut your own effects on onChange; levels
for dense module geometry; userData.lod = false opts a mesh out).
- The Quest budget (<= 150 calls, <= 300k tris, <= 2 lights shadows off, no per-frame
allocations) and how to measure a game with scripts/perf-games.cjs.
- The four lessons this round measured: no THREE object inside a Svelte $state (the deep
read was 88% of a game's frame), no per-frame whole-scene lookups, no per-frame
allocation, lights/shadow casters/transmission are draw calls.
Before/after (all seven games, desktop + VR-emulated) is in the lane handover and
after-31/31-perf/perf-{baseline-1.17,after-31,vr-baseline-1.17,vr-after-31}.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Roadmap 31 lane 31-perf — contract K4; items P1, plus the core half of D1 / F1 / W2. Do not merge here; the integrator merges.
What it does
scripts/perf-games.cjsprints the roadmap-31 Performance-protocol table for the seven Games-tab games. It installs the real zips and loads the real .tpscene through the Games-tab path, then plays with real Play + HUD start, at CPU throttle x4. Flags:--vr(fake XR session walking, post off),--profile(CPU self time + the app frame calling it + allocation incl. collected).$state, so Svelte walked the whole scene graph on every refresh. 88% of Waves' CPU; Dungeon Realms' 1 Hz hitch.getObjectByName.pokeScenewhile easing remote physics.lod.js/lodCore.js.api.lod; Settings "Simplify distant models".api.quality {level, max, labels, vr, onChange}.THE TABLES (
node scripts/perf-games.cjs, 10 s of Play, CPU x4, 1280x720, real GPU)Raw:
/home/deck/.code/lanes-30/after-31/31-perf/perf-{baseline-1.17,after-31,vr-baseline-1.17,vr-after-31}.{md,json}(+
perf-profile-1.17.md/perf-p1a.md= the CPU/alloc profiles with callers).Desktop (post stack on), 1.17 -> feat/31-perf:
VR-emulated (
--vr: fake XR session walking the left stick, post + AO off = the headset's directrender; still ONE eye at 1280x720, not a Quest number), 1.17 -> feat/31-perf:
Frames are vsync-quantised (16.7 / 33.3 / 50); heap Δ is noisy per 10 s window (a GC inside it reads
negative) — the churn signal is the 95-97 MB/10 s Dungeon rows collapsing. Against the Quest budget
(<= 150 calls, <= 300k tris, <= 2 lights): football 299-336 calls + 7 lights, jam-room 882-910 calls +
4 lights, stars/untangle/dungeon 3-5 lights are still over — all module/scene content (below).
In a real headset core now also turns shadows off at entry (level 1), which removes the shadow pass.
LOD column: auto LOD engaged on 1 mesh (untangle) — the seven games' meshes are small or skinned, so
LOD is headroom for authored/imported content, not a win on today's games.
MERGE NOTES (integrator)
src/lib/qualityGovernor.jsis also edited by 31-game-shell (gameForcedLevel/applyGameQuality/gameQualityState, an early return in decideNow). Hunks kept apart. Aftermerging add, at the top of the floor block in
startXRQuality:if (gameForcedLevel !== null)skip the floor (their pin wins; agreed with that lane). Their preset levels: high 0 / medium 3 / low 7.
src/lib/moduleSDK.js: I addlod+qualityright afterstorage:; game-shell addsapi.game.*only (agreed, no competing
quality).lodat the END (219/219/219 here) — re-count after the union.decimateCore.simplifyMeshgained an optionalerrorCap(absent = byte-identical).Tests
Held battery green on the pushed head: lod 22, perf-stutter 11, perf-governor 47, scene-budget 34, module-toolbox 46;
npm run buildgreen.New suites:
lod(22),perf-stutter(every guard RED with its fix reverted: 9 deep reads, 20451 edge reads, 100/100 traversals, 59 searches + 30 pokes, 8 panel rebuilds).perf-governorgains section 6b, the headset (8 checks, incl. an opt-out counterfactual). vitest: lodCore (14) + governor XR (4), 293/293 total. check-storage clean; svelte-check 333/47 = the 1.17 baseline.Owed on a Quest 3
🤖 Generated with Claude Code