Skip to content

test(browser): say whether a slow warm channel switch was app work or a runner stall - #696

Draft
loganj wants to merge 1 commit into
mainfrom
larry/warm-switch-settle
Draft

loganj wants to merge 1 commit into
mainfrom
larry/warm-switch-settle

Conversation

@loganj

@loganj loganj commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

🤖

Summary

  • When a browser speed check for switching between already-opened channels (a "warm" switch) misses its 200 ms limit, the failure now says where the time went: app work in the browser, or a stall of the whole test machine.
  • Today the failure only reports a number. Each time it fails, someone must download the CI files and rebuild the timeline by hand to decide whether the app got slower or the shared CI machine paused.
  • The 200 ms limit, the four samples, and what the check measures are unchanged. Nothing is retried or relaxed.

Why now

Main 0e2272c1 failed this check once: the first Alpha switch took 319 ms (run). The evidence says this was the machine, not the app:

  • The same source tree (317b0c32, identical tree hash) passed an hour earlier on its PR run at 104–134 ms. A rerun of the failed job also passed (134–143 ms).
  • In 47 other recent main and PR runs (188 warm samples), every sample was 76–166 ms.
  • In the failed sample, the first message row was visible at 100 ms, which is normal. Then the browser drew no frame for about 136 ms. In the same moment, the test server, a separate process, received an ordinary request 263 ms after the request the browser sent with it (normally 1–2 ms apart). Two processes stopping together points at the runner.
  • Locally, with the browser's CPU slowed 4×, no switch had extra work after its first visible row.

Details

  • Browser side: during each sample the check records the browser's long animation frames (frames longer than 50 ms) and which scripts ran in them. This uses Chromium's Long Animation Frames API; WebKit has no such API, so it reports "long frames unavailable".
  • Machine side: the Playwright test process also serves the app and the relay broker, so the check records the longest event-loop stall of that process during each sample.
  • Both go into the sample's evidence.json record and into the failure message, for example: Alpha warm-switch regression ceiling (1 long frames, 281 ms of script; test process stalled up to 12 ms).
  • I checked both directions with temporary changes: a 250 ms busy loop in the page showed up as one long frame with 250 ms of timer script (and failed the limit); a 250 ms busy loop in the test process showed up as a 252 ms stall, with no long frames. Normal samples show 0 long frames and an 11–17 ms stall (the monitor's resolution).

…nner stall

When a warm channel switch misses its 200 ms ceiling, the failure now
says whether the time went to app work or to the runner. Each sample
records the browser's long animation frames, with the scripts that ran
in them, and the longest event-loop stall of the test process, which
also serves the app and the broker. The ceiling and the samples are
unchanged.

On 0e2272c the first Alpha sample took 319 ms: its first row was
visible at 100 ms, then no frame ran for about 136 ms, while the broker
received a request 263 ms late. Both processes stopped at once, but the
evidence could not show that directly.

Signed-off-by: Larry <627498bd4bd1f281a16431e3c6cce3b5c25b6692798c78672298aefbf2f8f8b5@buzz.block.builderlab.xyz>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant