Skip to content

audit(token_audit): count "bogus" and "it didn't run" as the same class - #507

Open
EdbertChan wants to merge 1 commit into
stack/EdbertChan/reflect/ui-input-guard-hook-freshness-20260908/script-handed-user-claim-runs--2f495bfefrom
stack/EdbertChan/reflect/ui-input-guard-hook-freshness-20260908/count-bogus-didn-t-run-same-class--065043b6
Open

audit(token_audit): count "bogus" and "it didn't run" as the same class#507
EdbertChan wants to merge 1 commit into
stack/EdbertChan/reflect/ui-input-guard-hook-freshness-20260908/script-handed-user-claim-runs--2f495bfefrom
stack/EdbertChan/reflect/ui-input-guard-hook-freshness-20260908/count-bogus-didn-t-run-same-class--065043b6

Conversation

@EdbertChan

@EdbertChan EdbertChan commented Sep 12, 2026

Copy link
Copy Markdown
Owner

The shortcut-naming kind matched "cheap way out" and "why would you" but
not the wording that actually followed a broken handoff. Two spellings join
it, and a neighbour fixture pins that a neutral "why did you choose X"
stays silent.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01KU2pPKob4MJ1NqjsfTNyYJ

Depends-On: #506


Note

Low Risk
Heuristic regex tuning in reflect audit tooling with unit tests; may increase false positives on casual “bogus” or “didn't work” wording but no production runtime or auth paths.

Overview
Extends the reflect token audit cheap-way-out frustration detector so broken-handoff phrasing counts toward intervention-must-automate, not only explicit shortcut complaints like “cheap way out” or “why would you.”

The cheap-way-out regex in token_audit.py now also matches bogus and didn't (even )?work|run, so a user calling out a bad script and a follow-up that it still failed can register two hits in the same intervention class. Tests add a Codex fixture for that “bogus script / still didn't run” path and tighten the negative case to a neutral “why did you choose …” question so product blame plus ordinary rationale questions stay intervention-must-automate: no.

Reviewed by Cursor Bugbot for commit b680e2b. Bugbot is set up for automated code reviews on this repo. Configure here.

The shortcut-naming kind matched "cheap way out" and "why would you" but
not the wording that actually followed a broken handoff. Two spellings join
it, and a neighbour fixture pins that a neutral "why did you choose X"
stays silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KU2pPKob4MJ1NqjsfTNyYJ
Change-Id: I065043b69ac2933654b24236fc372ab3cb9cb8cd
@cursor

cursor Bot commented Sep 12, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_b64a634f-5c2c-4ac2-a7d0-c69da4cc77c7)

@EdbertChan

Copy link
Copy Markdown
Owner Author

This pull request is part of a Mergify stack:

# Pull Request Link
1 hook: a script handed to the user is a claim that it runs #506
2 audit(token_audit): count "bogus" and "it didn't run" as the same class #507 👈
3 hook: refuse a quoted command string passed through a login shell #508
4 gate: a branch taken per operating system needs a test that injects one #509
5 audit(token_audit): one submission carrying two slash commands is one message #510
6 audit: use the transcript's own markers for what the human actually sent #511

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant