Skip to content

feat(adapters): systematic-debugging scenario pack - #254

Open
WODE25500 wants to merge 2 commits into
microsoft:mainfrom
WODE25500:feat/superpowers-systematic-debugging
Open

feat(adapters): systematic-debugging scenario pack#254
WODE25500 wants to merge 2 commits into
microsoft:mainfrom
WODE25500:feat/superpowers-systematic-debugging

Conversation

@WODE25500

Copy link
Copy Markdown
Contributor

Summary

Adds a systematic-debugging scenario pack to the existing Superpowers evaluation adapter (skillopt_sleep/adapters/superpowers.py), alongside verification-before-completion. Refines issue #132 by extending the adapter to a second checkable skill.

What it does

The three scenarios judge mechanically-detectable process discipline, reusing the existing rule-based judge ops (no change to the evidence machinery):

  • reproduce-and-verify-before-done — observe a failing run, then re-run and verify after editing (guards against fix-without-reproduce / no-verify).
  • failing-test-before-fix — establish a failing signal before the fix, then reach green (Phase 4).
  • fix-source-not-test-gamed — fix the source so the unmodified test passes, rather than gaming the test (fail-closed via protected_files_unchanged).

Honest boundaries

  • The scenario pack deliberately does NOT judge whether the agent truly understood the root cause — that is beyond a rule judge (the upstream Superpowers project uses an LLM verifier for skill compliance). It is documented in the module docstring.
  • The change was built and validated offline (16 deterministic unit tests over the judge logic). It was NOT run against a live Claude/Codex harness (the submitting environment has no working authenticated Claude CLI). An opt-in real-harness smoke is documented in the docstring:
    python -m skillopt_sleep.adapters.superpowers --skill systematic-debugging
    A reviewer/CI with a working claude CLI can run it to get real empirical evidence.

Scope

  • 2 files, purely additive (~161 insertions, 0 deletions): skillopt_sleep/adapters/superpowers.py + tests/test_systematic_debugging_scenarios.py.
  • Does not touch the existing verification-before-completion scenarios or the evidence machinery.

Refs #132.

Add a systematic-debugging skill scenario pack to the Superpowers
adapters.SuperpowersEvaluator, alongside verification-before-completion.

Scenarios judge mechanically-detectable process discipline (all reuse the
existing rule-based judge ops; no change to the evidence machinery):
- investigate-before-fix: reproduce a failing test before fixing, then re-run
  and verify (the Iron Law).
- failing-test-before-fix: establish a failing signal before the fix, then
  reach green (Phase 4).
- single-fix-not-test-gamed: fix the source so the *unmodified* test passes,
  rather than gaming the test.

Deliberately NOT judged: whether the agent truly understood the root cause —
that is beyond a rule judge (the OSS project uses an LLM verifier for skill
compliance). Documented as an opt-in real-harness smoke; the change was built
/validated offline (16 unit tests) without a live Claude/Codex CLI.

Refs microsoft#132.
Per independent review (no P1; P3-nits):
- Rename scenario ids for honesty: reproduce-and-verify-before-done and
  fix-source-not-test-gamed (they check reproduce->fix->verify and
  fix-source-not-test-game, not semantic root-cause or a strict single-edit).
- Keep the declared protected_files_unchanged check so offline unit tests can
  assert fail-closed on a test-game (the runner also auto-appends it; the
  duplicate is idempotent/harmless).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant