Skip to content

fix(skill): the report synthesis reads its baseline from the frontmatter, and three query scripts match their surface - #662

Merged
using-system merged 7 commits into
mainfrom
fix/report-synthesis-and-query-scripts
Sep 26, 2026
Merged

using-system merged 7 commits into
mainfrom
fix/report-synthesis-and-query-scripts

Conversation

@using-system

Copy link
Copy Markdown
Owner

Summary

  • The report synthesis no longer names a baseline the run did not use. odd_report.py new writes a baseline: frontmatter field — the mission's named report (--baseline), else the recall's first match (same shared matcher as odd_recall.py, same-service-set rule), else none — and persist / show read it (or verifies on a replay). Older reports keep the prose fallback, which no longer takes a .md name from a line that opens with none / no previous.
  • A re-measure's headline counts its baseline rulings, and the body contract names section 7's | Check | Before | After | Verdict | table.
  • azure-monitor-traces.py exemplars ranks per operation and returns a p50 exemplar per operation (--slow N is now per operation, default 1).
  • grafana-metrics.py labels retries once without the series an Adaptive Metrics rule aggregates, with a note.
  • grafana-traces.py ops falls back to increase() of the span-metrics counter after a reset instead of withholding the calls.
  • observe-run's zsh trap line gives one copy-pasteable sed -n 'A,Bp;C,Dp' form.

Choices amended against the issue text are recorded on the issue: #659 (comment)

Tests

  • New tests in test_odd_report.py (baseline field from the recall / none / named / refused cases / same service set only, older-report prose guard, re-measure headline, Verdict column read by synthesis), rewritten exemplars test on re-captured masked fixtures, new Adaptive Metrics and reset-fallback tests for the grafana scripts.
  • pytest tests/skills tests/hooks 1387 passed after the rebase on fix(agent): an environment hard stop persists nothing, and recall follows the status lineage #661; ruff 0.16.4 check/format clean; check_stack_reference.py green; apm install --target claude + apm audit green.
  • show, synthesis and check over all 18 stored reports are byte-identical to main's; get-status --full identical.
  • Live: Azure Monitor (exemplars p50/slow per operation) and Grafana Cloud (labels under Adaptive Metrics, ops after a real span-metrics reset) verified by the implementer and re-run by the reviewer.

Review

A separate reviewer sub-agent checked the whole branch under the bound-review rule:

  • Round 1 found one blocking defect: a baseline the mission names was replaced by the recall's first line in the new field (a regression against main). Fixed with new --baseline (0 commands on the normal path, one contract line) and a test.
  • Round 2: no finding. After the rebase on fix(agent): an environment hard stop persists nothing, and recall follows the status lineage #661 (the recall's equal-service-set rule carried into the shared matcher, fake-gcx env lists merged), round 3: no finding, ready to merge.

Harness measurement

test-plugin-harnessing, copilot / openai/gpt-5.6-luna / medium, /odd-observe drive mission on the local llms-benchmark stack, ABBA, both lab branches rebuilt right before the chain; every changed file in the deployed path.

Side (mean of 2) Turns Commands Tokens in Tokens out Wall Generation Preflight Observation
main 41.5 41.5 3.06 M 17.0 k 421 s 202 s 62 s 232 s
branch 31.5 39.5 2.30 M 15.4 k 320 s 139.5 s 32 s 158 s

Main's spread today over ten samples: turns 35–45, commands 29–54, tokens in 2.57–3.55 M, wall 310–431 s. No branch figure is above main's spread; several are below it (turns on both samples) — not claimed as a gain on two samples, this is a fix owing no degradation. 1 premium request per sample. Every sample persisted a report on the first try (branch: baseline: none and "no previous report" headline, consistent); no sample had a stored baseline, so the "vs baseline" path is covered by the unit tests only. An earlier chain was discarded (one main sample exited 1 at copilot shutdown, its replacement recalled a leftover report).

Closes #659

🤖 Generated with Claude Code

using-system and others added 7 commits September 26, 2026 21:34
…t a replay's rulings

`new` records the recall's first line as a `baseline` field outside a
replay (`none` without one), and `show`/`persist` read it instead of the
first report filename on section 1's prose line; a report predating the
field falls back to that line, where a value opening with none yields no
baseline. A re-measure with no check table counts its baseline rulings.
The body contract names section 7's Verdict column on a replay.

Closes #659

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… one

`azure-monitor-traces.py exemplars` took the slowest requests over the
union of the operations, so one slow operation filled every slot, and
had no p50 pick. It now partitions the slowest by operation (`--slow`
per operation, default 1) and adds, per operation, the request nearest
its p50. Fixtures re-captured live, masked.

Closes #659

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gation

`grafana-metrics.py labels --label` exited 1 on Grafana Cloud as soon as
one series under the selector was aggregated by an Adaptive Metrics
rule. It now retries once without those series (`__aggregation__=""`)
and says so in its output.

Closes #659

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`grafana-traces.py ops` withheld every operation's calls when the span
metrics counter reset inside the window. It now reads the counter's
increase() over the window for the reset operations, as `histogram` and
`counter` already do, and marks the row.

Closes #659

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…anges

Every observe-run agent of the field campaign tripped on `echo ====` as a
separator between reads. The warning is replaced by the form to use:
`sed -n 'A,Bp;C,Dp' <file>`, no separator line.

Closes #659

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`new` took the `baseline` field from the recall's first line even when
the mission named another report, which the agent then diffed against,
so `show` and `persist` named the wrong baseline. `new --baseline <file>`
records the named report as given (refused on a replay, whose baseline
is `--verifies`, and when nothing stored carries that name); the recall
fills the field only without it.

Closes #659

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ce-set rule

Closes #659

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@using-system
using-system merged commit 868d752 into main Sep 26, 2026
12 checks passed
@using-system
using-system deleted the fix/report-synthesis-and-query-scripts branch September 26, 2026 20:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(skill): the report synthesis names a rejected baseline and counts no replay rulings, and three query scripts miss their documented surface

1 participant