Operational reference for common tasks. Each section is a self-contained
procedure. See CLAUDE.md for authoring rules and SPEC.md for the
composition model.
- Identify the parent template(s) — check
manifest.yamlfor the closest existing stack and trace itsdepends_onchain - Create
stack/<prefix>-<name>.md:- First line:
# Stack — <Full Name> - Second line:
[DEPENDS ON: <parent1>, <parent2>, ...] - One section per concern, each tagged
[ID: <name>] - Use
[EXTEND: <id>]to add rules on top of a parent section - Use
[OVERRIDE: <id>]to replace a parent section entirely - Follow the canonical section structure (ADR-017): MUST sections
are Stack, Commands, Project structure (pure libraries exempt
from the last); name the language section
<Language> conventions. SYS-06 gates the MUST tier on the resolved chain, so a derived stack may inherit a MUST section from its parent
- First line:
- Review for terminology carry-over from the source template:
- All framework and runtime names match the target stack
- CLI commands reference the correct package manager and tools
- Server/runtime terms are accurate (e.g. WSGI vs ASGI vs process manager)
- Language-specific vocabulary is correct (e.g. "packages" in Go, "crates" in Rust, "modules" in Python)
- Register in
manifest.yaml:- id: stack-<name> file: stack/<prefix>-<name>.md depends_on: - <parent-id>
- Add to the stack list in
SPEC.md(alphabetical within category) - Add a row to the stacks table in
README.md - Validate: attach
INTERVIEW.md+ new stack to an agent and review output
- Create the file in the correct layer directory:
base/<name>.md— cross-cutting, applies to all projectsbackend/<name>.md— backend services onlyfrontend/<name>.md— frontend/UI projects only
- Tag the file root with
[ID: <layer>-<name>] - Tag every section with a unique
[ID: <layer>-<name>-<section>] - Register in
manifest.yamlunder the correct layer key:- id: <layer>-<name> file: <layer>/<name>.md
- Update
docs/SPEC.md— add to the directory listing for the relevant layer - Add
backend/<name>.md(or frontend/) references in dependent stack[DEPENDS ON: ...]headers as appropriate
A lesson sourced from one downstream project (a bug write-up, a session retro) carries that project's domain in its framing. Vet it for genericity before it becomes template content:
-
Extract the generic kernel — the rule that holds regardless of stack or domain. Strip project-specific nouns (entity, tool, and metric names)
-
Pick the target by measured chain reach, not by the home the issue suggests — and know which kind of template you are measuring:
# how many of the 37 roots resolve a candidate file? for s in $(py tools/resolve.py --roots); do py tools/resolve.py "$s" | grep -q 'core/review.md' && echo "$s" done | wc -l
--rootsand not--list: a project resolves its stack, then its extras, then its platform, each as its own root, so the stacks are only part of what carries a file. Measuring over--listcounts the stack chains and silently omits every opt-in root.A chain template carries rules the generated project follows. Reach is the criterion for it, and a rule in one that no root resolves reaches no generated context file.
An opt-in template is a root of its own. Its reach is 1 — itself — and that is the point rather than a low score: what it may rely on is the core tier plus its own
depends_ontree, because the stack it is paired with is unknown when it is authored.platform/templates are the clearest case, since a project picks one regardless of stack.A pipeline template carries rules about generating or consuming a context file, and
templates/INTERVIEW.mdreads it directly rather than any root resolving it into a project. Reach says nothing about whether the rule lands:base/core/agents.mdis carried by a single root — its own — and shapes every file the pipeline generates. Recognise one by asking what reads it: a stack, an opt-in root, or INTERVIEW.md.A filed issue's suggested home is a hypothesis — it is usually written from the downstream project, without running this. Where a rule needs universal reach but applies only sometimes, put it in a high-reach template behind an
(if applicable)heading rather than in a low-reach one. Where a rule is one case of a contrast whose other cases live in a low-reach file, it stays with them: a case that cannot be read without the ones it is defined against is not made more useful by moving it somewhere better resolved -
If the kernel is genuinely cross-cutting, add it to the relevant
base/or layer template per "Add a new base or layer template" -
If a fragment only applies to one stack or domain, fold it as an example inside a generic rule — do not add it as standalone base content. Standalone one-stack content in a core-tier template loads into every chain and dilutes attention
-
Prefer extending an existing related rule over adding a parallel section — duplicate rules restated across templates are the same attention-dilution failure
-
Validate per "Validate a template change"
git mv <old-path> <new-path>- Update every
[DEPENDS ON: ...]header that references the old path - Update
manifest.yaml— change thefile:field for the entry - Update
SPEC.md,README.md,INTERVIEW.md— search for the old filename and replace - Update any
examples/files that reference the old path - Verify with
git statusthat no old references remain
- Change
[ID: <old>]to[ID: <new>]in the source template - Search all template files for
[EXTEND: <old>]and[OVERRIDE: <old>]— update every occurrence - Update
manifest.yamlif the ID is referenced independs_onlists - Update
manifest.yamlif the concept maps to a template entry
Use this workflow only after applying the decision threshold in
templates/base/core/docs.md: a record is owed when the decision changes
what a consuming project observes without reading this repository — the
composition model. Prose style, tool choice, release procedure, naming and
file layout need no ADR, however consequential they are. One coherent
architectural decision may cover related work across several issues or PRs.
ADRs live in docs/decisions/; ADR-010 records the frontmatter schema.
- Copy the template:
Slug is kebab-case, sentence-meaningful (e.g.
NNN=$(printf "%03d" $(($(ls docs/decisions/[0-9]*.md | wc -l) + 1))) cp docs/decisions/TEMPLATE.md docs/decisions/${NNN}-<slug>.md
011-provenance-principle). - Fill in the frontmatter —
id(quoted string matchingNNN),status: Proposed(orAcceptedif the decision is already ratified),date(today,YYYY-MM-DD),categoryfrom the closed set in ADR-010, andsupersedes/superseded_bylists. - Replace
# ADR-NNN: Title in sentence casewith the real title (colon form, sentence case). - Write the four sections in order — Context (why this decision is needed), Decision (what was decided, with RFC 2119 keywords), Alternatives considered (what was rejected and why), Consequences (downstream effects). Non-trivial decisions SHOULD include an inline ASCII diagram in the Decision section.
- If this ADR supersedes an existing one: update the old ADR's
frontmatter in the same PR — set its
status: Superseded, refreshdate, and add this ADR's id to itssuperseded_bylist. Preserve historical claims; format-only changes are also allowed under the decision-log rules intemplates/base/core/docs.md. - Run
py tests/run_smoke.py ADR-01to validate the frontmatter schema before opening the PR. - If the ADR removes a concept, sweep for prose that still
assumes it — in
templates/and in open issue bodies:A queued issue proposing template text that reintroduces the removed concept passes review on its own terms and undoes the ADR when implemented. Fix the issue body, do not rely on catching it at implementation time.grep -rn "<concept>" templates/ docs/ gh issue list --state open --limit 100 --json number,title,body \ --jq '.[] | select(.body | test("<concept>")) | "#\(.number) \(.title)"'
- Open the PR. Preserve the merged decision as history. Routine later refinements update current docs and the PR; only a material architectural replacement needs a superseding ADR.
A newer tag alone creates no work. Review a template update for a named project
need or material risk, applying "Adopting shared rules" in
templates/base/core/docs.md. Keep existing adopted conventions until the
project deliberately changes them; declining a candidate needs no ADR or ticket.
For consumers adopting the selective-adoption and ADR-threshold update:
- Inline the adoption boundary in the root context file before applying newly read template rules. Update copied inline rules as well as the submodule pin.
- Replace old directory-move, paragraph-length, and per-issue ADR triggers in the local context, wrap-up checklist, and authoring instructions, along with any threshold phrased as a judgement about how consequential a decision is. The test is what a user of the thing observes without reading the repository's internals.
- Keep existing ADRs as history. Put routine refinements in current docs and the PR; do not create a consolidation project or a decline register.
- Reconcile the reference list if a chosen update changes dependencies, and run the project's relevant existing checks. Select a released tag when the policy ships; until then the current pin remains usable.
- Open your agent (Claude Code recommended)
- Attach
templates/INTERVIEW.md - The agent explores what you want to build, asks a few clarifying questions, proposes a stack, and generates the file once you confirm
- Place the generated file at the project root
- Open your agent
- Attach the relevant stack template (e.g.
templates/stack/python-flask.md) - Provide your answers inline:
Generate a CLAUDE.md. Name: X, owner: Y, repo: Z, database: PostgreSQL, auth: JWT. - Place the generated file at the project root
If the agent cannot run scripts, use the pre-resolved files in
generated/. Each file contains the full template chain for one stack:
- Attach
generated/stack-flask.md(or the relevant stack) - Attach
templates/base/core/agents.md(output format) - Provide your answers and ask the agent to generate the file
| Interface | Recommended path | Why |
|---|---|---|
| Local agent (Claude Code, Codex CLI) | Interview or direct | Agent reads files from disk |
| Web portal (Claude.ai, ChatGPT) | Pre-resolved | Upload one file, no shell needed |
| REST API | Pre-resolved | Include generated/<stack>.md in prompt |
Prompt sizes and the minimum context window per stack category are in
README's "Model limitations" section, generated from the resolved
chains by py tools/sync.py. A copy here would be a second set of
figures with no generator behind it, and the copy is what ages.
- Output token limit < 16K: generate section by section
- Output token limit 32K+: full inline file fits in one pass
-
Smoke check: run the automated structural checks — no agent required:
py tests/run_smoke.py
This verifies all
[DEPENDS ON: ...]paths, unique IDs,[EXTEND: ...]/[OVERRIDE: ...]references, andmanifest.yamlconsistency in one pass.A rule added to a widely resolved file is read by every project on every chain carrying it, on every turn, and nothing refuses that size.
README.md's model-limits table reports it per stack category beside the smallest context window that still holds it, andpy tools/sync.py --checkfails when the table drifts from the tree. Read the table's diff before merging, and say in the pull request what the addition is worth.A category crossing to a larger window is a change to what a consumer needs to run the stack at all, and that is refused rather than reported: SYS-16 fails until the new tier is recorded in
tests/context-tiers.txtin the same change.Before adding a paragraph, try stating the rule inside the one that already carries the narrow version of it. A widened rule folded in place is often shorter than the text it replaces: the release-proposal record went from +472 characters as its own paragraph to 98 characters shorter than the narrow rule it replaced. Measured 2026-09-03 on
base/core/git.md. -
Agent check: attach
INTERVIEW.md+ the changed template to an agent and review the output for coherence; or run the relevant E2E test if one exists:py tests/run_e2e.py STK-01 # example — replace with the relevant IDReports are written to
tests/reports/after every run that has a result — an e2e dry run writes none — named for the time and the tree —<timestamp>-<runner>-<short-hash>.md, with-dirtywhere the working tree was not that commit, and no suffix where there is no repository to read.
A rule stated in two active sections of the same resolved chain dilutes
the agent's attention (see docs/design/template-content-quality.md). The
audit scans every root a project can pick -- the stacks and the
orthogonal templates alike -- and reports duplicates, excluding those the
override model legitimately produces (one section [OVERRIDE]s the
other):
py tools/audit_redundancy.py # exact in-chain duplicates
py tools/audit_redundancy.py --near # also paraphrase near-duplicates
py tools/audit_redundancy.py --check # exit 1 if any exact duplicate--check gates the exact tier only and prints the near count beside
its verdict, so a passing run states what it did not gate (ADR-038).
A near pair cannot fail CI; the count is what makes it visible.
Run it before opening a template PR and when consolidating rules. A finding is not automatically a defect — a duplicate may be intentional (e.g. a rule a base template needs when used standalone). Judge each against the single-source principle, then either trim the restatement (prefer the parent/owner template) or accept it with a reason.
--check runs in CI as a ratchet: it fails on any exact duplicate not
in the tool's BASELINE allowlist, so new redundancy is blocked while
the known duplicates clear through their owning issues (currently the
frontend pair, #624). When you resolve a baselined duplicate, remove
its BASELINE entry in tools/audit_redundancy.py — --check reports
any entry that no longer applies.
py tests/run_smoke.py # structural checks — seconds
py tests/run_conformance.py # the templates' own checks, run here
py tests/run_conformance.py --list # dispositions only, run nothing
py tests/run_e2e.py # canary test (python-lib)
py tests/run_e2e.py --all # all agent tests
py tests/run_e2e.py STK-01 FMT-01 # specific tests only
py tests/run_e2e.py --dry-run # build prompts, call no model, write no reportThe efficacy benchmark is a fourth runner and deliberately not in that
list: it costs quota, takes hours, and answers a different question from
whether this repository is sound. tests/efficacy/README.md owns its
procedure, and docs/design/efficacy-benchmark.md owns the method.
Its judge has a control of its own, which builds damaged and improved copies of two pinned applications and scores every tree. Run it whenever the judge model, its effort or the rubric prompt changes, because each of those invalidates every reading the benchmark has taken:
py tests/efficacy/control/control.py --root C:/efficacy/control-<date>
py tests/efficacy/judge.py --root C:/efficacy/control-<date>A row that does not fall on the damaged tree cannot detect damage; one that
does not rise on the improved tree cannot detect improvement, so no arm can
win on it. tests/efficacy/control/README.md carries the reading rule and
what the control found when it was first run.
See tests/CODIFICATION.md for the ID scheme and tests/INDEX.md for the
full list of specs. Requires py -m pip install pyyaml for the manifest
check.
run_e2e.py is the one runner that calls a model and costs money, so it
runs on a cadence rather than on every change: the STK-15 canary runs live
at every release cut, recorded in the cut's pull request, and the full
suite runs live at each periodic review. It is also the only runner that
reads what the templates generate — smoke and conformance read the
templates themselves. Name the provider and the model in the shell, which
takes precedence over .env:
E2E_PROVIDER=anthropic ANTHROPIC_MODEL=claude-opus-5 py tests/run_e2e.pyA live report names its mode, provider and model. A dry run calls no model
and writes no report, so every e2e report in tests/reports/ is a run. A
case that fails is filed as a bug with the strings it missed, not fixed
inside the cut that ran it.
run_smoke.py prints what each check inspected under its verdict — the
files scanned, the chains resolved, the directives compared. Those lines are
not findings; read them when a count moves without the tree moving, which is
how a check that silently stopped reaching its corpus shows up. A count of
zero fails the check, and a check that passes while reporting nothing is
failed by the runner itself.
run_smoke.py checks that the templates COMPOSE. run_conformance.py
checks that this repository OBEYS them — it extracts every fenced check in
templates/ and either runs it here or records why it does not apply.
Adding a fenced check to a template without a disposition in
tests/conformance.py fails the run, which is what stops a new check from
arriving unexamined. Checks reaching GitHub need gh authenticated.
A check whose verdict is a judgement reports REVIEW rather than PASS,
and the summary counts it as awaiting a reading. It does not fail the run
— the exit status answers whether a verdict was reached and was negative —
so a run ending 0 failed 2 not applicable 6 awaiting a reading is not
a clean run until someone has read those six. The middle count is not
among them: a check whose moment is not in progress exits 3 and has
answered, which is why it is counted apart. That is different from a
SKIP, which a person decided in advance about this repository and which
runs nothing at all. Recording them as passes is what let four
over-long changelog entries sit inside a green report for two sessions.
A SKIP carries a reason, and the runner refuses one that does not —
the skips are the largest population in the registry, and each is a check
that does not run, justified by prose. Where the only obstacle is a
placeholder the template left for its consumer to fill in, the entry takes
a substitute map instead and the check runs. That distinction is not
cosmetic: the release-pipeline check was skipped for naming
<release-workflow>.yml while .github/workflows/release-gates.yml had
been shipping for weeks, and nothing re-read the reason.
A check only earns that status where nothing in its output can be decided.
The changelog bound declares a limit and counts against it, so it is a
scored verdict now, and each check that stays a judgement carries a
reason in tests/conformance.py naming what about it takes a person.
The runner refuses a judgement disposition that states none.
Which checks may reach that status is fixed in tests/reading-budget.txt,
one title per line above the reason it cannot reach a verdict. A reading
the file does not name fails the run, so the set grows only in a diff that
says why; a named check that reaches a verdict or reports it does not apply
passes freely, which is what keeps the release moment from failing the run
— the ordering check answers "does not apply" while HEAD is the tag. An
empty budget is refused rather than obeyed. On a hosted runner the summary
is written to the run summary as well as the log.
Adding a judgement disposition therefore means two edits in the same
change: the reason beside the disposition, and the title in the budget.
Renaming a check means moving its budget line too — the runner reports the
stale entry as an unbudgeted reading rather than ignoring it.
A check that stopped reporting a finding and a check that can no longer report one produce the same clean result. Before believing a narrowed or newly-forked check, drive it through a fixture that MUST be reported and one that must not.
Where the fixture can live outside templates/, write it, run the single
check by a string from its body, and delete it:
py tests/run_conformance.py "<string from the check body>"Where the fixture must live UNDER templates/ — anything the check scans
by walking that tree — the runner refuses to execute: an unregistered
fenced block in the fixture is a missing disposition, which it reports
before running anything. Call its extraction directly instead, which skips
the registry reconciliation:
py - <<'EOF'
import sys, tempfile
sys.path.insert(0, "tests")
from run_conformance import iter_blocks, run_block
body = next(b[3] for b in iter_blocks()
if b[0] == "base/core/docs.md" and any("<find>" in l for l in b[3]))
code, out = run_block(body, "bash", tempfile.mkdtemp())
print("exit %d" % code)
for line in out:
print(" " + line)
EOFPass condition: the fixture that must be reported is reported, the one
that must not is not, and the check's own corpus counts move between the
two runs. testing-control-corpus-moves in
templates/base/core/testing.md is the rule, and says why the third is
what makes the first two readable.
For a check that takes an argument — a milestone, a module, a field —
replace the placeholder line in body before calling run_block. The
placeholder is not always what the issue or the prose calls it:
MILESTONE = None rather than an empty string cost a first attempt here.
Mutate the check's INPUT, never the string its disposition uses to find
it. The two are the same text more often than it looks: the CRLF check is
located by the extension list it filters on, so emptying its corpus by
editing that list stops the registry from finding the block at all. The
run then reports drift between the registry and the templates — a
different defect, in a different file, that says nothing about the check.
Empty the corpus from the other end instead, with a pathspec or a glob
the find string does not overlap.
Changing what a check exempts is a documentation edit as well as a
code edit. Sweep the documents for the check's name before merging and
re-read every hit against the new exemption set: the check's own tests
exercise the exemptions rather than the sentence describing them, so a
document's account of them decays with nothing reporting it.
quality-exemption-doc-duty in templates/base/core/quality.md is the
rule, and ships the sweep.
Adding a constraint owes a sweep of its own. Run the new rule or check
against the corpus that already exists before merging, and record in the
pull request where it looked, what it found and what it left out. Where the
corpus does not comply, say whether the instances were fixed or frozen, or,
for a corpus that cannot be edited, the boundary the rule binds forward
from. quality-constraint-corpus-sweep in templates/base/core/quality.md
is the rule.
After editing any template or manifest.yaml, regenerate the cached
files in generated/:
py tools/resolve.py --generate # regenerate all cached files
py tools/resolve.py --check # verify they are up to date
py tools/sync.py # also regenerates via --checkThe generated/ files are committed to the repo so agents without
shell access can use them directly.
Per ADR-016, the files in examples/*/CLAUDE.md are agent-generated
outputs, not hand-maintained. When a template in an example's chain
changes the output shape, regenerate the example — do not patch it by
hand.
- Attach three inputs to a local agent:
- the existing
examples/<name>/CLAUDE.md— for the project brief (name, owner, repo, deployment, stack choices, architecture) to preserve generated/<stack>.md— the resolved chain (the rules to apply); read any addon templates the example's stack source names but the chain does not includetemplates/base/core/agents.md— the output models (inline, reference, or hybrid — keep the example's existing model)
- the existing
- Ask the agent to regenerate
examples/<name>/CLAUDE.mdin the six-section structure, keeping the concrete project identity and recording the generation inputs in a note near the top. - Verify:
py tests/run_smoke.py(SYS-05 gates the §6.3 audit) and review the diff for coherence. Byte-for-byte reproduction is not expected — generation is non-deterministic.
Per templates/base/workflow/360.md, a 360 assesses the whole project
from independent stakeholder perspectives as parallel, context-isolated
subagents (the headless adaptation re-projects Quality into engineering
dimensions). Run one on demand — before a launch, after a milestone or a
major feature, or when a stakeholder needs the whole picture. A release
does not trigger one by itself: the pre-release check asks only whether
the newest record still falls inside the 90-day review interval, so a
release whose record is current owes neither the audit nor a record
declining it, whatever version it moves.
- Store each audit as a dated report at
docs/audits/YYYY-MM-DD-360.md(per §360-tracking) — this is the only audit location; never use a single-filedocs/360-audit.mdhistory. All audit history lives in the folder. - Each report carries a scores table, the issues created, the current bottleneck, and per-dimension findings tables with a grade rationale.
- File a labelled issue for every actionable finding (CLAUDE.md §2.2) and reference it from the report.
- Run the e2e suite live with
py tests/run_e2e.py --alland record its summary, provider and model in the report — the review is its cadence.
- Ensure you are on a feature branch — never commit to
maindirectly - Run the validation steps above for every changed template
- Update all affected documents (
SPEC.md,README.md,manifest.yaml) before committing - Commit with a conventional message:
feat(stack): add go-echo template docs(spec): update backend layer listing - Push and open a PR — one concern per PR
- Merge with the default squash subject —
gh pr merge <number> --squash, with no--subject. GitHub composes that subject from the title and appends the pull request number; supplying one replaces both. The milestone-coverage gate reads those numbers back out of the log, so a hand-typed subject removes the reference it resolves, and the gate reports on the commits it can still read - After merge: delete the branch and pull
main
Where a cut carries more than one pull request, every branch after the
first goes BEHIND. Each merge regenerates generated/, so the second
branch carries pre-resolved chains built from a tree that lacks the
first one's change. Git reports the branch MERGEABLE — the edits fall
in different regions of the same files — so nothing warns that the
chains are stale, and smoke does not read generated/ either. CI is the
first signal.
Recover it with gh pr update-branch, never a rebase and force-push.
Then re-run both staleness gates on the updated branch, because
mergeable is not the same claim as current:
gh pr update-branch <number>
git pull
py tools/sync.py --checkPass condition: the gate prints All files in sync after the
update, having listed every generated chain it compared. sync.py --check is the whole gate — it covers generated/ as well as the three
generated documents, so there is no second command to run. Run it before
the update instead and it reports the previous state, which reads as a
pass.
tools/resolve.py --check is a real gate and a narrower one: it compares
generated/ alone, prints the stale chains and exits 1. CI runs it as a
required step and every file in generated/ opens with a banner naming it.
The gate above subsumes it, so this step still runs one command — and an
unknown flag on that tool exits 1, so a typo there cannot read as a
clean gate.
Run this before scoping a cut. The procedure and the reasons behind its
order are in templates/base/workflow/issues.md, which this repository
consumes; what follows is only what is particular here.
List the backlog — an unmilestoned issue is triaged, not untriaged, so no view reports it:
gh issue list --state open --limit 200 --json number,title,milestone --jq '.[] | select(.milestone == null) | [.number, .title] | @tsv'- Verify a claim with the tool the claim is about:
py tests/run_smoke.pyfor a structural claim,py tools/resolve.pyfor a reach or chain claim,py tools/audit_redundancy.pyfor a duplication claim. Where the claim counts fenced blocks, usetests/conformance.py's own extractor rather than a grep overtemplates/— a grep counts occurrences no runner reads - Measure reach in roots, with
py tools/resolve.py --roots, before siting a rule that could go in either of two templates - Title the milestone
vA.B.C — <theme>. The annotated tag message quotes the theme, so a bare version leaves the tag with nothing to quote - The merged-work sweep is step 9 of the release sequence in
base/core/git.md;py tests/run_conformance.pyruns it here. Run it at the groom too, where a hit is cheapest to act on
Where a milestone already exists but its theme describes work that has not
happened, renumber it rather than dissolving it. A scoped milestone carries
an ordering and a rationale that took a groom to produce, and emptying it to
reuse the version number discards both. Move it to the next version, create
the version being cut under a theme matching what actually shipped, and
assign the closed issues to that. Met on 2026-09-03: v2.73 was themed for
six unstarted issues while five unrelated ones had already merged, and one
issue had been put in it solely to satisfy the release gate.
Run this before proposing a change to a document base/core/docs.md requires
— docs/PLAYBOOK.md, docs/ONBOARDING.md, README.md, CHANGELOG.md,
SECURITY.md, CONTRIBUTING.md. tools/sync.py reaches this repository and
nothing else, so the cost of renaming or restructuring one of them sits
downstream, in repositories no gate here can see.
-
Enumerate every account that generates from these templates. Both are the same owner, so no third party is affected, but each repository still needs its own pull request under its own branch protection:
{ gh repo list braboj --limit 100 --json nameWithOwner --jq '.[].nameWithOwner' gh repo list Imbra-Ltd --limit 100 --json nameWithOwner --jq '.[].nameWithOwner'; } | sort -
Fetch each default branch's tree once and test exact paths against it. Fetching the tree is one call per repository and answers every path question; probing paths one at a time is both slower and wrong, for the reason in step 3:
br=$(gh api "repos/$full" --jq '.default_branch') gh api "repos/$full/git/trees/$br?recursive=1" --jq '.tree[].path' \ | grep -c '^docs/PLAYBOOK\.md$'
-
Carry a control path through the same probe — one that MUST be absent everywhere, such as
docs/THIS-MUST-NOT-EXIST.md. A probe reporting a hit for every repository has usually failed rather than succeeded. The earlier form of this survey did exactly that:gh api "repos/$full/contents/docs/PLAYBOOK.md" --jq '.name' 2>/dev/null
It prints the 404 body to stdout, where the redirect cannot suppress it, and
--jq '.name'reads that body as a non-empty string. Sixteen of sixteen repositories reported a hit, including ones with nodocs/directory at all. Measured 2026-09-08 with the tree form: 31 repositories probed, 14 carrydocs/PLAYBOOK.md, 13 carrydocs/ONBOARDING.md, 0 were unreadable, and the control path is found in 0 -
Count mentions over tracked files, and count occurrences rather than matching lines. Reconcile the per-area sum against the total, so an area the enumeration missed shows up as a gap rather than as a smaller number nobody questions:
git ls-files -z | xargs -0 grep -io 'PLAYBOOK\.md' | wc -l
Measured here on 2026-09-08, the four forms disagree: occurrences over tracked files 137, matching lines over tracked files 133, occurrences over the working tree 491, matching lines over the working tree 215. The working-tree walk reads
.git/and inflates by three and a half times;grep -ccounts a line carrying two mentions once. Only the first figure is the number of places a rename has to edit -
Record the result in the issue proposing the change, with the date it was taken, before the decision is made. A survey quoted later without its date asserts about today's tree what was true on the day it ran
This repo has no version manifest (plain Markdown), so it follows the
no-build release variant from ADR-006. What that variant removes is the
version-bump commit and the chore: release branch carrying it — there is
no manifest to bump, so nothing needs one. It does not remove every pull
request: the changelog cut at step 4 is its own, and it is the release
commit the tag names.
Run base/core/git.md's pre-release checks first — that sequence is the
source, and the nine steps below are this repository's release procedure
proper, not a restatement of it. Two of those checks bind here and neither
has a step below: the periodic-review-scope check, which reads the age of
the newest record in docs/audits/ and takes no release version, and the
pipeline-history check.
Read each gate's output as well as its exit status. Every gate below now
carries its verdict in the status, so a finding fails the run rather than
printing and exiting zero. That held for five of them only from v2.82, and
for thirteen more only from v2.88, and this procedure was written against
the older behaviour. What a status still cannot carry is a reading: three
of the checks reach no automatic verdict and address their output to
whoever the failure summons, so a run reporting 0 failed has not been
read until those are. Measured during the v2.77.0 cut, when the status
carried nothing: milestone-coverage printed a subject names closed issue 1456, which carries no milestone and exited 0. That gate exits 1 on the
same finding now.
Each gate below reads its parameter from the environment as well as from
the constant, so RELEASE=v2.83.0 py tests/run_conformance.py runs the
whole set in one pass without editing a template. The variable is not
always spelled like the constant it fills: the milestone-coverage gate
reads RELEASE_MILESTONE, and setting MILESTONE instead leaves it
reporting that the release is not scoped to a milestone — the one answer
that looks like a clean pass. Read the constant's own line before
exporting. The Release gates
workflow does exactly that on every tag push, resolving the milestone
titled with the tag's full version, or failing that its v<major>.<minor>
line, and can be dispatched with the version as an input to run them at
the release commit.
That automatic run is the backstop, not the mechanism. It happens after the tag, which cannot be taken back cleanly, so it reports what the steps below exist to prevent. Run them here anyway.
- Confirm the milestone's issues are all closed and
mainis green (py tests/run_smoke.pypasses) and up to date (git pull) - Confirm the inverse — every issue closed since the previous tag
carries the milestone being released. Step 1 reads the milestone
and cannot see work merged without one. Run the milestone-coverage
check in
templates/base/core/git.md, exportingRELEASE_MILESTONEas the milestone being released. Left unset it reports that the check does not apply — correct for a routine release on a project that scopes some cuts and not others, and a silent pass here, where every cut is milestoned - Confirm nothing else is ready to merge, or decide which side of the
tag it lands on. Run the release-ordering check in
templates/base/core/git.md— anything merged between the release commit and the tag ships inside the release with no note naming it, and both pull requests stay green in either order - Cut
CHANGELOG.md— reconcile theUnreleasedsection against the commits the release carries with the changelog-completeness check intemplates/base/core/git.md, setting itsRELEASEto the version being cut. Left empty it reports that no release is in preparation and does not apply, which is correct on any ordinary day and a silent pass here. Choose the version from the section before naming it: a patch when every entry sits under### Fixedor### Security, a minor otherwise, perbase/core/git.mdVersioning — the release-documentation check fails a version that disagrees. A patch's milestone is retitledvA.B.C — <theme>before the gates run, or milestone-coverage reads the minor line's milestone instead. Then rename that section to## [A.B.C] - YYYY-MM-DDand open an emptyUnreleasedabove it. This repo has no version manifest to carry the cut, so it is its own pull request, and it MUST merge before the tag below — a tag placed first names a tree whose changelog does not mention the release. That pull request's body is this repository's release proposal: it names every step ofbase/core/git.md's pre-release sequence and the result it produced, including the checks that carry no step number here. It also records the STK-15 canary run live against the release commit: the report's tree, provider, model and verdict lines. The check's two counts are not required to match: a commit touching no template carries no entry, so a journal or tooling change is expected to appear in the carried list with nothing answering it - Tag and push, per
base/core/git.md's no-build sequence — the tag is annotated and names a commit, and that commit is step 4's changelog cut, notmain. The theme in the tag message is the milestone's, which is where this repository keeps it - Publish the release from that tag, per the same sequence. Locally
tag-guard.ymlis what fails a lightweight tag, so the guard the sequence prescribes is enforced here rather than merely advised - Close the release's milestone once the release is published — titled
vA.B — <theme>for a minor andvA.B.C — <theme>for a patch - Record that the session owes a
docs/dev-journal.mdentry — do not write it here. The entry is written at the end-of-session audit item that owns it, which is the last item for a reason: it is the only one whose output is a record of the others, and items above it file issues and open pull requests. Written at this point instead, the entry names none of them, andbase-docsfixes its account once written, so the repair is a second entry for one session. When the entry is written it is still separate — its owndocs(journal): ...PR with no milestone, not part of the release. A release cut without a wrap-up still owes the entry; publishing the release does not discharge it - Verify the release exists before closing the session. This is last
rather than part of step 6 on purpose: a check inside a step is
skipped whenever the step is:
Pass condition: prints the tag, the bare-version title, and
gh release view vA.B.C --json tagName,name,isDraft --jq '.tagName, .name, .isDraft'false.release not foundmeans step 6 did not happen — and every other artifact of the release, the tag and the closed milestone and the journal entry, is present either way, so nothing else surfaces it
Which of those steps are actually enforced, audited per step as
quality-gates-procedure-steps requires:
| Step | Enforced by |
|---|---|
| 1 | py tests/run_smoke.py, plus the milestone's own issue list |
| 2 | the milestone-coverage check in base/core/git.md |
| 3 | the release-ordering check in base/core/git.md |
| 4 | the changelog-completeness check in base/core/git.md, and its release-documentation check, which fails a version whose bump disagrees with the section's headings. The canary record is enforced by nothing |
| 5 | tag-guard.yml, which fails a pushed lightweight v* tag |
| 6 | step 9, and nothing before it |
| 7 | nothing — an open milestone after a published release is silent |
| 8 | nothing — a missing entry surfaces at the next wrap-up, or never. The entry itself is gated by the audit item that writes it |
| 9 | nothing; it is the closing check, so it is the one to run by hand |
The pre-release checks above the table are enforced by base/core/git.md
rather than by a row here, which is why they carry no step number: a
number would claim they are part of this sequence and drift from the one
that owns them. What keeps them from being skipped is not a number but
the record at step 4 — a check nobody ran is a line the release proposal
does not carry. Four minor cuts shipped before that record was required,
each owing a periodic review; ADR-037 holds their disposition.
Step 6 is the one to protect first, because its omission cannot be
repaired afterwards: publishing the release later dates it after the
release that followed it, and a mis-ordered release list is worse than a
visibly absent entry. v2.57.0 is the instance — tagged, milestoned and
journalled, never published, and deliberately left absent for that
reason.
Projects with a version manifest (package.json, pyproject.toml,
etc.) instead follow the branch → bump → PR → merge → tag flow; see
ADR-006 and base/core/git.md.