Skip to content

docs(vocabulary): the collaboration protocol and its task-unit sub-protocol get names - #71

Merged
wefio merged 33 commits into
mainfrom
docs/vocabulary-task-unit-protocol
Sep 25, 2026
Merged

wefio merged 33 commits into
mainfrom
docs/vocabulary-task-unit-protocol

Conversation

@wefio

@wefio wefio commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

What this is

A naming decision, so the main line can be discussed without re-deriving what it means each time.

Two levels, one meaning each:

  • Protocol-governed collaboration is the umbrella: agents coordinate through a published
    protocol, while the mechanics of that coordination stay with the program rather than the model's
    discussion.
  • Task-Unit Protocol is the sub-protocol this line builds: a task is declared as a unit (inputs,
    dependencies, acceptance, capability, budget), and the protocol fixes how it is adopted into one
    run, claimed, delivered and independently judged, and which facts the runtime decides and owns.

The parts keep the names they already have: Task Board is the sub-protocol's surface, managed
entry
is what a governed entry is, run is the frozen object, adopt is the binding
transition, and the caller that decides to adopt is called the adopter.

Why these words

Everything that came to hand was already taken, and each collision is recorded as considered and
rejected rather than quietly reused: fusion already names one mechanism (several units sharing a
session), a scheduler is what the task-unit semantics decision deliberately does not enable,
governance is the board addressing and readability line, ledger already names three things
(assumptions, budget, disclosure), and executor names the bounded context adapters and the Pi SDK
execution path.

What it deliberately does not do

  • No identifier or file name changes: run, adopt, managed, dispatch, the ooo- prefix and
    every path stay as they are, because the term and the paths are cited by dated measurement records
    and frozen run archives.
  • docs/glossary.yaml is untouched: it owns the repository and process vocabulary, and these are
    product concepts, which live in the concept map.
  • It builds nothing. The sub-protocol's mechanism exists (the run surface, the managed-write fence,
    the shared dispatch loop); the product path still has no adopter, which is the next piece of work
    rather than a consequence of the name.

Verification

  • npm run docs:check: 278 files, 0 errors, 0 warnings (58 implemented decisions).
  • npm run glossary:check: 14 terms, 0 errors (unchanged).
  • Documentation route only: no code, no test and no schema changed, and no paid model run was made.

Follow-up to #70 (already merged). This branch was cut from main at a9f56759, so it carries one
commit.

Added in the same branch

  • The naming record now carries an index section: one line of meaning and a pointer per name,
    covering both levels, the board verbs and the channel rules, the run and its facts, the legal action
    set and the dispatch loop, plus the two words deliberately not used (scheduler, OoO) with the owner
    of their real meaning. It says of itself that it is an index rather than a second specification and
    that the owner wins on any disagreement; the definitions were previously only reachable by opening
    the design, the obligations ledger and several decision records at once.

  • The concept map's Task Board row pointed at an anchor that does not exist in memory-graphs.md
    (#21-task-board-outside-the-three-memory-graphs). docs:check validates that a link's file
    exists, not its anchor, so the link read as working; it now names the heading that is really there,
    in both languages.

  • The naming record also carries the conceptual history now, because the names only make sense in
    order: the board becoming a protocol (2026-08-13), governance and addressing (09-06), out-of-order
    execution borrowed from the CPU with its limits written down (09-09..09-11), the task unit's
    semantics (09-13), the paid arms measuring both borrowed halves (09-18..09-19), and the board
    absorbing the execution half (through 09-20). Commit-level lineage stays in implementation-lineage.md.

  • Corrected while writing it: ooo-execution-bootstrap.md still said the speculation decision was
    proposed and that the no-speculation rule held until it was implemented. It has been implemented
    since 2026-09-11; the sentence now states what is actually in force.

  • Adds a proposed record for the question the naming decision unblocked:
    docs/decisions/proposed/2026-09-20-the-program-answers-legality.md (+ zh). The primitives have names and
    the shared layer already computes the legal set, but nothing stated whose decision each step is. The record
    gives the program exactly one share (answer legality: the ordered legal set, cut to the declared slot budget,
    with a reason per unit - choosing nobody, adopting nothing, judging nothing, waking nobody), leaves the plan,
    the order, the adopter and the judge to the agents as board facts, makes constraints named by the protocol and
    enabled by the plan, and keeps the query free of state because a published handoff occupies a serial slot. It
    is proposed, not implemented: the readable answer and the promotion of repair-first out of shared planning
    policy are its plan steps 2 and 3.

  • Adds the umbrella's parts inventory as a design document:
    docs/design/protocol-governed-collaboration.md (+ zh), and points the naming record's history section at
    it. Fifteen parts with owner and state (entries/addressing, identity, ownership/time, termination integrity,
    attention, scope, truth vs coordination state, the task unit, legality, ordering, budget, cancellation,
    recovery, handoff, evidence), seven gaps with the distributed-systems concept that fits each (commit boundary,
    orphan reaping, admission control, declarative policy, monotonic reads, conflict detection, fenced mutex,
    in-doubt transaction), what is deliberately not needed (consensus, leader election, quorum, 2PC, exactly-once,
    distributed transactions), and three analogies that mislead (a lease is a suspicion, not a death certificate;
    the board is not a queue; agents are not replicas).

  • Adds a second proposed record, the frame itself:
    docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md (+ zh). The primitives have names and the
    program's job is proposed, but nothing said which fields the frame has, how values are encoded, where they
    live in the store, or how a peer protocol sits beside the Task-Unit Protocol. Researching the finished
    specifications changed the shape of the problem: every layer but one is already defined by somebody, so the
    record chooses and names instead of inventing. It borrows per layer (CloudEvents for the envelope, A2A for the
    task model and paging, NLIP/ECMA-430 for the payload discriminator and the progressive shape, MCP for
    LLM-facing results and errors, RFC 8941 for a one-line text type system, FIPA-ACL for the act and threading
    fields, RFC 6709 for the unknown-field choice, ECMA-434 for security) and states that the one layer nobody
    defines is ownership, lease, fencing and legality, because every agent protocol assumes a task belongs to one
    agent. Three decisions carry it: the frame is a state record rather than a message, so FIPA's performative
    is a rendering of existing columns; the header stays in columns because the claim is one atomic
    compare-and-set UPDATE and a header inside JSON would turn that into read-modify-write; and the payload is
    one opaque document whose invariant is that every payload can be NULL while T0/T1 still render and claim,
    wake, expiry and compact read still work. A payload field a protocol wants to query is indexed by a generated
    column plus an expression index instead of a new board column, which is what keeps the schema from growing
    with the number of protocols; the protocol is a versioned value and a peer protocol supplies declaration
    validation, legality, projection and acceptance through the existing DispatchBoard and RoundQueryPort
    seams. Data format is answered as three separate questions (existing JSON RPC wire; flat key/value for T0/T1
    with natural language for negotiation, since format restrictions measurably degrade reasoning; typed columns
    plus one JSON payload document), and TOON and CBOR are deliberately not adopted. Security is named as the
    empty layer, with two rules adopted as text from ECMA-434 and a threat list owed. The record ends with an
    explicit scope ceiling so the change cannot grow: one group of nullable columns, one boundary adapter, one
    rendering rule, one test - anything needing a new table, tool or channel is outside it.

  • The frame record answered its first review in a second commit. Four points, all accepted. (1) The field
    mapping was missing and it changes what the invariant means: the record now maps every existing field to
    its place in the frame, records that content stays the single source of the entry body
    (readTaskBoardPreviews projects taskBoardPreview(entry.content)), that the frozen declaration was
    never a body (it lives in task_run_manifest/task_run_tasks), that task_run_facts.payload is the
    precedent for the payload shape, and that a payload may not restate a field that already has a column -
    which is how the second task body would appear. (2) "The insertion points already exist" was false:
    RoundQueryPort carries two typed reads and the dispatch loop requires a patch, so the four-role seam
    is a plan step, not a fact. (3) Two storage arguments are withdrawn with the measurement that kills them

    • a conditional UPDATE claims over a JSON payload exactly once, and a generated column is a column, so
      peer protocols are not schema-free; what remains is clarity, indexing, CHECK constraints, additive
      protocol-owned schema, and no field stored twice. (4) Three borrowings were too wide: source_session_id
      is provenance rather than a credential, the refusal rule keeps MCP's protocol-error/execution-error
      split (which the dispatch loop already makes as refused versus failure), and the CloudEvents naming
      constraint is stated as borrowing with protocol_version's underscore recorded as a deviation.
  • Adds a third proposed record that names the division the other two each assumed:
    docs/decisions/proposed/2026-09-21-mechanism-not-policy.md (+ zh). The legality record says what the program
    does and refuses to decide; the frame record says what the core knows and never parses; they are one division
    seen from the decision side and the data side. The record states it as two contracts (mechanism: one claim with
    its lease and fence, an opaque declaration with a digest, an opaque artifact with a digest, a verdict from
    someone else, the lifecycle, and the legal set with a reason per unit; policy: inputs, artifact class,
    judgement, how the work runs, concurrency preconditions, the wording and enablement of legality rules, and who
    is chosen, adopts or judges), adds a decision procedure that settles arguments rather than serving as a slogan
    (a rule's check is mechanism, its name and enablement are policy), and adds a mechanically checkable rule: no
    policy word may appear in the mechanism layer's code, types or schema paths, from a maintained list, with the
    path carried because a word can be policy in one place and mechanism in another. It also states what this makes
    of work shapes - a work shape is not a core concept but the policy layer's name for the declaration-and-artifact
    pair the core carries opaquely - classifies the six sites where the patch assumption lives, and keeps three rows
    deliberately undecided so they are not settled by preference. The legality and frame records now point here.

The board entry has two faces, and the read face now has one home

The layer model in docs/design/mechanism-in-the-middle.md was retired and replaced in the same document by
a model that came from measurement rather than from analogy: a board entry is a medium with two faces - a
write face that a producer puts a body on and a read face that a reader is shown - and the compositor is
brought by a protocol
, not by the board. The Task-Unit Protocol is the one that brings ours, which is the
same statement as "the program only answers legality": no declared legality, no judgement. Of the seven entry
kinds only handoff and result are ever claimed, delivered and judged; a note, a goal and a question travel
from the write face straight to the read face.

Two facts the drawing has to survive, both measured: the roles are not files (src/core/store/base.ts is
the medium, the compositor and the reader's rule at once), and feedback is a replay, not a reversal (a
reader becomes the next writer and enters through the write face again; the middle is never traversed
backwards).

The rule that decides where a seam goes needs no taste, only a count: a boundary is built when it has a
second implementation, and only declared with one.

  • Read face - built. It had three truncation rules: the store's 200-character preview, a generic helper in
    the pi adapter, and a bare 140-character slice that did not collapse whitespace at all, so a multi-line body
    reached a broadcast as a multi-line broadcast. src/core/board-entry-view.ts owns the rule now; the store's
    compact read and the adapter's broadcast call it; the private copy and the bare slice are gone;
    tests/core/task-board.test.ts pins the rule.
  • Write face - declared. A body has no shape of its own yet, and the three fillers are three kinds of
    entry rather than one shape drawn three ways. The trigger for a type is a fourth filler.
  • Middle - declared. One protocol brings one middle. The trigger is a second protocol.

Check (c): the seam is prepaid, and the compiler cannot see the coupling

Making the loop shape-agnostic (DispatchTicket.patch?: PatchWork replaced by an opaque
declaration?: {shape, digest, frozen}, PlanWorker pointed at it, preparePatchWork deleted from the unit
dispatcher) is one file, fifteen insertions and eighteen deletions. Measured on a throwaway branch and
reverted, it gave two results, neither of them the one the prediction was aiming at: tsc --noEmit reported
zero errors, because the field is optional and structural typing still lets the board's ticket satisfy the
port; and the suite reported 25 failing tests, every one with the same cause - the claim admits no work.
A second shape today would not fail as a new shape; it would refuse every unit. The coupling's surface by
grep is three source files, the board (which must change although it never reads the field), seven test files
and the driver's three worker kinds. So the seam is deferred with a named trigger - the first non-patch unit
an adopter declares - rather than built against a single shape.

Docs and code land in one commit, because the read face's home is what the two-face model implements.

The port's third implementation: the store's own board

src/integration/ooo-runner.ts is DispatchBoard implemented over the product's record - the probe
board and the dispatch test's stub are the other two. It carries no legality rule: the plan is
compiled by compileTaskUnits, the legal set is dispatchTasks + selectableTasks, cancellations come
from the coordinator's own reader, and the verdict belongs to the caller's acceptance. A rule that turns
out to be needed there belongs in task-semantics.ts; needing one is how the module would show that the
abstraction leaked.

Two inputs stay the caller's, and both are existing divisions rather than new ones: a task's declared file
contents (the store freezes paths) and an acceptance whose identity is independent of the deliverer (the
store already refuses a deliverer that judges its own delivery - pinned by the second test).

One contract is easy to get wrong, so the projection writes it down: the shared acceptance rule binds a
verdict to the artifact value the run carries (it compares the verdict's digest against the artifact),
not to the store's own deliverable hash. Reporting the hash there makes every accepted unit read as
unaccepted, and the symptom is silent - the next unit is never released and nothing reports an error.

The port's "put a result" is the product's delivery on the entry the unit was claimed on, so a unit's
work is recorded once rather than once on that entry and again on a result entry nobody would read; and
the unit is resolved from the claim rather than from the body, because a body's shape belongs to its owner
while a claim is the board's own fact.

Verification: npm run check clean; test:product 1527/1527 pass (2 new - a two-unit plan driven through
the product's store with an independent judge, and the self-judging refusal); build clean; lint clean;
format:check clean; docs:check 288 files / 0 errors / 0 warnings; agent:context:check valid once the
route claims the new file; complexity:gate ok.

…otocol get names

Two levels with one meaning each. Protocol-governed collaboration is the umbrella: agents
coordinate through a published protocol and the mechanics stay with the program. The
Task-Unit Protocol is the sub-protocol this line builds - a task is declared as a unit and
the protocol fixes how it is adopted into one run, claimed, delivered and independently
judged, and which facts the runtime decides.

The parts keep the names they already have: Task Board is the sub-protocol's surface,
managed entry is what a governed entry is, run is the frozen object, adopt is the binding
transition, and the caller that decides to adopt is called the adopter. The words that came
to hand were already taken - fusion (units sharing one session), scheduler (which the
task-unit semantics decision deliberately does not enable), governance (the board
addressing line), ledger (three of them) and executor (the bounded context adapters) - so
each is recorded as considered and rejected rather than reused.

No identifier or file name changes, for the reason the check-runner renaming recorded:
`ooo` and the existing paths are cited by dated measurement records and frozen run
archives. The name lives in the concept map with its aliases, and docs/glossary.yaml is
untouched because it owns the repository and process vocabulary, not product concepts.
The definitions were spread over the design, the obligations ledger and several decision
records, so reading them meant opening four files at once. The naming record now carries an
index section: one line of meaning and a pointer per name, for both levels, the board verbs
and the channel rules, the run and its facts, the legal action set and the dispatch loop.
It ends with the two words deliberately not used for this protocol - scheduler and OoO -
each with the owner of its real meaning, because reusing either is what made the discussion
ambiguous in the first place.

The section states its own status: an index, not a second specification, and where a row
and its owner disagree the owner wins. Same content in the zh pair.

Found while checking the links: the concept map's Task Board row points at an anchor that
does not exist in memory-graphs.md, and docs:check does not catch it because it validates
file existence rather than headings. The index uses the heading that is really there.
The row's anchor named "2.1 Task Board outside the three memory graphs", which is not a heading in
memory-graphs.md: the section there is "Shared Task Board (cross-Agent coordination, not a memory
graph)". Clicking the old link still opened the file, so it read as working. docs:check validates
that a link's file exists rather than its anchor, which is why this survived; the row now names the
heading that is really there, in both languages.
…bout speculation

The naming record now carries the conceptual history the names only make sense inside: the board
becoming a protocol on 2026-08-13 (identity and discovery that wake no LLM, directed delivery, serial
admission decided by claim/resolve/expiry), governance and addressing on 09-06, out-of-order
execution borrowed from the CPU with its limits written down on 09-09..09-11, the task unit's
semantics on 09-13, the paid arms measuring both borrowed halves on 09-18..09-19, and the board
absorbing the execution half through 09-20.

The OoO step is stated as the CPU correspondence rather than as "parallelism": out-of-order execution
with in-order commit, where the plan and its tickets are the reorder buffer, the coordinator is the
only retire stage, and the acceptance rules define what counts as commit. Its two limits are recorded
with it, because the second one is the point the fusion and speculation arms then measured: the CPU's
premise that a wrong guess wastes resources already committed or free does not transfer to agents,
where a wrong guess spends tokens and wall time (a worker call measured at 40-56 s). Hence the
decision adopts only the class whose wrong guess costs no tokens and refuses the class that hides a
wait by guessing it, and the arms record "cost with no gain" at the shape tried: 43k tokens over six
units, one false fact wasting 20 332 tokens, 0 of 3 prepared candidates publishable, and 175 ms of
verification against 6.2 s of work once the fact holds.

Commit-level lineage stays where it already lives (implementation-lineage.md, from its Task Board
row) rather than being copied into the record.

Also corrected: the bootstrap design still said the speculation decision was proposed and that the
no-speculation rule therefore stayed in force until it was implemented. The decision has been
implemented since 2026-09-11, so the sentence now states what is actually in force - the host-side
free class may be done, hiding a wait by guessing stays forbidden.
Adds a proposed record for the question the naming decision unblocked: the primitives have names and the
shared layer can already compute the legal set (an ordered legal set, that set cut to the declared slot
budget, and a refusal reason per unit), but no document states whose decision each step is.

The proposal splits the work three ways and gives the program exactly one share. The protocol owns what
the primitives are, the definition and criteria of legality, the named constraint list, and the two
structural rules that cannot be negotiated because they are what "legal" means here: a claim (lease plus
attempt fence) is the only arbiter of who holds a unit, and a deliverer never judges its own delivery.
The accompanying program computes the ordered legal set from the declared plan and current facts, cuts it
to the declared slot budget, and states for each unit why it is legal or not - it chooses nobody, adopts
nothing, judges nothing, wakes nobody. Agents own the arrangement: the plan's content, the order they
agree on, who takes which unit, who adopts a run, who judges - all of it as board facts, so the
arrangement is readable instead of being implied by a program's choice.

Constraints are named by the protocol and enabled by the plan, so a preference such as repair-first is
neither program policy nor advice; and legality is asked rather than published, because a published
handoff occupies a serial slot by the board's own protocol, so fusing the two would make a question cost
a slot.

The plan is three steps and only the first lands here: this record, then a readable answer through an
existing board action, then promoting the preferences that still live as shared planning policy
(repair-first first) into constraints a plan declares - which is what makes the "not enabled means
absent from the answer" criterion true. The record names that third step as the one place the proposal
is not yet true.

Deliberately not done here: no product surface change, no driver, no wake, no new tool or channel.
The proposed record now relates to board governance and capability addressing (2026-09-06), and its
Problem says what that record does not cover: the board's correctness line - reviewable finalize,
authentic content, scope isolation, capability addressing, and the deliver/judge/attempt-fencing slice -
governs how an entry is treated once it exists, not who decides what happens next.

The two records are two halves of one protocol, and reading them together is what makes the
distributed-systems half visible: claimTaskBoardEntry is a single atomic compare-and-set ("lease-based
claiming ... on a losing CAS the failure is diagnosed against a fresh read"), lease expiry is enforced
lazily with no background sweeper - that is a suspecting failure detector, not a death certificate -
attempt fencing exists precisely because a lapsed lease cannot be trusted ("without this, work
reassigned to another agent could still be read through a stale artifact"), and the deliverable/verdict
pair binds an acceptance to an artifact identity. Idempotent acknowledgements and the never-recycled
monotonic identity counter are the other two primitives already in place.

No prose in either record names this, so every discussion of the boundary re-derives it.
…istributed concepts behind them

Your read was that part of this is a distributed-systems problem, and reading the legality record
together with board governance showed the framing has no home: the governance record is an implemented
decision with a fixed format, and the 2026-08-13 document is board mechanics. This adds the inventory as
a design document, and the naming record now points at it from its history section.

It lists fifteen parts with owner and state: entries and addressing, identity and authenticity, ownership
and time, termination integrity, attention and wake, scope and visibility, truth versus coordination
state, the task unit, legality and admission, ordering and constraints, budget and accounting,
cancellation and fencing, recovery and replay, handoff, and evidence. The task-unit protocol is the
complete one, and no single part completes a task: a unit's declaration, ownership and time, termination
integrity, attention, legality, budget, truth separation and handoff have to hold at once.

It then names seven gaps, each with the concept that fits it rather than a new vocabulary: the adopter has
no clause (a commit boundary, and orphan reaping for the restart responsibility), legality has no readable
answer (admission control with a read-only precondition check), constraints have no declaration
(declarative policy, plus deterministic replay for comparability), read guarantees are unwritten
(monotonic reads, read-your-writes, causal consistency across one agent's sessions), disagreement has no
arbiter beyond veto and racing a claim (optimistic concurrency with conflict detection, or a plain
single-writer decision since one arbiter already exists), cross-run resources have no mutual exclusion
(a fenced mutex, not only the slot budget and the throwaway-worktree practice), and in-doubt work has no
resolution (an in-doubt transaction: the fact is known, the outcome is not).

It also records what is deliberately not needed - consensus, leader election, quorum, two-phase commit,
exactly-once delivery, distributed transactions, because truth lives in one store written by one daemon -
and three analogies that mislead: a lease expiry is a failure detector's suspicion and not a death
certificate, which is why attempt fencing must stay; the board is not a message queue, so re-reading is
normal and re-doing is the problem; and agents are not replicas, their independent judgement being
exactly why a delivery is recorded with a judge instead of a self-report trusted.
The primitives have names and the program's job is proposed, but the frame itself was undefined: which
fields exist, how values are encoded, where they live in the store, and how a peer protocol sits beside
the Task-Unit Protocol. Researching the finished specifications changed the shape of the problem - every
layer but one is already defined by somebody - so this record chooses and names rather than invents.

What it borrows, per layer: CloudEvents 1.0 for the envelope (only id/source/specversion/type required,
context attributes inspected without deserializing the event data, lowercase names of at most 20
characters, data reserved); A2A 1.0.0 for the task model (terminal states, contextId, cursor paging
chosen explicitly over offsets, historyLength, includeArtifacts false by default); NLIP/ECMA-430 for the
payload discriminator and the progressive shape (messagetype/format/subformat/content/submessages/label,
the first submessage inlined because most messages carry none, structured with subformat uri carrying a
URI instead of bytes); MCP 2025-06-18 for LLM-facing results (content blocks with annotations, structured
content beside text, isError so a failure is data); RFC 8941 for a one-line text type system; FIPA-ACL for
the act and conversation fields; RFC 6709 for the unknown-field choice; ECMA-434 for the security layer.
The one layer nobody defines is ownership, lease, fencing and legality, because every agent protocol
assumes a task belongs to one agent.

Decisions it records: the frame is a state record and not a message, so FIPA's performative is a
rendering of existing columns rather than a column; the header stays in columns because the claim is one
atomic compare-and-set UPDATE and a header inside JSON would turn that into read-modify-write; the
payload is one opaque document, with the invariant that every payload can be NULL and T0/T1 still render
while claim, wake, expiry and compact read still work (verified on the store's SQLite 3.53.3); a payload
field a protocol wants to query is indexed by a generated column plus an expression index rather than a
new board column, which is what keeps the schema from growing with the number of protocols; the protocol
is a value with a version, a peer protocol supplies declaration validation, legality, projection and
acceptance, and it plugs into the existing DispatchBoard and RoundQueryPort rather than into the board;
unknown protocol means readable and not actionable, and unknown fields are decided by the protocol's own
version rule, stated rather than assumed.

Data format, in three separate questions: JSON on the existing RPC wire (protobuf-as-normative-source is
not worth it with one binding); flat key/value for T0/T1 and natural language for negotiation, because
format restrictions measurably degrade reasoning and escaping is a failure source, with tool schemas
shaped for strict mode (additionalProperties false, every property required, absence as nullable); typed
columns plus one JSON payload document in the store. TOON and CBOR are deliberately not adopted: TOON's
saving applies to flat uniform data pushed into context, which is what pointer-first avoids, and CBOR
waits for a transport with a bandwidth constraint.

Security is the layer that is empty. ECMA-434 is mandatory for NLIP conformance and its companion
guidelines score fifteen threats with prompt injection highest; two rules are adopted as text now (an
in-band session token is a session credential and must not reach the model, and a caller's token is never
forwarded), and a threat list of our own is owed.

The scope ceiling is stated so the change cannot grow: one group of nullable columns, one boundary
adapter, one rendering rule, one test. Anything needing a new table, a new tool or a new channel is
outside this decision.
…o retractions

Four review points, all accepted, plus the code checks behind them.

1. The field mapping was missing, and it changes what the invariant means. The record now carries an
   existing-to-proposed table: content stays the single source of the entry body (readTaskBoardPreviews
   projects taskBoardPreview(entry.content) - whitespace collapsed, 200 characters, a memory=<id> body
   returned verbatim), so "payload NULL still renders T1" is a property of the board's read path and not
   something the storage format produces by itself; the frozen declaration was never a body, it lives in
   task_run_manifest and task_run_tasks reached by run_id and task_id, with task_run_facts' own payload
   column as the precedent; and the proposed payload carries nothing the Task-Unit Protocol needs, so it
   stays NULL for Task-Unit entries. A rule follows from it: a payload may not restate a field that
   already has a column or a table, which is how the second task body would appear.

2. "The insertion points already exist" was not true. RoundQueryPort carries two typed reads, cancelled()
   and accepted(); the dispatch loop fails a claim whose ticket has no patch, so patch work is the only
   work shape today and a peer protocol must bring its own. The record now says what is reusable
   material and what is still a plan step, and adds that step: widen the query port or add one beside
   DispatchBoard.

3. Two storage arguments are withdrawn. Atomicity does not require columns: measured on the store's
   SQLite 3.53.3, a conditional UPDATE over a JSON payload (WHERE json_extract(payload,'$.state')='open')
   succeeded once and changed zero rows on the second claim, identical to the column form - so the reason
   is clarity, direct indexing, CHECK constraints and PRAGMA table_info migration guards. And a generated
   column is a column, an index is a schema object: peer protocols are not schema-free. What is bought is
   narrower and still worth it - the added schema is purely additive, belongs to the protocol rather than
   the board, and no field is stored twice.

4. Three borrowings were applied too widely. source_session_id is provenance, not a credential, so the
   security rule is now stated as a constraint on what may be added rather than a defence of what exists.
   The refusal rule now keeps MCP's split: a domain refusal becomes data, a request-level error stays an
   error - which the dispatch loop already does by separating refused from failure. And the CloudEvents
   naming constraint is stated as borrowing (lowercase letters and digits; twenty characters
   recommended), with protocol_version's underscore recorded as a deviation rather than as compliance.
…eded

The legality record and the frame record each needed the same sentence and neither stated it: the legality
record says what the program does and what it refuses to decide, the frame record says what the core knows
and what it never parses. Those are one division seen from the side of the decision and from the side of the
data. Because nothing named it, it had to be re-argued every time a field or a rule appeared, and the cost
was already visible: "work means a patch" had reached six places, one of them a legality rule written in the
kernel's own vocabulary (refuseWidening decides permission closure by reading parent.patch.editable).

This record takes Hydra's principle - a kernel provides mechanisms and refuses policy - and states it as two
contracts. Mechanism: one claim (compare-and-set, lease, attempt fence), an opaque declaration with a digest,
an opaque artifact with a digest, a verdict from someone else, the lifecycle, and the legal set (which units
are legal now, in order, cut to the declared slot budget, with a reason per unit). Policy: what a unit's
inputs are, the artifact class, how the artifact is judged, how the work runs, its concurrency
preconditions, the wording and enablement of legality rules, and who is chosen, who adopts, who judges.

It also supplies the two things the slogan lacked. A decision procedure - for each field, type and code path,
ask mechanism or policy, with a rule's check counting as mechanism while its name, wording and enablement
count as policy. And a mechanically checkable rule: no policy word may appear in the mechanism layer's code,
type declarations or schema paths, from a maintained list (patch, editable, files, instruction, checks,
repair-first), with the path carried because a word can be policy in one place and mechanism in another.

The consequence for work shapes is stated: a work shape is not a core concept, it is the policy layer's name
for the declaration-and-artifact pair the core carries opaquely. The six leak sites are classified in a table,
and the legality rule found today is the refinement the record adds - its check belongs to the program while
permission closure's name and enablement belong to a declaration, so the fix is not to move the check out of
the program but to stop hard-wiring the rule in the kernel's vocabulary. That is the same shape as plan step 3
of the legality record, where repair-first becomes a declared constraint instead of shared planning policy;
two independent fixes taking the same shape is the evidence that the classification is right.

Three rows are deliberately left undecided so that they are not settled by preference: task_run_tasks.effect
is mechanism only if the mechanism must enforce write-set disjointness, the granularity of input and
dependencies is a separate question from whether dependencies are mechanism, and operation may be policy.
Alternatives rejected: keeping the slogan and deciding case by case, putting the principle in either record
(it is wider than both), a plugin registry above the core with no second policy to justify it, and deciding
the undecided rows now. Risks named: gutting the core (so the practical form is one default policy the core
may carry but must not require), an ossifying word list, an undecided row becoming permanent, a policy word
legitimately remaining for a release, and grep being gameable by synonyms.

The legality record and the frame record now point here, so the division has one home.
… earns a place

Records the model that came out of the last exchange as a draft rather than as a record, because it was
reached by stacking analogies and nothing about it has been measured.

The model: a drawing pipeline has four stages and so does ours - producer (a work-shape adapter), surface
(the board entry: payload plus a discriminator and a digest, delivery as commit, resolve and expiry as
release, wake as the frame callback), compositor (the board and the program: claim, lease, fence, serial and
wake, expiry and reaping, the legal set), and display (the presentation end: tiers, flat lines, compact read,
the tool descriptions an agent reads). From it come three ownership classes: content and semantics belong to
the owner and are semantically opaque to the middle (a compositor reads pixels but cannot read what a button
means, so it asks the owner; our board reads declared structure and never payload semantics); description is
the owner's obligation (geometry, level, shape, dirty regions, terminal state, identity, a snapshot able to
stand in for the owner - ours is the discriminator, digest, dependencies, scope, and the protocol's
projection); and arbitration is only the middle's, because stacking, occlusion, visibility, focus, capture,
timing and reaping are global properties. The boundary test that replaces a slogan: no reasoning needed means
the object stays opaque; structure needed means the owner supplies a description and the middle reads the
description but not the semantics; a global judgement means only the middle can make it, which is why the
description must be cheap and fresh and why staleness needs a meaning.

The document also carries its own case against itself: the analogies did the reasoning and were each partly
overruled by a fuller view, one of the four stages is a name with no seam or owner or test, the only
measurable claim in this arc (that the seam was narrow) was falsified by a grep within a minute, and nothing
got smaller - no code changed, no concept removed, three prose records added.

So it states an eligibility rule - an analogy earns a place in a record only when a check can falsify it -
and records its predictions before measuring: (a) policy words in the middle, predicted at fifteen or more
hits concentrated in the board and the legality module; (b) the two ends, predicted to be concentrated and
separable at the producer end and scattered with no seam at the presentation end; (c) a second shape through
the dispatch loop, predicted to change five or more files. The reduction test is named in advance too: less
code (the middle must lose a type and a freeze method, or the model failed), fewer concepts (the three
classes must absorb the two lists they replace rather than sit beside them), fewer change points (measured by
how many places a new shape touches), and better maintainability (a new shape adds files under one owner and
changes no test in the middle).
…he counting exposed

Both checks were run against the prediction recorded before them.

(a) Policy words in the middle: predicted fifteen or more hits, measured 195 raw - patch 106, files 39,
checks 23, editable 14, instruction 12, repair-first 1 - concentrated exactly where predicted, in
src/integration/ooo-board.ts (53 patch hits), src/integration/task-semantics.ts (27 patch, 9 editable) and
src/core/store/base.ts (13 patch). Two corrections forced by the measurement. The raw count overstates the
case, which is the caveat the check was written with: files in src/core/store/writes.ts and retrieval.ts is a
mechanism word (a path inside a store) and checks in src/integration/ooo-candidate.ts names the check runner,
also mechanism, so the word list drops files and checks and keeps patch, editable and instruction. What is
left is still about 120 hits, so "the middle is policy-free" is false as a description of today, an order of
magnitude past the prediction, and the model's support is directional: the leak is real, large and
concentrated in three files.

(b) The two ends: the producer end was predicted concentrated and separable and measured four to six
functions over three files (preparePatchWork, patchPrompt, patchCandidate, patchSubmission, snapshotText,
runTestFile); the presentation end was predicted scattered with no seam and measured six sites (the preview
text, the entry and preview types, the wire shape, the service, the agent-facing renderer, and the generated
tool descriptions). Both held. The measurement also corrected this document: the display row claimed T0-T3
tiers, and no such thing exists for a board entry - tieredDisclosure and tier belong to memory retrieval, and
a board entry has exactly one presentation, the 200-character preview plus raw fields. That row now says what
exists, so the second borrowed vocabulary is recorded rather than left in the table.

(c) A second shape is not run yet, and the reduction test is unchanged: nothing has got smaller, the middle's
three files hold the available reduction, and whether it pays is what (c) measures. The counting earned its
keep twice - it produced the word-and-path rule the earlier prose could not state, and it caught this document
borrowing the memory side's vocabulary as if it were the board's.
The model in mechanism-in-the-middle.md is retired and replaced: a board entry is a medium with a
write face and a read face, the compositor is brought by a protocol rather than by the board, and the
roles are roles rather than files - src/core/store/base.ts is the medium, the compositor and the
reader's rule at once, and src/cli/service.ts is both the wire face and a formatter.

The rule that decides where a seam goes needs no taste, only a count: a boundary is built when it has
a second implementation, and only declared with one. The read face has three, so it lands here.
src/core/board-entry-view.ts owns the rule, the store's compact read and the host adapter's broadcast
both call it, and the store's private copy and the adapter's bare slice are gone. One of the three was
not merely duplicated but wrong - the adapter's 140-character slice did not collapse whitespace, so a
multi-line body reached a broadcast as a multi-line broadcast.

The write face's type and the middle's seam are declared and not built, each with a named trigger.
Check (c) measured why: making the loop shape-agnostic is one file with fifteen insertions and
eighteen deletions, and tsc then reports zero errors while twenty-five tests fail, every one of them
because the claim admits no work. The coupling is invisible to the compiler and fatal at runtime, so
the seam is prepaid cost until an adopter declares a non-patch unit.

Docs and code land together, because the read face's home is what the two-face model implements.
The port now has a third implementation, and it is the product's.
src/integration/ooo-runner.ts projects the run's frozen table and its board entries
into the facts the shared rules read, and carries no legality rule of its own: the
plan is compiled by compileTaskUnits, the legal set is dispatchTasks + selectableTasks,
cancellations are read through the coordinator's own reader, and the verdict belongs to
the caller's acceptance. A rule that turns out to be needed here belongs in
task-semantics.ts instead - needing one is how this module would show that the
abstraction leaked.

Two things the projection cannot read from the store stay the caller's, and both are
existing divisions rather than new ones: a task's declared file contents (the store
freezes paths, and preparing the workspace is the patch path's caller's job) and an
acceptance whose identity is independent of the deliverer (the store already refuses a
deliverer that judges its own delivery, which the second test pins).

One contract is easy to get wrong, so the projection writes it down: the shared
acceptance rule binds a verdict to the artifact value the run carries - it compares the
verdict's digest against the artifact - and not to the store's own deliverable hash.
Reporting the hash there makes every accepted unit read as unaccepted, and the symptom
is silent: the next unit is never released and nothing reports an error.

The port's "put a result" is the product's delivery on the entry the unit was claimed
on, so a unit's work is recorded once rather than once on that entry and again on a
result entry nobody would read; and the unit is resolved from the claim rather than
from the body, because a body's shape belongs to its owner while a claim is the board's
own fact.

Verification: npm run check clean; test:product 1527/1527 pass (2 new - a two-unit plan
driven through the product's store with an independent judge, and the self-judging
refusal); build clean; lint clean; prettier clean; docs:check 288 files, 0 errors,
0 warnings; agent:context:check valid once the route claims the new file.
CI failed "Product tests and coverage" on one case of 1527 while the same suite passed 30
times in a row locally: "a value that expired a moment ago is still current" stamped an
expiry half a grace (25 ms) in the past in one statement and read it in another, so it
asserted that the read happens within the half that is left. A sweep between the stamp and
the read says where it flips - 0 ms and 10 ms answered "current", 20 ms and beyond answered
"not active" - and a loaded runner has no trouble spending 20 ms reopening a store. Sharing
a clock source is not sharing a clock read.

The case is now a relation: the store's own notExpired predicate is evaluated against a
boundary stamped in the same statement, so both share one 'now' and the expectation is
about the window rather than about the machine. It also pins the boundary instead of a
point inside it - at exactly one grace an expiry is already excluded. 8000 evaluated
assertions under deliberate load, no flip. A row-level form of that case cannot exist,
because a row's boundary is stamped by one statement and compared by another, so the store
level keeps the cases that do not depend on a duration: a future stamp inside the grace, a
minute either way, and 400 write-then-read rounds.

The mutant that names this case as its catcher is pointed at the new name and still fails
it: 4 of 4 caught, restored byte-identically. Postmortem 0004's guardrail claim and its
lessons are corrected with the recurrence: its "deterministic" test was deterministic on
three of its four boundary cases, and the fix for a timing dependency was itself one.
Three were in the store board's own test, which `npm run check` cannot see because tests
are outside its project. One of them mattered: the self-judging tooth read a digest off the
ticket, where the port's PatchWork is the pre-freeze declaration and carries none - so it
delivered an artifact with no digest at all, and the store took it, because a delivery is a
digest and a reference while the wire shape is the host's check rather than the board's.
That test now freezes and builds its artifact the way the loop does, through one helper the
worker shares instead of a second copy of the artifact shape.

Two more were the store's private single-entry reader reached from tests, in this file and
in tests/core/task-board.test.ts where they had been reported for a while. Both now use
getTaskBoardEntryById, which is the same reader plus the channel guard the store asks for:
the test says which channel it reads, and the private method stays private.

The five in evals/ooo-execution were declared shapes that did not match what the code uses.
The two evidence drivers refused a missing flag in a loop over an object, which narrows
nothing, so one of them asserted non-null at its use site and the other did not compile at
all; both now call a flag() helper that returns the value it checked. live-continuation.ts
declared its per-attempt result as the metrics shape alone, which made `artifact` - the
thing a continuation continues from - a property the file did not have; its type now
follows the call that produces it.

Verified: lsp_diagnostics reports 0 across the five files; check, lint, format:check,
docs:check and agent:context:check clean; test:product 1527/1527 pass.
A verdict binds to a digest rather than to bytes, so the digest's convention is a rule a
stored record depends on - and it had four module-private homes. Two of them were the
identical two lines (`createHash("sha256").update(JSON.stringify(value)).digest("hex")`)
in `task-semantics.ts` and `ooo-board.ts`, and two callers had folded their own shortening
into their copy, so a twelve-character report identity and a sixteen-character branch
identity read as if the shortening were part of the rule. `src/integration/work-identity.ts`
now owns it: `workDigest(bytes)` for text or raw bytes, `workDigestOf(value)` over JSON
text, sha256 in hex at full length.

Canonicalisation stays the caller's, and the module says why: JSON key order is whatever
the caller built, so a caller that needs one identity per shape rather than one per
serialisation freezes the shape first - which is exactly what `preparePatchWork` does when
it sorts the files, serialises once, and returns a digest rather than a serialisation. The
shortening two callers want stays theirs, and it now says so by slicing.

Sites that differ on purpose keep their own convention, and the module names them so a
later sweep cannot unify three different rules: a search index's base64url content hash, a
session identifier, and a protocol-visible `sha256:` prefixed identity.

Two named mutants make the convention checkable - another encoding, and a JSON variant
that does not digest JSON - both caught by the named case, restored byte-identically.

The write face's type did not land, and this commit records the measurement that says so
rather than building it: the artifact body inside a result already has a home
(`artifactEnvelope`/`artifactFromText` in `ooo-session-mechanism.ts`), while the entry body
wrapping it has one writer and two readers that are two *different* rules - the probe parses
the loop's envelope and binds a verdict to the ticket's digest, the product board reads the
artifact of a body that may be the artifact itself. Two rules with one implementation each
are not a seam.

Verified: check, lint, format:check, docs:check and agent:context:check clean (the new file
is claimed by the ooo-execution route); test:product 1529/1529 pass twice; the work-identity
mutant target 2 of 2 caught.
Asked whether the design is complete, the authority to answer with is
docs/design/task-unit-semantics-obligations.md, and it contradicted itself in two
places - one in each direction.

Its counts paragraph said "the only arm still unrun is the E arm itself", while the
E arm's own section, the row table and the closing paragraph all say it ran twice on
2026-09-19 with no gain claimed. A reader trusting the opening paragraph would think
the measurement phase was one run short of closed; a reader trusting the closing one
would think it was closed. The sentence now points at the two records that settle it.

Its verification section recorded the LSP diagnostics as "recorded rather than
repaired", naming one error each in board-deliver.ts and board-judge.ts plus a
{ pid: 0 } fallback in an evidence-driver test. Those are repaired now: the drivers'
flag narrowing and live-continuation.ts's declared result shape landed in e31fe77,
and this commit fixes the test's fallback, which needed the startedAt that ServerState
requires (a pid of zero and an empty start time still says "no server", which is what
the code means).

What the LSP still reports is stated as the number it is: 31 diagnostics in seven
evals/ files this arc does not own - benchmarks/run.ts 7, longmemeval/run.ts 13,
controller/run.ts 3, natural-maintenance/audit.ts 3, hierarchy-scale/run.ts 2,
longmemeval/score.ts 2, omnimemeval/bridge.ts 1. The sentence claims zero only for the
files this arc touched, because that is what was measured; the tests/ directory as a
whole timed out rather than answering, and an unmeasured claim is not written down.

Verified: docs:check 288 files, 0 errors, 0 warnings; format:check clean; lint clean;
tests/integration/ooo-evidence-drivers.test.ts 8 pass, 0 fail.
`selection` now returns its own reasoning beside the legal set: every gate it
applies names itself, and `unitLegality` reports the ordered legal set with a
cause list per unit. `deriveStatus` carries the same causes, so the status query
and dependency release keep sharing one predicate and now share its explanation.

The causes are computed inside `selection`, where the gates are, rather than
beside it: a second copy of a gate's condition is how an answer starts
disagreeing with the decision it explains. A plan-level gate is attached to the
units it holds back, so an empty legal set is never a bare empty set.

Two cases pin it: one names each gate in turn, one walks every combination of
the flags a plan's facts can carry (64 shapes, two budgets) and refuses to
accept a silent refusal. The mutant target grew two teeth for the new rule -
`a-refused-unit-is-silent` and `a-stale-input-is-not-named` - and the two
selection gates whose lines moved were re-pointed, including the budget mutant's
named case, which now names the case that actually distinguishes it. 24 of 24
mutants are caught and both files were restored byte-identically.

This is the shared half of the legality proposal's step 2. The board action that
lets a caller ask the question through a surface it already has is still to
come, so nothing here claims the answer is readable yet from outside the process.
The port gains the query the legality proposal's step 2 asks for: `legality()`
returns the ordered legal set, the room the run has left, and a named cause per
unit the rules do not have on offer. It calls the same `selection`, so a caller
cannot read a unit as legal here while the rule refuses it there, and
`candidates()` - the port's other read - is now that answer's `legal` on the
store's board, which is what keeps the two from drifting apart.

All three boards answer it: the store's, the probe board's (whose projection
moved into one private method so `ordered()` and the answer cannot disagree), and
the dispatch test's stub, which hands its own fields to the shared reading rather
than keeping a second set of rules. The asker is the process that owns the run's
workspace: a plan compiles from the caller's files and the daemon holds only
their frozen paths, so the answer cannot come from the store alone.

The product's case asks twice on an unchanged store, asserts the two answers and
the recorded facts are identical, that the answer names nobody (the store's own
refusal names the holder; this read does not), that a claim moves a unit to the
named cause while the claim itself stays the store's, and that the answer follows
the facts it reads rather than a cached plan.

Docs land with it: the two-faces document's read-face bullets, the proposal's
step 2 marked landed with step 3 named as the place it is still not true, and the
umbrella's gap 2 narrowed to the asker that has no workspace. `ooo-board`'s and
`ooo-execution`'s mutants were re-run - 23 of 23 and 24 of 24 caught, both files
restored byte-identically - and the product suite is 1532 of 1532.
Repair-first was the shared planner's own default, so the meaning of a plan
depended on which planner read it: the online move that prefers continuing a
session lived in `nextSessionMove`, and the run's recorded `policy` was a name
the probe board compared rather than a switch anything read. The legality
record's third step is that a constraint the protocol names and a plan enables,
and it is what makes "a constraint that is not enabled has no effect" true
rather than intended.

The protocol's side is a closed list. `PLAN_CONSTRAINTS` holds the names that
exist - `repair-first` is the only one so far - and `enabledConstraints` refuses
a name outside it by name, because reading an unknown name as "nothing was
asked for" is exactly the decoration the rule forbids. A plan that enables
nothing is a setting rather than a fallback: it runs one unit per session, which
is the baseline the arms' control cells measure.

The plan's side is a declaration. `SessionPlan` gained `constraints`, the move
is computed under it, and `optimisticPlan` copies it, so the offline projection
prices the same plan the online half decides with.

Every caller that wants a fused run now declares it: the arm driver's spec
carries the declared set, and `validateSpecFile` refuses a spec that declares a
bound above one without enabling the constraint - such a run would be the
control arm while its file said fusion, which is the one confusion that driver
must not create. The trial spec generator writes it, and three test fixtures
carry it.

What is demonstrable is a different move at a session boundary under an enabled
constraint, not a different legal set: repair-first orders a session, it does
not gate membership. The record's criterion said legal sets and was corrected
rather than quietly restated, and the record moves from proposed to implemented
in the same commit as the code, with the umbrella's rows 9 and 10 updated.

Evidence: two new mutants on `src/integration/ooo-fusion-plan.ts` (2 of 2 caught
by the named cases: the continuation not declared, an unknown constraint
ignored), the four re-run targets 37 of 37, `npm run test:product` 1534 of 1534,
`npm run verify:static` exit 0, `npm run docs:check` 288 files with 0 errors.
The sweep refuses to claim a check when a mutant's marker no longer matches the
source, and two markers had drifted away from the code they were written
against: `the-status-query-ignores-the-declared-budget` still named a
`startableTasks(dispatchTasks(units, facts), slots)` call that `deriveStatus` now
splits, and `a-cancelled-run-still-accepts-writes` still named a two-line
refusal that `managedWriteRefusal` now assembles from a subject. Both were
reported as "marker not found, refusing to claim a check" rather than as caught,
which is the honest failure and also a hole: nothing in the build notices a tooth
that stopped touching its rule.

Re-pointed, not rewritten: the budget mutant now hardcodes one slot in the call
the module makes, and the cancellation mutant inverts the guard inside
`managedWriteRefusal`. Measured after the fix: 27 of 27 caught by the named
cases, 2 of 2 files restored byte-identically.
89% of the mutation ledger's teeth (133 of 149) are a site plus one of a handful of operators, and 52%
of them (78) are a bare byte fragment matched anywhere in the file. That is the shape that dies first:
both stale teeth repaired in dbf10ac were of it. This lands the first two slices of the plan to derive
the mutant instead of storing it.

The vocabulary (tools/mutation-anchor.ts, new): six operators over a member-scoped selector -
condition-never, condition-holds, neutralize-term, replace-argument, replace-property, drop-statement.
Every selector resolves to exactly one site or refuses, and the refusals are the interesting half: a
member named twice, a call matched twice, a fragment that fits two decision positions (the innermost one
wins; two disjoint ones are refused), an operator that replaces a whole condition given a fragment that
names only part of it (that would silently widen the mutant into "every reason this rule has"), a
statement that does not own its line. A condition is any expression in a decision position - a test, a
returned value, an arrow's body, a value bound to a name - because `return a && b` and
`array.filter((x) => x.y)` are the same rule written without an `if`. The resolver holds no state, so it
moved out of the sweep script: it is 16 tests over source strings, with no filesystem.

The tool (tools/mutation-teeth.ts) keeps `from`/`to` and `ast` forms working and gains `--anchors-only`:
every selected anchor resolved, nothing written, no suite run, 0.6s for all 149 - the pass that answers
"is the tooth still aimed at something", which nothing answered before.

The pilot (src/integration/ooo-execution.ts, 24 teeth over two target entries): 19 converted, 5 staying
hand-written and named in the record, each still a single expression or value rather than a statement or
a message. Sweep after conversion: 24 of 24 caught, restored byte-identically 2 of 2. 23 are caught by
the case their `expect` names; the stale name fusion-continues-from-an-unverified-answer was already
that way and this conversion neither caused nor fixed it.

The demonstration, and the prediction that was wrong: a rename plus `spent || severalWaits` lifted into
`const closed` retired two sites - the hand anchor, and the derived a-live-claim-does-not-block-selection,
because a condition bound to a name was not a decision position yet. The prediction written before
running it had been "one failure". The selector was widened to include a variable's initializer, and the
corrected prediction held exactly: anchors: 148 of 149 resolve, the one failure the hand anchor, file
restored byte-identically. The limitation is real and recorded rather than hidden: a derived selector is
scoped to a decision position, so a rule rewritten into a different shape can still retire its tooth -
but now it does so visibly, naming the tooth and the fragment.

Not in this commit: the anchors-only pass as an atomic check of the ci-and-tests route, and the
retirement pass over the ~18-20 teeth whose rule already has a relational or enumerative check. The
record (docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md, with its zh-CN pair)
marks both, and holds the measurement behind the operator list.

Readings at this revision: npm run mutation:teeth -- --anchors-only -> 149 of 149 resolve over 23
targets, exit 0; the target sweep -> 24 of 24 caught; tests/tools/mutation-anchor.test.ts -> 16 pass,
0 fail; npm run verify:static -> exit 0.
A tooth that stopped matching its rule is invisible today: no CI job runs mutation:teeth, and in a sweep
it is reported as "not applicable" and excluded from the denominator - which is how the two dead teeth
repaired in dbf10ac stayed dead without anyone noticing.

npm run mutation:anchors runs tools/mutation-teeth.ts --anchors-only: every selected mutant's site is
resolved and nothing is written, no suite runs, and a site that cannot be resolved fails the run. It is
now one of verify:static's checks and one of the ci-and-tests route's atomic checks, in the same order
the route-contract test enforces (the route's list must equal the chain, test:product included).

Measured on this revision: verify:static exit 0 in 41 s; a full npm run agent:verify exit 0 in 112 s
inside its 150-second budget, with the anchors pass itself 0.98 s, all 149 anchors resolving over 23
targets.

The pass proves a site still resolves, not that the mutant is still caught - a stale expect still passes
it - so the full sweep stays the standing rule before a push and stays out of the gate.

docs/design/ci-cd-and-quality.md gains the check and the reason it is in the gate; the record
(docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md, with its zh-CN pair) marks its
third slice landed.
The retirement pass, first group. The criterion, applied one tooth at a time: a tooth may go when a
named case's assertion fails for exactly the violation the tooth introduces, and that case's wording
states the rule - so the ledger row names the check instead of the mutant. Judged by reading the
assertion, then by re-running the target's sweep with the tooth gone.

- the-board-read-path-stops-calling-the-predicate, the-board-decides-acceptance-on-its-own (row B5):
  tests/integration/ooo-acceptance-one-predicate.test.ts counts the call sites itself - one definition of
  the predicate in src/, one `acceptedFact({` in each reader, zero `verdict ===` comparisons in
  ooo-board.ts - so a rename that stops the call and a second decision inside it each fail a count. A
  behavioural differential would have caught neither; a count does. This pair is what made the criterion
  worth writing down.
- a-refused-unit-is-silent: the case enumerates the 32 flag combinations a plan's facts can carry and
  asserts every refused unit has a reason.
- the-merge-enumerates-one-order (row A5): the case asserts the multinomial count (10), which a merge
  returning one order fails.
- a-second-entry-rebinds-the-task (row D12): the case asserts a retry is the same binding and a second
  entry is refused by name.

Docs and code in one commit: the three rows now name the check that carries the rule and say when the
tooth was retired (A5, B5, D12 in the obligations ledger, whose reading list gains the dated readings
below), and the record's fourth plan item is marked landed with the criterion and the finding.

Finding carried forward: a-refused-unit-is-silent is named by no ledger row, and the case that catches it
is unowned too. Retiring it removed an orphan rather than a row's pin.

Readings at this revision: mutation:anchors -> anchors: 144 of 144 resolve, over 23 targets, exit 0 (the
register is 144 teeth, not 149); the four targets the pass touched -> mutants: 70 of 70 caught by the
named test, restored byte-identically 5 of 5; test:product -> 1562 pass, 0 fail, exit 0; verify:static ->
exit 0.
Second group of the retirement pass, same criterion as the first: a tooth may go when a named case's
assertion fails for exactly the violation the tooth introduces, and that case's wording states the rule.

- The five rules of a declared budget - the licence is the budget's part and not its head, a budget above
  one must name each handoff's target, every startable task gets a handoff, a startable handoff is not
  retired by a republish, and a multi-slot run directs its handoffs - are all stated by one case,
  evals/ooo-execution/board-slots.test.ts's "a declared budget holds two claims at once, and the store is
  why each handoff is directed". Its assertions name each rule: the refusal message, one handoff per
  startable task, serialState null for each, the same ids across a republish, and the non-head claimed
  first. So the-licence-is-the-head-whatever-the-budget, a-second-slot-is-declared-without-a-target,
  only-the-heads-handoff-is-published, a-startable-handoff-is-retired-as-unselected and
  a-multi-slot-handoff-is-published-un-directed are retired. The F2b-slot row keeps seven teeth, and the
  rules themselves keep a witness that says them.
- claim-is-not-scoped-to-its-run (row B1) is the tooth the namespace experiment first watched survive:
  tests/integration/ooo-run-namespace.test.ts was strengthened until it caught it, and that record names
  the assertion that did it (a raw-row read of the other run's owner). A case whose history is exactly
  this tooth is a better witness than the tooth's own name.

Docs and code in one commit: rows B1 and F2b-slot now name the checks that carry their rules and say when
each tooth was retired, the obligations ledger gains these dated readings, and the record's fourth plan
item is marked with this group, the two candidates judged not retirable so far, and the two stale expects
still outstanding.

Readings at this revision: mutation:anchors -> anchors: 138 of 138 resolve, over 23 targets, exit 0;
--targets=src/integration/ooo-board.ts -> mutants: 15 of 15 caught by the named test, restored
byte-identically; test:product -> 1562 pass, 0 fail, exit 0; verify:static -> exit 0.

One reading to carry, not a code failure: the first test:product run of this pass failed
tests/core/graph-cycles.test.ts's first case on EPERM from rmSync of its %TEMP% scratch directory, and the
file then passed 8 of 8 alone and the suite 1562 of 1562. That is the Windows temp-tree removal class the
ledger already records as undiagnosed, and the ledger now says so with this instance.
Two retirements, same criterion as the first two groups: a-retried-run-fact-is-appended-twice and
a-frozen-task-is-replaced-by-a-different-definition (row D11) go, because that row's prose already rested
on the two cases for those rules and the cases state them outright ("appending the same fact twice records
it once and keeps the first sequence", "freezing a task twice is a no-op, and a different definition for it
is refused").

The group's larger finding is what the first whole-register sweep said: 136 of 136 caught, but only 134
caught by the case each expect named. Four links were wrong, each in its own way, and all four are
repaired:

- fusion-continues-from-an-unverified-answer: its named case refused the pair for a second reason as well
  (the successor declared the predecessor as a dependency, so the dependency rule could refuse it alone),
  and a case that passes for another reason pins nothing. The pair now declares no dependency, which
  leaves the predecessor's verdict as the only condition that can refuse it, and the tooth's own expect
  case now fails under the mutant.
- next-is-not-the-head-of-the-ordered-candidates: its expect named a case in another target. The board's
  budget case now asserts next() is the head of the ordered set - the rule in its own words - and that is
  the name the tooth carries.
- the-caller-rebuilds-the-shared-floor: its expect was a paraphrase that named no case at all. Re-pointed
  to "a declining route's narrow run verifies on its own tests and nothing else", which fails under it.
- the-grace-is-zero: its named case derived its fixture from the constant under test
  (half = CLOCK_GRACE_MS / 2000), so zeroing the grace moved the stamp onto now and the case passed while
  the bug was live. The fixture is a literal now and the case fails under the mutant again.

Re-run after the repairs: mutants: 136 of 136 caught by the named test, all 136 by the case their expect
names, 23 of 23 targets restored byte-identically, exit 0 in 217 s. The class is the one this arc opened
with - a link that goes stale while the artifact still looks right - so the reading worth keeping is not
the count of teeth but the count of teeth whose named case is the one that fails.

One behaviour recorded rather than changed: a named case that does not finish inside its bound counts as
caught, with the reason printed (the-pass-asks-a-unit-it-already-failed-again, 30 s). The tool's own
comment says a run that never ends proves nothing; the code counts it as caught and says why. Resolving
that tension changes what the ledger's proven means, so it is left for a decision rather than a fix.

Docs and code in one commit: row D11 names the cases for the two retired rules and says when the teeth
went; the obligations ledger gains the whole-register reading with all four repairs named; the record's
third group notes the same, the two candidates still judged not retirable, and the timeout tension.

Readings at this revision: whole register -> 136 of 136 caught, 136 by name, 23 of 23 restored, 217 s;
mutation:anchors -> anchors: 136 of 136 resolve; test:product -> 1562 pass, 0 fail, exit 0; verify:static
-> exit 0; lint -> 0 findings.
The 27 teeth whose anchor is a place where code must not appear are the ones a derived selector cannot
express, so each of them was either going to stay hand-written forever or give way to a check that says
the same thing. 26 give way, and one is kept with its reason.

The method, applied per tooth rather than in bulk: the register-wide sweep first says which case fails
under the tooth; that case's assertion is then read to see whether it fails for exactly this violation and
states the rule in its own words. What qualified:
- a count over the store's own source - "the store runs its transaction boundary in exactly one place"
  counts one BEGIN, one COMMIT and one ROLLBACK, all three inside writeTransaction (a-method-opens-its-
  own-transaction);
- an enumeration of a set that must not change - the accepted prefix, the fallbacks by source id, the
  outcome of a scoring order, the orders the ordered mode allows, a run's tasks after a refused freeze
  (reopening-one-task-clears-every-acceptance, a-failing-source-takes-the-decision-with-it,
  the-ordering-adds-a-task-to-the-set, ordered-mode-becomes-any-topological-order,
  the-plan-freezes-one-task-per-transaction);
- a refusal by name plus a raw read back - the second reader of a ready task is refused and the owner is
  then read out of the table, the status of an unknown run stays unknown, a driver without a daemon refuses
  and no driver imports the store (a-live-claim-can-be-taken-by-another-agent, status-registers-the-run-it-
  cannot-find, a-driver-falls-back-to-opening-the-store);
- the file's own hash and a file that must not be created (the-read-only-factory-opens-a-writable-handle).

One is kept: every-task-is-frozen-at-position-zero. Its named case is a register/freeze/adopt/read-back
round-trip, weaker than "the position comes from the array order, not from the request" - the same reason
a-binding-does-not-record-its-channel stays, as recorded when the earlier groups were judged.

Seven of the 26 turned out to be orphans: no row in the obligations ledger names
round-publication-opens-its-own-transaction, round-never-releases-its-pin,
ordered-mode-becomes-any-topological-order, the-loop-awaits-each-unit-instead-of-the-batch,
a-unit-is-dispatched-twice-in-one-batch, a-unit-nothing-checks-is-still-a-unit or
the-parent-check-ignores-its-own-verdict, so each of those rules has a check and no row. The checks are
named in the record. Whether those seven rules deserve rows is a question about the ledger, and it is
left as one rather than answered by inventing rows here.

The target tools/agent-verify.ts had a single tooth and goes with it, so the register is 110 teeth over 22
entries covering 21 files.

Docs and code in one commit: the obligations ledger's rows stop naming the retired teeth as evidence and
name the case that states each rule instead (B3, B6, C3, D3, D6, D11, D13, D14, E1, E3, F2b-slot, F5, the
G rows and the prose that listed them); its readings gain the whole-register run for this pass and its
"how a row earns proven" paragraph gains the retirement criterion, including the two ways a case can look
like it states a rule without doing so (a fixture derived from the constant under test, a case that
refuses for a second reason as well) - both of which this arc found by reading assertions; the record's
item 3 is marked landed with the numbers, the kept tooth, the seven orphans and the next question.

Readings at this revision: whole register -> mutants: 110 of 110 caught by the named test, all 110 by the
case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 162 s; mutation:anchors ->
anchors: 110 of 110 resolve, over 22 targets; test:product -> 1562 pass, 0 fail, exit 0; verify:static ->
exit 0; lint -> 0 findings.
…ry close

The register-wide conversion the pilot started, taken target by target: 59 more hand-written teeth are now
a name, an operator and a selector, so a rename or a moved line cannot retire them. Per operator:
condition-never 44 (a guard or a variable initializer), neutralize-term 9, replace-property 9,
replace-argument 6, drop-statement 6, condition-holds 3. The register is 77 derived of 110.

Reading after the conversion: mutants: 110 of 110 caught by the named test, all 110 by the case their
expect names, 22 of 22 targets restored byte-identically, exit 0 in 171 s; anchors: 110 of 110 resolve.
The conversion does not make a tooth bite - the operator realises the same violation the hand anchor did -
so the reading that matters is the second one: every tooth is still caught by the case it names.

Three sites the operators could not name, each closed with a case in tests/tools/mutation-anchor.test.ts
(16 cases to 19) rather than argued about:

- A class constructor is a member. BoardAdmission's constructor refuses a second plan while it opens the
  store, and uniqueMember knew only methods and function declarations, so that site had no `within` to
  write. The name is `constructor`; matching it is what let a-second-plan-silently-adopts-the-run convert.
- A condition written across lines is one condition. The candidate filter compared raw text while the
  whole-condition check compared whitespace-normalized text, so a fragment of a wrapped `if` found no
  candidate and then matched nothing. Both normalize now, and the filter no longer demands the fragment be
  unique inside the candidate: `b` appears three times in `a && (b || !b)` and that is still the condition
  a selector naming `b` means.
- A guard clause is a statement. `if (...) throw ...;` was not among the statements drop-statement would
  remove, because an `if` is an IfStatement rather than an expression, a declaration, a `return` or a
  `throw`. It is now, and statements resolve to the innermost one containing the fragment - the rule
  conditions already followed, and what keeps a fragment from matching both a guard and the statement
  inside it.

One wrong `within` was shipped in this pass and the sweep is what caught it, which is the argument for
sweeping after a conversion instead of trusting the anchors pass: the-completion-ignores-a-cancelled-unit
was aimed at checkDispatch rather than checkCompletion (the fragment occurs in both members), so the
anchors pass resolved, the mutant was still caught - and it was caught by the suite, not by the case that
names a completion of a cancelled unit. Re-aimed, that case fails again.

What is left hand-written is 33 teeth, and they cluster by the operator they would need: a comparison
rewritten (3), a fragment inside a template or SQL string (5), an iterable emptied or a filter dropped (4),
a call or `new` replaced or unwrapped (6), an index moved (2), an initializer replaced by a different
expression (3), a condition negated (2), a statement rewritten into another statement (3), two statements
sharing one line (2), a literal swapped (1), a site at module top level where there is no member to name
(1), and one tooth that changes two things at once. Four clusters look worth an operator
(replace-comparison, replace-fragment, empty-iterable, replace-callee); the record says so and leaves the
decision open rather than adding operators nobody has asked for yet.

Docs and code in one commit: the record's item 2 carries the numbers, the three gaps and the wrong
`within`; its remaining list is the 33-tooth classification; the obligations ledger says 77 of 110 teeth
are derived, names the four candidate operators, and gains this pass's whole-register reading beside the
one it already had. The tools' own headers say a constructor's name and which statements drop-statement
covers.

Readings at this revision: whole register -> 110 of 110 caught, all by name, 22 of 22 restored, 171 s;
mutation:anchors -> 110 of 110 resolve, over 22 targets; mutation-anchor cases -> 19 pass, 0 fail;
test:product -> 1565 pass, 0 fail, exit 0; verify:static -> exit 0; lint -> 0 findings.
The 33 teeth still hand-written after the register-wide conversion were classified by what they would
need, and the classification was checked against what mutation tools actually name (Stryker's supported
mutators, pitest's mutator list, cargo-mutants' patterns, Cosmic Ray's operator concept). Thirteen of
those sites fall into mutation classes every catalogue carries, so the vocabulary grew instead of the
class being declared impossible:

  negate-condition      pitest NEGATE_CONDITIONALS / Stryker boolean literals        2 teeth
  negate-comparison     pitest NEGATE_CONDITIONALS / Stryker EqualityOperator        2 teeth
  remove-conditionals   pitest REMOVE_CONDITIONALS                                   1 tooth
  remove-call           Stryker filter/slice/sort removals / pitest VOID_METHOD_CALLS 2 teeth
  replace-call          Stryker MethodExpression / pitest CONSTRUCTOR_CALLS          1 tooth
  replace-initializer   pitest PRIMITIVE_RETURNS, INLINE_CONSTS / literal mutators   5 teeth
  replace-iterable      Stryker ArrayDeclaration / pitest EMPTY_RETURNS              3 teeth

Each computes its own bytes where the mutation determines them (false, true, the negated operator, the
call's receiver, the guard's body) and takes the mutant's `to` only where the new value is a choice -
the rule the earlier operators already followed. Operators are now a table of resolvers over one
context, so adding one is an entry plus its name in the union, and the dispatch's complexity stopped
growing with the vocabulary.

Two widenings, each shown by a real tooth that could not convert:

- A function bound to a name is a member. `const count = (label) => {...}` is the whole of
  board-worker.ts's logic and it is not a function declaration, so its only tooth had no scope to name;
  the selector now accepts a variable whose initializer is an arrow function or a function expression.
- Two identical calls need a holder to tell them apart. dispatchPlan calls board.candidates() on offer
  and again filtered; `in` now names a fragment of the statement the call sits in, as it already named
  the holding object literal for replace-property and the call's own text for replace-argument.

A wrong `within` was shipped again, and this time the anchors pass caught it loudly: the guard is in
runParentCheck and I named instrumentCommit, so it refused with "0 guards in instrumentCommit match the
selector". The first pass's wrong member was the quiet kind - checkDispatch instead of checkCompletion -
where the pass resolved, the mutant was still caught, and only the sweep showed it was caught by the
suite rather than by the case that names the rule. Wrong member is loud when it finds nothing, silent
when it finds the same shape twice: the sweep after a conversion is not optional.

Reading: anchors 110 of 110 resolve over 22 targets; mutants 110 of 110 caught by the named test, all
110 by the case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 176 s; the
anchor cases 30 pass; test:product 1576 pass; verify:static exit 0; lint 0 findings; complexity gate ok.
95 of the 110 teeth are derived now.

The residue is 15 teeth and its classification is measured, not guessed: three fragments inside a string
(two SQL clauses, one SQLite unit - catalogues mutate whole literals, and SQL has its own catalogue);
two at module top level where there is no member to name; two moving an array index, which corrects this
record's own earlier claim that the catalogues cover that class (FirstToLast is a method named first);
two writing two statements on one line; one rewriting a comparison into a range test, where the
catalogue's negation would not realise the violation the name states; four bespoke expression rewrites;
and one that changes two things at once. Three clusters would repay an operator and are recorded as open
(clause-level SQL mutation, a file-scope selector, an index operator); none is a widening, and the
measured cost of not having them is seven teeth.

The catalogue check also supplies the rationale this record was missing: pitest's "Less is more" names
subsumption - a mutant subsumed by others adds runtime, not confidence - which is what the retirement
criterion measures rather than guesses at; cargo-mutants' unviable mutants and its documented refusals
to generate some mutations are the same argument used here for leaving rare shapes alone. And one thing
is deliberately not borrowed: those tools count a mutant killed when any test fails, while this register
counts it only when the case it names is the one that fails, which is why "caught by the suite, not by
the case that names the rule" is a defect it reports and they cannot.

Docs and code in one commit: the record carries the operator table, the two widenings, both wrong-member
episodes, the measured residue and the two catalogue lessons; the ledger says 95 of 110 are derived and
gains this pass's readings; the teeth header says which names a member can have and that the operators
are the catalogues' mutation classes.
… no hand-written anchor left

Asked why the remaining 15 were not converted, the honest answer turned out to be that they could be -
and that three of the reasons recorded for leaving them were wrong. Measured, one at a time, each
conversion keeping its name and its named case:

- the file scope: `within` is optional, and without it the whole file is the scope with the selector
  required to be unique in it (a top-level `CLOCK_GRACE_MS`, a guard in a module-level script);
- `replace-index`: the element access the selector identifies gets the declared index (2 teeth);
- `replace-literal-fragment`: a piece of text written inside a literal becomes the declared fragment,
  with the same holder-statement disambiguator the call selectors have, because one SQL predicate is
  written seven times in one member (3 teeth: two clauses and one SQLite time unit);
- the call selectors accept `new X(...)` as a call site (1 tooth, the round store opened as a
  constructor call);
- `uniqueMember` accepts a function bound to an object property, `claim: (store, parsed) => ...` (1
  tooth, the daemon's handler table);
- `replace-property` writes a shorthand property out, `position` to `position: 0`, because a value
  cannot go where the name was (1 tooth);
- five more converted with operators that already existed: a call replaced by a boolean, a value
  replaced by a stub fact, a wrapper unwrapped to its inner call, an iterable shortened, a comparison
  negated.

Three recorded reasons were wrong, and the sweep is what said so:

- "the catalogue's negation would not realise the violation" (speculation-guesses-several-facts-at-once):
  it does, and the named case fails under `!== 1`;
- "it changes two things at once" (every-task-is-frozen-at-position-zero): only one of the two is
  load-bearing, because `noUnusedParameters` is not set, so the callback keeps the parameter it no
  longer reads;
- "two statements sharing one line" (two teeth): the hand anchors had to span two statements, but the
  mutations are single-statement deletions and both convert to `drop-statement`.

Two mistakes made and caught inside this pass, both worth keeping:

- a hand-typed `expect` reads as a broken mutant. Retyping two case names from memory instead of copying
  them made both teeth report "survived" - the tool filters the suite by the named case, so a name that
  matches no case is indistinguishable from a mutant nothing catches. Loud, not a false pass, and the
  conversion scripts now copy the string rather than retyping it;
- rerunning a conversion script after the vocabulary had grown rewrote two teeth back to their earlier
  form; the anchors pass refused both at once, which is the case for the pass being in the static
  contract seen from the other side - it catches edits to the register, not only drift in the code.

Reading: anchors 110 of 110 resolve over 22 targets; mutants 110 of 110 caught by the named test, all
110 by the case their expect names, 22 of 22 restored byte-identically, exit 0 in 160 s; anchor cases 34
pass; test:product 1580 pass; verify:static exit 0; lint 0 findings; complexity gate ok; format check
clean. The register is 110 teeth over 22 targets covering 21 files, every one a name plus an operator
plus a selector, with 26 operators - each a mutation class the catalogues name or a slot they name.

What is not claimed: a derived tooth is not stronger evidence than the byte anchor it replaced. It is
the same violation caught by the same case; what it buys is that the tooth stays aimed at its rule
across renames and reflows, which is the failure this record exists for. The operators are also not the
whole catalogue, only the classes this register's rules need.

Docs and code in one commit: the record's residue section is replaced by the conversion table, the three
corrected reasons and the two mistakes; the ledger says all 110 teeth are derived and gains this
reading; the teeth header says which names a scope can have, module level included.
The four Plan items landed, and the acceptance criterion is met more strongly than
it asked: the hand-written residue is zero, not small. The record proposed a
doctrine; the doctrine is now the register's only form, so it moves to
docs/decisions/implemented/ and takes the implemented format.

- Problem keeps the motivation and the measurement that removed an option (the
  three-way split does not hold: every caught mutant is externally observable).
- Proposal and Plan become Decision - the durable part is a name, an operator, a
  selector and the case that must fail - and Implementation state: the three
  waves, the widenings the residue forced, the corrections a run caught rather
  than a review, the readings, and what is not claimed.
- Acceptance criteria is gone (the implemented format forbids it); its content is
  where the readings are. The old Risks become Consequences, together with the
  costs of the new form and the two questions left open on purpose.
- Both languages in this commit, and the seven inbound links in
  docs/design/task-unit-semantics-obligations.md follow the move.
- docs/design/ci-cd-and-quality.md: the anchors pass resolves 110 teeth now, not
  the 149 the sentence was written about, and it takes about a second.

Readings at this revision: anchors 110 of 110 resolve over 22 targets; the whole
register 110 of 110 caught, all 110 by the case each expect names, 22 of 22
targets restored byte-identically, exit 0 in 160 s; verify:static exit 0.
@wefio
wefio merged commit 818228a into main Sep 25, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant