docs(vocabulary): the collaboration protocol and its task-unit sub-protocol get names - #71
Merged
Merged
Conversation
…otocol get names Two levels with one meaning each. Protocol-governed collaboration is the umbrella: agents coordinate through a published protocol and the mechanics stay with the program. The Task-Unit Protocol is the sub-protocol this line builds - a task is declared as a unit and the protocol fixes how it is adopted into one run, claimed, delivered and independently judged, and which facts the runtime decides. The parts keep the names they already have: Task Board is the sub-protocol's surface, managed entry is what a governed entry is, run is the frozen object, adopt is the binding transition, and the caller that decides to adopt is called the adopter. The words that came to hand were already taken - fusion (units sharing one session), scheduler (which the task-unit semantics decision deliberately does not enable), governance (the board addressing line), ledger (three of them) and executor (the bounded context adapters) - so each is recorded as considered and rejected rather than reused. No identifier or file name changes, for the reason the check-runner renaming recorded: `ooo` and the existing paths are cited by dated measurement records and frozen run archives. The name lives in the concept map with its aliases, and docs/glossary.yaml is untouched because it owns the repository and process vocabulary, not product concepts.
The definitions were spread over the design, the obligations ledger and several decision records, so reading them meant opening four files at once. The naming record now carries an index section: one line of meaning and a pointer per name, for both levels, the board verbs and the channel rules, the run and its facts, the legal action set and the dispatch loop. It ends with the two words deliberately not used for this protocol - scheduler and OoO - each with the owner of its real meaning, because reusing either is what made the discussion ambiguous in the first place. The section states its own status: an index, not a second specification, and where a row and its owner disagree the owner wins. Same content in the zh pair. Found while checking the links: the concept map's Task Board row points at an anchor that does not exist in memory-graphs.md, and docs:check does not catch it because it validates file existence rather than headings. The index uses the heading that is really there.
The row's anchor named "2.1 Task Board outside the three memory graphs", which is not a heading in memory-graphs.md: the section there is "Shared Task Board (cross-Agent coordination, not a memory graph)". Clicking the old link still opened the file, so it read as working. docs:check validates that a link's file exists rather than its anchor, which is why this survived; the row now names the heading that is really there, in both languages.
…bout speculation The naming record now carries the conceptual history the names only make sense inside: the board becoming a protocol on 2026-08-13 (identity and discovery that wake no LLM, directed delivery, serial admission decided by claim/resolve/expiry), governance and addressing on 09-06, out-of-order execution borrowed from the CPU with its limits written down on 09-09..09-11, the task unit's semantics on 09-13, the paid arms measuring both borrowed halves on 09-18..09-19, and the board absorbing the execution half through 09-20. The OoO step is stated as the CPU correspondence rather than as "parallelism": out-of-order execution with in-order commit, where the plan and its tickets are the reorder buffer, the coordinator is the only retire stage, and the acceptance rules define what counts as commit. Its two limits are recorded with it, because the second one is the point the fusion and speculation arms then measured: the CPU's premise that a wrong guess wastes resources already committed or free does not transfer to agents, where a wrong guess spends tokens and wall time (a worker call measured at 40-56 s). Hence the decision adopts only the class whose wrong guess costs no tokens and refuses the class that hides a wait by guessing it, and the arms record "cost with no gain" at the shape tried: 43k tokens over six units, one false fact wasting 20 332 tokens, 0 of 3 prepared candidates publishable, and 175 ms of verification against 6.2 s of work once the fact holds. Commit-level lineage stays where it already lives (implementation-lineage.md, from its Task Board row) rather than being copied into the record. Also corrected: the bootstrap design still said the speculation decision was proposed and that the no-speculation rule therefore stayed in force until it was implemented. The decision has been implemented since 2026-09-11, so the sentence now states what is actually in force - the host-side free class may be done, hiding a wait by guessing stays forbidden.
Adds a proposed record for the question the naming decision unblocked: the primitives have names and the shared layer can already compute the legal set (an ordered legal set, that set cut to the declared slot budget, and a refusal reason per unit), but no document states whose decision each step is. The proposal splits the work three ways and gives the program exactly one share. The protocol owns what the primitives are, the definition and criteria of legality, the named constraint list, and the two structural rules that cannot be negotiated because they are what "legal" means here: a claim (lease plus attempt fence) is the only arbiter of who holds a unit, and a deliverer never judges its own delivery. The accompanying program computes the ordered legal set from the declared plan and current facts, cuts it to the declared slot budget, and states for each unit why it is legal or not - it chooses nobody, adopts nothing, judges nothing, wakes nobody. Agents own the arrangement: the plan's content, the order they agree on, who takes which unit, who adopts a run, who judges - all of it as board facts, so the arrangement is readable instead of being implied by a program's choice. Constraints are named by the protocol and enabled by the plan, so a preference such as repair-first is neither program policy nor advice; and legality is asked rather than published, because a published handoff occupies a serial slot by the board's own protocol, so fusing the two would make a question cost a slot. The plan is three steps and only the first lands here: this record, then a readable answer through an existing board action, then promoting the preferences that still live as shared planning policy (repair-first first) into constraints a plan declares - which is what makes the "not enabled means absent from the answer" criterion true. The record names that third step as the one place the proposal is not yet true. Deliberately not done here: no product surface change, no driver, no wake, no new tool or channel.
The proposed record now relates to board governance and capability addressing (2026-09-06), and its
Problem says what that record does not cover: the board's correctness line - reviewable finalize,
authentic content, scope isolation, capability addressing, and the deliver/judge/attempt-fencing slice -
governs how an entry is treated once it exists, not who decides what happens next.
The two records are two halves of one protocol, and reading them together is what makes the
distributed-systems half visible: claimTaskBoardEntry is a single atomic compare-and-set ("lease-based
claiming ... on a losing CAS the failure is diagnosed against a fresh read"), lease expiry is enforced
lazily with no background sweeper - that is a suspecting failure detector, not a death certificate -
attempt fencing exists precisely because a lapsed lease cannot be trusted ("without this, work
reassigned to another agent could still be read through a stale artifact"), and the deliverable/verdict
pair binds an acceptance to an artifact identity. Idempotent acknowledgements and the never-recycled
monotonic identity counter are the other two primitives already in place.
No prose in either record names this, so every discussion of the boundary re-derives it.
…istributed concepts behind them Your read was that part of this is a distributed-systems problem, and reading the legality record together with board governance showed the framing has no home: the governance record is an implemented decision with a fixed format, and the 2026-08-13 document is board mechanics. This adds the inventory as a design document, and the naming record now points at it from its history section. It lists fifteen parts with owner and state: entries and addressing, identity and authenticity, ownership and time, termination integrity, attention and wake, scope and visibility, truth versus coordination state, the task unit, legality and admission, ordering and constraints, budget and accounting, cancellation and fencing, recovery and replay, handoff, and evidence. The task-unit protocol is the complete one, and no single part completes a task: a unit's declaration, ownership and time, termination integrity, attention, legality, budget, truth separation and handoff have to hold at once. It then names seven gaps, each with the concept that fits it rather than a new vocabulary: the adopter has no clause (a commit boundary, and orphan reaping for the restart responsibility), legality has no readable answer (admission control with a read-only precondition check), constraints have no declaration (declarative policy, plus deterministic replay for comparability), read guarantees are unwritten (monotonic reads, read-your-writes, causal consistency across one agent's sessions), disagreement has no arbiter beyond veto and racing a claim (optimistic concurrency with conflict detection, or a plain single-writer decision since one arbiter already exists), cross-run resources have no mutual exclusion (a fenced mutex, not only the slot budget and the throwaway-worktree practice), and in-doubt work has no resolution (an in-doubt transaction: the fact is known, the outcome is not). It also records what is deliberately not needed - consensus, leader election, quorum, two-phase commit, exactly-once delivery, distributed transactions, because truth lives in one store written by one daemon - and three analogies that mislead: a lease expiry is a failure detector's suspicion and not a death certificate, which is why attempt fencing must stay; the board is not a message queue, so re-reading is normal and re-doing is the problem; and agents are not replicas, their independent judgement being exactly why a delivery is recorded with a judge instead of a self-report trusted.
The primitives have names and the program's job is proposed, but the frame itself was undefined: which fields exist, how values are encoded, where they live in the store, and how a peer protocol sits beside the Task-Unit Protocol. Researching the finished specifications changed the shape of the problem - every layer but one is already defined by somebody - so this record chooses and names rather than invents. What it borrows, per layer: CloudEvents 1.0 for the envelope (only id/source/specversion/type required, context attributes inspected without deserializing the event data, lowercase names of at most 20 characters, data reserved); A2A 1.0.0 for the task model (terminal states, contextId, cursor paging chosen explicitly over offsets, historyLength, includeArtifacts false by default); NLIP/ECMA-430 for the payload discriminator and the progressive shape (messagetype/format/subformat/content/submessages/label, the first submessage inlined because most messages carry none, structured with subformat uri carrying a URI instead of bytes); MCP 2025-06-18 for LLM-facing results (content blocks with annotations, structured content beside text, isError so a failure is data); RFC 8941 for a one-line text type system; FIPA-ACL for the act and conversation fields; RFC 6709 for the unknown-field choice; ECMA-434 for the security layer. The one layer nobody defines is ownership, lease, fencing and legality, because every agent protocol assumes a task belongs to one agent. Decisions it records: the frame is a state record and not a message, so FIPA's performative is a rendering of existing columns rather than a column; the header stays in columns because the claim is one atomic compare-and-set UPDATE and a header inside JSON would turn that into read-modify-write; the payload is one opaque document, with the invariant that every payload can be NULL and T0/T1 still render while claim, wake, expiry and compact read still work (verified on the store's SQLite 3.53.3); a payload field a protocol wants to query is indexed by a generated column plus an expression index rather than a new board column, which is what keeps the schema from growing with the number of protocols; the protocol is a value with a version, a peer protocol supplies declaration validation, legality, projection and acceptance, and it plugs into the existing DispatchBoard and RoundQueryPort rather than into the board; unknown protocol means readable and not actionable, and unknown fields are decided by the protocol's own version rule, stated rather than assumed. Data format, in three separate questions: JSON on the existing RPC wire (protobuf-as-normative-source is not worth it with one binding); flat key/value for T0/T1 and natural language for negotiation, because format restrictions measurably degrade reasoning and escaping is a failure source, with tool schemas shaped for strict mode (additionalProperties false, every property required, absence as nullable); typed columns plus one JSON payload document in the store. TOON and CBOR are deliberately not adopted: TOON's saving applies to flat uniform data pushed into context, which is what pointer-first avoids, and CBOR waits for a transport with a bandwidth constraint. Security is the layer that is empty. ECMA-434 is mandatory for NLIP conformance and its companion guidelines score fifteen threats with prompt injection highest; two rules are adopted as text now (an in-band session token is a session credential and must not reach the model, and a caller's token is never forwarded), and a threat list of our own is owed. The scope ceiling is stated so the change cannot grow: one group of nullable columns, one boundary adapter, one rendering rule, one test. Anything needing a new table, a new tool or a new channel is outside this decision.
…o retractions Four review points, all accepted, plus the code checks behind them. 1. The field mapping was missing, and it changes what the invariant means. The record now carries an existing-to-proposed table: content stays the single source of the entry body (readTaskBoardPreviews projects taskBoardPreview(entry.content) - whitespace collapsed, 200 characters, a memory=<id> body returned verbatim), so "payload NULL still renders T1" is a property of the board's read path and not something the storage format produces by itself; the frozen declaration was never a body, it lives in task_run_manifest and task_run_tasks reached by run_id and task_id, with task_run_facts' own payload column as the precedent; and the proposed payload carries nothing the Task-Unit Protocol needs, so it stays NULL for Task-Unit entries. A rule follows from it: a payload may not restate a field that already has a column or a table, which is how the second task body would appear. 2. "The insertion points already exist" was not true. RoundQueryPort carries two typed reads, cancelled() and accepted(); the dispatch loop fails a claim whose ticket has no patch, so patch work is the only work shape today and a peer protocol must bring its own. The record now says what is reusable material and what is still a plan step, and adds that step: widen the query port or add one beside DispatchBoard. 3. Two storage arguments are withdrawn. Atomicity does not require columns: measured on the store's SQLite 3.53.3, a conditional UPDATE over a JSON payload (WHERE json_extract(payload,'$.state')='open') succeeded once and changed zero rows on the second claim, identical to the column form - so the reason is clarity, direct indexing, CHECK constraints and PRAGMA table_info migration guards. And a generated column is a column, an index is a schema object: peer protocols are not schema-free. What is bought is narrower and still worth it - the added schema is purely additive, belongs to the protocol rather than the board, and no field is stored twice. 4. Three borrowings were applied too widely. source_session_id is provenance, not a credential, so the security rule is now stated as a constraint on what may be added rather than a defence of what exists. The refusal rule now keeps MCP's split: a domain refusal becomes data, a request-level error stays an error - which the dispatch loop already does by separating refused from failure. And the CloudEvents naming constraint is stated as borrowing (lowercase letters and digits; twenty characters recommended), with protocol_version's underscore recorded as a deviation rather than as compliance.
…eded The legality record and the frame record each needed the same sentence and neither stated it: the legality record says what the program does and what it refuses to decide, the frame record says what the core knows and what it never parses. Those are one division seen from the side of the decision and from the side of the data. Because nothing named it, it had to be re-argued every time a field or a rule appeared, and the cost was already visible: "work means a patch" had reached six places, one of them a legality rule written in the kernel's own vocabulary (refuseWidening decides permission closure by reading parent.patch.editable). This record takes Hydra's principle - a kernel provides mechanisms and refuses policy - and states it as two contracts. Mechanism: one claim (compare-and-set, lease, attempt fence), an opaque declaration with a digest, an opaque artifact with a digest, a verdict from someone else, the lifecycle, and the legal set (which units are legal now, in order, cut to the declared slot budget, with a reason per unit). Policy: what a unit's inputs are, the artifact class, how the artifact is judged, how the work runs, its concurrency preconditions, the wording and enablement of legality rules, and who is chosen, who adopts, who judges. It also supplies the two things the slogan lacked. A decision procedure - for each field, type and code path, ask mechanism or policy, with a rule's check counting as mechanism while its name, wording and enablement count as policy. And a mechanically checkable rule: no policy word may appear in the mechanism layer's code, type declarations or schema paths, from a maintained list (patch, editable, files, instruction, checks, repair-first), with the path carried because a word can be policy in one place and mechanism in another. The consequence for work shapes is stated: a work shape is not a core concept, it is the policy layer's name for the declaration-and-artifact pair the core carries opaquely. The six leak sites are classified in a table, and the legality rule found today is the refinement the record adds - its check belongs to the program while permission closure's name and enablement belong to a declaration, so the fix is not to move the check out of the program but to stop hard-wiring the rule in the kernel's vocabulary. That is the same shape as plan step 3 of the legality record, where repair-first becomes a declared constraint instead of shared planning policy; two independent fixes taking the same shape is the evidence that the classification is right. Three rows are deliberately left undecided so that they are not settled by preference: task_run_tasks.effect is mechanism only if the mechanism must enforce write-set disjointness, the granularity of input and dependencies is a separate question from whether dependencies are mechanism, and operation may be policy. Alternatives rejected: keeping the slogan and deciding case by case, putting the principle in either record (it is wider than both), a plugin registry above the core with no second policy to justify it, and deciding the undecided rows now. Risks named: gutting the core (so the practical form is one default policy the core may carry but must not require), an ossifying word list, an undecided row becoming permanent, a policy word legitimately remaining for a release, and grep being gameable by synonyms. The legality record and the frame record now point here, so the division has one home.
… earns a place Records the model that came out of the last exchange as a draft rather than as a record, because it was reached by stacking analogies and nothing about it has been measured. The model: a drawing pipeline has four stages and so does ours - producer (a work-shape adapter), surface (the board entry: payload plus a discriminator and a digest, delivery as commit, resolve and expiry as release, wake as the frame callback), compositor (the board and the program: claim, lease, fence, serial and wake, expiry and reaping, the legal set), and display (the presentation end: tiers, flat lines, compact read, the tool descriptions an agent reads). From it come three ownership classes: content and semantics belong to the owner and are semantically opaque to the middle (a compositor reads pixels but cannot read what a button means, so it asks the owner; our board reads declared structure and never payload semantics); description is the owner's obligation (geometry, level, shape, dirty regions, terminal state, identity, a snapshot able to stand in for the owner - ours is the discriminator, digest, dependencies, scope, and the protocol's projection); and arbitration is only the middle's, because stacking, occlusion, visibility, focus, capture, timing and reaping are global properties. The boundary test that replaces a slogan: no reasoning needed means the object stays opaque; structure needed means the owner supplies a description and the middle reads the description but not the semantics; a global judgement means only the middle can make it, which is why the description must be cheap and fresh and why staleness needs a meaning. The document also carries its own case against itself: the analogies did the reasoning and were each partly overruled by a fuller view, one of the four stages is a name with no seam or owner or test, the only measurable claim in this arc (that the seam was narrow) was falsified by a grep within a minute, and nothing got smaller - no code changed, no concept removed, three prose records added. So it states an eligibility rule - an analogy earns a place in a record only when a check can falsify it - and records its predictions before measuring: (a) policy words in the middle, predicted at fifteen or more hits concentrated in the board and the legality module; (b) the two ends, predicted to be concentrated and separable at the producer end and scattered with no seam at the presentation end; (c) a second shape through the dispatch loop, predicted to change five or more files. The reduction test is named in advance too: less code (the middle must lose a type and a freeze method, or the model failed), fewer concepts (the three classes must absorb the two lists they replace rather than sit beside them), fewer change points (measured by how many places a new shape touches), and better maintainability (a new shape adds files under one owner and changes no test in the middle).
…he counting exposed Both checks were run against the prediction recorded before them. (a) Policy words in the middle: predicted fifteen or more hits, measured 195 raw - patch 106, files 39, checks 23, editable 14, instruction 12, repair-first 1 - concentrated exactly where predicted, in src/integration/ooo-board.ts (53 patch hits), src/integration/task-semantics.ts (27 patch, 9 editable) and src/core/store/base.ts (13 patch). Two corrections forced by the measurement. The raw count overstates the case, which is the caveat the check was written with: files in src/core/store/writes.ts and retrieval.ts is a mechanism word (a path inside a store) and checks in src/integration/ooo-candidate.ts names the check runner, also mechanism, so the word list drops files and checks and keeps patch, editable and instruction. What is left is still about 120 hits, so "the middle is policy-free" is false as a description of today, an order of magnitude past the prediction, and the model's support is directional: the leak is real, large and concentrated in three files. (b) The two ends: the producer end was predicted concentrated and separable and measured four to six functions over three files (preparePatchWork, patchPrompt, patchCandidate, patchSubmission, snapshotText, runTestFile); the presentation end was predicted scattered with no seam and measured six sites (the preview text, the entry and preview types, the wire shape, the service, the agent-facing renderer, and the generated tool descriptions). Both held. The measurement also corrected this document: the display row claimed T0-T3 tiers, and no such thing exists for a board entry - tieredDisclosure and tier belong to memory retrieval, and a board entry has exactly one presentation, the 200-character preview plus raw fields. That row now says what exists, so the second borrowed vocabulary is recorded rather than left in the table. (c) A second shape is not run yet, and the reduction test is unchanged: nothing has got smaller, the middle's three files hold the available reduction, and whether it pays is what (c) measures. The counting earned its keep twice - it produced the word-and-path rule the earlier prose could not state, and it caught this document borrowing the memory side's vocabulary as if it were the board's.
The model in mechanism-in-the-middle.md is retired and replaced: a board entry is a medium with a write face and a read face, the compositor is brought by a protocol rather than by the board, and the roles are roles rather than files - src/core/store/base.ts is the medium, the compositor and the reader's rule at once, and src/cli/service.ts is both the wire face and a formatter. The rule that decides where a seam goes needs no taste, only a count: a boundary is built when it has a second implementation, and only declared with one. The read face has three, so it lands here. src/core/board-entry-view.ts owns the rule, the store's compact read and the host adapter's broadcast both call it, and the store's private copy and the adapter's bare slice are gone. One of the three was not merely duplicated but wrong - the adapter's 140-character slice did not collapse whitespace, so a multi-line body reached a broadcast as a multi-line broadcast. The write face's type and the middle's seam are declared and not built, each with a named trigger. Check (c) measured why: making the loop shape-agnostic is one file with fifteen insertions and eighteen deletions, and tsc then reports zero errors while twenty-five tests fail, every one of them because the claim admits no work. The coupling is invisible to the compiler and fatal at runtime, so the seam is prepaid cost until an adopter declares a non-patch unit. Docs and code land together, because the read face's home is what the two-face model implements.
The port now has a third implementation, and it is the product's. src/integration/ooo-runner.ts projects the run's frozen table and its board entries into the facts the shared rules read, and carries no legality rule of its own: the plan is compiled by compileTaskUnits, the legal set is dispatchTasks + selectableTasks, cancellations are read through the coordinator's own reader, and the verdict belongs to the caller's acceptance. A rule that turns out to be needed here belongs in task-semantics.ts instead - needing one is how this module would show that the abstraction leaked. Two things the projection cannot read from the store stay the caller's, and both are existing divisions rather than new ones: a task's declared file contents (the store freezes paths, and preparing the workspace is the patch path's caller's job) and an acceptance whose identity is independent of the deliverer (the store already refuses a deliverer that judges its own delivery, which the second test pins). One contract is easy to get wrong, so the projection writes it down: the shared acceptance rule binds a verdict to the artifact value the run carries - it compares the verdict's digest against the artifact - and not to the store's own deliverable hash. Reporting the hash there makes every accepted unit read as unaccepted, and the symptom is silent: the next unit is never released and nothing reports an error. The port's "put a result" is the product's delivery on the entry the unit was claimed on, so a unit's work is recorded once rather than once on that entry and again on a result entry nobody would read; and the unit is resolved from the claim rather than from the body, because a body's shape belongs to its owner while a claim is the board's own fact. Verification: npm run check clean; test:product 1527/1527 pass (2 new - a two-unit plan driven through the product's store with an independent judge, and the self-judging refusal); build clean; lint clean; prettier clean; docs:check 288 files, 0 errors, 0 warnings; agent:context:check valid once the route claims the new file.
CI failed "Product tests and coverage" on one case of 1527 while the same suite passed 30 times in a row locally: "a value that expired a moment ago is still current" stamped an expiry half a grace (25 ms) in the past in one statement and read it in another, so it asserted that the read happens within the half that is left. A sweep between the stamp and the read says where it flips - 0 ms and 10 ms answered "current", 20 ms and beyond answered "not active" - and a loaded runner has no trouble spending 20 ms reopening a store. Sharing a clock source is not sharing a clock read. The case is now a relation: the store's own notExpired predicate is evaluated against a boundary stamped in the same statement, so both share one 'now' and the expectation is about the window rather than about the machine. It also pins the boundary instead of a point inside it - at exactly one grace an expiry is already excluded. 8000 evaluated assertions under deliberate load, no flip. A row-level form of that case cannot exist, because a row's boundary is stamped by one statement and compared by another, so the store level keeps the cases that do not depend on a duration: a future stamp inside the grace, a minute either way, and 400 write-then-read rounds. The mutant that names this case as its catcher is pointed at the new name and still fails it: 4 of 4 caught, restored byte-identically. Postmortem 0004's guardrail claim and its lessons are corrected with the recurrence: its "deterministic" test was deterministic on three of its four boundary cases, and the fix for a timing dependency was itself one.
Three were in the store board's own test, which `npm run check` cannot see because tests are outside its project. One of them mattered: the self-judging tooth read a digest off the ticket, where the port's PatchWork is the pre-freeze declaration and carries none - so it delivered an artifact with no digest at all, and the store took it, because a delivery is a digest and a reference while the wire shape is the host's check rather than the board's. That test now freezes and builds its artifact the way the loop does, through one helper the worker shares instead of a second copy of the artifact shape. Two more were the store's private single-entry reader reached from tests, in this file and in tests/core/task-board.test.ts where they had been reported for a while. Both now use getTaskBoardEntryById, which is the same reader plus the channel guard the store asks for: the test says which channel it reads, and the private method stays private. The five in evals/ooo-execution were declared shapes that did not match what the code uses. The two evidence drivers refused a missing flag in a loop over an object, which narrows nothing, so one of them asserted non-null at its use site and the other did not compile at all; both now call a flag() helper that returns the value it checked. live-continuation.ts declared its per-attempt result as the metrics shape alone, which made `artifact` - the thing a continuation continues from - a property the file did not have; its type now follows the call that produces it. Verified: lsp_diagnostics reports 0 across the five files; check, lint, format:check, docs:check and agent:context:check clean; test:product 1527/1527 pass.
A verdict binds to a digest rather than to bytes, so the digest's convention is a rule a
stored record depends on - and it had four module-private homes. Two of them were the
identical two lines (`createHash("sha256").update(JSON.stringify(value)).digest("hex")`)
in `task-semantics.ts` and `ooo-board.ts`, and two callers had folded their own shortening
into their copy, so a twelve-character report identity and a sixteen-character branch
identity read as if the shortening were part of the rule. `src/integration/work-identity.ts`
now owns it: `workDigest(bytes)` for text or raw bytes, `workDigestOf(value)` over JSON
text, sha256 in hex at full length.
Canonicalisation stays the caller's, and the module says why: JSON key order is whatever
the caller built, so a caller that needs one identity per shape rather than one per
serialisation freezes the shape first - which is exactly what `preparePatchWork` does when
it sorts the files, serialises once, and returns a digest rather than a serialisation. The
shortening two callers want stays theirs, and it now says so by slicing.
Sites that differ on purpose keep their own convention, and the module names them so a
later sweep cannot unify three different rules: a search index's base64url content hash, a
session identifier, and a protocol-visible `sha256:` prefixed identity.
Two named mutants make the convention checkable - another encoding, and a JSON variant
that does not digest JSON - both caught by the named case, restored byte-identically.
The write face's type did not land, and this commit records the measurement that says so
rather than building it: the artifact body inside a result already has a home
(`artifactEnvelope`/`artifactFromText` in `ooo-session-mechanism.ts`), while the entry body
wrapping it has one writer and two readers that are two *different* rules - the probe parses
the loop's envelope and binds a verdict to the ticket's digest, the product board reads the
artifact of a body that may be the artifact itself. Two rules with one implementation each
are not a seam.
Verified: check, lint, format:check, docs:check and agent:context:check clean (the new file
is claimed by the ooo-execution route); test:product 1529/1529 pass twice; the work-identity
mutant target 2 of 2 caught.
Asked whether the design is complete, the authority to answer with is
docs/design/task-unit-semantics-obligations.md, and it contradicted itself in two
places - one in each direction.
Its counts paragraph said "the only arm still unrun is the E arm itself", while the
E arm's own section, the row table and the closing paragraph all say it ran twice on
2026-09-19 with no gain claimed. A reader trusting the opening paragraph would think
the measurement phase was one run short of closed; a reader trusting the closing one
would think it was closed. The sentence now points at the two records that settle it.
Its verification section recorded the LSP diagnostics as "recorded rather than
repaired", naming one error each in board-deliver.ts and board-judge.ts plus a
{ pid: 0 } fallback in an evidence-driver test. Those are repaired now: the drivers'
flag narrowing and live-continuation.ts's declared result shape landed in e31fe77,
and this commit fixes the test's fallback, which needed the startedAt that ServerState
requires (a pid of zero and an empty start time still says "no server", which is what
the code means).
What the LSP still reports is stated as the number it is: 31 diagnostics in seven
evals/ files this arc does not own - benchmarks/run.ts 7, longmemeval/run.ts 13,
controller/run.ts 3, natural-maintenance/audit.ts 3, hierarchy-scale/run.ts 2,
longmemeval/score.ts 2, omnimemeval/bridge.ts 1. The sentence claims zero only for the
files this arc touched, because that is what was measured; the tests/ directory as a
whole timed out rather than answering, and an unmeasured claim is not written down.
Verified: docs:check 288 files, 0 errors, 0 warnings; format:check clean; lint clean;
tests/integration/ooo-evidence-drivers.test.ts 8 pass, 0 fail.
`selection` now returns its own reasoning beside the legal set: every gate it applies names itself, and `unitLegality` reports the ordered legal set with a cause list per unit. `deriveStatus` carries the same causes, so the status query and dependency release keep sharing one predicate and now share its explanation. The causes are computed inside `selection`, where the gates are, rather than beside it: a second copy of a gate's condition is how an answer starts disagreeing with the decision it explains. A plan-level gate is attached to the units it holds back, so an empty legal set is never a bare empty set. Two cases pin it: one names each gate in turn, one walks every combination of the flags a plan's facts can carry (64 shapes, two budgets) and refuses to accept a silent refusal. The mutant target grew two teeth for the new rule - `a-refused-unit-is-silent` and `a-stale-input-is-not-named` - and the two selection gates whose lines moved were re-pointed, including the budget mutant's named case, which now names the case that actually distinguishes it. 24 of 24 mutants are caught and both files were restored byte-identically. This is the shared half of the legality proposal's step 2. The board action that lets a caller ask the question through a surface it already has is still to come, so nothing here claims the answer is readable yet from outside the process.
The port gains the query the legality proposal's step 2 asks for: `legality()` returns the ordered legal set, the room the run has left, and a named cause per unit the rules do not have on offer. It calls the same `selection`, so a caller cannot read a unit as legal here while the rule refuses it there, and `candidates()` - the port's other read - is now that answer's `legal` on the store's board, which is what keeps the two from drifting apart. All three boards answer it: the store's, the probe board's (whose projection moved into one private method so `ordered()` and the answer cannot disagree), and the dispatch test's stub, which hands its own fields to the shared reading rather than keeping a second set of rules. The asker is the process that owns the run's workspace: a plan compiles from the caller's files and the daemon holds only their frozen paths, so the answer cannot come from the store alone. The product's case asks twice on an unchanged store, asserts the two answers and the recorded facts are identical, that the answer names nobody (the store's own refusal names the holder; this read does not), that a claim moves a unit to the named cause while the claim itself stays the store's, and that the answer follows the facts it reads rather than a cached plan. Docs land with it: the two-faces document's read-face bullets, the proposal's step 2 marked landed with step 3 named as the place it is still not true, and the umbrella's gap 2 narrowed to the asker that has no workspace. `ooo-board`'s and `ooo-execution`'s mutants were re-run - 23 of 23 and 24 of 24 caught, both files restored byte-identically - and the product suite is 1532 of 1532.
Repair-first was the shared planner's own default, so the meaning of a plan depended on which planner read it: the online move that prefers continuing a session lived in `nextSessionMove`, and the run's recorded `policy` was a name the probe board compared rather than a switch anything read. The legality record's third step is that a constraint the protocol names and a plan enables, and it is what makes "a constraint that is not enabled has no effect" true rather than intended. The protocol's side is a closed list. `PLAN_CONSTRAINTS` holds the names that exist - `repair-first` is the only one so far - and `enabledConstraints` refuses a name outside it by name, because reading an unknown name as "nothing was asked for" is exactly the decoration the rule forbids. A plan that enables nothing is a setting rather than a fallback: it runs one unit per session, which is the baseline the arms' control cells measure. The plan's side is a declaration. `SessionPlan` gained `constraints`, the move is computed under it, and `optimisticPlan` copies it, so the offline projection prices the same plan the online half decides with. Every caller that wants a fused run now declares it: the arm driver's spec carries the declared set, and `validateSpecFile` refuses a spec that declares a bound above one without enabling the constraint - such a run would be the control arm while its file said fusion, which is the one confusion that driver must not create. The trial spec generator writes it, and three test fixtures carry it. What is demonstrable is a different move at a session boundary under an enabled constraint, not a different legal set: repair-first orders a session, it does not gate membership. The record's criterion said legal sets and was corrected rather than quietly restated, and the record moves from proposed to implemented in the same commit as the code, with the umbrella's rows 9 and 10 updated. Evidence: two new mutants on `src/integration/ooo-fusion-plan.ts` (2 of 2 caught by the named cases: the continuation not declared, an unknown constraint ignored), the four re-run targets 37 of 37, `npm run test:product` 1534 of 1534, `npm run verify:static` exit 0, `npm run docs:check` 288 files with 0 errors.
The sweep refuses to claim a check when a mutant's marker no longer matches the source, and two markers had drifted away from the code they were written against: `the-status-query-ignores-the-declared-budget` still named a `startableTasks(dispatchTasks(units, facts), slots)` call that `deriveStatus` now splits, and `a-cancelled-run-still-accepts-writes` still named a two-line refusal that `managedWriteRefusal` now assembles from a subject. Both were reported as "marker not found, refusing to claim a check" rather than as caught, which is the honest failure and also a hole: nothing in the build notices a tooth that stopped touching its rule. Re-pointed, not rewritten: the budget mutant now hardcodes one slot in the call the module makes, and the cancellation mutant inverts the guard inside `managedWriteRefusal`. Measured after the fix: 27 of 27 caught by the named cases, 2 of 2 files restored byte-identically.
89% of the mutation ledger's teeth (133 of 149) are a site plus one of a handful of operators, and 52% of them (78) are a bare byte fragment matched anywhere in the file. That is the shape that dies first: both stale teeth repaired in dbf10ac were of it. This lands the first two slices of the plan to derive the mutant instead of storing it. The vocabulary (tools/mutation-anchor.ts, new): six operators over a member-scoped selector - condition-never, condition-holds, neutralize-term, replace-argument, replace-property, drop-statement. Every selector resolves to exactly one site or refuses, and the refusals are the interesting half: a member named twice, a call matched twice, a fragment that fits two decision positions (the innermost one wins; two disjoint ones are refused), an operator that replaces a whole condition given a fragment that names only part of it (that would silently widen the mutant into "every reason this rule has"), a statement that does not own its line. A condition is any expression in a decision position - a test, a returned value, an arrow's body, a value bound to a name - because `return a && b` and `array.filter((x) => x.y)` are the same rule written without an `if`. The resolver holds no state, so it moved out of the sweep script: it is 16 tests over source strings, with no filesystem. The tool (tools/mutation-teeth.ts) keeps `from`/`to` and `ast` forms working and gains `--anchors-only`: every selected anchor resolved, nothing written, no suite run, 0.6s for all 149 - the pass that answers "is the tooth still aimed at something", which nothing answered before. The pilot (src/integration/ooo-execution.ts, 24 teeth over two target entries): 19 converted, 5 staying hand-written and named in the record, each still a single expression or value rather than a statement or a message. Sweep after conversion: 24 of 24 caught, restored byte-identically 2 of 2. 23 are caught by the case their `expect` names; the stale name fusion-continues-from-an-unverified-answer was already that way and this conversion neither caused nor fixed it. The demonstration, and the prediction that was wrong: a rename plus `spent || severalWaits` lifted into `const closed` retired two sites - the hand anchor, and the derived a-live-claim-does-not-block-selection, because a condition bound to a name was not a decision position yet. The prediction written before running it had been "one failure". The selector was widened to include a variable's initializer, and the corrected prediction held exactly: anchors: 148 of 149 resolve, the one failure the hand anchor, file restored byte-identically. The limitation is real and recorded rather than hidden: a derived selector is scoped to a decision position, so a rule rewritten into a different shape can still retire its tooth - but now it does so visibly, naming the tooth and the fragment. Not in this commit: the anchors-only pass as an atomic check of the ci-and-tests route, and the retirement pass over the ~18-20 teeth whose rule already has a relational or enumerative check. The record (docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md, with its zh-CN pair) marks both, and holds the measurement behind the operator list. Readings at this revision: npm run mutation:teeth -- --anchors-only -> 149 of 149 resolve over 23 targets, exit 0; the target sweep -> 24 of 24 caught; tests/tools/mutation-anchor.test.ts -> 16 pass, 0 fail; npm run verify:static -> exit 0.
A tooth that stopped matching its rule is invisible today: no CI job runs mutation:teeth, and in a sweep it is reported as "not applicable" and excluded from the denominator - which is how the two dead teeth repaired in dbf10ac stayed dead without anyone noticing. npm run mutation:anchors runs tools/mutation-teeth.ts --anchors-only: every selected mutant's site is resolved and nothing is written, no suite runs, and a site that cannot be resolved fails the run. It is now one of verify:static's checks and one of the ci-and-tests route's atomic checks, in the same order the route-contract test enforces (the route's list must equal the chain, test:product included). Measured on this revision: verify:static exit 0 in 41 s; a full npm run agent:verify exit 0 in 112 s inside its 150-second budget, with the anchors pass itself 0.98 s, all 149 anchors resolving over 23 targets. The pass proves a site still resolves, not that the mutant is still caught - a stale expect still passes it - so the full sweep stays the standing rule before a push and stays out of the gate. docs/design/ci-cd-and-quality.md gains the check and the reason it is in the gate; the record (docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md, with its zh-CN pair) marks its third slice landed.
The retirement pass, first group. The criterion, applied one tooth at a time: a tooth may go when a
named case's assertion fails for exactly the violation the tooth introduces, and that case's wording
states the rule - so the ledger row names the check instead of the mutant. Judged by reading the
assertion, then by re-running the target's sweep with the tooth gone.
- the-board-read-path-stops-calling-the-predicate, the-board-decides-acceptance-on-its-own (row B5):
tests/integration/ooo-acceptance-one-predicate.test.ts counts the call sites itself - one definition of
the predicate in src/, one `acceptedFact({` in each reader, zero `verdict ===` comparisons in
ooo-board.ts - so a rename that stops the call and a second decision inside it each fail a count. A
behavioural differential would have caught neither; a count does. This pair is what made the criterion
worth writing down.
- a-refused-unit-is-silent: the case enumerates the 32 flag combinations a plan's facts can carry and
asserts every refused unit has a reason.
- the-merge-enumerates-one-order (row A5): the case asserts the multinomial count (10), which a merge
returning one order fails.
- a-second-entry-rebinds-the-task (row D12): the case asserts a retry is the same binding and a second
entry is refused by name.
Docs and code in one commit: the three rows now name the check that carries the rule and say when the
tooth was retired (A5, B5, D12 in the obligations ledger, whose reading list gains the dated readings
below), and the record's fourth plan item is marked landed with the criterion and the finding.
Finding carried forward: a-refused-unit-is-silent is named by no ledger row, and the case that catches it
is unowned too. Retiring it removed an orphan rather than a row's pin.
Readings at this revision: mutation:anchors -> anchors: 144 of 144 resolve, over 23 targets, exit 0 (the
register is 144 teeth, not 149); the four targets the pass touched -> mutants: 70 of 70 caught by the
named test, restored byte-identically 5 of 5; test:product -> 1562 pass, 0 fail, exit 0; verify:static ->
exit 0.
Second group of the retirement pass, same criterion as the first: a tooth may go when a named case's assertion fails for exactly the violation the tooth introduces, and that case's wording states the rule. - The five rules of a declared budget - the licence is the budget's part and not its head, a budget above one must name each handoff's target, every startable task gets a handoff, a startable handoff is not retired by a republish, and a multi-slot run directs its handoffs - are all stated by one case, evals/ooo-execution/board-slots.test.ts's "a declared budget holds two claims at once, and the store is why each handoff is directed". Its assertions name each rule: the refusal message, one handoff per startable task, serialState null for each, the same ids across a republish, and the non-head claimed first. So the-licence-is-the-head-whatever-the-budget, a-second-slot-is-declared-without-a-target, only-the-heads-handoff-is-published, a-startable-handoff-is-retired-as-unselected and a-multi-slot-handoff-is-published-un-directed are retired. The F2b-slot row keeps seven teeth, and the rules themselves keep a witness that says them. - claim-is-not-scoped-to-its-run (row B1) is the tooth the namespace experiment first watched survive: tests/integration/ooo-run-namespace.test.ts was strengthened until it caught it, and that record names the assertion that did it (a raw-row read of the other run's owner). A case whose history is exactly this tooth is a better witness than the tooth's own name. Docs and code in one commit: rows B1 and F2b-slot now name the checks that carry their rules and say when each tooth was retired, the obligations ledger gains these dated readings, and the record's fourth plan item is marked with this group, the two candidates judged not retirable so far, and the two stale expects still outstanding. Readings at this revision: mutation:anchors -> anchors: 138 of 138 resolve, over 23 targets, exit 0; --targets=src/integration/ooo-board.ts -> mutants: 15 of 15 caught by the named test, restored byte-identically; test:product -> 1562 pass, 0 fail, exit 0; verify:static -> exit 0. One reading to carry, not a code failure: the first test:product run of this pass failed tests/core/graph-cycles.test.ts's first case on EPERM from rmSync of its %TEMP% scratch directory, and the file then passed 8 of 8 alone and the suite 1562 of 1562. That is the Windows temp-tree removal class the ledger already records as undiagnosed, and the ledger now says so with this instance.
Two retirements, same criterion as the first two groups: a-retried-run-fact-is-appended-twice and
a-frozen-task-is-replaced-by-a-different-definition (row D11) go, because that row's prose already rested
on the two cases for those rules and the cases state them outright ("appending the same fact twice records
it once and keeps the first sequence", "freezing a task twice is a no-op, and a different definition for it
is refused").
The group's larger finding is what the first whole-register sweep said: 136 of 136 caught, but only 134
caught by the case each expect named. Four links were wrong, each in its own way, and all four are
repaired:
- fusion-continues-from-an-unverified-answer: its named case refused the pair for a second reason as well
(the successor declared the predecessor as a dependency, so the dependency rule could refuse it alone),
and a case that passes for another reason pins nothing. The pair now declares no dependency, which
leaves the predecessor's verdict as the only condition that can refuse it, and the tooth's own expect
case now fails under the mutant.
- next-is-not-the-head-of-the-ordered-candidates: its expect named a case in another target. The board's
budget case now asserts next() is the head of the ordered set - the rule in its own words - and that is
the name the tooth carries.
- the-caller-rebuilds-the-shared-floor: its expect was a paraphrase that named no case at all. Re-pointed
to "a declining route's narrow run verifies on its own tests and nothing else", which fails under it.
- the-grace-is-zero: its named case derived its fixture from the constant under test
(half = CLOCK_GRACE_MS / 2000), so zeroing the grace moved the stamp onto now and the case passed while
the bug was live. The fixture is a literal now and the case fails under the mutant again.
Re-run after the repairs: mutants: 136 of 136 caught by the named test, all 136 by the case their expect
names, 23 of 23 targets restored byte-identically, exit 0 in 217 s. The class is the one this arc opened
with - a link that goes stale while the artifact still looks right - so the reading worth keeping is not
the count of teeth but the count of teeth whose named case is the one that fails.
One behaviour recorded rather than changed: a named case that does not finish inside its bound counts as
caught, with the reason printed (the-pass-asks-a-unit-it-already-failed-again, 30 s). The tool's own
comment says a run that never ends proves nothing; the code counts it as caught and says why. Resolving
that tension changes what the ledger's proven means, so it is left for a decision rather than a fix.
Docs and code in one commit: row D11 names the cases for the two retired rules and says when the teeth
went; the obligations ledger gains the whole-register reading with all four repairs named; the record's
third group notes the same, the two candidates still judged not retirable, and the timeout tension.
Readings at this revision: whole register -> 136 of 136 caught, 136 by name, 23 of 23 restored, 217 s;
mutation:anchors -> anchors: 136 of 136 resolve; test:product -> 1562 pass, 0 fail, exit 0; verify:static
-> exit 0; lint -> 0 findings.
The 27 teeth whose anchor is a place where code must not appear are the ones a derived selector cannot express, so each of them was either going to stay hand-written forever or give way to a check that says the same thing. 26 give way, and one is kept with its reason. The method, applied per tooth rather than in bulk: the register-wide sweep first says which case fails under the tooth; that case's assertion is then read to see whether it fails for exactly this violation and states the rule in its own words. What qualified: - a count over the store's own source - "the store runs its transaction boundary in exactly one place" counts one BEGIN, one COMMIT and one ROLLBACK, all three inside writeTransaction (a-method-opens-its- own-transaction); - an enumeration of a set that must not change - the accepted prefix, the fallbacks by source id, the outcome of a scoring order, the orders the ordered mode allows, a run's tasks after a refused freeze (reopening-one-task-clears-every-acceptance, a-failing-source-takes-the-decision-with-it, the-ordering-adds-a-task-to-the-set, ordered-mode-becomes-any-topological-order, the-plan-freezes-one-task-per-transaction); - a refusal by name plus a raw read back - the second reader of a ready task is refused and the owner is then read out of the table, the status of an unknown run stays unknown, a driver without a daemon refuses and no driver imports the store (a-live-claim-can-be-taken-by-another-agent, status-registers-the-run-it- cannot-find, a-driver-falls-back-to-opening-the-store); - the file's own hash and a file that must not be created (the-read-only-factory-opens-a-writable-handle). One is kept: every-task-is-frozen-at-position-zero. Its named case is a register/freeze/adopt/read-back round-trip, weaker than "the position comes from the array order, not from the request" - the same reason a-binding-does-not-record-its-channel stays, as recorded when the earlier groups were judged. Seven of the 26 turned out to be orphans: no row in the obligations ledger names round-publication-opens-its-own-transaction, round-never-releases-its-pin, ordered-mode-becomes-any-topological-order, the-loop-awaits-each-unit-instead-of-the-batch, a-unit-is-dispatched-twice-in-one-batch, a-unit-nothing-checks-is-still-a-unit or the-parent-check-ignores-its-own-verdict, so each of those rules has a check and no row. The checks are named in the record. Whether those seven rules deserve rows is a question about the ledger, and it is left as one rather than answered by inventing rows here. The target tools/agent-verify.ts had a single tooth and goes with it, so the register is 110 teeth over 22 entries covering 21 files. Docs and code in one commit: the obligations ledger's rows stop naming the retired teeth as evidence and name the case that states each rule instead (B3, B6, C3, D3, D6, D11, D13, D14, E1, E3, F2b-slot, F5, the G rows and the prose that listed them); its readings gain the whole-register run for this pass and its "how a row earns proven" paragraph gains the retirement criterion, including the two ways a case can look like it states a rule without doing so (a fixture derived from the constant under test, a case that refuses for a second reason as well) - both of which this arc found by reading assertions; the record's item 3 is marked landed with the numbers, the kept tooth, the seven orphans and the next question. Readings at this revision: whole register -> mutants: 110 of 110 caught by the named test, all 110 by the case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 162 s; mutation:anchors -> anchors: 110 of 110 resolve, over 22 targets; test:product -> 1562 pass, 0 fail, exit 0; verify:static -> exit 0; lint -> 0 findings.
…ry close The register-wide conversion the pilot started, taken target by target: 59 more hand-written teeth are now a name, an operator and a selector, so a rename or a moved line cannot retire them. Per operator: condition-never 44 (a guard or a variable initializer), neutralize-term 9, replace-property 9, replace-argument 6, drop-statement 6, condition-holds 3. The register is 77 derived of 110. Reading after the conversion: mutants: 110 of 110 caught by the named test, all 110 by the case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 171 s; anchors: 110 of 110 resolve. The conversion does not make a tooth bite - the operator realises the same violation the hand anchor did - so the reading that matters is the second one: every tooth is still caught by the case it names. Three sites the operators could not name, each closed with a case in tests/tools/mutation-anchor.test.ts (16 cases to 19) rather than argued about: - A class constructor is a member. BoardAdmission's constructor refuses a second plan while it opens the store, and uniqueMember knew only methods and function declarations, so that site had no `within` to write. The name is `constructor`; matching it is what let a-second-plan-silently-adopts-the-run convert. - A condition written across lines is one condition. The candidate filter compared raw text while the whole-condition check compared whitespace-normalized text, so a fragment of a wrapped `if` found no candidate and then matched nothing. Both normalize now, and the filter no longer demands the fragment be unique inside the candidate: `b` appears three times in `a && (b || !b)` and that is still the condition a selector naming `b` means. - A guard clause is a statement. `if (...) throw ...;` was not among the statements drop-statement would remove, because an `if` is an IfStatement rather than an expression, a declaration, a `return` or a `throw`. It is now, and statements resolve to the innermost one containing the fragment - the rule conditions already followed, and what keeps a fragment from matching both a guard and the statement inside it. One wrong `within` was shipped in this pass and the sweep is what caught it, which is the argument for sweeping after a conversion instead of trusting the anchors pass: the-completion-ignores-a-cancelled-unit was aimed at checkDispatch rather than checkCompletion (the fragment occurs in both members), so the anchors pass resolved, the mutant was still caught - and it was caught by the suite, not by the case that names a completion of a cancelled unit. Re-aimed, that case fails again. What is left hand-written is 33 teeth, and they cluster by the operator they would need: a comparison rewritten (3), a fragment inside a template or SQL string (5), an iterable emptied or a filter dropped (4), a call or `new` replaced or unwrapped (6), an index moved (2), an initializer replaced by a different expression (3), a condition negated (2), a statement rewritten into another statement (3), two statements sharing one line (2), a literal swapped (1), a site at module top level where there is no member to name (1), and one tooth that changes two things at once. Four clusters look worth an operator (replace-comparison, replace-fragment, empty-iterable, replace-callee); the record says so and leaves the decision open rather than adding operators nobody has asked for yet. Docs and code in one commit: the record's item 2 carries the numbers, the three gaps and the wrong `within`; its remaining list is the 33-tooth classification; the obligations ledger says 77 of 110 teeth are derived, names the four candidate operators, and gains this pass's whole-register reading beside the one it already had. The tools' own headers say a constructor's name and which statements drop-statement covers. Readings at this revision: whole register -> 110 of 110 caught, all by name, 22 of 22 restored, 171 s; mutation:anchors -> 110 of 110 resolve, over 22 targets; mutation-anchor cases -> 19 pass, 0 fail; test:product -> 1565 pass, 0 fail, exit 0; verify:static -> exit 0; lint -> 0 findings.
The 33 teeth still hand-written after the register-wide conversion were classified by what they would
need, and the classification was checked against what mutation tools actually name (Stryker's supported
mutators, pitest's mutator list, cargo-mutants' patterns, Cosmic Ray's operator concept). Thirteen of
those sites fall into mutation classes every catalogue carries, so the vocabulary grew instead of the
class being declared impossible:
negate-condition pitest NEGATE_CONDITIONALS / Stryker boolean literals 2 teeth
negate-comparison pitest NEGATE_CONDITIONALS / Stryker EqualityOperator 2 teeth
remove-conditionals pitest REMOVE_CONDITIONALS 1 tooth
remove-call Stryker filter/slice/sort removals / pitest VOID_METHOD_CALLS 2 teeth
replace-call Stryker MethodExpression / pitest CONSTRUCTOR_CALLS 1 tooth
replace-initializer pitest PRIMITIVE_RETURNS, INLINE_CONSTS / literal mutators 5 teeth
replace-iterable Stryker ArrayDeclaration / pitest EMPTY_RETURNS 3 teeth
Each computes its own bytes where the mutation determines them (false, true, the negated operator, the
call's receiver, the guard's body) and takes the mutant's `to` only where the new value is a choice -
the rule the earlier operators already followed. Operators are now a table of resolvers over one
context, so adding one is an entry plus its name in the union, and the dispatch's complexity stopped
growing with the vocabulary.
Two widenings, each shown by a real tooth that could not convert:
- A function bound to a name is a member. `const count = (label) => {...}` is the whole of
board-worker.ts's logic and it is not a function declaration, so its only tooth had no scope to name;
the selector now accepts a variable whose initializer is an arrow function or a function expression.
- Two identical calls need a holder to tell them apart. dispatchPlan calls board.candidates() on offer
and again filtered; `in` now names a fragment of the statement the call sits in, as it already named
the holding object literal for replace-property and the call's own text for replace-argument.
A wrong `within` was shipped again, and this time the anchors pass caught it loudly: the guard is in
runParentCheck and I named instrumentCommit, so it refused with "0 guards in instrumentCommit match the
selector". The first pass's wrong member was the quiet kind - checkDispatch instead of checkCompletion -
where the pass resolved, the mutant was still caught, and only the sweep showed it was caught by the
suite rather than by the case that names the rule. Wrong member is loud when it finds nothing, silent
when it finds the same shape twice: the sweep after a conversion is not optional.
Reading: anchors 110 of 110 resolve over 22 targets; mutants 110 of 110 caught by the named test, all
110 by the case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 176 s; the
anchor cases 30 pass; test:product 1576 pass; verify:static exit 0; lint 0 findings; complexity gate ok.
95 of the 110 teeth are derived now.
The residue is 15 teeth and its classification is measured, not guessed: three fragments inside a string
(two SQL clauses, one SQLite unit - catalogues mutate whole literals, and SQL has its own catalogue);
two at module top level where there is no member to name; two moving an array index, which corrects this
record's own earlier claim that the catalogues cover that class (FirstToLast is a method named first);
two writing two statements on one line; one rewriting a comparison into a range test, where the
catalogue's negation would not realise the violation the name states; four bespoke expression rewrites;
and one that changes two things at once. Three clusters would repay an operator and are recorded as open
(clause-level SQL mutation, a file-scope selector, an index operator); none is a widening, and the
measured cost of not having them is seven teeth.
The catalogue check also supplies the rationale this record was missing: pitest's "Less is more" names
subsumption - a mutant subsumed by others adds runtime, not confidence - which is what the retirement
criterion measures rather than guesses at; cargo-mutants' unviable mutants and its documented refusals
to generate some mutations are the same argument used here for leaving rare shapes alone. And one thing
is deliberately not borrowed: those tools count a mutant killed when any test fails, while this register
counts it only when the case it names is the one that fails, which is why "caught by the suite, not by
the case that names the rule" is a defect it reports and they cannot.
Docs and code in one commit: the record carries the operator table, the two widenings, both wrong-member
episodes, the measured residue and the two catalogue lessons; the ledger says 95 of 110 are derived and
gains this pass's readings; the teeth header says which names a member can have and that the operators
are the catalogues' mutation classes.
… no hand-written anchor left Asked why the remaining 15 were not converted, the honest answer turned out to be that they could be - and that three of the reasons recorded for leaving them were wrong. Measured, one at a time, each conversion keeping its name and its named case: - the file scope: `within` is optional, and without it the whole file is the scope with the selector required to be unique in it (a top-level `CLOCK_GRACE_MS`, a guard in a module-level script); - `replace-index`: the element access the selector identifies gets the declared index (2 teeth); - `replace-literal-fragment`: a piece of text written inside a literal becomes the declared fragment, with the same holder-statement disambiguator the call selectors have, because one SQL predicate is written seven times in one member (3 teeth: two clauses and one SQLite time unit); - the call selectors accept `new X(...)` as a call site (1 tooth, the round store opened as a constructor call); - `uniqueMember` accepts a function bound to an object property, `claim: (store, parsed) => ...` (1 tooth, the daemon's handler table); - `replace-property` writes a shorthand property out, `position` to `position: 0`, because a value cannot go where the name was (1 tooth); - five more converted with operators that already existed: a call replaced by a boolean, a value replaced by a stub fact, a wrapper unwrapped to its inner call, an iterable shortened, a comparison negated. Three recorded reasons were wrong, and the sweep is what said so: - "the catalogue's negation would not realise the violation" (speculation-guesses-several-facts-at-once): it does, and the named case fails under `!== 1`; - "it changes two things at once" (every-task-is-frozen-at-position-zero): only one of the two is load-bearing, because `noUnusedParameters` is not set, so the callback keeps the parameter it no longer reads; - "two statements sharing one line" (two teeth): the hand anchors had to span two statements, but the mutations are single-statement deletions and both convert to `drop-statement`. Two mistakes made and caught inside this pass, both worth keeping: - a hand-typed `expect` reads as a broken mutant. Retyping two case names from memory instead of copying them made both teeth report "survived" - the tool filters the suite by the named case, so a name that matches no case is indistinguishable from a mutant nothing catches. Loud, not a false pass, and the conversion scripts now copy the string rather than retyping it; - rerunning a conversion script after the vocabulary had grown rewrote two teeth back to their earlier form; the anchors pass refused both at once, which is the case for the pass being in the static contract seen from the other side - it catches edits to the register, not only drift in the code. Reading: anchors 110 of 110 resolve over 22 targets; mutants 110 of 110 caught by the named test, all 110 by the case their expect names, 22 of 22 restored byte-identically, exit 0 in 160 s; anchor cases 34 pass; test:product 1580 pass; verify:static exit 0; lint 0 findings; complexity gate ok; format check clean. The register is 110 teeth over 22 targets covering 21 files, every one a name plus an operator plus a selector, with 26 operators - each a mutation class the catalogues name or a slot they name. What is not claimed: a derived tooth is not stronger evidence than the byte anchor it replaced. It is the same violation caught by the same case; what it buys is that the tooth stays aimed at its rule across renames and reflows, which is the failure this record exists for. The operators are also not the whole catalogue, only the classes this register's rules need. Docs and code in one commit: the record's residue section is replaced by the conversion table, the three corrected reasons and the two mistakes; the ledger says all 110 teeth are derived and gains this reading; the teeth header says which names a scope can have, module level included.
The four Plan items landed, and the acceptance criterion is met more strongly than it asked: the hand-written residue is zero, not small. The record proposed a doctrine; the doctrine is now the register's only form, so it moves to docs/decisions/implemented/ and takes the implemented format. - Problem keeps the motivation and the measurement that removed an option (the three-way split does not hold: every caught mutant is externally observable). - Proposal and Plan become Decision - the durable part is a name, an operator, a selector and the case that must fail - and Implementation state: the three waves, the widenings the residue forced, the corrections a run caught rather than a review, the readings, and what is not claimed. - Acceptance criteria is gone (the implemented format forbids it); its content is where the readings are. The old Risks become Consequences, together with the costs of the new form and the two questions left open on purpose. - Both languages in this commit, and the seven inbound links in docs/design/task-unit-semantics-obligations.md follow the move. - docs/design/ci-cd-and-quality.md: the anchors pass resolves 110 teeth now, not the 149 the sentence was written about, and it takes about a second. Readings at this revision: anchors 110 of 110 resolve over 22 targets; the whole register 110 of 110 caught, all 110 by the case each expect names, 22 of 22 targets restored byte-identically, exit 0 in 160 s; verify:static exit 0.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
A naming decision, so the main line can be discussed without re-deriving what it means each time.
Two levels, one meaning each:
protocol, while the mechanics of that coordination stay with the program rather than the model's
discussion.
dependencies, acceptance, capability, budget), and the protocol fixes how it is adopted into one
run, claimed, delivered and independently judged, and which facts the runtime decides and owns.
The parts keep the names they already have: Task Board is the sub-protocol's surface, managed
entry is what a governed entry is, run is the frozen object, adopt is the binding
transition, and the caller that decides to adopt is called the adopter.
Why these words
Everything that came to hand was already taken, and each collision is recorded as considered and
rejected rather than quietly reused: fusion already names one mechanism (several units sharing a
session), a scheduler is what the task-unit semantics decision deliberately does not enable,
governance is the board addressing and readability line, ledger already names three things
(assumptions, budget, disclosure), and executor names the bounded context adapters and the Pi SDK
execution path.
What it deliberately does not do
run,adopt,managed,dispatch, theooo-prefix andevery path stay as they are, because the term and the paths are cited by dated measurement records
and frozen run archives.
docs/glossary.yamlis untouched: it owns the repository and process vocabulary, and these areproduct concepts, which live in the concept map.
the shared dispatch loop); the product path still has no adopter, which is the next piece of work
rather than a consequence of the name.
Verification
npm run docs:check: 278 files, 0 errors, 0 warnings (58 implemented decisions).npm run glossary:check: 14 terms, 0 errors (unchanged).Follow-up to #70 (already merged). This branch was cut from
mainata9f56759, so it carries onecommit.
Added in the same branch
The naming record now carries an index section: one line of meaning and a pointer per name,
covering both levels, the board verbs and the channel rules, the run and its facts, the legal action
set and the dispatch loop, plus the two words deliberately not used (scheduler, OoO) with the owner
of their real meaning. It says of itself that it is an index rather than a second specification and
that the owner wins on any disagreement; the definitions were previously only reachable by opening
the design, the obligations ledger and several decision records at once.
The concept map's Task Board row pointed at an anchor that does not exist in
memory-graphs.md(
#21-task-board-outside-the-three-memory-graphs).docs:checkvalidates that a link's fileexists, not its anchor, so the link read as working; it now names the heading that is really there,
in both languages.
The naming record also carries the conceptual history now, because the names only make sense in
order: the board becoming a protocol (2026-08-13), governance and addressing (09-06), out-of-order
execution borrowed from the CPU with its limits written down (09-09..09-11), the task unit's
semantics (09-13), the paid arms measuring both borrowed halves (09-18..09-19), and the board
absorbing the execution half (through 09-20). Commit-level lineage stays in
implementation-lineage.md.Corrected while writing it:
ooo-execution-bootstrap.mdstill said the speculation decision wasproposed and that the no-speculation rule held until it was implemented. It has been implemented
since 2026-09-11; the sentence now states what is actually in force.
Adds a proposed record for the question the naming decision unblocked:
docs/decisions/proposed/2026-09-20-the-program-answers-legality.md(+ zh). The primitives have names andthe shared layer already computes the legal set, but nothing stated whose decision each step is. The record
gives the program exactly one share (answer legality: the ordered legal set, cut to the declared slot budget,
with a reason per unit - choosing nobody, adopting nothing, judging nothing, waking nobody), leaves the plan,
the order, the adopter and the judge to the agents as board facts, makes constraints named by the protocol and
enabled by the plan, and keeps the query free of state because a published handoff occupies a serial slot. It
is proposed, not implemented: the readable answer and the promotion of repair-first out of shared planning
policy are its plan steps 2 and 3.
Adds the umbrella's parts inventory as a design document:
docs/design/protocol-governed-collaboration.md(+ zh), and points the naming record's history section atit. Fifteen parts with owner and state (entries/addressing, identity, ownership/time, termination integrity,
attention, scope, truth vs coordination state, the task unit, legality, ordering, budget, cancellation,
recovery, handoff, evidence), seven gaps with the distributed-systems concept that fits each (commit boundary,
orphan reaping, admission control, declarative policy, monotonic reads, conflict detection, fenced mutex,
in-doubt transaction), what is deliberately not needed (consensus, leader election, quorum, 2PC, exactly-once,
distributed transactions), and three analogies that mislead (a lease is a suspicion, not a death certificate;
the board is not a queue; agents are not replicas).
Adds a second proposed record, the frame itself:
docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md(+ zh). The primitives have names and theprogram's job is proposed, but nothing said which fields the frame has, how values are encoded, where they
live in the store, or how a peer protocol sits beside the Task-Unit Protocol. Researching the finished
specifications changed the shape of the problem: every layer but one is already defined by somebody, so the
record chooses and names instead of inventing. It borrows per layer (CloudEvents for the envelope, A2A for the
task model and paging, NLIP/ECMA-430 for the payload discriminator and the progressive shape, MCP for
LLM-facing results and errors, RFC 8941 for a one-line text type system, FIPA-ACL for the act and threading
fields, RFC 6709 for the unknown-field choice, ECMA-434 for security) and states that the one layer nobody
defines is ownership, lease, fencing and legality, because every agent protocol assumes a task belongs to one
agent. Three decisions carry it: the frame is a state record rather than a message, so FIPA's
performativeis a rendering of existing columns; the header stays in columns because the claim is one atomic
compare-and-set
UPDATEand a header inside JSON would turn that into read-modify-write; and the payload isone opaque document whose invariant is that every payload can be NULL while T0/T1 still render and claim,
wake, expiry and compact read still work. A payload field a protocol wants to query is indexed by a generated
column plus an expression index instead of a new board column, which is what keeps the schema from growing
with the number of protocols; the protocol is a versioned value and a peer protocol supplies declaration
validation, legality, projection and acceptance through the existing
DispatchBoardandRoundQueryPortseams. Data format is answered as three separate questions (existing JSON RPC wire; flat key/value for T0/T1
with natural language for negotiation, since format restrictions measurably degrade reasoning; typed columns
plus one JSON payload document), and TOON and CBOR are deliberately not adopted. Security is named as the
empty layer, with two rules adopted as text from ECMA-434 and a threat list owed. The record ends with an
explicit scope ceiling so the change cannot grow: one group of nullable columns, one boundary adapter, one
rendering rule, one test - anything needing a new table, tool or channel is outside it.
The frame record answered its first review in a second commit. Four points, all accepted. (1) The field
mapping was missing and it changes what the invariant means: the record now maps every existing field to
its place in the frame, records that
contentstays the single source of the entry body(
readTaskBoardPreviewsprojectstaskBoardPreview(entry.content)), that the frozen declaration wasnever a body (it lives in
task_run_manifest/task_run_tasks), thattask_run_facts.payloadis theprecedent for the payload shape, and that a payload may not restate a field that already has a column -
which is how the second task body would appear. (2) "The insertion points already exist" was false:
RoundQueryPortcarries two typed reads and the dispatch loop requires apatch, so the four-role seamis a plan step, not a fact. (3) Two storage arguments are withdrawn with the measurement that kills them
UPDATEclaims over a JSON payload exactly once, and a generated column is a column, sopeer protocols are not schema-free; what remains is clarity, indexing,
CHECKconstraints, additiveprotocol-owned schema, and no field stored twice. (4) Three borrowings were too wide:
source_session_idis provenance rather than a credential, the refusal rule keeps MCP's protocol-error/execution-error
split (which the dispatch loop already makes as
refusedversusfailure), and the CloudEvents namingconstraint is stated as borrowing with
protocol_version's underscore recorded as a deviation.Adds a third proposed record that names the division the other two each assumed:
docs/decisions/proposed/2026-09-21-mechanism-not-policy.md(+ zh). The legality record says what the programdoes and refuses to decide; the frame record says what the core knows and never parses; they are one division
seen from the decision side and the data side. The record states it as two contracts (mechanism: one claim with
its lease and fence, an opaque declaration with a digest, an opaque artifact with a digest, a verdict from
someone else, the lifecycle, and the legal set with a reason per unit; policy: inputs, artifact class,
judgement, how the work runs, concurrency preconditions, the wording and enablement of legality rules, and who
is chosen, adopts or judges), adds a decision procedure that settles arguments rather than serving as a slogan
(a rule's check is mechanism, its name and enablement are policy), and adds a mechanically checkable rule: no
policy word may appear in the mechanism layer's code, types or schema paths, from a maintained list, with the
path carried because a word can be policy in one place and mechanism in another. It also states what this makes
of work shapes - a work shape is not a core concept but the policy layer's name for the declaration-and-artifact
pair the core carries opaquely - classifies the six sites where the patch assumption lives, and keeps three rows
deliberately undecided so they are not settled by preference. The legality and frame records now point here.
The board entry has two faces, and the read face now has one home
The layer model in
docs/design/mechanism-in-the-middle.mdwas retired and replaced in the same document bya model that came from measurement rather than from analogy: a board entry is a medium with two faces - a
write face that a producer puts a body on and a read face that a reader is shown - and the compositor is
brought by a protocol, not by the board. The Task-Unit Protocol is the one that brings ours, which is the
same statement as "the program only answers legality": no declared legality, no judgement. Of the seven entry
kinds only
handoffandresultare ever claimed, delivered and judged; a note, a goal and a question travelfrom the write face straight to the read face.
Two facts the drawing has to survive, both measured: the roles are not files (
src/core/store/base.tsisthe medium, the compositor and the reader's rule at once), and feedback is a replay, not a reversal (a
reader becomes the next writer and enters through the write face again; the middle is never traversed
backwards).
The rule that decides where a seam goes needs no taste, only a count: a boundary is built when it has a
second implementation, and only declared with one.
the pi adapter, and a bare 140-character slice that did not collapse whitespace at all, so a multi-line body
reached a broadcast as a multi-line broadcast.
src/core/board-entry-view.tsowns the rule now; the store'scompact read and the adapter's broadcast call it; the private copy and the bare slice are gone;
tests/core/task-board.test.tspins the rule.entry rather than one shape drawn three ways. The trigger for a type is a fourth filler.
Check (c): the seam is prepaid, and the compiler cannot see the coupling
Making the loop shape-agnostic (
DispatchTicket.patch?: PatchWorkreplaced by an opaquedeclaration?: {shape, digest, frozen},PlanWorkerpointed at it,preparePatchWorkdeleted from the unitdispatcher) is one file, fifteen insertions and eighteen deletions. Measured on a throwaway branch and
reverted, it gave two results, neither of them the one the prediction was aiming at:
tsc --noEmitreportedzero errors, because the field is optional and structural typing still lets the board's ticket satisfy the
port; and the suite reported 25 failing tests, every one with the same cause - the claim admits no work.
A second shape today would not fail as a new shape; it would refuse every unit. The coupling's surface by
grep is three source files, the board (which must change although it never reads the field), seven test files
and the driver's three worker kinds. So the seam is deferred with a named trigger - the first non-patch unit
an adopter declares - rather than built against a single shape.
Docs and code land in one commit, because the read face's home is what the two-face model implements.
The port's third implementation: the store's own board
src/integration/ooo-runner.tsisDispatchBoardimplemented over the product's record - the probeboard and the dispatch test's stub are the other two. It carries no legality rule: the plan is
compiled by
compileTaskUnits, the legal set isdispatchTasks+selectableTasks, cancellations comefrom the coordinator's own reader, and the verdict belongs to the caller's acceptance. A rule that turns
out to be needed there belongs in
task-semantics.ts; needing one is how the module would show that theabstraction leaked.
Two inputs stay the caller's, and both are existing divisions rather than new ones: a task's declared file
contents (the store freezes paths) and an acceptance whose identity is independent of the deliverer (the
store already refuses a deliverer that judges its own delivery - pinned by the second test).
One contract is easy to get wrong, so the projection writes it down: the shared acceptance rule binds a
verdict to the artifact value the run carries (it compares the verdict's digest against the artifact),
not to the store's own deliverable hash. Reporting the hash there makes every accepted unit read as
unaccepted, and the symptom is silent - the next unit is never released and nothing reports an error.
The port's "put a result" is the product's delivery on the entry the unit was claimed on, so a unit's
work is recorded once rather than once on that entry and again on a result entry nobody would read; and
the unit is resolved from the claim rather than from the body, because a body's shape belongs to its owner
while a claim is the board's own fact.
Verification:
npm run checkclean;test:product1527/1527 pass (2 new - a two-unit plan driven throughthe product's store with an independent judge, and the self-judging refusal);
buildclean;lintclean;format:checkclean;docs:check288 files / 0 errors / 0 warnings;agent:context:checkvalid once theroute claims the new file;
complexity:gateok.