From d6a0720e739c08d43bffec7d088f3cacc9d9977c Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Sun, 20 Sep 2026 22:59:42 +0800 Subject: [PATCH 01/32] docs(vocabulary): the collaboration protocol and its task-unit sub-protocol get names Two levels with one meaning each. Protocol-governed collaboration is the umbrella: agents coordinate through a published protocol and the mechanics stay with the program. The Task-Unit Protocol is the sub-protocol this line builds - a task is declared as a unit and the protocol fixes how it is adopted into one run, claimed, delivered and independently judged, and which facts the runtime decides. The parts keep the names they already have: Task Board is the sub-protocol's surface, managed entry is what a governed entry is, run is the frozen object, adopt is the binding transition, and the caller that decides to adopt is called the adopter. The words that came to hand were already taken - fusion (units sharing one session), scheduler (which the task-unit semantics decision deliberately does not enable), governance (the board addressing line), ledger (three of them) and executor (the bounded context adapters) - so each is recorded as considered and rejected rather than reused. No identifier or file name changes, for the reason the check-runner renaming recorded: `ooo` and the existing paths are cited by dated measurement records and frozen run archives. The name lives in the concept map with its aliases, and docs/glossary.yaml is untouched because it owns the repository and process vocabulary, not product concepts. --- ...6-09-20-name-the-collaboration-protocol.md | 69 +++++++++++++++++++ ...0-name-the-collaboration-protocol.zh-CN.md | 55 +++++++++++++++ docs/design/task-unit-semantics.md | 2 + docs/guides/concept-map.md | 2 + docs/guides/concept-map.zh-CN.md | 2 + 5 files changed, 130 insertions(+) create mode 100644 docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md create mode 100644 docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md new file mode 100644 index 00000000..2cdc294d --- /dev/null +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md @@ -0,0 +1,69 @@ +# Name the collaboration protocol and its task-unit sub-protocol + +[中文](2026-09-20-name-the-collaboration-protocol.zh-CN.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [Task unit semantics](../../design/task-unit-semantics.md), [The dispatch loop is shared](2026-09-19-dispatch-loop-is-shared.md) + +## Problem + +The arrangement this line is building had no name, so every discussion had to re-derive what it +meant: whether the subject was the board, the run object, the verbs an Agent uses, or the rules that +decide. The words that came to hand were already taken in ways that invite the wrong reading. +_Fusion_ names one mechanism (several units sharing a session: `sharedSessionLegal`, +`fusionAccounting`), and a _scheduler_ is what the task-unit semantics decision explicitly does not +enable. _Governance_ is the board addressing and readability line, _ledger_ already names three +things (assumptions, budget, disclosure), and _executor_ names the bounded context adapters and the +Pi SDK execution path. Naming the new thing after any of those would make the conversation shorter +and the understanding worse. + +## Decision + +Two levels, each with one meaning: + +- **Protocol-governed collaboration** is the umbrella: agents coordinate through a published + protocol, and the mechanics of that coordination belong to the program rather than to the model's + discussion. It is the larger idea the board's verbs are one instance of. +- **Task-Unit Protocol** is the sub-protocol this line builds: a task is declared as a unit (inputs, + dependencies, acceptance, capability, budget), and the protocol fixes how that unit is adopted into + one run, claimed, delivered and independently judged, and which facts the runtime decides and owns. + "Task unit" is the word the design already uses; the protocol half is what was missing. + +The parts keep their existing names instead of acquiring new ones: **Task Board** is the +sub-protocol's surface, **managed entry** is what an entry becomes once a run governs it, **run** is +the object whose plan, tasks and facts are frozen, and **adopt** is the transition that binds an +entry to a run. The one thing with no name is the caller that decides to adopt; this record calls it +the **adopter**. + +## Alternatives considered + +- **Keep saying OoO.** Rejected as the umbrella: `ooo` is the project's own name for the scheduling + model it implements and is load-bearing in more than a hundred documents, but it names the model, + not the two-part arrangement of who discusses and who decides, so it does not remove the ambiguity + the naming exists for. +- **Board-governed execution.** Rejected: `governance` is already the board addressing and + readability line, and the phrase reads as the board deciding, which is the opposite of the point + that the program decides. +- **Task-unit lifecycle protocol.** Rejected as too long to say; the lifecycle reading is recoverable + from the definition without being carried in the name. +- **Managed run protocol.** Rejected because it drops the semantics half, which is the larger part of + the design and the half the board cannot check by itself. +- **A name built on "fusion".** Rejected: execution fusion already names one mechanism, and adoption + and dispatch happen whether or not any session is fused. +- **Rename code or the design document to match.** Rejected for the reason the check-runner renaming + recorded: the `ooo` term and the existing paths are cited by dated measurement records and frozen + run archives, so a rename would leave evidence pointing at paths that no longer exist. + +## Consequences + +- A discussion can now name its level: protocol-governed collaboration for the whole arrangement, the + Task-Unit Protocol for the semantics-plus-runtime sub-protocol, the adopter for the missing caller, + and the existing words (board, run, adopt, managed entry) for the parts. +- No identifier changes. `run`, `adopt`, `managed`, `dispatch`, the `ooo-` file prefix and every file + name stay as they are, so no historical record needs repairing. +- The name lives in the concept map with aliases for search. `docs/glossary.yaml` is untouched: it + owns the repository and process vocabulary, and this is a product concept. +- The naming builds nothing. The sub-protocol's mechanism exists (the run surface, the managed-write + fence, the shared dispatch loop) and the product path still has no adopter, which is the next piece + of work rather than a consequence of the name. diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md new file mode 100644 index 00000000..27cfb315 --- /dev/null +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md @@ -0,0 +1,55 @@ +# 给协作协议与它的任务单元子协议定名 + +[English](2026-09-20-name-the-collaboration-protocol.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [任务单元语义](../../design/task-unit-semantics.md)、[派发循环是共享的](2026-09-19-dispatch-loop-is-shared.zh-CN.md) + +## 问题 + +这条线在做的东西一直没有名字,于是每次讨论都要重新推导指的是什么:是黑板、是 run 这个对象、是 +Agent 用的那几个动词,还是拍板的那些规则。顺手能用的词又都已经被占,而且占法会引导出错误理解: +**融合** 已经指一个具体机制(多单元共享一个会话:`sharedSessionLegal`、`fusionAccounting`); +**调度器** 是任务单元语义那条决策明确说不启用的东西;**治理** 已经是黑板地址与可读性那一线; +**台账** 已经指三样(假设、预算、披露);**executor** 已经指受限的 context executor 适配器和 Pi +SDK 执行路径。用其中任何一个给新东西命名,讨论是短了,理解会更差。 + +## 决策 + +分两层,每一层一个含义: + +- **协议化协作(Protocol-Governed Collaboration)** 是大类:agent 之间通过一份公开的协议协商, + 而协商的机制属于程序,不属于模型的讨论。黑板的那些动词只是它的一个实例。 +- **任务单元协议(Task-Unit Protocol)** 是这条线在做的子协议:把任务声明成单元(输入、依赖、 + 验收、能力、预算),并规定它如何被收编进一次 run、如何被认领、交付与独立裁决,以及哪些事实由 + 运行时拍板并负责。"任务单元"是设计里已有的词,缺的是"协议"的那一半。 + +各组成部分沿用已有名字,不另起:**黑板(Task Board)** 是这个子协议的面;**受管条目(managed +entry)** 是条目被某次 run 治理之后的状态;**run** 是计划、任务与事实被冻结的那个对象; +**adopt(收编)** 是把条目绑到一次 run 的那次转移。唯一没有名字的是"决定收编的调用方",本记录叫它 +**收编者(adopter)**。 + +## 考虑过的替代方案 + +- **继续叫 OoO。** 作为大类被拒:`ooo` 是项目给自己实现的调度模型起的名字、在 100 多份文档里承重, + 但它指的是模型,不是"谁讨论、谁拍板"这个二分,所以消不掉定名要消的歧义。 +- **黑板治理的执行(Board-Governed Execution)。** 被拒:`治理` 已经是黑板地址与可读性那一线,而且 + 这个说法读起来像"黑板在拍板",与"程序拍板"正好相反。 +- **任务单元生命周期协议。** 被拒:太长,不好说;生命周期那层意思从定义里读得出来,不必写进名字。 +- **受管运行协议(Managed Run Protocol)。** 被拒:丢了语义那一半,而那是设计的主体,也是黑板自己 + 查不了的那一半。 +- **名字里带"融合"。** 被拒:执行融合已经指一个具体机制(多单元共享一个会话),而收编与派发无论有 + 没有融合会话都会发生。 +- **为了让名字成立去改代码或设计文档名。** 被拒,理由与"给检查身份与检查运行器定名"那条记录一致: + `ooo` 这个词与现有路径被带日期的测量记录和冻结的运行归档引用,改名会让证据指向不存在的路径。 + +## 后果 + +- 讨论可以指名层次了:整体叫协议化协作,语义加运行时那一层叫任务单元协议,缺的调用方叫收编者, + 其余部分沿用黑板、run、adopt、受管条目这些已有词。 +- 不改任何标识符。`run`、`adopt`、`managed`、`dispatch`、`ooo-` 文件前缀以及所有文件名保持原样, + 因此没有历史记录需要修补。 +- 名字住在概念图里并带别名以便检索。`docs/glossary.yaml` 不动:它负责仓库与流程词汇,而这是产品概念。 +- 定名本身不建任何东西。子协议的机制已经在(run 面、受管写入围栏、共享派发循环),产品路径仍然没有 + 收编者——那是下一步工作,不是这个名字的后果。 diff --git a/docs/design/task-unit-semantics.md b/docs/design/task-unit-semantics.md index f9f83e48..2aa49f69 100644 --- a/docs/design/task-unit-semantics.md +++ b/docs/design/task-unit-semantics.md @@ -5,6 +5,8 @@ ## 问题与目标 +本仓库的术语里,本文描述的协议叫**任务单元协议(Task-Unit Protocol)**,是**协议化协作(Protocol-Governed Collaboration)**的子协议;定名记录见[命名决策](../decisions/implemented/2026-09-20-name-the-collaboration-protocol.md)。 + 任务单元是可以交接、验证和独立作废的工作;Agent 是执行这些工作的资源。两者不必一一对应。先表达工作成立的条件,再由运行时决定分配给谁、是否连续执行,才可能同时获得细粒度并发与上下文复用。 本文提出现有任务契约的内部编译视图(Task IR),不新增用户输入语言或独立持久化 schema。首个检验对象是冻结仓库上的提案型工作。自然语言仍描述意图;结构化字段描述输入、产物、权限和可检查义务。无法形式化的意图保留人工或独立评审,不伪装成编译器能证明的事实。 diff --git a/docs/guides/concept-map.md b/docs/guides/concept-map.md index 9541afd8..8b8ce21a 100644 --- a/docs/guides/concept-map.md +++ b/docs/guides/concept-map.md @@ -45,6 +45,8 @@ directory; `get` opens selected exact evidence. | QPP | Optional prediction of whether retrieval is broad and complete enough | Progressive recall needs a reason to stop, expand, or fold noise | [retrieval confidence controller](../design/retrieval-confidence-controller.md) | | Memory chain | A bounded ordered view over existing memory IDs | Some tasks need event order or explicit dependencies without copying evidence | [design.md §7.6](../design/design.md#76-static-temporal-and-logical-memory-chains) | | Task Board | Attributed, expiring, task-scoped Agent coordination outside semantic memory | Private AGs cannot communicate directly across Agents | [memory-graphs.md §2.1](../design/memory-graphs.md#21-task-board-outside-the-three-memory-graphs) | +| Protocol-governed collaboration | Agents coordinate through a published protocol, while the mechanics and the deciding stay with the program | Coordination needs a name that separates who discusses from who decides | [Name the collaboration protocol](../decisions/implemented/2026-09-20-name-the-collaboration-protocol.md) | +| Task-Unit Protocol | A task is declared as a unit and adopted into one run, where it is claimed, delivered and independently judged | The board alone cannot say what a unit's inputs, dependencies and acceptance are | [task-unit-semantics.md](../design/task-unit-semantics.md) | | Maintenance and consolidation | Deterministic index work plus evidence-gated semantic promotion or topology proposals | Write cost must stay bounded and repeated use must not manufacture truth | [design.md §10](../design/design.md#10-incremental-storage-and-index-maintenance) | | Learnable controller | An optional numeric policy over hard-bounded allocation, fold, and rerank decisions | Natural outcome evidence may improve control without making memory graphs differentiable | [design.md §12](../design/design.md#12-learnable-routing-and-minimal-differentiable-query-graphs) | | Lab | Explicitly leased optional capabilities such as reasoning workspace and graph reasoner | Experimental mechanisms must be usable without silently becoming defaults | [design.md §12ter](../design/design.md#12ter-session-reasoning-workspace-and-compaction-checkpoint) | diff --git a/docs/guides/concept-map.zh-CN.md b/docs/guides/concept-map.zh-CN.md index 3c8eb671..1e70c62b 100644 --- a/docs/guides/concept-map.zh-CN.md +++ b/docs/guides/concept-map.zh-CN.md @@ -41,6 +41,8 @@ flowchart LR | QPP | 可选的检索广度与充分性预测 | 渐进式回忆需要决定停止、扩展或折叠噪声 | [检索置信度控制器](../design/retrieval-confidence-controller.md) | | Memory chain | 对现有 memory ID 的有界有序视图 | 时间顺序和显式依赖不应复制证据 | [design.md §7.6](../design/design.md#76-static-temporal-and-logical-memory-chains) | | Task Board | 位于语义记忆之外、有归因、有过期时间的任务级协作区 | 私有 AG 不能直接完成跨 Agent 通讯 | [memory-graphs.md §2.1](../design/memory-graphs.md#21-task-board-outside-the-three-memory-graphs) | +| 协议化协作 | agent 通过一份公开协议协商,机制与拍板留在程序侧 | 协作需要一个把“谁讨论、谁拍板”分开的名字 | [命名决策](../decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md) | +| 任务单元协议 | 任务被声明为单元并收编进一次 run,在其中被认领、交付与独立裁决 | 黑板自己说不清一个单元的输入、依赖与验收是什么 | [task-unit-semantics.md](../design/task-unit-semantics.md) | | 维护与巩固 | 确定性索引维护,加上有证据门控的语义晋升或拓扑 proposal | 写入成本必须有界,重复检索不能制造“事实” | [design.md §10](../design/design.md#10-incremental-storage-and-index-maintenance) | | 可学习控制器 | 在硬预算内可选地学习 allocate、fold 和 rerank 的数值策略 | 可利用自然结果改进控制,但无需让记忆图本身可微 | [design.md §12](../design/design.md#12-learnable-routing-and-minimal-differentiable-query-graphs) | | Lab | 通过显式 lease 使用的 reasoning workspace、graph reasoner 等可选能力 | 实验机制可以被使用,但不能静默成为默认行为 | [design.md §12ter](../design/design.md#12ter-session-reasoning-workspace-and-compaction-checkpoint) | From ebc053d6c264cb6a1afd97e1d775cdbcc96c34cf Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Sun, 20 Sep 2026 23:22:54 +0800 Subject: [PATCH 02/32] docs(vocabulary): the names get one index, with their owners The definitions were spread over the design, the obligations ledger and several decision records, so reading them meant opening four files at once. The naming record now carries an index section: one line of meaning and a pointer per name, for both levels, the board verbs and the channel rules, the run and its facts, the legal action set and the dispatch loop. It ends with the two words deliberately not used for this protocol - scheduler and OoO - each with the owner of its real meaning, because reusing either is what made the discussion ambiguous in the first place. The section states its own status: an index, not a second specification, and where a row and its owner disagree the owner wins. Same content in the zh pair. Found while checking the links: the concept map's Task Board row points at an anchor that does not exist in memory-graphs.md, and docs:check does not catch it because it validates file existence rather than headings. The index uses the heading that is really there. --- ...6-09-20-name-the-collaboration-protocol.md | 35 +++++++++++++++++++ ...0-name-the-collaboration-protocol.zh-CN.md | 34 ++++++++++++++++++ 2 files changed, 69 insertions(+) diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md index 2cdc294d..d599a932 100644 --- a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md @@ -36,6 +36,41 @@ the object whose plan, tasks and facts are frozen, and **adopt** is the transiti entry to a run. The one thing with no name is the caller that decides to adopt; this record calls it the **adopter**. +## The names, in one place + +This is an index, not a second specification: one line of meaning and a pointer per name, and where a +row and its owner disagree, the owner wins. It exists because the definitions used to be reachable +only by opening the design, the obligations ledger and the decision records side by side. + +The two levels and what they are made of: + +| Name | One line | Contract owner | +| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Protocol-governed collaboration | Agents coordinate through a published protocol, while the mechanics and the deciding stay with the program | this record; [concept map](../../guides/concept-map.md) | +| Task-Unit Protocol | A task is declared as a unit and the protocol fixes its adoption, claim, delivery, judging and the facts the runtime owns | [task-unit-semantics.md](../../design/task-unit-semantics.md), [its obligations ledger](../../design/task-unit-semantics-obligations.md) | +| task unit | A unit of work that can be handed off, verified and independently voided, declared by inputs, dependencies, acceptance, capability and budget | [task-unit-semantics.md](../../design/task-unit-semantics.md) | +| Task Board | Attributed, expiring, task-scoped Agent coordination outside semantic memory | [memory-graphs.md §2](../../design/memory-graphs.md#shared-task-board-cross-agent-coordination-not-a-memory-graph) | +| entry | One board item - goal, question, handoff, blocker, result, note or decision - carrying a claim lease, deliveries and verdicts | [board-find-serial-a2a-compat](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| wake | A directed entry notifies that Agent's session; the board does not decide who works, only that someone was addressed | [board-find-serial-a2a-compat](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| serial channel | At most one un-directed actionable entry is pushed at a time; the next is promoted when the outstanding one is claimed or resolved | [board-find-serial-a2a-compat](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| claim, release, resolve | The lease verbs: one holder at a time, expiry returns the entry to the pool, a resolve closes it | [board governance and addressing](2026-09-06-board-governance-addressing.md) | +| deliver, judge | A claim ends in a digest-bound deliverable, judged by a different Agent (accepted, rejected, undecidable) | [board governance and addressing](2026-09-06-board-governance-addressing.md) | +| managed entry | An entry a run has adopted: its lifecycle verbs are refused outside that run's coordinated scope | [obligations ledger, B6](../../design/task-unit-semantics-obligations.md) | +| run | The frozen object: a registered run, one frozen plan, bound entries and the fact log, reached over `taskRun` | [obligations ledger, D11 and D13](../../design/task-unit-semantics-obligations.md) | +| adopt | The transition that binds a board entry to a run, recorded as the run fact `entry-bound` | [obligations ledger, D12](../../design/task-unit-semantics-obligations.md) | +| adopter | The caller that decides to adopt; the one name here with no product implementation yet | this record; [obligations ledger, what is left](../../design/task-unit-semantics-obligations.md) | +| run fact | One recorded transition of a run: `entry-bound`, `board-claim`, `board-deliver`, `board-judge`, `run-cancelled` | `src/integration/task-coordinator.ts` | +| legal action set | The deterministically computed set a source may rank within, cut to the declared slot budget | [declared slot budget](2026-09-18-declared-slot-budget.md), `src/integration/ooo-execution.ts` | +| dispatch loop | The shared sequential loop that drives a plan unit by unit, with the board behind a port | [the dispatch loop is shared](2026-09-19-dispatch-loop-is-shared.md) | +| session, fusion | A unit's session is keyed on the board rather than the harness; fusion is several units sharing one session, a mechanism and not this protocol's name | [session identity comes from the board](2026-09-19-session-identity-comes-from-the-board.md), [fusion legality and accounting](2026-09-18-fusion-legality-and-accounting.md) | + +Two words that are deliberately not used for this protocol: + +| Word | Why not | Owner of the real meaning | +| ----------- | ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ | +| scheduler | The semantics deliberately does not enable one: ordering is deterministic rules over the legal set plus a declared budget | [task-unit semantics](2026-09-13-task-unit-semantics.md) | +| OoO, `ooo-` | The project's own name for the scheduling model it implements, and the file prefix; it names the model, not who discusses and who decides | [ooo-execution-bootstrap](../../design/ooo-execution-bootstrap.md) | + ## Alternatives considered - **Keep saying OoO.** Rejected as the umbrella: `ooo` is the project's own name for the scheduling diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md index 27cfb315..347ab1af 100644 --- a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md @@ -30,6 +30,40 @@ entry)** 是条目被某次 run 治理之后的状态;**run** 是计划、 **adopt(收编)** 是把条目绑到一次 run 的那次转移。唯一没有名字的是"决定收编的调用方",本记录叫它 **收编者(adopter)**。 +## 命名一览 + +这是索引,不是第二份规范:每个名字一行含义加一个指针,行与 owner 冲突时 owner 胜。它存在的原因是这些定义 +原本只能把设计文档、义务台账和几条决策记录摊在一起才看得全。 + +两个层次与它们的组成: + +| 名字 | 一句话含义 | 契约 owner | +| ------------------ | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 协议化协作 | agent 通过一份公开协议协商,机制与拍板留在程序侧 | 本记录;[概念图](../../guides/concept-map.zh-CN.md) | +| 任务单元协议 | 把任务声明为单元,并规定它被收编、认领、交付、裁决的方式以及运行时拥有哪些事实 | [task-unit-semantics.md](../../design/task-unit-semantics.md)、[义务台账](../../design/task-unit-semantics-obligations.md) | +| 任务单元 | 可以交接、验证和独立作废的工作,由输入、依赖、验收、能力与预算声明 | [task-unit-semantics.md](../../design/task-unit-semantics.md) | +| 黑板(Task Board) | 位于语义记忆之外、有归因、有过期时间的任务级协作区 | [memory-graphs.md §2](../../design/memory-graphs.md#shared-task-board-cross-agent-coordination-not-a-memory-graph) | +| 条目 entry | 一条黑板项(goal / question / handoff / blocker / result / note / decision),带认领租约、交付与裁决 | [board-find-serial-a2a-compat](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| 唤醒 wake | 定向条目通知那个 agent 的会话;黑板只负责把人叫到,不决定谁干活 | [board-find-serial-a2a-compat](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| 串行通道 | 同一时刻只推送一条未定向的 actionable;它被认领或关闭时晋升下一条 | [board-find-serial-a2a-compat](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| 认领 / 释放 / 关闭 | 租约动词:同一时刻一个持有者,过期回池,关闭即结束 | [黑板治理与能力寻址](2026-09-06-board-governance-addressing.zh-CN.md) | +| 交付 / 裁决 | 认领以 digest 绑定的交付物收尾,由另一个 agent 裁决(accepted / rejected / undecidable) | [黑板治理与能力寻址](2026-09-06-board-governance-addressing.zh-CN.md) | +| 受管条目 | 已被某次 run 收编的条目:在 run 的协调范围之外,它的生命周期动词被拒 | [义务台账 B6](../../design/task-unit-semantics-obligations.md) | +| run(运行) | 冻结的对象:已注册的 run、一份冻结计划、绑定的条目与事实日志,经 `taskRun` 到达 | [义务台账 D11、D13](../../design/task-unit-semantics-obligations.md) | +| adopt(收编) | 把黑板条目绑到一次 run 的那次转移,记为运行事实 `entry-bound` | [义务台账 D12](../../design/task-unit-semantics-obligations.md) | +| 收编者 adopter | 决定收编的调用方;本表里唯一还没有产品实现的名字 | 本记录;[义务台账“还剩什么”](../../design/task-unit-semantics-obligations.md) | +| 运行事实 run fact | 一次已记录的运行转移:`entry-bound`、`board-claim`、`board-deliver`、`board-judge`、`run-cancelled` | `src/integration/task-coordinator.ts` | +| 合法动作集合 | 确定性算出的、来源只能在其中排序的集合,按声明的槽预算切分 | [声明的槽预算](2026-09-18-declared-slot-budget.zh-CN.md)、`src/integration/ooo-execution.ts` | +| 派发循环 | 共享的顺序循环,逐单元驱动一份计划,黑板在端口之后 | [派发循环是共享的](2026-09-19-dispatch-loop-is-shared.zh-CN.md) | +| 会话 / 融合 | 单元的会话按黑板取键而非按 harness;融合是多单元共享一个会话——一个机制,不是本协议的名字 | [会话身份来自黑板](2026-09-19-session-identity-comes-from-the-board.zh-CN.md)、[融合的合法性与记账](2026-09-18-fusion-legality-and-accounting.zh-CN.md) | + +刻意不用的两个词: + +| 词 | 为什么不用 | 真实含义的 owner | +| ---------------- | ---------------------------------------------------------------------------------- | ------------------------------------------------------------------ | +| 调度器 scheduler | 语义层刻意不启用:排序是合法集合上的确定性规则加声明的预算 | [任务单元语义](2026-09-13-task-unit-semantics.zh-CN.md) | +| OoO、`ooo-` | 项目给自己实现的调度模型起的名字,也是文件前缀;它指的是模型,不是“谁讨论、谁拍板” | [ooo-execution-bootstrap](../../design/ooo-execution-bootstrap.md) | + ## 考虑过的替代方案 - **继续叫 OoO。** 作为大类被拒:`ooo` 是项目给自己实现的调度模型起的名字、在 100 多份文档里承重, From 44318cb2ad322f7f3b812580c0f73f238fc3c52c Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Sun, 20 Sep 2026 23:23:25 +0800 Subject: [PATCH 03/32] docs(concept-map): the Task Board row points at a heading that exists The row's anchor named "2.1 Task Board outside the three memory graphs", which is not a heading in memory-graphs.md: the section there is "Shared Task Board (cross-Agent coordination, not a memory graph)". Clicking the old link still opened the file, so it read as working. docs:check validates that a link's file exists rather than its anchor, which is why this survived; the row now names the heading that is really there, in both languages. --- docs/guides/concept-map.md | 2 +- docs/guides/concept-map.zh-CN.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/guides/concept-map.md b/docs/guides/concept-map.md index 8b8ce21a..3754ebd7 100644 --- a/docs/guides/concept-map.md +++ b/docs/guides/concept-map.md @@ -44,7 +44,7 @@ directory; `get` opens selected exact evidence. | `activeGraphId` | The stable ID of one retrieval projection, passed from `search` to `get` | Exact disclosure must be budgeted, session-owned, and attributable to its search | [design.md §2.1](../design/design.md#21-cli-and-resident-service) | | QPP | Optional prediction of whether retrieval is broad and complete enough | Progressive recall needs a reason to stop, expand, or fold noise | [retrieval confidence controller](../design/retrieval-confidence-controller.md) | | Memory chain | A bounded ordered view over existing memory IDs | Some tasks need event order or explicit dependencies without copying evidence | [design.md §7.6](../design/design.md#76-static-temporal-and-logical-memory-chains) | -| Task Board | Attributed, expiring, task-scoped Agent coordination outside semantic memory | Private AGs cannot communicate directly across Agents | [memory-graphs.md §2.1](../design/memory-graphs.md#21-task-board-outside-the-three-memory-graphs) | +| Task Board | Attributed, expiring, task-scoped Agent coordination outside semantic memory | Private AGs cannot communicate directly across Agents | [memory-graphs.md §2](../design/memory-graphs.md#shared-task-board-cross-agent-coordination-not-a-memory-graph) | | Protocol-governed collaboration | Agents coordinate through a published protocol, while the mechanics and the deciding stay with the program | Coordination needs a name that separates who discusses from who decides | [Name the collaboration protocol](../decisions/implemented/2026-09-20-name-the-collaboration-protocol.md) | | Task-Unit Protocol | A task is declared as a unit and adopted into one run, where it is claimed, delivered and independently judged | The board alone cannot say what a unit's inputs, dependencies and acceptance are | [task-unit-semantics.md](../design/task-unit-semantics.md) | | Maintenance and consolidation | Deterministic index work plus evidence-gated semantic promotion or topology proposals | Write cost must stay bounded and repeated use must not manufacture truth | [design.md §10](../design/design.md#10-incremental-storage-and-index-maintenance) | diff --git a/docs/guides/concept-map.zh-CN.md b/docs/guides/concept-map.zh-CN.md index 1e70c62b..3f9d94a3 100644 --- a/docs/guides/concept-map.zh-CN.md +++ b/docs/guides/concept-map.zh-CN.md @@ -40,7 +40,7 @@ flowchart LR | `activeGraphId` | 一次检索投影的稳定 ID,由 `search` 传给 `get` | 精确披露需要受预算、会话所有权和检索归因约束 | [design.md §2.1](../design/design.md#21-cli-and-resident-service) | | QPP | 可选的检索广度与充分性预测 | 渐进式回忆需要决定停止、扩展或折叠噪声 | [检索置信度控制器](../design/retrieval-confidence-controller.md) | | Memory chain | 对现有 memory ID 的有界有序视图 | 时间顺序和显式依赖不应复制证据 | [design.md §7.6](../design/design.md#76-static-temporal-and-logical-memory-chains) | -| Task Board | 位于语义记忆之外、有归因、有过期时间的任务级协作区 | 私有 AG 不能直接完成跨 Agent 通讯 | [memory-graphs.md §2.1](../design/memory-graphs.md#21-task-board-outside-the-three-memory-graphs) | +| Task Board | 位于语义记忆之外、有归因、有过期时间的任务级协作区 | 私有 AG 不能直接完成跨 Agent 通讯 | [memory-graphs.md §2](../design/memory-graphs.md#shared-task-board-cross-agent-coordination-not-a-memory-graph) | | 协议化协作 | agent 通过一份公开协议协商,机制与拍板留在程序侧 | 协作需要一个把“谁讨论、谁拍板”分开的名字 | [命名决策](../decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md) | | 任务单元协议 | 任务被声明为单元并收编进一次 run,在其中被认领、交付与独立裁决 | 黑板自己说不清一个单元的输入、依赖与验收是什么 | [task-unit-semantics.md](../design/task-unit-semantics.md) | | 维护与巩固 | 确定性索引维护,加上有证据门控的语义晋升或拓扑 proposal | 写入成本必须有界,重复检索不能制造“事实” | [design.md §10](../design/design.md#10-incremental-storage-and-index-maintenance) | From a16571bbd0f1e17f865e56b3b3d6d2ba8ed528ab Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Sun, 20 Sep 2026 23:34:34 +0800 Subject: [PATCH 04/32] docs(vocabulary): where the words came from, and one stale sentence about speculation The naming record now carries the conceptual history the names only make sense inside: the board becoming a protocol on 2026-08-13 (identity and discovery that wake no LLM, directed delivery, serial admission decided by claim/resolve/expiry), governance and addressing on 09-06, out-of-order execution borrowed from the CPU with its limits written down on 09-09..09-11, the task unit's semantics on 09-13, the paid arms measuring both borrowed halves on 09-18..09-19, and the board absorbing the execution half through 09-20. The OoO step is stated as the CPU correspondence rather than as "parallelism": out-of-order execution with in-order commit, where the plan and its tickets are the reorder buffer, the coordinator is the only retire stage, and the acceptance rules define what counts as commit. Its two limits are recorded with it, because the second one is the point the fusion and speculation arms then measured: the CPU's premise that a wrong guess wastes resources already committed or free does not transfer to agents, where a wrong guess spends tokens and wall time (a worker call measured at 40-56 s). Hence the decision adopts only the class whose wrong guess costs no tokens and refuses the class that hides a wait by guessing it, and the arms record "cost with no gain" at the shape tried: 43k tokens over six units, one false fact wasting 20 332 tokens, 0 of 3 prepared candidates publishable, and 175 ms of verification against 6.2 s of work once the fact holds. Commit-level lineage stays where it already lives (implementation-lineage.md, from its Task Board row) rather than being copied into the record. Also corrected: the bootstrap design still said the speculation decision was proposed and that the no-speculation rule therefore stayed in force until it was implemented. The decision has been implemented since 2026-09-11, so the sentence now states what is actually in force - the host-side free class may be done, hiding a wait by guessing stays forbidden. --- ...6-09-20-name-the-collaboration-protocol.md | 21 +++++++++++++++++++ ...0-name-the-collaboration-protocol.zh-CN.md | 16 ++++++++++++++ docs/design/ooo-execution-bootstrap.md | 2 +- 3 files changed, 38 insertions(+), 1 deletion(-) diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md index d599a932..90c63b23 100644 --- a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md @@ -71,6 +71,27 @@ Two words that are deliberately not used for this protocol: | scheduler | The semantics deliberately does not enable one: ordering is deterministic rules over the legal set plus a declared budget | [task-unit semantics](2026-09-13-task-unit-semantics.md) | | OoO, `ooo-` | The project's own name for the scheduling model it implements, and the file prefix; it names the model, not who discusses and who decides | [ooo-execution-bootstrap](../../design/ooo-execution-bootstrap.md) | +## Where the words came from + +The vocabulary is the residue of six steps, none of them planned as a naming exercise. Commit-level +lineage lives in [implementation-lineage.md](../../design/implementation-lineage.md); this is the +conceptual one, because the names only make sense in this order. + +| When | What happened | Where it is owned | +| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 2026-08-13 | The board was already an explicit multi-Agent coordination surface shared by several adapters, and it became a **protocol**: identity registration and discovery that wake no LLM, directed delivery (`to=`), and serial admission decided by claim, resolve and expiry rather than by arrival or acknowledgement - taking over is the claim, so a delivery is not a handover. A2A compatibility was researched at the same time. | [board find/direct and serial](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| 2026-09-06 | Governance and addressing: capability keys, compact reads, writer attribution. Still coordination - nothing here executes anything. | [board governance and addressing](2026-09-06-board-governance-addressing.md) | +| 2026-09-09 to 09-11 | **Out-of-order execution is borrowed from the CPU, with its limits written down.** The correspondence that holds is out-of-order execution _with in-order commit_: workers run out of order, the plan and its tickets are the reorder buffer, the coordinator is the only retire stage, and the acceptance rules define what counts as commit. Two limits are recorded with it. This design is stricter than the field's default of unbounded parallelism: only a real external wait may pass the queue head, and no sleep may manufacture reordering. And speculation is gated on measured payoff, because the CPU's premise does not transfer - there a wrong guess wastes resources already committed or free, while here it spends tokens and wall time (a worker call measured at 40-56 s). The decision adopts only the class whose wrong guess costs no tokens and refuses the class that hides a wait by guessing it; where a guess would almost always hold, the recorded advice is to remove the dependency rather than speculate it. | [bootstrap design](../../design/ooo-execution-bootstrap.md), [gate speculation on measured payoff](2026-09-11-ooo-speculation.md) | +| 2026-09-13 | Borrowing OoO makes the decomposition the bottleneck, and that half gets its own semantics: a **task unit** declared by inputs, dependencies, acceptance, capability and budget, compiled into an internal view (Task IR). Deliberately not a new user-facing language and not a second schema. | [task-unit-semantics.md](../../design/task-unit-semantics.md) | +| 2026-09-18 to 09-19 | **The paid arms measure both borrowed halves.** Fusing sessions buys wall time and costs tokens. Bounded speculation measures as cost with no gain at the shape tried: 43k tokens over six units, a false fact wasting 20 332 tokens, prepared candidates publishable in 0 of 3 holds, and 175 ms of verification against 6.2 s of work once the fact holds. The half of out-of-order execution that assumes a guess is cheap is therefore not available to agents, and the records say so instead of claiming a gain. | [arms archive](../../experiments/execution/archive/ooo-arms-2026-09-19/README.md), [obligations ledger F5](../../design/task-unit-semantics-obligations.md) | +| 2026-09-12 to 09-20 | **The board absorbs the execution half.** The first shape gave OoO its own session surface and its own tool; that was abandoned in favour of changing the board protocol itself. The board gained a deliverable bound to a digest, an independent judge, attempt fencing, then the run as a frozen object, adoption as a fact, the managed-write fence and the daemon's run surface - and the round's own tool left the product tool directory, so an ordinary handoff now carries semantics, execution and acceptance. | [collaboration absorbs OoO](../../design/task-unit-semantics-obligations.md), [the dispatch loop is shared](2026-09-19-dispatch-loop-is-shared.md) | + +The sequence is what the two names describe: protocol-governed collaboration, because the board was +already a protocol before it became an execution surface, and the Task-Unit Protocol, because the fine +decomposition is the part a task's semantics had to supply. It also explains the words recorded as +rejected above: a scheduler is what this model deliberately did not become, OoO is the borrowing's own +name, and fusion is one measured mechanism inside it rather than the whole. + ## Alternatives considered - **Keep saying OoO.** Rejected as the umbrella: `ooo` is the project's own name for the scheduling diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md index 347ab1af..c4324ab5 100644 --- a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md @@ -64,6 +64,22 @@ entry)** 是条目被某次 run 治理之后的状态;**run** 是计划、 | 调度器 scheduler | 语义层刻意不启用:排序是合法集合上的确定性规则加声明的预算 | [任务单元语义](2026-09-13-task-unit-semantics.zh-CN.md) | | OoO、`ooo-` | 项目给自己实现的调度模型起的名字,也是文件前缀;它指的是模型,不是“谁讨论、谁拍板” | [ooo-execution-bootstrap](../../design/ooo-execution-bootstrap.md) | +## 这些词是怎么长出来的 + +这套词是后面六步的沉积,没有一步是为了命名而做的。提交级谱系在 +[implementation-lineage.md](../../design/implementation-lineage.md)(它的第 44 行起就是黑板/A2A);这里写概念史,因为这些名字只有按这个顺序看才成立。 + +| 时间 | 发生了什么 | 归属 | +| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | +| 2026-08-13 | 黑板当时已经是多个适配器共享的、显式的多 Agent 协作面,这一步把它变成一份**协议**:身份注册与发现不叫醒任何 LLM、定向投递(`to=`)、串行放行由 claim / resolve / 过期决定而不是由“已投递”或“已读”决定——接手就是认领,所以投递不等于交接;同时做了 A2A 兼容研究。 | [先广播后定向 + 串行交接](../../design/board-find-serial-a2a-compat-2026-08-13.md) | +| 2026-09-06 | 治理与寻址:能力键、紧凑读取、写入者归因。仍然只是协作——这一步里没有任何东西在执行。 | [黑板治理与能力寻址](2026-09-06-board-governance-addressing.zh-CN.md) | +| 2026-09-09 至 09-11 | **从 CPU 借乱序执行,同时把它的限度写下来。**真正成立的对应是**乱序执行 + 顺序提交**:worker 乱序运行,计划与票据充当 reorder buffer,协调器是唯一的 retire 阶段,验收规则就是“什么算提交”的定义。写下时带了两条限度:本设计比领域的默认(无界并行)更严——只有真实外部等待可以越过队首,也不许用 sleep 造出乱序;以及推测必须按实测盈亏门控,因为 CPU 的前提在这里不成立:那边猜错只浪费已经提交或本来就空闲的资源,这边猜错要花 token 和墙钟时间(实测一次 worker 调用 40–56 s)。因此该决策只采纳“猜错不花 token”的那一类,拒绝“靠猜掩盖等待”的那一类;若某个依赖的假设长期几乎必中,记录下来的做法是取消这个依赖,而不是推测它。 | [自举设计](../../design/ooo-execution-bootstrap.md)、[按实测盈亏门控推测](2026-09-11-ooo-speculation.zh-CN.md) | +| 2026-09-13 | 借了乱序执行,拆分就成了瓶颈,于是这一半有了自己的语义:**任务单元**由输入、依赖、验收、能力和预算声明,编译成内部视图(Task IR)。刻意不是新的用户语言,也不是第二份 schema。 | [task-unit-semantics.md](../../design/task-unit-semantics.md) | +| 2026-09-18 至 09-19 | **付费臂量了借来的两半。** 会话融合买到墙钟时间、付出更多 token;有界推测在试过的形状上是“成本真实、收益为零”:六个单元 43k token,一次假事实浪费 20 332 token,准备好的候选在 3 次成立中 0 次可发布,而事实成立后校验只要 175 ms(对应 6.2 s 的活)。所以“猜对了不亏”那一半对 agent 不可用,记录照实写,不声称收益。 | [臂归档](../../experiments/execution/archive/ooo-arms-2026-09-19/README.md)、[义务台账 F5](../../design/task-unit-semantics-obligations.md) | +| 2026-09-12 至 09-20 | **黑板吞下执行那一半。** 最早的形状给 OoO 单独的会话面和单独的工具;这被改成“改黑板协议本身”。黑板先后得到 digest 绑定的交付物、独立裁决、attempt 围栏,然后是冻结的 run、作为事实的收编、受管写入围栏和 daemon 的 run 面——而轮次自己的工具离开了产品工具目录,于是一次普通交接就带着语义、执行与验收。 | [协作吞下 OoO](../../design/task-unit-semantics-obligations.md)、[派发循环是共享的](2026-09-19-dispatch-loop-is-shared.zh-CN.md) | + +这个顺序就是两个名字的来历:大类叫“协议化协作”,因为黑板在成为执行面之前就已经是一份协议;子协议叫“任务单元协议”,因为更细的拆分是任务语义必须供上的那一半。它也解释了上面被记为拒绝的那些词:调度器是这个模型刻意没有变成的东西,OoO 是借来的那个思想自身的名字,而融合是它内部一个已实测的机制,不是整体。 + ## 考虑过的替代方案 - **继续叫 OoO。** 作为大类被拒:`ooo` 是项目给自己实现的调度模型起的名字、在 100 多份文档里承重, diff --git a/docs/design/ooo-execution-bootstrap.md b/docs/design/ooo-execution-bootstrap.md index 0848567b..4abb4f14 100644 --- a/docs/design/ooo-execution-bootstrap.md +++ b/docs/design/ooo-execution-bootstrap.md @@ -150,7 +150,7 @@ S2 的退出条件“故障注入后无重复完成、错误解锁或无限等 ### 推测的落点(规则由决策记录拥有,本节只列可落地的位置) -是否启用、以及启用条件,由[推测盈亏平衡决策](../decisions/implemented/2026-09-11-ooo-speculation.md)拥有;本节只回答“哪些点适合”。该决策仍为 proposed,所以上面“不推测”的现行规则继续有效,直到它进入 implemented。按**猜错花掉什么**分三类,这一点决定了一切。 +是否启用、以及启用条件,由[推测盈亏平衡决策](../decisions/implemented/2026-09-11-ooo-speculation.md)拥有;本节只回答“哪些点适合”。该决策已 implemented(2026-09-11):宿主侧免费类可做,靠猜测掩盖等待仍然禁止,直到有实测盈亏支持它。按**猜错花掉什么**分三类,这一点决定了一切。 **已有前提**。安全推测需要的三件事已经就位:(a)错误的推测产物必须能丢弃且不得生效——补丁契约只允许产出提案、由宿主应用,worker 没有不可逆副作用,所以 squash 无需回滚任何状态;(b)提交点必须能廉价判定猜测对错——真实检查终态带检查身份与 digest,正好是那个判定点;(c)提交顺序不变。`inputDigest` 会把假设绑定进输入,因此**猜错的产物自动失效**(fenced)而不是被误接受:这正是 ROB 保证精确状态的那件事在我们这里的对应物。 From ef33abc3f49730ab17455331e7a499c3e7fdb9ce Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Mon, 21 Sep 2026 00:28:42 +0800 Subject: [PATCH 05/32] docs(decisions): the program answers legality, and nothing else Adds a proposed record for the question the naming decision unblocked: the primitives have names and the shared layer can already compute the legal set (an ordered legal set, that set cut to the declared slot budget, and a refusal reason per unit), but no document states whose decision each step is. The proposal splits the work three ways and gives the program exactly one share. The protocol owns what the primitives are, the definition and criteria of legality, the named constraint list, and the two structural rules that cannot be negotiated because they are what "legal" means here: a claim (lease plus attempt fence) is the only arbiter of who holds a unit, and a deliverer never judges its own delivery. The accompanying program computes the ordered legal set from the declared plan and current facts, cuts it to the declared slot budget, and states for each unit why it is legal or not - it chooses nobody, adopts nothing, judges nothing, wakes nobody. Agents own the arrangement: the plan's content, the order they agree on, who takes which unit, who adopts a run, who judges - all of it as board facts, so the arrangement is readable instead of being implied by a program's choice. Constraints are named by the protocol and enabled by the plan, so a preference such as repair-first is neither program policy nor advice; and legality is asked rather than published, because a published handoff occupies a serial slot by the board's own protocol, so fusing the two would make a question cost a slot. The plan is three steps and only the first lands here: this record, then a readable answer through an existing board action, then promoting the preferences that still live as shared planning policy (repair-first first) into constraints a plan declares - which is what makes the "not enabled means absent from the answer" criterion true. The record names that third step as the one place the proposal is not yet true. Deliberately not done here: no product surface change, no driver, no wake, no new tool or channel. --- ...2026-09-20-the-program-answers-legality.md | 105 ++++++++++++++++++ ...9-20-the-program-answers-legality.zh-CN.md | 60 ++++++++++ 2 files changed, 165 insertions(+) create mode 100644 docs/decisions/proposed/2026-09-20-the-program-answers-legality.md create mode 100644 docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md new file mode 100644 index 00000000..bb8e95e8 --- /dev/null +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md @@ -0,0 +1,105 @@ +# The program answers legality, and nothing else + +[中文](2026-09-20-the-program-answers-legality.zh-CN.md) + +**Status:** proposed +**Relates to:** [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Task unit semantics](../../design/task-unit-semantics.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md) + +## Problem + +The primitives have names, and the shared layer can already compute the legal set: an ordered legal set, +that set cut to the declared slot budget, and a refusal reason per unit when a claim is checked. What no +document states is **whose decision each step is**. The cost of that gap is observable: every caller +that wants to drive a run has to redraw the boundary for itself, so both recorded failure modes come +back - a program that picks the person for the agents, which is the scheduler the naming decision +records as the word this model deliberately did not become, and agents that cannot tell what the program +guarantees, so they ask for what they actually need in the only terms available: more slots, stranger +ordering. + +[The obligations ledger](../../design/task-unit-semantics-obligations.md) says where the boundary must +not move - no second task body, no dedicated tool, no dedicated channel - but it does not say what the +program's answer is. + +## Proposal + +Three owners, three kinds of statement. The program's share is exactly one: **it answers legality**. + +- **The protocol owns** what the primitives are (already named; this record does not restate them), the + definition and criteria of legality, which constraints exist **by name**, and the two structural + rules that are what "legal" means here and therefore cannot be negotiated: a claim (lease plus + attempt fence) is the only arbiter of who holds a unit, and a deliverer never judges its own + delivery. +- **The accompanying program owns** computing the ordered legal set from the declared plan and current + facts, cutting it to the declared slot budget, and stating for each unit why it is legal or not - + reusing the reasons it already refuses claims with. It chooses nobody, adopts nothing, judges + nothing, and wakes nobody. +- **Agents own the arrangement**: the plan's content, the order they agree on, who takes which unit, + who adopts a run, who judges. All of it lands as board facts, so the arrangement is readable and + auditable instead of being implied by a program's choice. + +**Constraints are named by the protocol and enabled by the plan.** A preference such as repair-first is +neither a program policy nor merely advice: it is a named constraint that a plan - or the adopter +declaring it - enables, and the legality answer is computed under the enabled set. A constraint that is +not enabled must be genuinely absent from the answer, or the declaration is decoration. + +**Legality is asked, not published.** It is a function of current facts, so a query returns the ordered +legal set with a reason per unit, changing no state. A published handoff on the board is a different +thing: by the board's own protocol it occupies a serial slot. Fusing the two into one action would make +asking a question consume a slot. + +## Plan + +1. This record: write down the division of labour and the standing of constraints. No product surface. +2. Make legality readable through an existing board action (the action name lives in the shared + tool contract, its description is generated from the prompt source), keeping the read free of state. +3. Promote the preferences that currently live as shared planning policy - repair-first first - into + named constraints a plan declares, which is what makes the "disabled means absent" criterion true. + +Each step is verifiable on its own; none of them requires a driver, a wake, or a new tool. + +## Alternatives considered + +- **Let the program select the next unit** (today's `next()` used as policy). Rejected: a bound is not a + decision - it does not choose among legal successors - and a program that picks the person is a + scheduler again, which the naming decision records as the word this model deliberately did not become. +- **Let the program store state only, with agents asserting their own legality.** Rejected: "legal" + would then be whatever the loudest caller says, and the two structural rules (one holder per claim, + no self-judging) would lose their home. +- **Demote preferences to advice.** Rejected: the arms compare runs on the premise that the same plan + plus the same facts yield the same answer, and advice that may be ignored removes exactly that. +- **Publish the legal set as a board offering** (make the query a handoff). Rejected: it makes a + question consume a serial slot, and withdrawing an offer is the board protocol's business. +- **Give the arrangement its own tool or channel.** Rejected: the ledger forbids it; the primitives must + take effect through an ordinary handoff. +- **Keep repair-first as hidden program policy.** Rejected: same class as a program that picks the + person. + +## Acceptance criteria + +- The answer is per unit, ordered, and carries the reason for each legal or refused unit. +- The answer names no person, no adopter and no judge. +- Asking changes no state: no entry, no claim, no slot, no wake. Asking twice on an unchanged store + returns the same answer, and the store's version is unchanged. +- A constraint that is not enabled has no effect on the answer: enabling and disabling the same named + constraint on one plan yields different legal sets. +- The declared slot budget still cuts the answer, and declaring more than one slot still requires a + declared handoff target. +- The two structural rules hold in the answer's own terms: a second claim on a held unit is refused, + and a deliverer cannot judge. +- Who adopted, who claimed and who judged exist only as board facts; no program field asserts them. +- The ledger's prohibition still holds: no new tool, no dedicated channel, no second task body. + +## Risks + +- With no constraint declared, a wide legal set can read as permission to fan out, recreating the + "no discussion, everything at once" problem somewhere new. Mitigation: the slot budget stays a + declared cap rather than a social one. +- Repair-first currently lives as policy inside the shared planning code, so making it a declared + constraint is a change rather than a rename; until it moves, the answer is computed under an implicit + constraint, which is the one place this proposal is not yet true. +- Reasons become an interface: a caller that parses a reason string creates a second source of truth. + Reasons are for people and logs; any caller decision must go through a claim or an explicit field. +- With agents free to choose order, the arms' equal-instrument premise depends on each run recording + which constraints were enabled; an unrecorded set makes two runs incomparable. +- A query invites polling. Polling is cheap, and a caller that asks and then acts can still race - the + claim remains the arbiter, which is the point. diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md new file mode 100644 index 00000000..698b6f1c --- /dev/null +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md @@ -0,0 +1,60 @@ +# 程序只回答合法性,别的都不管 + +[English](2026-09-20-the-program-answers-legality.md) + +**Status:** proposed +**Relates to:** [给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md)、[契约的义务](../../design/task-unit-semantics-obligations.md) + +## 问题 + +元语有了名字,共享层也已经能算出合法集合:有序合法集合、按声明槽位预算切一刀、认领被拒时给出每个单元的理由。但**没有任何一份文档在说"每一步是谁的决定"**。这个缺口有可观察的代价:任何想驱动一次 run 的调用者都得自己重新划一遍边界,于是两种已被记录在案的失败模式都回来了——程序顺手替 agent 挑人(也就是命名决策里记为"这个模型刻意没有变成"的那个词:调度器),以及 agent 不知道程序到底给什么保证,只能用唯一能用的说法去要它真正需要的东西:更多槽位、更怪的顺序。 + +[义务台账](../../design/task-unit-semantics-obligations.md)写明了边界不能往哪边动——不得再有第二个任务体、专用工具、专用频道——但它没有回答程序给出的那个答案是什么。 + +## 提案 + +三个所有者,三类陈述。程序那一份只有一条:**它回答合法性**。 + +- **协议拥有**:元语是什么(已被命名,本文不复述)、合法性的定义与判据、**约束的具名清单**(有哪些约束可以启用),以及两条结构性规则——它们就是"合法"在这里的含义,因此不可协商:认领(租约 + attempt 围栏)是"谁持有某个单元"的唯一仲裁;交付者不裁决自己的交付。 +- **配套程序拥有**:按声明的计划与当下事实算出有序合法集合,按声明的槽位预算切一刀,并给出每个单元合法或不合法的理由——复用它现在拒绝认领时用的那些理由。它不挑人、不收编、不裁决、不唤醒。 +- **agent 拥有安排本身**:计划的内容、他们商定的顺序、谁拿哪个单元、谁收编一次 run、谁裁决。所有这些都落成黑板上的事实,所以安排是可读、可审的,而不是由程序的选择暗示出来的。 + +**约束由协议具名、由计划启用。** 修复优先这类偏好既不是程序策略,也不只是建议:它是一条具名约束,由计划(或声明它的收编者)启用,合法性答案在启用集合下计算。没被启用的约束必须在答案里真的不生效,否则这个声明就是装饰。 + +**合法性被问,不被发布。** 它是当下事实的函数,所以查询返回有序合法集合,外加每个单元的理由,且不改变任何状态。黑板上发布出去的 handoff 是另一回事:按黑板自己的协议,它占用一个串行槽位。把两个合成一个动作,就会让"问一句"花掉一个槽位。 + +## 计划 + +1. 本文:写下分工与约束的地位。不动产品面。 +2. 让合法性可读——走已有的黑板动作(动作名在共享 tool contract 里,描述由 prompt 源生成),并保持这次读取不写状态。 +3. 把目前活在共享规划策略里的偏好,首先是修复优先,提升为由计划声明的具名约束;这一步才让"未启用即不生效"这条验收标准成立。 + +每一步都能单独验证;没有一步需要驱动器、唤醒或新工具。 + +## 考虑过的替代方案 + +- **让程序挑下一个单元**(把现在的 `next()` 当策略用)。拒绝:一个上限不是决定——它不在合法后继里做选择;而挑人的程序就是调度器,命名决策已把"调度器"记为这个模型刻意没有变成的词。 +- **程序只存状态,合法性由 agent 各自声明**。拒绝:那么"合法"就是嗓门最大的调用者说的;两条结构性规则(一个认领只有一个持有者、交付者不能自裁)会失去家。 +- **把偏好降成建议**。拒绝:arm 之间的可比性建立在"同一计划加同样事实给出同一答案"上,而可被无视的建议恰好拿走这一点。 +- **把合法集合作为黑板 offer 发布**(把查询做成 handoff)。拒绝:这样"问一句"要占串行槽位,而且撤回 offer 是黑板协议的事。 +- **为这套安排新开工具或频道**。拒绝:台账已禁;元语必须经普通交接生效。 +- **保留修复优先作为隐藏的程序策略**。拒绝:与"程序挑人"同属一类。 + +## 验收标准 + +- 答案按单元给出、有序,并带每个单元合法或不合格的理由。 +- 答案里不出现人、收编者或裁决者。 +- 问不改变状态:不产生条目、认领、槽位或唤醒。在未变更的 store 上问两次得到同一答案,且版本未变。 +- 未启用的约束对答案没有影响:同一计划上启用与停用同一条具名约束,合法集合不同。 +- 声明的槽位预算仍然切这个答案;声明多于一个槽位仍然要求声明 handoff 目标。 +- 两条结构性规则在答案自身的术语里成立:对已持有的单元再次认领被拒;交付者不能裁决。 +- 谁收编、谁认领、谁裁决只作为黑板事实存在,没有任何程序字段声称它们。 +- 台账的禁令仍然成立:没有新工具、没有专用频道、没有第二个任务体。 + +## 风险 + +- 一条约束都不声明时,宽的合法集合可能被读成"可以一起上"的许可,把"没商量同时并行"的问题换个地方重演。缓解:槽位预算仍是声明的上限,而不是社会约定。 +- 修复优先现在住在共享规划代码里当策略,所以把它变成声明式约束是一次改动,不只是改名;搬完之前,答案是在一条隐含约束下算出来的——这是本提案唯一还不成立的地方。 +- 理由一旦成为接口,解析理由字符串的调用者就造出了第二个事实源。理由是给人看和进日志的;调用者的任何决定都必须走认领或一个显式字段。 +- agent 自由选序之后,arm 的 equal-instrument 前提取决于每次 run 记录了自己启用了哪些约束;没记录约束集合,两次 run 就不可比。 +- 查询会招来轮询。轮询便宜,而"先问后动"仍可能撞车——认领仍是仲裁者,这正是设计意图。 From 4db614034e449018117c05f4542d760ffc05a770 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Mon, 21 Sep 2026 00:34:23 +0800 Subject: [PATCH 06/32] docs(decisions): read the legality record together with board governance The proposed record now relates to board governance and capability addressing (2026-09-06), and its Problem says what that record does not cover: the board's correctness line - reviewable finalize, authentic content, scope isolation, capability addressing, and the deliver/judge/attempt-fencing slice - governs how an entry is treated once it exists, not who decides what happens next. The two records are two halves of one protocol, and reading them together is what makes the distributed-systems half visible: claimTaskBoardEntry is a single atomic compare-and-set ("lease-based claiming ... on a losing CAS the failure is diagnosed against a fresh read"), lease expiry is enforced lazily with no background sweeper - that is a suspecting failure detector, not a death certificate - attempt fencing exists precisely because a lapsed lease cannot be trusted ("without this, work reassigned to another agent could still be read through a stale artifact"), and the deliverable/verdict pair binds an acceptance to an artifact identity. Idempotent acknowledgements and the never-recycled monotonic identity counter are the other two primitives already in place. No prose in either record names this, so every discussion of the boundary re-derives it. --- .../proposed/2026-09-20-the-program-answers-legality.md | 7 +++++-- .../2026-09-20-the-program-answers-legality.zh-CN.md | 4 ++-- 2 files changed, 7 insertions(+), 4 deletions(-) diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md index bb8e95e8..88344263 100644 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md @@ -3,7 +3,7 @@ [中文](2026-09-20-the-program-answers-legality.zh-CN.md) **Status:** proposed -**Relates to:** [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Task unit semantics](../../design/task-unit-semantics.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md) +**Relates to:** [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Task unit semantics](../../design/task-unit-semantics.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md) ## Problem @@ -18,7 +18,10 @@ ordering. [The obligations ledger](../../design/task-unit-semantics-obligations.md) says where the boundary must not move - no second task body, no dedicated tool, no dedicated channel - but it does not say what the -program's answer is. +program's answer is. Neither does the board's correctness line - reviewable finalize, authentic +content, scope isolation, capability addressing ([board governance and capability +addressing](../implemented/2026-09-06-board-governance-addressing.md)) - which governs how an entry is +treated once it exists, not who decides what happens next. ## Proposal diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md index 698b6f1c..4e378a8e 100644 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md @@ -3,13 +3,13 @@ [English](2026-09-20-the-program-answers-legality.md) **Status:** proposed -**Relates to:** [给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md)、[契约的义务](../../design/task-unit-semantics-obligations.md) +**Relates to:** [黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md)、[契约的义务](../../design/task-unit-semantics-obligations.md) ## 问题 元语有了名字,共享层也已经能算出合法集合:有序合法集合、按声明槽位预算切一刀、认领被拒时给出每个单元的理由。但**没有任何一份文档在说"每一步是谁的决定"**。这个缺口有可观察的代价:任何想驱动一次 run 的调用者都得自己重新划一遍边界,于是两种已被记录在案的失败模式都回来了——程序顺手替 agent 挑人(也就是命名决策里记为"这个模型刻意没有变成"的那个词:调度器),以及 agent 不知道程序到底给什么保证,只能用唯一能用的说法去要它真正需要的东西:更多槽位、更怪的顺序。 -[义务台账](../../design/task-unit-semantics-obligations.md)写明了边界不能往哪边动——不得再有第二个任务体、专用工具、专用频道——但它没有回答程序给出的那个答案是什么。 +[义务台账](../../design/task-unit-semantics-obligations.md)写明了边界不能往哪边动——不得再有第二个任务体、专用工具、专用频道——但它没有回答程序给出的那个答案是什么。黑板那条正确性线也没有:可复核的终结、内容真实性、作用域隔离、能力寻址([黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md))管的是一个条目在存在之后怎么被对待,不是下一步该由谁定。 ## 提案 From 9832bf6a4c28c2cf4523afddc73932af2869877f Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Mon, 21 Sep 2026 00:40:41 +0800 Subject: [PATCH 07/32] docs(design): inventory the umbrella - its parts, the gaps, and the distributed concepts behind them Your read was that part of this is a distributed-systems problem, and reading the legality record together with board governance showed the framing has no home: the governance record is an implemented decision with a fixed format, and the 2026-08-13 document is board mechanics. This adds the inventory as a design document, and the naming record now points at it from its history section. It lists fifteen parts with owner and state: entries and addressing, identity and authenticity, ownership and time, termination integrity, attention and wake, scope and visibility, truth versus coordination state, the task unit, legality and admission, ordering and constraints, budget and accounting, cancellation and fencing, recovery and replay, handoff, and evidence. The task-unit protocol is the complete one, and no single part completes a task: a unit's declaration, ownership and time, termination integrity, attention, legality, budget, truth separation and handoff have to hold at once. It then names seven gaps, each with the concept that fits it rather than a new vocabulary: the adopter has no clause (a commit boundary, and orphan reaping for the restart responsibility), legality has no readable answer (admission control with a read-only precondition check), constraints have no declaration (declarative policy, plus deterministic replay for comparability), read guarantees are unwritten (monotonic reads, read-your-writes, causal consistency across one agent's sessions), disagreement has no arbiter beyond veto and racing a claim (optimistic concurrency with conflict detection, or a plain single-writer decision since one arbiter already exists), cross-run resources have no mutual exclusion (a fenced mutex, not only the slot budget and the throwaway-worktree practice), and in-doubt work has no resolution (an in-doubt transaction: the fact is known, the outcome is not). It also records what is deliberately not needed - consensus, leader election, quorum, two-phase commit, exactly-once delivery, distributed transactions, because truth lives in one store written by one daemon - and three analogies that mislead: a lease expiry is a failure detector's suspicion and not a death certificate, which is why attempt fencing must stay; the board is not a message queue, so re-reading is normal and re-doing is the problem; and agents are not replicas, their independent judgement being exactly why a delivery is recorded with a judge instead of a self-report trusted. --- ...6-09-20-name-the-collaboration-protocol.md | 4 + ...0-name-the-collaboration-protocol.zh-CN.md | 2 + .../design/protocol-governed-collaboration.md | 91 +++++++++++++++++++ .../protocol-governed-collaboration.zh-CN.md | 55 +++++++++++ 4 files changed, 152 insertions(+) create mode 100644 docs/design/protocol-governed-collaboration.md create mode 100644 docs/design/protocol-governed-collaboration.zh-CN.md diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md index 90c63b23..536b835e 100644 --- a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.md @@ -92,6 +92,10 @@ decomposition is the part a task's semantics had to supply. It also explains the rejected above: a scheduler is what this model deliberately did not become, OoO is the borrowing's own name, and fusion is one measured mechanism inside it rather than the whole. +The umbrella's parts are inventoried separately - which sub-protocols exist, which are complete, which +are still missing, and which of them already have a name in distributed systems: +[protocol-governed collaboration](../../design/protocol-governed-collaboration.md). + ## Alternatives considered - **Keep saying OoO.** Rejected as the umbrella: `ooo` is the project's own name for the scheduling diff --git a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md index c4324ab5..db2c2b12 100644 --- a/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md +++ b/docs/decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md @@ -80,6 +80,8 @@ entry)** 是条目被某次 run 治理之后的状态;**run** 是计划、 这个顺序就是两个名字的来历:大类叫“协议化协作”,因为黑板在成为执行面之前就已经是一份协议;子协议叫“任务单元协议”,因为更细的拆分是任务语义必须供上的那一半。它也解释了上面被记为拒绝的那些词:调度器是这个模型刻意没有变成的东西,OoO 是借来的那个思想自身的名字,而融合是它内部一个已实测的机制,不是整体。 +大类有哪些组成部分另有一份清单:哪些子协议已有、哪些是齐的、哪些还缺,以及其中哪些在分布式领域已经有名字:[协议化协作](../../design/protocol-governed-collaboration.zh-CN.md)。 + ## 考虑过的替代方案 - **继续叫 OoO。** 作为大类被拒:`ooo` 是项目给自己实现的调度模型起的名字、在 100 多份文档里承重, diff --git a/docs/design/protocol-governed-collaboration.md b/docs/design/protocol-governed-collaboration.md new file mode 100644 index 00000000..fcc43069 --- /dev/null +++ b/docs/design/protocol-governed-collaboration.md @@ -0,0 +1,91 @@ +# Protocol-governed collaboration: the parts, the gaps, and the distributed systems behind them + +[中文](protocol-governed-collaboration.zh-CN.md) + +The umbrella name is [protocol-governed collaboration](../decisions/implemented/2026-09-20-name-the-collaboration-protocol.md). +This document is an inventory of it: which sub-protocols exist, which one is complete, which are missing, +and which of those - existing or missing - already have a name in distributed systems. + +It is an index, not a second specification. Each part's owner is named in its row, and where this +document and an owner disagree, the owner wins. + +## What the umbrella is made of + +| Part | What it governs | State | Owner | +| ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| Entries and addressing | Typed, attributed, expiring entries; compact reads, cursors, inbox; directed delivery (`to=`), discovery, capability addressing (`need`) | Complete | [board governance and capability addressing](../decisions/implemented/2026-09-06-board-governance-addressing.md), [board find/direct and serial](board-find-serial-a2a-compat-2026-08-13.md) | +| Identity and authenticity | Writer attribution, content hash by default, signing across a trust boundary | Complete | [board governance and capability addressing](../decisions/implemented/2026-09-06-board-governance-addressing.md) | +| Ownership and time | Atomic compare-and-set claim, lease with lazy expiry, renewal, release, serial promotion of the next pending entry | Mechanism complete; the responsibility to restart work is not assigned | `src/core/store/base.ts`, [board find/direct and serial](board-find-serial-a2a-compat-2026-08-13.md) | +| Termination integrity | Reviewable finalize and veto, `deliver` with a digest, independent `judge`, attempt fencing, `undecidable` as a first-class outcome | Complete | [board governance and capability addressing](../decisions/implemented/2026-09-06-board-governance-addressing.md) | +| Attention and wake | Directed wake of one session, membership by subscription, silent notes, acknowledgement suppressing re-notification | Complete | [board find/direct and serial](board-find-serial-a2a-compat-2026-08-13.md) | +| Scope and visibility | An entry targets an authorized agent subset, so boards are not all-to-all | Complete | [board governance and capability addressing](../decisions/implemented/2026-09-06-board-governance-addressing.md) | +| Truth versus coordination state | The board is a temporary coordination medium; durable memory is the truth store; `memory=` pointers | Complete | [board governance](../decisions/implemented/2026-09-06-board-governance-addressing.md), [memory graphs](memory-graphs.md) | +| The task unit | A unit declared by inputs, dependencies, acceptance, capability and budget; decomposition; the internal compile view; the run, its adoption, and the facts the runtime owns | Complete | [task unit semantics](task-unit-semantics.md), [the obligations ledger](task-unit-semantics-obligations.md) | +| Legality and admission | The ordered legal set, the cut to the declared slot budget, the reason each unit is or is not legal | Mechanism complete; the readable answer is missing | `src/integration/ooo-board.ts`, [the program answers legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md) | +| Ordering and constraints | Determinism (one plan plus one set of facts yields one answer), repair-first, tie-breaking by plan order | Mechanism complete; the declaration is missing | [fusion planning: repair-first](../decisions/implemented/2026-09-19-fusion-planning-repair-first.md), [the program answers legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md) | +| Budget and accounting | Declared slot budget, token and cache accounting, the offline cost ceiling | Complete | [declared slot budget](../decisions/implemented/2026-09-18-declared-slot-budget.md), [the cost model](../experiments/execution/ooo-cost-model-2026-09-17.md) | +| Cancellation and fencing | Explicit cancellation, no orphan worker or check after it, late artifacts fenced | Complete | [long checks run detached](../decisions/implemented/2026-09-18-detached-long-checks.md), [the bootstrap design](ooo-execution-bootstrap.md) | +| Recovery and replay | Log plus replay of a round, the transactional outbox drained after commit and on restart, no stealing another's work after a crash | Complete inside the arms; the product side is not wired | [the bootstrap design](ooo-execution-bootstrap.md) | +| Handoff | Session identity comes from the board; a continuation surface carries what the next holder needs | Complete | [session identity comes from the board](../decisions/implemented/2026-09-19-session-identity-comes-from-the-board.md), [session continuation obligations](task-unit-semantics-obligations.md) | +| Evidence and audit | Run facts, verification evidence, traceability aggregation | Complete | [rtm evidence aggregation](../decisions/implemented/2026-09-20-rtm-evidence-aggregation.md) | + +No single part completes a task. A task completes when the unit's declaration, ownership and time, +termination integrity, attention, legality, budget, truth separation and handoff all hold at once; the +rows above are the ones a given piece of work cannot do without. + +## The gaps, and the concepts available for them + +1. **The adopter has no protocol clause.** Who may adopt a discussion into a run, and who is responsible + for restarting work after a lease lapses, is not stated. The mechanism exists (`expiry`, serial + promotion, the dispatch loop), the responsibility does not. Distributed systems call the first one a + **commit boundary** - the single point that makes a discussed arrangement binding - and the second an + **orphan reaper / restart responsibility**. +2. **Legality has no readable answer.** A caller can claim and be refused, but cannot ask what is legal + and why. That is **admission control** with a **read-only precondition check**. +3. **Constraints have no declaration.** Repair-first currently lives as policy inside shared planning + code. Making preferences named constraints a plan enables is **declarative policy**, and the + requirement that one plan plus one set of facts yields one answer is **deterministic replay**. +4. **Read guarantees are unwritten.** Board reads are cursor-based and may lag; nothing states whether a + reader may assume monotonic reads or read-your-writes. The concepts are **monotonic reads**, + **read-your-writes** and, across one agent's own sequence of sessions, **causal consistency**. +5. **Disagreement has no arbiter.** Two agents that disagree about what to do next can only contest a + completion (`veto`) or race a claim. There is no clause for settling the disagreement itself. The + closest concepts are **optimistic concurrency with conflict detection** and, because one arbiter + already exists, a plain **single-writer decision**. Competitive allocation was already refused + ([board governance](../decisions/implemented/2026-09-06-board-governance-addressing.md)). +6. **Cross-run resources have no mutual exclusion.** Several runs competing for one worktree or one file + are handled by practice (a throwaway worktree per arm) and by the slot budget, not by a rule. The + concept is a **fenced mutex** - a lock whose holder cannot be believed after its lease lapses. +7. **In-doubt work has no resolution.** A delivery that nobody judged is `undecidable` when the run ends, + but an artifact left in doubt across runs has no clause. This is precisely an **in-doubt transaction**: + the fact is known, the outcome is not, and somebody has to resolve it or write down why not. + +## The concepts this system deliberately does not need + +- **Consensus, leader election, quorum, two-phase commit.** Truth lives in one store written by one + daemon, so there are no replicas to agree; the only fact needing agreement is who holds a claim, and a + single atomic compare-and-set provides it. +- **Exactly-once delivery.** What is delivered is an artifact plus a judgement, not a message. Repeating + work costs tokens and wall time, and that difference from a CPU, where replay is nearly free, is why + speculation measured as cost without gain in [the arms](../experiments/execution/archive/ooo-arms-2026-09-19/README.md). +- **Distributed transactions.** One run commits once; there is no atomic commit spanning parts. + +## Three analogies that mislead + +1. **Lease expiry is not death.** Expiry is a failure detector, and a detector reports a suspicion, never + a fact. The system already answers this correctly with attempt fencing - a new holder starts attempt + N+1, which voids the previous attempt's artifact and verdict - and that property must not be traded + away for convenience. +2. **The board is not a message queue.** It is a shared, attributed, expiring entry space: entries persist + and are re-read, so re-reading is normal and re-doing is the problem. Treating it as a queue would + make every reader a consumer and every re-read a duplicate. +3. **Agents are not replicas.** They do not agree by construction, they each judge independently, and + that independence is the point - it is why the board records what was delivered and who judged it + rather than trusting a self-report. + +## What would make this inventory complete + +Rows 9 and 10 are the interfaces the legality proposal covers: the readable answer and the declared +constraint set. Rows 1, 4, 5, 6 and 7 are gaps with no owner yet. Until each gap either gets a clause or +is written down as deliberately absent, "complete" for this umbrella means the parts that have owners, +not the parts a task needs. diff --git a/docs/design/protocol-governed-collaboration.zh-CN.md b/docs/design/protocol-governed-collaboration.zh-CN.md new file mode 100644 index 00000000..480178eb --- /dev/null +++ b/docs/design/protocol-governed-collaboration.zh-CN.md @@ -0,0 +1,55 @@ +# 协议化协作:组成部分、空缺,以及它们背后的分布式概念 + +[English](protocol-governed-collaboration.md) + +大类的名字是[协议化协作](../decisions/implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)。本文是它的一份清单:有哪些子协议、哪一个是全的、还缺哪些,以及这些(已有的和缺的)在分布式领域里有哪些已经有了名字。 + +它是一份索引,不是第二份规格。每一部分的归属写在它那一行里;本文与归属方不一致时,以归属方为准。 + +## 大类由什么组成 + +| 部分 | 管什么 | 状态 | 归属 | +| ------------------ | -------------------------------------------------------------------------------------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 条目与寻址 | 有类型、有归因、会过期的条目;紧凑读、游标、收件箱;定向投递(`to=`)、发现、能力寻址(`need`) | 完整 | [黑板治理与能力寻址](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[黑板先广播后定向与串行](board-find-serial-a2a-compat-2026-08-13.md) | +| 身份与真实性 | 写入者归因、默认内容哈希、跨信任边界签名 | 完整 | [黑板治理与能力寻址](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md) | +| 拥有与时间 | 原子比较交换认领、租约与惰性过期、续租、释放、下一个待处理条目的串行晋升 | 机制完整;"谁来重活"这份责任没有归属 | `src/core/store/base.ts`、[黑板先广播后定向与串行](board-find-serial-a2a-compat-2026-08-13.md) | +| 终止完整性 | 可复核的终结与否决、带 digest 的 `deliver`、独立 `judge`、attempt 围栏、`undecidable` 作为一等结果 | 完整 | [黑板治理与能力寻址](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md) | +| 注意力与唤醒 | 只唤醒被点名的一个会话、按订阅确定成员、note 静默、ack 抑制重通知 | 完整 | [黑板先广播后定向与串行](board-find-serial-a2a-compat-2026-08-13.md) | +| 作用域与可见性 | 条目的可见范围是授权过的 agent 子集,黑板不是 all-to-all | 完整 | [黑板治理与能力寻址](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md) | +| 真值与协调状态分离 | 黑板是临时协调面,真值在 durable memory,`memory=` 做指针 | 完整 | [黑板治理](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[记忆图](memory-graphs.md) | +| 任务单元 | 由输入、依赖、验收、能力、预算声明的单元;拆分;内部编译视图;run、它的收编、以及运行时拥有的事实 | 完整 | [任务单元语义](task-unit-semantics.md)、[义务台账](task-unit-semantics-obligations.md) | +| 合法性与准入 | 有序合法集合、按声明槽位预算切一刀、每个单元合法与否的理由 | 机制完整;可读的答案缺失 | `src/integration/ooo-board.ts`、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md) | +| 顺序与约束 | 确定性(同一计划加同一组事实只产出一个答案)、修复优先、按计划顺序破平 | 机制完整;声明方式缺失 | [融合规划的修复优先](../decisions/implemented/2026-09-19-fusion-planning-repair-first.zh-CN.md)、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md) | +| 预算与计量 | 声明槽位预算、token 与缓存计量、离线成本上限 | 完整 | [声明式槽位预算](../decisions/implemented/2026-09-18-declared-slot-budget.zh-CN.md)、[成本模型](../experiments/execution/ooo-cost-model-2026-09-17.md) | +| 取消与围栏 | 显式取消、取消后不留孤儿 worker 或检查、迟到产物被围栏挡住 | 完整 | [长检查分离运行](../decisions/implemented/2026-09-18-detached-long-checks.zh-CN.md)、[自举设计](ooo-execution-bootstrap.md) | +| 恢复与重放 | 轮次的日志加重放、事务性 outbox 在提交后与重启时排空、崩溃后不偷别人的活 | 臂内完整;产品面尚未接线 | [自举设计](ooo-execution-bootstrap.md) | +| 交接 | 会话身份来自黑板;延续面把下一位持有者需要的东西带过去 | 完整 | [会话身份来自黑板](../decisions/implemented/2026-09-19-session-identity-comes-from-the-board.zh-CN.md)、[会话延续义务](task-unit-semantics-obligations.md) | +| 证据与审计 | run 事实、验证证据、可追溯性汇总 | 完整 | [rtm 证据汇总](../decisions/implemented/2026-09-20-rtm-evidence-aggregation.zh-CN.md) | + +没有任何单独一部分能完成任务。一次任务完成,需要单元的声明、拥有与时间、终止完整性、注意力、合法性、预算、真值分离、交接**同时**成立;上面这些行就是一件活不能缺的那几样。 + +## 空缺,以及可以为它们借用的概念 + +1. **收编者没有协议条款。** 谁有资格把一次讨论收编成一次 run,以及租约失效后谁负责把活重新推出去,都没有写。机制是有的(`expiry`、串行晋升、派发循环),责任没有。分布式里把前者叫做**提交边界**——让一份商定的安排生效的那个唯一点;把后者叫做**孤儿回收/重启责任**。 +2. **合法性没有可读的答案。** 调用者可以去认领然后被拒,但无法问"现在什么合法、为什么"。这是**准入控制**加一次**只读的前置条件检查**。 +3. **约束没有声明方式。** 修复优先目前活在共享规划代码里当策略。把偏好变成计划可启用的具名约束,是**声明式策略**;而"同一计划加同一组事实只产出一个答案"这个要求,是**确定性重放**。 +4. **读的保证没有写下来。** 黑板读基于游标、可能落后;没有一处说明读者是否可以假设单调读或读己所写。概念是**单调读**、**读己所写**,以及跨同一个 agent 自己一串会话的**因果一致性**。 +5. **分歧没有仲裁者。** 两个 agent 对下一步做什么意见不合时,只能争夺一次完成(`veto`)或抢认领。对分歧本身没有条款。最近的概念是**乐观并发加冲突检测**,以及——因为仲裁者本来就存在——干脆的**单写者决定**。竞价式分配早就被拒绝了([黑板治理](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md))。 +6. **跨 run 的资源没有互斥。** 多个 run 抢同一个工作树或同一个文件,现在靠实践(每条臂一个一次性 worktree)和槽位预算应付,没有规则。概念是**带围栏的互斥锁**——租约失效之后,它的持有者就不能再被相信。 +7. **悬而未决的活没有归属。** 没人裁决的交付在 run 结束时算 `undecidable`,但跨 run 之后仍悬着的产物没有条款。这正是一个**存疑事务(in-doubt transaction)**:事实已知、结果未知,必须有人解掉它,或者写下为什么不。 + +## 这套系统刻意不需要的概念 + +- **共识、领导者选举、quorum、两阶段提交。** 真值在一个由唯一 daemon 写入的 store 里,没有副本需要达成一致;唯一需要一致的只有"谁持有认领",而一次原子比较交换就给了它。 +- **恰好一次投递。** 交付的是产物加一次判断,不是消息。重复做活的代价是 token 和墙钟时间,而这个与 CPU(重放几乎免费)的差别,正是[臂的实测](../experiments/execution/archive/ooo-arms-2026-09-19/README.md)里推测呈现为"成本真实、收益为零"的原因。 +- **分布式事务。** 一次 run 只提交一次,没有跨部分的原子提交。 + +## 三个会误导的类比 + +1. **租约过期不是死亡。** 过期是故障检测器,而检测器给出的是疑似,从不是事实。系统已经用 attempt 围栏正确回答了这一点——新的持有者从第 N+1 次尝试开始,作废上一代的产物与裁决——这条性质不能为了方便让掉。 +2. **黑板不是消息队列。** 它是一个共享的、带归因的、会过期的条目空间:条目留着、会被重读,所以重读是常态,重做才是问题。把它当队列,会让每个读者都变成消费者、每次重读都变成重复。 +3. **agent 不是副本。** 他们不靠构造达成一致,各自独立判断,而这种独立正是关键——所以黑板记录"交付了什么、谁判的",而不是相信自报。 + +## 这份清单要怎样才算完整 + +第 9、10 行就是合法性提案覆盖的两个接口:可读的答案与可声明的约束集合。第 1、4、5、6、7 行是还没有主的空缺。在每个空缺或者拿到条款、或者被明确写成"刻意没有"之前,这个大类的"完整"指的只是那些有归属的部分,而不是一件任务真正需要的那几样。 From d93ae396ac0d321435c3cb69eecece646d01dac4 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Mon, 21 Sep 2026 22:33:19 +0800 Subject: [PATCH 08/32] docs(decisions): the frame, its data format, and its storage The primitives have names and the program's job is proposed, but the frame itself was undefined: which fields exist, how values are encoded, where they live in the store, and how a peer protocol sits beside the Task-Unit Protocol. Researching the finished specifications changed the shape of the problem - every layer but one is already defined by somebody - so this record chooses and names rather than invents. What it borrows, per layer: CloudEvents 1.0 for the envelope (only id/source/specversion/type required, context attributes inspected without deserializing the event data, lowercase names of at most 20 characters, data reserved); A2A 1.0.0 for the task model (terminal states, contextId, cursor paging chosen explicitly over offsets, historyLength, includeArtifacts false by default); NLIP/ECMA-430 for the payload discriminator and the progressive shape (messagetype/format/subformat/content/submessages/label, the first submessage inlined because most messages carry none, structured with subformat uri carrying a URI instead of bytes); MCP 2025-06-18 for LLM-facing results (content blocks with annotations, structured content beside text, isError so a failure is data); RFC 8941 for a one-line text type system; FIPA-ACL for the act and conversation fields; RFC 6709 for the unknown-field choice; ECMA-434 for the security layer. The one layer nobody defines is ownership, lease, fencing and legality, because every agent protocol assumes a task belongs to one agent. Decisions it records: the frame is a state record and not a message, so FIPA's performative is a rendering of existing columns rather than a column; the header stays in columns because the claim is one atomic compare-and-set UPDATE and a header inside JSON would turn that into read-modify-write; the payload is one opaque document, with the invariant that every payload can be NULL and T0/T1 still render while claim, wake, expiry and compact read still work (verified on the store's SQLite 3.53.3); a payload field a protocol wants to query is indexed by a generated column plus an expression index rather than a new board column, which is what keeps the schema from growing with the number of protocols; the protocol is a value with a version, a peer protocol supplies declaration validation, legality, projection and acceptance, and it plugs into the existing DispatchBoard and RoundQueryPort rather than into the board; unknown protocol means readable and not actionable, and unknown fields are decided by the protocol's own version rule, stated rather than assumed. Data format, in three separate questions: JSON on the existing RPC wire (protobuf-as-normative-source is not worth it with one binding); flat key/value for T0/T1 and natural language for negotiation, because format restrictions measurably degrade reasoning and escaping is a failure source, with tool schemas shaped for strict mode (additionalProperties false, every property required, absence as nullable); typed columns plus one JSON payload document in the store. TOON and CBOR are deliberately not adopted: TOON's saving applies to flat uniform data pushed into context, which is what pointer-first avoids, and CBOR waits for a transport with a bandwidth constraint. Security is the layer that is empty. ECMA-434 is mandatory for NLIP conformance and its companion guidelines score fifteen threats with prompt injection highest; two rules are adopted as text now (an in-band session token is a session credential and must not reach the model, and a caller's token is never forwarded), and a threat list of our own is owed. The scope ceiling is stated so the change cannot grow: one group of nullable columns, one boundary adapter, one rendering rule, one test. Anything needing a new table, a new tool or a new channel is outside this decision. --- .../2026-09-21-the-frame-and-its-storage.md | 182 ++++++++++++++++++ ...6-09-21-the-frame-and-its-storage.zh-CN.md | 120 ++++++++++++ 2 files changed, 302 insertions(+) create mode 100644 docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md create mode 100644 docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md new file mode 100644 index 00000000..49a50676 --- /dev/null +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md @@ -0,0 +1,182 @@ +# The frame, its data format, and its storage + +[中文](2026-09-21-the-frame-and-its-storage.zh-CN.md) + +**Status:** proposed +**Relates to:** [The program answers legality](2026-09-20-the-program-answers-legality.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Task unit semantics](../../design/task-unit-semantics.md) + +## Problem + +The primitives are named, the umbrella is inventoried, and the program's job - answer legality - is +proposed. What is still undefined is the frame itself: which fields exist, how their values are encoded, +where they live in the store, and how a **peer** protocol sits beside the Task-Unit Protocol. Left open, +every addition re-opens the same questions, and the natural answers are the expensive ones: parse the +payload in order to route, add a column per protocol, invent a fresh vocabulary per protocol. + +A survey of finished specifications changes the shape of the problem. Every layer but one is already +defined by somebody, so the work is choosing and naming, not inventing: + +| Layer | Already defined by | +| ----------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Event envelope: identity, time, version, content type, mode | CloudEvents 1.0 - required are only `id`/`source`/`specversion`/`type`; context attributes are "inspected at the destination without having to deserialize the event data"; attribute names are lowercase ASCII, at most 20 characters, never `data`; the type set is closed | +| Task, parts, states, paging | A2A 1.0.0 - terminal-state semantics, `contextId`, cursor paging (`pageToken`/`nextPageToken`, chosen explicitly over offsets), `historyLength`, and `includeArtifacts` defaulting false "to reduce payload size" | +| Payload discriminator and progressive shape | NLIP / ECMA-430 - `messagetype`, `format`, `subformat`, `content`, `submessages[]`, `label`; the first submessage is inlined at the top because most messages carry none; `format: structured` with `subformat: uri` carries a URI instead of bytes | +| LLM-facing results and errors | MCP 2025-06-18 - content blocks with `annotations.audience`, `structuredContent` beside text, and `isError` in the result, so a failure is data | +| One-line text encoding | RFC 8941 Structured Fields - List/Dictionary/Item with Parameters, intentionally strict | +| Acts and conversation threading | FIPA-ACL - `performative` is the only required field, beside `protocol`, `conversation-id`, `in-reply-to`, `reply-by` | +| Unknown-field policy | RFC 6709 - ignoring and refusing by name are both legitimate; the choice has to be stated | +| Security | ECMA-434 and the NLIP security guidelines - mandatory for NLIP conformance, with a scored threat list | +| **Ownership, lease, fencing, legality** | **Nobody.** Every agent protocol assumes a task belongs to one agent, so "who may hold it" does not arise | + +## Proposal + +### 1. The frame + +The frame is a board entry, and it is a **state record, not a message**. That distinction decides the +field set: the acts are already state columns (a claim, a delivery, a verdict, a resolution), so FIPA's +`performative` is a rendering of those columns rather than a column of its own. + +- **Header columns** (the board acts on these alone): `id`, `channel`, `kind` (the seven collaboration + kinds, unchanged), `state`, `author`, `source_session`, `to`, `need`, `scope`, `created_at`, + `expires_at`, `lease{holder, until}`, `attempt`, `digest`, `verdict`, `acked`, `resolution`, plus the + protocol selectors `protocol` and `protocol_version`. +- **One payload document**: whatever the named protocol needs, opaque to the board. +- **Rendering, not storage**: the act an entry shows (FIPA's `performative`), and the progressive tiers + T0 header, T1 summary, T2 structure, T3 truth. + +### 2. The data format, three separate questions + +- **Wire bytes.** JSON over the daemon's existing RPC. Unchanged; no second binding exists to justify + code generation, so A2A's protobuf-as-normative-source answer is not adopted yet. +- **LLM tokens.** T0 and T1 render as flat key/value lines, not as JSON with braces and quotes. The + negotiation text stays natural language: forcing a model to emit JSON for prose costs reasoning + (format constraints measurably degrade it, and stricter constraints degrade it more), and escaping is + a failure source. Structured content is never inlined in the body - it is reached through the + discriminator and a pointer. An agent's own writes are tool arguments validated by the harness, and + those schemas follow strict mode: `additionalProperties: false` on every object, every property listed + in `required`, absence expressed as nullable rather than by omission. +- **Storage.** Typed columns for the header, one JSON document for the payload (section 3). + +Deliberately not adopted at this layer: TOON (its 30-60% saving applies to flat uniform data that has to +be pushed into context, which is exactly what pointer-first avoids), CBOR with a JSON fallback (ECMA-432 +deferred until a transport with a bandwidth constraint exists), and protobuf as a normative source. + +### 3. Storage layout and its invariant + +The header stays in columns because the claim is one atomic compare-and-set `UPDATE`. A header inside a +JSON document would turn atomic claiming into read-modify-write. That is a mechanism constraint, not a +preference. + +The payload is one document beside those columns, and the board never parses it. The store already keeps +this shape: `claims_json`, `markers_json`, `scope_json`, `evidence_ids_json` and others sit beside typed +columns in the same tables. + +**Invariant (testable):** with `payload` set to NULL for every row, T0 and T1 still render and claim, +wake, expiry and compact read still work. Verified against the store's own SQLite (3.53.3). + +When a protocol needs to query a field inside its own payload, it does not ask the board for a column: a +generated column plus an expression index answers it (also verified on the same SQLite). This is what +keeps the schema from growing with the number of protocols. + +Cost, stated plainly: JSON columns give up database-level typing, a trade the store already makes, and +cross-protocol queries depend on generated columns or the protocol's own tables. + +### 4. Peer protocols and replacement + +The protocol is a **value**, not a schema. The Task-Unit Protocol is one value; a peer protocol is +another, and the board's coordination layer stays protocol-agnostic (envelope, addressing, time and +lease, claim and fencing, delivery digest and verdict, wake, scope, retention). + +A peer protocol supplies four narrow things: + +1. **Declaration validation** - refuse by name when a field it cannot map is missing, never downgrade it + to free text (the semantics design already states this rule). +2. **Legality** - the definition of which units are legal under the current facts. This is the party that + defines the answer "the program answers legality" is about. +3. **Projection** - the T1 summary line and the T2 structure. Required because the payload is opaque to + the board, and the shape MCP and A2A already use for the same reason. +4. **Acceptance** - the criteria for accepting a delivery. The form (digest plus independent verdict) + does not change. + +The insertion points already exist: `DispatchBoard` and `RoundQueryPort`. A protocol plugs in on the +semantics side, and replacing one means another implementation plus another `protocol` value; the payloads +of entries written under the old value stay exactly as they are - readable, not actionable, and no +migration. + +Unknown handling is stated rather than assumed: an unknown `protocol` makes an entry readable and not +actionable, and an unknown field inside a known protocol is decided by that protocol's version rule, +which must say whether it is ignored or refused by name. + +### 5. Security, the layer that is empty + +ECMA-434 is a conformance requirement for NLIP, and its companion security guidelines score fifteen +threats - prompt injection highest, then confused-deputy and token passthrough, session hijack, memory +injection. Two rules are adoptable now as text: + +- An in-band session token is a session credential: it must not be handed to the agent's model or exposed + in application-level payloads. The board's session identity is in-band and the board renders into model + context, so this rule applies to us directly. +- A caller's token is never forwarded; a scoped token is minted instead. This becomes load-bearing only + when entries cross a trust boundary, which today they do not. + +What is owed is a threat list of our own, with the ones we accept and the ones we do not defend. + +### 6. Scope ceiling + +The change surface is one group of nullable columns, one boundary adapter, one rendering rule and one +test. Anything that needs a new table, a new tool or a new channel is outside this decision. + +## Plan + +1. This record. No code. +2. One migration: `protocol`, `protocol_version`, `payload_format`, `payload`, plus the invariant test. +3. Refusals become data at the RPC and tool boundary; the store keeps throwing. +4. Tool schemas shaped for strict mode; flat T0/T1 rendering. +5. A threat list, with the two rules carried as text until a trust boundary exists. + +## Alternatives considered + +- **One JSON document for the whole entry, header included.** Rejected: atomic claiming becomes + read-modify-write, and every reader would have to parse to find out whether an entry is claimable. +- **A column group per protocol.** Rejected: the schema grows with the number of protocols and every new + protocol has to touch the board, which is the rework this record exists to avoid. +- **protobuf as the normative source with generated bindings** (A2A's answer). Rejected for now: there is + one wire and no second binding, so code generation buys nothing yet. +- **A single-line RFC 8941 encoding for the header.** Rejected for storage - our header lives in columns + - but kept as an option if a compact textual read is ever needed. +- **TOON.** Rejected: the saving appears on flat uniform data pushed into context, and the frame is flat + with pointers instead. +- **CBOR with a JSON text fallback.** Deferred until a transport with a bandwidth constraint exists. +- **Storing the structured payload as JSON inside the entry body.** Rejected: escaping failures, the + reasoning cost of format constraints, and MCP's own note that structured content is a different thing + from schema-constrained model generation. + +## Acceptance criteria + +- With every `payload` NULL, the board renders T0/T1 and completes claim, wake, expiry and compact read. +- Adding a peer protocol changes nothing in the board's claim, wake or expiry paths, demonstrated by a + second protocol value with its own payload shape and its own projection. +- An entry whose protocol the store does not know is readable, is not actionable, and its payload is never + parsed by the board. +- A payload field a protocol wants to query is indexed by a generated column and an expression index, + with no new board column. +- Header field names follow the borrowed naming constraints: lowercase ASCII, at most 20 characters, and + `data` is not used as a name. +- A refusal reaches the caller as data with a reason, not as an exception, at the RPC and tool boundary. +- LLM-facing schemas are strict-mode shaped, and no negotiation text is required to be a JSON string. +- Every new field names the specification it came from, or records why none fitted. + +## Risks + +- One document per entry invites over-stuffing: the header can drift into the payload until the board has + to parse it again. The invariant test and the never-parse rule are the mitigation. +- A protocol that projects its own summary can summarize dishonestly. Mitigation: projection is separate + from the verdict, and only `accepted` may be depended on. +- Generated columns and expression indexes are SQLite-specific; a portable store would have to re-examine + them. +- Refusals-as-data changes the shape existing callers see, so the arms' drivers have to move in the same + slice. +- Strict-mode schemas make absence explicit, which costs tokens per call and can make a nullable field + read as mandatory. +- A borrowed vocabulary can ossify. If a borrowed name does not fit, deviating is allowed - and the + deviation is recorded rather than renamed silently. diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md new file mode 100644 index 00000000..168a2320 --- /dev/null +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md @@ -0,0 +1,120 @@ +# 帧、它的数据格式与它的存储 + +[English](2026-09-21-the-frame-and-its-storage.md) + +**Status:** proposed +**Relates to:** [程序只回答合法性](2026-09-20-the-program-answers-legality.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) + +## 问题 + +元语有了名字,大类有了清单,程序的职责(只回答合法性)也有了提案。还没定的是**帧本身**:有哪些字段、取值怎么编码、存在库里的哪里、以及一个**平级**协议怎样与任务单元协议并排。放着不定,每次新增都会重开同一批问题,而最自然的答案恰好都是贵的:为了让路由能动就得解析载荷、每加一个协议就加一列、每个协议各造一套词。 + +对已定规范的调研改变了问题的形状:除了一层之外,每一层都已经有人定过,所以这件事是**选择与命名**,不是发明: + +| 层 | 已由谁定 | +| ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| 事件信封:身份、时间、版本、内容类型、模式 | CloudEvents 1.0——必填只有 `id`/`source`/`specversion`/`type`;属性原话是"在目的地被检查而无需反序列化事件数据";属性名小写 ASCII、最多 20 字符、不得叫 `data`;类型集合封闭 | +| 任务、部件、状态、分页 | A2A 1.0.0——终态语义、`contextId`、游标分页(`pageToken`/`nextPageToken`,并明确不用 offset)、`historyLength`、`includeArtifacts` 默认 false("缩减载荷大小") | +| 载荷判别与渐进形状 | NLIP / ECMA-430——`messagetype`、`format`、`subformat`、`content`、`submessages[]`、`label`;首子消息平铺在顶层,因为多数消息不带子消息;`format: structured` 配 `subformat: uri` 时内容是一个 URI 而不是字节 | +| 面向 LLM 的结果与错误 | MCP 2025-06-18——内容块带 `annotations.audience`、`structuredContent` 与文本并存、`isError` 放在结果里,于是失败也是数据 | +| 单行文本编码 | RFC 8941 Structured Fields——List/Dictionary/Item 加 Parameters,故意严格 | +| 行为与对话线程 | FIPA-ACL——`performative` 是唯一必填,旁边是 `protocol`、`conversation-id`、`in-reply-to`、`reply-by` | +| 未知字段策略 | RFC 6709——忽略与按名字拒绝都合法,但**必须写明选了哪个** | +| 安全 | ECMA-434 与 NLIP 安全指南——对 NLIP 是符合性要求,并给出打分的威胁清单 | +| **拥有权、租约、围栏、合法性** | **没有人。** 所有 agent 协议都假设任务属于某一个 agent,"谁能持有"不成立 | + +## 提案 + +### 1. 帧 + +帧是一个黑板条目,而且它是**状态记录,不是消息**。这个区别决定了字段集:行为本身已经是状态列(一次认领、一次交付、一个裁决、一次了结),所以 FIPA 的 `performative` 是这些列的**渲染**,不是它自己的列。 + +- **头列**(板子只靠这些行动):`id`、`channel`、`kind`(七个协作用途,不变)、`state`、`author`、`source_session`、`to`、`need`、`scope`、`created_at`、`expires_at`、`lease{holder, until}`、`attempt`、`digest`、`verdict`、`acked`、`resolution`,加上协议选择子 `protocol` 与 `protocol_version`。 +- **一条载荷文档**:具名协议自己需要的东西,对板子不透明。 +- **渲染,不是存储**:条目呈现的行为(FIPA 的 `performative`),以及分层的 T0 头、T1 摘要、T2 结构、T3 真值。 + +### 2. 数据格式是三个分开的问题 + +- **线上字节**:JSON,走 daemon 现有的 RPC。不变;既然只有一条线而没有第二个绑定,就不值得为代码生成付费,所以 A2A 的"proto 作规范源"暂不采用。 +- **给 LLM 的 token**:T0 与 T1 渲染成扁平 key/value 行,不是带花括号和引号的 JSON。协商内容保持自然语言:逼模型为散文产出 JSON 会付出推理代价(格式约束会可测地降低推理,约束越严越低),而转义本身就是失败源。结构化内容永远不内联进正文——它经判别位加指针抵达。agent 自己写的东西是工具参数,由 harness 校验,而那些 schema 按 strict 模式写:每个 object 都 `additionalProperties: false`、每个属性都列进 `required`、缺省用 nullable 表达而不是省略。 +- **存储**:头进列,载荷进一条 JSON 文档(第 3 节)。 + +这一层刻意不采用:TOON(它 30–60% 的节省出现在"必须把扁平均匀数据推进上下文"的场景,而这正是指针优先要避开的)、CBOR 加 JSON 回退(ECMA-432 留到真出现有带宽约束的传输时再谈)、proto 作规范源。 + +### 3. 存储布局与它的不变量 + +头留在列里,因为认领是**一条原子比较交换的 `UPDATE`**。把头放进 JSON 文档会把原子认领变成读-改-写。这是机制约束,不是偏好。 + +载荷是与这些列并排的一条文档,板子从不解析它。这个形状库里已有:`claims_json`、`markers_json`、`scope_json`、`evidence_ids_json` 等,就是与 typed 列同表并存的。 + +**不变量(可测)**:把所有行的 `payload` 置为 NULL,T0 与 T1 仍渲染,认领、唤醒、过期、紧凑读仍工作。已在 store 自己的 SQLite(3.53.3)上验证。 + +当某个协议要查自己载荷里的字段时,它**不要**向板子要一列:生成列加表达式索引就够了(同样在该 SQLite 上验证过)。这一点正是让 schema 不随协议数量增长的机制。 + +代价说明白:JSON 列放弃数据库级类型检查(这是仓库早已做过的交换),跨协议查询依赖生成列或协议自己的表。 + +### 4. 平级协议与替换 + +协议是**值**,不是模式。任务单元协议是一个取值;平级协议是另一个取值,而板子的协调层保持协议无关(信封、寻址、时间与租约、认领与围栏、交付 digest 与裁决、唤醒、作用域、保留)。 + +一个平级协议要提供四样窄的东西: + +1. **声明校验**——它无法映射的字段缺失时按名字拒绝,绝不降级为自由文本(语义设计已经立了这条)。 +2. **合法性**——在当下事实上哪些单元合法。这就是"程序只回答合法性"里那个合法性的**定义方**。 +3. **投影**——T1 的一行摘要与 T2 的结构。必须有,因为载荷对板子不透明;MCP 与 A2A 出于同样理由也是这个形状。 +4. **验收**——接受一次交付的判据。形式(digest 加独立裁决)不变。 + +插入点已经存在:`DispatchBoard` 与 `RoundQueryPort`。协议接在语义侧,替换一个协议就是换一个实现加换一个 `protocol` 取值;旧取值下写入的载荷原样留着——可读、不可行动,且不需要迁移。 + +未知的处理要**写明**而不是默认:未知 `protocol` 使条目可读但不可行动;已知协议内的未知字段由该协议的版本规则决定,而那条规则必须说明是忽略还是按名字拒绝。 + +### 5. 安全,那个还空着的层 + +ECMA-434 对 NLIP 是符合性要求,配套安全指南给十五类威胁打分:prompt injection 最高,其次是 confused-deputy 与令牌透传、会话劫持、记忆投毒。两条现在就能以文字形式采纳: + +- 带内会话令牌就是会话凭据:不得交给 agent 的模型,也不得暴露在应用层载荷里。黑板的会话身份是带内的,而黑板会渲染进模型上下文,所以这条直接适用于我们。 +- 绝不转发调用方的令牌,改为铸一个有范围的令牌。这条只有在条目跨信任边界时才承重,而今天还没有那样的边界。 + +欠的是一份我们自己的威胁清单,写明接受哪些、不防御哪些。 + +### 6. 范围上限 + +改动面是一组可空列、一处边界适配、一条渲染规则、一条测试。**任何需要新表、新工具、新频道的东西都在本决策之外。** + +## 计划 + +1. 本文。不动代码。 +2. 一次迁移:`protocol`、`protocol_version`、`payload_format`、`payload`,加那条不变量测试。 +3. 拒绝在 RPC 与工具边界变成数据;store 继续抛异常。 +4. 工具 schema 按 strict 模式写;T0/T1 扁平渲染。 +5. 一份威胁清单,两条规则先以文字承载,等出现信任边界再谈强制。 + +## 考虑过的替代方案 + +- **整个条目一条 JSON 文档(含头)**。拒绝:原子认领会退化成读-改-写,而且每个读者都得先解析才知道条目能不能认领。 +- **每个协议一组列**。拒绝:schema 随协议数量增长,每个新协议都要动板子——那正是本文要避免的返工。 +- **proto 作规范源并生成绑定**(A2A 的答案)。暂时拒绝:只有一条线、没有第二个绑定,代码生成现在买不到东西。 +- **头用 RFC 8941 的单行编码**。存储上拒绝——我们的头活在列里——但若将来需要紧凑文本读,保留为选项。 +- **TOON**。拒绝:节省只出现在被推进上下文的扁平均匀数据上,而帧是扁平的且有指针。 +- **CBOR 加 JSON 文本回退**。留到真出现有带宽约束的传输。 +- **把结构化载荷以 JSON 塞进条目正文**。拒绝:转义失败、格式约束的推理代价,以及 MCP 自己写明的一点——结构化内容与"schema 约束的模型生成"是两回事。 + +## 验收标准 + +- 所有 `payload` 为 NULL 时,板子仍渲染 T0/T1,并仍能完成认领、唤醒、过期与紧凑读。 +- 增加一个平级协议不改变板子的认领、唤醒、过期路径;用一个自带载荷形状与投影的第二协议取值演示。 +- store 不认识的协议:条目可读、不可行动,且板子从不解析它的载荷。 +- 协议想查的载荷字段由生成列加表达式索引支撑,板子不加新列。 +- 头字段名遵守借来的命名约束:小写 ASCII、最多 20 字符、不使用 `data` 作为名字。 +- 拒绝在 RPC 与工具边界以带理由的数据抵达调用者,而不是异常。 +- 面向 LLM 的 schema 按 strict 模式成型,且没有任何协商文本被要求写成 JSON 字符串。 +- 每个新字段都指名它来自哪份规范,或者记录下为什么没有一份合适。 + +## 风险 + +- 每条一条文档会诱发"什么都往里塞":头可能慢慢漂进载荷,直到板子又得解析它。缓解不变量测试与"板子从不解析"这条规则。 +- 自己给自己做投影的协议可以说得不诚实。缓解:投影与裁决分离,且只有 `accepted` 可被依赖。 +- 生成列与表达式索引是 SQLite 特有的;换可移植存储时要重新审视。 +- "拒绝即数据"改变现有调用者看到的形状,所以臂的 driver 必须在同一个切片里一起动。 +- strict 模式的 schema 让缺省显式化,每次调用多花 token,而且可能让读者把可空字段读成必填。 +- 借来的词汇会僵化。若某个借来的名字不合身,允许偏离——但要**记录**偏离,而不是悄悄改名。 From 9e1041d7ff1d28d3c342d1b31efe0a3b1b38c737 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Mon, 21 Sep 2026 22:46:19 +0800 Subject: [PATCH 09/32] docs(decisions): answer the frame record's review - mapping, seam, two retractions Four review points, all accepted, plus the code checks behind them. 1. The field mapping was missing, and it changes what the invariant means. The record now carries an existing-to-proposed table: content stays the single source of the entry body (readTaskBoardPreviews projects taskBoardPreview(entry.content) - whitespace collapsed, 200 characters, a memory= body returned verbatim), so "payload NULL still renders T1" is a property of the board's read path and not something the storage format produces by itself; the frozen declaration was never a body, it lives in task_run_manifest and task_run_tasks reached by run_id and task_id, with task_run_facts' own payload column as the precedent; and the proposed payload carries nothing the Task-Unit Protocol needs, so it stays NULL for Task-Unit entries. A rule follows from it: a payload may not restate a field that already has a column or a table, which is how the second task body would appear. 2. "The insertion points already exist" was not true. RoundQueryPort carries two typed reads, cancelled() and accepted(); the dispatch loop fails a claim whose ticket has no patch, so patch work is the only work shape today and a peer protocol must bring its own. The record now says what is reusable material and what is still a plan step, and adds that step: widen the query port or add one beside DispatchBoard. 3. Two storage arguments are withdrawn. Atomicity does not require columns: measured on the store's SQLite 3.53.3, a conditional UPDATE over a JSON payload (WHERE json_extract(payload,'$.state')='open') succeeded once and changed zero rows on the second claim, identical to the column form - so the reason is clarity, direct indexing, CHECK constraints and PRAGMA table_info migration guards. And a generated column is a column, an index is a schema object: peer protocols are not schema-free. What is bought is narrower and still worth it - the added schema is purely additive, belongs to the protocol rather than the board, and no field is stored twice. 4. Three borrowings were applied too widely. source_session_id is provenance, not a credential, so the security rule is now stated as a constraint on what may be added rather than a defence of what exists. The refusal rule now keeps MCP's split: a domain refusal becomes data, a request-level error stays an error - which the dispatch loop already does by separating refused from failure. And the CloudEvents naming constraint is stated as borrowing (lowercase letters and digits; twenty characters recommended), with protocol_version's underscore recorded as a deviation rather than as compliance. --- .../2026-09-21-the-frame-and-its-storage.md | 119 ++++++++++++------ ...6-09-21-the-frame-and-its-storage.zh-CN.md | 71 +++++++---- 2 files changed, 129 insertions(+), 61 deletions(-) diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md index 49a50676..06ecb919 100644 --- a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md @@ -16,17 +16,17 @@ payload in order to route, add a column per protocol, invent a fresh vocabulary A survey of finished specifications changes the shape of the problem. Every layer but one is already defined by somebody, so the work is choosing and naming, not inventing: -| Layer | Already defined by | -| ----------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Event envelope: identity, time, version, content type, mode | CloudEvents 1.0 - required are only `id`/`source`/`specversion`/`type`; context attributes are "inspected at the destination without having to deserialize the event data"; attribute names are lowercase ASCII, at most 20 characters, never `data`; the type set is closed | -| Task, parts, states, paging | A2A 1.0.0 - terminal-state semantics, `contextId`, cursor paging (`pageToken`/`nextPageToken`, chosen explicitly over offsets), `historyLength`, and `includeArtifacts` defaulting false "to reduce payload size" | -| Payload discriminator and progressive shape | NLIP / ECMA-430 - `messagetype`, `format`, `subformat`, `content`, `submessages[]`, `label`; the first submessage is inlined at the top because most messages carry none; `format: structured` with `subformat: uri` carries a URI instead of bytes | -| LLM-facing results and errors | MCP 2025-06-18 - content blocks with `annotations.audience`, `structuredContent` beside text, and `isError` in the result, so a failure is data | -| One-line text encoding | RFC 8941 Structured Fields - List/Dictionary/Item with Parameters, intentionally strict | -| Acts and conversation threading | FIPA-ACL - `performative` is the only required field, beside `protocol`, `conversation-id`, `in-reply-to`, `reply-by` | -| Unknown-field policy | RFC 6709 - ignoring and refusing by name are both legitimate; the choice has to be stated | -| Security | ECMA-434 and the NLIP security guidelines - mandatory for NLIP conformance, with a scored threat list | -| **Ownership, lease, fencing, legality** | **Nobody.** Every agent protocol assumes a task belongs to one agent, so "who may hold it" does not arise | +| Layer | Already defined by | +| ----------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Event envelope: identity, time, version, content type, mode | CloudEvents 1.0 - required are only `id`/`source`/`specversion`/`type`; context attributes are "inspected at the destination without having to deserialize the event data"; attribute names are lowercase letters and digits, where the twenty-character limit is a recommendation rather than a rule, and `data` is not used as a name; the type set is closed | +| Task, parts, states, paging | A2A 1.0.0 - terminal-state semantics, `contextId`, cursor paging (`pageToken`/`nextPageToken`, chosen explicitly over offsets), `historyLength`, and `includeArtifacts` defaulting false "to reduce payload size" | +| Payload discriminator and progressive shape | NLIP / ECMA-430 - `messagetype`, `format`, `subformat`, `content`, `submessages[]`, `label`; the first submessage is inlined at the top because most messages carry none; `format: structured` with `subformat: uri` carries a URI instead of bytes | +| LLM-facing results and errors | MCP 2025-06-18 - content blocks with `annotations.audience`, `structuredContent` beside text, and `isError` for a tool execution error, which is a different thing from a protocol error and stays an error | +| One-line text encoding | RFC 8941 Structured Fields - List/Dictionary/Item with Parameters, intentionally strict | +| Acts and conversation threading | FIPA-ACL - `performative` is the only required field, beside `protocol`, `conversation-id`, `in-reply-to`, `reply-by` | +| Unknown-field policy | RFC 6709 - ignoring and refusing by name are both legitimate; the choice has to be stated | +| Security | ECMA-434 and the NLIP security guidelines - mandatory for NLIP conformance, with a scored threat list | +| **Ownership, lease, fencing, legality** | **Nobody.** Every agent protocol assumes a task belongs to one agent, so "who may hold it" does not arise | ## Proposal @@ -44,6 +44,18 @@ field set: the acts are already state columns (a claim, a delivery, a verdict, a - **Rendering, not storage**: the act an entry shows (FIPA's `performative`), and the progressive tiers T0 header, T1 summary, T2 structure, T3 truth. +What already holds the data, existing field to proposed field - nothing below moves: + +| Existing | In this frame | +| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `content` (NOT NULL) | The single source of the entry body. `readTaskBoardPreviews` projects `taskBoardPreview(entry.content)`: whitespace collapsed, cut to 200 characters, and a body that is only a `memory=` pointer returned verbatim | +| `kind`, `status`, `serial_state`, `created_at`, `expires_at`, `resolved_at`, `resolved_by`, `resolution` | Header columns, unchanged | +| `agent_id`, `to`, `source_session_id` | Header columns, unchanged. `source_session_id` is provenance - which session wrote the entry - not a credential | +| `claimed_by`, `claimed_at`, `claim_expires_at`, `attempt` | The lease and the attempt fence, unchanged | +| `delivered_by`, `delivered_at`, `deliverable_digest`, `deliverable_ref`, `deliverable_summary`, `judged_by`, `judged_at`, `judged_digest`, `verdict`, `verdict_reason`, `vetoed_by`, `vetoed_at`, `veto_reason` | Delivery reference, verdict and veto, unchanged | +| `task_run_manifest`, `task_run_tasks` (`input`, `dependencies`, `effect`, `wait_event`, `operation`, `kind`, `patch_files`, `patch_editable`) | The frozen task declaration, already typed, reached by `run_id` and `task_id`; per-attempt facts append to `task_run_facts`, whose own `payload` column is this record's precedent | +| - | The proposed `payload` carries nothing the Task-Unit Protocol needs. It is the slot a peer protocol's declaration uses, and for Task-Unit entries it stays NULL | + ### 2. The data format, three separate questions - **Wire bytes.** JSON over the daemon's existing RPC. Unchanged; no second binding exists to justify @@ -63,20 +75,32 @@ deferred until a transport with a bandwidth constraint exists), and protobuf as ### 3. Storage layout and its invariant -The header stays in columns because the claim is one atomic compare-and-set `UPDATE`. A header inside a -JSON document would turn atomic claiming into read-modify-write. That is a mechanism constraint, not a -preference. +The header stays in columns for clarity, direct indexing, `CHECK` constraints on enums and migration +guards that can read `PRAGMA table_info` - not because atomicity demands it. A conditional `UPDATE` +claims over a JSON payload just as well: measured on the store's SQLite 3.53.3, `UPDATE ... WHERE +json_extract(payload,'$.state')='open'` succeeded once and changed zero rows on the second claim, +identical to the column form. An earlier draft of this record argued from atomicity; that argument is +withdrawn. The payload is one document beside those columns, and the board never parses it. The store already keeps this shape: `claims_json`, `markers_json`, `scope_json`, `evidence_ids_json` and others sit beside typed columns in the same tables. -**Invariant (testable):** with `payload` set to NULL for every row, T0 and T1 still render and claim, -wake, expiry and compact read still work. Verified against the store's own SQLite (3.53.3). +**Invariant (testable):** the board's own reads and actions never consult `payload` - rendering T0/T1, +claiming, waking, expiry and compact read all work with it absent, and with it present they are +unaffected. This is a property of the board's read path, not something the storage format produces by +itself: T1 comes from `content`, as the mapping above shows. Under the Task-Unit Protocol `payload` is +NULL by construction, because the frozen declaration already has typed tables of its own. + +A payload may not restate a field that already has a column or a table - no digest, no verdict, no +holder, no instruction. Restating one is how the second task body this record exists to prevent would +come into being. -When a protocol needs to query a field inside its own payload, it does not ask the board for a column: a -generated column plus an expression index answers it (also verified on the same SQLite). This is what -keeps the schema from growing with the number of protocols. +When a protocol needs to query a field inside its own payload, it adds a generated column and an +expression index (both verified on the same SQLite). A generated column is a column and an index is a +schema object, so this does not make peer protocols schema-free; an earlier draft claimed it did, and +that claim is withdrawn. What it buys is narrower and still worth it: the added schema is purely +additive, it belongs to the protocol rather than to the board, and no field is stored twice. Cost, stated plainly: JSON columns give up database-level typing, a trade the store already makes, and cross-protocol queries depend on generated columns or the protocol's own tables. @@ -98,10 +122,16 @@ A peer protocol supplies four narrow things: 4. **Acceptance** - the criteria for accepting a delivery. The form (digest plus independent verdict) does not change. -The insertion points already exist: `DispatchBoard` and `RoundQueryPort`. A protocol plugs in on the -semantics side, and replacing one means another implementation plus another `protocol` value; the payloads -of entries written under the old value stay exactly as they are - readable, not actionable, and no -migration. +Reusable material exists; an interface with all four roles does not. `RoundQueryPort` carries two typed +reads - `cancelled()` and `accepted()` - and the dispatch loop fails a claim whose ticket has no `patch`, +so patch work is the only work shape today and a peer protocol has to bring its own. Widening that seam +(or adding a port beside `DispatchBoard`) is a step of this record, not a fact that already holds. What is +already in place is the legality computation the board calls, the read-only query port pattern, and the +dispatch loop's own separation of a refusal to use a slot from a failed unit. + +Once that seam exists, replacing a protocol means another implementation plus another `protocol` value; +the payloads of entries written under the old value stay exactly as they are - readable, not actionable, +and no migration. Unknown handling is stated rather than assumed: an unknown `protocol` makes an entry readable and not actionable, and an unknown field inside a known protocol is decided by that protocol's version rule, @@ -113,12 +143,16 @@ ECMA-434 is a conformance requirement for NLIP, and its companion security guide threats - prompt injection highest, then confused-deputy and token passthrough, session hijack, memory injection. Two rules are adoptable now as text: -- An in-band session token is a session credential: it must not be handed to the agent's model or exposed - in application-level payloads. The board's session identity is in-band and the board renders into model - context, so this rule applies to us directly. +- If an in-band session token is ever introduced, it is a session credential: it must not be handed to an + agent's model or exposed in application-level payloads. No such token exists in the frame today, so this + rule is a constraint on what may be added, not a defence of what is there. - A caller's token is never forwarded; a scoped token is minted instead. This becomes load-bearing only when entries cross a trust boundary, which today they do not. +The frame's `source_session_id` is provenance - which session wrote the entry - and not a credential. +Treating it as one because of its name would be exactly the over-extended borrowing this record has to +avoid. + What is owed is a threat list of our own, with the ones we accept and the ones we do not defend. ### 6. Scope ceiling @@ -130,14 +164,19 @@ test. Anything that needs a new table, a new tool or a new channel is outside th 1. This record. No code. 2. One migration: `protocol`, `protocol_version`, `payload_format`, `payload`, plus the invariant test. -3. Refusals become data at the RPC and tool boundary; the store keeps throwing. +3. Domain refusals become data at the RPC and tool boundary while request-level errors stay errors; the + store keeps throwing. 4. Tool schemas shaped for strict mode; flat T0/T1 rendering. -5. A threat list, with the two rules carried as text until a trust boundary exists. +5. The four-role protocol seam: widen the query port or add one beside `DispatchBoard`, and let a peer + protocol bring its own work shape instead of requiring `patch`. +6. A threat list, with the two rules carried as text until a trust boundary exists. ## Alternatives considered -- **One JSON document for the whole entry, header included.** Rejected: atomic claiming becomes - read-modify-write, and every reader would have to parse to find out whether an entry is claimable. +- **One JSON document for the whole entry, header included.** Rejected: every reader would have to parse + the header to learn whether an entry is claimable, `CHECK` constraints and the `PRAGMA table_info` + migration guards would no longer see the fields, and every read path would depend on JSON extraction. + Atomicity is not the reason - a conditional `UPDATE` over a JSON field claims exactly once. - **A column group per protocol.** Rejected: the schema grows with the number of protocols and every new protocol has to touch the board, which is the rework this record exists to avoid. - **protobuf as the normative source with generated bindings** (A2A's answer). Rejected for now: there is @@ -153,23 +192,31 @@ test. Anything that needs a new table, a new tool or a new channel is outside th ## Acceptance criteria -- With every `payload` NULL, the board renders T0/T1 and completes claim, wake, expiry and compact read. +- With every `payload` NULL, and again with payloads present, the board renders T0/T1 identically and + completes claim, wake, expiry and compact read. - Adding a peer protocol changes nothing in the board's claim, wake or expiry paths, demonstrated by a second protocol value with its own payload shape and its own projection. - An entry whose protocol the store does not know is readable, is not actionable, and its payload is never parsed by the board. - A payload field a protocol wants to query is indexed by a generated column and an expression index, - with no new board column. -- Header field names follow the borrowed naming constraints: lowercase ASCII, at most 20 characters, and - `data` is not used as a name. -- A refusal reaches the caller as data with a reason, not as an exception, at the RPC and tool boundary. + with nothing added to the board's own read and write paths. +- No field is stored twice: no header column is also carried inside a payload. +- Header field names borrow CloudEvents' attribute convention (lowercase letters and digits, twenty + characters recommended) or record the deviation. `protocol_version` has an underscore and so does not + follow it; that is recorded as a deviation rather than presented as compliance. +- A domain refusal - a unit outside the legal set, no published handoff - reaches the caller as data with a + reason, while a request-level error - a malformed call, an unknown protocol, an unknown tool - stays an + error. The dispatch loop already separates `refused` from `failure` in this spirit. - LLM-facing schemas are strict-mode shaped, and no negotiation text is required to be a JSON string. - Every new field names the specification it came from, or records why none fitted. ## Risks - One document per entry invites over-stuffing: the header can drift into the payload until the board has - to parse it again. The invariant test and the never-parse rule are the mitigation. + to parse it again. The invariant test and the never-parse rule are the mitigation, and the mapping table + is where duplication would start - a field carried in both a column and a payload is the failure mode. +- Withdrawn arguments invite re-use: atomicity does not require columns, and peer protocols do not avoid + schema changes. Both retractions are in section 3 so a later reader does not restate them. - A protocol that projects its own summary can summarize dishonestly. Mitigation: projection is separate from the verdict, and only `accepted` may be depended on. - Generated columns and expression indexes are SQLite-specific; a portable store would have to re-examine diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md index 168a2320..67128a6b 100644 --- a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md @@ -11,17 +11,17 @@ 对已定规范的调研改变了问题的形状:除了一层之外,每一层都已经有人定过,所以这件事是**选择与命名**,不是发明: -| 层 | 已由谁定 | -| ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| 事件信封:身份、时间、版本、内容类型、模式 | CloudEvents 1.0——必填只有 `id`/`source`/`specversion`/`type`;属性原话是"在目的地被检查而无需反序列化事件数据";属性名小写 ASCII、最多 20 字符、不得叫 `data`;类型集合封闭 | -| 任务、部件、状态、分页 | A2A 1.0.0——终态语义、`contextId`、游标分页(`pageToken`/`nextPageToken`,并明确不用 offset)、`historyLength`、`includeArtifacts` 默认 false("缩减载荷大小") | -| 载荷判别与渐进形状 | NLIP / ECMA-430——`messagetype`、`format`、`subformat`、`content`、`submessages[]`、`label`;首子消息平铺在顶层,因为多数消息不带子消息;`format: structured` 配 `subformat: uri` 时内容是一个 URI 而不是字节 | -| 面向 LLM 的结果与错误 | MCP 2025-06-18——内容块带 `annotations.audience`、`structuredContent` 与文本并存、`isError` 放在结果里,于是失败也是数据 | -| 单行文本编码 | RFC 8941 Structured Fields——List/Dictionary/Item 加 Parameters,故意严格 | -| 行为与对话线程 | FIPA-ACL——`performative` 是唯一必填,旁边是 `protocol`、`conversation-id`、`in-reply-to`、`reply-by` | -| 未知字段策略 | RFC 6709——忽略与按名字拒绝都合法,但**必须写明选了哪个** | -| 安全 | ECMA-434 与 NLIP 安全指南——对 NLIP 是符合性要求,并给出打分的威胁清单 | -| **拥有权、租约、围栏、合法性** | **没有人。** 所有 agent 协议都假设任务属于某一个 agent,"谁能持有"不成立 | +| 层 | 已由谁定 | +| ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 事件信封:身份、时间、版本、内容类型、模式 | CloudEvents 1.0--必填只有 `id`/`source`/`specversion`/`type`;属性原话是"在目的地被检查而无需反序列化事件数据";属性名只允许小写字母与数字(20 字符是建议上限而非规则),且不使用 `data` 作为名字;类型集合封闭 | +| 任务、部件、状态、分页 | A2A 1.0.0——终态语义、`contextId`、游标分页(`pageToken`/`nextPageToken`,并明确不用 offset)、`historyLength`、`includeArtifacts` 默认 false("缩减载荷大小") | +| 载荷判别与渐进形状 | NLIP / ECMA-430——`messagetype`、`format`、`subformat`、`content`、`submessages[]`、`label`;首子消息平铺在顶层,因为多数消息不带子消息;`format: structured` 配 `subformat: uri` 时内容是一个 URI 而不是字节 | +| 面向 LLM 的结果与错误 | MCP 2025-06-18--内容块带 `annotations.audience`、`structuredContent` 与文本并存、`isError` 用于工具执行错误,它与协议错误是两回事,后者仍是错误 | +| 单行文本编码 | RFC 8941 Structured Fields——List/Dictionary/Item 加 Parameters,故意严格 | +| 行为与对话线程 | FIPA-ACL——`performative` 是唯一必填,旁边是 `protocol`、`conversation-id`、`in-reply-to`、`reply-by` | +| 未知字段策略 | RFC 6709——忽略与按名字拒绝都合法,但**必须写明选了哪个** | +| 安全 | ECMA-434 与 NLIP 安全指南——对 NLIP 是符合性要求,并给出打分的威胁清单 | +| **拥有权、租约、围栏、合法性** | **没有人。** 所有 agent 协议都假设任务属于某一个 agent,"谁能持有"不成立 | ## 提案 @@ -33,6 +33,18 @@ - **一条载荷文档**:具名协议自己需要的东西,对板子不透明。 - **渲染,不是存储**:条目呈现的行为(FIPA 的 `performative`),以及分层的 T0 头、T1 摘要、T2 结构、T3 真值。 +现有字段到拟议字段的映射——下面没有一项要搬: + +| 现有 | 在本帧里的位置 | +| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `content`(NOT NULL) | 条目正文的唯一来源。`readTaskBoardPreviews` 投影的是 `taskBoardPreview(entry.content)`:空白折叠、截到 200 字符,而整条正文只是一个 `memory=` 指针时原样返回 | +| `kind`、`status`、`serial_state`、`created_at`、`expires_at`、`resolved_at`、`resolved_by`、`resolution` | 头列,不变 | +| `agent_id`、`to`、`source_session_id` | 头列,不变。`source_session_id` 是**来源归因**(哪个会话写的),不是凭据 | +| `claimed_by`、`claimed_at`、`claim_expires_at`、`attempt` | 租约与 attempt 围栏,不变 | +| `delivered_by`、`delivered_at`、`deliverable_digest`、`deliverable_ref`、`deliverable_summary`、`judged_by`、`judged_at`、`judged_digest`、`verdict`、`verdict_reason`、`vetoed_by`、`vetoed_at`、`veto_reason` | 交付物引用、裁决与否决,不变 | +| `task_run_manifest`、`task_run_tasks`(`input`、`dependencies`、`effect`、`wait_event`、`operation`、`kind`、`patch_files`、`patch_editable`) | 冻结任务声明,已经是 typed 表,由 `run_id` 与 `task_id` 抵达;每次尝试的事实追加进 `task_run_facts`,而后者自己的 `payload` 列正是本文形状的先例 | +| —— | 拟议的 `payload` **不承载任务单元协议需要的任何东西**。它是平级协议声明用的槽位,对任务单元条目保持 NULL | + ### 2. 数据格式是三个分开的问题 - **线上字节**:JSON,走 daemon 现有的 RPC。不变;既然只有一条线而没有第二个绑定,就不值得为代码生成付费,所以 A2A 的"proto 作规范源"暂不采用。 @@ -43,13 +55,15 @@ ### 3. 存储布局与它的不变量 -头留在列里,因为认领是**一条原子比较交换的 `UPDATE`**。把头放进 JSON 文档会把原子认领变成读-改-写。这是机制约束,不是偏好。 +头留在列里是为了清晰、可直接索引、对枚举加 `CHECK` 约束,以及能用 `PRAGMA table_info` 读的迁移守卫——**不是**因为原子性要求如此。条件 `UPDATE` 对 JSON 载荷一样能认领:在 store 的 SQLite 3.53.3 上实测,`UPDATE ... WHERE json_extract(payload,'$.state')='open'` 第一次成功、第二次更新零行,与列表单形式一致。本文早先的草稿从原子性论证,那个论证**撤回**。 载荷是与这些列并排的一条文档,板子从不解析它。这个形状库里已有:`claims_json`、`markers_json`、`scope_json`、`evidence_ids_json` 等,就是与 typed 列同表并存的。 -**不变量(可测)**:把所有行的 `payload` 置为 NULL,T0 与 T1 仍渲染,认领、唤醒、过期、紧凑读仍工作。已在 store 自己的 SQLite(3.53.3)上验证。 +**不变量(可测)**:板子自己的读与行动从不查 `payload`——`payload` 缺席时 T0/T1 渲染、认领、唤醒、过期、紧凑读照旧;`payload` 在场时不受影响。这是**读路径的性质,不是存储格式自然产生的性质**:如上面的映射所示,T1 来自 `content`。在任务单元协议下 `payload` 按构造就是 NULL,因为冻结声明已经有自己的 typed 表。 + +载荷**不得重述**已经有列或表承载的字段——不含 digest、不含裁决、不含持有人、不含指令。重述一个,就是让本文要防的那种“第二份任务体”出现的方式。 -当某个协议要查自己载荷里的字段时,它**不要**向板子要一列:生成列加表达式索引就够了(同样在该 SQLite 上验证过)。这一点正是让 schema 不随协议数量增长的机制。 +当某个协议要查自己载荷里的字段时,它加生成列与表达式索引(同样在该 SQLite 上验证过)。生成列**是列**,索引**是 schema 对象**,所以这并不能让平级协议“无 schema 变更”;本文早先的草稿这么声称过,那个说法**撤回**。它买到的是更窄但仍然值的东西:新增的 schema 是纯追加的、属于该协议而不属于板子,且没有字段被存两遍。 代价说明白:JSON 列放弃数据库级类型检查(这是仓库早已做过的交换),跨协议查询依赖生成列或协议自己的表。 @@ -64,7 +78,9 @@ 3. **投影**——T1 的一行摘要与 T2 的结构。必须有,因为载荷对板子不透明;MCP 与 A2A 出于同样理由也是这个形状。 4. **验收**——接受一次交付的判据。形式(digest 加独立裁决)不变。 -插入点已经存在:`DispatchBoard` 与 `RoundQueryPort`。协议接在语义侧,替换一个协议就是换一个实现加换一个 `protocol` 取值;旧取值下写入的载荷原样留着——可读、不可行动,且不需要迁移。 +可复用的材料存在,但四项职责齐全的接口不存在。`RoundQueryPort` 只有两个 typed 读——`cancelled()` 与 `accepted()`;派发循环对没有 `patch` 的票直接判失败,所以今天唯一的“工作形态”就是 patch,平级协议必须自带它自己的。把这条缝拓宽(或在 `DispatchBoard` 旁边加一个 port)是本文的**计划步骤**,不是既成事实。已经到位的是板子调用的合法性计算、只读查询端口的模式,以及派发循环里已有的“拒绝使用槽位”与“单元失败”之分。 + +那条缝一旦存在,替换一个协议就是换一个实现加换一个 `protocol` 取值;旧取值下写入的载荷原样留着——可读、不可行动,且不需要迁移。 未知的处理要**写明**而不是默认:未知 `protocol` 使条目可读但不可行动;已知协议内的未知字段由该协议的版本规则决定,而那条规则必须说明是忽略还是按名字拒绝。 @@ -72,9 +88,11 @@ ECMA-434 对 NLIP 是符合性要求,配套安全指南给十五类威胁打分:prompt injection 最高,其次是 confused-deputy 与令牌透传、会话劫持、记忆投毒。两条现在就能以文字形式采纳: -- 带内会话令牌就是会话凭据:不得交给 agent 的模型,也不得暴露在应用层载荷里。黑板的会话身份是带内的,而黑板会渲染进模型上下文,所以这条直接适用于我们。 +- 若将来引入带内会话令牌,它就是会话凭据:不得交给 agent 的模型,也不得暴露在应用层载荷里。今天帧里没有这样的令牌,所以这条是对“可新增什么”的约束,而不是对已有东西的防御。 - 绝不转发调用方的令牌,改为铸一个有范围的令牌。这条只有在条目跨信任边界时才承重,而今天还没有那样的边界。 +帧里的 `source_session_id` 是**来源归因**(哪个会话写的),不是凭据。仅因为名字里有 session 就把它当凭据,正是本文必须避免的那类过度套用。 + 欠的是一份我们自己的威胁清单,写明接受哪些、不防御哪些。 ### 6. 范围上限 @@ -85,13 +103,14 @@ ECMA-434 对 NLIP 是符合性要求,配套安全指南给十五类威胁打 1. 本文。不动代码。 2. 一次迁移:`protocol`、`protocol_version`、`payload_format`、`payload`,加那条不变量测试。 -3. 拒绝在 RPC 与工具边界变成数据;store 继续抛异常。 +3. 域内拒绝在 RPC 与工具边界变成数据,而请求级错误仍是错误;store 继续抛异常。 4. 工具 schema 按 strict 模式写;T0/T1 扁平渲染。 -5. 一份威胁清单,两条规则先以文字承载,等出现信任边界再谈强制。 +5. 四项职责的协议缝:拓宽查询端口,或在 `DispatchBoard` 旁边加一个;让平级协议自带工作形态,而不是要求 `patch`。 +6. 一份威胁清单,两条规则先以文字承载,等出现信任边界再谈强制。 ## 考虑过的替代方案 -- **整个条目一条 JSON 文档(含头)**。拒绝:原子认领会退化成读-改-写,而且每个读者都得先解析才知道条目能不能认领。 +- **整个条目一条 JSON 文档(含头)**。拒绝:每个读者都得先解析头才知道条目能不能认领;`CHECK` 约束与读 `PRAGMA table_info` 的迁移守卫都看不到字段;每条读路径都要依赖 JSON 抽取。**理由不是原子性**——对 JSON 字段的条件 `UPDATE` 恰好只认领一次。 - **每个协议一组列**。拒绝:schema 随协议数量增长,每个新协议都要动板子——那正是本文要避免的返工。 - **proto 作规范源并生成绑定**(A2A 的答案)。暂时拒绝:只有一条线、没有第二个绑定,代码生成现在买不到东西。 - **头用 RFC 8941 的单行编码**。存储上拒绝——我们的头活在列里——但若将来需要紧凑文本读,保留为选项。 @@ -101,20 +120,22 @@ ECMA-434 对 NLIP 是符合性要求,配套安全指南给十五类威胁打 ## 验收标准 -- 所有 `payload` 为 NULL 时,板子仍渲染 T0/T1,并仍能完成认领、唤醒、过期与紧凑读。 +- 所有 `payload` 为 NULL 时、以及载荷在场时,板子渲染的 T0/T1 一致,并都能完成认领、唤醒、过期与紧凑读。 - 增加一个平级协议不改变板子的认领、唤醒、过期路径;用一个自带载荷形状与投影的第二协议取值演示。 - store 不认识的协议:条目可读、不可行动,且板子从不解析它的载荷。 -- 协议想查的载荷字段由生成列加表达式索引支撑,板子不加新列。 -- 头字段名遵守借来的命名约束:小写 ASCII、最多 20 字符、不使用 `data` 作为名字。 -- 拒绝在 RPC 与工具边界以带理由的数据抵达调用者,而不是异常。 +- 协议想查的载荷字段由生成列加表达式索引支撑,板子自己的读写路径不加东西。 +- 没有字段被存两遍:没有哪个头列同时也在载荷里。 +- 头字段名**借鉴** CloudEvents 的属性约定(只小写字母与数字,20 字符是建议),或者记录偏离。`protocol_version` 带下划线、并不遵守它;这记作偏离,而不是声称遵守。 +- 域内拒绝(单元在合法集之外、没有公布交接)以带理由的数据抵达调用者;请求级错误(畸形的调用、未知协议、未知工具)仍是错误。派发循环已经在这个意义上区分 `refused` 与 `failure`。 - 面向 LLM 的 schema 按 strict 模式成型,且没有任何协商文本被要求写成 JSON 字符串。 - 每个新字段都指名它来自哪份规范,或者记录下为什么没有一份合适。 ## 风险 -- 每条一条文档会诱发"什么都往里塞":头可能慢慢漂进载荷,直到板子又得解析它。缓解不变量测试与"板子从不解析"这条规则。 +- 每条一条文档会诱发“什么都往里塞”:头可能慢慢漂进载荷,直到板子又得解析它。缓解是不变量测试与“板子从不解析”这条规则;而**映射表就是重复开始的地方**——一个字段同时躺在列里和载荷里,就是那个失败模式。 +- 撤回的论证会被重新用起来:原子性不要求列,平级协议也躲不开 schema 变更。两条撤回都写在第三节,免得后来的读者再把它们复述一遍。 - 自己给自己做投影的协议可以说得不诚实。缓解:投影与裁决分离,且只有 `accepted` 可被依赖。 - 生成列与表达式索引是 SQLite 特有的;换可移植存储时要重新审视。 -- "拒绝即数据"改变现有调用者看到的形状,所以臂的 driver 必须在同一个切片里一起动。 +- “域内拒绝即数据”改变现有调用者看到的形状,所以臂的 driver 必须在同一个切片里一起动。 - strict 模式的 schema 让缺省显式化,每次调用多花 token,而且可能让读者把可空字段读成必填。 - 借来的词汇会僵化。若某个借来的名字不合身,允许偏离——但要**记录**偏离,而不是悄悄改名。 From 41867e5c83adb316489f802f89b5da3b5b079456 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Mon, 21 Sep 2026 23:33:07 +0800 Subject: [PATCH 10/32] docs(decisions): mechanism, not policy - the division both records needed The legality record and the frame record each needed the same sentence and neither stated it: the legality record says what the program does and what it refuses to decide, the frame record says what the core knows and what it never parses. Those are one division seen from the side of the decision and from the side of the data. Because nothing named it, it had to be re-argued every time a field or a rule appeared, and the cost was already visible: "work means a patch" had reached six places, one of them a legality rule written in the kernel's own vocabulary (refuseWidening decides permission closure by reading parent.patch.editable). This record takes Hydra's principle - a kernel provides mechanisms and refuses policy - and states it as two contracts. Mechanism: one claim (compare-and-set, lease, attempt fence), an opaque declaration with a digest, an opaque artifact with a digest, a verdict from someone else, the lifecycle, and the legal set (which units are legal now, in order, cut to the declared slot budget, with a reason per unit). Policy: what a unit's inputs are, the artifact class, how the artifact is judged, how the work runs, its concurrency preconditions, the wording and enablement of legality rules, and who is chosen, who adopts, who judges. It also supplies the two things the slogan lacked. A decision procedure - for each field, type and code path, ask mechanism or policy, with a rule's check counting as mechanism while its name, wording and enablement count as policy. And a mechanically checkable rule: no policy word may appear in the mechanism layer's code, type declarations or schema paths, from a maintained list (patch, editable, files, instruction, checks, repair-first), with the path carried because a word can be policy in one place and mechanism in another. The consequence for work shapes is stated: a work shape is not a core concept, it is the policy layer's name for the declaration-and-artifact pair the core carries opaquely. The six leak sites are classified in a table, and the legality rule found today is the refinement the record adds - its check belongs to the program while permission closure's name and enablement belong to a declaration, so the fix is not to move the check out of the program but to stop hard-wiring the rule in the kernel's vocabulary. That is the same shape as plan step 3 of the legality record, where repair-first becomes a declared constraint instead of shared planning policy; two independent fixes taking the same shape is the evidence that the classification is right. Three rows are deliberately left undecided so that they are not settled by preference: task_run_tasks.effect is mechanism only if the mechanism must enforce write-set disjointness, the granularity of input and dependencies is a separate question from whether dependencies are mechanism, and operation may be policy. Alternatives rejected: keeping the slogan and deciding case by case, putting the principle in either record (it is wider than both), a plugin registry above the core with no second policy to justify it, and deciding the undecided rows now. Risks named: gutting the core (so the practical form is one default policy the core may carry but must not require), an ossifying word list, an undecided row becoming permanent, a policy word legitimately remaining for a release, and grep being gameable by synonyms. The legality record and the frame record now point here, so the division has one home. --- ...2026-09-20-the-program-answers-legality.md | 2 +- ...9-20-the-program-answers-legality.zh-CN.md | 2 +- .../2026-09-21-mechanism-not-policy.md | 129 ++++++++++++++++++ .../2026-09-21-mechanism-not-policy.zh-CN.md | 91 ++++++++++++ .../2026-09-21-the-frame-and-its-storage.md | 2 +- ...6-09-21-the-frame-and-its-storage.zh-CN.md | 2 +- 6 files changed, 224 insertions(+), 4 deletions(-) create mode 100644 docs/decisions/proposed/2026-09-21-mechanism-not-policy.md create mode 100644 docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md index 88344263..bf8fb31f 100644 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md @@ -3,7 +3,7 @@ [中文](2026-09-20-the-program-answers-legality.zh-CN.md) **Status:** proposed -**Relates to:** [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Task unit semantics](../../design/task-unit-semantics.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md) +**Relates to:** [Mechanism, not policy](2026-09-21-mechanism-not-policy.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Task unit semantics](../../design/task-unit-semantics.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md) ## Problem diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md index 4e378a8e..b0366c34 100644 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md @@ -3,7 +3,7 @@ [English](2026-09-20-the-program-answers-legality.md) **Status:** proposed -**Relates to:** [黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md)、[契约的义务](../../design/task-unit-semantics-obligations.md) +**Relates to:** [机制,不是策略](2026-09-21-mechanism-not-policy.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md)、[契约的义务](../../design/task-unit-semantics-obligations.md) ## 问题 diff --git a/docs/decisions/proposed/2026-09-21-mechanism-not-policy.md b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.md new file mode 100644 index 00000000..85734743 --- /dev/null +++ b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.md @@ -0,0 +1,129 @@ +# Mechanism, not policy + +[中文](2026-09-21-mechanism-not-policy.zh-CN.md) + +**Status:** proposed +**Relates to:** [The program answers legality](2026-09-20-the-program-answers-legality.md), [The frame, its data format, and its storage](2026-09-21-the-frame-and-its-storage.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Task unit semantics](../../design/task-unit-semantics.md) + +## Problem + +Two proposed records each need the same sentence, and neither states it. The legality record says what the +program does and what it refuses to decide. The frame and storage record says what the core knows and what +it never parses. Those are the same division, seen from the side of the decision and from the side of the +data, and because nothing names it, the division has to be re-argued every time a field or a rule appears. + +The cost is already visible. The assumption that work means a patch reached six places, and one of them is a +legality rule written in the kernel's own vocabulary: `refuseWidening` decides permission closure by reading +`parent.patch.editable`, so a rule about write sets is expressed in terms of a work shape the kernel is not +supposed to know. + +The split also needs to be sharp enough to settle an argument rather than to serve as a slogan. "Keep policy +out of the core" does not decide whether `effect` or `operation` may be a column, or whether a legality rule +may live in the program at all. + +## Proposal + +Take Hydra's principle - a kernel provides mechanisms and refuses policy - and state it as two contracts. + +**Mechanism contract** (the core: the board, the store, the program): + +- one claim: compare-and-set, lease, attempt fence +- an opaque declaration with a digest +- an opaque artifact with a digest +- a verdict given by someone else +- the lifecycle: expiry, reaping, wake delivery, compact read +- the legal set: which units are legal now, in order, cut to the declared slot budget, with a reason per unit + +**Policy contract** (above the core: the protocol, the plan, the agents): + +- what a unit's inputs are, the artifact class, how the artifact is judged, how the work runs (synchronous + or detached, interruptibility, whether abandoning it midway is safe), and its concurrency preconditions +- the wording and the enablement of legality rules: named by the protocol, enabled by the plan +- who is chosen, in what order, who adopts, who judges + +**The decision procedure.** For each field, type and code path, ask whether it is mechanism or policy. A +rule's _check_ is mechanism; the rule's _name, wording and enablement_ are policy. + +**The mechanically checkable rule.** No policy word may appear in the mechanism layer's code, type +declarations or schema paths. The list is maintained - `patch`, `editable`, `files`, `instruction`, +`checks`, `repair-first` - and a hit is a leak rather than a style question. The list may grow, and adding +a word is recorded. A word alone is not enough to judge: `files` is a policy word in a ticket type and a +mechanism word at a filesystem boundary, so the check carries the path. + +**Rows deliberately left undecided**, so that they are not settled by accident: + +- `task_run_tasks.effect`, the declared write set, is mechanism only if the mechanism must enforce that + write sets do not overlap. If enforcement is a protocol obligation, it is policy data. +- the granularity of `input` and `dependencies`: dependencies affect legality and order, which is mechanism, + but whether an input carries content or only a digest is a separate question. +- `operation`, which may be policy. + +## What this makes of the work shape + +A work shape is not a core concept. It is the policy layer's name for the declaration-and-artifact pair +that the core carries without understanding it. The six sites where the patch assumption lives classify as +follows. + +| Site | Mechanism or policy | Where it belongs | +| ---------------------------------------------------------------------------- | ----------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | +| `BoardTicket.patch?: FrozenPatchTask & { digest }` | policy | the ticket carries an opaque `declaration` with a digest | +| `patchFrozen()` calling `preparePatchWork`, and the `not a patch task` throw | policy inside the core | freezing and field validation move into the shape adapter; the core only hands back the opaque declaration | +| `refuseWidening` reading `parent.patch.editable` | check is mechanism, wording and enablement are policy | permission closure becomes a declared constraint the program enforces and reports reasons for | +| the ten `FrozenPatchWork` signatures in the session mechanism | mixed | rendering a declaration into a prompt and submitting an artifact are policy; session lifecycle, metrics and cancellation are mechanism | +| `task_run_tasks.patch_files` / `patch_editable`, parsed in the store | policy in the schema | belongs in the payload document; the mechanism part is the run, task, revision, input, dependency and operation columns | +| drivers asserting `ticket.patch!.digest` | policy assertion | the mechanism assertion is that a declaration carries a digest, that a claim binds the attempt, and that a delivery binds the digest | + +The legality rule found today is the clearest case of the refinement this record adds: its check belongs to +the program, while permission closure's name and enablement belong to a declaration. The fix is therefore not +to move the check out of the program but to stop hard-wiring the rule in the kernel's vocabulary - the same +shape of fix as plan step 3 of the legality record, where repair-first becomes a declared constraint rather +than shared planning policy. Two independent fixes taking the same shape is evidence that the classification +is the right one. + +## Alternatives considered + +- **Keep the slogan and decide case by case.** Rejected: that is what produced six sites, and it offers no + test to apply. +- **Put the principle in the frame record.** Rejected: the principle is wider than the frame - it also + governs the program's share of decisions - so the frame record would become the owner of a rule about the + program. +- **Put it in the legality record.** Rejected for the same reason in reverse: that record is narrower, being + about the program's decisions, and what the core may know is not a legality question. +- **A plugin framework with a registry above the core.** Rejected: no second policy exists to justify the + machinery; one default adapter plus one other shape tests the seam. +- **Decide the undecided rows now.** Rejected: choosing them by preference is precisely what this record + replaces with evidence. + +## Acceptance criteria + +- The classification covers every site that mentions a policy word, each row marked mechanism, policy, or + undecided. +- No policy word appears in the mechanism layer's read and write paths, type declarations, or schema paths, + from a maintained list and checked mechanically. +- A legality rule's name and enablement come from a declaration, while its check runs in the program and + returns a reason per unit. +- Adding a second work shape changes no mechanism code, demonstrated by the value work the data-check + runner already performs. +- Every undecided row is marked as undecided and is resolved by an experiment rather than by preference. +- The legality record and the frame record both point here, so the division has one home. + +## Risks + +- **Gutting the core.** Hydra's lesson cuts both ways: a kernel with no default policy is unusable. The + practical form is that the core may carry one default policy, patch work, and must not require it. +- **The policy-word list can ossify.** A word may be mechanism in one place and policy in another, so the + check needs the path rather than the word alone, and the list has to accept additions. +- **An undecided row can become permanent.** A row marked undecided without an experiment behind it is a + policy nobody declared. +- **A policy word may legitimately remain for a release.** This record does not require a big-bang rename; it + requires that each remaining occurrence is listed. +- **A grep check can be gamed by synonyms.** The check is a floor, not a proof of separation. + +## Plan + +1. This record. No code. +2. The classification table, with the undecided rows marked. +3. The check: a maintained policy-word list plus a script, wired into the static set. +4. Both records point here. +5. Then the work-shape slice: the mechanism side carries an opaque declaration, with patch as the default + adapter rather than a type. diff --git a/docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md new file mode 100644 index 00000000..b1bdb670 --- /dev/null +++ b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md @@ -0,0 +1,91 @@ +# 机制,不是策略 + +[English](2026-09-21-mechanism-not-policy.md) + +**Status:** proposed +**Relates to:** [程序只回答合法性](2026-09-20-the-program-answers-legality.zh-CN.md)、[帧、它的数据格式与它的存储](2026-09-21-the-frame-and-its-storage.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) + +## 问题 + +两份提案各需要同一句话,而谁都没把它说出来。合法性记录讲的是**程序做什么**、以及它拒绝决定什么;帧与存储记录讲的是**核心知道什么**、以及它从不解析什么。这是同一个划分,一面从决定看、一面从数据看;因为没有名字,每次出现一个新字段或一条新规则,这个划分都得重新论证一遍。 + +代价已经看得见:**"工作就是补丁"这个假设走到了六处**,其中一处是用核心自己的词汇写出来的合法性规则——`refuseWidening` 靠读 `parent.patch.editable` 来判定 permission closure,于是"关于写集"的规则被写成了"关于某个核心不该知道的工作形态"的规则。 + +而且这个划分必须锐到能**终结争论**,而不是当口号。"别把策略放进核心"这句话,并不能决定 `effect` 或 `operation` 能不能是一列,也不能决定一条合法性规则到底能不能住在程序里。 + +## 提案 + +取 Hydra 的原则——**内核提供机制、拒绝策略**——把它写成两条契约。 + +**机制契约**(核心:黑板、存储、程序): + +- 一次认领:比较交换、租约、attempt 围栏 +- 一份**不透明**声明 + 摘要 +- 一份**不透明**产物 + 摘要 +- 一个**由别人给出**的裁决 +- 生命周期:过期、回收、唤醒投递、紧凑读 +- 合法集:此刻哪些单元合法、有序、按声明的槽位预算裁剪、每个单元给理由 + +**策略契约**(核心之上:协议、计划、各 agent): + +- 单元的输入是什么、产物属于哪一类、产物怎么被判、工作怎么跑(同步还是分离式检查、可否打断、中途放弃是否安全)、以及它的并发前提 +- 合法性规则的**措辞与启用**:由协议具名、由计划声明启用 +- 选谁、按什么顺序、谁采纳、谁裁决 + +**判定程序**:对每个字段、每个类型、每条代码路径问一次——**这是机制还是策略?** 一条规则的**检查**是机制,规则的**名字、措辞与启用**是策略。 + +**可机械检查的规则**:**机制层(核心与程序)的代码、类型声明与 schema 路径里不得出现策略词汇。** 清单是维护的——`patch`、`editable`、`files`、`instruction`、`checks`、`repair-first`——命中即泄漏,不是风格问题。清单可以加词,加词要记录。光看词不足以判定:`files` 在票的类型里是策略词,在文件系统边界上是机制词,所以检查要带上**路径**。 + +**刻意留作待判的行**(免得被顺手定掉): + +- `task_run_tasks.effect`(声明的写集):只有当**机制必须强制写集不相交**时它才是机制;若强制是协议的义务,它就是策略数据。 +- `input` 与 `dependencies` 的粒度:依赖影响合法性与顺序,这是机制;但"输入"是携带内容还是只携带摘要,是另一个问题。 +- `operation`:可能是策略。 + +## 这给"工作形态"定了位 + +**工作形态不是核心概念**,它是策略层对"核心不透明携带的那对声明与产物"的叫法。六处站点的分类如下。 + +| 站点 | 机制还是策略 | 归属 | +| ---------------------------------------------------------------------- | ---------------------------- | ----------------------------------------------------------------------------------- | +| `BoardTicket.patch?: FrozenPatchTask & { digest }` | 策略 | 票上带一份不透明的 `declaration` 加摘要 | +| `patchFrozen()` 调 `preparePatchWork`、以及 `throw "not a patch task"` | 策略(在核心里) | 冻结与字段校验移进形态适配器;核心只交出不透明声明 | +| `refuseWidening` 读 `parent.patch.editable` | 检查是机制;措辞与启用是策略 | permission closure 变成程序**执行并给理由**的一条声明约束 | +| 会话机制里十处 `FrozenPatchWork` | 混合 | 把声明渲染成提示、提交产物是策略;会话生命周期、指标、取消是机制 | +| `task_run_tasks.patch_files` / `patch_editable`(store 里解析) | 策略进了 schema | 该进载荷文档;机制部分是 run、task、revision、input、dependencies、operation 那几列 | +| 驱动断言 `ticket.patch!.digest` | 策略断言 | 机制断言是:声明带摘要、认领绑 attempt、交付绑摘要 | + +今天那条合法性规则正是本文新增的那点修正的最清楚的例子:**它的检查属于程序,而 permission closure 的名字与启用属于声明。** 所以修法不是把检查搬出程序,而是**别再用核心的词汇硬编码这条规则**——这与合法性记录计划第 3 步(把 repair-first 从共享规划策略变成声明的约束)是**同一形状的修法**。两处各自独立的修法落在同一形状上,是"这个分类是对的"的证据。 + +## 考虑过的替代方案 + +- **留着口号,逐案判断。** 拒绝:那就正是产出六处站点的方式,而且它不提供可以套用的检验。 +- **把这条原则放进帧记录。** 拒绝:原则比帧更宽——它也管程序那半决定——那样帧记录会变成"关于程序的规则"的拥有者。 +- **放进合法性记录。** 反向的同样理由拒绝:那份记录更窄(讲程序的决定),而"核心可以知道什么"不是合法性问题的。 +- **核心之上搞一个带注册表的插件框架。** 拒绝:还没有第二个策略来支撑这套机械;一个默认适配器加另一个形态就足以检验这条缝。 +- **现在就定掉待判的行。** 拒绝:凭偏好挑一边,正是本文要用证据替换掉的做法。 + +## 验收标准 + +- 分类覆盖所有出现策略词汇的站点,每行标为机制、策略或待判。 +- 机制层的读写路径、类型声明与 schema 路径里不出现策略词汇(清单维护,机械检查)。 +- 合法性规则的名字与启用来自声明,检查在程序里跑,并按单元返回理由。 +- 增加第二个工作形态不改变任何机制代码;用数据式检查运行器已经在做的"值工作"演示。 +- 每一行待判都标为待判,并且由**实验**而不是偏好来定。 +- 合法性记录与帧记录都指到这里,让这个划分只有一个出处。 + +## 风险 + +- **把核心做空。** Hydra 的教训是双向的:一个不带默认策略的内核不可用。务实形式是核心**可以带一个默认策略**(补丁工作),但**不得要求**它。 +- **策略词清单会僵化。** 同一个词在一处是机制、在另一处是策略,所以检查要带路径而不只是词,而清单必须接受加词。 +- **待判的行可能永久待判。** 一行标着待判却没有实验在背后,就是一个没人声明的策略。 +- **某个策略词可能合理地再留一个版本。** 本文不要求一次性重命名;它要求每一处残留都被**列出来**。 +- **grep 检查可以被同义词绕过。** 检查是地板,不是"已经分离"的证明。 + +## 计划 + +1. 本文。不动代码。 +2. 分类表,标出待判的行。 +3. 那条检查:维护的策略词清单加一个脚本,接进静态集合。 +4. 两份记录都指过来。 +5. 然后才是工作形态切片:机制侧携带不透明声明,`patch` 作为默认适配器而不再是类型。 diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md index 06ecb919..9111fca4 100644 --- a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md @@ -3,7 +3,7 @@ [中文](2026-09-21-the-frame-and-its-storage.zh-CN.md) **Status:** proposed -**Relates to:** [The program answers legality](2026-09-20-the-program-answers-legality.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Task unit semantics](../../design/task-unit-semantics.md) +**Relates to:** [Mechanism, not policy](2026-09-21-mechanism-not-policy.md), [The program answers legality](2026-09-20-the-program-answers-legality.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Task unit semantics](../../design/task-unit-semantics.md) ## Problem diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md index 67128a6b..71c966ba 100644 --- a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md @@ -3,7 +3,7 @@ [English](2026-09-21-the-frame-and-its-storage.md) **Status:** proposed -**Relates to:** [程序只回答合法性](2026-09-20-the-program-answers-legality.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) +**Relates to:** [机制,不是策略](2026-09-21-mechanism-not-policy.zh-CN.md)、[程序只回答合法性](2026-09-20-the-program-answers-legality.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) ## 问题 From 4cfbe115c0bd4d07724bb0322ae27cb943387812 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Tue, 22 Sep 2026 21:25:29 +0800 Subject: [PATCH 11/32] docs(design): a layered model, with the checks that decide whether it earns a place Records the model that came out of the last exchange as a draft rather than as a record, because it was reached by stacking analogies and nothing about it has been measured. The model: a drawing pipeline has four stages and so does ours - producer (a work-shape adapter), surface (the board entry: payload plus a discriminator and a digest, delivery as commit, resolve and expiry as release, wake as the frame callback), compositor (the board and the program: claim, lease, fence, serial and wake, expiry and reaping, the legal set), and display (the presentation end: tiers, flat lines, compact read, the tool descriptions an agent reads). From it come three ownership classes: content and semantics belong to the owner and are semantically opaque to the middle (a compositor reads pixels but cannot read what a button means, so it asks the owner; our board reads declared structure and never payload semantics); description is the owner's obligation (geometry, level, shape, dirty regions, terminal state, identity, a snapshot able to stand in for the owner - ours is the discriminator, digest, dependencies, scope, and the protocol's projection); and arbitration is only the middle's, because stacking, occlusion, visibility, focus, capture, timing and reaping are global properties. The boundary test that replaces a slogan: no reasoning needed means the object stays opaque; structure needed means the owner supplies a description and the middle reads the description but not the semantics; a global judgement means only the middle can make it, which is why the description must be cheap and fresh and why staleness needs a meaning. The document also carries its own case against itself: the analogies did the reasoning and were each partly overruled by a fuller view, one of the four stages is a name with no seam or owner or test, the only measurable claim in this arc (that the seam was narrow) was falsified by a grep within a minute, and nothing got smaller - no code changed, no concept removed, three prose records added. So it states an eligibility rule - an analogy earns a place in a record only when a check can falsify it - and records its predictions before measuring: (a) policy words in the middle, predicted at fifteen or more hits concentrated in the board and the legality module; (b) the two ends, predicted to be concentrated and separable at the producer end and scattered with no seam at the presentation end; (c) a second shape through the dispatch loop, predicted to change five or more files. The reduction test is named in advance too: less code (the middle must lose a type and a freeze method, or the model failed), fewer concepts (the three classes must absorb the two lists they replace rather than sit beside them), fewer change points (measured by how many places a new shape touches), and better maintainability (a new shape adds files under one owner and changes no test in the middle). --- docs/design/mechanism-in-the-middle.md | 106 +++++++++++++++++++ docs/design/mechanism-in-the-middle.zh-CN.md | 65 ++++++++++++ 2 files changed, 171 insertions(+) create mode 100644 docs/design/mechanism-in-the-middle.md create mode 100644 docs/design/mechanism-in-the-middle.zh-CN.md diff --git a/docs/design/mechanism-in-the-middle.md b/docs/design/mechanism-in-the-middle.md new file mode 100644 index 00000000..d692dac1 --- /dev/null +++ b/docs/design/mechanism-in-the-middle.md @@ -0,0 +1,106 @@ +# Mechanism in the middle, adapters at both ends + +**Status:** draft +**Created:** 2026-09-21 + +A layered model of the collaboration protocol, and the checks that decide whether the model earns a place in +the records. This document owns the model, the boundary test, and the checks. It does not restate the parts +inventory, which lives in [protocol-governed-collaboration.md](protocol-governed-collaboration.md), nor the +decisions, which live in three proposed records: [mechanism, not policy](../decisions/proposed/2026-09-21-mechanism-not-policy.md), +[the frame and its storage](../decisions/proposed/2026-09-21-the-frame-and-its-storage.md), and +[the program answers legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md). + +## The model + +A drawing pipeline has four stages, and so does ours. + +| Stage | What it does | What it is here | Where it lives today | +| ---------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | +| Producer | turns something into a buffer | a work-shape adapter: validate the declaration, prepare the inputs, run, submit | `ooo-patch.ts`, `ooo-session-mechanism.ts`, `evals/ooo-execution/data-check-runner.ts` | +| Surface | the shared, opaque exchange object with a role, commit, release and frame callbacks | the board entry: `payload` plus a discriminator and a digest, delivery as commit, resolve and expiry as release, wake as the frame callback | `task_board_entries` with its added columns | +| Compositor | stacking, occlusion, blending, timing, culling, and playing a snapshot when the owner is away | the board and the program: claim, lease, fence, serial and wake, expiry and reaping, the legal set | `src/integration/ooo-board.ts`, `src/integration/ooo-dispatch.ts`, `task-semantics.ts` | +| Display | mode, refresh rate, colour space, scaling - the same frame presented differently per device | the presentation end: T0-T3 tiers, flat lines, compact read, the tool descriptions an agent reads | `taskBoardPreview`, the tier rendering, `src/prompts/nmg-prompts.yaml` | + +Three ownership classes follow, and they are the model's real content: + +1. **Content and semantics** belong to the owner and are semantically opaque to the middle. The compositor + reads pixels to blur or to know an occlusion shape, but it cannot read what a button means; for meaning it + asks the owner. Our board may read declared structure and never reads payload semantics. +2. **Description** is the owner's obligation: geometry, level and parent relationships, shape and opaque and + dirty regions, the terminal state of an animation, an identity for restoration, and a snapshot able to + stand in for the owner. Ours is the discriminator, the digest, dependencies, scope and need, and the + projection a protocol supplies. +3. **Arbitration** is only the middle's, because only the middle has the global view: stacking and occlusion + are global properties, and so are visibility, focus, capture permission, timing and reaping. Ours is the + legal set, claim and lease, wake, expiry. + +The boundary test that replaces a slogan: + +- the middle does not need to reason about the object, so the object stays opaque and belongs to its owner; +- the middle needs structure, so the owner must supply a description and the middle reads the description + but not the semantics; +- the judgement is global, so only the middle can make it - which requires the description to be cheap and + fresh, and requires staleness to have a meaning (an expired lease is a suspicion, not a death certificate). + +Two consequences are already load-bearing elsewhere. The middle must compose in a form that is independent +of how it will be shown, so it never learns a reader's format. And presentation is not verifiable - a display +applies its own colour, crop and scaling - so an acceptance can only ever rest on the artifact and its digest, +never on how it was shown. + +## Why this is a draft and not a record + +- It was reached by stacking analogies, and each analogy was already partly overruled by a fuller view: first + a kernel split, then a window split into frame and content, now four stages. The analogies are doing the + reasoning, and that is the smell this section exists to record. +- One of the four stages is a name and nothing else. The presentation end has no seam, no owner and no tests; + it is a description of scattered code. +- The only measurable claim made in this arc - that the mechanism/policy seam was narrow - was falsified by a + grep within a minute: the patch assumption sits in six places, one of them a legality rule. +- Nothing has got smaller. No code changed, no concept was removed, and three prose records were added. On + today's evidence the model is strictly more concepts than the thing it describes. + +## The eligibility rule + +**An analogy earns a place in a record only when a check can falsify it.** The model above is allowed to stay +in this document while its checks are pending, and it may be promoted into a record only if the checks come +out in its favour. If they do not, this document is archived rather than promoted. + +## The checks + +Predictions are recorded before the measurement, so a surprise is visible rather than rationalised. + +**(a) Policy words in the middle.** Grep the middle layer - `src/core/store/`, `src/integration/ooo-board.ts`, +`src/integration/ooo-dispatch.ts`, `src/integration/task-semantics.ts`, `src/integration/ooo-execution.ts`, +`src/integration/ooo-candidate.ts` - for the policy word list: `patch`, `editable`, `instruction`, `checks`, +`files`, `repair-first`. Report every hit with its path, not just a count, because a word can be mechanism at +one path and policy at another. Prediction: fifteen or more hits, concentrated in the board and the legality +module. Expected finding: the board's own ticket type and the freeze method are the centre of it. + +**(b) The two ends.** Enumerate every site that translates between the board's canonical form and a producer +or reader format. Prediction: the producer end is concentrated and separable (about four sites), while the +presentation end is scattered across the store's preview function, the tier rendering, the prompt file, the +CLI and the extension, with no seam at all. If the presentation end cannot be separated, the fourth stage is a +name for a fact rather than a design, and it should be dropped rather than built. + +**(c) A second shape.** Run one unit through the dispatch loop whose work is not a patch, with a minimal +resolver (patch stays the default; an unknown shape is refused by name rather than failed). Count the files +that have to change. Prediction: five or more. + +## The reduction test + +The model earns its place only by making something smaller, and the four things it must make smaller are +named in advance: + +- **Less code:** taking the seam means the middle loses a type and a freeze method, and the shape's field + validation moves to the adapter that owns the shape. If the net is an increase, the model failed. +- **Fewer concepts:** the three ownership classes must _absorb_ the two lists they replace - the mechanism + list and the policy list - rather than being added beside them. If a reader has to hold both, the model is + a restatement, not an abstraction. +- **Fewer change points:** the six sites must become fewer, or the same six in one place instead of six. The + number that matters is how many places a _new_ shape must touch. +- **Better maintainability:** a new shape adds files under one owner and changes no test in the middle; the + middle's suite passes untouched with a shape it has never seen. + +## Results + +Pending. diff --git a/docs/design/mechanism-in-the-middle.zh-CN.md b/docs/design/mechanism-in-the-middle.zh-CN.md new file mode 100644 index 00000000..f0fcdd7e --- /dev/null +++ b/docs/design/mechanism-in-the-middle.zh-CN.md @@ -0,0 +1,65 @@ +# 机制在中间,适配器在两端 + +**Status:** draft +**Created:** 2026-09-21 + +协作协议的分层模型,以及决定这个模型有没有资格进记录的那几项检查。本文拥有模型、那条界线测试、以及这些检查。它不复述大类清单(那在 [protocol-governed-collaboration.md](protocol-governed-collaboration.zh-CN.md)),也不复述决策(那在三份提案里:[机制,不是策略](../decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md)、[帧、它的数据格式与它的存储](../decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md)、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md))。 + +## 模型 + +一条绘图流水线有四段,我们也是。 + +| 段 | 它做什么 | 我们这边的对应 | 今天在哪 | +| ------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | +| 生产者 | 把某样东西变成缓冲 | 工作形态适配器:校验声明、准备输入、跑、提交 | `ooo-patch.ts`、`ooo-session-mechanism.ts`、`evals/ooo-execution/data-check-runner.ts` | +| surface | 共享的、不透明的交换物:role、commit、release、帧回调 | 黑板条目:`payload` 加判别位与摘要;交付是 commit,resolve 与过期是 release,唤醒是帧回调 | `task_board_entries` 及其追加列 | +| 合成器 | 堆叠、遮挡、混合、时序、剔除、以及拥有者不在场时替它演快照 | 板子与程序:认领、租约、围栏、串行与唤醒、过期与回收、合法集 | `src/integration/ooo-board.ts`、`src/integration/ooo-dispatch.ts`、`task-semantics.ts` | +| 显示器 | 模式、刷新率、色彩空间、缩放——同一帧在不同设备上呈现不同 | 呈现端:T0–T3 分层、扁平行、紧凑读、agent 读到的工具描述 | `taskBoardPreview`、分层渲染、`src/prompts/nmg-prompts.yaml` | + +由这条流水线出来三类归属,而它们才是模型真正的内容: + +1. **内容与语义**归拥有者,对中间层**语义不透明**。合成器为了模糊或遮挡形状会读像素,但读不出"这颗按钮是什么意思";要含义它得问拥有者。我们的板子可以读声明的结构,从不读载荷的语义。 +2. **描述**是拥有者的义务:几何、层级与父子关系、形状与不透明区与脏区、动画的**终态**、用于恢复的身份、以及能替拥有者顶一阵的快照。我们这边是判别位、摘要、依赖、scope 与 need、以及协议提供的投影。 +3. **裁定**只归中间层,因为只有它有全局视野:堆叠与遮挡是全局性质,可见性、焦点、捕获权限、时序与回收同理。我们这边是合法集、认领与租约、唤醒、过期。 + +取代口号的那条界线测试: + +- 中间层**不需要**就这个对象推理 → 对象保持不透明,归拥有者; +- 中间层**需要结构** → 拥有者必须交描述,中间层只读描述、不读语义; +- 这个判断**是全局的** → 只能中间层做——因此描述必须廉价且够新,并且"陈旧"要有明确语义(租约过期是怀疑,不是死亡证明)。 + +两条推论在别处已经承重:中间层必须在**与将来怎么呈现无关的形式**里做合成,所以它永远不该学会某个读者的格式;而**呈现不可验证**(显示器会自己调色、裁切、缩放),所以验收只能依赖产物与摘要,绝不能依赖呈现方式。 + +## 为什么它现在还只是草稿 + +- 它是靠**叠比喻**得到的,而每个比喻都已经被更完整的视角部分否掉过:先是内核那套,再是窗口的框与内容,现在是四段。**是比喻在替设计推理**——这一节存在的意义就是记下这个味道。 +- 四段里有一段**只是个名字**。呈现端没有缝、没有归属、没有测试;它是对一堆散落代码的描述。 +- 这一段弧里唯一可测的判断——机制/策略那条缝很窄——**在一分钟内被 grep 否证**:补丁假设坐在六处,其中一处是合法性规则。 +- **没有任何东西变小。** 没改代码、没消除概念,反而加了三份散文记录。按今天的证据,这个模型比它描述的对象**多了**概念。 + +## 资格规则 + +**一个比喻只有在能被一次检查否证时,才有资格进记录。** 在上面的检查完成之前,这个模型只能留在本文;只有当检查结果对它有利,它才可能被提升进记录。若结果不利,本文进归档,而不是被提升。 + +## 检查 + +预测先写下来再测量,这样"意外"是可见的,而不是事后被圆过去的。 + +**(a)中间层里的策略词。** 对中间层——`src/core/store/`、`src/integration/ooo-board.ts`、`src/integration/ooo-dispatch.ts`、`src/integration/task-semantics.ts`、`src/integration/ooo-execution.ts`、`src/integration/ooo-candidate.ts`——grep 这份策略词清单:`patch`、`editable`、`instruction`、`checks`、`files`、`repair-first`。**每一处命中都要带路径列出,不能只给计数**,因为同一个词在一条路径上是机制、在另一条路径上是策略。预测:十五处以上,集中在板子与合法性模块。预期的发现:板子自己的票类型与那个冻结方法就是中心。 + +**(b)两端。** 清点所有"在板子的规范形式与某个生产者或读者格式之间翻译"的站点。预测:生产端**集中且可分**(约四处),呈现端**散落**在 store 的预览函数、分层渲染、提示词文件、CLI 与扩展里,**完全没有缝**。如果呈现端分离不出来,那第四段就只是一个事实的名字,而不是一个设计——那就该丢掉,而不是去建它。 + +**(c)第二个形态。** 让一个不是补丁的单元跑同一条派发循环,用一个最小 resolver(补丁仍是默认;未知形态按名字拒绝,而不是判失败)。数要改几个文件。预测:五个以上。 + +## 变小测试 + +这个模型只能靠"让某样东西变小"来赢得位置,而它必须变小的四样,事先写清: + +- **更少的代码**:接上这条缝意味着中间层**丢掉**一个类型和一个冻结方法,而那套字段校验搬去拥有该形态的适配器。如果净增,模型没通过。 +- **更少的概念**:三类归属必须**吸收**它们所替代的两张清单(机制清单与策略清单),而不是被并排加上去。如果读者还得同时握两套,那这个模型是复述,不是抽象。 +- **更少的改动点**:六处要么变少,要么从六处收成一处。真正重要的数字是:**加一个新形态要碰几处**。 +- **更好的可维护性**:加一个形态只在**一个归属者**名下加文件,中间层的测试一个都不改;中间层那套测试在一个它从没见过的形态下原样通过。 + +## 结果 + +待填。 From 9ea3296ab33070ba6a3830b692e6057061d033aa Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Tue, 22 Sep 2026 21:26:45 +0800 Subject: [PATCH 12/32] docs(design): fill in checks a and b, and correct the display stage the counting exposed Both checks were run against the prediction recorded before them. (a) Policy words in the middle: predicted fifteen or more hits, measured 195 raw - patch 106, files 39, checks 23, editable 14, instruction 12, repair-first 1 - concentrated exactly where predicted, in src/integration/ooo-board.ts (53 patch hits), src/integration/task-semantics.ts (27 patch, 9 editable) and src/core/store/base.ts (13 patch). Two corrections forced by the measurement. The raw count overstates the case, which is the caveat the check was written with: files in src/core/store/writes.ts and retrieval.ts is a mechanism word (a path inside a store) and checks in src/integration/ooo-candidate.ts names the check runner, also mechanism, so the word list drops files and checks and keeps patch, editable and instruction. What is left is still about 120 hits, so "the middle is policy-free" is false as a description of today, an order of magnitude past the prediction, and the model's support is directional: the leak is real, large and concentrated in three files. (b) The two ends: the producer end was predicted concentrated and separable and measured four to six functions over three files (preparePatchWork, patchPrompt, patchCandidate, patchSubmission, snapshotText, runTestFile); the presentation end was predicted scattered with no seam and measured six sites (the preview text, the entry and preview types, the wire shape, the service, the agent-facing renderer, and the generated tool descriptions). Both held. The measurement also corrected this document: the display row claimed T0-T3 tiers, and no such thing exists for a board entry - tieredDisclosure and tier belong to memory retrieval, and a board entry has exactly one presentation, the 200-character preview plus raw fields. That row now says what exists, so the second borrowed vocabulary is recorded rather than left in the table. (c) A second shape is not run yet, and the reduction test is unchanged: nothing has got smaller, the middle's three files hold the available reduction, and whether it pays is what (c) measures. The counting earned its keep twice - it produced the word-and-path rule the earlier prose could not state, and it caught this document borrowing the memory side's vocabulary as if it were the board's. --- docs/design/mechanism-in-the-middle.md | 48 +++++++++++++++++--- docs/design/mechanism-in-the-middle.zh-CN.md | 24 +++++++--- 2 files changed, 58 insertions(+), 14 deletions(-) diff --git a/docs/design/mechanism-in-the-middle.md b/docs/design/mechanism-in-the-middle.md index d692dac1..7a0fc009 100644 --- a/docs/design/mechanism-in-the-middle.md +++ b/docs/design/mechanism-in-the-middle.md @@ -14,12 +14,12 @@ decisions, which live in three proposed records: [mechanism, not policy](../deci A drawing pipeline has four stages, and so does ours. -| Stage | What it does | What it is here | Where it lives today | -| ---------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | -| Producer | turns something into a buffer | a work-shape adapter: validate the declaration, prepare the inputs, run, submit | `ooo-patch.ts`, `ooo-session-mechanism.ts`, `evals/ooo-execution/data-check-runner.ts` | -| Surface | the shared, opaque exchange object with a role, commit, release and frame callbacks | the board entry: `payload` plus a discriminator and a digest, delivery as commit, resolve and expiry as release, wake as the frame callback | `task_board_entries` with its added columns | -| Compositor | stacking, occlusion, blending, timing, culling, and playing a snapshot when the owner is away | the board and the program: claim, lease, fence, serial and wake, expiry and reaping, the legal set | `src/integration/ooo-board.ts`, `src/integration/ooo-dispatch.ts`, `task-semantics.ts` | -| Display | mode, refresh rate, colour space, scaling - the same frame presented differently per device | the presentation end: T0-T3 tiers, flat lines, compact read, the tool descriptions an agent reads | `taskBoardPreview`, the tier rendering, `src/prompts/nmg-prompts.yaml` | +| Stage | What it does | What it is here | Where it lives today | +| ---------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Producer | turns something into a buffer | a work-shape adapter: validate the declaration, prepare the inputs, run, submit | `ooo-patch.ts`, `ooo-session-mechanism.ts`, `evals/ooo-execution/data-check-runner.ts` | +| Surface | the shared, opaque exchange object with a role, commit, release and frame callbacks | the board entry: `payload` plus a discriminator and a digest, delivery as commit, resolve and expiry as release, wake as the frame callback | `task_board_entries` with its added columns | +| Compositor | stacking, occlusion, blending, timing, culling, and playing a snapshot when the owner is away | the board and the program: claim, lease, fence, serial and wake, expiry and reaping, the legal set | `src/integration/ooo-board.ts`, `src/integration/ooo-dispatch.ts`, `task-semantics.ts` | +| Display | mode, refresh rate, colour space, scaling - the same frame presented differently per device | the presentation end: what a reader is shown | one presentation of a board entry exists - `taskBoardPreview`, 200 characters, plus raw fields - at six sites; the tier vocabulary lives only on the memory side (`tieredDisclosure`), never for board entries | Three ownership classes follow, and they are the model's real content: @@ -103,4 +103,38 @@ named in advance: ## Results -Pending. +**(a) Policy words in the middle.** The prediction was fifteen or more hits. Measured 195 raw across the +listed files: `patch` 106, `files` 39, `checks` 23, `editable` 14, `instruction` 12, `repair-first` 1. The +concentration is where it was predicted: `src/integration/ooo-board.ts` carries 53 `patch` hits, +`src/integration/task-semantics.ts` 27 `patch` and 9 `editable`, `src/core/store/base.ts` 13 `patch`. + +The measurement forced two corrections. First, the raw count overstates the case, which is the caveat this +check was written with: `files` in `src/core/store/writes.ts` and `src/core/store/retrieval.ts` is a mechanism +word - a path inside a store - and `checks` in `src/integration/ooo-candidate.ts` names the check _runner_, +which is mechanism too. The word list should therefore drop `files` and `checks` and keep `patch`, `editable` +and `instruction`. Second, what remains is still about 120 hits, so "the middle is policy-free" is false as a +description of today, at a scale an order of magnitude past the prediction. What the model has is directional +support: the leak is real, large, and concentrated in three files, which is exactly what a seam would have to +remove. + +**(b) The two ends.** The producer end was predicted concentrated and separable, and measured four to six +functions over three files: `preparePatchWork`, `patchPrompt`, `patchCandidate` and `patchSubmission` in +`src/integration/ooo-patch.ts`, `snapshotText` in `src/integration/ooo-session-mechanism.ts`, and `runTestFile` +in `evals/ooo-execution/data-check-runner.ts`. The prediction held. The presentation end was predicted +scattered with no seam, and measured six sites: the preview text in `src/core/store/base.ts`, the entry and +preview types in `src/core/types.ts`, the wire shape in `src/cli/protocol.ts`, the service in +`src/cli/service.ts`, the agent-facing renderer in `.pi/extensions/nmg/index.ts`, and the generated tool +descriptions from `src/prompts/nmg-prompts.yaml`. That prediction held too, with one correction against this +document - see the display row above. A board entry has exactly one presentation, so the fourth stage is +thinner than the table first claimed, and the tier vocabulary in it had been borrowed from the memory side +rather than found on the board side. + +**(c) A second shape.** Not run yet. + +## What the two checks say about the reduction test + +Nothing has got smaller yet, and the checks say where a reduction would have to come from: three files in the +middle hold about a hundred and twenty policy-sense hits, and six sites present a board entry with no owner +between them. The reduction available today is in the middle's three files; whether it pays is what check (c) +measures. The counting earned its keep twice: it produced the word-and-path rule that the earlier prose could +not state, and it caught this document borrowing another subsystem's vocabulary as if it were the board's. diff --git a/docs/design/mechanism-in-the-middle.zh-CN.md b/docs/design/mechanism-in-the-middle.zh-CN.md index f0fcdd7e..9034c7da 100644 --- a/docs/design/mechanism-in-the-middle.zh-CN.md +++ b/docs/design/mechanism-in-the-middle.zh-CN.md @@ -9,12 +9,12 @@ 一条绘图流水线有四段,我们也是。 -| 段 | 它做什么 | 我们这边的对应 | 今天在哪 | -| ------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | -| 生产者 | 把某样东西变成缓冲 | 工作形态适配器:校验声明、准备输入、跑、提交 | `ooo-patch.ts`、`ooo-session-mechanism.ts`、`evals/ooo-execution/data-check-runner.ts` | -| surface | 共享的、不透明的交换物:role、commit、release、帧回调 | 黑板条目:`payload` 加判别位与摘要;交付是 commit,resolve 与过期是 release,唤醒是帧回调 | `task_board_entries` 及其追加列 | -| 合成器 | 堆叠、遮挡、混合、时序、剔除、以及拥有者不在场时替它演快照 | 板子与程序:认领、租约、围栏、串行与唤醒、过期与回收、合法集 | `src/integration/ooo-board.ts`、`src/integration/ooo-dispatch.ts`、`task-semantics.ts` | -| 显示器 | 模式、刷新率、色彩空间、缩放——同一帧在不同设备上呈现不同 | 呈现端:T0–T3 分层、扁平行、紧凑读、agent 读到的工具描述 | `taskBoardPreview`、分层渲染、`src/prompts/nmg-prompts.yaml` | +| 段 | 它做什么 | 我们这边的对应 | 今天在哪 | +| ------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | +| 生产者 | 把某样东西变成缓冲 | 工作形态适配器:校验声明、准备输入、跑、提交 | `ooo-patch.ts`、`ooo-session-mechanism.ts`、`evals/ooo-execution/data-check-runner.ts` | +| surface | 共享的、不透明的交换物:role、commit、release、帧回调 | 黑板条目:`payload` 加判别位与摘要;交付是 commit,resolve 与过期是 release,唤醒是帧回调 | `task_board_entries` 及其追加列 | +| 合成器 | 堆叠、遮挡、混合、时序、剔除、以及拥有者不在场时替它演快照 | 板子与程序:认领、租约、围栏、串行与唤醒、过期与回收、合法集 | `src/integration/ooo-board.ts`、`src/integration/ooo-dispatch.ts`、`task-semantics.ts` | +| 显示器 | 模式、刷新率、色彩空间、缩放——同一帧在不同设备上呈现不同 | 呈现端:读者被展示到什么 | 黑板条目的呈现只有一种——`taskBoardPreview`,200 字符,加原始字段——被六个站点消费;分层的词汇只活在记忆那一侧(`tieredDisclosure`),从不用于黑板条目 | 由这条流水线出来三类归属,而它们才是模型真正的内容: @@ -62,4 +62,14 @@ ## 结果 -待填。 +**(a)中间层里的策略词。** 预测十五处以上。实测六个文件里共 195 处:`patch` 106、`files` 39、`checks` 23、`editable` 14、`instruction` 12、`repair-first` 1。集中处与预测一致:`src/integration/ooo-board.ts` 带 53 处 `patch`,`src/integration/task-semantics.ts` 27 处 `patch` 与 9 处 `editable`,`src/core/store/base.ts` 13 处 `patch`。 + +测量迫使两条修正。**第一,原始计数高估了情况**——而这正是写这条检查时要带上的那个前提:`src/core/store/writes.ts` 与 `src/core/store/retrieval.ts` 里的 `files` 是机制词(存储里的路径),`src/integration/ooo-candidate.ts` 里的 `checks` 指的是检查的**运行器**,也是机制。所以词表应**去掉 `files` 与 `checks`**,只留 `patch`、`editable`、`instruction`。**第二,剩下的仍有约 120 处**,所以「中间层是无策略的」作为**今天的描述**是假的,而且超出一个数量级。模型拿到的是**方向性支持**:泄漏真实、量大、集中在三个文件里——那正是接上一条缝要拿掉的东西。 + +**(b)两端。** 生产端预测「集中且可分」,实测三到六个函数散在三个文件:`src/integration/ooo-patch.ts` 里的 `preparePatchWork`、`patchPrompt`、`patchCandidate`、`patchSubmission`,`src/integration/ooo-session-mechanism.ts` 里的 `snapshotText`,`evals/ooo-execution/data-check-runner.ts` 里的 `runTestFile`。预测成立。呈现端预测「散落、没有缝」,实测六个站点:`src/core/store/base.ts` 的预览文本、`src/core/types.ts` 的条目与预览类型、`src/cli/protocol.ts` 的线上形状、`src/cli/service.ts` 的服务、`.pi/extensions/nmg/index.ts` 里面向 agent 的渲染、以及由 `src/prompts/nmg-prompts.yaml` 生成的工具描述。这条预测也成立,但有**一处对本文的修正**——见上表「显示器」一行:黑板条目只有一种呈现,所以第四段比表里最初写的更薄,而它里面的分层词汇是从记忆那一侧**借**来的,不是黑板这一侧**找到**的。 + +**(c)第二个形态。** 还没跑。 + +## 这两条检查对变小测试说了什么 + +现在还什么都没变小,而检查指出了变小只能从哪里来:中间层三个文件里约 120 处策略语义命中,以及六个站点共同呈现一个黑板条目、彼此之间没有归属者。今天唯一可取的变小是中间层那三个文件;它是否值得,由检查(c)测。这次清点两次证明了它的价值:它产出了散文说不出的「词 + 路径」规则,并且抓出本文把另一个子系统的词汇当黑板自己的用。 From d56270d0710448fdcf896fd4ceab8a893f4d8786 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Tue, 22 Sep 2026 22:49:19 +0800 Subject: [PATCH 13/32] Give the board entry two faces, and the read face one home The model in mechanism-in-the-middle.md is retired and replaced: a board entry is a medium with a write face and a read face, the compositor is brought by a protocol rather than by the board, and the roles are roles rather than files - src/core/store/base.ts is the medium, the compositor and the reader's rule at once, and src/cli/service.ts is both the wire face and a formatter. The rule that decides where a seam goes needs no taste, only a count: a boundary is built when it has a second implementation, and only declared with one. The read face has three, so it lands here. src/core/board-entry-view.ts owns the rule, the store's compact read and the host adapter's broadcast both call it, and the store's private copy and the adapter's bare slice are gone. One of the three was not merely duplicated but wrong - the adapter's 140-character slice did not collapse whitespace, so a multi-line body reached a broadcast as a multi-line broadcast. The write face's type and the middle's seam are declared and not built, each with a named trigger. Check (c) measured why: making the loop shape-agnostic is one file with fifteen insertions and eighteen deletions, and tsc then reports zero errors while twenty-five tests fail, every one of them because the claim admits no work. The coupling is invisible to the compiler and fatal at runtime, so the seam is prepaid cost until an adopter declares a non-patch unit. Docs and code land together, because the read face's home is what the two-face model implements. --- .pi/extensions/nmg/index.ts | 3 +- docs/design/mechanism-in-the-middle.md | 167 ++++++++++++------- docs/design/mechanism-in-the-middle.zh-CN.md | 74 +++++--- src/core/board-entry-view.ts | 18 ++ src/core/store/base.ts | 13 +- tests/core/task-board.test.ts | 12 ++ 6 files changed, 193 insertions(+), 94 deletions(-) create mode 100644 src/core/board-entry-view.ts diff --git a/.pi/extensions/nmg/index.ts b/.pi/extensions/nmg/index.ts index cfc9e275..239dd179 100644 --- a/.pi/extensions/nmg/index.ts +++ b/.pi/extensions/nmg/index.ts @@ -38,6 +38,7 @@ import { loadPrompts, renderDisclosure } from "../../../src/prompts/load.ts"; import { memoryDisclosureEntries } from "../../../src/integration/search-projection.ts"; import { PI_BOARD_ACTIONS, PI_REMEMBER_ACTIONS } from "../../../src/integration/tool-contract.ts"; import { TASK_BOARD_VERDICTS } from "../../../src/core/types.ts"; +import { boardEntryView } from "../../../src/core/board-entry-view.ts"; import { resolveSkillOptPolicyChannels } from "../../../src/lab/skillopt-policy.ts"; import type { ActiveGraphBudget, @@ -2998,7 +2999,7 @@ export async function maybeBroadcastToWorld(input: { entryIds: [entry.id], })) as { delivered: string[]; suppressed: boolean }; if (worldCheck.delivered.includes(entry.id)) return false; - const excerpt = entry.content.length > 140 ? `${entry.content.slice(0, 140)}…` : entry.content; + const excerpt = boardEntryView(entry.content, 140); const label = kindLabel(entry.kind); const broadcast = `[NMG board 协作广播] 频道 ${entry.taskId} 有 #${entry.id} 未认领的${label}(open):${excerpt}。有空的 agent 可用 nmg_board read taskId=${entry.taskId} 查看详情、claim 认领处理。`; await invoke("taskBoard", { diff --git a/docs/design/mechanism-in-the-middle.md b/docs/design/mechanism-in-the-middle.md index 7a0fc009..d87af946 100644 --- a/docs/design/mechanism-in-the-middle.md +++ b/docs/design/mechanism-in-the-middle.md @@ -2,62 +2,93 @@ **Status:** draft **Created:** 2026-09-21 +**Updated:** 2026-09-21 -A layered model of the collaboration protocol, and the checks that decide whether the model earns a place in -the records. This document owns the model, the boundary test, and the checks. It does not restate the parts -inventory, which lives in [protocol-governed-collaboration.md](protocol-governed-collaboration.md), nor the -decisions, which live in three proposed records: [mechanism, not policy](../decisions/proposed/2026-09-21-mechanism-not-policy.md), +A model of how a board entry travels, and the checks that decide whether the model earns a place in the +records. This document owns the model, the rule that says when a boundary earns a seam, and the checks. It +does not restate the parts inventory, which lives in [protocol-governed-collaboration.md](protocol-governed-collaboration.md), +nor the decisions, which live in three proposed records: [mechanism, not policy](../decisions/proposed/2026-09-21-mechanism-not-policy.md), [the frame and its storage](../decisions/proposed/2026-09-21-the-frame-and-its-storage.md), and [the program answers legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md). ## The model -A drawing pipeline has four stages, and so does ours. - -| Stage | What it does | What it is here | Where it lives today | -| ---------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Producer | turns something into a buffer | a work-shape adapter: validate the declaration, prepare the inputs, run, submit | `ooo-patch.ts`, `ooo-session-mechanism.ts`, `evals/ooo-execution/data-check-runner.ts` | -| Surface | the shared, opaque exchange object with a role, commit, release and frame callbacks | the board entry: `payload` plus a discriminator and a digest, delivery as commit, resolve and expiry as release, wake as the frame callback | `task_board_entries` with its added columns | -| Compositor | stacking, occlusion, blending, timing, culling, and playing a snapshot when the owner is away | the board and the program: claim, lease, fence, serial and wake, expiry and reaping, the legal set | `src/integration/ooo-board.ts`, `src/integration/ooo-dispatch.ts`, `task-semantics.ts` | -| Display | mode, refresh rate, colour space, scaling - the same frame presented differently per device | the presentation end: what a reader is shown | one presentation of a board entry exists - `taskBoardPreview`, 200 characters, plus raw fields - at six sites; the tier vocabulary lives only on the memory side (`tieredDisclosure`), never for board entries | - -Three ownership classes follow, and they are the model's real content: - -1. **Content and semantics** belong to the owner and are semantically opaque to the middle. The compositor - reads pixels to blur or to know an occlusion shape, but it cannot read what a button means; for meaning it - asks the owner. Our board may read declared structure and never reads payload semantics. -2. **Description** is the owner's obligation: geometry, level and parent relationships, shape and opaque and - dirty regions, the terminal state of an animation, an identity for restoration, and a snapshot able to - stand in for the owner. Ours is the discriminator, the digest, dependencies, scope and need, and the - projection a protocol supplies. -3. **Arbitration** is only the middle's, because only the middle has the global view: stacking and occlusion - are global properties, and so are visibility, focus, capture permission, timing and reaping. Ours is the - legal set, claim and lease, wake, expiry. - -The boundary test that replaces a slogan: - -- the middle does not need to reason about the object, so the object stays opaque and belongs to its owner; -- the middle needs structure, so the owner must supply a description and the middle reads the description - but not the semantics; -- the judgement is global, so only the middle can make it - which requires the description to be cheap and - fresh, and requires staleness to have a meaning (an expired lease is a suspicion, not a death certificate). - -Two consequences are already load-bearing elsewhere. The middle must compose in a form that is independent -of how it will be shown, so it never learns a reader's format. And presentation is not verifiable - a display -applies its own colour, crop and scaling - so an acceptance can only ever rest on the artifact and its digest, -never on how it was shown. +A board entry is a medium with two faces, and the roles around it are three. + +``` + producer ──[ producer-side interface · write ]──┐ ┌──[ presentation interface · read ]──▶ reader + ▼ ▲ + ┌────────────────────────────────────┐ + │ board entry (the medium) │ + │ write face read face │ + └────────────────────────────────────┘ + │ ▲ + Task-Unit Protocol ──▶ compositor +``` + +| Edge | What it is | Where it lives today | State | +| -------------- | ------------------------------------- | ------------------------------------------------------ | ------------------------------------------------------------------------------------------- | +| the write face | what a producer puts on the medium | `put` and `deliver`; a body arrives as `content` | three fillers - a patch artifact, a human broadcast, a `memory=` pointer - and no shape | +| the read face | what a reader is shown | `boardEntryView` in `src/core/board-entry-view.ts` | one rule, one home, one test | +| the middle | the compositor, brought by a protocol | `ooo-board.ts`, `ooo-dispatch.ts`, `task-semantics.ts` | one protocol brings it: the Task-Unit Protocol | + +Three facts about the shape, each measured rather than assumed. + +1. **The middle comes from a protocol, not from the board.** The board offers two faces and nothing else; a + protocol names the constraints that put a compositor between them, and the Task-Unit Protocol is the one + that brings ours. That is why "the program only answers legality" says the same thing: no declared + legality, no judgement. Of the seven entry kinds, only `handoff` and `result` are ever claimed, delivered + and judged - a note, a goal and a question travel from the write face straight to the read face. +2. **The roles are not files.** `src/core/store/base.ts` is the medium (columns and transactions), the + compositor (the lease-based claim compare-and-set and the clock) and the reader's rule at once, and + `src/cli/service.ts` is both the wire face and a formatter. So this is a map of roles, and a check can sit + only on the artifacts that cross an edge, never on module imports. +3. **Feedback is a replay, not a reversal.** A reader becomes the next writer and enters through the write + face again; the middle is never traversed backwards. The broadcast path in `.pi/extensions/nmg/index.ts` + (it reads an entry, then writes a new one) is a place where two roles live side by side, and rejection + followed by re-issue moves forward as well - the middle writes a new declaration onto the medium and the + producer reads it from the write face. + +## Which boundary earns a seam + +**A boundary is built when it has a second implementation; with one, it is only declared.** The rule needs no +taste, only a count. + +| Boundary | Implementations today | Consequence | +| -------------- | ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | +| the write face | three fillers for one body | declared: a type is worth writing when a fourth filler appears | +| the read face | three truncation rules - 200 characters collapsing, a generic helper, and a bare 140 slice that collapsed nothing | **built**: one rule, one home, one test | +| the middle | one protocol | declared only: a seam is built when a second protocol brings a second middle | + +The alternative - building every seam now, each against a single implementation - is the mistake this +document already made once by stacking analogies. A seam designed against one shape is a seam designed +against an imagined second shape, and the imagined one is the one that gets built wrong. + +## What landed with this document + +- **The read face has one home.** `src/core/board-entry-view.ts` owns the rule: a lone `memory=` pointer + is returned whole, anything longer is collapsed to one bounded line, and the bound belongs to the caller, + because how much of a body fits is the reader's decision. The store's compact read and the host adapter's + broadcast both call it; the store's private copy and the adapter's bare slice are gone. One of the three + rules was also wrong rather than merely duplicated: the 140-character slice in `.pi/extensions/nmg/index.ts` + did not collapse whitespace, so a multi-line body reached a broadcast as a multi-line broadcast. + `tests/core/task-board.test.ts` pins the rule directly. The adapter's `excerpt` helper stays for memory + results, which are not board entries and keep their own rule. +- **The write face's type did not land.** The three fillers are three kinds of entry rather than one shape + drawn three ways, so the trigger for a type is a fourth filler. +- **The middle's seam did not land.** Check (c) says why. ## Why this is a draft and not a record -- It was reached by stacking analogies, and each analogy was already partly overruled by a fuller view: first - a kernel split, then a window split into frame and content, now four stages. The analogies are doing the - reasoning, and that is the smell this section exists to record. -- One of the four stages is a name and nothing else. The presentation end has no seam, no owner and no tests; - it is a description of scattered code. +- The pipeline form of this document - four stages borrowed from a drawing pipeline - was reached by stacking + analogies and was **retired on 2026-09-21** in favour of the two-faces model above, which came from + measurement instead: roles cohabit files, feedback replays rather than reverses, and a protocol is what + brings the middle. - The only measurable claim made in this arc - that the mechanism/policy seam was narrow - was falsified by a grep within a minute: the patch assumption sits in six places, one of them a legality rule. -- Nothing has got smaller. No code changed, no concept was removed, and three prose records were added. On - today's evidence the model is strictly more concepts than the thing it describes. +- One thing has got smaller: the read face's three rules are now one, in one home, with one test. The middle's + hundred-and-twenty policy-sense hits and the write face's unshaped body have not. +- The model is still more concepts than the code it describes, so it stays a draft. ## The eligibility rule @@ -125,16 +156,36 @@ scattered with no seam, and measured six sites: the preview text in `src/core/st preview types in `src/core/types.ts`, the wire shape in `src/cli/protocol.ts`, the service in `src/cli/service.ts`, the agent-facing renderer in `.pi/extensions/nmg/index.ts`, and the generated tool descriptions from `src/prompts/nmg-prompts.yaml`. That prediction held too, with one correction against this -document - see the display row above. A board entry has exactly one presentation, so the fourth stage is -thinner than the table first claimed, and the tier vocabulary in it had been borrowed from the memory side -rather than found on the board side. - -**(c) A second shape.** Not run yet. - -## What the two checks say about the reduction test - -Nothing has got smaller yet, and the checks say where a reduction would have to come from: three files in the -middle hold about a hundred and twenty policy-sense hits, and six sites present a board entry with no owner -between them. The reduction available today is in the middle's three files; whether it pays is what check (c) -measures. The counting earned its keep twice: it produced the word-and-path rule that the earlier prose could -not state, and it caught this document borrowing another subsystem's vocabulary as if it were the board's. +document: a board entry has exactly one presentation, and the tier vocabulary in the old table had been +borrowed from the memory side rather than found on the board side. A later pass over the same sites found +what the count alone could not: the six sites are a chain, not six formatters, and the duplication sits in +three of them - the store's preview rule, the adapter's generic helper, and a bare slice in the adapter's +broadcast path. + +**(c) A second shape.** The prediction was five or more files. The change was made on a throwaway branch and +then reverted: the loop was made shape-agnostic by replacing `DispatchTicket.patch?: PatchWork` with an opaque +`declaration?: {shape, digest, frozen}`, pointing `PlanWorker` at that declaration, and deleting the +`preparePatchWork` call from the unit dispatcher - one file, fifteen insertions and eighteen deletions. Two +results followed, and neither was the result the prediction was aiming at. + +- `tsc --noEmit` reported **zero errors**. The coupling is invisible to the compiler: the field is optional, + and structural typing still lets the board's ticket, which carries `patch`, satisfy the port, so the change + produced no compile-time signal at all. +- The suite reported **25 failing tests**, every one with the same cause - "the claim admits no work". A second + shape today would not fail as a new shape; it would refuse every unit. + +The coupling's surface, counted by grep: three source files name the frozen patch work (`ooo-dispatch.ts`, +`ooo-patch.ts`, `ooo-session-mechanism.ts`, the last with about ten uses), the board must change although it +never reads the field, seven test files read `ticket.patch` directly, and the worker side couples through the +`PlanWorker` type, which breaks the driver's three worker kinds. The prediction held and then some, and the +decision it forced is the opposite of building: the seam is prepaid cost, deferred with a named trigger - the +first non-patch unit an adopter declares - rather than built now against a single shape. + +## What the checks say about the reduction test + +One thing is smaller: three rules for showing a board entry became one rule in one home with one test, and the +rule that disappeared was the wrong one. Everything else stands. The middle's three files still hold about a +hundred and twenty policy-sense hits, and the write face's body still has no shape. The counting earned its +keep three times: it produced the word-and-path rule that the earlier prose could not state, it caught this +document borrowing another subsystem's vocabulary as if it were the board's, and it turned "should the seam +exist" from a matter of taste into a measured file count. diff --git a/docs/design/mechanism-in-the-middle.zh-CN.md b/docs/design/mechanism-in-the-middle.zh-CN.md index 9034c7da..a5ce8203 100644 --- a/docs/design/mechanism-in-the-middle.zh-CN.md +++ b/docs/design/mechanism-in-the-middle.zh-CN.md @@ -2,40 +2,61 @@ **Status:** draft **Created:** 2026-09-21 +**Updated:** 2026-09-21 -协作协议的分层模型,以及决定这个模型有没有资格进记录的那几项检查。本文拥有模型、那条界线测试、以及这些检查。它不复述大类清单(那在 [protocol-governed-collaboration.md](protocol-governed-collaboration.zh-CN.md)),也不复述决策(那在三份提案里:[机制,不是策略](../decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md)、[帧、它的数据格式与它的存储](../decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md)、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md))。 +一个黑板条目怎么走,以及决定这个模型有没有资格进记录的那几项检查。本文拥有模型、"一条边界什么时候够资格建缝"的那条规则、以及这些检查。它不复述大类清单(那在 [protocol-governed-collaboration.md](protocol-governed-collaboration.zh-CN.md)),也不复述决策(那在三份提案里:[机制,不是策略](../decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md)、[帧、它的数据格式与它的存储](../decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md)、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md))。 ## 模型 -一条绘图流水线有四段,我们也是。 +一个黑板条目是一个**有正反两面的介质**,围着它的角色有三个。 -| 段 | 它做什么 | 我们这边的对应 | 今天在哪 | -| ------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | -| 生产者 | 把某样东西变成缓冲 | 工作形态适配器:校验声明、准备输入、跑、提交 | `ooo-patch.ts`、`ooo-session-mechanism.ts`、`evals/ooo-execution/data-check-runner.ts` | -| surface | 共享的、不透明的交换物:role、commit、release、帧回调 | 黑板条目:`payload` 加判别位与摘要;交付是 commit,resolve 与过期是 release,唤醒是帧回调 | `task_board_entries` 及其追加列 | -| 合成器 | 堆叠、遮挡、混合、时序、剔除、以及拥有者不在场时替它演快照 | 板子与程序:认领、租约、围栏、串行与唤醒、过期与回收、合法集 | `src/integration/ooo-board.ts`、`src/integration/ooo-dispatch.ts`、`task-semantics.ts` | -| 显示器 | 模式、刷新率、色彩空间、缩放——同一帧在不同设备上呈现不同 | 呈现端:读者被展示到什么 | 黑板条目的呈现只有一种——`taskBoardPreview`,200 字符,加原始字段——被六个站点消费;分层的词汇只活在记忆那一侧(`tieredDisclosure`),从不用于黑板条目 | +``` + 生产者 ──[ 生产者端 · 写 ]──┐ ┌──[ 呈现端 · 读 ]──▶ 读者 + ▼ ▲ + ┌──────────────────────────────────┐ + │ 黑板条目(介质) │ + │ 写面 读面 │ + └──────────────────────────────────┘ + │ ▲ + 任务单元协议 ──▶ 裁定者(合成) +``` -由这条流水线出来三类归属,而它们才是模型真正的内容: +| 边 | 是什么 | 今天在哪 | 状态 | +| ---- | ------------------------ | ------------------------------------------------------ | -------------------------------------------------------------- | +| 写面 | 生产者放到介质上的东西 | `put` 与 `deliver`;内容以 `content` 到达 | 三种填法——补丁产物、人类广播、`memory=` 指针——**没有形状** | +| 读面 | 读者被展示到什么 | `src/core/board-entry-view.ts` 的 `boardEntryView` | **一条规则、一个家、一个测试** | +| 中间 | 裁定者,由**协议**带进来 | `ooo-board.ts`、`ooo-dispatch.ts`、`task-semantics.ts` | 只有一份协议带它:任务单元协议 | -1. **内容与语义**归拥有者,对中间层**语义不透明**。合成器为了模糊或遮挡形状会读像素,但读不出"这颗按钮是什么意思";要含义它得问拥有者。我们的板子可以读声明的结构,从不读载荷的语义。 -2. **描述**是拥有者的义务:几何、层级与父子关系、形状与不透明区与脏区、动画的**终态**、用于恢复的身份、以及能替拥有者顶一阵的快照。我们这边是判别位、摘要、依赖、scope 与 need、以及协议提供的投影。 -3. **裁定**只归中间层,因为只有它有全局视野:堆叠与遮挡是全局性质,可见性、焦点、捕获权限、时序与回收同理。我们这边是合法集、认领与租约、唤醒、过期。 +关于这个形状的三条事实,每一条都是**量出来的**,不是假定的。 -取代口号的那条界线测试: +1. **中间来自协议,不来自黑板。** 黑板只提供两个面;是**协议**点名了那些约束,才把裁定者接在两面之间,而带进我们这一段的是**任务单元协议**。所以"程序只回答合法性"说的是同一件事:没有声明的合法性,就没有裁定。七种条目类型里只有 `handoff` 与 `result` 会被认领、交付与裁决——`note`、`goal`、`question` 从写面直接到读面。 +2. **角色不是文件。** `src/core/store/base.ts` 同时是介质(列与事务)、裁定者(基于租约的认领 compare-and-set 与时钟)和读者的那条规则;`src/cli/service.ts` 同时是线上面与格式化器。所以这是一张**角色图**,检查只能落在**跨边的产物**上,绝不能落在 import 邻接上。 +3. **反馈是重放,不是倒放。** 读者成为下一个写者,再次从**写面**进来;中间从不被反向穿过。`.pi/extensions/nmg/index.ts` 的广播那条路(先读一个条目、再写一个新条目)就是两个角色并排住在一处的地方;而拒绝之后的重发同样是向前走——中间把新声明写到介质上,生产者从写面把它读走。 -- 中间层**不需要**就这个对象推理 → 对象保持不透明,归拥有者; -- 中间层**需要结构** → 拥有者必须交描述,中间层只读描述、不读语义; -- 这个判断**是全局的** → 只能中间层做——因此描述必须廉价且够新,并且"陈旧"要有明确语义(租约过期是怀疑,不是死亡证明)。 +## 哪条边界够资格建缝 -两条推论在别处已经承重:中间层必须在**与将来怎么呈现无关的形式**里做合成,所以它永远不该学会某个读者的格式;而**呈现不可验证**(显示器会自己调色、裁切、缩放),所以验收只能依赖产物与摘要,绝不能依赖呈现方式。 +**一条边界只有在出现第二个实现时才建;只有一个实现时,只声明。** 这条规则不需要品味,只需要一个计数。 + +| 边界 | 今天的实现数 | 后果 | +| ---- | ------------------------------------------------------------------------------- | ------------------------------------------ | +| 写面 | 一个条目内容有三种填法 | 只声明:等到出现**第四种**填法再写类型 | +| 读面 | 三条截断规则——200 字符折叠、一个通用助手、以及一条**连空白都不折叠**的 140 裸切 | **建**:一条规则、一个家、一个测试 | +| 中间 | 一份协议 | 只声明:等第二份协议带来第二段中间,再建缝 | + +另一种做法——现在就把每条缝都建起来、每条缝对着一个实现——正是本文已经犯过一次的错误:**叠比喻**。对着一个实现设计出来的缝,是**对着想象出来的第二个实现**设计的,而被想错的那个,恰恰就是真会被建错的那个。 + +## 随本文一起落地的东西 + +- **读面有了一个家。** `src/core/board-entry-view.ts` 拥有这条规则:孤立的 `memory=` 指针原样返回;更长的内容折叠成**一行有界文本**;界**归调用方**,因为"塞得下多少"是读者的决定。store 的紧凑读与宿主适配器的广播都调它;store 里那份私有副本和适配器里那条裸切都删掉了。三条规则里有一条不只是重复,而是**错的**:`.pi/extensions/nmg/index.ts` 那条 140 字符裸切不折叠空白,于是一个多行内容会以多行广播的形态发出去。`tests/core/task-board.test.ts` 直接钉住这条规则。适配器里的 `excerpt` 助手留给记忆结果——那不是黑板条目,保留它自己的规则。 +- **写面的类型没有落地。** 三种填法是三种**条目**,不是同一个形状画了三遍;所以写类型的触发条件是出现第四种填法。 +- **中间的缝没有落地。** 检查(c)说明了为什么。 ## 为什么它现在还只是草稿 -- 它是靠**叠比喻**得到的,而每个比喻都已经被更完整的视角部分否掉过:先是内核那套,再是窗口的框与内容,现在是四段。**是比喻在替设计推理**——这一节存在的意义就是记下这个味道。 -- 四段里有一段**只是个名字**。呈现端没有缝、没有归属、没有测试;它是对一堆散落代码的描述。 +- 本文的**流水线形态**(从绘图流水线借来的四段)是靠**叠比喻**得到的,并于 2026-09-21 被上面那个**两个面**的模型**退役**——后者来自测量而不是比喻:角色会同住文件、反馈是重放不是倒放、中间是协议带进来的。 - 这一段弧里唯一可测的判断——机制/策略那条缝很窄——**在一分钟内被 grep 否证**:补丁假设坐在六处,其中一处是合法性规则。 -- **没有任何东西变小。** 没改代码、没消除概念,反而加了三份散文记录。按今天的证据,这个模型比它描述的对象**多了**概念。 +- **有一件事变小了**:读面的三条规则变成了一条,一个家,一个测试。中间层约 120 处策略语义命中、以及写面没有形状的条目内容,都还没变。 +- 这个模型仍然比它描述的代码**多了**概念,所以它还是草稿。 ## 资格规则 @@ -66,10 +87,15 @@ 测量迫使两条修正。**第一,原始计数高估了情况**——而这正是写这条检查时要带上的那个前提:`src/core/store/writes.ts` 与 `src/core/store/retrieval.ts` 里的 `files` 是机制词(存储里的路径),`src/integration/ooo-candidate.ts` 里的 `checks` 指的是检查的**运行器**,也是机制。所以词表应**去掉 `files` 与 `checks`**,只留 `patch`、`editable`、`instruction`。**第二,剩下的仍有约 120 处**,所以「中间层是无策略的」作为**今天的描述**是假的,而且超出一个数量级。模型拿到的是**方向性支持**:泄漏真实、量大、集中在三个文件里——那正是接上一条缝要拿掉的东西。 -**(b)两端。** 生产端预测「集中且可分」,实测三到六个函数散在三个文件:`src/integration/ooo-patch.ts` 里的 `preparePatchWork`、`patchPrompt`、`patchCandidate`、`patchSubmission`,`src/integration/ooo-session-mechanism.ts` 里的 `snapshotText`,`evals/ooo-execution/data-check-runner.ts` 里的 `runTestFile`。预测成立。呈现端预测「散落、没有缝」,实测六个站点:`src/core/store/base.ts` 的预览文本、`src/core/types.ts` 的条目与预览类型、`src/cli/protocol.ts` 的线上形状、`src/cli/service.ts` 的服务、`.pi/extensions/nmg/index.ts` 里面向 agent 的渲染、以及由 `src/prompts/nmg-prompts.yaml` 生成的工具描述。这条预测也成立,但有**一处对本文的修正**——见上表「显示器」一行:黑板条目只有一种呈现,所以第四段比表里最初写的更薄,而它里面的分层词汇是从记忆那一侧**借**来的,不是黑板这一侧**找到**的。 +**(b)两端。** 生产端预测「集中且可分」,实测三到六个函数散在三个文件:`src/integration/ooo-patch.ts` 里的 `preparePatchWork`、`patchPrompt`、`patchCandidate`、`patchSubmission`,`src/integration/ooo-session-mechanism.ts` 里的 `snapshotText`,`evals/ooo-execution/data-check-runner.ts` 里的 `runTestFile`。预测成立。呈现端预测「散落、没有缝」,实测六个站点:`src/core/store/base.ts` 的预览文本、`src/core/types.ts` 的条目与预览类型、`src/cli/protocol.ts` 的线上形状、`src/cli/service.ts` 的服务、`.pi/extensions/nmg/index.ts` 里面向 agent 的渲染、以及由 `src/prompts/nmg-prompts.yaml` 生成的工具描述。这条预测也成立,但有**一处对本文的修正**:黑板条目只有一种呈现,而旧表里的分层词汇是从记忆那一侧**借**来的,不是黑板这一侧**找到**的。对同一批站点后来的一遍检查看到计数看不到的东西:那六个站点是一条**链**,不是六个格式化器,而重复就坐在其中三处——store 的预览规则、适配器的通用助手、以及适配器广播路径里的那条裸切。 + +**(c)第二个形态。** 预测五个以上文件。改动在一个**用完即弃的分支**上做完,然后退回:把循环改成形态无关——`DispatchTicket.patch?: PatchWork` 换成不透明的 `declaration?: {shape, digest, frozen}`,`PlanWorker` 收这个声明,并从单元派发里删掉 `preparePatchWork` 那次调用——一个文件,15 增 18 删。得到两个结果,而两个都不是预测瞄着的那个。 + +- `tsc --noEmit` 报 **0 个错误**。这个耦合**对编译器隐形**:字段是**可选**的,而结构类型让带着 `patch` 的产品板 ticket 依然满足这个端口,所以这次改动**没有任何编译期信号**。 +- 测试报 **25 个用例失败**,每一个的原因都相同——"the claim admits no work"。今天加第二个形态,不会是"新形态失败",而是**每一个单元一起拒绝工作**。 -**(c)第二个形态。** 还没跑。 +耦合的面,按 grep 数:三个源文件点名那份冻结的补丁工作(`ooo-dispatch.ts`、`ooo-patch.ts`、`ooo-session-mechanism.ts`,最后一个约十处),板子**虽然从不读这个字段也必须改**,七个测试文件直接读 `ticket.patch`,而 worker 一侧通过 `PlanWorker` 类型耦合——它打断驱动里的三种 worker。预测成立,而且不止于此;它逼出来的决定与"建"相反:这条缝是**预付成本**,被**带着一个点名的触发条件**推迟——adopter 声明的第一个非补丁单元——而不是现在对着一个实现把它建起来。 -## 这两条检查对变小测试说了什么 +## 这些检查对变小测试说了什么 -现在还什么都没变小,而检查指出了变小只能从哪里来:中间层三个文件里约 120 处策略语义命中,以及六个站点共同呈现一个黑板条目、彼此之间没有归属者。今天唯一可取的变小是中间层那三个文件;它是否值得,由检查(c)测。这次清点两次证明了它的价值:它产出了散文说不出的「词 + 路径」规则,并且抓出本文把另一个子系统的词汇当黑板自己的用。 +**有一件事变小了**:展示一个黑板条目的三条规则,变成一条规则、一个家、一个测试,而且消失的那条正是错的那条。其余照旧。中间层那三个文件里仍有约 120 处策略语义命中,写面的条目内容仍然没有形状。这次清点三度证明了它的价值:它产出了散文说不出的「词 + 路径」规则;它抓出本文把另一个子系统的词汇当黑板自己的用;它把"这条缝该不该存在"从品味之争变成了一个量出来的文件数。 diff --git a/src/core/board-entry-view.ts b/src/core/board-entry-view.ts new file mode 100644 index 00000000..46889b4b --- /dev/null +++ b/src/core/board-entry-view.ts @@ -0,0 +1,18 @@ +/** + * The read face's one rule for a board entry's body. + * + * An entry's body is opaque to the middle and may be long: a memory pointer, a note, a delivered + * artifact, a broadcast that quotes another entry. Whoever shows an entry to a reader shapes that body + * the same way, so the rule lives here once, and both the store's compact read and the host adapters use + * it instead of each writing its own truncation. A lone `memory=` pointer is already the intended + * low-context form and is returned whole; anything longer is collapsed to a single bounded line, so a + * read never carries a body its reader did not ask for. + * + * The bound is the caller's, because how much of a body fits is the reader's decision. What a body looks + * like when it is shown is this rule's. + */ +export function boardEntryView(content: string, max = 200): string { + if (content.trim().startsWith("memory=")) return content.trim(); + const single = content.replace(/\s+/g, " ").trim(); + return single.length <= max ? single : `${single.slice(0, max - 1)}…`; +} diff --git a/src/core/store/base.ts b/src/core/store/base.ts index cb3560ed..2dd4f016 100644 --- a/src/core/store/base.ts +++ b/src/core/store/base.ts @@ -37,6 +37,7 @@ import type { VectorEmbedder, } from "../types.ts"; import { TASK_BOARD_VERDICTS, WORLD_BOARD_ID } from "../types.ts"; +import { boardEntryView } from "../board-entry-view.ts"; import { currentlyValid, notExpired } from "./clock.ts"; import { histogramAdd } from "../perf.ts"; import { Router } from "../router.ts"; @@ -591,7 +592,7 @@ export class NmgStoreBase { ackCount: entry.ackedBy.length, createdAt: entry.createdAt, resolvedAt: entry.resolvedAt, - preview: taskBoardPreview(entry.content), + preview: boardEntryView(entry.content), })); return { previews, nextCursor }; } @@ -3034,16 +3035,6 @@ export class NmgStoreBase { } } -/** Bounded preview used by the compact read (readTaskBoardPreviews): a lone - * memory pointer is already the intended low-context form and is returned whole; - * anything longer is collapsed to a single bounded line so a sync never carries - * a full body it did not ask for. */ -function taskBoardPreview(content: string, max = 200): string { - if (content.trim().startsWith("memory=")) return content.trim(); - const single = content.replace(/\s+/g, " ").trim(); - return single.length <= max ? single : `${single.slice(0, max - 1)}…`; -} - function mapTaskBoardEntry(row: Row): TaskBoardEntry { return { id: String(row.id), diff --git a/tests/core/task-board.test.ts b/tests/core/task-board.test.ts index 0dc6860a..c4b39b2e 100644 --- a/tests/core/task-board.test.ts +++ b/tests/core/task-board.test.ts @@ -5,6 +5,7 @@ import { join } from "node:path"; import test from "node:test"; import { NmgStore } from "../../src/core/store.ts"; +import { boardEntryView } from "../../src/core/board-entry-view.ts"; function withStore(run: (store: NmgStore) => void): void { const directory = mkdtempSync(join(tmpdir(), "nmg-task-board-")); @@ -103,6 +104,17 @@ test("task board supports cursor reads, cross-agent resolution, and expiry", () }); }); +test("the read face has one rule for a board entry's body", () => { + // A pointer is already the low-context form the reader wants, so it is returned whole and never cut. + assert.equal(boardEntryView("memory=abc123"), "memory=abc123"); + // Anything else becomes one bounded line, so a read never carries a body it did not ask for. + assert.equal(boardEntryView("first line\n\nsecond line"), "first line second line"); + assert.equal(boardEntryView("x".repeat(10), 10), "x".repeat(10)); + const bounded = boardEntryView("y".repeat(11), 10); + assert.equal(bounded.length, 10); + assert.equal(bounded, `${"y".repeat(9)}…`); +}); + test("compact preview read omits long bodies but keeps ordering and cursor", () => { withStore((store) => { const first = store.putTaskBoardEntry({ From 23d86446b181810d949d4b261589c892ec6e1356 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 19:43:54 +0800 Subject: [PATCH 14/32] The store gets its own board The port now has a third implementation, and it is the product's. src/integration/ooo-runner.ts projects the run's frozen table and its board entries into the facts the shared rules read, and carries no legality rule of its own: the plan is compiled by compileTaskUnits, the legal set is dispatchTasks + selectableTasks, cancellations are read through the coordinator's own reader, and the verdict belongs to the caller's acceptance. A rule that turns out to be needed here belongs in task-semantics.ts instead - needing one is how this module would show that the abstraction leaked. Two things the projection cannot read from the store stay the caller's, and both are existing divisions rather than new ones: a task's declared file contents (the store freezes paths, and preparing the workspace is the patch path's caller's job) and an acceptance whose identity is independent of the deliverer (the store already refuses a deliverer that judges its own delivery, which the second test pins). One contract is easy to get wrong, so the projection writes it down: the shared acceptance rule binds a verdict to the artifact value the run carries - it compares the verdict's digest against the artifact - and not to the store's own deliverable hash. Reporting the hash there makes every accepted unit read as unaccepted, and the symptom is silent: the next unit is never released and nothing reports an error. The port's "put a result" is the product's delivery on the entry the unit was claimed on, so a unit's work is recorded once rather than once on that entry and again on a result entry nobody would read; and the unit is resolved from the claim rather than from the body, because a body's shape belongs to its owner while a claim is the board's own fact. Verification: npm run check clean; test:product 1527/1527 pass (2 new - a two-unit plan driven through the product's store with an independent judge, and the self-judging refusal); build clean; lint clean; prettier clean; docs:check 288 files, 0 errors, 0 warnings; agent:context:check valid once the route claims the new file. --- agent-context.yaml | 1 + docs/design/mechanism-in-the-middle.md | 16 + docs/design/mechanism-in-the-middle.zh-CN.md | 3 + src/core/store/base.ts | 9 +- src/integration/ooo-runner.ts | 375 ++++++++++++++++++ tests/integration/ooo-store-run-board.test.ts | 240 +++++++++++ 6 files changed, 643 insertions(+), 1 deletion(-) create mode 100644 src/integration/ooo-runner.ts create mode 100644 tests/integration/ooo-store-run-board.test.ts diff --git a/agent-context.yaml b/agent-context.yaml index 43182b95..54213d38 100644 --- a/agent-context.yaml +++ b/agent-context.yaml @@ -217,6 +217,7 @@ routes: - src/integration/ooo-fusion-plan.ts - src/integration/ooo-mutation.ts - src/integration/ooo-patch.ts + - src/integration/ooo-runner.ts - src/integration/check-ticket.ts - src/integration/check-runner.ts - src/integration/ooo-session-facts.ts diff --git a/docs/design/mechanism-in-the-middle.md b/docs/design/mechanism-in-the-middle.md index d87af946..e0183189 100644 --- a/docs/design/mechanism-in-the-middle.md +++ b/docs/design/mechanism-in-the-middle.md @@ -77,6 +77,22 @@ against an imagined second shape, and the imagined one is the one that gets buil - **The write face's type did not land.** The three fillers are three kinds of entry rather than one shape drawn three ways, so the trigger for a type is a fourth filler. - **The middle's seam did not land.** Check (c) says why. +- **The port has a third implementation, and it is the product's.** `src/integration/ooo-runner.ts` is the + store's own board: it projects the run's frozen table and its board entries into the facts the shared + rules read, and it carries no legality rule of its own. Three implementations satisfy `DispatchBoard` now + - the probe board, the dispatch test's stub, and this one - and only the projection differs between them, + which is the abstraction working rather than a claim that it does. +- **The projection has one contract that is easy to get wrong.** The shared acceptance rule binds a verdict + to the artifact a run carries: it compares the verdict's digest against the artifact value, so a verdict + about a different artifact cannot pass as acceptance. Reporting the store's own deliverable hash in that + field instead makes every accepted unit read as unaccepted, and the symptom is silent - the next unit is + never released and nothing reports an error. Both identities are legitimate; only the artifact value + belongs in that field. +- **Where the port's verbs meet the product's lifecycle.** A product run keeps a deliverable on the entry a + unit was claimed on, so the port's "put a result" is that delivery and its "submit" is the judgement of + it: one entry per unit, not two, because a second result entry would be a record nothing reads. The unit + is resolved from the claim the loop has just taken rather than from the body, because a body's shape + belongs to its owner while a claim is the board's own fact. ## Why this is a draft and not a record diff --git a/docs/design/mechanism-in-the-middle.zh-CN.md b/docs/design/mechanism-in-the-middle.zh-CN.md index a5ce8203..a2a27aa8 100644 --- a/docs/design/mechanism-in-the-middle.zh-CN.md +++ b/docs/design/mechanism-in-the-middle.zh-CN.md @@ -50,6 +50,9 @@ - **读面有了一个家。** `src/core/board-entry-view.ts` 拥有这条规则:孤立的 `memory=` 指针原样返回;更长的内容折叠成**一行有界文本**;界**归调用方**,因为"塞得下多少"是读者的决定。store 的紧凑读与宿主适配器的广播都调它;store 里那份私有副本和适配器里那条裸切都删掉了。三条规则里有一条不只是重复,而是**错的**:`.pi/extensions/nmg/index.ts` 那条 140 字符裸切不折叠空白,于是一个多行内容会以多行广播的形态发出去。`tests/core/task-board.test.ts` 直接钉住这条规则。适配器里的 `excerpt` 助手留给记忆结果——那不是黑板条目,保留它自己的规则。 - **写面的类型没有落地。** 三种填法是三种**条目**,不是同一个形状画了三遍;所以写类型的触发条件是出现第四种填法。 - **中间的缝没有落地。** 检查(c)说明了为什么。 +- **端口有了第三个实现,而且是产品自己的那个。** `src/integration/ooo-runner.ts` 是 store 自己的黑板:它把运行的冻结表与黑板条目**投影**成共享规则要读的那些事实,自己不带任何合法性规则。现在有三个实现满足 `DispatchBoard`——探针黑板、dispatch 测试里的桩、和这一个——它们之间**只有投影不同**,这是抽象在起作用,而不是关于它的说法。 +- **投影里有一条容易搞错的契约。** 共享的验收规则把判词**绑在运行携带的产物上**:它拿判词的摘要与产物**值**比较,所以关于另一个产物的判词不能冒充验收。把 store 自己的交付物哈希填进那个字段,会让每个已验收单元都读成未验收,而症状是**无声的**——下一个单元永不释放,且任何地方都不报错。两种身份都是正当的,但那个字段只属于产物值。 +- **端口的动词落在产品生命周期的哪里。** 产品的一次运行把交付物存在**单元被认领时的那条条目**上,所以端口的"投一个结果"就是那次交付、"提交"就是对它的判定:一个单元一条条目,不是两条,因为第二条结果条目会是一条没人读的记录。单元是从循环刚拿到的**认领**解出来的,而不是从正文解出来的——正文的形状归它的主人,而认领是黑板自己的事实。 ## 为什么它现在还只是草稿 diff --git a/src/core/store/base.ts b/src/core/store/base.ts index 2dd4f016..907720da 100644 --- a/src/core/store/base.ts +++ b/src/core/store/base.ts @@ -775,7 +775,7 @@ export class NmgStoreBase { } /** True when a board entry carries a live claim (holder set, lease not expired). */ private taskBoardClaimLive(entry: TaskBoardEntry, now: string): boolean { - return entry.claimedBy !== null && entry.claimExpiresAt !== null && entry.claimExpiresAt > now; + return taskBoardClaimIsLive(entry, now); } /** * Reply-gated serial handoff: promote the earliest pending actionable to @@ -3035,6 +3035,13 @@ export class NmgStoreBase { } } +/** True when a board entry carries a live claim (holder set, lease not expired). One home for the + * predicate: eligibility outside this module reads it too, so a second copy of it would be a second + * opinion about whether a task is already being worked on. */ +export function taskBoardClaimIsLive(entry: TaskBoardEntry, now: string): boolean { + return entry.claimedBy !== null && entry.claimExpiresAt !== null && entry.claimExpiresAt > now; +} + function mapTaskBoardEntry(row: Row): TaskBoardEntry { return { id: String(row.id), diff --git a/src/integration/ooo-runner.ts b/src/integration/ooo-runner.ts new file mode 100644 index 00000000..7740c3d6 --- /dev/null +++ b/src/integration/ooo-runner.ts @@ -0,0 +1,375 @@ +/** + * The store's own board: one run driven through the product's record. + * + * This is the third implementation of `DispatchBoard` - the probe board and the dispatch test's stub + * are the other two - and it is deliberately the thinnest of the three. No legality rule lives here: + * the plan is compiled by `compileTaskUnits`, the legal set is `dispatchTasks` and `selectableTasks`, + * cancellations are read through the coordinator's own reader, and the verdict belongs to the + * caller's acceptance. What this module adds is a **projection** from the product's run tables and + * board entries into the facts those pure functions read, plus the six board operations. A rule that + * turns out to be needed here belongs in `task-semantics.ts` instead: needing one is how this module + * would show that the abstraction leaked. + * + * Two things the projection cannot read from the store are the caller's, and both are existing + * divisions rather than new ones: a task's declared file contents (the store freezes paths, and + * preparing the workspace belongs to the patch path's caller) and an acceptance whose identity is + * independent of the deliverer (the store refuses a deliverer that judges its own delivery). + */ +import { createHash } from "node:crypto"; + +import type { NmgStore } from "../core/store.ts"; +import { taskBoardClaimIsLive } from "../core/store/base.ts"; +import { TASK_BOARD_VERDICTS } from "../core/types.ts"; +import type { DispatchBoard, DispatchEntry, DispatchTicket } from "./ooo-dispatch.ts"; +import type { PatchTaskSpec, ProbeOperation, ProbePlan } from "./ooo-board.ts"; +import { + patchSubmission, + preparePatchWork, + type PatchSubmission, + type PatchWork, +} from "./ooo-patch.ts"; +import { selectableTasks } from "./ooo-execution.ts"; +import { compileTaskUnits, dispatchTasks, type RecordedFacts } from "./task-semantics.ts"; +import { coordinatedBoardWrite, taskCancellation, taskRunStatus } from "./task-coordinator.ts"; + +/** The words a board verdict may use, as the store defines them. */ +export type RunVerdict = (typeof TASK_BOARD_VERDICTS)[number]; + +/** The words an acceptance may answer with, as a patch spec defines them. */ +export type AcceptanceAnswer = "accept" | "reject" | "undecidable"; + +export interface RunBoardOptions { + runId: string; + /** The channel this run's entries live on. The caller declares it, because an entry is looked up by + * the channel it was put on rather than by the run owning exactly one. */ + channel: string; + /** How many units of the legal set may be in flight at once. Declared by the run, read as given. */ + slots: number; + /** The declared files' contents: the store froze the paths a task declares, and the bytes are the + * caller's, which is the division the patch work already assumes. */ + workspace: (taskId: string, files: readonly string[]) => Readonly>; + /** The acceptance, and a name that need not be the deliverer's: the store refuses a deliverer that + * judges its own delivery, so a run whose judge is its worker is refused by the board and not by + * this module. One acceptance is read twice - the compiler freezes it into each spec and the board + * asks it again when a delivery arrives - so a unit's declared acceptance and its verdict cannot + * come to mean two things. */ + acceptance: { + agentId: string; + verify: (input: { + taskId: string; + submission: PatchSubmission; + }) => Promise | AcceptanceAnswer; + }; + /** The clock the loop stamps entries with. Omitted, the process clock. */ + now?: () => number; +} + +/** One task of the frozen plan, as the store has it. */ +type RunTask = ReturnType[number]; + +/** One binding: the entry a task's work is carried by. */ +type RunBinding = ReturnType["bindings"][number]; + +/** The store's board for one run, satisfying the loop's port and nothing else. */ +export class StoreRunBoard implements DispatchBoard { + readonly channel: string; + readonly #store: NmgStore; + readonly #options: RunBoardOptions; + + constructor(store: NmgStore, options: RunBoardOptions) { + this.#store = store; + this.#options = options; + this.channel = options.channel; + } + + get now(): number { + return this.#options.now?.() ?? Date.now(); + } + + /** What the loop may claim right now, in the shared rule's own order. */ + candidates(): readonly string[] { + const { compiled, facts } = this.#projection(); + return selectableTasks(dispatchTasks(compiled.units, facts), this.#options.slots); + } + + /** The accepted artifact per task id: the same map the rules read, so the two cannot disagree. */ + accepted(): Readonly> { + return this.#projection().facts.artifacts ?? {}; + } + + /** Take one unit. The claim is the store's compare-and-set on the entry its binding names. */ + claim(taskId: string, owner: string): DispatchTicket { + const binding = this.#bindingFor(taskId); + const entry = coordinatedBoardWrite(this.#store, { + runId: this.#options.runId, + entryId: binding.entryId, + verb: "claim", + actorId: owner, + apply: () => + this.#store.claimTaskBoardEntry({ + taskId: this.channel, + entryId: binding.entryId, + agentId: owner, + }), + }).entry; + const attempt = entry.attempt ?? binding.attempt; + const task = this.#task(taskId); + const accepted = this.accepted(); + return { + attempt, + dependencies: Object.fromEntries( + task.dependencies + .filter((dependency) => dependency in accepted) + .map((dependency) => [dependency, accepted[dependency]!]), + ), + patch: this.#work(task, attempt) ?? undefined, + }; + } + + /** + * The port's "put a result" is the product's delivery: the artifact lands on the entry the unit was + * claimed on, which is where this board keeps deliverables and where a judge looks for one. The unit + * is resolved from the claim - the identity the loop claimed with - and not from the content, because + * how a body is shaped is the shape owner's business while a claim is the board's own fact. + * + * The returned identity is therefore the claimed entry rather than a new one: this board records a + * unit's work once, on the entry that carries the unit, instead of once on that entry and again on a + * result entry nothing would read. + */ + putTaskBoardEntry(input: { + taskId: string; + agentId: string; + kind: "result"; + content: string; + expiresAt: string; + }): DispatchEntry { + const binding = this.#claimedBy(input.agentId); + const artifact = artifactOf(input.content); + coordinatedBoardWrite(this.#store, { + runId: this.#options.runId, + entryId: binding.entryId, + verb: "deliver", + actorId: input.agentId, + apply: () => + this.#store.deliverTaskBoardEntry({ + taskId: this.channel, + entryId: binding.entryId, + agentId: input.agentId, + digest: createHash("sha256").update(artifact).digest("hex"), + ref: artifact, + }), + }); + return { id: binding.entryId }; + } + + /** + * Let the acceptance decide the verdict of what was delivered, under the judge's own name. The + * frozen envelope is rebuilt here for the same reason the loop rebuilds it: the digest is the + * artifact's identity, and recomputing it is how a caller holding only an entry id gets it back. + */ + async submit(entryId: string): Promise { + const binding = this.#store.taskRunForEntry(entryId); + if (!binding) throw new Error(`entry ${entryId} is not adopted by a run, so it has no verdict`); + const entry = this.#store.getTaskBoardEntryById(this.channel, entryId); + if (!entry) throw new Error(`no board entry ${entryId} on channel ${this.channel}`); + if (entry.deliverableRef === null && entry.deliverableDigest === null) + throw new Error(`entry ${entryId} holds no deliverable, so there is nothing to judge`); + const task = this.#task(binding.taskId); + const artifact = entry.deliverableRef ?? entry.deliverableDigest!; + const verdict = verdictOf( + await this.#options.acceptance.verify({ + taskId: task.taskId, + submission: patchSubmission( + frozenOf(this.#declarationOnly(task, entry.attempt ?? binding.attempt)), + artifact, + ), + }), + ); + coordinatedBoardWrite(this.#store, { + runId: this.#options.runId, + entryId, + verb: "judge", + actorId: this.#options.acceptance.agentId, + apply: () => + this.#store.judgeTaskBoardEntry({ + taskId: this.channel, + entryId, + agentId: this.#options.acceptance.agentId, + verdict, + }), + }); + return verdict; + } + + /** The frozen plan compiled by the shared compiler, and the facts the run's entries record. */ + #projection(): { compiled: ReturnType; facts: RecordedFacts } { + const tasks = this.#tasks(); + const specs: Record = {}; + for (const task of tasks) { + const declared = task.patchFiles ?? []; + if (declared.length === 0 && task.patchEditable === null) continue; + specs[task.taskId] = { + instruction: task.input, + files: this.#options.workspace(task.taskId, declared), + editable: [...(task.patchEditable ?? [])], + visible: [...declared], + verify: async (submission) => + this.#options.acceptance.verify({ taskId: task.taskId, submission }), + }; + } + const compiled = compileTaskUnits({ plan: this.#plan(tasks), specs }); + if (!compiled.legal) { + // A plan the shared compiler refuses is refused here by name, with the refusals that say why. + const first = compiled.refusals[0]; + throw new Error( + `plan refused by the shared semantics: ${compiled.refusals.length} refusal(s); first is ` + + `${String(first?.task)}/${String(first?.field)}: ${String(first?.reason)}`, + ); + } + return { compiled, facts: this.#recordedFacts() }; + } + + /** The frozen plan in the row shape the shared compiler reads. */ + #plan(tasks: readonly RunTask[]): ProbePlan { + return tasks.map((task) => [ + task.taskId, + task.input, + [...task.dependencies], + task.effect, + task.waitEvent, + (task.operation === "" ? null : task.operation) as ProbeOperation | null, + ]); + } + + /** + * The run's record as the pure rules read it. Every fact comes from where it is authoritative - the + * entries for claims and verdicts, the run's log for cancellations - and nothing is derived here. + */ + #recordedFacts(): RecordedFacts { + const artifacts: Record = {}; + const verdicts: Record = {}; + const claimed: string[] = []; + const now = new Date(this.now).toISOString(); + for (const binding of this.#bindings()) { + const entry = binding.entryId + ? this.#store.getTaskBoardEntryById(this.channel, binding.entryId) + : null; + if (!entry) continue; + if (entry.verdict !== null && entry.deliverableDigest !== null) { + // Two identities meet here and only one of them is the rule's. The store keeps a hash of the + // deliverable in its own column, while the shared acceptance rule binds a verdict to the + // artifact value the run carries: it compares the verdict's digest against the artifact, so a + // verdict about another artifact cannot pass as acceptance. This projection therefore reports + // the artifact as the digest, and the store's hash stays where it is authoritative - the + // judge's own guard, which refuses a verdict once the deliverable has changed. + const artifact = entry.deliverableRef ?? artifactOf(entry.content); + verdicts[binding.taskId] = { digest: artifact, verdict: entry.verdict as RunVerdict }; + if (entry.verdict === "accepted") artifacts[binding.taskId] = artifact; + } + if (taskBoardClaimIsLive(entry, now)) claimed.push(binding.taskId); + } + const cancellations = this.#tasks() + .filter((task) => taskCancellation(this.#store, this.#options.runId, task.taskId) !== null) + .map((task) => task.taskId); + return { artifacts, verdicts, claimed, cancellations }; + } + + /** The entry each task of the plan is bound to, from the run's own record of its bindings. */ + #bindings(): readonly RunBinding[] { + return taskRunStatus(this.#store, this.#options.runId).bindings; + } + + /** The entry a claim holder's work belongs to: the one live claim that is theirs, by name. */ + #claimedBy(owner: string): { taskId: string; attempt: number; entryId: string } { + const now = new Date(this.now).toISOString(); + const held = this.#bindings().filter((binding) => { + const entry = binding.entryId + ? this.#store.getTaskBoardEntryById(this.channel, binding.entryId) + : null; + return entry !== null && entry.claimedBy === owner && taskBoardClaimIsLive(entry, now); + }); + const only = held.length === 1 ? held[0] : undefined; + if (!only?.entryId) + throw new Error( + held.length === 0 + ? `${owner} holds no live claim on channel ${this.channel}, so a result has no unit to belong to` + : `${owner} holds ${held.length} live claims on channel ${this.channel}, so a result cannot be placed`, + ); + return { taskId: only.taskId, attempt: only.attempt, entryId: only.entryId }; + } + + #bindingFor(taskId: string): { taskId: string; attempt: number; entryId: string } { + const binding = [...this.#bindings()] + .reverse() + .find((candidate) => candidate.taskId === taskId); + if (!binding?.entryId) + throw new Error( + `no entry is adopted for task ${taskId} of run ${this.#options.runId} on channel ${this.channel}`, + ); + return { taskId: binding.taskId, attempt: binding.attempt, entryId: binding.entryId }; + } + + #tasks(): readonly RunTask[] { + return [...this.#store.taskRunTasks(this.#options.runId)].sort( + (left, right) => left.position - right.position, + ); + } + + #task(taskId: string): RunTask { + const task = this.#tasks().find((candidate) => candidate.taskId === taskId); + if (!task) throw new Error(`run ${this.#options.runId} freezes no task ${taskId}`); + return task; + } + + /** + * The work a claim admits: the declaration the store froze, with the caller's bytes in it. Null when + * the task declares no patch files, which is the store's own reading of a non-patch unit. + */ + #work(task: RunTask, attempt: number): PatchWork | null { + const declared = task.patchFiles ?? []; + if (declared.length === 0 && task.patchEditable === null) return null; + return this.#declarationOnly(task, attempt); + } + + #declarationOnly(task: RunTask, attempt: number): PatchWork { + const declared = task.patchFiles ?? []; + return { + taskId: task.taskId, + attempt, + instruction: task.input, + files: this.#options.workspace(task.taskId, declared), + editable: [...(task.patchEditable ?? [])], + visible: [...declared], + }; + } +} + +/** The frozen envelope of a declaration, which is what the shape's own parser needs. */ +function frozenOf(work: PatchWork) { + return preparePatchWork({ + taskId: work.taskId, + attempt: work.attempt, + instruction: work.instruction, + files: work.files, + editable: work.editable, + visible: work.visible, + }); +} + +/** The artifact a worker's entry carries: the loop's own put shape, read back without interpretation. */ +function artifactOf(content: string): string { + try { + const parsed = JSON.parse(content) as { artifact?: unknown }; + if (typeof parsed.artifact === "string") return parsed.artifact; + } catch { + // Not the loop's shape: the content is the artifact itself. + } + return content; +} + +/** One acceptance's answer as the board's word for it. */ +function verdictOf(answer: AcceptanceAnswer): RunVerdict { + if (answer === "accept") return "accepted"; + if (answer === "reject") return "rejected"; + return "undecidable"; +} diff --git a/tests/integration/ooo-store-run-board.test.ts b/tests/integration/ooo-store-run-board.test.ts new file mode 100644 index 00000000..48e3ad9c --- /dev/null +++ b/tests/integration/ooo-store-run-board.test.ts @@ -0,0 +1,240 @@ +/** + * The store's board is the third implementation of the loop's port, and the first that reads the + * product's own record instead of a probe's. + * + * What is asserted here is the projection - what the run's frozen table and its board entries add up + * to once the shared rules read them - and the two ends the store owns: a claim is the store's + * compare-and-set, and a delivery is judged under a name other than the deliverer's. No legality rule + * is asserted because none lives in the board; the plan order and the dependency gate come from + * `task-semantics.ts`, which the probe board reads too. + */ +import assert from "node:assert/strict"; +import { mkdtempSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import test from "node:test"; + +import { NmgStore } from "../../src/core/store.ts"; +import { dispatchPlan, type PlanWorker } from "../../src/integration/ooo-dispatch.ts"; +import type { SessionPlan } from "../../src/integration/ooo-execution.ts"; +import { StoreRunBoard, type AcceptanceAnswer } from "../../src/integration/ooo-runner.ts"; +import type { PatchSubmission } from "../../src/integration/ooo-patch.ts"; +import { + coordinatedBoardWrite, + createBoundEntry, + freezeRunPlan, + registerRun, +} from "../../src/integration/task-coordinator.ts"; + +const RUN = "run-1"; +const CHANNEL = "run-1-board"; +const FILE = "src/work.ts"; +const baseline = { [FILE]: "export const value = 1;\n" }; + +/** Two units of one plan, the second depending on the first, both patch work. */ +function tasks() { + return [ + { + taskId: "P", + revision: "r1", + input: "add the first line", + dependencies: [], + effect: "isolated-artifact", + kind: "patch", + patchFiles: [FILE], + patchEditable: [FILE], + }, + { + taskId: "T", + revision: "r1", + input: "add the second line", + dependencies: ["P"], + effect: "isolated-artifact", + kind: "patch", + patchFiles: [FILE], + patchEditable: [FILE], + }, + ]; +} + +/** A run registered, frozen and adopted: what the runner owns before any unit runs. */ +function openRun(store: NmgStore, taskIds: readonly string[] = ["P", "T"]): void { + registerRun(store, { + runId: RUN, + planDigest: "d1", + policy: "repair-first", + revision: "r1", + retention: "run", + }); + freezeRunPlan(store, { + runId: RUN, + tasks: tasks().filter((task) => taskIds.includes(task.taskId)), + }); + for (const taskId of taskIds) { + createBoundEntry(store, { + entry: { + taskId: CHANNEL, + agentId: "runner", + kind: "handoff", + content: `work on ${taskId}`, + expiresAt: new Date(Date.now() + 86_400_000).toISOString(), + }, + runId: RUN, + taskId, + attempt: 0, + }); + } +} + +/** The caller's bytes for a task's declared files: the store froze paths, not contents. */ +function workspace(_taskId: string, files: readonly string[]): Readonly> { + return Object.fromEntries(files.map((path) => [path, baseline[path] ?? ""])); +} + +/** The artifact a worker produces: the wire shape the host validates, with this task's own marker. */ +const worker: PlanWorker = async (taskId, frozen) => + JSON.stringify({ + digest: frozen.digest, + files: [{ path: FILE, content: `${baseline[FILE]}// ${taskId}\n` }], + }); + +/** The acceptance reads the submission and answers; its name is not the deliverer's. */ +const acceptance = { + agentId: "judge", + verify: ({ + taskId, + submission, + }: { + taskId: string; + submission: PatchSubmission; + }): AcceptanceAnswer => { + if (submission.kind !== "patch") return "undecidable"; + return (submission.files[FILE] ?? "").includes(`// ${taskId}`) ? "accept" : "reject"; + }, +}; + +/** The session view the loop asks when a chain continues. This run declares no sessions. */ +function legality(): SessionPlan { + return { tasks: [], declarations: {} }; +} + +async function withStore(run: (store: NmgStore) => Promise): Promise { + const directory = mkdtempSync(join(tmpdir(), "nmg-run-board-")); + const store = new NmgStore(join(directory, "nmg.sqlite")); + try { + await run(store); + } finally { + store.close(); + rmSync(directory, { recursive: true, force: true }); + } +} + +test("one plan runs through the product's own board, judged by someone other than the deliverer", async () => { + await withStore(async (store) => { + openRun(store); + const board = new StoreRunBoard(store, { + runId: RUN, + channel: CHANNEL, + slots: 1, + workspace, + acceptance, + }); + + // Before anything runs, exactly the plan's first unit is legal: the dependency gate is the + // shared rule's, read through this board's projection. + assert.deepEqual(board.candidates(), ["P"], "a dependent unit is not offered early"); + + const outcome = await dispatchPlan({ + board, + plan: ["P", "T"], + slots: 1, + legality, + worker, + ownerOf: (taskId) => `worker:${taskId}`, + }); + + assert.deepEqual(outcome.failures, [], "no unit failed"); + assert.deepEqual(outcome.order, ["P", "T"], "the plan order held"); + assert.deepEqual( + outcome.units.map((unit) => unit.verdict), + ["accepted", "accepted"], + "the acceptance accepted both units", + ); + + // The run's own record says the same: both units are accepted, and the judge is not the worker. + const accepted = board.accepted(); + assert.deepEqual(Object.keys(accepted).sort(), ["P", "T"]); + assert.match(accepted.P!, /\/\/ P/u); + for (const taskId of ["P", "T"]) { + const entry = store.taskBoardEntry(bindingEntry(store, taskId)); + assert.ok(entry, `task ${taskId} has an entry`); + assert.equal(entry.verdict, "accepted"); + assert.equal(entry.judgedBy, acceptance.agentId); + assert.equal(entry.deliveredBy, `worker:${taskId}`); + assert.ok(entry.deliverableDigest, "the deliverable the judge bound to is recorded"); + } + }); +}); + +test("the deliverer cannot judge its own delivery, so a run cannot self-accept", async () => { + await withStore(async (store) => { + openRun(store, ["P"]); + const board = new StoreRunBoard(store, { + runId: RUN, + channel: CHANNEL, + slots: 1, + workspace, + acceptance, + }); + const entryId = bindingEntry(store, "P"); + const ticket = board.claim("P", "worker:P"); + const artifact = JSON.stringify({ + digest: ticket.patch!.digest, + files: [{ path: FILE, content: `${baseline[FILE]}// P\n` }], + }); + coordinatedBoardWrite(store, { + runId: RUN, + entryId, + verb: "deliver", + actorId: "worker:P", + apply: () => + store.deliverTaskBoardEntry({ + taskId: CHANNEL, + entryId, + agentId: "worker:P", + digest: "d", + ref: artifact, + }), + }); + // The lifecycle write of a managed entry goes through the run's coordinated transition, so the + // refusal below is the board's own check reached the way a caller reaches it. + assert.throws( + () => + coordinatedBoardWrite(store, { + runId: RUN, + entryId, + verb: "judge", + actorId: "worker:P", + apply: () => + store.judgeTaskBoardEntry({ + taskId: CHANNEL, + entryId, + agentId: "worker:P", + verdict: "accepted", + }), + }), + /deliverer cannot judge its own deliverable/u, + "the board refuses a self-judged delivery, which is why the acceptance has its own name", + ); + }); +}); + +/** The entry a task's binding holds, read from the run's own record. */ +function bindingEntry(store: NmgStore, taskId: string): string { + const fact = store + .taskRunFacts(RUN) + .filter((candidate) => candidate.taskId === taskId && candidate.entryId !== null) + .at(-1); + assert.ok(fact?.entryId, `task ${taskId} has a bound entry`); + return fact.entryId; +} From dcc611a0d2af1af873a1400e661f930e92224479 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 19:59:38 +0800 Subject: [PATCH 15/32] The current-value case was a stopwatch CI failed "Product tests and coverage" on one case of 1527 while the same suite passed 30 times in a row locally: "a value that expired a moment ago is still current" stamped an expiry half a grace (25 ms) in the past in one statement and read it in another, so it asserted that the read happens within the half that is left. A sweep between the stamp and the read says where it flips - 0 ms and 10 ms answered "current", 20 ms and beyond answered "not active" - and a loaded runner has no trouble spending 20 ms reopening a store. Sharing a clock source is not sharing a clock read. The case is now a relation: the store's own notExpired predicate is evaluated against a boundary stamped in the same statement, so both share one 'now' and the expectation is about the window rather than about the machine. It also pins the boundary instead of a point inside it - at exactly one grace an expiry is already excluded. 8000 evaluated assertions under deliberate load, no flip. A row-level form of that case cannot exist, because a row's boundary is stamped by one statement and compared by another, so the store level keeps the cases that do not depend on a duration: a future stamp inside the grace, a minute either way, and 400 write-then-read rounds. The mutant that names this case as its catcher is pointed at the new name and still fails it: 4 of 4 caught, restored byte-identically. Postmortem 0004's guardrail claim and its lessons are corrected with the recurrence: its "deterministic" test was deterministic on three of its four boundary cases, and the fix for a timing dependency was itself one. --- .../0004-flaky-was-a-clock-boundary.md | 39 ++++++++++----- .../0004-flaky-was-a-clock-boundary.zh-CN.md | 20 ++++++-- tests/core/store/current-value-window.test.ts | 47 +++++++++++++++---- tools/mutation-teeth.ts | 2 +- 4 files changed, 81 insertions(+), 27 deletions(-) diff --git a/docs/postmortem/0004-flaky-was-a-clock-boundary.md b/docs/postmortem/0004-flaky-was-a-clock-boundary.md index ba890773..dbde9317 100644 --- a/docs/postmortem/0004-flaky-was-a-clock-boundary.md +++ b/docs/postmortem/0004-flaky-was-a-clock-boundary.md @@ -7,7 +7,7 @@ ## Executive summary `npm run test:product` failed once, under load, in `demoteMemory: demotes LTG memory to STG`. A previous -session recorded it in the ledger as *flaky, not fixed* and moved on; the entry name was the whole +session recorded it in the ledger as _flaky, not fixed_ and moved on; the entry name was the whole diagnosis, and it closed the question for a session. Run under a loop instead of once, it reproduced in about 1 write per 1500: a memory was written with `valid_from = …38.468Z` and read back while SQLite's `now` said `…38.467Z`, so the read path's `valid_from <= now` was false, the row was invisible, and the @@ -19,13 +19,13 @@ disagree at millisecond granularity. The current-value predicate compares a stored timestamp against SQLite's own clock, four times in `src/core/store/base.ts` and `src/core/store/retrieval.ts`. Writes stamp from JavaScript, so the two -sources race: whenever the row's stamp lands within ~1-2 ms *after* the reading connection's `now`, a +sources race: whenever the row's stamp lands within ~1-2 ms _after_ the reading connection's `now`, a predicate that means "this value is in force" answers "no" for a row that was written microseconds ago. Under a loaded machine the skew widens; the loop that found it needed 3000 iterations to see 2-4 hits. The fix keeps the comparison but widens the window by a named grace, in one new home, `src/core/store/clock.ts`: `CLOCK_GRACE_MS = 50`, with `clockNow("later" | "earlier")` feeding the four -predicates. The grace only ever *widens* what counts as current - it never narrows - so the change cannot +predicates. The grace only ever _widens_ what counts as current - it never narrows - so the change cannot hide an expiry, and the two directions are separate so the asymmetry is visible at each site. Measured after the change: 3000 iterations, 0 failures (was 2-4). @@ -36,7 +36,7 @@ fractional seconds (`'+0.050 seconds'`) and why a mutant now pins that choice. ## Impact A memory written within about a millisecond of being read could be invisible to one read: the write -succeeded, the row was stored, and the *next* read saw it. No durable loss and no corruption - the failure +succeeded, the row was stored, and the _next_ read saw it. No durable loss and no corruption - the failure mode is a single false "not active" answer, which surfaces as an error from `requireActiveMemory` or as a missing row in maintenance, demotion, dedup and search paths. It reached a product test only because that test reads a just-written memory in a tight loop; under load a user-visible `not active` error for a @@ -56,10 +56,15 @@ product - and a labelled failure is a failure nobody has to look at again. iterations, both `memory … is not active` for a row that exists. - Dumping the row and SQLite's `now` side by side shows the 1 ms order: stamp `…38.468Z`, `now` `…38.467Z`. - Fix: the window's grace, in `src/core/store/clock.ts`, wired into the four predicates. Loop: 0 of 3000. -- A deterministic test (`tests/core/store/current-value-window.test.ts`, 6 cases) pins both boundaries, and - 4 named mutants (2 on the boundaries, 1 on the grace, 1 on the SQLite time unit) make the pin checkable: - 4 of 4 caught. +- A test (`tests/core/store/current-value-window.test.ts`, 6 cases) pins both boundaries, and 4 named + mutants (2 on the boundaries, 1 on the grace, 1 on the SQLite time unit) make the pin checkable: 4 of 4 + caught. Three of its four boundary cases were deterministic from the start; the freshness one was not, + and the correction is below. - The ledger row is corrected from "flaky, not fixed" to the fixed defect with its reproduction rate. +- 2026-09: that same case fails on a loaded CI runner (`Product tests and coverage`, one failure of 1527) + while the suite passes 30 times in a row locally. A sweep between the stamp and the read locates the + flip - 0 ms and 10 ms after the boundary answered "current", 20 ms and beyond answered "not active" - + and the case becomes a relation. The mutant that names it as its catcher still fails it, 4 of 4. ## Root cause @@ -72,7 +77,7 @@ millisecond in the wrong direction makes a fresh row look like a future one. Not which clock had authority, because both were "the clock". The process one: an intermittent failure that passes on re-run is easy to file as flakiness, and the -ledger made that filing *look* like a result. "Flaky, not fixed" has no reproduction attempt attached, no +ledger made that filing _look_ like a result. "Flaky, not fixed" has no reproduction attempt attached, no rate, and no hypothesis - it is a label wearing a diagnosis's clothes. The failure reappeared later in the session, which is the evidence that the label was wrong; without the user's push it would have been labelled a second time. @@ -81,9 +86,14 @@ labelled a second time. - `src/core/store/clock.ts` is the single home for the current-value window's grace, so a future predicate has one place to read it from instead of hand-writing another comparison. -- `tests/core/store/current-value-window.test.ts` (6 cases) makes the boundary deterministic: a value - stamped half a grace in the future is current; a minute in the future is not; the same on the expiry - side; 400 write-then-read rounds never fail; and the window widens on both boundaries and never narrows. +- `tests/core/store/current-value-window.test.ts` (6 cases) pins the boundary: a value stamped half a grace + in the future is current; a minute in the future is not; 400 write-then-read rounds never fail; the window + widens on both boundaries and never narrows; and the freshness half widens by **one** grace, evaluated + against a boundary stamped in the same statement. That last one is the corrected shape of "a value that + expired a moment ago is still current", which stamped a boundary in one statement and read it in another - + a stopwatch whose margin was the half of the grace that was left. A row-level form of it cannot exist: a + row's boundary is stamped by one statement and compared by another, so what is testable is the relation, + not the duration. - 4 named mutants in `tools/mutation-teeth.ts` (own target for `src/core/store/clock.ts`) make the test's teeth checkable, including the SQLite time unit that silently excluded every row. - The ledger's `test:product` row no longer says "flaky": an intermittent failure is recorded with its @@ -96,8 +106,13 @@ labelled a second time. is a claim about a test; it is not a diagnosis, and it must not be recorded as one. - When two readers of the same quantity disagree, the code has to say which one is authoritative - or, as here, deliberately widen the comparison so that neither is asked for an impossible precision. -- A wrong answer from a *silently* NULL expression is worse than an exception: `strftime` with an unknown +- A wrong answer from a _silently_ NULL expression is worse than an exception: `strftime` with an unknown modifier excluded every row and reported `ok: null`. A mutant is what makes that reachable-in-theory mistake permanent. +- Sharing a clock _source_ is not sharing a clock _read_. Stamping a boundary in SQL removes the + JavaScript-versus-SQLite skew and leaves the other race untouched: two statements are two instants, so a + case that expects a value to be inside a 50 ms grace has 50 ms worth of stopwatch in it. This is the + first lesson one level down - the fix for a timing dependency was itself a timing dependency, and only a + relation is free of it. - Labelling is cheap, and that is exactly why it is dangerous: the cost of a wrong label shows up in a later session, under someone else's deadline. diff --git a/docs/postmortem/0004-flaky-was-a-clock-boundary.zh-CN.md b/docs/postmortem/0004-flaky-was-a-clock-boundary.zh-CN.md index ec4f30b8..1870ada8 100644 --- a/docs/postmortem/0004-flaky-was-a-clock-boundary.zh-CN.md +++ b/docs/postmortem/0004-flaky-was-a-clock-boundary.zh-CN.md @@ -7,7 +7,7 @@ ## Executive summary `npm run test:product` 在负载下失败过一次,用例是 `demoteMemory: demotes LTG memory to STG`。上一个会话 -在 ledger 里把它记成 *flaky, not fixed* 就走开了:那个标签本身就是全部诊断,并且让这个问题沉寂了一整个 +在 ledger 里把它记成 _flaky, not fixed_ 就走开了:那个标签本身就是全部诊断,并且让这个问题沉寂了一整个 会话。改用循环而不是跑一次之后,它在约 1/1500 次写入中复现:一条记忆以 `valid_from = …38.468Z` 写入, 读回时 SQLite 的 `now` 是 `…38.467Z`,于是读路径的 `valid_from <= now` 为假,行不可见,调用方为一条**刚写 入**的记忆收到 `memory is not active`。这是真实缺陷,位于产品的 current-value 窗口:两个时钟读者 @@ -48,9 +48,13 @@ SQLite 没有 `milliseconds` 日期修饰符:`'+50 milliseconds'` 会让 `strf `memory … is not active`。 - 把行与 SQLite 的 `now` 并列打印,看到那 1 ms 的先后:时间戳 `…38.468Z`,`now` `…38.467Z`。 - 修复:窗口宽限,落在 `src/core/store/clock.ts`,接入四个谓词。循环:3000 次 0 失败。 -- 一个确定性测试(`tests/core/store/current-value-window.test.ts`,6 个用例)钉住两个边界;4 个具名 mutant - (2 个钉边界、1 个钉宽限、1 个钉 SQLite 时间单位)让这个钉子可被检查:4/4 被抓住。 +- 一个测试(`tests/core/store/current-value-window.test.ts`,6 个用例)钉住两个边界;4 个具名 mutant + (2 个钉边界、1 个钉宽限、1 个钉 SQLite 时间单位)让这个钉子可被检查:4/4 被抓住。它的四条边界用例里 + 有三条一开始就是确定性的,过期侧那条不是——更正见下。 - ledger 中那一行从 "flaky, not fixed" 更正为已修复的缺陷及其复现率。 +- 2026-09:同一条用例在负载中的 CI 上失败(`Product tests and coverage`,1527 中 1 失败),而本地连跑 + 30 次全过。在写时间戳与读之间做扫描,定位到翻转点——边界后 0 ms 与 10 ms 答 "current",20 ms 及以后答 + "not active"——用例改成关系式。点名抓它的那个 mutant 仍然能抓住它,4/4。 ## Root cause @@ -67,8 +71,11 @@ SQLite 在读时提供 `now`。同一个墙上时钟的两个读者不会返回 ## Guardrails added - `src/core/store/clock.ts` 是 current-value 窗口宽限的唯一归属,未来的谓词有一处可读,而不必再手写一个比较。 -- `tests/core/store/current-value-window.test.ts`(6 个用例)让边界确定化:时间戳在未来半个宽限内的算生效; - 未来一分钟的算不生效;过期侧同理;400 轮"写后立刻读"永不失败;窗口在两个边界都放宽、从不收紧。 +- `tests/core/store/current-value-window.test.ts`(6 个用例)钉住边界:时间戳在未来半个宽限内的算生效; + 未来一分钟的算不生效;400 轮"写后立刻读"永不失败;窗口在两个边界都放宽、从不收紧;过期侧**恰好**放宽一个 + 宽限,且与同一条语句里写下的边界相比。最后这一条是"刚刚过期一刻的值仍算生效"更正后的形状——旧写法 + 在一条语句里写边界、在另一条语句里读它,于是它是一条秒表,余量就是剩下那半个宽限。这条用例不可能有行级 + 写法:行的边界由一条语句写下、由另一条语句比较,可测的是那条**关系**,不是那个时长。 - `tools/mutation-teeth.ts` 中 4 个具名 mutant(`src/core/store/clock.ts` 拥有独立 target)让这个测试的"牙齿" 可被检查,其中包括那个会静默排除每一行的 SQLite 时间单位。 - ledger 的 `test:product` 行不再写 "flaky":间歇失败要么连同复现尝试与比率一起记录,要么记为 open。 @@ -82,4 +89,7 @@ SQLite 在读时提供 `now`。同一个墙上时钟的两个读者不会返回 要求不可能达到的精度。 - 来自**静默** NULL 表达式的错误答案比异常更糟:`strftime` 遇到未知修饰符会排除每一行并报 `ok: null`。 一个 mutant 才能让这种"理论上可能"的失误变成永久可见。 +- 共享一个时钟**来源**不等于共享一次时钟**读取**。把边界写进 SQL 消掉了 JavaScript 与 SQLite 的偏差,却 + 没有碰到另一个竞态:两条语句就是两个瞬间,所以一条期望"值落在 50 ms 宽限内"的用例里装着 50 ms 的秒表。 + 这是第一条教训往下一层的同一个形状——为一个时序依赖做的修复本身又是个时序依赖,只有关系式不含秒表。 - 贴标签很便宜,这正是它危险的原因:错误标签的代价会在之后的会话里、在别人的截止日期下出现。 diff --git a/tests/core/store/current-value-window.test.ts b/tests/core/store/current-value-window.test.ts index d2cf3bde..9964ad4e 100644 --- a/tests/core/store/current-value-window.test.ts +++ b/tests/core/store/current-value-window.test.ts @@ -9,8 +9,15 @@ * named grace on both boundaries and never narrowed, and a value dated well into the future is still * excluded, because that is what the comparison is for. * - * Every boundary below is stamped **in SQL**, so the stamp and the read share one clock source and the - * expectation does not depend on how long the test takes. + * Every boundary below is stamped **in SQL**, so the stamp and the read share one clock source: no read + * here compares JavaScript's clock against SQLite's. Sharing a clock *source* is not the same as sharing + * one clock *read*, though, and that difference cost this file a case. "An expiry a moment ago is still + * current" stamped a boundary half a grace in the past in one statement and read it in another, so it + * asserted that the read happens within the half that is left - measured, a delay of 20 ms after the + * store was reopened already answered "not active", while 0 and 10 ms answered "current". It failed on a + * loaded CI runner and not once in 30 local runs. A row's boundary is stamped by one statement and + * compared by another, so that case cannot be asked at row level at all; it is asked below as a relation, + * where the boundary and the predicate share one `'now'` in one statement. */ import assert from "node:assert/strict"; import { mkdtempSync, rmSync } from "node:fs"; @@ -70,13 +77,35 @@ test("a value dated well into the future is still not current", () => { }); }); -test("a value that expired a moment ago is still current", () => { - withStamped( - `expires_at = strftime('%Y-%m-%dT%H:%M:%fZ', 'now', '-${half} seconds')`, - (store, id) => { - assert.equal(store.demoteMemory(id, "expired a clock tick ago").residence, "stg"); - }, - ); +test("the freshness half widens one grace into the past and no further", () => { + // One statement, one clock read, and the store's own predicate rather than a copy of it: this pins the + // widening itself instead of a stopwatch, and it pins the boundary - at exactly one grace an expiry is + // already excluded. Every assertion holds whether or not SQLite hands out the same `'now'` twice inside + // one statement: the two boundaries that could disagree with it are a whole grace apart. + const raw = new DatabaseSync(":memory:"); + const isCurrent = (offsetSeconds: number): boolean => { + const boundary = `strftime('%Y-%m-%dT%H:%M:%fZ', 'now', '${offsetSeconds.toFixed(3)} seconds')`; + const row = raw + .prepare(`SELECT ${notExpired("r")} AS is_current FROM (SELECT ${boundary} AS expires_at) r`) + .get() as { is_current: number }; + return row.is_current === 1; + }; + try { + assert.equal(isCurrent(0), true, "an expiry stamped now is current"); + assert.equal( + isCurrent(-(CLOCK_GRACE_MS / 2000)), + true, + "an expiry inside the grace is current", + ); + assert.equal( + isCurrent(-(CLOCK_GRACE_MS / 1000)), + false, + "at exactly the grace it is not current", + ); + assert.equal(isCurrent(-60), false, "an expiry a minute ago is not current"); + } finally { + raw.close(); + } }); test("a value that expired a minute ago is not current", () => { diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 2d824151..d06b942c 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -1381,7 +1381,7 @@ const TARGETS: readonly Target[] = [ name: "the-window-does-not-grace-expiry", from: ' return `(${alias}.expires_at IS NULL OR ${alias}.expires_at > ${clockNow("earlier")})`;', to: ' return `(${alias}.expires_at IS NULL OR ${alias}.expires_at > ${clockNow("later")})`;', - expect: "a value that expired a moment ago is still current", + expect: "the freshness half widens one grace into the past and no further", }, { name: "the-grace-is-zero", From e31fe776d1c2a3acba5878565cb5c018d7f43bf9 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 19:59:46 +0800 Subject: [PATCH 16/32] The type gate had ten diagnostics; it now has none Three were in the store board's own test, which `npm run check` cannot see because tests are outside its project. One of them mattered: the self-judging tooth read a digest off the ticket, where the port's PatchWork is the pre-freeze declaration and carries none - so it delivered an artifact with no digest at all, and the store took it, because a delivery is a digest and a reference while the wire shape is the host's check rather than the board's. That test now freezes and builds its artifact the way the loop does, through one helper the worker shares instead of a second copy of the artifact shape. Two more were the store's private single-entry reader reached from tests, in this file and in tests/core/task-board.test.ts where they had been reported for a while. Both now use getTaskBoardEntryById, which is the same reader plus the channel guard the store asks for: the test says which channel it reads, and the private method stays private. The five in evals/ooo-execution were declared shapes that did not match what the code uses. The two evidence drivers refused a missing flag in a loop over an object, which narrows nothing, so one of them asserted non-null at its use site and the other did not compile at all; both now call a flag() helper that returns the value it checked. live-continuation.ts declared its per-attempt result as the metrics shape alone, which made `artifact` - the thing a continuation continues from - a property the file did not have; its type now follows the call that produces it. Verified: lsp_diagnostics reports 0 across the five files; check, lint, format:check, docs:check and agent:context:check clean; test:product 1527/1527 pass. --- evals/ooo-execution/board-deliver.ts | 11 +++++--- evals/ooo-execution/board-judge.ts | 25 ++++++++--------- evals/ooo-execution/live-continuation.ts | 6 ++-- tests/core/task-board.test.ts | 19 +++++++------ tests/integration/ooo-store-run-board.test.ts | 28 ++++++++++++------- 5 files changed, 50 insertions(+), 39 deletions(-) diff --git a/evals/ooo-execution/board-deliver.ts b/evals/ooo-execution/board-deliver.ts index 1180ba61..3a28a5a4 100644 --- a/evals/ooo-execution/board-deliver.ts +++ b/evals/ooo-execution/board-deliver.ts @@ -32,12 +32,15 @@ const { values } = parseArgs({ }, }); -const channel = values.channel; -const entryId = values.entry; -const agentId = values.agent; +const channel = flag("channel", values.channel); +const entryId = flag("entry", values.entry); +const agentId = flag("agent", values.agent); const ref = values.ref; -for (const [name, value] of Object.entries({ channel, entry: entryId, agent: agentId })) { +/** One required flag, refused by name. A loop over an object of them cannot narrow any of them, which is + * why this returns the value it checked instead of only rejecting a missing one. */ +function flag(name: string, value: string | undefined): string { if (!value) throw new Error(`--${name} is required`); + return value; } if (!values.daemon) { throw new Error("--daemon is required: the store path whose daemon serves this round's board"); diff --git a/evals/ooo-execution/board-judge.ts b/evals/ooo-execution/board-judge.ts index c24b032c..49468fbd 100644 --- a/evals/ooo-execution/board-judge.ts +++ b/evals/ooo-execution/board-judge.ts @@ -31,21 +31,18 @@ const { values } = parseArgs({ }, }); -const channel = values.channel; -const entryId = values.entry; -const agentId = values.agent; -const verdict = values.verdict as "accepted" | "rejected" | "undecidable" | undefined; -const reason = values.reason; -for (const [name, value] of Object.entries({ - channel, - entry: entryId, - agent: agentId, - verdict, - reason, -})) { +const channel = flag("channel", values.channel); +const entryId = flag("entry", values.entry); +const agentId = flag("agent", values.agent); +const verdict = flag("verdict", values.verdict) as "accepted" | "rejected" | "undecidable"; +const reason = flag("reason", values.reason); +/** One required flag, refused by name. A loop over an object of them cannot narrow any of them, which is + * why this returns the value it checked instead of only rejecting a missing one. */ +function flag(name: string, value: string | undefined): string { if (!value) throw new Error(`--${name} is required`); + return value; } -if (!["accepted", "rejected", "undecidable"].includes(verdict!)) { +if (!["accepted", "rejected", "undecidable"].includes(verdict)) { throw new Error(`--verdict must be accepted | rejected | undecidable, got ${verdict}`); } if (!values.daemon) { @@ -54,7 +51,7 @@ if (!values.daemon) { const state = roundDaemon(resolve(values.daemon)); { - const read = await boardCall(state, { action: "read", taskId: channel!, agentId, limit: 200 }); + const read = await boardCall(state, { action: "read", taskId: channel, agentId, limit: 200 }); if (read.action !== "read") throw new Error("the board did not answer a read with entries"); const entry = read.entries.find((candidate) => candidate.id === entryId); if (!entry) throw new Error(`no entry ${entryId} in ${channel}`); diff --git a/evals/ooo-execution/live-continuation.ts b/evals/ooo-execution/live-continuation.ts index fdec285a..0fd9b7af 100644 --- a/evals/ooo-execution/live-continuation.ts +++ b/evals/ooo-execution/live-continuation.ts @@ -484,8 +484,10 @@ if (role === "part1" || role === "part2") { const store = openStore(); // Declared outside the try so a failure can still report what the call cost: a failing attempt // that leaves no cost behind biases exactly the comparison this check exists to make. - let execution: { tokens?: number; turns?: number; reads?: number; sessionId?: string } | null = - null; + // The declared type follows the call that produces it: the metrics fields are read from here too, and + // an inline shape that lists only them made `artifact` - the thing a continuation actually continues + // from - a property this file did not have. + let execution: Awaited> | null = null; let entryId: string | undefined; let startedAt = 0; try { diff --git a/tests/core/task-board.test.ts b/tests/core/task-board.test.ts index c4b39b2e..81779fef 100644 --- a/tests/core/task-board.test.ts +++ b/tests/core/task-board.test.ts @@ -758,14 +758,15 @@ test("task board serial handoff promotes pending on claim, resolve, and expiry", content, expiresAt, }); - const stateOf = (store: NmgStore, id: string) => store.taskBoardEntry(id)!.serialState; + const stateOf = (store: NmgStore, channel: string, id: string) => + store.getTaskBoardEntryById(channel, id)?.serialState ?? null; // — Claim drives promotion ("回复=接手"): claiming the outstanding lets the // next pending promote; the claimed entry stops occupying the serial slot. const c1 = putHandoff("serial-promo-1", "first", "2099-01-01T00:00:00.000Z"); const c2 = putHandoff("serial-promo-1", "second", "2099-01-01T00:00:00.000Z"); - assert.equal(stateOf(store, c1.id), "outstanding"); - assert.equal(stateOf(store, c2.id), "pending"); + assert.equal(stateOf(store, "serial-promo-1", c1.id), "outstanding"); + assert.equal(stateOf(store, "serial-promo-1", c2.id), "pending"); assert.throws( () => store.claimTaskBoardEntry({ @@ -782,8 +783,8 @@ test("task board serial handoff promotes pending on claim, resolve, and expiry", agentId: "worker", leaseSeconds: 60, }); - assert.equal(stateOf(store, c1.id), null); // claimed → no longer the slot - assert.equal(stateOf(store, c2.id), "outstanding"); // promoted + assert.equal(stateOf(store, "serial-promo-1", c1.id), null); // claimed → no longer the slot + assert.equal(stateOf(store, "serial-promo-1", c2.id), "outstanding"); // promoted // — Resolve drives promotion: closing the outstanding promotes the next. const r1 = putHandoff("serial-promo-2", "first", "2099-01-01T00:00:00.000Z"); @@ -793,7 +794,7 @@ test("task board serial handoff promotes pending on claim, resolve, and expiry", entryId: r1.id, agentId: "worker", }); - assert.equal(stateOf(store, r2.id), "outstanding"); + assert.equal(stateOf(store, "serial-promo-2", r2.id), "outstanding"); // — Expiry drives promotion: pruning the expired outstanding moves the // pending up (RAII on the serial slot). @@ -801,15 +802,15 @@ test("task board serial handoff promotes pending on claim, resolve, and expiry", const e1 = putHandoff("serial-promo-3", "first", past); const e2 = putHandoff("serial-promo-3", "second", "2099-01-01T00:00:00.000Z"); store.pruneExpiredTaskBoardEntries(new Date().toISOString()); - assert.equal(store.taskBoardEntry(e1.id), null); // pruned - assert.equal(stateOf(store, e2.id), "outstanding"); + assert.equal(store.getTaskBoardEntryById("serial-promo-3", e1.id), null); // pruned + assert.equal(stateOf(store, "serial-promo-3", e2.id), "outstanding"); // — No pending to promote is a safe no-op (last outstanding resolved). const n1 = putHandoff("serial-promo-4", "only", "2099-01-01T00:00:00.000Z"); store.resolveTaskBoardEntry({ taskId: "serial-promo-4", entryId: n1.id, agentId: "worker" }); // No throw, no phantom outstanding in a fresh channel. const fresh = putHandoff("serial-promo-5", "fresh", "2099-01-01T00:00:00.000Z"); - assert.equal(stateOf(store, fresh.id), "outstanding"); + assert.equal(stateOf(store, "serial-promo-5", fresh.id), "outstanding"); }); }); diff --git a/tests/integration/ooo-store-run-board.test.ts b/tests/integration/ooo-store-run-board.test.ts index 48e3ad9c..21a6a6da 100644 --- a/tests/integration/ooo-store-run-board.test.ts +++ b/tests/integration/ooo-store-run-board.test.ts @@ -18,7 +18,11 @@ import { NmgStore } from "../../src/core/store.ts"; import { dispatchPlan, type PlanWorker } from "../../src/integration/ooo-dispatch.ts"; import type { SessionPlan } from "../../src/integration/ooo-execution.ts"; import { StoreRunBoard, type AcceptanceAnswer } from "../../src/integration/ooo-runner.ts"; -import type { PatchSubmission } from "../../src/integration/ooo-patch.ts"; +import { + preparePatchWork, + type FrozenPatchWork, + type PatchSubmission, +} from "../../src/integration/ooo-patch.ts"; import { coordinatedBoardWrite, createBoundEntry, @@ -29,7 +33,7 @@ import { const RUN = "run-1"; const CHANNEL = "run-1-board"; const FILE = "src/work.ts"; -const baseline = { [FILE]: "export const value = 1;\n" }; +const baseline: Readonly> = { [FILE]: "export const value = 1;\n" }; /** Two units of one plan, the second depending on the first, both patch work. */ function tasks() { @@ -91,12 +95,15 @@ function workspace(_taskId: string, files: readonly string[]): Readonly [path, baseline[path] ?? ""])); } -/** The artifact a worker produces: the wire shape the host validates, with this task's own marker. */ -const worker: PlanWorker = async (taskId, frozen) => - JSON.stringify({ +/** The artifact a task's worker produces: the wire shape the host validates, with the task's marker. */ +function artifactFor(taskId: string, frozen: FrozenPatchWork): string { + return JSON.stringify({ digest: frozen.digest, files: [{ path: FILE, content: `${baseline[FILE]}// ${taskId}\n` }], }); +} + +const worker: PlanWorker = async (taskId, frozen) => artifactFor(taskId, frozen); /** The acceptance reads the submission and answers; its name is not the deliverer's. */ const acceptance = { @@ -166,7 +173,7 @@ test("one plan runs through the product's own board, judged by someone other tha assert.deepEqual(Object.keys(accepted).sort(), ["P", "T"]); assert.match(accepted.P!, /\/\/ P/u); for (const taskId of ["P", "T"]) { - const entry = store.taskBoardEntry(bindingEntry(store, taskId)); + const entry = store.getTaskBoardEntryById(CHANNEL, bindingEntry(store, taskId)); assert.ok(entry, `task ${taskId} has an entry`); assert.equal(entry.verdict, "accepted"); assert.equal(entry.judgedBy, acceptance.agentId); @@ -188,10 +195,11 @@ test("the deliverer cannot judge its own delivery, so a run cannot self-accept", }); const entryId = bindingEntry(store, "P"); const ticket = board.claim("P", "worker:P"); - const artifact = JSON.stringify({ - digest: ticket.patch!.digest, - files: [{ path: FILE, content: `${baseline[FILE]}// P\n` }], - }); + // The artifact the loop would have delivered, frozen the way the loop freezes it. Reading a digest + // off the ticket wrote an artifact with no digest at all, and the store took it: a delivery is a + // digest and a reference, and the wire shape is the host's check rather than the board's. + const frozen = preparePatchWork({ ...ticket.patch!, attempt: ticket.attempt }); + const artifact = artifactFor("P", frozen); coordinatedBoardWrite(store, { runId: RUN, entryId, From 9fb3beb6a8c334d2ec0ef726d3c7b4072ef02b94 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 20:15:31 +0800 Subject: [PATCH 17/32] One home for the identity of work A verdict binds to a digest rather than to bytes, so the digest's convention is a rule a stored record depends on - and it had four module-private homes. Two of them were the identical two lines (`createHash("sha256").update(JSON.stringify(value)).digest("hex")`) in `task-semantics.ts` and `ooo-board.ts`, and two callers had folded their own shortening into their copy, so a twelve-character report identity and a sixteen-character branch identity read as if the shortening were part of the rule. `src/integration/work-identity.ts` now owns it: `workDigest(bytes)` for text or raw bytes, `workDigestOf(value)` over JSON text, sha256 in hex at full length. Canonicalisation stays the caller's, and the module says why: JSON key order is whatever the caller built, so a caller that needs one identity per shape rather than one per serialisation freezes the shape first - which is exactly what `preparePatchWork` does when it sorts the files, serialises once, and returns a digest rather than a serialisation. The shortening two callers want stays theirs, and it now says so by slicing. Sites that differ on purpose keep their own convention, and the module names them so a later sweep cannot unify three different rules: a search index's base64url content hash, a session identifier, and a protocol-visible `sha256:` prefixed identity. Two named mutants make the convention checkable - another encoding, and a JSON variant that does not digest JSON - both caught by the named case, restored byte-identically. The write face's type did not land, and this commit records the measurement that says so rather than building it: the artifact body inside a result already has a home (`artifactEnvelope`/`artifactFromText` in `ooo-session-mechanism.ts`), while the entry body wrapping it has one writer and two readers that are two *different* rules - the probe parses the loop's envelope and binds a verdict to the ticket's digest, the product board reads the artifact of a body that may be the artifact itself. Two rules with one implementation each are not a seam. Verified: check, lint, format:check, docs:check and agent:context:check clean (the new file is claimed by the ooo-execution route); test:product 1529/1529 pass twice; the work-identity mutant target 2 of 2 caught. --- .pi/extensions/nmg/ooo-execution.ts | 4 +-- agent-context.yaml | 1 + docs/design/mechanism-in-the-middle.md | 18 ++++++++++- docs/design/mechanism-in-the-middle.zh-CN.md | 3 +- evals/ooo-execution/admission.ts | 8 ++--- evals/ooo-execution/board-deliver.ts | 4 +-- evals/ooo-execution/board-judge.ts | 4 +-- evals/ooo-execution/board-worker.ts | 4 +-- evals/ooo-execution/live-continuation.ts | 5 +-- evals/ooo-execution/speculation-pilot.ts | 4 +-- src/integration/ooo-board.ts | 12 +++---- src/integration/ooo-patch.ts | 4 +-- src/integration/ooo-runner.ts | 5 ++- src/integration/task-semantics.ts | 8 ++--- src/integration/work-identity.ts | 34 ++++++++++++++++++++ tests/integration/work-identity.test.ts | 33 +++++++++++++++++++ tools/mutation-teeth.ts | 22 +++++++++++++ 17 files changed, 138 insertions(+), 35 deletions(-) create mode 100644 src/integration/work-identity.ts create mode 100644 tests/integration/work-identity.test.ts diff --git a/.pi/extensions/nmg/ooo-execution.ts b/.pi/extensions/nmg/ooo-execution.ts index 7177b8c6..704b77f9 100644 --- a/.pi/extensions/nmg/ooo-execution.ts +++ b/.pi/extensions/nmg/ooo-execution.ts @@ -7,9 +7,9 @@ import { SessionManager, SettingsManager, } from "@earendil-works/pi-coding-agent"; -import { createHash } from "node:crypto"; import { Type } from "typebox"; import { snapshotPrompt, type SnapshotInput } from "../../../src/integration/ooo-execution.ts"; +import { workDigestOf } from "../../../src/integration/work-identity.ts"; import type { FrozenPatchWork } from "../../../src/integration/ooo-patch.ts"; import { ARTIFACT_TOOL, @@ -36,7 +36,7 @@ import { * that agree on the spec can still differ here, and an instrument version stops being a matter of prose: * the report carries the digest the run actually used. */ function digestOf(input: SessionRunInput): string { - return createHash("sha256").update(JSON.stringify(input)).digest("hex").slice(0, 12); + return workDigestOf(input).slice(0, 12); } /** Pi-only execution adapter. Selection, ownership and acceptance are not model decisions. diff --git a/agent-context.yaml b/agent-context.yaml index 54213d38..b0f26d23 100644 --- a/agent-context.yaml +++ b/agent-context.yaml @@ -227,6 +227,7 @@ routes: - src/integration/task-semantics-interleavings.ts - src/integration/task-semantics-model.ts - src/integration/task-semantics.ts + - src/integration/work-identity.ts owners: - docs/design/ooo-execution-bootstrap.md - docs/design/task-unit-semantics.md diff --git a/docs/design/mechanism-in-the-middle.md b/docs/design/mechanism-in-the-middle.md index e0183189..b447d4dd 100644 --- a/docs/design/mechanism-in-the-middle.md +++ b/docs/design/mechanism-in-the-middle.md @@ -75,7 +75,23 @@ against an imagined second shape, and the imagined one is the one that gets buil `tests/core/task-board.test.ts` pins the rule directly. The adapter's `excerpt` helper stays for memory results, which are not board entries and keep their own rule. - **The write face's type did not land.** The three fillers are three kinds of entry rather than one shape - drawn three ways, so the trigger for a type is a fourth filler. + drawn three ways, so the trigger for a type is a fourth filler. Measured again here: the _artifact_ body + inside a result already has a home (`artifactEnvelope` and `artifactFromText` in + `src/integration/ooo-session-mechanism.ts`, which is why the worker's tool and its text channel obey one + contract), while the _entry_ body that wraps it has one writer and two readers that are two different + rules - the probe parses the loop's envelope and binds a verdict to the ticket's digest, and the product + board reads the artifact of a body that may be the artifact itself, because an agent that delivers through + the board tool writes the bytes it produced. Two rules with one implementation each are not a seam. +- **The identity of work has one home.** `src/integration/work-identity.ts` owns the rule a verdict binds + to: sha256, hex, full length, over exactly the bytes or the JSON text the caller froze. Four modules had + written the same two lines - twice as an identical private `digestOf(value)` - and two of those callers + had folded their own shortening into it, so a twelve-character report identity and a sixteen-character + branch identity read as part of the rule instead of a caller's choice. Canonicalisation stays the caller's + and the module says so: JSON key order is whatever the caller built, which is why freezing a patch is what + makes its digest stable. Two named mutants make the convention checkable - another encoding, and a JSON + variant that does not digest JSON. Sites that differ on purpose (a base64url search content hash, a + session id, a protocol-visible `sha256:` prefix) keep their own convention, and the module names them, so + a later sweep cannot unify three different rules. - **The middle's seam did not land.** Check (c) says why. - **The port has a third implementation, and it is the product's.** `src/integration/ooo-runner.ts` is the store's own board: it projects the run's frozen table and its board entries into the facts the shared diff --git a/docs/design/mechanism-in-the-middle.zh-CN.md b/docs/design/mechanism-in-the-middle.zh-CN.md index a2a27aa8..2ba177f2 100644 --- a/docs/design/mechanism-in-the-middle.zh-CN.md +++ b/docs/design/mechanism-in-the-middle.zh-CN.md @@ -48,7 +48,8 @@ ## 随本文一起落地的东西 - **读面有了一个家。** `src/core/board-entry-view.ts` 拥有这条规则:孤立的 `memory=` 指针原样返回;更长的内容折叠成**一行有界文本**;界**归调用方**,因为"塞得下多少"是读者的决定。store 的紧凑读与宿主适配器的广播都调它;store 里那份私有副本和适配器里那条裸切都删掉了。三条规则里有一条不只是重复,而是**错的**:`.pi/extensions/nmg/index.ts` 那条 140 字符裸切不折叠空白,于是一个多行内容会以多行广播的形态发出去。`tests/core/task-board.test.ts` 直接钉住这条规则。适配器里的 `excerpt` 助手留给记忆结果——那不是黑板条目,保留它自己的规则。 -- **写面的类型没有落地。** 三种填法是三种**条目**,不是同一个形状画了三遍;所以写类型的触发条件是出现第四种填法。 +- **写面的类型没有落地。** 三种填法是三种**条目**,不是同一个形状画了三遍;所以写类型的触发条件是出现第四种填法。本文重新量了一次:结果里的**产物正文**早有一个家(`src/integration/ooo-session-mechanism.ts` 里的 `artifactEnvelope` 与 `artifactFromText`,这也是 worker 的工具通道与文本通道遵守同一契约的原因),而包裹它的**条目正文**有一个写者、两个读者,且这两个读者是两条不同的规则——探针解析循环的信封并把判词绑在票据的摘要上,产品黑板读的是"一个可能本身就是产物的正文"里的产物,因为用黑板工具交付的 agent 写下的就是它产出的字节。两条规则各有各的一个实现,那不是一条缝。 +- **工作的身份有了一个家。** `src/integration/work-identity.ts` 拥有判词所绑的那条规则:sha256、hex、全长,作用在调用方**冻结**过的那些字节或 JSON 文本上。此前有四个模块写过同两行——其中两处是逐字相同的私有 `digestOf(value)`——而其中两个调用方把自己的截断也塞了进去,于是十二字符的报告身份和十六字符的分支身份读起来像是规则的一部分,而不是调用方的选择。规范化仍归调用方,模块明说这一点:JSON 的键顺序是调用方自己选出来的,所以**冻结**一个补丁才是让它的摘要稳定的那一步。两个具名 mutant 让这条约定可被检查——换成另一种编码,以及一个不消化 JSON 的 JSON 变体。那些**故意**不同(搜索索引的 base64url 内容哈希、会话 id、协议可见的 `sha256:` 前缀)的站点保留自己的约定,模块把它们点名,以免后来的一次清扫把三条不同的规则合成一条。 - **中间的缝没有落地。** 检查(c)说明了为什么。 - **端口有了第三个实现,而且是产品自己的那个。** `src/integration/ooo-runner.ts` 是 store 自己的黑板:它把运行的冻结表与黑板条目**投影**成共享规则要读的那些事实,自己不带任何合法性规则。现在有三个实现满足 `DispatchBoard`——探针黑板、dispatch 测试里的桩、和这一个——它们之间**只有投影不同**,这是抽象在起作用,而不是关于它的说法。 - **投影里有一条容易搞错的契约。** 共享的验收规则把判词**绑在运行携带的产物上**:它拿判词的摘要与产物**值**比较,所以关于另一个产物的判词不能冒充验收。把 store 自己的交付物哈希填进那个字段,会让每个已验收单元都读成未验收,而症状是**无声的**——下一个单元永不释放,且任何地方都不报错。两种身份都是正当的,但那个字段只属于产物值。 diff --git a/evals/ooo-execution/admission.ts b/evals/ooo-execution/admission.ts index c699974d..1a7d2368 100644 --- a/evals/ooo-execution/admission.ts +++ b/evals/ooo-execution/admission.ts @@ -1,4 +1,6 @@ -import { createHash, randomUUID } from "node:crypto"; +import { randomUUID } from "node:crypto"; + +import { workDigestOf } from "../../src/integration/work-identity.ts"; export interface Task { id: string; @@ -59,9 +61,7 @@ export function createAdmissionGate(plan: Task[]) { if (!dependenciesReady(task)) throw new Error("unfulfilled dependencies"); const attempt = (attempts.get(taskId)?.attempt ?? 0) + 1; if (!Number.isSafeInteger(attempt)) throw new Error("attempt exhausted"); - const inputDigest = createHash("sha256") - .update(JSON.stringify([task.revision, task.input, dependencyResults(task)])) - .digest("hex"); + const inputDigest = workDigestOf([task.revision, task.input, dependencyResults(task)]); const ticket = Object.freeze({ runId, taskId, diff --git a/evals/ooo-execution/board-deliver.ts b/evals/ooo-execution/board-deliver.ts index 3a28a5a4..e404baa1 100644 --- a/evals/ooo-execution/board-deliver.ts +++ b/evals/ooo-execution/board-deliver.ts @@ -11,12 +11,12 @@ * --daemon --channel --entry --agent \ * --digest [--ref ] [--summary ] */ -import { createHash } from "node:crypto"; import { readFileSync, statSync } from "node:fs"; import { resolve } from "node:path"; import { parseArgs } from "node:util"; import { boardCall, roundDaemon } from "./round-client.ts"; +import { workDigest } from "../../src/integration/work-identity.ts"; // node:util owns flag parsing; an unknown flag or a repeated one is an error rather // than something this script silently ignores. @@ -62,7 +62,7 @@ const state = roundDaemon(resolve(values.daemon)); // If the artifact is a readable file, never trust the caller's digest: recompute. const digest = ref && statSync(ref, { throwIfNoEntry: false })?.isFile() - ? createHash("sha256").update(readFileSync(ref)).digest("hex") + ? workDigest(readFileSync(ref)) : values.digest; if (!digest) throw new Error("--digest is required when --ref is not a readable file"); if (values.digest && values.digest !== digest) { diff --git a/evals/ooo-execution/board-judge.ts b/evals/ooo-execution/board-judge.ts index 49468fbd..d0fe63d8 100644 --- a/evals/ooo-execution/board-judge.ts +++ b/evals/ooo-execution/board-judge.ts @@ -11,12 +11,12 @@ * --daemon --channel ooo-probe: --entry --agent coordinator \ * --verdict accepted|rejected|undecidable --reason "..." */ -import { createHash } from "node:crypto"; import { readFileSync, statSync } from "node:fs"; import { resolve } from "node:path"; import { parseArgs } from "node:util"; import { boardCall, roundDaemon } from "./round-client.ts"; +import { workDigest } from "../../src/integration/work-identity.ts"; // node:util owns flag parsing; an unknown flag or a repeated one is an error rather // than something this script silently ignores. @@ -65,7 +65,7 @@ const state = roundDaemon(resolve(values.daemon)); let observed = "unavailable"; if (ref) { const bytes = statSync(ref).size; - observed = createHash("sha256").update(readFileSync(ref)).digest("hex"); + observed = workDigest(readFileSync(ref)); console.log(`[judge] ${ref} bytes=${bytes} digest=${observed}`); } if (verdict === "accepted" && observed !== entry.deliverableDigest) { diff --git a/evals/ooo-execution/board-worker.ts b/evals/ooo-execution/board-worker.ts index 93873d8f..5aed1637 100644 --- a/evals/ooo-execution/board-worker.ts +++ b/evals/ooo-execution/board-worker.ts @@ -18,13 +18,13 @@ * it holds the claim, that the artifact file exists and is non-empty, that the digest recomputes, * and that the daemon recorded exactly the digest it reported. */ -import { createHash } from "node:crypto"; import { existsSync, mkdirSync, readFileSync, statSync, writeFileSync } from "node:fs"; import { dirname, resolve } from "node:path"; import { spawnSync } from "node:child_process"; import { parseArgs } from "node:util"; import { boardCall, roundDaemon } from "./round-client.ts"; +import { workDigest } from "../../src/integration/work-identity.ts"; const DEFAULT_SUITES = [ "tests/core/task-board-deliverable.test.ts", @@ -113,7 +113,7 @@ const suites = (values.suites ?? DEFAULT_SUITES.join(",")).split(",").filter(Boo if (!existsSync(out) || statSync(out).size === 0) throw new Error(`artifact ${out} is empty`); // 4. The artifact's identity, recomputed from the bytes that were just written. - const digest = createHash("sha256").update(readFileSync(out)).digest("hex"); + const digest = workDigest(readFileSync(out)); // Two reporter shapes exist in the wild: TAP (`# pass 27`) and node's spec reporter // (`ℹ pass 27`). Reading only one of them made this worker report 0/0 for a 27/27 run — // a self-report that contradicted its own artifact. Parse both, and refuse to deliver diff --git a/evals/ooo-execution/live-continuation.ts b/evals/ooo-execution/live-continuation.ts index 0fd9b7af..9f9069c8 100644 --- a/evals/ooo-execution/live-continuation.ts +++ b/evals/ooo-execution/live-continuation.ts @@ -22,12 +22,13 @@ * Refusals: no --live means no model call; a missing provider, model, run directory or task id is * named rather than guessed; a judge refuses to judge its own delivery. */ -import { createHash } from "node:crypto"; import { appendFileSync, mkdirSync, readFileSync, writeFileSync } from "node:fs"; import { join, resolve } from "node:path"; import { pathToFileURL } from "node:url"; import { parseArgs } from "node:util"; +import { workDigest } from "../../src/integration/work-identity.ts"; + const { NmgStoreBase } = await import("../../src/core/store/base.ts"); const { preparePatchWork, patchCandidate } = await import("../../src/integration/ooo-patch.ts"); const { executePiPatch } = await import("../../.pi/extensions/nmg/ooo-execution.ts"); @@ -356,7 +357,7 @@ function record(entry: Record): void { } function digestOf(bytes: Buffer | string): string { - return createHash("sha256").update(bytes).digest("hex"); + return workDigest(bytes); } function taskOf(id: string | undefined): Task { diff --git a/evals/ooo-execution/speculation-pilot.ts b/evals/ooo-execution/speculation-pilot.ts index ce3f062d..5d87ed94 100644 --- a/evals/ooo-execution/speculation-pilot.ts +++ b/evals/ooo-execution/speculation-pilot.ts @@ -14,11 +14,11 @@ // A published candidate is verified by running the unit's own frozen check against it, in a copy of the // fixture directory, so the quality term is a real check result and not the model's own claim. import { cpSync, mkdirSync, readFileSync, writeFileSync } from "node:fs"; -import { createHash } from "node:crypto"; import { execFileSync } from "node:child_process"; import { randomUUID } from "node:crypto"; import { patchCandidate, preparePatchWork } from "../../src/integration/ooo-patch.ts"; +import { workDigest } from "../../src/integration/work-identity.ts"; import { speculationOutcome, type ResolvedPredicate, @@ -71,7 +71,7 @@ writeFileSync( "failures unexplainable.", ].join("\n") + "\n", ); -const digestOf = (value: string) => createHash("sha256").update(value).digest("hex").slice(0, 16); +const digestOf = (value: string) => workDigest(value).slice(0, 16); const interfaceDigest = digestOf(readFileSync(`${FIXTURE}/interface.ts`, "utf8")); /** The frozen work for one attempt, always under a fresh ticket: a discarded branch's ticket is never diff --git a/src/integration/ooo-board.ts b/src/integration/ooo-board.ts index 2873bcae..fdbb10c6 100644 --- a/src/integration/ooo-board.ts +++ b/src/integration/ooo-board.ts @@ -1,4 +1,4 @@ -import { createHash, randomUUID } from "node:crypto"; +import { randomUUID } from "node:crypto"; import { existsSync } from "node:fs"; import { DatabaseSync } from "node:sqlite"; import { NmgStore } from "../../src/core/store.ts"; @@ -15,6 +15,7 @@ import { type PatchLimits, type PatchSubmission, } from "./ooo-patch.ts"; +import { workDigest, workDigestOf } from "./work-identity.ts"; export type ProbeOperation = SnapshotWork["operation"]; @@ -77,10 +78,9 @@ export function roundChannel(runId: string): string { const RETENTION_OWNER = "coordinator"; export function artifactDigest(commit: string): string { - return createHash("sha256").update(commit).digest("hex"); + return workDigest(commit); } const policy = "narrow-snapshot-work/v3"; -const digest = (value: unknown) => createHash("sha256").update(JSON.stringify(value)).digest("hex"); interface Row { id: string; @@ -372,7 +372,7 @@ export class BoardAdmission extends NmgStore { // that held it. It was a pointer to the entry whose verdict accepted an artifact, and the // verdict is looked up by artifact digest now — a migrated store keeps the artifact plus the // entries that decide it, and keeps no stale pointer that a reader could mistake for one. - const wanted = `${policy}:${digest(plan)}`; + const wanted = `${policy}:${workDigestOf(plan)}`; const runs = this.db .prepare("SELECT run_id, policy FROM ooo_probe_runs ORDER BY created_at") .all() as unknown as { run_id: string; policy: string }[]; @@ -635,7 +635,7 @@ export class BoardAdmission extends NmgStore { private inputDigest(row: Row): string { if (this.patchSpec(row)) return this.patchFrozen(row, row.attempt > 0 ? row.attempt : 1).digest; - return digest([ + return workDigestOf([ policy, row.revision, row.source_revision, @@ -826,7 +826,7 @@ export class BoardAdmission extends NmgStore { // Deliberately NOT run-scoped: a check id is what a round log records, and a replay runs // under a new runId while reproducing the same attempts. Uniqueness in the store is the // (run_id, task_id) key, so two runs may share a check id without colliding. - checkId: digest([id, attempt]).slice(0, 32), + checkId: workDigestOf([id, attempt]).slice(0, 32), attempt, inputDigest: this.inputDigest(row), owner, diff --git a/src/integration/ooo-patch.ts b/src/integration/ooo-patch.ts index 3cfcaaa5..e89e4ba5 100644 --- a/src/integration/ooo-patch.ts +++ b/src/integration/ooo-patch.ts @@ -1,4 +1,4 @@ -import { createHash } from "node:crypto"; +import { workDigest } from "./work-identity.ts"; export interface PatchBudget { perFile: number; @@ -217,7 +217,7 @@ export function preparePatchWork(input: PatchWork) { }); const serialized = JSON.stringify(work); if (Buffer.byteLength(serialized, "utf8") > 256_000) throw new Error("snapshot budget exceeded"); - const digest = createHash("sha256").update(serialized).digest("hex"); + const digest = workDigest(serialized); return Object.freeze({ work, digest }); } diff --git a/src/integration/ooo-runner.ts b/src/integration/ooo-runner.ts index 7740c3d6..69e8d68d 100644 --- a/src/integration/ooo-runner.ts +++ b/src/integration/ooo-runner.ts @@ -15,8 +15,7 @@ * preparing the workspace belongs to the patch path's caller) and an acceptance whose identity is * independent of the deliverer (the store refuses a deliverer that judges its own delivery). */ -import { createHash } from "node:crypto"; - +import { workDigest } from "./work-identity.ts"; import type { NmgStore } from "../core/store.ts"; import { taskBoardClaimIsLive } from "../core/store/base.ts"; import { TASK_BOARD_VERDICTS } from "../core/types.ts"; @@ -155,7 +154,7 @@ export class StoreRunBoard implements DispatchBoard { taskId: this.channel, entryId: binding.entryId, agentId: input.agentId, - digest: createHash("sha256").update(artifact).digest("hex"), + digest: workDigest(artifact), ref: artifact, }), }); diff --git a/src/integration/task-semantics.ts b/src/integration/task-semantics.ts index c28c2943..290eff9d 100644 --- a/src/integration/task-semantics.ts +++ b/src/integration/task-semantics.ts @@ -16,8 +16,8 @@ * (execution limits re-expressed as wall clock), or an unknown requirement * kind gets a refusal naming the task and the field. */ -import { createHash } from "node:crypto"; import { startableTasks, type DispatchTask } from "./ooo-execution.ts"; +import { workDigestOf } from "./work-identity.ts"; import { CONCLUSION_KINDS, MAX_PATCH_BUDGET, @@ -128,10 +128,6 @@ const UNIT_OBLIGATIONS = [ "stoppable-voidable", ] as const; -function digestOf(value: unknown): string { - return createHash("sha256").update(JSON.stringify(value)).digest("hex"); -} - function unknownKeys(value: unknown, allowed: ReadonlySet): string[] { if (!value || typeof value !== "object") return []; return Object.keys(value as Record).filter((key) => !allowed.has(key)); @@ -435,7 +431,7 @@ export function compileTaskUnits(input: CompileInput): CompiledTasks { units, refusals, legal: refusals.length === 0, - digest: digestOf({ plan, requires: input.requires ?? {} }), + digest: workDigestOf({ plan, requires: input.requires ?? {} }), }; } diff --git a/src/integration/work-identity.ts b/src/integration/work-identity.ts new file mode 100644 index 00000000..3d180f9c --- /dev/null +++ b/src/integration/work-identity.ts @@ -0,0 +1,34 @@ +/** + * What names a piece of work, in one place. + * + * A verdict binds to a digest rather than to bytes: the board records the digest of what was delivered, + * and a second delivery of other bytes cannot inherit the first one's judgement. That digest is sha256 in + * hex, and it is written here rather than at each site because the same two lines had been written in four + * modules - twice as an identical private `digestOf(value)` - while two of those callers had folded their + * own shortening into it, so a twelve-character report identity and a sixteen-character branch identity + * looked like part of the rule instead of a caller's choice. Only the algorithm and the encoding live here, + * because those are the part a reader of a stored digest depends on. + * + * Sites that need another encoding on purpose keep their own convention: a search index's base64url content + * hash, a session identifier, and a CLI-protocol identity that carries a visible `sha256:` prefix are three + * different rules about three different values, not three spellings of this one. + * + * Canonicalisation is deliberately the caller's job, and that is why the JSON variant takes a value: JSON + * key order is whatever the caller built, so a caller that needs one identity per shape rather than one per + * serialisation has to shape or sort the value first. Freezing a patch does exactly that before digesting + * it, which is why `preparePatchWork` returns a digest and not a serialisation. Shortening a digest is the + * caller's choice too, and it says so by slicing - two callers want short identities for logs and reports. + */ +import { createHash } from "node:crypto"; + +/** The identity of work bytes: sha256, hex, full length. Text and raw bytes are the same rule, because an + * artifact is bytes and a stored digest does not care which reader produced them. */ +export function workDigest(bytes: string | Uint8Array): string { + return createHash("sha256").update(bytes).digest("hex"); +} + +/** The identity of a value, over its JSON text. The caller owns key order: two serialisations of one shape + * are two identities unless the caller froze the shape first. */ +export function workDigestOf(value: unknown): string { + return workDigest(JSON.stringify(value)); +} diff --git a/tests/integration/work-identity.test.ts b/tests/integration/work-identity.test.ts new file mode 100644 index 00000000..526de2be --- /dev/null +++ b/tests/integration/work-identity.test.ts @@ -0,0 +1,33 @@ +/** + * One home for the identity of work, so a stored digest cannot come to mean two things. + * + * The rule these cases pin is not the hash itself - node owns that - but the convention a reader of a + * stored digest depends on: sha256, hex, full length, over exactly the bytes or the JSON text the caller + * froze. The mutants that catch a drift in that convention are named in `tools/mutation-teeth.ts` for + * `src/integration/work-identity.ts`. + */ +import assert from "node:assert/strict"; +import test from "node:test"; + +import { workDigest, workDigestOf } from "../../src/integration/work-identity.ts"; + +test("the identity of work bytes is sha256 in hex, at full length", () => { + const bytes = Buffer.from("deliver these bytes\n", "utf8"); + // The known answer, from `printf 'deliver these bytes\n' | sha256sum` at the time of writing: a constant + // rather than another call to the same hash, so the case fails if the encoding changes rather than only + // if both sides change together. Text and bytes are asserted to be the same rule. + assert.equal( + workDigest(bytes), + "33f9ff34371d8a7e7cd9810ca147946bfdcd1894160bd7c93eeb4a650205fd44", + ); + assert.equal(workDigest("deliver these bytes\n"), workDigest(bytes)); + assert.match(workDigest(bytes), /^[0-9a-f]{64}$/u); + assert.notEqual(workDigest(bytes), workDigest("deliver these bytes")); +}); + +test("the JSON variant digests JSON text, so key order is the caller's", () => { + assert.equal(workDigestOf({ a: 1, b: 2 }), workDigest(JSON.stringify({ a: 1, b: 2 }))); + // Not a defect to fix here: a caller that needs one identity per shape freezes the shape first, which is + // what `preparePatchWork` does before it digests. This case exists so that rule is written down. + assert.notEqual(workDigestOf({ a: 1, b: 2 }), workDigestOf({ b: 2, a: 1 })); +}); diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index d06b942c..e37493ef 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -1397,6 +1397,28 @@ const TARGETS: readonly Target[] = [ }, ], }, + { + // The identity of work. Each mutant is one of the ways a stored digest can stop naming the bytes a + // verdict was about: another encoding (which changes every digest without changing any bytes), and a + // different serialisation of the JSON variant. The shortening two callers apply stays out of it - it + // is their choice, and a mutant there would be testing a caller rather than the rule. + target: "src/integration/work-identity.ts", + suites: ["tests/integration/work-identity.test.ts"], + mutants: [ + { + name: "the-identity-is-not-hex", + from: ' return createHash("sha256").update(bytes).digest("hex");', + to: ' return createHash("sha256").update(bytes).digest("base64url");', + expect: "the identity of work bytes is sha256 in hex, at full length", + }, + { + name: "the-json-variant-does-not-digest-json", + from: " return workDigest(JSON.stringify(value));", + to: " return workDigest(String(value));", + expect: "the JSON variant digests JSON text, so key order is the caller's", + }, + ], + }, ]; interface MutantOutcome { From 95fb3bcf363e91b2ea4ae08958033c1da4dd6a2e Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 20:36:44 +0800 Subject: [PATCH 18/32] The completeness ledger stops over-claiming in both directions Asked whether the design is complete, the authority to answer with is docs/design/task-unit-semantics-obligations.md, and it contradicted itself in two places - one in each direction. Its counts paragraph said "the only arm still unrun is the E arm itself", while the E arm's own section, the row table and the closing paragraph all say it ran twice on 2026-09-19 with no gain claimed. A reader trusting the opening paragraph would think the measurement phase was one run short of closed; a reader trusting the closing one would think it was closed. The sentence now points at the two records that settle it. Its verification section recorded the LSP diagnostics as "recorded rather than repaired", naming one error each in board-deliver.ts and board-judge.ts plus a { pid: 0 } fallback in an evidence-driver test. Those are repaired now: the drivers' flag narrowing and live-continuation.ts's declared result shape landed in e31fe776, and this commit fixes the test's fallback, which needed the startedAt that ServerState requires (a pid of zero and an empty start time still says "no server", which is what the code means). What the LSP still reports is stated as the number it is: 31 diagnostics in seven evals/ files this arc does not own - benchmarks/run.ts 7, longmemeval/run.ts 13, controller/run.ts 3, natural-maintenance/audit.ts 3, hierarchy-scale/run.ts 2, longmemeval/score.ts 2, omnimemeval/bridge.ts 1. The sentence claims zero only for the files this arc touched, because that is what was measured; the tests/ directory as a whole timed out rather than answering, and an unmeasured claim is not written down. Verified: docs:check 288 files, 0 errors, 0 warnings; format:check clean; lint clean; tests/integration/ooo-evidence-drivers.test.ts 8 pass, 0 fail. --- docs/design/task-unit-semantics-obligations.md | 4 ++-- tests/integration/ooo-evidence-drivers.test.ts | 6 +++++- 2 files changed, 7 insertions(+), 3 deletions(-) diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index dac01369..5e97c0cb 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -3,7 +3,7 @@ **Authority:** living ledger for `docs/design/task-unit-semantics.md` — each row is one obligation from that design; progress is counted in rows moved to `proven`, not in edits made. -Counts at this revision, computed from the rows below rather than from memory: **A** 5 proven, 0 partly, 0 owed (5 rows); **B** 9 proven, 0 partly, 0 owed (9 rows, B7 nothing to fail); **C** 4 proven, 0 partly, 0 owed (4 rows); **D** 13 proven, 0 partly, 0 owed, 1 not applicable (14 rows; D10's subject, `runCycle`, was retired by [the retirement decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); **E** 9 proven as a seam with offline proofs and no source wired (9 rows); **G** 7 proven, 0 partly, 0 owed (7 rows). **F** complete except the fusion slice: its offline layer, the plan's single home, the arms' driver and its slot budget, the two task families and the real-model pilot have all landed and are recorded below, the last one in [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md) with its own sample-size caveat; **F4** (execution fusion) has landed its offline half (legality, accounting, the driver's session policy) **and** the live mechanism it needed - `createPiSessionRunner` in the extension, `piSessionWorker` + `--session-runner` in the driver - and the paid D arm is measured: fused `unitsPerSession: 2` against the `1` control, 3 reps each, one session of two units against two sessions of one, quality parity, ~1.9 s faster per run and token-neutral on that plan shape - a two-rep trend whose own spreads (1.2 s each) exceeded the 1.9 s gap, so it does not price the session-startup term the cost model had left `unmeasured`; on the four-unit fine plan the same cap experiment measures 3.2-4.2 s of wall clock per avoided session with the saving five times its within-cell spread, while its token columns settle nothing at three reps ([the D arm](../experiments/execution/ooo-arms-pilot-2026-09-18.md#d-arm-fusion-added-2026-09-19)). The A, B and C cells were then run on the plain path with the same fixture, worker, envelope and parent check (coarse one rep: 9 726 ms / 9 018 tokens; fine at one slot two reps: 21 885 and 23 273 ms, median 22 579; the same spec at two slots two reps: 18 565 and 19 336 ms, median 18 950, and cheaper in money than the one-slot pair), which also made the chain surface's own price visible: declaring fusion at a bound of one costs 3 607 ms and 15 412 tokens more than the plain path for the same plan ([the A-D cells](../experiments/execution/archive/ooo-arms-2026-09-19/README.md)). Their reports record `inputTokens`/`outputTokens`/`cacheRead`/`cost` per unit and the commit they ran from, so a comparison can name its instrument instead of arguing about it. **F5** (speculation lifecycle) landed offline in the same shared module: the design's `assumptions=[{predicateId, version, expected}]` as a declaration, the first experiment's bounds (exactly one pending fact, no speculative successor, no irreversible operation) refused by name, and the three outcomes read from authoritative evidence - true publishes, false discards the candidate and closes its branch session, unknown waits, because a missing reading is not permission to publish - with 9 cases and 5 named mutants. The only arm still unrun is the **E arm** itself (budgeted speculation against a real model), which needs an instrument that decides the guessed fact and re-runs the real path under a new ticket. The counts sentence above is what this line used to over-claim: it read "**F** complete" while the F4 row said otherwise. F1 built the advisory cost model the design orders before any paid call (`evals/ooo-execution/cost-model.ts`, derived plan graphs, self-checks, no quality term) and recorded its sweep in [the experiment record](../experiments/execution/ooo-cost-model-2026-09-17.md); the sweep says one slot buys nothing (so B is predicted worse than A on cost), a chain or a two-unit refinement loses at every granularity, and the host check queue is what caps fine granularity. F2a then made the plan a value with one home and made the round's log name the plan it was given, which the runner's byte-identical copy of the default made worth doing on its own; the decision that the arms get their own research-side driver, with the 42 couplings measured behind it, is [recorded here](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md). F2b then built that driver (`evals/ooo-execution/plan-driver.ts`, one slot at a time, with the ordered legal set taken from `BoardAdmission.candidates()` rather than re-derived), and **building it measured that the C arm has no mechanism**: a run can hold exactly one claim, so `slots: 4` runs with `slotsUsed: 1` and names the refusal instead of reporting a time. The operator's decision on that finding is [recorded here](../decisions/implemented/2026-09-18-declared-slot-budget.md) — a run declares its **slot budget** (`slots`, default 1), the ordered legal set is cut to it after ordering, and the status query reports the same cut — and it has landed (row `F2b-slot` below, with the measurement kept in "What F2b measured"). Before it, E landed: `src/integration/task-advisers.ts` is the port a source may speak to and the rules it cannot break, `selectableTasks` is the legal set the shared rules already decided (with `nextTask` as its head), and `BoardAdmission` takes optional advisers whose absence is the rule policy - nine obligations proven offline with nine registered mutants, no gate changed and no HA or MGR implementation wired. Before it, D14 gave the drivers a daemon to be clients of and made that client boundary pessimistic (a bounded, named call, and no client calling the endpoint it serves). +Counts at this revision, computed from the rows below rather than from memory: **A** 5 proven, 0 partly, 0 owed (5 rows); **B** 9 proven, 0 partly, 0 owed (9 rows, B7 nothing to fail); **C** 4 proven, 0 partly, 0 owed (4 rows); **D** 13 proven, 0 partly, 0 owed, 1 not applicable (14 rows; D10's subject, `runCycle`, was retired by [the retirement decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); **E** 9 proven as a seam with offline proofs and no source wired (9 rows); **G** 7 proven, 0 partly, 0 owed (7 rows). **F** complete except the fusion slice: its offline layer, the plan's single home, the arms' driver and its slot budget, the two task families and the real-model pilot have all landed and are recorded below, the last one in [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md) with its own sample-size caveat; **F4** (execution fusion) has landed its offline half (legality, accounting, the driver's session policy) **and** the live mechanism it needed - `createPiSessionRunner` in the extension, `piSessionWorker` + `--session-runner` in the driver - and the paid D arm is measured: fused `unitsPerSession: 2` against the `1` control, 3 reps each, one session of two units against two sessions of one, quality parity, ~1.9 s faster per run and token-neutral on that plan shape - a two-rep trend whose own spreads (1.2 s each) exceeded the 1.9 s gap, so it does not price the session-startup term the cost model had left `unmeasured`; on the four-unit fine plan the same cap experiment measures 3.2-4.2 s of wall clock per avoided session with the saving five times its within-cell spread, while its token columns settle nothing at three reps ([the D arm](../experiments/execution/ooo-arms-pilot-2026-09-18.md#d-arm-fusion-added-2026-09-19)). The A, B and C cells were then run on the plain path with the same fixture, worker, envelope and parent check (coarse one rep: 9 726 ms / 9 018 tokens; fine at one slot two reps: 21 885 and 23 273 ms, median 22 579; the same spec at two slots two reps: 18 565 and 19 336 ms, median 18 950, and cheaper in money than the one-slot pair), which also made the chain surface's own price visible: declaring fusion at a bound of one costs 3 607 ms and 15 412 tokens more than the plain path for the same plan ([the A-D cells](../experiments/execution/archive/ooo-arms-2026-09-19/README.md)). Their reports record `inputTokens`/`outputTokens`/`cacheRead`/`cost` per unit and the commit they ran from, so a comparison can name its instrument instead of arguing about it. **F5** (speculation lifecycle) landed offline in the same shared module: the design's `assumptions=[{predicateId, version, expected}]` as a declaration, the first experiment's bounds (exactly one pending fact, no speculative successor, no irreversible operation) refused by name, and the three outcomes read from authoritative evidence - true publishes, false discards the candidate and closes its branch session, unknown waits, because a missing reading is not permission to publish - with 9 cases and 5 named mutants. The **E arm** itself (budgeted speculation against a real model) has since run twice, on 2026-09-19, through the instrument it needed - one that decides the guessed fact and re-runs the real path under a new ticket - with no gain claimed: [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md#e-arm-bounded-speculation-first-run-2026-09-19) and [the archive README](../experiments/execution/archive/ooo-arms-2026-09-19/README.md). This sentence is what this paragraph used to over-claim the other way: it read as though the arm programme were one run short, while the measurement phase had closed. The counts sentence above is what this line used to over-claim: it read "**F** complete" while the F4 row said otherwise. F1 built the advisory cost model the design orders before any paid call (`evals/ooo-execution/cost-model.ts`, derived plan graphs, self-checks, no quality term) and recorded its sweep in [the experiment record](../experiments/execution/ooo-cost-model-2026-09-17.md); the sweep says one slot buys nothing (so B is predicted worse than A on cost), a chain or a two-unit refinement loses at every granularity, and the host check queue is what caps fine granularity. F2a then made the plan a value with one home and made the round's log name the plan it was given, which the runner's byte-identical copy of the default made worth doing on its own; the decision that the arms get their own research-side driver, with the 42 couplings measured behind it, is [recorded here](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md). F2b then built that driver (`evals/ooo-execution/plan-driver.ts`, one slot at a time, with the ordered legal set taken from `BoardAdmission.candidates()` rather than re-derived), and **building it measured that the C arm has no mechanism**: a run can hold exactly one claim, so `slots: 4` runs with `slotsUsed: 1` and names the refusal instead of reporting a time. The operator's decision on that finding is [recorded here](../decisions/implemented/2026-09-18-declared-slot-budget.md) — a run declares its **slot budget** (`slots`, default 1), the ordered legal set is cut to it after ordering, and the status query reports the same cut — and it has landed (row `F2b-slot` below, with the measurement kept in "What F2b measured"). Before it, E landed: `src/integration/task-advisers.ts` is the port a source may speak to and the rules it cannot break, `selectableTasks` is the legal set the shared rules already decided (with `nextTask` as its head), and `BoardAdmission` takes optional advisers whose absence is the rule policy - nine obligations proven offline with nine registered mutants, no gate changed and no HA or MGR implementation wired. Before it, D14 gave the drivers a daemon to be clients of and made that client boundary pessimistic (a bounded, named call, and no client calling the endpoint it serves). Verification commands, run in the worktree that holds this branch, with the values they returned at this revision (re-run them rather than trusting the numbers; the harness writes no log file). Readings taken at @@ -19,7 +19,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run test:product` -> 1457 pass, 0 fail, exit 0. **A row that used to sit here said "one full run first reported a single failure under parallel load, then passed 1433/1433 on re-run; recorded as flaky, not fixed" - that label was wrong, and it hid a product defect.** The failure was `demoteMemory: demotes LTG memory to STG`, and it was a clock boundary: a memory written with `valid_from` a moment _after_ the reading connection's `strftime('now')` read as not current (measured 2 of 3000 write-then-read rounds, stamp `…38.468Z` against `now` `…38.467Z`). Fixed by a named grace in `src/core/store/clock.ts`, pinned by `tests/core/store/current-value-window.test.ts` (6 cases) and 4 named mutants, decided in [the clock-grace record](../decisions/implemented/2026-09-18-clock-grace-window.md), recorded as [post-mortem 0004](../postmortem/0004-flaky-was-a-clock-boundary.md). The count moved 1446 -> 1457 with the fixed window and the cases added since - `npm run mutation:teeth` -> 136 of 136 caught by the named test, 20 of 20 targets restored byte-identically, exit 0. Run on 2026-09-19 as four lanes, one sweep per tree, using `git worktree add --detach` on the same commit for three of them: a sweep is sequential _within_ a tree because its mutants substitute into the same file, and parallel across trees, where each lane also gets the isolation property that no lane's suites can read another lane's mutant. Lanes: 42 of 42 (`ooo-board`, `task-coordinator`), 41 of 41 (`base`, `ooo-execution`'s 17, `task-semantics-interleavings`), 42 of 42 (thirteen small targets) and 11 of 11 (`plan-driver`) - the last serialised into its own lane because its suite has a 25 s case and does real candidate verification (~92 s per run, against ~2 s for the cheap suites). **That lane's cost has since changed**: the arms' checks are data checks now, so at `c01d3fe` its clean run is 5.5 s and its five mutants are caught by the case each names in 0.8-1.2 s on the same target; the numbers in this paragraph describe the 2026-09-19 instrument. Three things the run itself taught, all fixed and pinned afterwards: the lock's `live` flag was never written on substitution (a multi-hunk edit failed as a whole and only the restore half was reapplied), so the field lied about a running sweep; `NODE_TEST_CONTEXT` inherited when a sweep is started from inside a `node --test` process made the nested runner exit 0 having run no test at all, which the harness reported as "the suite passed" and which turned every mutant of that target into a false "not caught"; and the refusal in `agent:verify` fired on `--dry-run` too, which made two of the verifier's own tests fail while a sweep held the tree (a dry run reads the plan and the route config, not the mutated file, so it is exempt now). The lane that reported a clean-run failure (`tools/agent-verify.ts`) had found the last of those three. A sweep is also refused while any lock is present, including one whose owner died, because a killed sweep leaves its mutant in the target (post-mortem 0003). Interruption note: two lane processes were killed by the console that launched them and were relaunched; the JSON each run writes at its end survived even when the buffered stdout summary was lost, so the lane results above were read from those files rather than from stdout. (the retirement pass added the driver's interleaving mutant; was 111 of 111 before this pass: `src/integration/task-semantics-interleavings.ts` gained three budget mutants and `evals/ooo-execution/plan-driver.ts` three for the per-unit checks, the canned worker and the parent composition). How these runs are scheduled (scoped during a change, full before a push, detached with a collected result) is a standing rule of the repository now, in [`skills/repo-development/SKILL.md`](../../skills/repo-development/SKILL.md) with its measured costs in [the decision](../decisions/implemented/2026-09-18-detached-long-checks.md) - `npm run complexity:gate` -> exit 0, 18 methods above 15 unchanged from baseline. It caught the E pass's first version (adding the advisers option pushed `BoardAdmission`'s constructor to 17, so the options check moved into `admissionAdvice()`) rather than the threshold being raised; the slot pass added its option check the same way (`admissionSlots`), and the ordered-set/`publishReady` reads stayed under it. -- `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP this pass and not fixed**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files - it reports 3 diagnostics there, all in files this pass did not touch: one type error each in `board-deliver.ts` and `board-judge.ts`, plus `tests/integration/ooo-evidence-drivers.test.ts:115` (a `{ pid: 0 }` fallback passed where a `ServerState` is required), and the 15 duplicate-key ones it used to report in `evals/ooo-execution/cycle.test.ts` went with that file when the round was retired. Recorded rather than repaired: the two remaining are outside this slice +- `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. - `node --experimental-strip-types --test --test-concurrency=4 tests/integration/ooo-ordinary-failure.test.ts tests/integration/ooo-managed-fence.test.ts tests/integration/ooo-read-paths-agree.test.ts tests/integration/ooo-round-query.test.ts tests/integration/ooo-task-tables.test.ts` -> 5, 3, 1, 2 and 4 pass, 0 fail, exit 0 How a row earns `proven`: it names a test that fails when the code satisfying it is broken. Where diff --git a/tests/integration/ooo-evidence-drivers.test.ts b/tests/integration/ooo-evidence-drivers.test.ts index e18cb8cb..39dae7b1 100644 --- a/tests/integration/ooo-evidence-drivers.test.ts +++ b/tests/integration/ooo-evidence-drivers.test.ts @@ -112,7 +112,11 @@ async function startHost( // Ask rather than kill, so the host's release path is what runs: a host that died holding // its lease would leave the next host on this store with a lease it cannot take. try { - await httpCall(readServerState(serverStatePath(storePath)) ?? { pid: 0 }, "shutdown"); + // The state file names no server when it is gone, and the value passed here says exactly that: + // a pid of zero and no start time. The call is expected to fail, and the catch below is the + // path a dead host takes. + const state = readServerState(serverStatePath(storePath)) ?? { pid: 0, startedAt: "" }; + await httpCall(state, "shutdown"); } catch { // Already gone; the wait below still applies. } From 996e33973805f6c31f8675a22693795074fb1c69 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 20:57:12 +0800 Subject: [PATCH 19/32] The shared answer says why, not just which `selection` now returns its own reasoning beside the legal set: every gate it applies names itself, and `unitLegality` reports the ordered legal set with a cause list per unit. `deriveStatus` carries the same causes, so the status query and dependency release keep sharing one predicate and now share its explanation. The causes are computed inside `selection`, where the gates are, rather than beside it: a second copy of a gate's condition is how an answer starts disagreeing with the decision it explains. A plan-level gate is attached to the units it holds back, so an empty legal set is never a bare empty set. Two cases pin it: one names each gate in turn, one walks every combination of the flags a plan's facts can carry (64 shapes, two budgets) and refuses to accept a silent refusal. The mutant target grew two teeth for the new rule - `a-refused-unit-is-silent` and `a-stale-input-is-not-named` - and the two selection gates whose lines moved were re-pointed, including the budget mutant's named case, which now names the case that actually distinguishes it. 24 of 24 mutants are caught and both files were restored byte-identically. This is the shared half of the legality proposal's step 2. The board action that lets a caller ask the question through a surface it already has is still to come, so nothing here claims the answer is readable yet from outside the process. --- src/integration/ooo-execution.ts | 112 ++++++++++++++++--- src/integration/task-semantics.ts | 24 ++++- tests/integration/ooo-advisers.test.ts | 143 ++++++++++++++++++++++++- tools/mutation-teeth.ts | 32 +++++- 4 files changed, 287 insertions(+), 24 deletions(-) diff --git a/src/integration/ooo-execution.ts b/src/integration/ooo-execution.ts index ffbb573e..48e1d563 100644 --- a/src/integration/ooo-execution.ts +++ b/src/integration/ooo-execution.ts @@ -83,6 +83,29 @@ export interface DispatchTask { cancelled?: boolean; } +/** Why a unit of the plan is not in the legal set. Every cause names one of the gates `selection` + * applies, so the answer a caller asks for is the rule's own reasoning rather than a second reading + * of the rules. The three plan-level causes are attached to the units they hold back. */ +export type DispatchCause = + | "claimed" + | "cancelled" + | "stale-input" + | "external-wait" + | "effect-not-startable" + | "dependency-not-accepted" + | "delivered-not-accepted" + | "budget-spent" + | "several-waits-pending" + | "earlier-unit-blocked"; + +/** One unit as the rule reads it: on offer or not, and why not. `reasons` is empty exactly for a unit + * that is legal or already accepted, so no refusal is silent. */ +export interface UnitLegality { + id: string; + legal: boolean; + reasons: readonly DispatchCause[]; +} + /** Input order and declarations belong to the coordinator, never the worker. * This checks eligibility, not whether arbitrary worker code is actually safe. */ /** @@ -178,12 +201,24 @@ function acceptedDependency( return task.dependencies.every((dependency) => acceptedDependency(byId, dependency, path)); } +/** The effects the rule hands out. Named once: `ready` and the cause that reports a unit held back + * for its effect read the same list, so an effect cannot be startable in one and not in the other. */ +const STARTABLE_EFFECTS = ["read-only", "isolated-artifact"]; + +/** The rule's own reading: the ordered legal set, the budget it left, and one cause list per unit it + * does not have on offer. */ +interface Selection { + legal: readonly string[]; + room: number; + causes: ReadonlyMap; +} + /** The one implementation both readings share, so the budget cannot come to mean two things: `legal` is - * the ordered candidate set, and `room` is how much of it the run's remaining budget pays for. */ -function selection( - plan: readonly DispatchTask[], - slots: number, -): { legal: readonly string[]; room: number } { + * the ordered candidate set, and `room` is how much of it the run's remaining budget pays for. + * + * The causes are computed here rather than beside the rule: a cause names a gate below, and a second + * copy of a gate's condition is how an answer starts disagreeing with the decision it explains. */ +function selection(plan: readonly DispatchTask[], slots: number): Selection { checkedSlots(slots); const byId = taskIndex(plan); const valid = (id: string, visiting = new Set()) => @@ -193,7 +228,7 @@ function selection( current(task) && !task.cancelled && !waiting(task) && - ["read-only", "isolated-artifact"].includes(task.effect) && + STARTABLE_EFFECTS.includes(task.effect) && task.dependencies.every((id) => valid(id)); // A task holding bytes nobody can claim is not a task to select, and it is not a reason to // select nothing either: it is dropped from the plan, which is what the coordinator's reopen() @@ -201,17 +236,68 @@ function selection( const selectable = plan.filter((task) => task.accepted || !task.delivered); const pending = selectable.filter((task) => !valid(task.id)); const room = slots - claimedInFlight(pending); - const none = { legal: [], room: 0 } as const; - // ponytail: scan the bounded experiment plan; no learned priorities or preemption. - if (room < 1 || pending.filter(waiting).length > 1) return none; + const spent = room < 1; + const severalWaits = pending.filter(waiting).length > 1; const first = pending[0]; - if (!first) return none; + // A unit's own gates, in the order `ready` applies them. A cause names the gate; it never restates + // the condition, so the answer cannot come to mean something the rule does not. + const ownCauses = (task: DispatchTask): DispatchCause[] => { + const causes: DispatchCause[] = []; + if (!selectable.includes(task)) causes.push("delivered-not-accepted"); + if (task.claimed) causes.push("claimed"); + if (task.cancelled) causes.push("cancelled"); + if (!current(task)) causes.push("stale-input"); + if (waiting(task)) causes.push("external-wait"); + if (!STARTABLE_EFFECTS.includes(task.effect)) causes.push("effect-not-startable"); + if (task.dependencies.some((id) => !valid(id))) causes.push("dependency-not-accepted"); + return causes; + }; + // The answer: what is on offer, and for every other unit the gates it failed - or, when it is a + // plan-level gate that holds it, that gate's name. An accepted unit owes nothing and stays empty. + const answer = (legal: readonly string[], legalRoom: number, held?: DispatchCause): Selection => { + const offered = new Set(legal); + const causes = new Map(); + for (const task of plan) { + if (offered.has(task.id) || task.accepted) continue; + const own = ownCauses(task); + causes.set(task.id, own.length > 0 ? own : held ? [held] : []); + } + return { legal, room: legalRoom, causes }; + }; + // ponytail: scan the bounded experiment plan; no learned priorities or preemption. + if (spent || severalWaits) return answer([], 0, spent ? "budget-spent" : "several-waits-pending"); + if (!first) return answer([], 0); const ids = (tasks: readonly DispatchTask[]) => tasks.filter((task) => !task.claimed && ready(task)).map((task) => task.id); - if (ready(first)) return { legal: ids(pending), room }; + if (ready(first)) return answer(ids(pending), room); // A stale/missing input or undeclared dependency is not an external wait license. - if (!current(first) || !waiting(first)) return none; - return { legal: ids(pending.slice(1)), room }; + if (!current(first) || !waiting(first)) return answer([], 0, "earlier-unit-blocked"); + return answer(ids(pending.slice(1)), room); +} + +/** + * The legality answer: the ordered legal set, the run's remaining claim budget, and a reading per unit + * in plan order. It is a query - it writes nothing, claims nothing and wakes nobody - and it calls + * `selection`, so a caller cannot read a unit as legal here while the shared rule refuses it there. + * + * The cut to the budget is `startableTasks`'; this answer reports the whole legal set and the room + * left, so a caller with its own order still cuts the same way instead of re-deriving what is pending. + */ +export function unitLegality( + plan: readonly DispatchTask[], + slots = 1, +): { legal: readonly string[]; room: number; units: readonly UnitLegality[] } { + const { legal, room, causes } = selection(plan, slots); + const offered = new Set(legal); + return { + legal, + room, + units: plan.map((task) => ({ + id: task.id, + legal: offered.has(task.id), + reasons: causes.get(task.id) ?? [], + })), + }; } /** What fusion legality needs and a `DispatchTask` does not carry: which executor may run a unit and diff --git a/src/integration/task-semantics.ts b/src/integration/task-semantics.ts index 290eff9d..2c08bbe0 100644 --- a/src/integration/task-semantics.ts +++ b/src/integration/task-semantics.ts @@ -16,7 +16,12 @@ * (execution limits re-expressed as wall clock), or an unknown requirement * kind gets a refusal naming the task and the field. */ -import { startableTasks, type DispatchTask } from "./ooo-execution.ts"; +import { + startableTasks, + unitLegality, + type DispatchCause, + type DispatchTask, +} from "./ooo-execution.ts"; import { workDigestOf } from "./work-identity.ts"; import { CONCLUSION_KINDS, @@ -526,24 +531,35 @@ export function dispatchTasks(units: readonly TaskUnit[], facts: RecordedFacts): /** Status and dependency release share the derived dispatch state and the existing eligibility rule; * neither gets its own notion of "accepted". `ready` is what a run may start now: one task at the - * default budget, and the run's remaining claim budget's worth when it declared more. */ + * default budget, and the run's remaining claim budget's worth when it declared more. + * + * A blocked unit carries the rule's own causes beside the declared dependencies it is waiting on + * (`waitingFor`, direct; the causes are transitive where the rule is), so "why is this not offered" + * is answered by the same `selection` that decided it rather than by a reader's guess. */ export function deriveStatus( units: readonly TaskUnit[], facts: RecordedFacts, slots = 1, ): { ready: readonly string[]; - blocked: readonly { id: string; waitingFor: readonly string[] }[]; + blocked: readonly { + id: string; + waitingFor: readonly string[]; + reasons: readonly DispatchCause[]; + }[]; accepted: readonly string[]; } { const accepted = units.filter((unit) => isAccepted(unit, facts)).map((unit) => unit.id); const acceptedIds = new Set(accepted); - const ready = startableTasks(dispatchTasks(units, facts), slots); + const tasks = dispatchTasks(units, facts); + const ready = startableTasks(tasks, slots); + const legality = new Map(unitLegality(tasks, slots).units.map((unit) => [unit.id, unit.reasons])); const blocked = units .filter((unit) => !isAccepted(unit, facts)) .map((unit) => ({ id: unit.id, waitingFor: unit.inputs.dependencies.filter((dependency) => !acceptedIds.has(dependency)), + reasons: legality.get(unit.id) ?? [], })) .filter((entry) => entry.waitingFor.length > 0 || !ready.includes(entry.id)); return { ready, blocked, accepted }; diff --git a/tests/integration/ooo-advisers.test.ts b/tests/integration/ooo-advisers.test.ts index d1853ba6..16c8ab7e 100644 --- a/tests/integration/ooo-advisers.test.ts +++ b/tests/integration/ooo-advisers.test.ts @@ -9,6 +9,10 @@ * The two cases at the end are the wiring rather than the rules: an admission with no source answers * exactly as the rule does, and one with a source can only reorder what the shared rules already * made legal. + * + * The last two are the answer itself: what a caller that asks "what is legal, and why not" gets back. + * The causes are the selector's own gates, so the verdict and the reasons cannot come apart - and the + * second case is the invariant that makes that worth anything: no refusal is silent. */ import assert from "node:assert/strict"; import { mkdtempSync } from "node:fs"; @@ -21,7 +25,7 @@ import { type PatchTaskSpec, type ProbePlan, } from "../../src/integration/ooo-board.ts"; -import { nextTask, selectableTasks } from "../../src/integration/ooo-execution.ts"; +import { nextTask, selectableTasks, unitLegality } from "../../src/integration/ooo-execution.ts"; import { orderCandidates, revalidateSuggestion, @@ -30,7 +34,11 @@ import { type SuggestionProvenance, type SuggestionSource, } from "../../src/integration/task-advisers.ts"; -import { compileTaskUnits, dispatchTasks } from "../../src/integration/task-semantics.ts"; +import { + compileTaskUnits, + dispatchTasks, + type RecordedFacts, +} from "../../src/integration/task-semantics.ts"; const SCOPE = { sessionId: "session-1", @@ -325,3 +333,134 @@ test("an admission with no source answers exactly as the rule does, and a source assert.equal(advised.refuseStaleRanking("B", ["A", "B"]), null); assert.match(advised.refuseStaleRanking("B", ["A"])!.reason, /no longer a legal candidate/u); }); + +test("the answer names the gate a unit was refused through, and the gates are the rule's own", () => { + const REV = "input-v1"; + const units = compileTaskUnits({ + plan: [ + ["A", REV, [], "read-only", null, null], + ["B", REV, [], "read-only", null, null], + ], + specs: {}, + }); + assert.ok(units.legal, "the fixture plan is legal"); + const read = (facts: RecordedFacts, slots = 2) => + unitLegality(dispatchTasks(units.units, facts), slots); + const reasons = (answer: ReturnType, id: string) => + answer.units.find((unit) => unit.id === id)!.reasons; + + // A claim in flight is named on its own unit and spends the budget for the others: a plan-level + // cause is attached to the unit it holds back, so an empty legal set is never a bare empty set. + const claimed = read({ claimed: ["A"] }, 1); + assert.deepEqual(claimed.legal, [], "the declared budget is spent"); + assert.deepEqual(reasons(claimed, "A"), ["claimed"]); + assert.deepEqual(reasons(claimed, "B"), ["budget-spent"]); + + // Each of these is one gate of the rule, and the cause is that gate's name. + assert.deepEqual(reasons(read({ cancellations: ["A"] }), "A"), ["cancelled"]); + assert.deepEqual(reasons(read({ revisions: { A: "other" } }), "A"), ["stale-input"]); + assert.deepEqual(reasons(read({ artifacts: { A: "artifact-a" } }), "A"), [ + "delivered-not-accepted", + ]); + + const writeEffect = compileTaskUnits({ + plan: [["A", REV, [], "workspace-write", null, null]], + specs: {}, + }); + assert.ok(writeEffect.legal, "the fixture plan is legal"); + assert.deepEqual(unitLegality(dispatchTasks(writeEffect.units, {}), 1).units[0]!.reasons, [ + "effect-not-startable", + ]); + + const dependent = compileTaskUnits({ + plan: [ + ["B", REV, [], "read-only", null, null], + ["A", REV, ["B"], "read-only", null, null], + ], + specs: {}, + }); + assert.ok(dependent.legal, "the fixture plan is legal"); + const unmet = unitLegality(dispatchTasks(dependent.units, {}), 2); + assert.deepEqual(unmet.legal, ["B"], "a dependency nobody accepted is not a candidate"); + assert.deepEqual(unmet.units.find((unit) => unit.id === "A")!.reasons, [ + "dependency-not-accepted", + ]); + + // A declared wait is licensed: the unit that declares it is named, and what follows stays legal. + const waited = compileTaskUnits({ + plan: [ + ["W", REV, [], "read-only", "fact-ready", null], + ["X", REV, [], "read-only", null, null], + ], + specs: {}, + }); + assert.ok(waited.legal, "the fixture plan is legal"); + const waiting = unitLegality(dispatchTasks(waited.units, {}), 2); + assert.deepEqual(waiting.legal, ["X"], "a declared wait licenses what follows it"); + assert.deepEqual(waiting.units.find((unit) => unit.id === "W")!.reasons, ["external-wait"]); + + // Two pending waits are not one licence, and the unit that has no cause of its own says so. + const twoWaits = compileTaskUnits({ + plan: [ + ["W1", REV, [], "read-only", "one", null], + ["W2", REV, [], "read-only", "two", null], + ["Y", REV, [], "read-only", null, null], + ], + specs: {}, + }); + assert.ok(twoWaits.legal, "the fixture plan is legal"); + const several = unitLegality(dispatchTasks(twoWaits.units, {}), 2); + assert.deepEqual(several.legal, [], "two pending waits are not a licence"); + assert.deepEqual(several.units.find((unit) => unit.id === "Y")!.reasons, [ + "several-waits-pending", + ]); + + // The head rule holds the rest back: a stale head is not skipped, and the units behind it say so. + const behind = read({ revisions: { A: "other" } }); + assert.deepEqual(behind.legal, []); + assert.deepEqual(behind.units.find((unit) => unit.id === "A")!.reasons, ["stale-input"]); + assert.deepEqual(behind.units.find((unit) => unit.id === "B")!.reasons, ["earlier-unit-blocked"]); +}); + +test("no refusal in the answer is silent, over the flags a plan's facts can carry", () => { + const REV = "input-v1"; + const units = compileTaskUnits({ + plan: [ + ["A", REV, [], "read-only", null, null], + ["B", REV, [], "isolated-artifact", null, null], + ], + specs: {}, + }); + assert.ok(units.legal, "the fixture plan is legal"); + let refusals = 0; + for (const claimed of [false, true]) + for (const cancelled of [false, true]) + for (const stale of [false, true]) + for (const delivered of [false, true]) + for (const slots of [1, 2]) { + const facts: RecordedFacts = { + claimed: claimed ? ["A"] : [], + cancellations: cancelled ? ["A"] : [], + revisions: stale ? { A: "other" } : {}, + artifacts: delivered ? { A: "artifact-a" } : {}, + sourceRevisions: { A: REV }, + }; + const tasks = dispatchTasks(units.units, facts); + const answer = unitLegality(tasks, slots); + assert.deepEqual( + answer.legal, + selectableTasks(tasks, slots), + "the answer's verdict is the rule's, whatever the flags", + ); + for (const unit of answer.units) { + const accepted = tasks.find((task) => task.id === unit.id)!.accepted; + if (unit.legal || accepted) continue; + refusals += 1; + assert.ok( + unit.reasons.length > 0, + `${unit.id} is refused with no reason (slots ${String(slots)})`, + ); + } + } + assert.ok(refusals > 0, "the family refuses something, so the invariant was exercised"); +}); diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index e37493ef..7586971b 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -549,9 +549,10 @@ const TARGETS: readonly Target[] = [ // decide whether selection is open: this is the rule the C arm's slot count was once absent from. name: "a-live-claim-does-not-block-selection", ast: { within: "selection" }, - from: " if (room < 1 || pending.filter(waiting).length > 1) return none;", - to: " if (pending.filter(waiting).length > 1) return none;", - expect: "with no fusion point the plan falls back to its declared order", + from: ' if (spent || severalWaits) return answer([], 0, spent ? "budget-spent" : "several-waits-pending");', + to: ' if (severalWaits) return answer([], 0, spent ? "budget-spent" : "several-waits-pending");', + expect: + "the answer names the gate a unit was refused through, and the gates are the rule's own", }, { // A task someone is working is not on offer, whatever the budget. Without this, a run with @@ -585,10 +586,31 @@ const TARGETS: readonly Target[] = [ // likely to bypass by accident, so it has its own tooth. name: "a-head-blocked-by-a-stale-input-is-skipped", ast: { within: "selection" }, - from: " if (!current(first) || !waiting(first)) return none;", - to: " if (false) return none;", + from: ' if (!current(first) || !waiting(first)) return answer([], 0, "earlier-unit-blocked");', + to: ' if (false) return answer([], 0, "earlier-unit-blocked");', expect: "the round's own answer is the shared rule's answer, not an ordering's", }, + { + // The answer a caller asks for carries the rule's own reasoning, so a refused unit that has + // no cause of its own is still named by the plan-level gate that held it back. This mutant + // silences every refusal at once: an empty legal set with no reasons is exactly the answer + // the design says a caller cannot be given. + name: "a-refused-unit-is-silent", + ast: { within: "selection" }, + from: " causes.set(task.id, own.length > 0 ? own : held ? [held] : []);", + to: " causes.set(task.id, []);", + expect: "no refusal in the answer is silent, over the flags a plan's facts can carry", + }, + { + // A cause names the gate; it never restates the condition. Dropping one leaves a refusal + // whose reason the rule can no longer give, even when another gate would still be true. + name: "a-stale-input-is-not-named", + ast: { within: "selection" }, + from: ' if (!current(task)) causes.push("stale-input");', + to: ' if (false) causes.push("stale-input");', + expect: + "the answer names the gate a unit was refused through, and the gates are the rule's own", + }, { // `nextTask` returns the head of what it decided, so an ordering cannot disagree with the // shared rule about what may be selected: dropping the head moves both. From f11fdb40e992c5386d8ff3e1062dc682d9f453d1 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 21:09:56 +0800 Subject: [PATCH 20/32] The board's read answers why, not only which The port gains the query the legality proposal's step 2 asks for: `legality()` returns the ordered legal set, the room the run has left, and a named cause per unit the rules do not have on offer. It calls the same `selection`, so a caller cannot read a unit as legal here while the rule refuses it there, and `candidates()` - the port's other read - is now that answer's `legal` on the store's board, which is what keeps the two from drifting apart. All three boards answer it: the store's, the probe board's (whose projection moved into one private method so `ordered()` and the answer cannot disagree), and the dispatch test's stub, which hands its own fields to the shared reading rather than keeping a second set of rules. The asker is the process that owns the run's workspace: a plan compiles from the caller's files and the daemon holds only their frozen paths, so the answer cannot come from the store alone. The product's case asks twice on an unchanged store, asserts the two answers and the recorded facts are identical, that the answer names nobody (the store's own refusal names the holder; this read does not), that a claim moves a unit to the named cause while the claim itself stays the store's, and that the answer follows the facts it reads rather than a cached plan. Docs land with it: the two-faces document's read-face bullets, the proposal's step 2 marked landed with step 3 named as the place it is still not true, and the umbrella's gap 2 narrowed to the asker that has no workspace. `ooo-board`'s and `ooo-execution`'s mutants were re-run - 23 of 23 and 24 of 24 caught, both files restored byte-identically - and the product suite is 1532 of 1532. --- ...2026-09-20-the-program-answers-legality.md | 12 +++ ...9-20-the-program-answers-legality.zh-CN.md | 4 + docs/design/mechanism-in-the-middle.md | 9 ++ docs/design/mechanism-in-the-middle.zh-CN.md | 2 + .../design/protocol-governed-collaboration.md | 13 ++- .../protocol-governed-collaboration.zh-CN.md | 4 +- src/integration/ooo-board.ts | 42 +++++++-- src/integration/ooo-dispatch.ts | 15 ++-- src/integration/ooo-execution.ts | 13 ++- src/integration/ooo-runner.ts | 16 +++- tests/integration/ooo-dispatch.test.ts | 36 ++++++-- tests/integration/ooo-store-run-board.test.ts | 86 +++++++++++++++++++ 12 files changed, 221 insertions(+), 31 deletions(-) diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md index bf8fb31f..d2e83c83 100644 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md @@ -60,6 +60,18 @@ asking a question consume a slot. Each step is verifiable on its own; none of them requires a driver, a wake, or a new tool. +**Step 2, landed 2026-09-23.** The board port's read is `legality()`: the ordered legal set, the room the +run has left, and a named cause for every unit the rules do not have on offer, computed inside the rules' +own `selection` so a cause names a gate rather than restating it - and `candidates()` is that answer's +`legal`, so the two cannot disagree. The asker is the process that owns the run's workspace: a plan +compiles from the caller's files while the daemon holds only their frozen paths, so the shared tool +contract's wording is unchanged, because the read it describes is the same read. + +**Step 3 is not built**, which is the one place this proposal is not yet true: the online move that +prefers continuing a session still lives as the shared planner's own ordering, and the run's recorded +`policy` is a name the board checks rather than a switch the planner reads, so the answer is computed +under an implicit constraint. + ## Alternatives considered - **Let the program select the next unit** (today's `next()` used as policy). Rejected: a bound is not a diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md index b0366c34..b13c58cd 100644 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md +++ b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md @@ -31,6 +31,10 @@ 每一步都能单独验证;没有一步需要驱动器、唤醒或新工具。 +**第 2 步已于 2026-09-23 落地。** 黑板端口的读就是 `legality()`:有序合法集合、运行剩下的预算,以及每个不在候选里的单元各自一条**具名原因**;原因在规则自己的 `selection` 里算出来,所以一条原因点的是某道闸门而不是把它复述一遍——而 `candidates()` 就是这份答案的 `legal`,两者不可能不一致。提问者是拥有这次运行工作区的那个进程:计划是从调用方的文件编译出来的,而守护进程只持有那些文件的冻结路径,所以共享 tool contract 的措辞没变,因为它描述的仍是同一次读。 + +**第 3 步尚未建。** 这正是这份提案目前唯一还不成立的地方:偏好延续会话的那个在线行动仍然活在共享规划器自己的排序里,而运行里记下的 `policy` 是一个**黑板会校验的名字**、不是规划器会读的开关,所以这份答案是在一条隐式约束下算出来的。 + ## 考虑过的替代方案 - **让程序挑下一个单元**(把现在的 `next()` 当策略用)。拒绝:一个上限不是决定——它不在合法后继里做选择;而挑人的程序就是调度器,命名决策已把"调度器"记为这个模型刻意没有变成的词。 diff --git a/docs/design/mechanism-in-the-middle.md b/docs/design/mechanism-in-the-middle.md index b447d4dd..64731bf6 100644 --- a/docs/design/mechanism-in-the-middle.md +++ b/docs/design/mechanism-in-the-middle.md @@ -98,6 +98,15 @@ against an imagined second shape, and the imagined one is the one that gets buil rules read, and it carries no legality rule of its own. Three implementations satisfy `DispatchBoard` now - the probe board, the dispatch test's stub, and this one - and only the projection differs between them, which is the abstraction working rather than a claim that it does. +- **The port's read answers why, not only which.** `legality()` is that read: the ordered legal set, the + room the run has left, and a named cause for every unit the rules do not have on offer. The causes are + computed inside the rules' own `selection`, so a cause names a gate rather than restating its condition, + and a plan-level gate is attached to the units it holds back; `candidates()` is that answer's `legal`, so + the two cannot disagree. The answer names nobody - who claimed, delivered or judged stays a board fact - + and asking changes nothing: no entry, no claim, no budget, no wake. +- **Who can be asked.** The process that owns the run's workspace, through its board. A plan is compiled + from the caller's files and the daemon holds only their frozen paths, so this read cannot be answered from + the store alone; that is a property of where the plan lives, not a surface the daemon is missing. - **The projection has one contract that is easy to get wrong.** The shared acceptance rule binds a verdict to the artifact a run carries: it compares the verdict's digest against the artifact value, so a verdict about a different artifact cannot pass as acceptance. Reporting the store's own deliverable hash in that diff --git a/docs/design/mechanism-in-the-middle.zh-CN.md b/docs/design/mechanism-in-the-middle.zh-CN.md index 2ba177f2..7ad39207 100644 --- a/docs/design/mechanism-in-the-middle.zh-CN.md +++ b/docs/design/mechanism-in-the-middle.zh-CN.md @@ -52,6 +52,8 @@ - **工作的身份有了一个家。** `src/integration/work-identity.ts` 拥有判词所绑的那条规则:sha256、hex、全长,作用在调用方**冻结**过的那些字节或 JSON 文本上。此前有四个模块写过同两行——其中两处是逐字相同的私有 `digestOf(value)`——而其中两个调用方把自己的截断也塞了进去,于是十二字符的报告身份和十六字符的分支身份读起来像是规则的一部分,而不是调用方的选择。规范化仍归调用方,模块明说这一点:JSON 的键顺序是调用方自己选出来的,所以**冻结**一个补丁才是让它的摘要稳定的那一步。两个具名 mutant 让这条约定可被检查——换成另一种编码,以及一个不消化 JSON 的 JSON 变体。那些**故意**不同(搜索索引的 base64url 内容哈希、会话 id、协议可见的 `sha256:` 前缀)的站点保留自己的约定,模块把它们点名,以免后来的一次清扫把三条不同的规则合成一条。 - **中间的缝没有落地。** 检查(c)说明了为什么。 - **端口有了第三个实现,而且是产品自己的那个。** `src/integration/ooo-runner.ts` 是 store 自己的黑板:它把运行的冻结表与黑板条目**投影**成共享规则要读的那些事实,自己不带任何合法性规则。现在有三个实现满足 `DispatchBoard`——探针黑板、dispatch 测试里的桩、和这一个——它们之间**只有投影不同**,这是抽象在起作用,而不是关于它的说法。 +- **端口的读回答“为什么”,而不只是“是哪一个”。** `legality()` 就是这次读:有序合法集合、运行剩下的预算,以及每个不在候选里的单元各自一条**具名原因**。原因在规则自己的 `selection` 里算出来,所以一条原因点的是某道闸门,而不是复述它的条件;计划级闸门被挂在被它挡住的那些单元上。`candidates()` 就是这份答案的 `legal`,两者不可能不一致。答案里不出现任何人——谁认领、谁交付、谁判定都仍是黑板事实——而且**问不改变任何东西**:不产生条目、不认领、不花预算、不唤醒。 +- **谁可以去问。** 拥有这次运行工作区的那个进程,通过它自己的黑板去问。计划是从调用方的文件编译出来的,而守护进程只持有那些文件的冻结路径,所以单靠 store 答不了这次读;这是“计划住在哪里”的性质,不是守护进程缺了一个面。 - **投影里有一条容易搞错的契约。** 共享的验收规则把判词**绑在运行携带的产物上**:它拿判词的摘要与产物**值**比较,所以关于另一个产物的判词不能冒充验收。把 store 自己的交付物哈希填进那个字段,会让每个已验收单元都读成未验收,而症状是**无声的**——下一个单元永不释放,且任何地方都不报错。两种身份都是正当的,但那个字段只属于产物值。 - **端口的动词落在产品生命周期的哪里。** 产品的一次运行把交付物存在**单元被认领时的那条条目**上,所以端口的"投一个结果"就是那次交付、"提交"就是对它的判定:一个单元一条条目,不是两条,因为第二条结果条目会是一条没人读的记录。单元是从循环刚拿到的**认领**解出来的,而不是从正文解出来的——正文的形状归它的主人,而认领是黑板自己的事实。 diff --git a/docs/design/protocol-governed-collaboration.md b/docs/design/protocol-governed-collaboration.md index fcc43069..e9211bd2 100644 --- a/docs/design/protocol-governed-collaboration.md +++ b/docs/design/protocol-governed-collaboration.md @@ -40,8 +40,12 @@ rows above are the ones a given piece of work cannot do without. promotion, the dispatch loop), the responsibility does not. Distributed systems call the first one a **commit boundary** - the single point that makes a discussed arrangement binding - and the second an **orphan reaper / restart responsibility**. -2. **Legality has no readable answer.** A caller can claim and be refused, but cannot ask what is legal - and why. That is **admission control** with a **read-only precondition check**. +2. **Legality's readable answer now exists; asking it from a store alone does not.** The board port + answers `legality()` - the ordered legal set, the room left, and a named cause per unit - and that read + changes no state, so admission control with a read-only precondition check now has a home. What is still + missing is the asker without a workspace: a plan is compiled from the caller's files, so a process that + holds only the store cannot compute the answer. The concepts are unchanged; the gap now narrows to where + the plan lives. 3. **Constraints have no declaration.** Repair-first currently lives as policy inside shared planning code. Making preferences named constraints a plan enables is **declarative policy**, and the requirement that one plan plus one set of facts yields one answer is **deterministic replay**. @@ -85,7 +89,8 @@ rows above are the ones a given piece of work cannot do without. ## What would make this inventory complete -Rows 9 and 10 are the interfaces the legality proposal covers: the readable answer and the declared -constraint set. Rows 1, 4, 5, 6 and 7 are gaps with no owner yet. Until each gap either gets a clause or +Row 9's readable answer landed on 2026-09-23 as the board port's own read - see [The program answers +legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md), whose step 3 is the row 10 it +still owes. Rows 1, 4, 5, 6 and 7 are gaps with no owner yet. Until each gap either gets a clause or is written down as deliberately absent, "complete" for this umbrella means the parts that have owners, not the parts a task needs. diff --git a/docs/design/protocol-governed-collaboration.zh-CN.md b/docs/design/protocol-governed-collaboration.zh-CN.md index 480178eb..e63479ea 100644 --- a/docs/design/protocol-governed-collaboration.zh-CN.md +++ b/docs/design/protocol-governed-collaboration.zh-CN.md @@ -31,7 +31,7 @@ ## 空缺,以及可以为它们借用的概念 1. **收编者没有协议条款。** 谁有资格把一次讨论收编成一次 run,以及租约失效后谁负责把活重新推出去,都没有写。机制是有的(`expiry`、串行晋升、派发循环),责任没有。分布式里把前者叫做**提交边界**——让一份商定的安排生效的那个唯一点;把后者叫做**孤儿回收/重启责任**。 -2. **合法性没有可读的答案。** 调用者可以去认领然后被拒,但无法问"现在什么合法、为什么"。这是**准入控制**加一次**只读的前置条件检查**。 +2. **合法性现在有了可读的答案;但只握着 store 的问不出来。** 黑板端口回答 `legality()`——有序合法集合、剩下的预算、以及每个单元各自一条**具名原因**——而这次读不改变任何状态,所以**准入控制**加一次**只读的前置条件检查**现在有了家。仍然缺的是**没有工作区的提问者**:计划是从调用方的文件编译出来的,所以只持有 store 的进程算不出这份答案。概念没变,缺口现在窄到“计划住在哪里”这一点上。 3. **约束没有声明方式。** 修复优先目前活在共享规划代码里当策略。把偏好变成计划可启用的具名约束,是**声明式策略**;而"同一计划加同一组事实只产出一个答案"这个要求,是**确定性重放**。 4. **读的保证没有写下来。** 黑板读基于游标、可能落后;没有一处说明读者是否可以假设单调读或读己所写。概念是**单调读**、**读己所写**,以及跨同一个 agent 自己一串会话的**因果一致性**。 5. **分歧没有仲裁者。** 两个 agent 对下一步做什么意见不合时,只能争夺一次完成(`veto`)或抢认领。对分歧本身没有条款。最近的概念是**乐观并发加冲突检测**,以及——因为仲裁者本来就存在——干脆的**单写者决定**。竞价式分配早就被拒绝了([黑板治理](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md))。 @@ -52,4 +52,4 @@ ## 这份清单要怎样才算完整 -第 9、10 行就是合法性提案覆盖的两个接口:可读的答案与可声明的约束集合。第 1、4、5、6、7 行是还没有主的空缺。在每个空缺或者拿到条款、或者被明确写成"刻意没有"之前,这个大类的"完整"指的只是那些有归属的部分,而不是一件任务真正需要的那几样。 +第 9 行的可读答案已于 2026-09-23 以黑板端口自己的读落地——见[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md),它还欠的第 10 行正是那份提案的第 3 步。第 1、4、5、6、7 行是还没有主的空缺。在每个空缺或者拿到条款、或者被明确写成"刻意没有"之前,这个大类的"完整"指的只是那些有归属的部分,而不是一件任务真正需要的那几样。 diff --git a/src/integration/ooo-board.ts b/src/integration/ooo-board.ts index fdbb10c6..8ab8fbd0 100644 --- a/src/integration/ooo-board.ts +++ b/src/integration/ooo-board.ts @@ -38,6 +38,8 @@ import { remainingSlots, selectableTasks, snapshotAnswer, + unitLegality, + type LegalityAnswer, type SnapshotWork, } from "./ooo-execution.ts"; import { compileTaskUnits, dispatchTasks, type RecordedFacts } from "./task-semantics.ts"; @@ -927,13 +929,27 @@ export class BoardAdmission extends NmgStore { return order.slice(0, room); } - private ordered(): { order: readonly string[]; room: number } { - // A cancelled round selects nothing: the successor is not "the next task", it is - // the explicit terminal decision the caller asked for. - if (this.cancelled() !== null) { - this.lastAdvice = null; - return { order: [], room: 0 }; - } + /** + * The legality answer: the ordered legal set, the room left, and a cause for every unit the rules + * do not have on offer. A read - it claims nothing, publishes nothing and wakes nobody - and it + * comes from the same projection `candidates()` and `next()` do, so a caller can ask why a unit is + * not on offer instead of inferring it from the refusals it collects. + */ + legality(): LegalityAnswer { + return unitLegality(this.projection().dispatch, this.slots); + } + + /** + * The shared reading of this board's plan and its recorded facts, computed in one place so the + * ordered set, the budget and the legality answer cannot come to different conclusions about the + * same plan. A plan the compiler refuses is refused here by name, not scheduled on a hand-rolled + * reading of the same rows. + */ + private projection(): { + compiled: ReturnType; + dispatch: ReturnType; + rows: readonly Row[]; + } { const rows = this.db .prepare("SELECT * FROM ooo_probe_task_view WHERE run_id=? ORDER BY position") .all(this.runId) as unknown as Row[]; @@ -950,7 +966,17 @@ export class BoardAdmission extends NmgStore { `first is ${String(first?.task)}/${String(first?.field)}: ${String(first?.reason)}`, ); } - const dispatch = dispatchTasks(compiled.units, this.recordedFacts(rows)); + return { compiled, dispatch: dispatchTasks(compiled.units, this.recordedFacts(rows)), rows }; + } + + private ordered(): { order: readonly string[]; room: number } { + // A cancelled round selects nothing: the successor is not "the next task", it is + // the explicit terminal decision the caller asked for. + if (this.cancelled() !== null) { + this.lastAdvice = null; + return { order: [], room: 0 }; + } + const { compiled, rows, dispatch } = this.projection(); const legal = selectableTasks(dispatch, this.slots); const room = remainingSlots(dispatch, this.slots); if (this.advisers.length === 0) { diff --git a/src/integration/ooo-dispatch.ts b/src/integration/ooo-dispatch.ts index 3f1749d7..158c797b 100644 --- a/src/integration/ooo-dispatch.ts +++ b/src/integration/ooo-dispatch.ts @@ -11,10 +11,10 @@ * * Two things are deliberately *not* decided here. * - * The board is a port (`DispatchBoard`) of operations - candidates, accepted, claim, put, submit - - * and not a class. The product's board satisfies it by being one; the research instrument's board - * satisfies it structurally too. Naming either here is what would make the shared layer depend on - * one of them, and a shared loop that only the instrument can call is not shared. + * The board is a port (`DispatchBoard`) of operations - candidates, legality, accepted, claim, put, + * submit - and not a class. The product's board satisfies it by being one; the research instrument's + * board satisfies it structurally too. Naming either here is what would make the shared layer depend + * on one of them, and a shared loop that only the instrument can call is not shared. * * The worker is a port for the same reason, plus one of its own: only the caller knows how a * candidate is produced. The arms supply a live model or a recorded one; a host supplies the session @@ -28,7 +28,7 @@ */ import type { PatchWork, FrozenPatchWork } from "./ooo-patch.ts"; import { preparePatchWork } from "./ooo-patch.ts"; -import type { SessionPlan } from "./ooo-execution.ts"; +import type { LegalityAnswer, SessionPlan } from "./ooo-execution.ts"; import { nextSessionMove } from "./ooo-fusion-plan.ts"; import { decideSessionMove } from "./ooo-session-facts.ts"; import type { NmgStore } from "../core/store.ts"; @@ -104,6 +104,11 @@ export interface DispatchBoard { readonly now: number; /** Which of the plan's tasks may be claimed right now, in the order they should run. */ candidates(): readonly string[]; + /** The same reading, and why: the ordered legal set, the room the run has left, and a cause for + * every unit the rules do not have on offer. A read - it takes no claim, puts no entry, spends no + * budget and wakes nobody - so a caller may ask it as often as it likes instead of inferring the + * rules from refusals. `candidates()` is this answer's `legal`. */ + legality(): LegalityAnswer; /** The accepted artifact per task id: the rule every dependency and every parent check reads. */ accepted(): Readonly>; /** Take one unit. The board re-checks legality here, so a stale answer becomes a refusal. */ diff --git a/src/integration/ooo-execution.ts b/src/integration/ooo-execution.ts index 48e1d563..34d764b0 100644 --- a/src/integration/ooo-execution.ts +++ b/src/integration/ooo-execution.ts @@ -275,6 +275,14 @@ function selection(plan: readonly DispatchTask[], slots: number): Selection { return answer(ids(pending.slice(1)), room); } +/** The answer a caller asks for: the ordered legal set, the budget it left, and one reading per unit + * in plan order. Plain data on purpose - a caller may print two of them and diff them. */ +export interface LegalityAnswer { + legal: readonly string[]; + room: number; + units: readonly UnitLegality[]; +} + /** * The legality answer: the ordered legal set, the run's remaining claim budget, and a reading per unit * in plan order. It is a query - it writes nothing, claims nothing and wakes nobody - and it calls @@ -283,10 +291,7 @@ function selection(plan: readonly DispatchTask[], slots: number): Selection { * The cut to the budget is `startableTasks`'; this answer reports the whole legal set and the room * left, so a caller with its own order still cuts the same way instead of re-deriving what is pending. */ -export function unitLegality( - plan: readonly DispatchTask[], - slots = 1, -): { legal: readonly string[]; room: number; units: readonly UnitLegality[] } { +export function unitLegality(plan: readonly DispatchTask[], slots = 1): LegalityAnswer { const { legal, room, causes } = selection(plan, slots); const offered = new Set(legal); return { diff --git a/src/integration/ooo-runner.ts b/src/integration/ooo-runner.ts index 69e8d68d..58091744 100644 --- a/src/integration/ooo-runner.ts +++ b/src/integration/ooo-runner.ts @@ -3,7 +3,7 @@ * * This is the third implementation of `DispatchBoard` - the probe board and the dispatch test's stub * are the other two - and it is deliberately the thinnest of the three. No legality rule lives here: - * the plan is compiled by `compileTaskUnits`, the legal set is `dispatchTasks` and `selectableTasks`, + * the plan is compiled by `compileTaskUnits`, the legal set is `dispatchTasks` and `unitLegality`, * cancellations are read through the coordinator's own reader, and the verdict belongs to the * caller's acceptance. What this module adds is a **projection** from the product's run tables and * board entries into the facts those pure functions read, plus the six board operations. A rule that @@ -27,7 +27,7 @@ import { type PatchSubmission, type PatchWork, } from "./ooo-patch.ts"; -import { selectableTasks } from "./ooo-execution.ts"; +import { unitLegality, type LegalityAnswer } from "./ooo-execution.ts"; import { compileTaskUnits, dispatchTasks, type RecordedFacts } from "./task-semantics.ts"; import { coordinatedBoardWrite, taskCancellation, taskRunStatus } from "./task-coordinator.ts"; @@ -87,8 +87,18 @@ export class StoreRunBoard implements DispatchBoard { /** What the loop may claim right now, in the shared rule's own order. */ candidates(): readonly string[] { + return this.legality().legal; + } + + /** + * The same answer, and why: the ordered legal set, the room the run has left, and a cause for every + * unit the rules do not have on offer. It reads the run's recorded facts under the shared rules and + * writes nothing, so a caller can ask what is legal instead of inferring it from the refusals it + * collects. `candidates()` is that answer's `legal`, which is why it is computed from it. + */ + legality(): LegalityAnswer { const { compiled, facts } = this.#projection(); - return selectableTasks(dispatchTasks(compiled.units, facts), this.#options.slots); + return unitLegality(dispatchTasks(compiled.units, facts), this.#options.slots); } /** The accepted artifact per task id: the same map the rules read, so the two cannot disagree. */ diff --git a/tests/integration/ooo-dispatch.test.ts b/tests/integration/ooo-dispatch.test.ts index 04c6cdfd..70c3d835 100644 --- a/tests/integration/ooo-dispatch.test.ts +++ b/tests/integration/ooo-dispatch.test.ts @@ -27,11 +27,20 @@ import { recordedSessionMoves, } from "../../src/integration/ooo-session-facts.ts"; import { registerRun } from "../../src/integration/task-coordinator.ts"; -import type { DispatchTask, SessionPlan } from "../../src/integration/ooo-execution.ts"; +import { + unitLegality, + type DispatchTask, + type LegalityAnswer, + type SessionPlan, +} from "../../src/integration/ooo-execution.ts"; import type { PlanSession, PlanWorker } from "../../src/integration/ooo-dispatch.ts"; const REMOVE_TEMP_TREE = { recursive: true, force: true, maxRetries: 5, retryDelay: 100 }; +/** The revision this stub's plan declares. The stub models the loop's order of operations, not drift, + * so every task's inputs are the ones the plan declared and `current()` always holds. */ +const REVISION = "stub-v1"; + type BoardOptions = { /** Which legal unit to refuse, as the board does when another claim holds the handoff. */ refuse?: (taskId: string) => boolean; @@ -68,11 +77,28 @@ class StubBoard implements DispatchBoard { this.#options = options; } + /** + * The answer the port asks for, from this board's own fields: its order, dependencies, claims and + * accepted artifacts are exactly what the shared rules read, so the stub hands them to the shared + * reading instead of growing a second set of rules. No budget is declared here - the stub's job is + * the loop's order of operations - so the reading runs with a budget that never cuts. + */ + legality(): LegalityAnswer { + const tasks = this.#order.map((id) => ({ + id, + effect: "isolated-artifact", + sourceVersion: REVISION, + observedVersion: REVISION, + dependencies: [...(this.#dependencies[id] ?? [])], + accepted: this.#accepted[id] !== undefined, + claimed: this.#inFlight.has(id), + externalReady: false, + })); + return unitLegality(tasks, Number.MAX_SAFE_INTEGER); + } + candidates(): readonly string[] { - return this.#order.filter((id) => { - if (this.#accepted[id] !== undefined || this.#inFlight.has(id)) return false; - return (this.#dependencies[id] ?? []).every((needed) => this.#accepted[needed] !== undefined); - }); + return this.legality().legal; } accepted(): Readonly> { diff --git a/tests/integration/ooo-store-run-board.test.ts b/tests/integration/ooo-store-run-board.test.ts index 21a6a6da..5116abb3 100644 --- a/tests/integration/ooo-store-run-board.test.ts +++ b/tests/integration/ooo-store-run-board.test.ts @@ -237,6 +237,92 @@ test("the deliverer cannot judge its own delivery, so a run cannot self-accept", }); }); +test("the board answers what is legal and why, and asking changes nothing", async () => { + await withStore(async (store) => { + openRun(store); + const board = new StoreRunBoard(store, { + runId: RUN, + channel: CHANNEL, + slots: 1, + workspace, + acceptance, + }); + const entryId = bindingEntry(store, "P"); + // What the answer reads, so the read can be shown to leave it as it found it. + const recorded = () => ({ + facts: store.taskRunFacts(RUN).length, + entry: store.getTaskBoardEntryById(CHANNEL, entryId), + }); + + const before = recorded(); + const first = board.legality(); + assert.deepEqual( + board.legality(), + first, + "asking twice on an unchanged store returns the same answer", + ); + assert.deepEqual(recorded(), before, "the read wrote no run fact and moved no entry"); + + // Per unit, in plan order, with the reason on the unit that is not on offer: the dependency gate + // is the shared rule's, read through this board's projection. + assert.deepEqual(first.legal, ["P"], "the dependent unit is not offered early"); + assert.equal(first.room, 1, "the declared budget is reported, and asking does not spend it"); + assert.deepEqual(first.units, [ + { id: "P", legal: true, reasons: [] }, + { id: "T", legal: false, reasons: ["dependency-not-accepted"] }, + ]); + + // The answer names nobody. The store's own refusal names the holder; this read does not, because + // who claimed, who delivered and who judged are board facts rather than answer fields. + const printed = JSON.stringify(first); + for (const name of ["worker", acceptance.agentId, "runner", "repair-first"]) + assert.ok(!printed.includes(name), `the answer does not name ${name}`); + + // The structural rule in the answer's own terms: a claim takes the unit out of the set and says + // which gate did it, while the claim itself stays the store's. + board.claim("P", "worker:P"); + const claimed = board.legality(); + assert.deepEqual(claimed.legal, [], "the single declared slot is spent"); + assert.equal(claimed.room, 0, "the claim spent the declared budget"); + assert.deepEqual( + claimed.units.map((unit) => [unit.id, unit.reasons]), + [ + ["P", ["claimed"]], + ["T", ["dependency-not-accepted"]], + ], + ); + assert.equal(store.getTaskBoardEntryById(CHANNEL, entryId)?.claimedBy, "worker:P"); + assert.throws(() => board.claim("P", "worker:other"), /already claimed by worker:P/u); + + // Bytes are delivered and nobody has decided the verdict yet, so the unit is still its + // claimant's: this board's facts carry acceptance rather than delivery (the probe's carry the + // other), which is why the gate named here is the claim. In-doubt work across runs is the + // umbrella's own open gap, not something this read may paper over. + const ticket = board.claim("P", "worker:P"); + const frozen = preparePatchWork({ ...ticket.patch!, attempt: ticket.attempt }); + coordinatedBoardWrite(store, { + runId: RUN, + entryId, + verb: "deliver", + actorId: "worker:P", + apply: () => + store.deliverTaskBoardEntry({ + taskId: CHANNEL, + entryId, + agentId: "worker:P", + digest: "d", + ref: artifactFor("P", frozen), + }), + }); + const delivered = board.legality(); + assert.deepEqual(delivered.legal, [], "a delivered unit nobody judged is still not on offer"); + assert.deepEqual(delivered.units.find((unit) => unit.id === "P")!.reasons, ["claimed"]); + assert.deepEqual(delivered.units.find((unit) => unit.id === "T")!.reasons, [ + "dependency-not-accepted", + ]); + }); +}); + /** The entry a task's binding holds, read from the run's own record. */ function bindingEntry(store: NmgStore, taskId: string): string { const fact = store From d7519e9366e80ab11583b4cc74d5c15bd9aa8723 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 23:19:55 +0800 Subject: [PATCH 21/32] The continuation is a constraint the plan declares Repair-first was the shared planner's own default, so the meaning of a plan depended on which planner read it: the online move that prefers continuing a session lived in `nextSessionMove`, and the run's recorded `policy` was a name the probe board compared rather than a switch anything read. The legality record's third step is that a constraint the protocol names and a plan enables, and it is what makes "a constraint that is not enabled has no effect" true rather than intended. The protocol's side is a closed list. `PLAN_CONSTRAINTS` holds the names that exist - `repair-first` is the only one so far - and `enabledConstraints` refuses a name outside it by name, because reading an unknown name as "nothing was asked for" is exactly the decoration the rule forbids. A plan that enables nothing is a setting rather than a fallback: it runs one unit per session, which is the baseline the arms' control cells measure. The plan's side is a declaration. `SessionPlan` gained `constraints`, the move is computed under it, and `optimisticPlan` copies it, so the offline projection prices the same plan the online half decides with. Every caller that wants a fused run now declares it: the arm driver's spec carries the declared set, and `validateSpecFile` refuses a spec that declares a bound above one without enabling the constraint - such a run would be the control arm while its file said fusion, which is the one confusion that driver must not create. The trial spec generator writes it, and three test fixtures carry it. What is demonstrable is a different move at a session boundary under an enabled constraint, not a different legal set: repair-first orders a session, it does not gate membership. The record's criterion said legal sets and was corrected rather than quietly restated, and the record moves from proposed to implemented in the same commit as the code, with the umbrella's rows 9 and 10 updated. Evidence: two new mutants on `src/integration/ooo-fusion-plan.ts` (2 of 2 caught by the named cases: the continuation not declared, an unknown constraint ignored), the four re-run targets 37 of 37, `npm run test:product` 1534 of 1534, `npm run verify:static` exit 0, `npm run docs:check` 288 files with 0 errors. --- ...2026-09-19-fusion-planning-repair-first.md | 9 +- ...9-19-fusion-planning-repair-first.zh-CN.md | 2 +- ...2026-09-20-the-program-answers-legality.md | 124 ++++++++++++++++++ ...9-20-the-program-answers-legality.zh-CN.md | 55 ++++++++ ...2026-09-20-the-program-answers-legality.md | 120 ----------------- ...9-20-the-program-answers-legality.zh-CN.md | 64 --------- .../2026-09-21-mechanism-not-policy.md | 6 +- .../2026-09-21-mechanism-not-policy.zh-CN.md | 4 +- .../2026-09-21-the-frame-and-its-storage.md | 2 +- ...6-09-21-the-frame-and-its-storage.zh-CN.md | 2 +- docs/design/mechanism-in-the-middle.md | 4 +- docs/design/mechanism-in-the-middle.zh-CN.md | 2 +- docs/design/ooo-fusion-planning.md | 16 ++- .../design/protocol-governed-collaboration.md | 24 ++-- .../protocol-governed-collaboration.zh-CN.md | 8 +- .../ooo-execution/make-fusion-trial-specs.mjs | 2 +- evals/ooo-execution/plan-driver.test.ts | 61 ++++++++- evals/ooo-execution/plan-driver.ts | 27 +++- src/integration/ooo-execution.ts | 5 + src/integration/ooo-fusion-plan.ts | 40 +++++- tests/integration/ooo-dispatch.test.ts | 3 + tests/integration/ooo-fusion-plan.test.ts | 42 +++++- tests/integration/ooo-session-facts.test.ts | 3 + tools/mutation-teeth.ts | 25 ++++ 24 files changed, 420 insertions(+), 230 deletions(-) create mode 100644 docs/decisions/implemented/2026-09-20-the-program-answers-legality.md create mode 100644 docs/decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md delete mode 100644 docs/decisions/proposed/2026-09-20-the-program-answers-legality.md delete mode 100644 docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md diff --git a/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.md b/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.md index c903cd0f..c378d7c4 100644 --- a/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.md +++ b/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.md @@ -32,9 +32,12 @@ Split fusion planning into two clocks, and write down neither as a plan. - **Online** is one move about the current session: _admit_ the next legal successor or _close_ the session, naming the condition that closed it. A fused session is irreversible, so no move may rewrite it. -- **Repair-first**: the default is to continue; only a declared change (rejected verdict, cancellation, - an unmet dependency, a declared external wait that is not ready) may end a session early. Repairing - keeps every commitment intact and decides only about what has not run. +- **Repair-first**: continuing is a constraint the plan declares rather than this design's default - a + plan that enables nothing runs one unit per session ([the program answers + legality](2026-09-20-the-program-answers-legality.md), step 3, landed 2026-09-23) - and, once enabled, + only a declared change (rejected verdict, cancellation, an unmet dependency, a declared external wait + that is not ready) may end a session early. Repairing keeps every commitment intact and decides only + about what has not run. - **Baseline**: a move is not revisited while its facts hold, and the same plan plus the same facts yield the same move, ties broken by plan order. Without determinism two runs of one plan are not comparable, which is what the arms need. diff --git a/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.zh-CN.md b/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.zh-CN.md index dc79b16b..fc45d671 100644 --- a/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.zh-CN.md +++ b/docs/decisions/implemented/2026-09-19-fusion-planning-repair-first.zh-CN.md @@ -16,7 +16,7 @@ 把融合规划拆成两个时钟,并且**两边都不写成一个计划**。 - **在线**只做关于当前会话的**一个动作**:*接纳*下一个合法后继,或*关闭*会话并说出关闭它的那个条件。已融合的会话**不可撤销**,所以任何动作都不能改写它。 -- **修复优先**:默认是继续;只有**声明的变化**(裁决被拒、取消、依赖未达成、声明的外部等待未就绪)才能提前结束会话。修复保留全部已作出的承诺,只决定**尚未运行**的部分。 +- **修复优先**:延续是一条**由计划声明的约束**,不是本设计的默认——什么都没启用的计划就是每个会话一个单元([程序只回答合法性](2026-09-20-the-program-answers-legality.zh-CN.md) 的第 3 步,2026-09-23 落地);启用后,只有**声明的变化**(裁决被拒、取消、依赖未达成、声明的外部等待未就绪)才能提前结束会话。修复保留全部已作出的承诺,只决定**尚未运行**的部分。 - **基线**:事实不变时,动作不重议;同样的计划加同样的事实得到同样的动作,平局按计划顺序打破。没有确定性,同一计划的两次运行就不可比——而两臂正需要可比。 - **成本模型在线只读**:离线拟合、按运行冻结。查询优化器也是在边界上拿已收集的统计**重新规划**,而**绝不在查询中途重新收集统计**。 - **离线**从一个**乐观投影**计算上限,并把投影喂给**同一个** `sharedSessionLegal`,让五条条件只有一个归属:下界来自**最小链覆盖**(Dilworth:等于最大反链,用 `单元数 − 最大匹配`(关系传递闭包上的二分图匹配)算出),可行界来自**按声明上限的贪心 list scheduling**。 diff --git a/docs/decisions/implemented/2026-09-20-the-program-answers-legality.md b/docs/decisions/implemented/2026-09-20-the-program-answers-legality.md new file mode 100644 index 00000000..5fb72b92 --- /dev/null +++ b/docs/decisions/implemented/2026-09-20-the-program-answers-legality.md @@ -0,0 +1,124 @@ +# The program answers legality, and nothing else + +[中文](2026-09-20-the-program-answers-legality.zh-CN.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [Mechanism, not policy](../proposed/2026-09-21-mechanism-not-policy.md), [Board governance and capability addressing](2026-09-06-board-governance-addressing.md), [Name the collaboration protocol and its task-unit sub-protocol](2026-09-20-name-the-collaboration-protocol.md), [Task unit semantics](../../design/task-unit-semantics.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md) + +## Problem + +The primitives have names, and the shared layer can already compute the legal set: an ordered legal set, +that set cut to the declared slot budget, and a refusal reason per unit when a claim is checked. What no +document states is **whose decision each step is**. The cost of that gap is observable: every caller +that wants to drive a run has to redraw the boundary for itself, so both recorded failure modes come +back - a program that picks the person for the agents, which is the scheduler the naming decision +records as the word this model deliberately did not become, and agents that cannot tell what the program +guarantees, so they ask for what they actually need in the only terms available: more slots, stranger +ordering. + +[The obligations ledger](../../design/task-unit-semantics-obligations.md) says where the boundary must +not move - no second task body, no dedicated tool, no dedicated channel - but it does not say what the +program's answer is. Neither does the board's correctness line - reviewable finalize, authentic +content, scope isolation, capability addressing ([board governance and capability +addressing](2026-09-06-board-governance-addressing.md)) - which governs how an entry is treated once it +exists, not who decides what happens next. + +## Decision + +Three owners, three kinds of statement. The program's share is exactly one: **it answers legality**. + +- **The protocol owns** what the primitives are (already named; this record does not restate them), the + definition and criteria of legality, which constraints exist **by name**, and the two structural + rules that are what "legal" means here and therefore cannot be negotiated: a claim (lease plus + attempt fence) is the only arbiter of who holds a unit, and a deliverer never judges its own + delivery. +- **The accompanying program owns** computing the ordered legal set from the declared plan and current + facts, cutting it to the declared slot budget, and stating for each unit why it is legal or not - + reusing the reasons it already refuses claims with. It chooses nobody, adopts nothing, judges + nothing, and wakes nobody. +- **Agents own the arrangement**: the plan's content, the order they agree on, who takes which unit, + who adopts a run, who judges. All of it lands as board facts, so the arrangement is readable and + auditable instead of being implied by a program's choice. + +**Constraints are named by the protocol and enabled by the plan.** A preference such as repair-first is +neither a program policy nor merely advice: it is a named constraint that a plan - or the adopter +declaring it - enables, and the answer is computed under the enabled set. A constraint that is not +enabled must be genuinely absent from the answer, or the declaration is decoration. The names the +protocol defines are a closed list (`PLAN_CONSTRAINTS`), and a plan naming one outside it is refused by +name, because reading an unknown name as "nothing was asked for" is exactly the decoration this rule +forbids. Enabling nothing is a setting rather than a fallback: a plan that enables nothing runs one unit +per session. + +**Legality is asked, not published.** It is a function of current facts, so a query returns the ordered +legal set with a reason per unit, changing no state. A published handoff on the board is a different +thing: by the board's own protocol it occupies a serial slot. Fusing the two into one action would make +asking a question consume a slot. + +## Implementation state + +**Step 1** is this record: the division of labour and the standing of constraints, with no product +surface. + +**Step 2, landed 2026-09-23.** The board port's read is `legality()`: the ordered legal set, the room the +run has left, and a named cause for every unit the rules do not have on offer. The causes are computed +inside the rules' own `selection`, so a cause names a gate rather than restating it, and `candidates()` +is that answer's `legal`, so the port's two read verbs cannot disagree. The invariant is mechanical: no +refusal is silent over 64 flag combinations times two slots. The asker is the process that owns the +run's workspace - a plan compiles from the caller's files while the daemon holds only their frozen paths +- so the shared tool contract's wording is unchanged, because the read it describes is the same read. + +**Step 3, landed 2026-09-23.** The continuation is a declared constraint, not the planner's default: +`src/integration/ooo-fusion-plan.ts` owns `PLAN_CONSTRAINTS`, reads the plan's declaration before it +continues a session, and refuses a name outside the list. Every caller that wants a fused run now +declares it: the arm driver's spec carries the declared set and is refused when a spec declares a bound +above one without enabling the constraint - such a run would be the control arm while its file said +fusion. + +What is demonstrable is a different **move** at a session boundary under an enabled constraint, not a +different legal set: repair-first orders a session, it does not gate membership, and a name that changed +which units are legal would be a different kind of answer. The criterion was corrected here rather than +quietly restated. + +## Alternatives considered + +- **Let the program select the next unit** (today's `next()` used as policy). Rejected: a bound is not a + decision - it does not choose among legal successors - and a program that picks the person is a + scheduler again, which the naming decision records as the word this model deliberately did not become. +- **Let the program store state only, with agents asserting their own legality.** Rejected: "legal" + would then be whatever the loudest caller says, and the two structural rules (one holder per claim, + no self-judging) would lose their home. +- **Demote preferences to advice.** Rejected: the arms compare runs on the premise that the same plan + plus the same facts yield the same answer, and advice that may be ignored removes exactly that. +- **Publish the legal set as a board offering** (make the query a handoff). Rejected: it makes a + question consume a serial slot, and withdrawing an offer is the board protocol's business. +- **Give the arrangement its own tool or channel.** Rejected: the ledger forbids it; the primitives must + take effect through an ordinary handoff. +- **Keep repair-first as hidden program policy.** Rejected: same class as a program that picks the + person. +- **Read an absent declaration as the old behaviour** (default a plan to repair-first). Rejected: that + is the same implicit policy with a default attached, and it leaves "disabled" with no way to be said. + +## Consequences + +- **Every fused run declares the constraint.** Three test fixtures, the arm driver's spec and the trial + spec generator carry it now, and the driver refuses a spec that declares a bound it did not enable. + The measurement of any fused arm depends on a spec that declares it, which is the point: two runs are + comparable only if each records which constraints it enabled. +- **A constraint is a protocol act.** The vocabulary is closed and an unknown name is refused by name, + so adding one is a change to the protocol, not a string a caller may invent. +- **What a constraint may change is still unwritten.** This rule orders a session; nothing yet says + whether a constraint may change which units are legal. That is what the umbrella's row 10 gap became + after this landed. +- **The run's recorded `policy` is a label, not the switch.** The probe board compares it with the + policy it expects; the planning decision is read from the plan's declaration. Two homes for one rule + would be the defect this decision exists to avoid, so the recorded name is deliberately not wired to + the move. +- **Reasons are for people and logs.** A caller that parses a reason string would create a second source + of truth; a caller's decision goes through a claim or an explicit field. +- **Asking stays free.** No entry, no claim, no slot, no wake; the store's version is unchanged, and the + same store yields the same answer. A caller that asks and then acts can still race, and the claim + remains the arbiter. +- **Nothing here answers two neighbouring gaps.** A process holding only the store cannot ask (a plan + compiles from the caller's files), and work left in doubt across runs still has no clause; the + umbrella keeps both. diff --git a/docs/decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md b/docs/decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md new file mode 100644 index 00000000..f8812e0b --- /dev/null +++ b/docs/decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md @@ -0,0 +1,55 @@ +# 程序只回答合法性,别的都不管 + +[English](2026-09-20-the-program-answers-legality.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [机制,不是策略](../proposed/2026-09-21-mechanism-not-policy.zh-CN.md)、[黑板治理与能力寻址](2026-09-06-board-governance-addressing.zh-CN.md)、[给协作协议及其任务单元子协议命名](2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md)、[契约的义务](../../design/task-unit-semantics-obligations.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md) + +## 问题 + +元语有了名字,共享层也已经能算出合法集合:有序合法集合、按声明槽位预算切一刀、认领被拒时给出每个单元的理由。但**没有任何一份文档在说“每一步是谁的决定”**。这个缺口有可观察的代价:任何想驱动一次 run 的调用者都得自己重新划一遍边界,于是两种已被记录在案的失败模式都回来了——程序顺手替 agent 挑人(也就是命名决策里记为“这个模型刻意没有变成”的那个词:调度器),以及 agent 不知道程序到底给什么保证,只能用唯一能用的说法去要它真正需要的东西:更多槽位、更怪的顺序。 + +[义务台账](../../design/task-unit-semantics-obligations.md)写明了边界不能往哪边动——不得再有第二个任务体、专用工具、专用频道——但它没有回答程序给出的那个答案是什么。黑板那条正确性线也没有:可复核的终结、内容真实性、作用域隔离、能力寻址([黑板治理与能力寻址](2026-09-06-board-governance-addressing.zh-CN.md))管的是一个条目在存在之后怎么被对待,不是下一步该由谁定。 + +## 决策 + +三个所有者,三类陈述。程序那一份只有一条:**它回答合法性**。 + +- **协议拥有**:元语是什么(已被命名,本文不复述)、合法性的定义与判据、**约束的具名清单**,以及两条结构性规则——它们就是“合法”在这里的含义,因此不可协商:认领(租约 + attempt 围栏)是“谁持有某个单元”的唯一仲裁;交付者不裁决自己的交付。 +- **配套程序拥有**:按声明的计划与当下事实算出有序合法集合,按声明的槽位预算切一刀,并给出每个单元合法或不合法的理由——复用它现在拒绝认领时用的那些理由。它不挑人、不收编、不裁决、不唤醒。 +- **agent 拥有安排本身**:计划的内容、他们商定的顺序、谁拿哪个单元、谁收编一次 run、谁裁决。所有这些都落成黑板上的事实,所以安排是可读、可审的,而不是由程序的选择暗示出来的。 + +**约束由协议具名、由计划启用。** 修复优先这类偏好既不是程序策略,也不只是建议:它是一条具名约束,由计划(或声明它的收编者)启用,答案在启用集合下计算。没被启用的约束必须在答案里真的不生效,否则这个声明就是装饰。协议定义的名单是**闭集**(`PLAN_CONSTRAINTS`),计划写了名单外的名字会被**点名拒绝**——把不认识的名字读成“什么都没要求”,正是这条规则要禁的那种装饰。什么都不启用是一种设置,不是缺省兜底:什么都不启用的计划就是每个会话跑一个单元。 + +**合法性被问,不被发布。** 它是当下事实的函数,所以查询返回有序合法集合,外加每个单元的理由,且不改变任何状态。黑板上发布出去的 handoff 是另一回事:按黑板自己的协议,它占用一个串行槽位。把两个合成一个动作,就会让“问一句”花掉一个槽位。 + +## 落地情况 + +**第 1 步**就是本文:写下分工与约束的地位,不动产品面。 + +**第 2 步于 2026-09-23 落地。** 黑板端口的读就是 `legality()`:有序合法集合、运行剩下的预算,以及每个不在候选里的单元各自一条**具名原因**。原因在规则自己的 `selection` 里算出来,所以一条原因点的是某道闸门,而不是把它复述一遍;而 `candidates()` 就是这份答案的 `legal`,所以端口的两个读动词不可能互相矛盾。这条不变式是机械的:64 种 flag 组合乘以两个槽位,没有一处拒绝是无声的。提问者是拥有这次运行工作区的那个进程——计划是从调用方的文件编译出来的,而守护进程只持有那些文件的冻结路径——所以共享 tool contract 的措辞没变,因为它描述的仍是同一次读。 + +**第 3 步于 2026-09-23 落地。** 延续会话是一条**声明的约束**,不再是规划器自己的默认:`src/integration/ooo-fusion-plan.ts` 拥有 `PLAN_CONSTRAINTS`,在决定是否延续会话之前先读计划的声明,并在名单外点名拒绝。每个想要融合 run 的调用者现在都得声明它:arm 驱动器的 spec 带着这份启用集合,而一个声明了大于一的预算却没启用这条约束的 spec 会被拒绝——那样的 run 会是控制臂,而它的文件写着融合。 + +可以证明的是:在同一条约束启用与不启用下,会话边界上算出的是**不同的移动**,而不是不同的合法集合。修复优先排的是会话顺序,它不管成员资格;一个会改变哪些单元合法的名字,是另一种性质的答案。这条判据在这里被修正,而不是照原样重述。 + +## 考虑过的替代方案 + +- **让程序挑下一个单元**(把现在的 `next()` 当策略用)。拒绝:一个上限不是决定——它不在合法后继里做选择;而挑人的程序就是调度器,命名决策已把“调度器”记为这个模型刻意没有变成的词。 +- **程序只存状态,合法性由 agent 各自声明**。拒绝:那么“合法”就是嗓门最大的调用者说的;两条结构性规则(一个认领只有一个持有者、交付者不能自裁)会失去家。 +- **把偏好降成建议**。拒绝:arm 之间的可比性建立在“同一计划加同样事实给出同一答案”上,而可被无视的建议恰好拿走这一点。 +- **把合法集合作为黑板 offer 发布**(把查询做成 handoff)。拒绝:这样“问一句”要占串行槽位,而且撤回 offer 是黑板协议的事。 +- **为这套安排新开工具或频道**。拒绝:台账已禁;元语必须经普通交接生效。 +- **保留修复优先作为隐藏的程序策略**。拒绝:与“程序挑人”同属一类。 +- **把“没有声明”读成旧行为**(计划没写就按修复优先)。拒绝:那还是同一条隐式策略,只是加了个默认值,而且让“停用”根本没有说法。 + +## 后果 + +- **每个融合 run 都要声明这条约束。** 现在三处测试夹具、arm 驱动器的 spec 和试验 spec 生成器都带着它,而驱动器会拒绝一个声明了预算却没启用它的 spec。任何融合 arm 的测量都依赖声明了它的 spec,这正是要点:只有每次 run 都记录自己启用了哪些约束,两次 run 才可比。 +- **加一条约束是协议行为。** 名单是闭集,名单外的名字被点名拒绝,所以添一条约束是改协议,而不是调用者随手编一个字符串。 +- **约束能改变什么还没写下来。** 这一条只排会话的顺序;目前还没有一处说约束可不可以改变哪些单元合法。这就是这次落地之后,大类清单第 10 行剩下的那个缺口。 +- **run 里记下的 `policy` 是标签,不是开关。** 探针板拿它跟自己期望的策略比对;规划上的决定是从计划的声明里读的。同一条规则有两个家,正是本决策要避免的缺陷,所以那个被记下的名字是刻意不接到这个移动上的。 +- **理由是给人看和进日志的。** 解析理由字符串的调用者会造出第二个事实源;调用者的决定走认领或一个显式字段。 +- **问仍然是免费的。** 不产生条目、认领、槽位或唤醒;store 版本不变,同一个 store 给出同一个答案。“先问后动”仍可能撞车——认领仍是仲裁者。 +- **这里没有回答邻近的两个缺口。** 只握着 store 的进程问不出来(计划是从调用方的文件编译的),跨 run 之后仍悬着的活仍然没有条款;两个都留在那份大类清单里。 diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md deleted file mode 100644 index d2e83c83..00000000 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.md +++ /dev/null @@ -1,120 +0,0 @@ -# The program answers legality, and nothing else - -[中文](2026-09-20-the-program-answers-legality.zh-CN.md) - -**Status:** proposed -**Relates to:** [Mechanism, not policy](2026-09-21-mechanism-not-policy.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Task unit semantics](../../design/task-unit-semantics.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md) - -## Problem - -The primitives have names, and the shared layer can already compute the legal set: an ordered legal set, -that set cut to the declared slot budget, and a refusal reason per unit when a claim is checked. What no -document states is **whose decision each step is**. The cost of that gap is observable: every caller -that wants to drive a run has to redraw the boundary for itself, so both recorded failure modes come -back - a program that picks the person for the agents, which is the scheduler the naming decision -records as the word this model deliberately did not become, and agents that cannot tell what the program -guarantees, so they ask for what they actually need in the only terms available: more slots, stranger -ordering. - -[The obligations ledger](../../design/task-unit-semantics-obligations.md) says where the boundary must -not move - no second task body, no dedicated tool, no dedicated channel - but it does not say what the -program's answer is. Neither does the board's correctness line - reviewable finalize, authentic -content, scope isolation, capability addressing ([board governance and capability -addressing](../implemented/2026-09-06-board-governance-addressing.md)) - which governs how an entry is -treated once it exists, not who decides what happens next. - -## Proposal - -Three owners, three kinds of statement. The program's share is exactly one: **it answers legality**. - -- **The protocol owns** what the primitives are (already named; this record does not restate them), the - definition and criteria of legality, which constraints exist **by name**, and the two structural - rules that are what "legal" means here and therefore cannot be negotiated: a claim (lease plus - attempt fence) is the only arbiter of who holds a unit, and a deliverer never judges its own - delivery. -- **The accompanying program owns** computing the ordered legal set from the declared plan and current - facts, cutting it to the declared slot budget, and stating for each unit why it is legal or not - - reusing the reasons it already refuses claims with. It chooses nobody, adopts nothing, judges - nothing, and wakes nobody. -- **Agents own the arrangement**: the plan's content, the order they agree on, who takes which unit, - who adopts a run, who judges. All of it lands as board facts, so the arrangement is readable and - auditable instead of being implied by a program's choice. - -**Constraints are named by the protocol and enabled by the plan.** A preference such as repair-first is -neither a program policy nor merely advice: it is a named constraint that a plan - or the adopter -declaring it - enables, and the legality answer is computed under the enabled set. A constraint that is -not enabled must be genuinely absent from the answer, or the declaration is decoration. - -**Legality is asked, not published.** It is a function of current facts, so a query returns the ordered -legal set with a reason per unit, changing no state. A published handoff on the board is a different -thing: by the board's own protocol it occupies a serial slot. Fusing the two into one action would make -asking a question consume a slot. - -## Plan - -1. This record: write down the division of labour and the standing of constraints. No product surface. -2. Make legality readable through an existing board action (the action name lives in the shared - tool contract, its description is generated from the prompt source), keeping the read free of state. -3. Promote the preferences that currently live as shared planning policy - repair-first first - into - named constraints a plan declares, which is what makes the "disabled means absent" criterion true. - -Each step is verifiable on its own; none of them requires a driver, a wake, or a new tool. - -**Step 2, landed 2026-09-23.** The board port's read is `legality()`: the ordered legal set, the room the -run has left, and a named cause for every unit the rules do not have on offer, computed inside the rules' -own `selection` so a cause names a gate rather than restating it - and `candidates()` is that answer's -`legal`, so the two cannot disagree. The asker is the process that owns the run's workspace: a plan -compiles from the caller's files while the daemon holds only their frozen paths, so the shared tool -contract's wording is unchanged, because the read it describes is the same read. - -**Step 3 is not built**, which is the one place this proposal is not yet true: the online move that -prefers continuing a session still lives as the shared planner's own ordering, and the run's recorded -`policy` is a name the board checks rather than a switch the planner reads, so the answer is computed -under an implicit constraint. - -## Alternatives considered - -- **Let the program select the next unit** (today's `next()` used as policy). Rejected: a bound is not a - decision - it does not choose among legal successors - and a program that picks the person is a - scheduler again, which the naming decision records as the word this model deliberately did not become. -- **Let the program store state only, with agents asserting their own legality.** Rejected: "legal" - would then be whatever the loudest caller says, and the two structural rules (one holder per claim, - no self-judging) would lose their home. -- **Demote preferences to advice.** Rejected: the arms compare runs on the premise that the same plan - plus the same facts yield the same answer, and advice that may be ignored removes exactly that. -- **Publish the legal set as a board offering** (make the query a handoff). Rejected: it makes a - question consume a serial slot, and withdrawing an offer is the board protocol's business. -- **Give the arrangement its own tool or channel.** Rejected: the ledger forbids it; the primitives must - take effect through an ordinary handoff. -- **Keep repair-first as hidden program policy.** Rejected: same class as a program that picks the - person. - -## Acceptance criteria - -- The answer is per unit, ordered, and carries the reason for each legal or refused unit. -- The answer names no person, no adopter and no judge. -- Asking changes no state: no entry, no claim, no slot, no wake. Asking twice on an unchanged store - returns the same answer, and the store's version is unchanged. -- A constraint that is not enabled has no effect on the answer: enabling and disabling the same named - constraint on one plan yields different legal sets. -- The declared slot budget still cuts the answer, and declaring more than one slot still requires a - declared handoff target. -- The two structural rules hold in the answer's own terms: a second claim on a held unit is refused, - and a deliverer cannot judge. -- Who adopted, who claimed and who judged exist only as board facts; no program field asserts them. -- The ledger's prohibition still holds: no new tool, no dedicated channel, no second task body. - -## Risks - -- With no constraint declared, a wide legal set can read as permission to fan out, recreating the - "no discussion, everything at once" problem somewhere new. Mitigation: the slot budget stays a - declared cap rather than a social one. -- Repair-first currently lives as policy inside the shared planning code, so making it a declared - constraint is a change rather than a rename; until it moves, the answer is computed under an implicit - constraint, which is the one place this proposal is not yet true. -- Reasons become an interface: a caller that parses a reason string creates a second source of truth. - Reasons are for people and logs; any caller decision must go through a claim or an explicit field. -- With agents free to choose order, the arms' equal-instrument premise depends on each run recording - which constraints were enabled; an unrecorded set makes two runs incomparable. -- A query invites polling. Polling is cheap, and a caller that asks and then acts can still race - the - claim remains the arbiter, which is the point. diff --git a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md b/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md deleted file mode 100644 index b13c58cd..00000000 --- a/docs/decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md +++ /dev/null @@ -1,64 +0,0 @@ -# 程序只回答合法性,别的都不管 - -[English](2026-09-20-the-program-answers-legality.md) - -**Status:** proposed -**Relates to:** [机制,不是策略](2026-09-21-mechanism-not-policy.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md)、[契约的义务](../../design/task-unit-semantics-obligations.md) - -## 问题 - -元语有了名字,共享层也已经能算出合法集合:有序合法集合、按声明槽位预算切一刀、认领被拒时给出每个单元的理由。但**没有任何一份文档在说"每一步是谁的决定"**。这个缺口有可观察的代价:任何想驱动一次 run 的调用者都得自己重新划一遍边界,于是两种已被记录在案的失败模式都回来了——程序顺手替 agent 挑人(也就是命名决策里记为"这个模型刻意没有变成"的那个词:调度器),以及 agent 不知道程序到底给什么保证,只能用唯一能用的说法去要它真正需要的东西:更多槽位、更怪的顺序。 - -[义务台账](../../design/task-unit-semantics-obligations.md)写明了边界不能往哪边动——不得再有第二个任务体、专用工具、专用频道——但它没有回答程序给出的那个答案是什么。黑板那条正确性线也没有:可复核的终结、内容真实性、作用域隔离、能力寻址([黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md))管的是一个条目在存在之后怎么被对待,不是下一步该由谁定。 - -## 提案 - -三个所有者,三类陈述。程序那一份只有一条:**它回答合法性**。 - -- **协议拥有**:元语是什么(已被命名,本文不复述)、合法性的定义与判据、**约束的具名清单**(有哪些约束可以启用),以及两条结构性规则——它们就是"合法"在这里的含义,因此不可协商:认领(租约 + attempt 围栏)是"谁持有某个单元"的唯一仲裁;交付者不裁决自己的交付。 -- **配套程序拥有**:按声明的计划与当下事实算出有序合法集合,按声明的槽位预算切一刀,并给出每个单元合法或不合法的理由——复用它现在拒绝认领时用的那些理由。它不挑人、不收编、不裁决、不唤醒。 -- **agent 拥有安排本身**:计划的内容、他们商定的顺序、谁拿哪个单元、谁收编一次 run、谁裁决。所有这些都落成黑板上的事实,所以安排是可读、可审的,而不是由程序的选择暗示出来的。 - -**约束由协议具名、由计划启用。** 修复优先这类偏好既不是程序策略,也不只是建议:它是一条具名约束,由计划(或声明它的收编者)启用,合法性答案在启用集合下计算。没被启用的约束必须在答案里真的不生效,否则这个声明就是装饰。 - -**合法性被问,不被发布。** 它是当下事实的函数,所以查询返回有序合法集合,外加每个单元的理由,且不改变任何状态。黑板上发布出去的 handoff 是另一回事:按黑板自己的协议,它占用一个串行槽位。把两个合成一个动作,就会让"问一句"花掉一个槽位。 - -## 计划 - -1. 本文:写下分工与约束的地位。不动产品面。 -2. 让合法性可读——走已有的黑板动作(动作名在共享 tool contract 里,描述由 prompt 源生成),并保持这次读取不写状态。 -3. 把目前活在共享规划策略里的偏好,首先是修复优先,提升为由计划声明的具名约束;这一步才让"未启用即不生效"这条验收标准成立。 - -每一步都能单独验证;没有一步需要驱动器、唤醒或新工具。 - -**第 2 步已于 2026-09-23 落地。** 黑板端口的读就是 `legality()`:有序合法集合、运行剩下的预算,以及每个不在候选里的单元各自一条**具名原因**;原因在规则自己的 `selection` 里算出来,所以一条原因点的是某道闸门而不是把它复述一遍——而 `candidates()` 就是这份答案的 `legal`,两者不可能不一致。提问者是拥有这次运行工作区的那个进程:计划是从调用方的文件编译出来的,而守护进程只持有那些文件的冻结路径,所以共享 tool contract 的措辞没变,因为它描述的仍是同一次读。 - -**第 3 步尚未建。** 这正是这份提案目前唯一还不成立的地方:偏好延续会话的那个在线行动仍然活在共享规划器自己的排序里,而运行里记下的 `policy` 是一个**黑板会校验的名字**、不是规划器会读的开关,所以这份答案是在一条隐式约束下算出来的。 - -## 考虑过的替代方案 - -- **让程序挑下一个单元**(把现在的 `next()` 当策略用)。拒绝:一个上限不是决定——它不在合法后继里做选择;而挑人的程序就是调度器,命名决策已把"调度器"记为这个模型刻意没有变成的词。 -- **程序只存状态,合法性由 agent 各自声明**。拒绝:那么"合法"就是嗓门最大的调用者说的;两条结构性规则(一个认领只有一个持有者、交付者不能自裁)会失去家。 -- **把偏好降成建议**。拒绝:arm 之间的可比性建立在"同一计划加同样事实给出同一答案"上,而可被无视的建议恰好拿走这一点。 -- **把合法集合作为黑板 offer 发布**(把查询做成 handoff)。拒绝:这样"问一句"要占串行槽位,而且撤回 offer 是黑板协议的事。 -- **为这套安排新开工具或频道**。拒绝:台账已禁;元语必须经普通交接生效。 -- **保留修复优先作为隐藏的程序策略**。拒绝:与"程序挑人"同属一类。 - -## 验收标准 - -- 答案按单元给出、有序,并带每个单元合法或不合格的理由。 -- 答案里不出现人、收编者或裁决者。 -- 问不改变状态:不产生条目、认领、槽位或唤醒。在未变更的 store 上问两次得到同一答案,且版本未变。 -- 未启用的约束对答案没有影响:同一计划上启用与停用同一条具名约束,合法集合不同。 -- 声明的槽位预算仍然切这个答案;声明多于一个槽位仍然要求声明 handoff 目标。 -- 两条结构性规则在答案自身的术语里成立:对已持有的单元再次认领被拒;交付者不能裁决。 -- 谁收编、谁认领、谁裁决只作为黑板事实存在,没有任何程序字段声称它们。 -- 台账的禁令仍然成立:没有新工具、没有专用频道、没有第二个任务体。 - -## 风险 - -- 一条约束都不声明时,宽的合法集合可能被读成"可以一起上"的许可,把"没商量同时并行"的问题换个地方重演。缓解:槽位预算仍是声明的上限,而不是社会约定。 -- 修复优先现在住在共享规划代码里当策略,所以把它变成声明式约束是一次改动,不只是改名;搬完之前,答案是在一条隐含约束下算出来的——这是本提案唯一还不成立的地方。 -- 理由一旦成为接口,解析理由字符串的调用者就造出了第二个事实源。理由是给人看和进日志的;调用者的任何决定都必须走认领或一个显式字段。 -- agent 自由选序之后,arm 的 equal-instrument 前提取决于每次 run 记录了自己启用了哪些约束;没记录约束集合,两次 run 就不可比。 -- 查询会招来轮询。轮询便宜,而"先问后动"仍可能撞车——认领仍是仲裁者,这正是设计意图。 diff --git a/docs/decisions/proposed/2026-09-21-mechanism-not-policy.md b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.md index 85734743..dd9b0dc5 100644 --- a/docs/decisions/proposed/2026-09-21-mechanism-not-policy.md +++ b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.md @@ -3,7 +3,7 @@ [中文](2026-09-21-mechanism-not-policy.zh-CN.md) **Status:** proposed -**Relates to:** [The program answers legality](2026-09-20-the-program-answers-legality.md), [The frame, its data format, and its storage](2026-09-21-the-frame-and-its-storage.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Task unit semantics](../../design/task-unit-semantics.md) +**Relates to:** [The program answers legality](../implemented/2026-09-20-the-program-answers-legality.md), [The frame, its data format, and its storage](2026-09-21-the-frame-and-its-storage.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Task unit semantics](../../design/task-unit-semantics.md) ## Problem @@ -76,8 +76,8 @@ follows. The legality rule found today is the clearest case of the refinement this record adds: its check belongs to the program, while permission closure's name and enablement belong to a declaration. The fix is therefore not to move the check out of the program but to stop hard-wiring the rule in the kernel's vocabulary - the same -shape of fix as plan step 3 of the legality record, where repair-first becomes a declared constraint rather -than shared planning policy. Two independent fixes taking the same shape is evidence that the classification +shape of fix as step 3 of the legality record, which landed on 2026-09-23: repair-first became a declared +constraint rather than shared planning policy. Two independent fixes taking the same shape is evidence that the classification is the right one. ## Alternatives considered diff --git a/docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md index b1bdb670..c2d70d13 100644 --- a/docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md +++ b/docs/decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md @@ -3,7 +3,7 @@ [English](2026-09-21-mechanism-not-policy.md) **Status:** proposed -**Relates to:** [程序只回答合法性](2026-09-20-the-program-answers-legality.zh-CN.md)、[帧、它的数据格式与它的存储](2026-09-21-the-frame-and-its-storage.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) +**Relates to:** [程序只回答合法性](../implemented/2026-09-20-the-program-answers-legality.zh-CN.md)、[帧、它的数据格式与它的存储](2026-09-21-the-frame-and-its-storage.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) ## 问题 @@ -55,7 +55,7 @@ | `task_run_tasks.patch_files` / `patch_editable`(store 里解析) | 策略进了 schema | 该进载荷文档;机制部分是 run、task、revision、input、dependencies、operation 那几列 | | 驱动断言 `ticket.patch!.digest` | 策略断言 | 机制断言是:声明带摘要、认领绑 attempt、交付绑摘要 | -今天那条合法性规则正是本文新增的那点修正的最清楚的例子:**它的检查属于程序,而 permission closure 的名字与启用属于声明。** 所以修法不是把检查搬出程序,而是**别再用核心的词汇硬编码这条规则**——这与合法性记录计划第 3 步(把 repair-first 从共享规划策略变成声明的约束)是**同一形状的修法**。两处各自独立的修法落在同一形状上,是"这个分类是对的"的证据。 +今天那条合法性规则正是本文新增的那点修正的最清楚的例子:**它的检查属于程序,而 permission closure 的名字与启用属于声明。** 所以修法不是把检查搬出程序,而是**别再用核心的词汇硬编码这条规则**——这与合法性记录第 3 步(把 repair-first 从共享规划策略变成由计划声明的约束,2026-09-23 已落地)是**同一形状的修法**。两处各自独立的修法落在同一形状上,是"这个分类是对的"的证据。 ## 考虑过的替代方案 diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md index 9111fca4..2c426f3c 100644 --- a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.md @@ -3,7 +3,7 @@ [中文](2026-09-21-the-frame-and-its-storage.zh-CN.md) **Status:** proposed -**Relates to:** [Mechanism, not policy](2026-09-21-mechanism-not-policy.md), [The program answers legality](2026-09-20-the-program-answers-legality.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Task unit semantics](../../design/task-unit-semantics.md) +**Relates to:** [Mechanism, not policy](2026-09-21-mechanism-not-policy.md), [The program answers legality](../implemented/2026-09-20-the-program-answers-legality.md), [Name the collaboration protocol and its task-unit sub-protocol](../implemented/2026-09-20-name-the-collaboration-protocol.md), [Protocol-governed collaboration: the parts, the gaps](../../design/protocol-governed-collaboration.md), [Board governance and capability addressing](../implemented/2026-09-06-board-governance-addressing.md), [Task unit semantics](../../design/task-unit-semantics.md) ## Problem diff --git a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md index 71c966ba..2df24bf3 100644 --- a/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md +++ b/docs/decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md @@ -3,7 +3,7 @@ [English](2026-09-21-the-frame-and-its-storage.md) **Status:** proposed -**Relates to:** [机制,不是策略](2026-09-21-mechanism-not-policy.zh-CN.md)、[程序只回答合法性](2026-09-20-the-program-answers-legality.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) +**Relates to:** [机制,不是策略](2026-09-21-mechanism-not-policy.zh-CN.md)、[程序只回答合法性](../implemented/2026-09-20-the-program-answers-legality.zh-CN.md)、[给协作协议及其任务单元子协议命名](../implemented/2026-09-20-name-the-collaboration-protocol.zh-CN.md)、[协议化协作:组成部分与空缺](../../design/protocol-governed-collaboration.zh-CN.md)、[黑板治理与能力寻址](../implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[任务单元语义](../../design/task-unit-semantics.md) ## 问题 diff --git a/docs/design/mechanism-in-the-middle.md b/docs/design/mechanism-in-the-middle.md index 64731bf6..0c714ab5 100644 --- a/docs/design/mechanism-in-the-middle.md +++ b/docs/design/mechanism-in-the-middle.md @@ -7,9 +7,9 @@ A model of how a board entry travels, and the checks that decide whether the model earns a place in the records. This document owns the model, the rule that says when a boundary earns a seam, and the checks. It does not restate the parts inventory, which lives in [protocol-governed-collaboration.md](protocol-governed-collaboration.md), -nor the decisions, which live in three proposed records: [mechanism, not policy](../decisions/proposed/2026-09-21-mechanism-not-policy.md), +nor the decisions, which live in three records: [mechanism, not policy](../decisions/proposed/2026-09-21-mechanism-not-policy.md), [the frame and its storage](../decisions/proposed/2026-09-21-the-frame-and-its-storage.md), and -[the program answers legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md). +[the program answers legality](../decisions/implemented/2026-09-20-the-program-answers-legality.md). ## The model diff --git a/docs/design/mechanism-in-the-middle.zh-CN.md b/docs/design/mechanism-in-the-middle.zh-CN.md index 7ad39207..536cd545 100644 --- a/docs/design/mechanism-in-the-middle.zh-CN.md +++ b/docs/design/mechanism-in-the-middle.zh-CN.md @@ -4,7 +4,7 @@ **Created:** 2026-09-21 **Updated:** 2026-09-21 -一个黑板条目怎么走,以及决定这个模型有没有资格进记录的那几项检查。本文拥有模型、"一条边界什么时候够资格建缝"的那条规则、以及这些检查。它不复述大类清单(那在 [protocol-governed-collaboration.md](protocol-governed-collaboration.zh-CN.md)),也不复述决策(那在三份提案里:[机制,不是策略](../decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md)、[帧、它的数据格式与它的存储](../decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md)、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md))。 +一个黑板条目怎么走,以及决定这个模型有没有资格进记录的那几项检查。本文拥有模型、"一条边界什么时候够资格建缝"的那条规则、以及这些检查。它不复述大类清单(那在 [protocol-governed-collaboration.md](protocol-governed-collaboration.zh-CN.md)),也不复述决策(那在三份记录里:[机制,不是策略](../decisions/proposed/2026-09-21-mechanism-not-policy.zh-CN.md)、[帧、它的数据格式与它的存储](../decisions/proposed/2026-09-21-the-frame-and-its-storage.zh-CN.md)、[程序只回答合法性](../decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md))。 ## 模型 diff --git a/docs/design/ooo-fusion-planning.md b/docs/design/ooo-fusion-planning.md index 5c9677b6..99ff0d22 100644 --- a/docs/design/ooo-fusion-planning.md +++ b/docs/design/ooo-fusion-planning.md @@ -39,12 +39,16 @@ So the online decision is not a plan, it is one move about the _current_ session Three properties make that move safe to make repeatedly: -- **Repair-first.** The default is to continue the current session; a re-decision may only _end_ it, on - a declared change: a rejected verdict, a cancellation, a dependency that did not become accepted, or - a declared external wait that is not ready. Repairing instead of re-planning is the documented - trade: reusing a plan saves work but risks acting on a stale one, and re-planning from scratch churns - (`plan repair versus full replanning`). Repair-first is the middle: keep what is committed, decide - only about what has not run. +- **Repair-first.** Continuing the current session is a constraint **the plan declares**, not this + design's default: a plan that enables `repair-first` continues the session, and a plan that enables + nothing runs one unit per session and closes at every unit boundary. Once enabled, a re-decision may + only _end_ the session, on a declared change: a rejected verdict, a cancellation, a dependency that + did not become accepted, or a declared external wait that is not ready. Repairing instead of + re-planning is the documented trade: reusing a plan saves work but risks acting on a stale one, and + re-planning from scratch churns (`plan repair versus full replanning`). Repair-first is the middle: + keep what is committed, decide only about what has not run. The names the protocol defines are a + closed list (`PLAN_CONSTRAINTS`), and a plan naming one outside it is refused by name rather than + read as having asked for nothing. - **Baseline.** A move, once made, is not revisited while its facts hold. Facts arrive at boundaries; between boundaries there is nothing to re-decide, so an unchanged fact set yields an unchanged set of moves. This is Oracle SQL Plan Management's plan-baseline idea: constrain the plan to accepted diff --git a/docs/design/protocol-governed-collaboration.md b/docs/design/protocol-governed-collaboration.md index e9211bd2..e1bcd1cc 100644 --- a/docs/design/protocol-governed-collaboration.md +++ b/docs/design/protocol-governed-collaboration.md @@ -21,8 +21,8 @@ document and an owner disagree, the owner wins. | Scope and visibility | An entry targets an authorized agent subset, so boards are not all-to-all | Complete | [board governance and capability addressing](../decisions/implemented/2026-09-06-board-governance-addressing.md) | | Truth versus coordination state | The board is a temporary coordination medium; durable memory is the truth store; `memory=` pointers | Complete | [board governance](../decisions/implemented/2026-09-06-board-governance-addressing.md), [memory graphs](memory-graphs.md) | | The task unit | A unit declared by inputs, dependencies, acceptance, capability and budget; decomposition; the internal compile view; the run, its adoption, and the facts the runtime owns | Complete | [task unit semantics](task-unit-semantics.md), [the obligations ledger](task-unit-semantics-obligations.md) | -| Legality and admission | The ordered legal set, the cut to the declared slot budget, the reason each unit is or is not legal | Mechanism complete; the readable answer is missing | `src/integration/ooo-board.ts`, [the program answers legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md) | -| Ordering and constraints | Determinism (one plan plus one set of facts yields one answer), repair-first, tie-breaking by plan order | Mechanism complete; the declaration is missing | [fusion planning: repair-first](../decisions/implemented/2026-09-19-fusion-planning-repair-first.md), [the program answers legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md) | +| Legality and admission | The ordered legal set, the cut to the declared slot budget, the reason each unit is or is not legal | Complete | `src/integration/ooo-board.ts`, [the program answers legality](../decisions/implemented/2026-09-20-the-program-answers-legality.md) | +| Ordering and constraints | Determinism (one plan plus one set of facts yields one answer), repair-first, tie-breaking by plan order | Complete | [fusion planning: repair-first](../decisions/implemented/2026-09-19-fusion-planning-repair-first.md), [the program answers legality](../decisions/implemented/2026-09-20-the-program-answers-legality.md) | | Budget and accounting | Declared slot budget, token and cache accounting, the offline cost ceiling | Complete | [declared slot budget](../decisions/implemented/2026-09-18-declared-slot-budget.md), [the cost model](../experiments/execution/ooo-cost-model-2026-09-17.md) | | Cancellation and fencing | Explicit cancellation, no orphan worker or check after it, late artifacts fenced | Complete | [long checks run detached](../decisions/implemented/2026-09-18-detached-long-checks.md), [the bootstrap design](ooo-execution-bootstrap.md) | | Recovery and replay | Log plus replay of a round, the transactional outbox drained after commit and on restart, no stealing another's work after a crash | Complete inside the arms; the product side is not wired | [the bootstrap design](ooo-execution-bootstrap.md) | @@ -46,9 +46,12 @@ rows above are the ones a given piece of work cannot do without. missing is the asker without a workspace: a plan is compiled from the caller's files, so a process that holds only the store cannot compute the answer. The concepts are unchanged; the gap now narrows to where the plan lives. -3. **Constraints have no declaration.** Repair-first currently lives as policy inside shared planning - code. Making preferences named constraints a plan enables is **declarative policy**, and the - requirement that one plan plus one set of facts yields one answer is **deterministic replay**. +3. **Constraints now have a declaration; what a constraint may change is still unwritten.** + `repair-first` is a named constraint a plan enables, the shared planner reads that declaration + instead of holding the preference itself, and a plan that enables nothing runs one unit per session. + What has no clause yet is who may enable a constraint, and what a constraint is allowed to change: + this one orders a session, and a name that changed which units are *legal* would be a different kind + of answer. The concepts are **declarative policy** and **deterministic replay**. 4. **Read guarantees are unwritten.** Board reads are cursor-based and may lag; nothing states whether a reader may assume monotonic reads or read-your-writes. The concepts are **monotonic reads**, **read-your-writes** and, across one agent's own sequence of sessions, **causal consistency**. @@ -89,8 +92,9 @@ rows above are the ones a given piece of work cannot do without. ## What would make this inventory complete -Row 9's readable answer landed on 2026-09-23 as the board port's own read - see [The program answers -legality](../decisions/proposed/2026-09-20-the-program-answers-legality.md), whose step 3 is the row 10 it -still owes. Rows 1, 4, 5, 6 and 7 are gaps with no owner yet. Until each gap either gets a clause or -is written down as deliberately absent, "complete" for this umbrella means the parts that have owners, -not the parts a task needs. +Row 9's readable answer landed on 2026-09-23 as the board port's own read, and row 10's declaration +landed the same day: a plan enables `repair-first` and the planner reads the declaration rather than +assuming it - see [The program answers +legality](../decisions/implemented/2026-09-20-the-program-answers-legality.md). Rows 1, 4, 5, 6 and 7 are +gaps with no owner yet. Until each gap either gets a clause or is written down as deliberately absent, +"complete" for this umbrella means the parts that have owners, not the parts a task needs. diff --git a/docs/design/protocol-governed-collaboration.zh-CN.md b/docs/design/protocol-governed-collaboration.zh-CN.md index e63479ea..71272e25 100644 --- a/docs/design/protocol-governed-collaboration.zh-CN.md +++ b/docs/design/protocol-governed-collaboration.zh-CN.md @@ -18,8 +18,8 @@ | 作用域与可见性 | 条目的可见范围是授权过的 agent 子集,黑板不是 all-to-all | 完整 | [黑板治理与能力寻址](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md) | | 真值与协调状态分离 | 黑板是临时协调面,真值在 durable memory,`memory=` 做指针 | 完整 | [黑板治理](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md)、[记忆图](memory-graphs.md) | | 任务单元 | 由输入、依赖、验收、能力、预算声明的单元;拆分;内部编译视图;run、它的收编、以及运行时拥有的事实 | 完整 | [任务单元语义](task-unit-semantics.md)、[义务台账](task-unit-semantics-obligations.md) | -| 合法性与准入 | 有序合法集合、按声明槽位预算切一刀、每个单元合法与否的理由 | 机制完整;可读的答案缺失 | `src/integration/ooo-board.ts`、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md) | -| 顺序与约束 | 确定性(同一计划加同一组事实只产出一个答案)、修复优先、按计划顺序破平 | 机制完整;声明方式缺失 | [融合规划的修复优先](../decisions/implemented/2026-09-19-fusion-planning-repair-first.zh-CN.md)、[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md) | +| 合法性与准入 | 有序合法集合、按声明槽位预算切一刀、每个单元合法与否的理由 | 完整 | `src/integration/ooo-board.ts`、[程序只回答合法性](../decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md) | +| 顺序与约束 | 确定性(同一计划加同一组事实只产出一个答案)、修复优先、按计划顺序破平 | 完整 | [融合规划的修复优先](../decisions/implemented/2026-09-19-fusion-planning-repair-first.zh-CN.md)、[程序只回答合法性](../decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md) | | 预算与计量 | 声明槽位预算、token 与缓存计量、离线成本上限 | 完整 | [声明式槽位预算](../decisions/implemented/2026-09-18-declared-slot-budget.zh-CN.md)、[成本模型](../experiments/execution/ooo-cost-model-2026-09-17.md) | | 取消与围栏 | 显式取消、取消后不留孤儿 worker 或检查、迟到产物被围栏挡住 | 完整 | [长检查分离运行](../decisions/implemented/2026-09-18-detached-long-checks.zh-CN.md)、[自举设计](ooo-execution-bootstrap.md) | | 恢复与重放 | 轮次的日志加重放、事务性 outbox 在提交后与重启时排空、崩溃后不偷别人的活 | 臂内完整;产品面尚未接线 | [自举设计](ooo-execution-bootstrap.md) | @@ -32,7 +32,7 @@ 1. **收编者没有协议条款。** 谁有资格把一次讨论收编成一次 run,以及租约失效后谁负责把活重新推出去,都没有写。机制是有的(`expiry`、串行晋升、派发循环),责任没有。分布式里把前者叫做**提交边界**——让一份商定的安排生效的那个唯一点;把后者叫做**孤儿回收/重启责任**。 2. **合法性现在有了可读的答案;但只握着 store 的问不出来。** 黑板端口回答 `legality()`——有序合法集合、剩下的预算、以及每个单元各自一条**具名原因**——而这次读不改变任何状态,所以**准入控制**加一次**只读的前置条件检查**现在有了家。仍然缺的是**没有工作区的提问者**:计划是从调用方的文件编译出来的,所以只持有 store 的进程算不出这份答案。概念没变,缺口现在窄到“计划住在哪里”这一点上。 -3. **约束没有声明方式。** 修复优先目前活在共享规划代码里当策略。把偏好变成计划可启用的具名约束,是**声明式策略**;而"同一计划加同一组事实只产出一个答案"这个要求,是**确定性重放**。 +3. **约束现在有了声明方式;约束能改变什么还没写下来。** `repair-first` 是计划可启用的具名约束,共享规划器读这份声明,而不是自己持有这个偏好;一个什么都没启用的计划就是每个会话一个单元。还没有条款的是:谁有权启用一条约束,以及一条约束允许改变什么——这一条只排会话的顺序,而某个会改变哪些单元**合法**的名字,会是一种不同性质的答案。概念是**声明式策略**和**确定性重放**。 4. **读的保证没有写下来。** 黑板读基于游标、可能落后;没有一处说明读者是否可以假设单调读或读己所写。概念是**单调读**、**读己所写**,以及跨同一个 agent 自己一串会话的**因果一致性**。 5. **分歧没有仲裁者。** 两个 agent 对下一步做什么意见不合时,只能争夺一次完成(`veto`)或抢认领。对分歧本身没有条款。最近的概念是**乐观并发加冲突检测**,以及——因为仲裁者本来就存在——干脆的**单写者决定**。竞价式分配早就被拒绝了([黑板治理](../decisions/implemented/2026-09-06-board-governance-addressing.zh-CN.md))。 6. **跨 run 的资源没有互斥。** 多个 run 抢同一个工作树或同一个文件,现在靠实践(每条臂一个一次性 worktree)和槽位预算应付,没有规则。概念是**带围栏的互斥锁**——租约失效之后,它的持有者就不能再被相信。 @@ -52,4 +52,4 @@ ## 这份清单要怎样才算完整 -第 9 行的可读答案已于 2026-09-23 以黑板端口自己的读落地——见[程序只回答合法性](../decisions/proposed/2026-09-20-the-program-answers-legality.zh-CN.md),它还欠的第 10 行正是那份提案的第 3 步。第 1、4、5、6、7 行是还没有主的空缺。在每个空缺或者拿到条款、或者被明确写成"刻意没有"之前,这个大类的"完整"指的只是那些有归属的部分,而不是一件任务真正需要的那几样。 +第 9 行的可读答案已于 2026-09-23 以黑板端口自己的读落地,第 10 行的声明方式同一天落地:计划启用 `repair-first`,规划器读这份声明而不是自己假定它——见[程序只回答合法性](../decisions/implemented/2026-09-20-the-program-answers-legality.zh-CN.md)。第 1、4、5、6、7 行是还没有主的空缺。在每个空缺或者拿到条款、或者被明确写成"刻意没有"之前,这个大类的"完整"指的只是那些有归属的部分,而不是一件任务真正需要的那几样。 diff --git a/evals/ooo-execution/make-fusion-trial-specs.mjs b/evals/ooo-execution/make-fusion-trial-specs.mjs index fa8fc0db..d8d10ad0 100644 --- a/evals/ooo-execution/make-fusion-trial-specs.mjs +++ b/evals/ooo-execution/make-fusion-trial-specs.mjs @@ -41,7 +41,7 @@ if (fixture.worker?.kind !== "canned") { const live = { ...fixture, worker: { kind: "pi", provider, model } }; const control = { ...live, id: "pipeline-control" }; -const fusion = { ...live, id: "pipeline-fusion", fusion: { unitsPerSession } }; +const fusion = { ...live, id: "pipeline-fusion", fusion: { unitsPerSession, constraints: ["repair-first"] } }; mkdirSync(OUT, { recursive: true }); writeFileSync(`${OUT}/control.spec.json`, `${JSON.stringify(control, null, 2)}\n`); diff --git a/evals/ooo-execution/plan-driver.test.ts b/evals/ooo-execution/plan-driver.test.ts index c66d1b64..4cf2c829 100644 --- a/evals/ooo-execution/plan-driver.test.ts +++ b/evals/ooo-execution/plan-driver.test.ts @@ -14,6 +14,7 @@ import { piWorker, runPlan, specFrom, + validateSpecFile, type PlanDriverSpec, type PlanWorker, type SpecFile, @@ -132,7 +133,11 @@ function sessionWorker(latencyMs = 20): PlanWorker { test("a fused run runs several units in one session, each with its own ticket and verdict", async () => { const run = await runPlan( - spec({ slots: 1, fusion: { unitsPerSession: 2 }, worker: sessionWorker(10) }), + spec({ + slots: 1, + fusion: { unitsPerSession: 2, constraints: ["repair-first"] }, + worker: sessionWorker(10), + }), ); // One session of two units, then a yield boundary and a second session: `summary` becomes a // candidate only once `third` is accepted, so the chain continues into it. @@ -164,7 +169,11 @@ test("a fused session ends where the next unit needs another capability", async const run = await runPlan( spec({ slots: 1, - fusion: { unitsPerSession: 4, declarations: { third: { capability: "other" } } }, + fusion: { + unitsPerSession: 4, + constraints: ["repair-first"], + declarations: { third: { capability: "other" } }, + }, worker: sessionWorker(10), }), ); @@ -187,7 +196,9 @@ test("a unit with no verdict ends the session it was running in", async () => { const produced = await sessionWorker(10)(taskId, frozen, dependencies, session); return produced; }; - const run = await runPlan(spec({ slots: 1, fusion: { unitsPerSession: 4 }, worker })); + const run = await runPlan( + spec({ slots: 1, fusion: { unitsPerSession: 4, constraints: ["repair-first"] }, worker }), + ); assert.equal(run.failures, 1); assert.ok(run.incomplete.some((entry) => entry.includes("second"))); assert.deepEqual(run.sessions, [["first"]], "the session ended at the unit with no verdict"); @@ -200,7 +211,11 @@ test("a fused session does not continue from a unit the host rejected", async () const bad: readonly DataCheck[] = [ dataCheck("unit check", { status: "failed", log: "the host refused this answer" }), ]; - const base = spec({ slots: 1, fusion: { unitsPerSession: 4 }, worker: sessionWorker(10) }); + const base = spec({ + slots: 1, + fusion: { unitsPerSession: 4, constraints: ["repair-first"] }, + worker: sessionWorker(10), + }); const run = await runPlan({ ...base, units: { ...base.units, second: { ...base.units["second"]!, checks: bad } }, @@ -225,7 +240,11 @@ test("a worker that starts its own session is not reported as fusion", async () // `recordingWorker` never echoes a session, which is what a worker that opens a fresh session per // unit looks like from the driver's side. The run must report boundaries, not fusion. const run = await runPlan( - spec({ slots: 1, fusion: { unitsPerSession: 4 }, worker: recordingWorker(10) }), + spec({ + slots: 1, + fusion: { unitsPerSession: 4, constraints: ["repair-first"] }, + worker: recordingWorker(10), + }), ); assert.deepEqual(run.sessions, [["first"], ["second"], ["third"], ["summary"]]); assert.ok( @@ -469,12 +488,12 @@ test("a spec file's fusion block reaches the run it describes", () => { }, }, worker: { kind: "stub" as const, latencyMs: 1 }, - fusion: { unitsPerSession: 2 }, + fusion: { unitsPerSession: 2, constraints: ["repair-first"] }, }; const declared = specFrom(file, recordingWorker(), 1); assert.deepEqual( declared.fusion, - { unitsPerSession: 2 }, + { unitsPerSession: 2, constraints: ["repair-first"] }, "a spec that asked for fusion must not be run as the control arm", ); const without: SpecFile = { ...file }; @@ -485,3 +504,31 @@ test("a spec file's fusion block reaches the run it describes", () => { "a spec that declared none has none: the two readings must not be the same run", ); }); + +test("a spec cannot declare a bound it does not enable the constraint for", () => { + const file: SpecFile = { + baseline: ["evals/ooo-execution/fixtures/report/alpha.ts"], + plan: [{ id: "first", effect: "isolated-artifact" }], + units: { + first: { + instruction: "work on first", + editable: ["evals/ooo-execution/fixtures/report/alpha.ts"], + checks: [{ label: "ok", test: "evals/ooo-execution/fixtures/report/alpha.test.ts" }], + }, + }, + worker: { kind: "stub", latencyMs: 1 }, + fusion: { unitsPerSession: 2 }, + }; + // Without the declaration the run would be the control arm while the file says fusion, so the spec + // is refused by name rather than believed. A bound of one asks for nothing and stays legal. + assert.throws( + () => validateSpecFile(file), + /declare fusion\.constraints = \["repair-first"\]/u, + ); + assert.equal(validateSpecFile({ ...file, fusion: { unitsPerSession: 1 } }).fusion?.unitsPerSession, 1); + assert.equal( + validateSpecFile({ ...file, fusion: { unitsPerSession: 2, constraints: ["repair-first"] } }).fusion + ?.unitsPerSession, + 2, + ); +}); diff --git a/evals/ooo-execution/plan-driver.ts b/evals/ooo-execution/plan-driver.ts index dd41c4bb..1cb95fda 100644 --- a/evals/ooo-execution/plan-driver.ts +++ b/evals/ooo-execution/plan-driver.ts @@ -82,6 +82,10 @@ export interface PlanDriverSpec { * chains, so this bound is what keeps a fused run from swallowing the plan. Omitted: no fusion. */ fusion?: { unitsPerSession: number; + /** The constraints the spec enables for that fusion, by the protocol's names. A fused cell declares + * `repair-first`: the move that continues a session is a declared constraint now, not the planner's + * default, so a spec that asks for fusion without it would run the control arm and be read as one. */ + constraints?: readonly string[]; /** Per-unit session declarations. Omitted: every unit shares the run's one capability and * authority, and only its own visibility decides. */ declarations?: Readonly>; @@ -253,6 +257,7 @@ export async function runPlan(spec: PlanDriverSpec): Promise { declarations: Object.fromEntries( spec.plan.map((row) => [String(row[0]), declarationOf(String(row[0]))]), ), + ...(spec.fusion?.constraints ? { constraints: spec.fusion.constraints } : {}), ...(pendingBranches.length ? { pendingBranches } : {}), }; }; @@ -434,6 +439,7 @@ export interface SpecFile { * the control arm. */ fusion?: { unitsPerSession: number; + constraints?: readonly string[]; declarations?: Readonly>; }; worker: @@ -462,7 +468,15 @@ function checkList(raw: readonly { label: string; test: string }[]): DataCheck[] } function readSpecFile(path: string): SpecFile { - const file = JSON.parse(readFileSync(path, "utf8")) as SpecFile; + return validateSpecFile(JSON.parse(readFileSync(path, "utf8")) as SpecFile); +} + +/** + * Refuse a spec file that cannot be run as written, by name, before any of it is believed. Exported + * because the refusals are the driver's contract with whoever writes a spec: a file that declares + * fusion without the constraint that continues a session would run the control arm. + */ +export function validateSpecFile(file: SpecFile): SpecFile { if (!Array.isArray(file.plan) || !file.plan.length) throw new Error("spec.plan must be non-empty"); if (!Array.isArray(file.baseline) || !file.baseline.length) @@ -473,6 +487,17 @@ function readSpecFile(path: string): SpecFile { for (const id of Object.keys(file.units)) if (!planIds.has(id)) throw new Error(`spec.units.${id} is not in the plan; the plan is the authority`); + // A spec that declares a bound above one and no constraint enabling the continuation would run the + // control arm while its file says fusion, which is the one confusion this driver must not create. + if ( + file.fusion && + (file.fusion.unitsPerSession ?? 1) > 1 && + !(file.fusion.constraints ?? []).includes("repair-first") + ) + throw new Error( + "spec.fusion declares more than one unit per session but enables no constraint that continues a " + + 'session; declare fusion.constraints = ["repair-first"]', + ); return file; } diff --git a/src/integration/ooo-execution.ts b/src/integration/ooo-execution.ts index 34d764b0..41f5ed56 100644 --- a/src/integration/ooo-execution.ts +++ b/src/integration/ooo-execution.ts @@ -320,6 +320,11 @@ export interface SessionPlan { tasks: readonly DispatchTask[]; declarations: Readonly>; pendingBranches?: readonly string[]; + /** The constraints this plan enables, by the protocol's names - the vocabulary is + * `PLAN_CONSTRAINTS` in `ooo-fusion-plan.ts`, and the move that reads them lives there too. + * Absent means none: a plan that declares nothing gets the baseline (one unit per session, in + * plan order) rather than whatever the planner happens to prefer. */ + constraints?: readonly string[]; } function subset(inner: readonly string[], outer: readonly string[]): boolean { diff --git a/src/integration/ooo-fusion-plan.ts b/src/integration/ooo-fusion-plan.ts index c81b8995..009faee9 100644 --- a/src/integration/ooo-fusion-plan.ts +++ b/src/integration/ooo-fusion-plan.ts @@ -42,6 +42,7 @@ export function optimisticPlan(plan: SessionPlan): SessionPlan { externalReady: true, })), declarations: plan.declarations, + ...(plan.constraints ? { constraints: plan.constraints } : {}), }; } @@ -201,13 +202,48 @@ export type SessionMove = | { readonly kind: "close"; readonly reason: string }; /** - * Repair-first: the default is to continue the session that is already running, and a close carries - * the condition that closed it. Only facts may end a session early - a rejected verdict shows up as + * The constraints the protocol defines, by name. This is the whole list: a preference that is not on it + * is not a constraint a plan may enable, and a plan that names one is refused rather than read. + * + * `repair-first` is the one the arms measured - continue the session that is already running instead of + * starting a fresh one at every unit boundary. It used to be this module's unstated default, which is + * what made a plan's meaning depend on which planner read it; now it is a name the plan declares, and a + * plan that declares nothing runs one unit per session. + */ +export const PLAN_CONSTRAINTS = ["repair-first"] as const; + +/** + * The constraints a plan enables, validated. A name outside `PLAN_CONSTRAINTS` is refused by name + * instead of ignored: a plan that enables a constraint the program does not know is asking for something + * nobody implements, and reading it as "nothing was asked" would make the declaration decoration. + */ +export function enabledConstraints( + plan: SessionPlan, +): readonly (typeof PLAN_CONSTRAINTS)[number][] { + const named = plan.constraints ?? []; + const unknown = named.filter((name) => !(PLAN_CONSTRAINTS as readonly string[]).includes(name)); + if (unknown.length > 0) + throw new Error( + `unknown constraint ${unknown.join(", ")}; the protocol defines ${PLAN_CONSTRAINTS.join(", ")}`, + ); + return named as readonly (typeof PLAN_CONSTRAINTS)[number][]; +} + +/** + * The move at a session boundary, computed under the plan's enabled constraints. A close carries the + * condition that closed it; only facts may end a session early - a rejected verdict shows up as * `current` not being accepted, a cancellation and an unmet dependency as `sharedSessionLegal`, and a * declared external wait as condition 4 - so this is a pure function of the plan and its facts, with * plan order breaking ties. Numeric optimisation belongs to the offline half and is not read here. */ export function nextSessionMove(input: SessionMoveInput): SessionMove { + // The continuation is a declared constraint rather than this function's default. Without it the plan + // runs one unit per session, which is the baseline the arms' control cells measure. + if (!enabledConstraints(input.plan).includes("repair-first")) + return { + kind: "close", + reason: "repair-first is not enabled, so this plan runs one unit per session", + }; if (input.size >= input.bound) return { kind: "close", reason: "the declared bound is reached" }; const successor = fusionSuccessors(input.current, input.plan).find((id) => input.onOffer.includes(id), diff --git a/tests/integration/ooo-dispatch.test.ts b/tests/integration/ooo-dispatch.test.ts index 70c3d835..b9b1cd50 100644 --- a/tests/integration/ooo-dispatch.test.ts +++ b/tests/integration/ooo-dispatch.test.ts @@ -184,6 +184,9 @@ function legality( declarations: Object.fromEntries( order.map((id) => [id, { capability: "patch", authority: "host", visible: [] }]), ), + // The cases that declare sessions here expect the chain to continue, so the plan enables the + // constraint that allows it. A case with no sessions never reaches the move at all. + constraints: ["repair-first"], }; } diff --git a/tests/integration/ooo-fusion-plan.test.ts b/tests/integration/ooo-fusion-plan.test.ts index 74e948cd..86589f87 100644 --- a/tests/integration/ooo-fusion-plan.test.ts +++ b/tests/integration/ooo-fusion-plan.test.ts @@ -3,12 +3,15 @@ * * The cases below pin the three properties the design claims - a move is a pure function of the plan * and its facts, an offline graph prices the best case rather than a run's state, and a floor is a - * floor - plus the two refusals that keep a bound from turning into a schedule. + * floor - plus the two refusals that keep a bound from turning into a schedule, and the declaration + * that keeps a preference from becoming a default. */ import assert from "node:assert/strict"; import test from "node:test"; import { + PLAN_CONSTRAINTS, chainCoverFloor, + enabledConstraints, fusionGraph, listScheduleSessions, nextSessionMove, @@ -48,6 +51,7 @@ function plan(count: number, over: Partial = {}): SessionPlan { return { tasks, declarations: Object.fromEntries(names.map((name) => [name, declaration()])), + constraints: ["repair-first"], ...over, }; } @@ -160,6 +164,7 @@ test("the move closes the session by name, never by guessing", () => { const incompatible: SessionPlan = { tasks: [task("one"), task("two")], declarations: { one: declaration(), two: declaration({ capability: "other" }) }, + constraints: ["repair-first"], }; assert.deepEqual( nextSessionMove({ plan: incompatible, current: "one", size: 1, bound: 2, onOffer: ["two"] }), @@ -171,3 +176,38 @@ test("the same plan and the same facts yield the same move", () => { const input = { plan: plan(3), current: "one", size: 1, bound: 3, onOffer: ["two", "three"] }; assert.deepEqual(nextSessionMove(input), nextSessionMove(input)); }); + +test("the continuation is a declared constraint, not the planner's default", () => { + const boundary = { current: "one", size: 1, bound: 2, onOffer: ["two"] }; + const declared = plan(3); + const silent: SessionPlan = { ...declared, constraints: [] }; + + // The same plan and the same facts, and the only difference is which constraints the plan enabled. + assert.deepEqual(nextSessionMove({ plan: declared, ...boundary }), { kind: "admit", unit: "two" }); + assert.deepEqual(nextSessionMove({ plan: silent, ...boundary }), { + kind: "close", + reason: "repair-first is not enabled, so this plan runs one unit per session", + }); + + // And the constraint that is not enabled is absent from the plan's own answer rather than assumed. + assert.deepEqual(enabledConstraints(declared), ["repair-first"]); + assert.deepEqual(enabledConstraints(silent), []); + assert.deepEqual(enabledConstraints({ tasks: [], declarations: {} }), []); + assert.deepEqual(PLAN_CONSTRAINTS, ["repair-first"], "the protocol's list is the whole list"); +}); + +test("a constraint the protocol does not define is refused by name, never ignored", () => { + const unknown: SessionPlan = { + tasks: [task("one"), task("two")], + declarations: { one: declaration(), two: declaration() }, + constraints: ["faster"], + }; + assert.throws( + () => enabledConstraints(unknown), + /unknown constraint faster; the protocol defines repair-first/u, + ); + assert.throws( + () => nextSessionMove({ plan: unknown, current: "one", size: 1, bound: 2, onOffer: ["two"] }), + /unknown constraint faster/u, + ); +}); diff --git a/tests/integration/ooo-session-facts.test.ts b/tests/integration/ooo-session-facts.test.ts index eab084ca..accd11ea 100644 --- a/tests/integration/ooo-session-facts.test.ts +++ b/tests/integration/ooo-session-facts.test.ts @@ -127,6 +127,9 @@ function linearPlan(): SessionPlan { { capability: "patch", authority: "host", visible: ["src/a.ts"] }, ]), ), + // Every case in this file asks about a boundary of a session that continues, so the plan enables + // the constraint that lets it: the decision, not the plan's declarations, is what is under test. + constraints: ["repair-first"], }; } diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 7586971b..6e65b2a1 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -665,6 +665,31 @@ const TARGETS: readonly Target[] = [ }, ], }, + { + target: "src/integration/ooo-fusion-plan.ts", + suites: ["tests/integration/ooo-fusion-plan.test.ts"], + mutants: [ + { + // The continuation is a constraint the plan declares rather than this planner's default. + // Dropping the declaration restores a preference no plan enabled, which is what made a plan's + // meaning depend on which planner read it. + name: "the-continuation-is-not-declared", + ast: { within: "nextSessionMove" }, + from: ' if (!enabledConstraints(input.plan).includes("repair-first"))', + to: " if (false)", + expect: "the continuation is a declared constraint, not the planner's default", + }, + { + // A name the protocol does not define is refused by name; reading it as "nothing was asked" + // is what would make the declaration decoration. + name: "an-unknown-constraint-is-ignored", + ast: { within: "enabledConstraints" }, + from: " const unknown = named.filter((name) => !(PLAN_CONSTRAINTS as readonly string[]).includes(name));", + to: " const unknown: readonly string[] = [];", + expect: "a constraint the protocol does not define is refused by name, never ignored", + }, + ], + }, { target: "src/integration/task-semantics-interleavings.ts", suites: ["tests/integration/ooo-publication-invariants.test.ts"], From dbf10ac539d3bfffea65d0743ed1c8397867f347 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Wed, 23 Sep 2026 23:22:43 +0800 Subject: [PATCH 22/32] Two teeth had stopped pointing at their rule The sweep refuses to claim a check when a mutant's marker no longer matches the source, and two markers had drifted away from the code they were written against: `the-status-query-ignores-the-declared-budget` still named a `startableTasks(dispatchTasks(units, facts), slots)` call that `deriveStatus` now splits, and `a-cancelled-run-still-accepts-writes` still named a two-line refusal that `managedWriteRefusal` now assembles from a subject. Both were reported as "marker not found, refusing to claim a check" rather than as caught, which is the honest failure and also a hole: nothing in the build notices a tooth that stopped touching its rule. Re-pointed, not rewritten: the budget mutant now hardcodes one slot in the call the module makes, and the cancellation mutant inverts the guard inside `managedWriteRefusal`. Measured after the fix: 27 of 27 caught by the named cases, 2 of 2 files restored byte-identically. --- tools/mutation-teeth.ts | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 6e65b2a1..9e6922fe 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -141,8 +141,8 @@ const TARGETS: readonly Target[] = [ // would report a ready set narrower than what the run declared, and the two answers would differ. name: "the-status-query-ignores-the-declared-budget", ast: { within: "deriveStatus" }, - from: " const ready = startableTasks(dispatchTasks(units, facts), slots);", - to: " const ready = startableTasks(dispatchTasks(units, facts), 1);", + from: " const ready = startableTasks(tasks, slots);", + to: " const ready = startableTasks(tasks, 1);", expect: "a run that declares more slots reports the tasks it may start, not just the head", }, ], @@ -1124,8 +1124,9 @@ const TARGETS: readonly Target[] = [ // A cancelled run is the end of its managed entries' lifecycle, and the fence is the only // thing that says so. name: "a-cancelled-run-still-accepts-writes", - from: " if (cancelled)\n return `run ${runId} was cancelled at sequence ${cancelled.sequence}; its managed entries take no further lifecycle writes`;", - to: " if (cancelled && false)\n return `run ${runId} was cancelled at sequence ${cancelled.sequence}; its managed entries take no further lifecycle writes`;", + ast: { within: "managedWriteRefusal" }, + from: " if (!cancelled) return null;", + to: " if (cancelled) return null;", expect: "a cancelled run takes no further lifecycle writes on what it adopted", }, { From f74859e0f0e1b44c330b82e5890390960acad76b Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Thu, 24 Sep 2026 21:59:25 +0800 Subject: [PATCH 23/32] A tooth is an operator over a selector, not a copy of a line 89% of the mutation ledger's teeth (133 of 149) are a site plus one of a handful of operators, and 52% of them (78) are a bare byte fragment matched anywhere in the file. That is the shape that dies first: both stale teeth repaired in dbf10ac5 were of it. This lands the first two slices of the plan to derive the mutant instead of storing it. The vocabulary (tools/mutation-anchor.ts, new): six operators over a member-scoped selector - condition-never, condition-holds, neutralize-term, replace-argument, replace-property, drop-statement. Every selector resolves to exactly one site or refuses, and the refusals are the interesting half: a member named twice, a call matched twice, a fragment that fits two decision positions (the innermost one wins; two disjoint ones are refused), an operator that replaces a whole condition given a fragment that names only part of it (that would silently widen the mutant into "every reason this rule has"), a statement that does not own its line. A condition is any expression in a decision position - a test, a returned value, an arrow's body, a value bound to a name - because `return a && b` and `array.filter((x) => x.y)` are the same rule written without an `if`. The resolver holds no state, so it moved out of the sweep script: it is 16 tests over source strings, with no filesystem. The tool (tools/mutation-teeth.ts) keeps `from`/`to` and `ast` forms working and gains `--anchors-only`: every selected anchor resolved, nothing written, no suite run, 0.6s for all 149 - the pass that answers "is the tooth still aimed at something", which nothing answered before. The pilot (src/integration/ooo-execution.ts, 24 teeth over two target entries): 19 converted, 5 staying hand-written and named in the record, each still a single expression or value rather than a statement or a message. Sweep after conversion: 24 of 24 caught, restored byte-identically 2 of 2. 23 are caught by the case their `expect` names; the stale name fusion-continues-from-an-unverified-answer was already that way and this conversion neither caused nor fixed it. The demonstration, and the prediction that was wrong: a rename plus `spent || severalWaits` lifted into `const closed` retired two sites - the hand anchor, and the derived a-live-claim-does-not-block-selection, because a condition bound to a name was not a decision position yet. The prediction written before running it had been "one failure". The selector was widened to include a variable's initializer, and the corrected prediction held exactly: anchors: 148 of 149 resolve, the one failure the hand anchor, file restored byte-identically. The limitation is real and recorded rather than hidden: a derived selector is scoped to a decision position, so a rule rewritten into a different shape can still retire its tooth - but now it does so visibly, naming the tooth and the fragment. Not in this commit: the anchors-only pass as an atomic check of the ci-and-tests route, and the retirement pass over the ~18-20 teeth whose rule already has a relational or enumerative check. The record (docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md, with its zh-CN pair) marks both, and holds the measurement behind the operator list. Readings at this revision: npm run mutation:teeth -- --anchors-only -> 149 of 149 resolve over 23 targets, exit 0; the target sweep -> 24 of 24 caught; tests/tools/mutation-anchor.test.ts -> 16 pass, 0 fail; npm run verify:static -> exit 0. --- ...-09-24-mutants-are-derived-not-anchored.md | 184 +++++++ ...-mutants-are-derived-not-anchored.zh-CN.md | 70 +++ .../design/task-unit-semantics-obligations.md | 10 + tests/tools/mutation-anchor.test.ts | 275 ++++++++++ tools/mutation-anchor.ts | 477 ++++++++++++++++++ tools/mutation-teeth.ts | 340 ++++++------- 6 files changed, 1179 insertions(+), 177 deletions(-) create mode 100644 docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md create mode 100644 docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md create mode 100644 tests/tools/mutation-anchor.test.ts create mode 100644 tools/mutation-anchor.ts diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md new file mode 100644 index 00000000..b0b8ed70 --- /dev/null +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -0,0 +1,184 @@ +# A tooth is an operator over a symbol, not a copy of a line + +[中文](2026-09-24-mutants-are-derived-not-anchored.zh-CN.md) + +**Status:** proposed +**Approved:** explicit +**Relates to:** [Tests do not need a filesystem](../implemented/2026-09-20-tests-need-no-filesystem.md), [The checks read a live mutant](../../postmortem/0003-checks-read-a-live-mutant.md), [Bound agent verification as one run](../implemented/2026-09-23-verification-whole-run-deadline.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md), [Mechanism, not policy](2026-09-21-mechanism-not-policy.md) + +## Problem + +Every rule this repository has decided to protect is protected the same way: a named mutant in +`tools/mutation-teeth.ts` that replaces a byte range with a hand-written replacement, plus the name of +the case that must fail when it does. That evidence is honest - a mutant is the only form that speaks +about the implementation's freedom to be wrong - and it is also the most fragile artifact in the tree, +because **the anchor is the mutant's identity**. It is a copy of a line, so the line moving, being +inlined, or having its message text extracted retires the tooth. + +Measured on 2026-09-24, at 149 mutants over 22 targets: + +| What the anchor points at | Mutants | Share | Already scoped to a member | +| ----------------------------------------------------- | ------- | ----- | -------------------------- | +| A guard or predicate term neutralised (G) | 64 | 43% | 30 | +| A value, argument, index or callee substituted (V) | 58 | 39% | 26 | +| A statement or call removed (R) | 11 | 7% | 6 | +| Code inserted: a statement, branch or second copy (I) | 8 | 5% | 4 | +| An expression replaced wholesale (E) | 8 | 5% | 5 | + +Two readings decide this record. **89% of the teeth (133 of 149) are a site plus one of a handful of +operators** - disable a condition, drop a term, substitute an argument, remove a statement - so the +operator is the intent and the site is incidental; storing the site as text is a choice, not a +requirement. And **78 of 149 (52%) have no structural scope at all**: they are a byte fragment matched +anywhere in the file, which is exactly the shape that dies first. Both failures this session were of +that shape: a statement whose inner call was inlined, and a refusal whose message text was extracted +into a `subject` expression. A third form showed up beside them: two teeth whose named case no longer +catches them (`fusion-continues-from-an-unverified-answer`, `next-is-not-the-head-of-the-ordered-candidates` +read as "caught by the suite, not the named case"), which is the same coupling one layer out - the +anchor held, the sentence "this case is the one that fails" went stale. + +A measurement taken while planning this record also has to be reported, because it removed an option: +the three-way split "derivable / expressible as input / needs an internal perturbation" **does not +hold**. A targeted mutant is by definition _observable from outside_ - that is what "caught" means - so +"input-expressible" is true of all 149. The three cases predicted to need an internal perturbation all +had an external observation: the status read path opening a writable handle is observed by the store +file not changing, a second copy of the acceptance rule by a verdict whose digest does not match the +artifact, and a transaction wrapper by a plan whose third task is refused while the first two stay +frozen. So the replacement direction is not "input checks instead of mutants"; it is **derive the +mutant instead of storing it**, with input-side checks as the form for the rules where a contrast or an +enumeration already says it better. + +## Proposal + +**A mutant's durable part is its name, its operator, its selector and the case that must fail. The site +and the replacement bytes are computed from the syntax tree on every run.** + +1. **Six operators over a named selector**, resolved inside a member: `condition-never` (the condition + the selector identifies never holds), `condition-holds` (it always holds - the mirror, because a rule + written as `return a && b` says the condition is true and replacing it with `false` would reverse the + rule instead of removing it), `neutralize-term` (one term of that condition becomes its identity - + `true` under `&&`, `false` under `||`), `replace-argument` (argument _n_ of a call becomes the + declared fragment), `replace-property` (the value of a named object property becomes it), and + `drop-statement` (the statement the selector identifies is removed with its line). Two rules keep + selectors honest: a selector that matches more than one site, or a fragment that fits two candidates, + is refused rather than guessed at; and the two whole-condition operators refuse a fragment that + names only part of a condition, because replacing all of it would silently widen the mutant into + "every reason this rule has" - a term has its own operator. Where a rule needs none of them, a + `within`-scoped minimal fragment stays available, and "minimal" is the point: no message text, no + sibling argument, no whole statement where a term does. +2. **Move the 133 G/V/R teeth to the derived form**, target by target, beginning with + `src/integration/ooo-execution.ts` as the pilot: it holds 24 teeth over two target entries (16 fusion- + legality and speculation, 8 session fusion), the same mutant names, the same `expect` cases, and the + same 24 of 24 caught after conversion - 19 of them derived, 5 staying hand-written and named below - + plus a real refactor in that file that leaves every derived site applying where a byte anchor + retires. + The five residues in the pilot file, each a single expression or value rather than a statement or a + message: `selection-ignores-a-withdrawn-acceptance` (a whole `const` initializer replaced), + `the-budget-is-not-cut-from-the-startable-set` (a returned expression's call removed), + `next-task-is-not-the-head-of-the-legal-set` (an index `[0]` becomes `[1]`), + `speculation-guesses-several-facts-at-once` (`length === 1` becomes `>= 1`), and + `a-guess-with-no-evidence-publishes` (an inserted branch plus a rewritten message). +3. **What a derived form cannot express stays hand-written and says so.** The 16 I/E teeth - insertions, + wrappers, a second copy of a rule - are the ones whose anchor is a _place where code must not + appear_, so they are the last to move and the first to be reconsidered: where the rule already has a + check whose form is relational or enumerative, the tooth is retired and the check is named in its + place. +4. **A check that a tooth still applies becomes part of the static contract**: an anchors-only pass that + resolves every mutant, without running a suite, in single-digit seconds, failing when a site cannot be + resolved, resolves more than once, or when a target claims more teeth than it can apply. The + `ci-and-tests` route then lists it beside the other atomic checks, and the whole-run deadline + ([bound agent verification](../implemented/2026-09-23-verification-whole-run-deadline.md)) is a + constraint on it: the pass is cheap, the full sweep stays out of the gate. +5. **The evidence standard is clarified, not loosened.** "How a row earns `proven`" already says a test + that fails when the rule is broken; the clarification is that the demonstration may be a code + perturbation (derived or hand-written) or an input-side case whose assertion contrasts two inputs or + enumerates a bounded space. A row that used to name a tooth may name a contrast check instead, in the + same commit that retires the tooth. + +## Plan + +1. This record, plus the tool's operator vocabulary and its resolver, proven by tests over source + strings rather than by a filesystem. **Landed:** the resolver is `tools/mutation-anchor.ts` (it held + no state, so it moved out of the sweep script and the tests can call it), with 16 cases over source + strings and no filesystem access. +2. The pilot target converted and swept: same names, same cases, same catches; then a refactor inside it + that demonstrates a derived tooth surviving what a byte anchor did not. **Landed**, with one + correction: the demo first refactor was a rename plus a condition lifted into `const closed = spent || +severalWaits`, and the prediction made before running it - one dead anchor, seven derived teeth + surviving - was **wrong**: two sites died, the hand anchor and the derived + `a-live-claim-does-not-block-selection`, because a condition bound to a name was not a decision + position for the selector. Lifting a condition into a local is a refactor a maintainer makes, so the + selector was widened to include a variable's initializer; the corrected prediction (only the hand + anchor dies) then held exactly: `anchors: 148 of 149 resolve`, one failure, and the file restored + byte-identically. +3. The anchors-only pass, wired into the static contract with its route and design updates. **Not + landed yet:** the pass exists (`--anchors-only`, 149 of 149 resolving in 0.6 s, no suite run, nothing + written) but is not yet one of the route's atomic checks. +4. Retirement pass over the ~18-20 teeth whose rule already has a relational or enumerative check, and + over the I/E teeth that can be replaced; the ledger's `proven` sentence updated in the same commit. + **Not started.** + +## Acceptance criteria + +- A mutant declared as an operator over a selector stores no byte range, and the tool resolves the site + and the replacement from the syntax tree. +- A refactor that inlines, renames, reflows or extracts code around a derived site leaves the tooth + applying and still caught by its named case; a byte-anchored tooth in the same place does not survive + the same refactor (the pilot names both). +- The anchors-only pass runs without suites in under ten seconds on this tree, fails on an unresolvable + or ambiguous site, reports the claimed-versus-applicable gap, and leaves the tree byte-identical. +- Adding it to the static contract keeps a full `agent:verify` inside its 150-second budget. +- Teeth retired in favour of a contrast or enumeration are named in the commit that retires them, and + the row they proved names that check afterwards. +- The residue of hand-written fragments is enumerated by name, and every one of them is a single + expression or term rather than a statement or a message. + +## Risks + +- **A derived selector is still scoped to a decision position.** Measured in the pilot: a condition lifted + out of its `if` into `const closed = spent || severalWaits` retired the tooth - the selector looked for + a condition in a decision position, and a name bound for a decision made a line later was not one + until it was added. A tooth can therefore still die when its rule is rewritten into a _different_ + shape. The difference is that it now dies visibly: the anchors-only pass names the tooth and the + fragment it could not resolve, which is exactly the failure that went unnoticed before. +- **A wrong selector mutates the wrong node.** A derived site is resolved by the tool, so a selector that + matches another call can produce a mutant nobody intended. Mitigation: exactly one match or refusal, + which the tool already does for `ast`, plus the existing two readings (a clean run must pass, and the + named case must be the one that fails - a mutant that does not compile is reported as "by the suite, + not the named case" rather than as a caught tooth). +- **Derivation hides the perturbation from the reader.** A reader of a `from`/`to` pair sees the change; + a reader of `condition-never(within: judgeTaskBoardEntry, condition: existing.deliveredBy)` has to + know the operator. Mitigation: the operators are few and fixed, they are documented beside the type, + and the tool prints the bytes it wrote when asked. +- **Retiring a tooth can lose a named case.** A relational check may catch a class where the tooth caught + a specific wrong behaviour, so the row would claim more than it shows. Mitigation: a retirement names + the check that replaces it, in the ledger row and in the commit. +- **The pass could be trusted instead of run.** An anchors-only pass proves the site still resolves, not + that the mutant is still caught; a tooth whose named case went stale still passes it. Mitigation: the + pass reports the stale-`expect` reading the sweep already produces, and the full sweep remains the + standing rule before a push. +- **Conversion is evidence surgery.** Every conversion touches the artifact that says the code is + protected, so a mistake reduces coverage silently. Mitigation: one target per commit, the same names + and cases, the catch results recorded before and after, and the post-mortem rule about a live mutant + (status the file under test before any run). + +## Alternatives considered + +- **Keep hand-written anchors and make them smaller.** Rejected as the whole answer: it is a discipline, + not a mechanism, and the discipline is what 52% of the teeth currently violate. Kept as a rule for the + residue. +- **Anchor by AST node only, keeping byte replacements** (today's `ast.within` plus a text pair). + Rejected as insufficient: measured, 71 of 149 teeth already have that, and two teeth died inside it + this session, because the _replacement_ and the _fragment_ are still copies of lines. +- **Replace mutants with input-side checks wholesale.** Rejected on the measurement above: it does not + discriminate (every caught mutant is externally observable), and it would drop the only evidence that + speaks about an unforeseen implementation slip rather than a declared violation. +- **Insert the mutation into the code behind a runtime switch** (mutant schemata, `mutation_active("...")` + guards). Rejected for this repository: the guard is real code in `src/`, and a switched-off mutation + path is a policy word in the mechanism layer, which [mechanism, not + policy](2026-09-21-mechanism-not-policy.md) forbids. +- **Auto-generate mutants from operators over the whole tree** (what PIT, StrykerJS and cargo-mutants + do). Rejected as the form here: it would replace named evidence with a score, and the ledger needs a + named tooth per row. The derived form keeps the operator idea and the naming. +- **Leave the sweep out of every gate and rely on the standing rule.** Rejected: the measured case for + this record is that nobody noticed two teeth had stopped biting, and a rule that depends on being + remembered is the shape this repository already replaced elsewhere. diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md new file mode 100644 index 00000000..8adb9c4f --- /dev/null +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -0,0 +1,70 @@ +# 牙是算子加符号,不是一行文本的副本 + +[English](2026-09-24-mutants-are-derived-not-anchored.md) + +**Status:** proposed +**Approved:** explicit +**Relates to:** [测试不需要文件系统](../implemented/2026-09-20-tests-need-no-filesystem.md)、[检查读到的是一个活着的 mutant](../../postmortem/0003-checks-read-a-live-mutant.md)、[把 agent 验证限制成一次运行](../implemented/2026-09-23-verification-whole-run-deadline.md)、[契约的义务](../../design/task-unit-semantics-obligations.md)、[机制,不是策略](2026-09-21-mechanism-not-policy.md) + +## 问题 + +这个仓库决定要保护的每一条规则,都用同一种方式保护:`tools/mutation-teeth.ts` 里一颗具名 mutant——把一段字节替换成手写的替换文本,外加“哪一个用例必须因此失败”的名字。这条证据是诚实的(只有 mutant 这种形式能谈论“实现自己有可能错”),同时它也是树里最脆的产物,因为**锚点就是 mutant 的身份**:它是某一行的副本,于是那一行搬家、被内联、或者消息文本被抽走,这颗牙就退役了。 + +2026-09-24 在 149 颗牙、22 个目标上实测: + +| 锚点指向什么 | 颗数 | 占比 | 其中已有成员级作用域 | +| ------------------------------------- | ---- | ---- | -------------------- | +| 守卫或谓词项被中和(G) | 64 | 43% | 30 | +| 值、实参、下标或被调者被替换(V) | 58 | 39% | 26 | +| 语句或调用被删(R) | 11 | 7% | 6 | +| 插入代码:语句、分支或第二份规则(I) | 8 | 5% | 4 | +| 整段表达式被替换(E) | 8 | 5% | 5 | + +两个读数为这份记录定性。**89% 的牙(149 里的 133)就是“一个位置 + 少数几种算子”**——让条件永不成立、去掉一项、替换一个实参、删掉一条语句——也就是说算子是意图,位置只是陪衬;把位置存成文本是一种选择,而不是必需。而 **149 里有 78 颗(52%)根本没有结构作用域**:一个在全文件里匹配的字节片段,正是死得最快的那种形状。本轮两次失效都是这个形状:一条语句里被内联掉的内部调用,以及一段拒绝消息被抽成 `subject` 表达式。旁边还出现了第三种形状:两颗牙的点名用例已经不再是抓住它们的那个用例(`fusion-continues-from-an-unverified-answer`、`next-is-not-the-head-of-the-ordered-candidates` 被读成“caught by the suite, not the named case”)——同一类耦合往外一层:锚点还在,“这个用例就是失败的那个”这句话过期了。 + +规划这份记录时做的另一次测量也必须写下来,因为它排除了一个选项:三分法“可派生 / 可用输入式反例 / 必须扰动内部”**不成立**。一颗被抓的 mutant 按定义就是**从外部可观察**的(这就是“被抓”的含义),所以“可用输入式反例”对 149 颗全都成立。三颗预测“必须扰动内部”的牙都找得到外部观察:状态读路径开了可写句柄,用“store 文件没有变化”观察;acceptance 规则的第二份副本,用一个 digest 与产物不匹配的裁决观察;事务包装,用“计划里第三个任务被拒时前两个仍冻结”观察。所以替换方向不是“用输入检查替代 mutant”,而是**派生 mutant,而不是把它存下来**;输入式检查留给那些“对比或穷举本来就说得更清楚”的规则。 + +## 提案 + +**一颗牙的耐久部分是它的名字、它的算子、它的选择器,以及必须失败的那个用例。位置与替换字节每次从语法树算出来。** + +1. **六个作用在具名选择器上的算子**,作用域限定在一个成员内:`condition-never`(选择器指到的那条条件永不成立)、`condition-holds`(它始终成立——镜像的那个,因为写成 `return a && b` 的规则说的是“这条条件为真”,把它换成 `false` 是把规则反过来,而不是去掉)、`neutralize-term`(该条件里的一项变成它的恒等元——`&&` 下为 `true`,`||` 下为 `false`)、`replace-argument`(某个调用的第 _n_ 个实参换成声明的片段)、`replace-property`(具名对象属性的值换成声明的片段)、`drop-statement`(选择器指到的那条语句连行带缩进被删掉)。两条规则让选择器保持诚实:匹配到多于一处、或者片段同时命中两个候选的,一律拒绝而不是猜;两个“整条件”算子拒绝只命名条件一部分的片段,因为整段替换会把 mutant 悄悄放大成“这条规则的每一条理由”——一项有它自己的算子。规则用不上这些时,仍可用“成员作用域 + 最小片段”的旧形式;而**“最小”是要点**:不要消息文本、不要同级实参、一项能表达的不要写成整条语句。 +2. **把 133 颗 G/V/R 牙逐目标改成派生形式**,从 `src/integration/ooo-execution.ts` 开始试点:该文件在两条目标条目里共 24 颗(16 颗融合合法性推演、8 颗会话融合);同样的名字、同样的 `expect` 用例,转换后同样 24/24 被抓——其中 19 颗派生、5 颗保留手写(下面具名)——然后在该文件里做一次真重构,证明派生出来的位置都还在,而字节锚点会在这次重构里退役。 + 试点文件里那 5 颗残余,每一颗都是单个表达式或单个值,而不是整条语句或消息:`selection-ignores-a-withdrawn-acceptance`(整个 `const` 初始化式被替换)、`the-budget-is-not-cut-from-the-startable-set`(被返回表达式里的调用被删)、`next-task-is-not-the-head-of-the-legal-set`(下标 `[0]` 变 `[1]`)、`speculation-guesses-several-facts-at-once`(`length === 1` 变 `>= 1`)、`a-guess-with-no-evidence-publishes`(插入一个分支并重写消息)。 +3. **派生形式表达不了的,保留手写,并且说明为什么。** 16 颗 I/E 牙——插入、包装、第二份规则——它们的锚点是“某段代码不该出现的地方”,所以它们最后才搬,也最先被重新考虑:凡是规则已经有一个关系式或穷举式检查的,就退役这颗牙,并在它原来的位置上点名那个检查。 +4. **“牙是否还咬得住”成为静态契约的一部分**:一趟只做解析、不跑用例的 anchors-only 检查,个位秒数内完成,在位置无法解析、解析出多于一处、或者某个目标声称的颗数多于能应用的颗数时失败。`ci-and-tests` 路由随后把它与其它原子检查并列;而[整轮预算](../implemented/2026-09-23-verification-whole-run-deadline.md)是它的约束:这趟检查必须便宜,全量 sweep 仍然不进闸门。 +5. **证据标准是被澄清,不是被放宽。** “一行如何挣得 `proven`”原文已经写的是“一个在规则被破坏时会失败的检查”;澄清之处在于:这个证明可以是一次**代码扰动**(派生的或手写的),也可以是一个**输入侧用例**,其断言对比两组输入、或穷举一个有界空间。某一行以前点名一颗牙,现在可以改为点名一个对比检查——在同一提交里退役那颗牙。 + +## 计划 + +1. 本文,加上工具的算子词汇表与解析器;用“对源码字符串”而不是对文件系统的测试来证明它。**已落地:**解析器是 `tools/mutation-anchor.ts`(它不持有任何状态,所以从 sweep 脚本里搬出来,测试可以直接调用),16 个用例全部作用在源码字符串上,不碰文件系统。 +2. 试点目标转换并重跑:同样的名字、同样的用例、同样的抓法;随后在该文件内做一次重构,证明派生出来的牙活过了字节锚点活不过去的那次改动。**已落地,带一处更正:**演示用的那次重构是“改一个局部变量名 + 把条件抬成 `const closed = spent || severalWaits`”,而运行前写下的预测——死一颗锚点、七个派生位置存活——**是错的**:死了两处,手写锚点和派生的 `a-live-claim-does-not-block-selection`,因为“绑定到名字上的条件”当时不算选择器眼里的一次决策位置。把条件抬成局部常量是维护者真会做的重构,于是选择器扩到了变量初始化式;更正后的预测(只有手写锚点会死)随后精确成立:`anchors: 148 of 149 resolve`,一处失败,文件逐字节还原。 +3. anchors-only 这趟检查接入静态契约,并同步路由与设计文档。**尚未落地:**这趟检查已经存在(`--anchors-only`,149/149 在 0.6 秒内解析完,不跑用例,不写任何字节),但还没有成为该路由的一条原子检查。 +4. 对“规则已有关系式/穷举式检查”的那约 18-20 颗,以及能被替代的 I/E 牙,做退役;台账里 `proven` 那句话在同一提交里更新。**尚未开始。** + +## 验收标准 + +- 以“算子 + 选择器”声明的 mutant 不存任何字节区间,位置与替换文本由工具从语法树解析。 +- 一次把周围代码内联、改名、重排或抽出的重构之后,派生位置仍能应用、且仍被点名用例抓住;同一处的字节锚点则不然(试点要把两者都点出来)。 +- anchors-only 检查在本地树上不跑用例、十秒内完成;无法解析或解析出多于一处时失败;报告“声称 vs 可应用”的差额;并让工作树逐字节不变。 +- 把它加进静态契约之后,一次完整的 `agent:verify` 仍留在 150 秒预算之内。 +- 因对比或穷举而退役的牙,在退役它的那次提交里被点名,它原来保护的那一行之后点名那个检查。 +- 手写片段的残余按名字列出,且每一颗都是单个表达式或单项,而不是一条语句或一段消息。 + +## 风险 + +- **派生选择器仍然被限定在“一次决策位置”上。** 试点实测:把条件从它的 `if` 里抬成 `const closed = spent || severalWaits` 会直接让那颗牙退役——选择器找的是决策位置上的条件,而“为下一行的决策先绑个名字”在扩展之前不算一处。所以当规则被改写成*另一种形状*时,牙仍可能死掉。区别在于它现在是**可见地**死:anchors-only 那趟检查会点出牙的名字和它解析不到的片段,而这恰恰是以前没人注意到的那种失效。 +- **选择器写错会改到错误的节点。** 派生位置由工具解析,所以匹配到另一个调用的选择器会产出一颗没人想要的 mutant。缓解:恰好一处否则拒绝(工具对 `ast` 已经如此),加上原有的两重读数(clean run 必须通过;必须是点名用例失败——编译不过的 mutant 会被读成“by the suite, not the named case”,而不是被算作抓到的牙)。 +- **派生把扰动藏起来了。** 读 `from`/`to` 的人看得见改动;读 `condition-never(within: judgeTaskBoardEntry, condition: existing.deliveredBy)` 的人得知道这个算子。缓解:算子少而固定、写在类型旁边;工具在被要求时打印它实际写入的字节。 +- **退役一颗牙可能丢掉一个点名用例。** 关系式检查抓的可能是一类,而那颗牙抓的是一个具体的错行为,于是那一行会宣称超过它所展示的。缓解:退役时在台账行与提交里点名替代它的检查。 +- **这趟检查可能被信任而不是被运行。** anchors-only 只证明位置还能解析,不证明牙还被抓住;点名用例过期的那类牙照样通过。缓解:这趟检查报告 sweep 已经会产出的“点名用例过期”读数,而全量 sweep 仍是推送前的常设规则。 +- **转换是在证据上动手术。** 每次转换都碰到“这份代码被保护着”的凭据本身,所以一个错误会静默地减少覆盖。缓解:一个提交一个目标;名字与用例不变;转换前后都记录抓取结果;并遵守“活着的 mutant”那条验尸规则(任何一次运行之前先看被测文件的状态)。 + +## 考虑过的替代方案 + +- **保留手写锚点,只把它们写小一点。** 作为全部答案被拒绝:那是一条纪律,不是一种机制,而 52% 的牙今天违反的正是这条纪律。作为残余部分的规则保留。 +- **只用 AST 定位,替换仍是字节**(即今天的 `ast.within` 加一对文本)。作为不够被拒绝:实测已有 71/149 是这样,本轮仍有两颗牙死在它里面,因为**替换文本**和**片段**依旧是行的副本。 +- **整体用输入侧检查替换 mutant。** 按上面的测量拒绝:它不具区分度(每颗被抓的 mutant 都是外部可观察的),而且会丢掉唯一一种谈论“未预料到的实现偏差”而不是“声明的违反”的证据。 +- **把变异插进代码、用运行时开关激活**(mutant schemata,`mutation_active("...")` 这类守卫)。本仓拒绝:守卫是 `src/` 里的真代码,而一条被关掉的变异路径就是机制层里的策略词,[机制,不是策略](2026-09-21-mechanism-not-policy.md) 已经禁止。 +- **全树按算子自动生成 mutant**(PIT、StrykerJS、cargo-mutants 的做法)。作为本仓的形式被拒绝:它会把具名证据换成一个分数,而台账每一行需要一个具名的牙。派生形式保留了算子这个主意,也保留了具名。 +- **把这趟检查留在闸门之外,只靠常设规则。** 拒绝:这份记录的实测理由就是没人发现两颗牙早已不再咬人,而一条依赖“有人记得”的规则,是仓库在别处已经替换掉的形状。 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index 5e97c0cb..caf0a454 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -19,6 +19,9 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run test:product` -> 1457 pass, 0 fail, exit 0. **A row that used to sit here said "one full run first reported a single failure under parallel load, then passed 1433/1433 on re-run; recorded as flaky, not fixed" - that label was wrong, and it hid a product defect.** The failure was `demoteMemory: demotes LTG memory to STG`, and it was a clock boundary: a memory written with `valid_from` a moment _after_ the reading connection's `strftime('now')` read as not current (measured 2 of 3000 write-then-read rounds, stamp `…38.468Z` against `now` `…38.467Z`). Fixed by a named grace in `src/core/store/clock.ts`, pinned by `tests/core/store/current-value-window.test.ts` (6 cases) and 4 named mutants, decided in [the clock-grace record](../decisions/implemented/2026-09-18-clock-grace-window.md), recorded as [post-mortem 0004](../postmortem/0004-flaky-was-a-clock-boundary.md). The count moved 1446 -> 1457 with the fixed window and the cases added since - `npm run mutation:teeth` -> 136 of 136 caught by the named test, 20 of 20 targets restored byte-identically, exit 0. Run on 2026-09-19 as four lanes, one sweep per tree, using `git worktree add --detach` on the same commit for three of them: a sweep is sequential _within_ a tree because its mutants substitute into the same file, and parallel across trees, where each lane also gets the isolation property that no lane's suites can read another lane's mutant. Lanes: 42 of 42 (`ooo-board`, `task-coordinator`), 41 of 41 (`base`, `ooo-execution`'s 17, `task-semantics-interleavings`), 42 of 42 (thirteen small targets) and 11 of 11 (`plan-driver`) - the last serialised into its own lane because its suite has a 25 s case and does real candidate verification (~92 s per run, against ~2 s for the cheap suites). **That lane's cost has since changed**: the arms' checks are data checks now, so at `c01d3fe` its clean run is 5.5 s and its five mutants are caught by the case each names in 0.8-1.2 s on the same target; the numbers in this paragraph describe the 2026-09-19 instrument. Three things the run itself taught, all fixed and pinned afterwards: the lock's `live` flag was never written on substitution (a multi-hunk edit failed as a whole and only the restore half was reapplied), so the field lied about a running sweep; `NODE_TEST_CONTEXT` inherited when a sweep is started from inside a `node --test` process made the nested runner exit 0 having run no test at all, which the harness reported as "the suite passed" and which turned every mutant of that target into a false "not caught"; and the refusal in `agent:verify` fired on `--dry-run` too, which made two of the verifier's own tests fail while a sweep held the tree (a dry run reads the plan and the route config, not the mutated file, so it is exempt now). The lane that reported a clean-run failure (`tools/agent-verify.ts`) had found the last of those three. A sweep is also refused while any lock is present, including one whose owner died, because a killed sweep leaves its mutant in the target (post-mortem 0003). Interruption note: two lane processes were killed by the console that launched them and were relaunched; the JSON each run writes at its end survived even when the buffered stdout summary was lost, so the lane results above were read from those files rather than from stdout. (the retirement pass added the driver's interleaving mutant; was 111 of 111 before this pass: `src/integration/task-semantics-interleavings.ts` gained three budget mutants and `evals/ooo-execution/plan-driver.ts` three for the per-unit checks, the canned worker and the parent composition). How these runs are scheduled (scoped during a change, full before a push, detached with a collected result) is a standing rule of the repository now, in [`skills/repo-development/SKILL.md`](../../skills/repo-development/SKILL.md) with its measured costs in [the decision](../decisions/implemented/2026-09-18-detached-long-checks.md) - `npm run complexity:gate` -> exit 0, 18 methods above 15 unchanged from baseline. It caught the E pass's first version (adding the advisers option pushed `BoardAdmission`'s constructor to 17, so the options check moved into `admissionAdvice()`) rather than the threshold being raised; the slot pass added its option check the same way (`admissionSlots`), and the ordered-set/`publishReady` reads stayed under it. +- `node --experimental-strip-types --test tests/tools/mutation-anchor.test.ts` -> 16 pass, 0 fail, exit 0 (2026-09-24: the resolver's six operators and its refusals, over source strings, no filesystem) +- `npm run mutation:teeth -- --anchors-only` -> `anchors: 149 of 149 resolve, over 23 targets`, exit 0 in 0.6 s (2026-09-24: every tooth still applies where it is aimed; one anchor is re-taken through whitespace normalization, `the-completion-does-not-bind-the-verdict-to-the-bytes`) +- `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) - `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. - `node --experimental-strip-types --test --test-concurrency=4 tests/integration/ooo-ordinary-failure.test.ts tests/integration/ooo-managed-fence.test.ts tests/integration/ooo-read-paths-agree.test.ts tests/integration/ooo-round-query.test.ts tests/integration/ooo-task-tables.test.ts` -> 5, 3, 1, 2 and 4 pass, 0 fail, exit 0 @@ -27,6 +30,13 @@ such a mutation is registered, the mutant's name is given, because a test that c description rather than a pin. Every mutant name below was read from `tools/mutation-teeth.ts`, not recalled. Rows follow the design's own order, which is the work order. +A tooth is a **name, an operator and a selector** - not a copy of a line: the site and the replacement +bytes are resolved from the syntax tree on every run, so a rename or a reflow does not retire the pin +while the rule still stands ([the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md), +landed so far on the pilot file). `npm run mutation:teeth -- --anchors-only` answers whether every +anchor still applies - under a second, no suite run, nothing written - because a tooth that stopped +matching its rule is otherwise only visible as a check that stopped counting. + ## Where a proof lives: two evidence bases, and neither stands for the other The rows cite test files, and those files split into two bases that **do not share storage** and whose diff --git a/tests/tools/mutation-anchor.test.ts b/tests/tools/mutation-anchor.test.ts new file mode 100644 index 00000000..eb392c91 --- /dev/null +++ b/tests/tools/mutation-anchor.test.ts @@ -0,0 +1,275 @@ +/** + * A derived mutant: the operator says what to do, the selector says where, and the bytes are computed + * from the source rather than stored beside it. + * + * The rule this file pins is that a tooth's identity is not its text. A byte anchor dies the first time + * the formatter, a rename or a lifted condition changes the bytes around the site - and it dies + * silently, which is how two teeth stopped pointing at their rule without a build noticing. So the + * selectors below are asserted twice: once on the source as written, and once on the same code after a + * refactor that preserves it. Nothing here touches the filesystem. + */ +import assert from "node:assert/strict"; +import test from "node:test"; + +import { locate, matchText, type Mutant } from "../../tools/mutation-anchor.ts"; + +const SOURCE = `function selection(tasks, slots) { + const spent = slots < 1; + const ready = (task) => current(task) && !task.cancelled && !task.waiting; + if (spent || severalWaits) return answer([], 0, "closed"); + const ids = (pending) => + tasks.filter((task) => !task.claimed && ready(task)).map((task) => task.id); + causes.set(task.id, own.length > 0 ? own : []); + return { outcome: "discard", sessionReusable: false }; +} +`; + +/** The bytes a mutant would replace: asserting the selection beats asserting offsets. */ +function selected(source: string, mutant: Mutant): string { + const site = locate(source, mutant); + assert.ok(!("reason" in site), `expected a site, got ${"reason" in site ? site.reason : ""}`); + return source.slice(site.start, site.end); +} + +function refusal(source: string, mutant: Mutant): string { + const site = locate(source, mutant); + assert.ok( + "reason" in site, + `expected a refusal, got a site: ${source.slice(site.start, site.end)}`, + ); + return site.reason; +} + +const named = (name: string, derive: Mutant["derive"], extra: Partial = {}): Mutant => ({ + name, + derive, + expect: "unused here", + ...extra, +}); + +test("a guard that never holds is the condition, and nothing around it", () => { + const above = `function selection(task) {\n if (!task.accepted) return false;\n return true;\n}\n`; + assert.equal( + selected( + above, + named("x", { within: "selection", operator: "condition-never", condition: "!task.accepted" }), + ), + "!task.accepted", + ); +}); + +test("a rule written as a return is a condition too, and it can hold or never hold", () => { + const above = `function selection(task) {\n return !task.cancelled && !task.waiting;\n}\n`; + const whole = "!task.cancelled && !task.waiting"; + for (const operator of ["condition-holds", "condition-never"] as const) + assert.equal( + selected(above, named("x", { within: "selection", operator, condition: whole })), + whole, + ); +}); + +test("an operator that replaces a whole condition refuses a fragment that names only part of it", () => { + const reason = refusal( + SOURCE, + named("x", { within: "selection", operator: "condition-holds", condition: "!task.cancelled" }), + ); + assert.match(reason, /names part of selection's condition.*use neutralize-term for a term/); +}); + +test("a term's identity is the literal its operator absorbs: true under &&, false under ||", () => { + const term = (condition: string, text: string): Mutant => + named("x", { within: "selection", operator: "neutralize-term", condition, term: text }); + assert.equal( + selected(SOURCE, term("!task.cancelled && !task.waiting", "!task.waiting")), + "!task.waiting", + ); + assert.equal(selected(SOURCE, term("spent || severalWaits", "severalWaits")), "severalWaits"); +}); + +test("a term's identity is read through the brackets that close it", () => { + const above = `function selection(a, b) {\n return a && (b || !b);\n}\n`; + const site = locate( + above, + named("x", { within: "selection", operator: "neutralize-term", condition: "b", term: "!b" }), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "!b"); +}); + +test("a term the code does not join with a logical operator is refused, not guessed", () => { + const above = `function selection(a, b) {\n if (a === b) return false;\n return true;\n}\n`; + const reason = refusal( + above, + named("x", { + within: "selection", + operator: "neutralize-term", + condition: "a === b", + term: "a", + }), + ); + assert.match(reason, /cannot tell which operator joins the term/); +}); + +test("an argument is selected by the call it belongs to and its position", () => { + const site = selected( + SOURCE, + named( + "x", + { within: "selection", operator: "replace-argument", call: "causes.set", arg: 1 }, + { to: "[]" }, + ), + ); + assert.equal(site, "own.length > 0 ? own : []"); +}); + +test("a call that appears twice in the member is refused, and so is an argument that is not there", () => { + const twice = `function selection(a) {\n causes.set(a.id, 1);\n causes.set(a.other, 2);\n}\n`; + assert.match( + refusal( + twice, + named( + "x", + { within: "selection", operator: "replace-argument", call: "causes.set", arg: 1 }, + { to: "[]" }, + ), + ), + /matched 2 sites/, + ); + assert.match( + refusal( + SOURCE, + named( + "x", + { within: "selection", operator: "replace-argument", call: "causes.set", arg: 3 }, + { to: "[]" }, + ), + ), + /has no argument 3/, + ); +}); + +test("a property is selected by name, and by the literal that holds it when the name repeats", () => { + const property: Mutant["derive"] = { + within: "selection", + operator: "replace-property", + property: "sessionReusable", + }; + assert.equal(selected(SOURCE, named("x", property, { to: "true" })), "false"); + const twice = `function selection() {\n if (a) return { outcome: "wait", sessionReusable: true };\n return { outcome: "discard", sessionReusable: false };\n}\n`; + assert.match(refusal(twice, named("x", property, { to: "true" })), /matched 2 sites/); + assert.equal( + selected(twice, named("x", { ...property, in: 'outcome: "discard"' }, { to: "true" })), + "false", + ); +}); + +test("a dropped statement takes its line and its indentation, and refuses to guess otherwise", () => { + const above = `function selection(task) {\n causes.push("stale-input");\n return task;\n}\n`; + // The indentation and the newline go with it: what is left is a hole where the line was. + assert.equal( + selected( + above, + named("x", { within: "selection", operator: "drop-statement", statement: "causes.push" }), + ), + ' causes.push("stale-input");\n', + ); + const shared = `function selection(task) {\n if (task.accepted) return task;\n return null;\n}\n`; + assert.match( + refusal( + shared, + named("x", { within: "selection", operator: "drop-statement", statement: "return task;" }), + ), + /does not begin its line/, + ); +}); + +test("the fragment is matched inside the selected condition, not anywhere in the member", () => { + const above = `function selection(task) {\n causes.push("stale-input");\n if (!task.accepted) return false;\n return true;\n}\n`; + assert.match( + refusal( + above, + named("x", { within: "selection", operator: "condition-never", condition: '"stale-input"' }), + ), + /0 conditions in selection mention/, + ); +}); + +test("a member the selector cannot tell apart is refused: none, several, or a fragment that fits two", () => { + const absent = refusal( + SOURCE, + named("x", { within: "elsewhere", operator: "condition-never", condition: "spent" }), + ); + assert.match(absent, /member elsewhere matched 0 members/); + const twice = `function selection(a) {\n if (!a) return 1;\n}\nfunction selection(b) {\n if (!b) return 2;\n}\n`; + assert.match( + refusal( + twice, + named("x", { within: "selection", operator: "condition-never", condition: "!b" }), + ), + /member selection matched 2 members/, + ); + assert.match( + refusal( + SOURCE, + named("x", { within: "selection", operator: "condition-never", condition: "task" }), + ), + /conditions in selection mention task/, + ); +}); + +test("when a fragment sits inside two decision positions, the innermost one is the site", () => { + const nested = named("x", { + within: "selection", + operator: "condition-holds", + condition: "!task.claimed && ready(task)", + }); + assert.equal(selected(SOURCE, nested), "!task.claimed && ready(task)"); +}); + +test("the same tooth applies after the code around it is refactored", () => { + const renamed = SOURCE.replaceAll("ready(task)", "isReady(task)"); + const lifted = SOURCE.replace( + " if (spent || severalWaits) return answer", + " const closed = spent || severalWaits;\n if (closed) return answer", + ); + const tooth = named("x", { + within: "selection", + operator: "neutralize-term", + condition: "spent || severalWaits", + term: "spent", + }); + for (const source of [SOURCE, renamed, lifted]) assert.equal(selected(source, tooth), "spent"); + // The byte anchors the same sites used to carry do not survive either edit - which is the point. + const renamedAnchor: Mutant = { + name: "old", + from: " tasks.filter((task) => !task.claimed && ready(task)).map((task) => task.id);", + to: " tasks.filter((task) => !task.claimed).map((task) => task.id);", + expect: "x", + }; + const liftedAnchor: Mutant = { + name: "old", + from: ' if (spent || severalWaits) return answer([], 0, "closed");', + to: ' if (severalWaits) return answer([], 0, "closed");', + expect: "x", + }; + assert.ok("reason" in locate(renamed, renamedAnchor)); + assert.ok("reason" in locate(lifted, liftedAnchor)); +}); + +test("a mutant with no derive operator and no bytes is refused, and so is one with bytes but no replacement", () => { + assert.match( + refusal(SOURCE, { name: "x", expect: "x" }), + /mutant has no replacement and no derive operator/, + ); + assert.match( + refusal(SOURCE, { name: "x", from: " return {", expect: "x" }), + /mutant has no replacement/, + ); +}); + +test("matchText still refuses a marker that occurs twice, and re-takes a reflowed one", () => { + assert.deepEqual(matchText("a && b", "a && b"), { start: 0, end: 6, retaken: false }); + assert.ok("reason" in matchText("x x", "x")); + const reflowed = matchText("a &&\n b", "a && b"); + assert.ok(!("reason" in reflowed) && reflowed.retaken); +}); diff --git a/tools/mutation-anchor.ts b/tools/mutation-anchor.ts new file mode 100644 index 00000000..ada74812 --- /dev/null +++ b/tools/mutation-anchor.ts @@ -0,0 +1,477 @@ +// Where a mutant applies, and what it writes there. +// +// This module is the anchor: a mutant may name its site by syntax tree, by bytes, or by a derived +// selector (`derive`: an operator over a member-scoped fragment), and the replacement bytes are +// returned with the range. It holds no state and reads no file, so the sweep and the anchors-only pass +// ask the same function the same question, and a test can ask it about a source string. +import ts from "typescript"; + +export interface Mutant { + /** What the wrong version does, in the words of the rule it breaks. */ + readonly name: string; + /** The exact bytes to replace. Omitted when `ast` or `derive` locates the site instead. */ + readonly from?: string; + /** A format-independent locator: find the site by syntax tree, not by text. Prettier reflows these + * files on every commit, and a text anchor silently stops applying the first time that happens. */ + readonly ast?: { + /** The method or function whose body is searched. The structurally scoped form: it survives + * reflow, and it refuses when the code it guards has moved out of the member it belongs to. */ + readonly within?: string; + /** The call or constructor to locate, by its callee name. */ + readonly call?: string; + /** How many arguments it takes, when the count is what tells the sites apart. */ + readonly argCount?: number; + }; + /** + * A derived site: the operator and the selector are the mutant's durable part, and the bytes are + * computed here on every run. This is what keeps a tooth alive across the edits that retire a text + * anchor - inlining a call, extracting a message, reflowing a statement. The name stays hand-written, + * because it is what the ledger and the named case refer to. + */ + readonly derive?: Derive; + readonly to?: string; + /** The test that must be the one to fail. */ + readonly expect: string; +} + +/** The operators a derived mutant may use, and the selector each one needs. */ +export interface Derive { + /** The member whose body holds the site. Exact, and required: the scope is what makes it unique. */ + readonly within: string; + readonly operator: + /** The condition the selector identifies never holds: it becomes `false`. */ + | "condition-never" + /** The condition the selector identifies always holds: it becomes `true`. The mirror of + * `condition-never`, because a rule written as `return a && b` says "this holds" - replacing it + * with `false` would reverse the rule instead of removing it. */ + | "condition-holds" + /** One term of that condition becomes its identity: `true` under `&&`, `false` under `||`. */ + | "neutralize-term" + /** Argument `arg` of the call to `call` becomes the declared `to` fragment. */ + | "replace-argument" + /** The value of the object property named `property` becomes the declared `to` fragment. */ + | "replace-property" + /** The statement the selector identifies is removed, with its line and its indentation. */ + | "drop-statement"; + /** Which guard: a fragment its condition's own text contains, e.g. `existing.deliveredBy`. */ + readonly condition?: string; + /** Which term of it, for `neutralize-term`: the term's own text, e.g. `!task.cancelled`. */ + readonly term?: string; + /** Which call, for `replace-argument`: its callee as written, e.g. `startableTasks`. */ + readonly call?: string; + /** Which argument, for `replace-argument`, counting from zero. */ + readonly arg?: number; + /** Which property, for `replace-property`: its name as written, e.g. `sessionReusable`. */ + readonly property?: string; + /** For `replace-property`: a fragment of the object literal that holds it, when the same property + * name is written in several literals of one member and only one of them is the site. */ + readonly in?: string; + /** Which statement, for `drop-statement`: a fragment of the statement's own text. */ + readonly statement?: string; +} + +/** The replacement bytes, absent only while `derive` computes them. The two value-substituting + * operators (`replace-argument`, `replace-property`) read it as the new value's text. */ + +/** A located site: the byte range to replace, and what to put there. */ +export interface Site { + readonly start: number; + readonly end: number; + readonly replacement: string; + readonly retaken: boolean; +} + +/** Whitespace-normalized text search: exact bytes first, then reflowed form. + * + * The commit hook runs prettier, so a reflowed anchor must not retire a tooth. More than one match + * is still refused, because replacing the first would leave the rule intact somewhere else. */ +export function matchText( + haystack: string, + anchor: string, +): { start: number; end: number; retaken: boolean } | { reason: string } { + const occurrences = haystack.split(anchor).length - 1; + if (occurrences === 1) { + const start = haystack.indexOf(anchor); + return { start, end: start + anchor.length, retaken: false }; + } + if (occurrences > 1) + return { reason: `marker occurs ${occurrences} times, refusing to claim a check` }; + // Built without a regex literal: one containing `${` confuses Node's type-stripping parser. + const special = ".*+?^$()[]{}|\\"; + const escaped = anchor + .trim() + .split(/\s+/u) + .map((part) => [...part].map((ch) => (special.includes(ch) ? "\\" + ch : ch)).join("")) + .join("\\s+"); + const matches = [...haystack.matchAll(new RegExp(escaped, "gu"))]; + if (matches.length !== 1) + return { + reason: `marker not found (${matches.length} matches once whitespace is normalized), refusing to claim a check`, + }; + const match = matches[0]!; + return { start: match.index, end: match.index + match[0].length, retaken: true }; +} + +/** Where a mutant applies: by selector when it derives one, by syntax tree when it says so, by bytes + * otherwise. + * + * A site that cannot be located is a failure, not an "not applicable": that verdict is reserved for + * a target file that is not on this branch at all. Files that are still being edited should carry a + * derived selector or an `ast` locator, because a text anchor in them retires itself the first time + * the formatter runs. */ +export function locate(text: string, mutant: Mutant): Site | { reason: string } { + if (mutant.derive) { + const derived = deriveSite(text, mutant.derive, mutant.to); + if ("reason" in derived) + return { reason: `derived ${mutant.derive.operator}: ${derived.reason}` }; + return derived; + } + if (mutant.to === undefined) + return { reason: "mutant has no replacement and no derive operator" }; + if (mutant.ast) { + const source = ts.createSourceFile("mutant.ts", text, ts.ScriptTarget.Latest, true); + if (mutant.ast.within !== undefined) { + const members: ts.Node[] = []; + const visit = (node: ts.Node): void => { + const named = + (ts.isMethodDeclaration(node) || ts.isFunctionDeclaration(node)) && + node.name?.getText(source) === mutant.ast!.within; + if (named) members.push(node); + ts.forEachChild(node, visit); + }; + visit(source); + if (members.length !== 1) + return { + reason: `ast scope ${mutant.ast.within} matched ${members.length} members, refusing to claim a check`, + }; + const member = members[0]!; + if (mutant.from === undefined) + return { reason: "an ast scope needs a from anchor to find inside it" }; + const inner = matchText(member.getText(source), mutant.from); + if ("reason" in inner) return { reason: `inside ${mutant.ast.within}: ${inner.reason}` }; + const offset = member.getStart(source); + return { + start: offset + inner.start, + end: offset + inner.end, + retaken: inner.retaken, + replacement: mutant.to, + }; + } + const found: ts.Node[] = []; + const walk = (node: ts.Node): void => { + if (ts.isCallExpression(node) || ts.isNewExpression(node)) { + const args = node.arguments?.length ?? 0; + if ( + node.expression.getText(source) === mutant.ast!.call && + (mutant.ast!.argCount === undefined || args === mutant.ast!.argCount) + ) + found.push(node); + } + ts.forEachChild(node, walk); + }; + walk(source); + if (found.length !== 1) + return { + reason: `ast locator ${mutant.ast.call} matched ${found.length} sites, refusing to claim a check`, + }; + return { + start: found[0]!.getStart(source), + end: found[0]!.getEnd(), + retaken: false, + replacement: mutant.to, + }; + } + if (mutant.from === undefined) + return { reason: "mutant has neither an ast locator nor a from anchor" }; + const found = matchText(text, mutant.from); + if ("reason" in found) return found; + if (found.retaken) + process.stdout.write( + ` re-taken anchor: ${mutant.name} (formatting reflowed it; ${String(found.end - found.start)} bytes)\n`, + ); + return { ...found, replacement: mutant.to }; +} + +/** Every node a predicate accepts, in source order. */ +function collect(root: ts.Node, isWanted: (node: ts.Node) => boolean): T[] { + const found: T[] = []; + const walk = (node: ts.Node): void => { + if (isWanted(node)) found.push(node as T); + ts.forEachChild(node, walk); + }; + walk(root); + return found; +} + +/** The one member with this name in the file, or why it is not one site. */ +function uniqueMember(source: ts.SourceFile, name: string): ts.Node | { reason: string } { + const named = (node: ts.Node): boolean => + (ts.isMethodDeclaration(node) || ts.isFunctionDeclaration(node)) && + node.name?.getText(source) === name; + const members = collect(source, named); + if (members.length !== 1) + return { + reason: `member ${name} matched ${members.length} members, refusing to claim a check`, + }; + return members[0]!; +} + +/** Argument `arg` of the one call to `call` inside the member. */ +function argumentSite( + source: ts.SourceFile, + member: ts.Node, + derive: Derive, + to: string | undefined, +): Site | { reason: string } { + if (derive.call === undefined || derive.arg === undefined) + return { reason: "replace-argument needs `call` and `arg`" }; + if (to === undefined) return { reason: "replace-argument needs the mutant's `to` fragment" }; + const calls = collect( + member, + (node) => ts.isCallExpression(node) && node.expression.getText(source) === derive.call, + ); + if (calls.length !== 1) + return { + reason: `call ${derive.call} matched ${calls.length} sites in ${derive.within}, refusing to claim a check`, + }; + const argument = calls[0]!.arguments[derive.arg]; + if (!argument) + return { + reason: `call ${derive.call} has no argument ${derive.arg}, refusing to claim a check`, + }; + return { + start: argument.getStart(source), + end: argument.getEnd(), + replacement: to, + retaken: false, + }; +} + +/** The text of the object literal a property sits in, which is what tells two same-named ones apart. */ +function holderText(source: ts.SourceFile, node: ts.Node): string { + let holder: ts.Node | undefined = node.parent; + while (holder && !ts.isObjectLiteralExpression(holder)) holder = holder.parent; + return holder ? holder.getText(source) : ""; +} + +/** The value of the one object property with this name inside the member. */ +function propertySite( + source: ts.SourceFile, + member: ts.Node, + derive: Derive, + to: string | undefined, +): Site | { reason: string } { + if (derive.property === undefined) return { reason: "replace-property needs a `property` name" }; + if (to === undefined) return { reason: "replace-property needs the mutant's `to` fragment" }; + const named = (node: ts.Node): boolean => + ts.isPropertyAssignment(node) && node.name.getText(source) === derive.property; + const properties = collect(member, named); + const held = + derive.in === undefined + ? properties + : properties.filter((node) => holderText(source, node).includes(derive.in!)); + if (held.length !== 1) + return { + reason: `property ${derive.property} matched ${held.length} sites in ${derive.within}, refusing to claim a check`, + }; + const value = held[0]!.initializer; + return { start: value.getStart(source), end: value.getEnd(), replacement: to, retaken: false }; +} + +/** The statement containing this fragment, deleted with its line and its indentation. */ +function statementSite( + source: ts.SourceFile, + member: ts.Node, + derive: Derive, + text: string, +): Site | { reason: string } { + if (derive.statement === undefined) + return { reason: "drop-statement needs a `statement` fragment" }; + const fragment = derive.statement; + const containing = (node: ts.Node): boolean => + (ts.isExpressionStatement(node) || + ts.isVariableStatement(node) || + ts.isReturnStatement(node) || + ts.isThrowStatement(node)) && + node.getText(source).includes(fragment); + const statements = collect(member, containing); + if (statements.length !== 1) + return { + reason: `${statements.length} statements in ${derive.within} contain the fragment, refusing to claim a check`, + }; + const statement = statements[0]!; + // The whole line goes: the indentation to its left and the newline to its right. A trailing comment + // on that line would be dropped with it, and a statement that shares its line with anything else + // would leave half of that line behind, so both are refused rather than silently damaged. + let start = statement.getStart(source); + while (start > 0 && (text[start - 1] === " " || text[start - 1] === "\t")) start -= 1; + if (start > 0 && text[start - 1] !== "\n") + return { + reason: "the statement does not begin its line, refusing to delete part of another one", + }; + const newline = text.slice(statement.getEnd()).match(/^[ \t]*(\r?\n)/u); + if (!newline) + return { reason: "the statement is not alone on its line, refusing to drop the line with it" }; + return { start, end: statement.getEnd() + newline[0].length, replacement: "", retaken: false }; +} + +/** The expression a decision position holds: a test, a returned value, an arrow's body, a bound value. */ +function decisionExpression(node: ts.Node): ts.Expression | undefined { + if (ts.isIfStatement(node) || ts.isWhileStatement(node) || ts.isDoStatement(node)) + return node.expression; + if (ts.isConditionalExpression(node)) return node.condition; + if (ts.isReturnStatement(node)) return node.expression; + if (ts.isForStatement(node)) return node.condition; + if (ts.isVariableDeclaration(node)) return node.initializer; + if (ts.isArrowFunction(node) && !ts.isBlock(node.body)) return node.body; + return undefined; +} + +/** Every decision expression in the member, in source order. */ +function decisionExpressions(member: ts.Node): ts.Expression[] { + const found: ts.Expression[] = []; + const walk = (node: ts.Node): void => { + const expression = decisionExpression(node); + if (expression) found.push(expression); + ts.forEachChild(node, walk); + }; + walk(member); + return found; +} + +/** + * The boolean expression the selector names: the member's decision positions, innermost first. + * + * Decision positions rather than only guards, because `return a && b` and `array.filter((x) => x.y)` + * are the same rule written without an `if`. A fragment can sit inside two of them at once, because + * an arrow's body may hold another arrow: + * `const ids = (pending) => tasks.filter((task) => !task.claimed && ready(task))`. The innermost one + * is the position the fragment names; two disjoint positions are two possible sites, and a site the + * selector cannot tell apart is refused. + */ +function conditionSite( + source: ts.SourceFile, + member: ts.Node, + derive: Derive, +): ts.Expression | { reason: string } { + const conditions = decisionExpressions(member); + const matching = + derive.condition === undefined + ? conditions + : conditions.filter((condition) => condition.getText(source).includes(derive.condition!)); + const innermost = matching.filter( + (condition) => + !matching.some( + (other) => + other !== condition && + other.getStart(source) >= condition.getStart(source) && + other.getEnd() <= condition.getEnd(), + ), + ); + if (innermost.length !== 1) + return { + reason: + derive.condition === undefined + ? `${conditions.length} conditions in ${derive.within}, refusing to choose one` + : `${innermost.length} conditions in ${derive.within} mention ${derive.condition}, refusing to claim a check`, + }; + return innermost[0]!; +} + +/** A condition replaced whole: it never holds, or it always holds. */ +function wholeCondition( + source: ts.SourceFile, + condition: ts.Expression, + derive: Derive, +): Site | { reason: string } { + if (derive.condition === undefined) + return { reason: `${derive.operator} needs a \`condition\` fragment` }; + // These two operators replace the condition whole, so a fragment that names only part of it would + // silently widen the mutant into "every reason this rule has". A term has its own operator. + const text = condition.getText(source); + const whole = matchText(text, derive.condition); + if ("reason" in whole) return { reason: `inside the condition: ${whole.reason}` }; + if (whole.start !== 0 || whole.end !== text.length) + return { + reason: `the fragment names part of ${derive.within}'s condition; ${derive.operator} replaces all of it - use neutralize-term for a term`, + }; + return { + start: condition.getStart(source), + end: condition.getEnd(), + replacement: derive.operator === "condition-never" ? "false" : "true", + retaken: whole.retaken, + }; +} + +/** + * The literal that makes a term's operator absorb it: `true` under `&&`, `false` under `||`. The + * literal is what keeps the rest of the condition - and its formatting - intact. `null` when the code + * does not say which operator joins the term, which is refused rather than guessed at. + */ +function identityLiteral(before: string, after: string): string | null { + // The fragment may be the last operand of a parenthesized group, so brackets between it and the + // joining operator are skipped before the operator is read, and the operator on the left is read + // only when there is nothing to the right. + const afterBare = after + .trimStart() + .replace(/^[)\]},;]+/u, "") + .trimStart(); + const beforeBare = before + .trimEnd() + .replace(/[[({]+$/u, "") + .trimEnd(); + if (afterBare.startsWith("&&")) return "true"; + if (afterBare.startsWith("||")) return "false"; + if (afterBare !== "") return null; + if (beforeBare.endsWith("&&")) return "true"; + if (beforeBare.endsWith("||")) return "false"; + return null; +} + +/** One term of a condition becomes its identity. */ +function neutralizedTerm( + source: ts.SourceFile, + condition: ts.Expression, + derive: Derive, +): Site | { reason: string } { + if (derive.term === undefined) return { reason: "neutralize-term needs a `term` fragment" }; + const text = condition.getText(source); + const term = matchText(text, derive.term); + if ("reason" in term) return { reason: `inside the condition: ${term.reason}` }; + const literal = identityLiteral(text.slice(0, term.start), text.slice(term.end)); + if (literal === null) + return { + reason: `cannot tell which operator joins the term in ${derive.within}, refusing to claim a check`, + }; + const offset = condition.getStart(source); + return { + start: offset + term.start, + end: offset + term.end, + replacement: literal, + retaken: term.retaken, + }; +} + +/** + * Resolve a derived mutant: the selector picks the code, the operator says what to do to it, and the + * bytes are computed here rather than stored. Every selector is scoped to one member, and every + * fragment is matched as exactly one occurrence or refused, so a reflow or an inlining that would + * retire a byte anchor leaves a derived tooth applying. + */ +function deriveSite( + text: string, + derive: Derive, + to: string | undefined, +): Site | { reason: string } { + const source = ts.createSourceFile("mutant.ts", text, ts.ScriptTarget.Latest, true); + const member = uniqueMember(source, derive.within); + if ("reason" in member) return member; + if (derive.operator === "replace-argument") return argumentSite(source, member, derive, to); + if (derive.operator === "replace-property") return propertySite(source, member, derive, to); + if (derive.operator === "drop-statement") return statementSite(source, member, derive, text); + const condition = conditionSite(source, member, derive); + if ("reason" in condition) return condition; + return derive.operator === "neutralize-term" + ? neutralizedTerm(source, condition, derive) + : wholeCondition(source, condition, derive); +} diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 9e6922fe..6a668344 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -32,6 +32,13 @@ * * Where a mutant says where it applies: * + * - `derive: { within, operator, ... }` is a **selector plus an operator**: the member is named, a + * fragment inside it is matched (the same matcher as below, so a reflow cannot break it), and the + * bytes to write are computed on every run. This is the form a tooth should have. Its identity is + * the rule and the site it names, never the text that happens to be there today, and it is what + * lets a rename, a reordered condition or a lifted line leave the tooth aimed at the same rule. + * Selectors that match more than one site - or a fragment that fits two conditions - are refused, + * never guessed at. * - `ast: { within: "" }` (or `ast: { call, argCount }`) locates the site through the syntax * tree. Use this in any file that is still being edited. It survives reformatting, and it refuses * when the code it guards has moved out of the member it belongs to - a move that a byte anchor @@ -41,14 +48,17 @@ * printed, so a reflow never retires a tooth without saying so. * * Usage: - * npm run mutation:teeth -- [--targets=[,...]] [--json ] + * npm run mutation:teeth -- [--targets=[,...]] [--mutant=[,...]] [--json ] + * npm run mutation:teeth -- --anchors-only [--targets=...] + * `--anchors-only` resolves every selected mutant and writes nothing: it answers "are the teeth still + * aimed at something", in under a second, without running a suite. Run it before a sweep, and in the + * static contract, because a dead anchor is otherwise only visible as a tooth that stopped counting. * Exit status is non-zero if the clean run fails, any mutant survives, any anchor is missing in a - * named target, or a restore is not byte-identical. + * named target, a restore is not byte-identical, or (in `--anchors-only`) any site no longer applies. */ import { execFileSync } from "node:child_process"; import { existsSync, readFileSync, writeFileSync } from "node:fs"; import { parseArgs } from "node:util"; -import ts from "typescript"; import { clearMutationLock, @@ -57,27 +67,7 @@ import { type MutationLock, } from "./mutation-lock.ts"; import { writeJsonAtomic } from "./parts/fs.ts"; - -interface Mutant { - /** What the wrong version does, in the words of the rule it breaks. */ - readonly name: string; - /** The exact bytes to replace. Omitted when `ast` locates the site instead. */ - readonly from?: string; - /** A format-independent locator: find the site by syntax tree, not by text. Prettier reflows these - * files on every commit, and a text anchor silently stops applying the first time that happens. */ - readonly ast?: { - /** The method or function whose body is searched. The structurally scoped form: it survives - * reflow, and it refuses when the code it guards has moved out of the member it belongs to. */ - readonly within?: string; - /** The call or constructor to locate, by its callee name. */ - readonly call?: string; - /** How many arguments it takes, when the count is what tells the sites apart. */ - readonly argCount?: number; - }; - readonly to: string; - /** The test that must be the one to fail. */ - readonly expect: string; -} +import { locate, type Mutant } from "./mutation-anchor.ts"; interface Target { readonly target: string; @@ -524,16 +514,22 @@ const TARGETS: readonly Target[] = [ // A cancellation is a fact about the task, so it gates dispatch the way a rejection gates // a dependent. This is the half acceptance already had and eligibility did not. name: "a-cancelled-task-is-still-dispatched", - ast: { within: "selection" }, - from: " current(task) &&\n !task.cancelled &&", - to: " current(task) &&", + derive: { + within: "selection", + operator: "neutralize-term", + condition: "!task.cancelled", + term: "!task.cancelled", + }, expect: "a cancelled task is not dispatched, and nothing reads one as a closed input", }, { name: "the-dispatch-does-not-require-a-cancelled-input-to-be-closed", - ast: { within: "acceptedDependency" }, - from: " if (!task || !task.accepted || task.cancelled || !current(task) || visiting.has(id)) return false;", - to: " if (!task || !task.accepted || !current(task) || visiting.has(id)) return false;", + derive: { + within: "acceptedDependency", + operator: "neutralize-term", + condition: "task.cancelled", + term: "task.cancelled", + }, expect: "nextTask refuses a task marked cancelled, whatever else the caller set", }, { @@ -548,9 +544,12 @@ const TARGETS: readonly Target[] = [ // The budget is what a claim spends, so a spent budget is an empty set. Nothing else may // decide whether selection is open: this is the rule the C arm's slot count was once absent from. name: "a-live-claim-does-not-block-selection", - ast: { within: "selection" }, - from: ' if (spent || severalWaits) return answer([], 0, spent ? "budget-spent" : "several-waits-pending");', - to: ' if (severalWaits) return answer([], 0, spent ? "budget-spent" : "several-waits-pending");', + derive: { + within: "selection", + operator: "neutralize-term", + condition: "spent || severalWaits", + term: "spent", + }, expect: "the answer names the gate a unit was refused through, and the gates are the rule's own", }, @@ -558,9 +557,12 @@ const TARGETS: readonly Target[] = [ // A task someone is working is not on offer, whatever the budget. Without this, a run with // slots to spare would hand the same task to a second worker. name: "a-claimed-task-stays-on-offer", - ast: { within: "selection" }, - from: " tasks.filter((task) => !task.claimed && ready(task)).map((task) => task.id);", - to: " tasks.filter((task) => ready(task)).map((task) => task.id);", + derive: { + within: "selection", + operator: "neutralize-term", + condition: "!task.claimed && ready(task)", + term: "!task.claimed", + }, expect: "a declared budget is spent by claims in flight, not by the next task's rank", }, { @@ -575,9 +577,12 @@ const TARGETS: readonly Target[] = [ { // Zero or half a slot is not a smaller budget, and rounding it would hide the caller's typo. name: "half-a-slot-is-a-smaller-budget", - ast: { within: "checkedSlots" }, - from: ' if (!Number.isSafeInteger(slots) || slots < 1) throw new Error("slots must be a positive integer");', - to: ' if (!Number.isSafeInteger(slots)) throw new Error("slots must be a positive integer");', + derive: { + within: "checkedSlots", + operator: "neutralize-term", + condition: "!Number.isSafeInteger(slots) || slots < 1", + term: "slots < 1", + }, expect: "a claim in flight does not release a dependent, and half a slot is not a budget", }, { @@ -585,9 +590,11 @@ const TARGETS: readonly Target[] = [ // is not skipped in favour of a later ready task. This is the rule an ordering step is most // likely to bypass by accident, so it has its own tooth. name: "a-head-blocked-by-a-stale-input-is-skipped", - ast: { within: "selection" }, - from: ' if (!current(first) || !waiting(first)) return answer([], 0, "earlier-unit-blocked");', - to: ' if (false) return answer([], 0, "earlier-unit-blocked");', + derive: { + within: "selection", + operator: "condition-never", + condition: "!current(first) || !waiting(first)", + }, expect: "the round's own answer is the shared rule's answer, not an ordering's", }, { @@ -596,18 +603,19 @@ const TARGETS: readonly Target[] = [ // silences every refusal at once: an empty legal set with no reasons is exactly the answer // the design says a caller cannot be given. name: "a-refused-unit-is-silent", - ast: { within: "selection" }, - from: " causes.set(task.id, own.length > 0 ? own : held ? [held] : []);", - to: " causes.set(task.id, []);", + derive: { within: "selection", operator: "replace-argument", call: "causes.set", arg: 1 }, + to: "[]", expect: "no refusal in the answer is silent, over the flags a plan's facts can carry", }, { // A cause names the gate; it never restates the condition. Dropping one leaves a refusal // whose reason the rule can no longer give, even when another gate would still be true. name: "a-stale-input-is-not-named", - ast: { within: "selection" }, - from: ' if (!current(task)) causes.push("stale-input");', - to: ' if (false) causes.push("stale-input");', + derive: { + within: "selection", + operator: "condition-never", + condition: "!current(task)", + }, expect: "the answer names the gate a unit was refused through, and the gates are the rule's own", }, @@ -642,25 +650,33 @@ const TARGETS: readonly Target[] = [ }, { name: "an-unattested-reading-counts-as-evidence", - ast: { within: "speculationOutcome" }, - from: " if (!fact.authoritative)", - to: " if (false)", + derive: { + within: "speculationOutcome", + operator: "condition-never", + condition: "!fact.authoritative", + }, expect: "a reading nobody attested is not evidence", }, { name: "evidence-about-another-version-is-the-same-fact", - ast: { within: "speculationOutcome" }, - from: " if (fact.version !== assumption.version)", - to: " if (false)", + derive: { + within: "speculationOutcome", + operator: "condition-never", + condition: "fact.version !== assumption.version", + }, expect: "evidence about another version is not evidence about this fact", }, { // The invalidation rule: a contradicted guess closes its branch session, so the real path // cannot take an answer from a model that has already been told the guess. name: "a-contradicted-guess-keeps-its-session", - ast: { within: "speculationOutcome" }, - from: ' outcome: "discard",\n sessionReusable: false,', - to: ' outcome: "discard",\n sessionReusable: true,', + derive: { + within: "speculationOutcome", + operator: "replace-property", + property: "sessionReusable", + in: 'outcome: "discard"', + }, + to: "true", expect: "a guess the evidence contradicts is discarded, and its branch session is closed", }, ], @@ -1324,58 +1340,77 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "fusion-shares-a-session-across-different-capabilities", - ast: { within: "compatibleDeclarations" }, - from: " first.capability === next.capability &&", - to: " first.capability === first.capability &&", + derive: { + within: "compatibleDeclarations", + operator: "neutralize-term", + condition: "next.capability", + term: "first.capability === next.capability", + }, expect: "fusion refuses a unit that needs a different execution capability", }, { name: "fusion-shares-a-session-across-different-authorities", - ast: { within: "compatibleDeclarations" }, - from: " first.authority === next.authority &&", - to: " first.authority === first.authority &&", + derive: { + within: "compatibleDeclarations", + operator: "neutralize-term", + condition: "next.authority", + term: "first.authority === next.authority", + }, expect: "fusion refuses a unit acting under a different authority", }, { name: "fusion-widens-what-a-unit-may-read", - ast: { within: "compatibleDeclarations" }, - from: " subset(next.visible, first.visible)", - to: " subset(first.visible, next.visible)", + derive: { + within: "compatibleDeclarations", + operator: "neutralize-term", + condition: "subset(next.visible, first.visible)", + term: "subset(next.visible, first.visible)", + }, expect: "fusion refuses a successor whose visibility the session would widen", }, { name: "fusion-continues-from-an-unverified-answer", - ast: { within: "sharedSessionLegal" }, - from: " if (!first.accepted) return false;", - to: " if (false && !first.accepted) return false;", + derive: { + within: "sharedSessionLegal", + operator: "condition-never", + condition: "!first.accepted", + }, expect: "fusion refuses to continue from a unit whose verdict is not accepted", }, { name: "fusion-starts-a-successor-whose-dependency-is-not-accepted", - ast: { within: "sharedSessionLegal" }, - from: " if (next.dependencies.some((id) => !acceptedDependency(byId, id))) return false;", - to: " if (false && next.dependencies.some((id) => !acceptedDependency(byId, id))) return false;", + derive: { + within: "sharedSessionLegal", + operator: "condition-never", + condition: "next.dependencies.some((id) => !acceptedDependency(byId, id))", + }, expect: "fusion refuses a successor whose dependency is delivered but not accepted", }, { name: "fusion-carries-a-cancelled-unit-into-its-next-unit", - ast: { within: "neitherCancelled" }, - from: " return !first.cancelled && !next.cancelled;", - to: " return true;", + derive: { + within: "neitherCancelled", + operator: "condition-holds", + condition: "!first.cancelled && !next.cancelled", + }, expect: "fusion refuses a cancelled unit, before or after", }, { name: "fusion-crosses-a-host-yield-boundary", - ast: { within: "sharedSessionLegal" }, - from: " if (!!first.externalEvent && !first.externalReady) return false;", - to: " if (false && !!first.externalEvent && !first.externalReady) return false;", + derive: { + within: "sharedSessionLegal", + operator: "condition-never", + condition: "!!first.externalEvent && !first.externalReady", + }, expect: "fusion ends the session at a declared external wait that is not ready", }, { name: "fusion-reuses-history-across-a-pending-branch", - ast: { within: "acrossAPendingBranch" }, - from: " return pending.includes(before) || pending.includes(after);", - to: " return false;", + derive: { + within: "acrossAPendingBranch", + operator: "condition-never", + condition: "pending.includes(before) || pending.includes(after)", + }, expect: "fusion never reuses the history across a fact whose branch is still pending", }, ], @@ -1493,101 +1528,6 @@ interface Outcome { readonly mutants: readonly MutantOutcome[]; } -/** Whitespace-normalized text search: exact bytes first, then reflowed form. - * - * The commit hook runs prettier, so a reflowed anchor must not retire a tooth. More than one match - * is still refused, because replacing the first would leave the rule intact somewhere else. */ -function matchText( - haystack: string, - anchor: string, -): { start: number; end: number; retaken: boolean } | { reason: string } { - const occurrences = haystack.split(anchor).length - 1; - if (occurrences === 1) { - const start = haystack.indexOf(anchor); - return { start, end: start + anchor.length, retaken: false }; - } - if (occurrences > 1) - return { reason: `marker occurs ${occurrences} times, refusing to claim a check` }; - // Built without a regex literal: one containing `${` confuses Node's type-stripping parser. - const special = ".*+?^$()[]{}|\\"; - const escaped = anchor - .trim() - .split(/\s+/u) - .map((part) => [...part].map((ch) => (special.includes(ch) ? "\\" + ch : ch)).join("")) - .join("\\s+"); - const matches = [...haystack.matchAll(new RegExp(escaped, "gu"))]; - if (matches.length !== 1) - return { - reason: `marker not found (${matches.length} matches once whitespace is normalized), refusing to claim a check`, - }; - const match = matches[0]!; - return { start: match.index, end: match.index + match[0].length, retaken: true }; -} - -/** Where a mutant applies: by syntax tree when it says so, by bytes otherwise. - * - * A site that cannot be located is a failure, not an "not applicable": that verdict is reserved for - * a target file that is not on this branch at all. Files that are still being edited should carry an - * `ast` locator, because a text anchor in them retires itself the first time the formatter runs. */ -function locate( - text: string, - mutant: Mutant, -): { start: number; end: number; retaken: boolean } | { reason: string } { - if (mutant.ast) { - const source = ts.createSourceFile("mutant.ts", text, ts.ScriptTarget.Latest, true); - if (mutant.ast.within !== undefined) { - const members: ts.Node[] = []; - const visit = (node: ts.Node): void => { - const named = - (ts.isMethodDeclaration(node) || ts.isFunctionDeclaration(node)) && - node.name?.getText(source) === mutant.ast!.within; - if (named) members.push(node); - ts.forEachChild(node, visit); - }; - visit(source); - if (members.length !== 1) - return { - reason: `ast scope ${mutant.ast.within} matched ${members.length} members, refusing to claim a check`, - }; - const member = members[0]!; - if (mutant.from === undefined) - return { reason: "an ast scope needs a from anchor to find inside it" }; - const inner = matchText(member.getText(source), mutant.from); - if ("reason" in inner) return { reason: `inside ${mutant.ast.within}: ${inner.reason}` }; - const offset = member.getStart(source); - return { start: offset + inner.start, end: offset + inner.end, retaken: inner.retaken }; - } - const found: ts.Node[] = []; - const walk = (node: ts.Node): void => { - if (ts.isCallExpression(node) || ts.isNewExpression(node)) { - const args = node.arguments?.length ?? 0; - if ( - node.expression.getText(source) === mutant.ast!.call && - (mutant.ast!.argCount === undefined || args === mutant.ast!.argCount) - ) - found.push(node); - } - ts.forEachChild(node, walk); - }; - walk(source); - if (found.length !== 1) - return { - reason: `ast locator ${mutant.ast.call} matched ${found.length} sites, refusing to claim a check`, - }; - return { start: found[0]!.getStart(source), end: found[0]!.getEnd(), retaken: false }; - } - if (mutant.from === undefined) - return { reason: "mutant has neither an ast locator nor a from anchor" }; - const found = matchText(text, mutant.from); - if ("reason" in found) return found; - if (found.retaken) - process.stdout.write( - ` re-taken anchor: ${mutant.name} (formatting reflowed it; ${String(found.end - found.start)} bytes) -`, - ); - return found; -} - /** The failure lines a suite run reported. Used for both verdicts: a surviving mutant * and a clean run that failed. The clean run is the one that proves nothing, so * naming its failing case is what turns "the harness proves nothing" into a fix. */ @@ -1677,6 +1617,10 @@ const { values } = parseArgs({ // One mutant at a time, so a single sweep command can be bounded by a caller with a time budget // instead of a whole target's worth of suite runs. Repeating it runs the mutants named. mutant: { type: "string", multiple: true }, + // Resolve every mutant's site and write nothing: the pass that says whether the teeth still bite + // where they are aimed, without running a suite. Cheap enough to sit in the static contract, and + // it is the only check that notices a tooth whose anchor stopped matching. + "anchors-only": { type: "boolean" }, }, }); const requested = (values.targets ?? []).flatMap((entry) => entry.split(",")).filter(Boolean); @@ -1698,6 +1642,48 @@ const selected = requested.length /** Naming a target is a claim that it can be exercised here; the default list is not. */ const strict = requested.length > 0; +/** + * Resolve every selected mutant and say which sites no longer apply, without running a suite and + * without writing a byte. A sweep proves the teeth bite; this proves they are still aimed at something, + * which is the reading that was missing when two teeth had silently stopped matching their rule. + */ +if (values["anchors-only"]) { + const hazard = mutationHazard(); + if (hazard) + throw new Error( + `refusing to read anchors: ${hazard} - a target file may hold a live mutant, so the sites on ` + + `disk are not the tree's`, + ); + const failures: string[] = []; + let claimed = 0; + let applicable = 0; + for (const { target, mutants } of selected) { + if (!existsSync(target)) { + failures.push(`${target}: not present in this checkout`); + continue; + } + claimed += mutants.length; + const text = readFileSync(target, "utf8"); + const wanted = onlyMutants.length + ? mutants.filter((mutant) => onlyMutants.includes(mutant.name)) + : mutants; + for (const mutant of wanted) { + const site = locate(text, mutant); + if ("reason" in site) failures.push(`${target} / ${mutant.name}: ${site.reason}`); + else applicable += 1; + } + } + process.stdout.write( + `anchors: ${String(applicable)} of ${String(claimed)} resolve, over ${String(selected.length)} targets\n`, + ); + for (const failure of failures) process.stderr.write(`${failure}\n`); + if (failures.length > 0) { + process.stdout.write(`anchors: ${String(failures.length)} sites no longer apply\n`); + process.exit(1); + } + process.exit(0); +} + const outcomes: Outcome[] = []; const problems: string[] = []; const skipped: string[] = []; @@ -1784,7 +1770,7 @@ for (const { target, suites, mutants: declared } of selected) { } sweep.live = true; writeMutationLock(sweep); - writeFileSync(target, text.slice(0, site.start) + mutant.to + text.slice(site.end)); + writeFileSync(target, text.slice(0, site.start) + site.replacement + text.slice(site.end)); // The named case first, and the whole suite only if it did not fail. Both readings mean the same // thing - `caught` still requires the named case to appear in a failing run - so the fallback // never weakens a tooth; it just avoids paying for a suite whose other cases cannot change the From fcbf87f7dd1c54625bb6195db25510c6fccc1c28 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:03:25 +0800 Subject: [PATCH 24/32] The anchors pass joins the static contract A tooth that stopped matching its rule is invisible today: no CI job runs mutation:teeth, and in a sweep it is reported as "not applicable" and excluded from the denominator - which is how the two dead teeth repaired in dbf10ac5 stayed dead without anyone noticing. npm run mutation:anchors runs tools/mutation-teeth.ts --anchors-only: every selected mutant's site is resolved and nothing is written, no suite runs, and a site that cannot be resolved fails the run. It is now one of verify:static's checks and one of the ci-and-tests route's atomic checks, in the same order the route-contract test enforces (the route's list must equal the chain, test:product included). Measured on this revision: verify:static exit 0 in 41 s; a full npm run agent:verify exit 0 in 112 s inside its 150-second budget, with the anchors pass itself 0.98 s, all 149 anchors resolving over 23 targets. The pass proves a site still resolves, not that the mutant is still caught - a stale expect still passes it - so the full sweep stays the standing rule before a push and stays out of the gate. docs/design/ci-cd-and-quality.md gains the check and the reason it is in the gate; the record (docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md, with its zh-CN pair) marks its third slice landed. --- agent-context.yaml | 1 + ...-09-24-mutants-are-derived-not-anchored.md | 9 ++++-- ...-mutants-are-derived-not-anchored.zh-CN.md | 2 +- docs/design/ci-cd-and-quality.md | 30 ++++++++++--------- package.json | 3 +- 5 files changed, 26 insertions(+), 19 deletions(-) diff --git a/agent-context.yaml b/agent-context.yaml index e0e63af8..d0eaa0cd 100644 --- a/agent-context.yaml +++ b/agent-context.yaml @@ -143,6 +143,7 @@ routes: - build - package:check - check + - mutation:anchors - check:lock - lint - format:check diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index b0b8ed70..c8d9a6f1 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -110,9 +110,12 @@ severalWaits`, and the prediction made before running it - one dead anchor, seve selector was widened to include a variable's initializer; the corrected prediction (only the hand anchor dies) then held exactly: `anchors: 148 of 149 resolve`, one failure, and the file restored byte-identically. -3. The anchors-only pass, wired into the static contract with its route and design updates. **Not - landed yet:** the pass exists (`--anchors-only`, 149 of 149 resolving in 0.6 s, no suite run, nothing - written) but is not yet one of the route's atomic checks. +3. The anchors-only pass, wired into the static contract with its route and design updates. **Landed:** + `npm run mutation:anchors` (`--anchors-only`) is one of `verify:static`'s checks and one of the + `ci-and-tests` route's atomic checks, in the same order the route-contract test enforces. Measured + on this revision: all 149 anchors resolve in 0.98 s inside a full `npm run agent:verify`, which is + 112 s end to end - inside its 150-second budget - and a standalone `verify:static` is 41 s. The full + sweep stays out of the gate. 4. Retirement pass over the ~18-20 teeth whose rule already has a relational or enumerative check, and over the I/E teeth that can be replaced; the ledger's `proven` sentence updated in the same commit. **Not started.** diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index 8adb9c4f..cd0cb7e6 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -39,7 +39,7 @@ 1. 本文,加上工具的算子词汇表与解析器;用“对源码字符串”而不是对文件系统的测试来证明它。**已落地:**解析器是 `tools/mutation-anchor.ts`(它不持有任何状态,所以从 sweep 脚本里搬出来,测试可以直接调用),16 个用例全部作用在源码字符串上,不碰文件系统。 2. 试点目标转换并重跑:同样的名字、同样的用例、同样的抓法;随后在该文件内做一次重构,证明派生出来的牙活过了字节锚点活不过去的那次改动。**已落地,带一处更正:**演示用的那次重构是“改一个局部变量名 + 把条件抬成 `const closed = spent || severalWaits`”,而运行前写下的预测——死一颗锚点、七个派生位置存活——**是错的**:死了两处,手写锚点和派生的 `a-live-claim-does-not-block-selection`,因为“绑定到名字上的条件”当时不算选择器眼里的一次决策位置。把条件抬成局部常量是维护者真会做的重构,于是选择器扩到了变量初始化式;更正后的预测(只有手写锚点会死)随后精确成立:`anchors: 148 of 149 resolve`,一处失败,文件逐字节还原。 -3. anchors-only 这趟检查接入静态契约,并同步路由与设计文档。**尚未落地:**这趟检查已经存在(`--anchors-only`,149/149 在 0.6 秒内解析完,不跑用例,不写任何字节),但还没有成为该路由的一条原子检查。 +3. anchors-only 这趟检查接入静态契约,并同步路由与设计文档。**已落地:**`npm run mutation:anchors`(`--anchors-only`)现在是 `verify:static` 的一条检查,也是 `ci-and-tests` 路由的一条原子检查,顺序与路由契约测试强制的一致。本版本实测:一次完整的 `npm run agent:verify` 里,149 颗锚点全部解析,用时 0.98 秒,整轮 112 秒——留在 150 秒预算之内;单独跑 `verify:static` 是 41 秒。全量 sweep 仍不进闸门。 4. 对“规则已有关系式/穷举式检查”的那约 18-20 颗,以及能被替代的 I/E 牙,做退役;台账里 `proven` 那句话在同一提交里更新。**尚未开始。** ## 验收标准 diff --git a/docs/design/ci-cd-and-quality.md b/docs/design/ci-cd-and-quality.md index e5ff94d4..de1ac9d5 100644 --- a/docs/design/ci-cd-and-quality.md +++ b/docs/design/ci-cd-and-quality.md @@ -43,15 +43,15 @@ exit_criteria: Replace with a stable contract test or remove after the redesign ## 2. 执行轨道 -| 命令 | 内容 | 用途 | -| --------------------------- | ------------------------------------------------------ | ---------------- | -| `npm run test:product` | core、CLI、adapter、docs、Skill、工具与 test support | 产品正确性 | -| `npm run test:coverage` | 与 product 相同的集合并生成覆盖率 | 阻塞 CI | -| `npm run test:research` | `tests/benchmarks`、`tests/evals`、`tests/official` | 非阻塞研究表征 | -| `npm run test:chaos` | Windows 故障注入与进程/文件锁生命周期 | 独立阻塞轨道 | -| `npm test` | 所有 `tests/**/*.test.ts` | 本地完整兼容入口 | -| `npm run verify:static` | build/package/type/lint/format/docs/context/complexity | 本地与 CI 共用 | -| `npm run verify:product-ci` | build + product coverage | 本地与 CI 共用 | +| 命令 | 内容 | 用途 | +| --------------------------- | --------------------------------------------------------------- | ---------------- | +| `npm run test:product` | core、CLI、adapter、docs、Skill、工具与 test support | 产品正确性 | +| `npm run test:coverage` | 与 product 相同的集合并生成覆盖率 | 阻塞 CI | +| `npm run test:research` | `tests/benchmarks`、`tests/evals`、`tests/official` | 非阻塞研究表征 | +| `npm run test:chaos` | Windows 故障注入与进程/文件锁生命周期 | 独立阻塞轨道 | +| `npm test` | 所有 `tests/**/*.test.ts` | 本地完整兼容入口 | +| `npm run verify:static` | build/package/type/牙齿锚点/lint/format/docs/context/complexity | 本地与 CI 共用 | +| `npm run verify:product-ci` | build + product coverage | 本地与 CI 共用 | 另有三个命令只用于本地,刻意不作为 CI 轨道:`npm run lint:fix`(与 `lint` 同一 ESLint 范围,加 `--fix`)、`npm run hotspot:modules`(对 store 各方法做静态调用点计数)、`npm run perf:hotspots`(从某个 store 的 `perf_aggregates` 打印分段延迟占比)。 @@ -59,6 +59,8 @@ exit_criteria: Replace with a stable contract test or remove after the redesign `npm run check:tests`(`tsc -p tsconfig.tests.json --noUnusedLocals --noUnusedParameters`,`tests/` 面第一次被类型检查)刻意不进入任何阻塞契约,只在 CI 以 advisory 步骤运行。 +`verify:static` 中的 `mutation:anchors`(`tools/mutation-teeth.ts --anchors-only`)只做一件事:把 149 颗具名 mutant 的位置全部解析一遍,不跑任何用例、不写任何字节,在 0.6 秒内回答“每一颗牙是否还瞄着东西”。它进入静态契约是因为**一颗锚点失效时没有别的检查会注意到**:全量 sweep 不跑(`mutation:teeth` 不在任何 CI 作业里),而一颗匹配不到位置的牙在 sweep 报告里只是“不可应用”并被排除出分母——本轮修掉的两颗牙就是这样悄无声息地停摆的。全量 sweep 仍然不进闸门:它是分钟级、要跑用例,属于推送前的常设规则([决策](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md))。 + `verify:static` 中的 `complexity:gate` 默认以 `git merge-base HEAD origin/main` 为基线(可用 `--base ` 显式覆盖)。基线必须是 merge base 而不是 `HEAD`:后者只比较未提交的工作树,于是已提交到分支的改动完全不可见 —— 在 CI 的干净检出上它永远报“无改动”,等于每个 PR 都没有被这条 gate 检查过。因此每次运行都会**陈述自己用了哪个基线**,并**点名它未能测量的改动文件**(ESLint 拒绝某路径、或文件根本无法解析,都会产出“零发现”,与“量过且干净”无法区分)。 `eval:*`、`benchmark:*`、真实 LLM/embedding 与官方大数据集运行不进入 keyless CI。研究测试可以验证 adapter 和计分契约,但不得访问外部密钥或把实验常量提升为产品默认值。 @@ -103,11 +105,11 @@ TestRuntime 测试资源按断言需观察的边界分为四类,按需取得最轻且隔离的资源([决策](../decisions/implemented/2026-09-24-resource-matched-test-fixtures.md)): -| 断言边界 | 测试资源 | 可省的准备工作 | -| --- | --- | --- | -| 纯数据或决策 | 进程内输入 | 工作区、数据库和外部进程 | -| 单个 SQLite 连接的读写 | 每项独立的 `:memory:` store | 临时数据库文件及其清理 | -| 重开、旧库迁移、WAL、独立连接或路径 | 每项唯一的文件数据库或工作区 | 与断言无关的额外文件和服务 | +| 断言边界 | 测试资源 | 可省的准备工作 | +| ----------------------------------------- | -------------------------------------------- | -------------------------- | +| 纯数据或决策 | 进程内输入 | 工作区、数据库和外部进程 | +| 单个 SQLite 连接的读写 | 每项独立的 `:memory:` store | 临时数据库文件及其清理 | +| 重开、旧库迁移、WAL、独立连接或路径 | 每项唯一的文件数据库或工作区 | 与断言无关的额外文件和服务 | | CLI、Git、HTTP、daemon lease 或跨进程共享 | 断言这些边界时使用实际进程、端口及所需工作区 | 不参与断言的重复启动和检查 | 观察文件或跨进程语义的测试保留真实边界;只断言控制平面策略、receipt 或 harness 编排时,可以注入固定的 `RepositoryProvider` 观察值,另由真实 Git 集成测试覆盖发现、dirty 范围、提交和 forge 绑定。集成 `testDatabase()` 使用真实文件路径供 daemon 共享。真实 Git fixture 保留 `init/add/commit`,提交身份只传给 `commit`,避免每项再启动两次 `git config`。资源类别只决定 fixture 和可安全减少的准备工作,不改变 Safety/Contract/Guardrail 分类、阻塞地位或 narrow/full 验证范围;后两者由既有 route 与验证契约决定。 diff --git a/package.json b/package.json index 0241f63a..7b1f3f25 100644 --- a/package.json +++ b/package.json @@ -127,9 +127,10 @@ "ci:uncovered-tests": "node --experimental-strip-types tools/ci-uncovered-tests.ts --check", "complexity:gate": "node --experimental-strip-types tools/complexity-gate.ts", "mutation:teeth": "node --experimental-strip-types tools/mutation-teeth.ts", + "mutation:anchors": "node --experimental-strip-types tools/mutation-teeth.ts --anchors-only", "verify:packages": "node --experimental-strip-types tools/verify-packages.ts", "check:lock": "node --experimental-strip-types tools/check-lock.ts", - "verify:static": "npm run build && npm run package:check && npm run check && npm run check:lock && npm run lint && npm run format:check && npm run docs:check && npm run agent:context:check && npm run complexity:gate && npm run verify:packages && npm run glossary:check && npm run rtm:check", + "verify:static": "npm run build && npm run package:check && npm run check && npm run mutation:anchors && npm run check:lock && npm run lint && npm run format:check && npm run docs:check && npm run agent:context:check && npm run complexity:gate && npm run verify:packages && npm run glossary:check && npm run rtm:check", "verify:product-ci": "npm run build && npm run test:coverage", "verify:research": "npm run prompts:generate && npm run test:research", "verify:node-compat": "npm run build && npm run check && npm run package:check", From ed5b9f8d17bc6ed34d55b3f698c6f2988dde3a6a Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:15:45 +0800 Subject: [PATCH 25/32] Five teeth retire where a check already states their rule The retirement pass, first group. The criterion, applied one tooth at a time: a tooth may go when a named case's assertion fails for exactly the violation the tooth introduces, and that case's wording states the rule - so the ledger row names the check instead of the mutant. Judged by reading the assertion, then by re-running the target's sweep with the tooth gone. - the-board-read-path-stops-calling-the-predicate, the-board-decides-acceptance-on-its-own (row B5): tests/integration/ooo-acceptance-one-predicate.test.ts counts the call sites itself - one definition of the predicate in src/, one `acceptedFact({` in each reader, zero `verdict ===` comparisons in ooo-board.ts - so a rename that stops the call and a second decision inside it each fail a count. A behavioural differential would have caught neither; a count does. This pair is what made the criterion worth writing down. - a-refused-unit-is-silent: the case enumerates the 32 flag combinations a plan's facts can carry and asserts every refused unit has a reason. - the-merge-enumerates-one-order (row A5): the case asserts the multinomial count (10), which a merge returning one order fails. - a-second-entry-rebinds-the-task (row D12): the case asserts a retry is the same binding and a second entry is refused by name. Docs and code in one commit: the three rows now name the check that carries the rule and say when the tooth was retired (A5, B5, D12 in the obligations ledger, whose reading list gains the dated readings below), and the record's fourth plan item is marked landed with the criterion and the finding. Finding carried forward: a-refused-unit-is-silent is named by no ledger row, and the case that catches it is unowned too. Retiring it removed an orphan rather than a row's pin. Readings at this revision: mutation:anchors -> anchors: 144 of 144 resolve, over 23 targets, exit 0 (the register is 144 teeth, not 149); the four targets the pass touched -> mutants: 70 of 70 caught by the named test, restored byte-identically 5 of 5; test:product -> 1562 pass, 0 fail, exit 0; verify:static -> exit 0. --- ...-09-24-mutants-are-derived-not-anchored.md | 21 ++++++++- ...-mutants-are-derived-not-anchored.zh-CN.md | 8 +++- .../design/task-unit-semantics-obligations.md | 22 +++++----- tools/mutation-teeth.ts | 43 ------------------- 4 files changed, 39 insertions(+), 55 deletions(-) diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index c8d9a6f1..aaa104fe 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -118,7 +118,26 @@ severalWaits`, and the prediction made before running it - one dead anchor, seve sweep stays out of the gate. 4. Retirement pass over the ~18-20 teeth whose rule already has a relational or enumerative check, and over the I/E teeth that can be replaced; the ledger's `proven` sentence updated in the same commit. - **Not started.** + **First group landed (five teeth, 2026-09-24), and the criterion is the point of it:** a tooth may go + when a named case's assertion _fails for exactly the violation the tooth introduces_ and that case's + wording states the rule - so the row can name the check instead of the mutant. Applied by reading the + assertion and then re-running the target's sweep with the tooth gone. + - `the-board-read-path-stops-calling-the-predicate` and `the-board-decides-acceptance-on-its-own` + (row B5): the case counts the call sites itself - one definition of the predicate in `src/`, one + `acceptedFact({` in each reader, zero `verdict ===` comparisons in `ooo-board.ts` - so both + violations fail a count. This is the pair that made the criterion worth writing down: a + behavioural differential would not have caught either, a count does. + - `a-refused-unit-is-silent`: the case enumerates the 32 flag combinations a plan's facts can carry + and asserts every refused unit has a reason. + - `the-merge-enumerates-one-order` (row A5): the case asserts the multinomial count (10), which a + merge returning one order fails. + - `a-second-entry-rebinds-the-task` (row D12): the case asserts a retry is the same binding and a + second entry is refused by name. + Measured after the retirements: `anchors: 144 of 144 resolve` (the register is 144 teeth, not 149), + and the four targets the pass touched sweep `70 of 70 caught`, restored byte-identically 5 of 5. + One finding to carry forward: `a-refused-unit-is-silent` was named by **no** ledger row, and the case + that catches it is unowned too - an orphan tooth whose retirement removed the orphan rather than a + row's pin. **Remaining:** the rest of the ~18-20 candidates, and the I/E teeth. ## Acceptance criteria diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index cd0cb7e6..53d76430 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -40,7 +40,13 @@ 1. 本文,加上工具的算子词汇表与解析器;用“对源码字符串”而不是对文件系统的测试来证明它。**已落地:**解析器是 `tools/mutation-anchor.ts`(它不持有任何状态,所以从 sweep 脚本里搬出来,测试可以直接调用),16 个用例全部作用在源码字符串上,不碰文件系统。 2. 试点目标转换并重跑:同样的名字、同样的用例、同样的抓法;随后在该文件内做一次重构,证明派生出来的牙活过了字节锚点活不过去的那次改动。**已落地,带一处更正:**演示用的那次重构是“改一个局部变量名 + 把条件抬成 `const closed = spent || severalWaits`”,而运行前写下的预测——死一颗锚点、七个派生位置存活——**是错的**:死了两处,手写锚点和派生的 `a-live-claim-does-not-block-selection`,因为“绑定到名字上的条件”当时不算选择器眼里的一次决策位置。把条件抬成局部常量是维护者真会做的重构,于是选择器扩到了变量初始化式;更正后的预测(只有手写锚点会死)随后精确成立:`anchors: 148 of 149 resolve`,一处失败,文件逐字节还原。 3. anchors-only 这趟检查接入静态契约,并同步路由与设计文档。**已落地:**`npm run mutation:anchors`(`--anchors-only`)现在是 `verify:static` 的一条检查,也是 `ci-and-tests` 路由的一条原子检查,顺序与路由契约测试强制的一致。本版本实测:一次完整的 `npm run agent:verify` 里,149 颗锚点全部解析,用时 0.98 秒,整轮 112 秒——留在 150 秒预算之内;单独跑 `verify:static` 是 41 秒。全量 sweep 仍不进闸门。 -4. 对“规则已有关系式/穷举式检查”的那约 18-20 颗,以及能被替代的 I/E 牙,做退役;台账里 `proven` 那句话在同一提交里更新。**尚未开始。** +4. 对“规则已有关系式/穷举式检查”的那约 18-20 颗,以及能被替代的 I/E 牙,做退役;台账里 `proven` 那句话在同一提交里更新。**第一组已落地(五颗,2026-09-24),而准则本身才是这件事的要点:**当一条具名用例的断言*恰好因为这颗牙引入的违规*而失败、且该用例的措辞陈述了那条规则时,这颗牙才可以退役——于是台账行点名的变成那条检查,而不是 mutant。判定方式是先读断言,再把牙删掉后重跑该目标的 sweep。 + - `the-board-read-path-stops-calling-the-predicate` 与 `the-board-decides-acceptance-on-its-own`(B5 行):用例自己数调用点——`src/` 里谓词只定义一次、两个读取方各有一个 `acceptedFact({`、`ooo-board.ts` 里 `verdict ===` 出现零次——所以两种违规都让某个计数失败。正是这两颗让这条准则值得写下来:一个行为差异对比两颗都抓不住,计数能。 + - `a-refused-unit-is-silent`:用例穷举了计划事实可能携带的 32 种 flag 组合,并断言每一个被拒单元都有理由。 + - `the-merge-enumerates-one-order`(A5 行):用例断言多重组合计数(10),而只返回一种顺序的合并会失败。 + - `a-second-entry-rebinds-the-task`(D12 行):用例断言重试是同一个绑定、且第二条 entry 被按名拒绝。 + 退役后实测:`anchors: 144 of 144 resolve`(登记处是 144 颗牙,不是 149),这趟碰到的四个目标 sweep `70 of 70 caught`,5 of 5 逐字节还原。 + 一个要带下去的发见:`a-refused-unit-is-silent` **没有任何**台账行点名它,抓住它的那条用例同样无人认领——一颗孤儿牙,退役它移除的是孤儿,而不是某一行的钉子。**尚未完成:**其余约 18-20 个候选,以及 I/E 牙。 ## 验收标准 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index caf0a454..93f26648 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -20,7 +20,9 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth` -> 136 of 136 caught by the named test, 20 of 20 targets restored byte-identically, exit 0. Run on 2026-09-19 as four lanes, one sweep per tree, using `git worktree add --detach` on the same commit for three of them: a sweep is sequential _within_ a tree because its mutants substitute into the same file, and parallel across trees, where each lane also gets the isolation property that no lane's suites can read another lane's mutant. Lanes: 42 of 42 (`ooo-board`, `task-coordinator`), 41 of 41 (`base`, `ooo-execution`'s 17, `task-semantics-interleavings`), 42 of 42 (thirteen small targets) and 11 of 11 (`plan-driver`) - the last serialised into its own lane because its suite has a 25 s case and does real candidate verification (~92 s per run, against ~2 s for the cheap suites). **That lane's cost has since changed**: the arms' checks are data checks now, so at `c01d3fe` its clean run is 5.5 s and its five mutants are caught by the case each names in 0.8-1.2 s on the same target; the numbers in this paragraph describe the 2026-09-19 instrument. Three things the run itself taught, all fixed and pinned afterwards: the lock's `live` flag was never written on substitution (a multi-hunk edit failed as a whole and only the restore half was reapplied), so the field lied about a running sweep; `NODE_TEST_CONTEXT` inherited when a sweep is started from inside a `node --test` process made the nested runner exit 0 having run no test at all, which the harness reported as "the suite passed" and which turned every mutant of that target into a false "not caught"; and the refusal in `agent:verify` fired on `--dry-run` too, which made two of the verifier's own tests fail while a sweep held the tree (a dry run reads the plan and the route config, not the mutated file, so it is exempt now). The lane that reported a clean-run failure (`tools/agent-verify.ts`) had found the last of those three. A sweep is also refused while any lock is present, including one whose owner died, because a killed sweep leaves its mutant in the target (post-mortem 0003). Interruption note: two lane processes were killed by the console that launched them and were relaunched; the JSON each run writes at its end survived even when the buffered stdout summary was lost, so the lane results above were read from those files rather than from stdout. (the retirement pass added the driver's interleaving mutant; was 111 of 111 before this pass: `src/integration/task-semantics-interleavings.ts` gained three budget mutants and `evals/ooo-execution/plan-driver.ts` three for the per-unit checks, the canned worker and the parent composition). How these runs are scheduled (scoped during a change, full before a push, detached with a collected result) is a standing rule of the repository now, in [`skills/repo-development/SKILL.md`](../../skills/repo-development/SKILL.md) with its measured costs in [the decision](../decisions/implemented/2026-09-18-detached-long-checks.md) - `npm run complexity:gate` -> exit 0, 18 methods above 15 unchanged from baseline. It caught the E pass's first version (adding the advisers option pushed `BoardAdmission`'s constructor to 17, so the options check moved into `admissionAdvice()`) rather than the threshold being raised; the slot pass added its option check the same way (`admissionSlots`), and the ordered-set/`publishReady` reads stayed under it. - `node --experimental-strip-types --test tests/tools/mutation-anchor.test.ts` -> 16 pass, 0 fail, exit 0 (2026-09-24: the resolver's six operators and its refusals, over source strings, no filesystem) -- `npm run mutation:teeth -- --anchors-only` -> `anchors: 149 of 149 resolve, over 23 targets`, exit 0 in 0.6 s (2026-09-24: every tooth still applies where it is aimed; one anchor is re-taken through whitespace normalization, `the-completion-does-not-bind-the-verdict-to-the-bytes`) +- `npm run mutation:teeth -- --anchors-only` -> `anchors: 149 of 149 resolve, over 23 targets`, exit 0 in 0.6 s +- `npm run mutation:teeth -- --anchors-only` -> `anchors: 144 of 144 resolve, over 23 targets`, exit 0 in 0.7 s (2026-09-24, after five teeth were retired in favour of the checks that already state their rules) +- `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) (2026-09-24: every tooth still applies where it is aimed; one anchor is re-taken through whitespace normalization, `the-completion-does-not-bind-the-verdict-to-the-bytes`) - `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) - `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. - `node --experimental-strip-types --test --test-concurrency=4 tests/integration/ooo-ordinary-failure.test.ts tests/integration/ooo-managed-fence.test.ts tests/integration/ooo-read-paths-agree.test.ts tests/integration/ooo-round-query.test.ts tests/integration/ooo-task-tables.test.ts` -> 5, 3, 1, 2 and 4 pass, 0 fail, exit 0 @@ -126,13 +128,13 @@ separate claims. ## A. Offline semantics (the design's first slice, already landed) -| node | obligation | state | evidence | -| ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| A1 | The compiler returns a legal plan or a refusal naming task, field and obligation | proven | `tests/integration/task-semantics.test.ts` | -| A2 | The offline model dispatches without publishing or storing run state | proven | `tests/integration/task-semantics-model.test.ts` | -| A3 | The design's six discriminating cases are executable | proven | `tests/integration/task-semantics-cases.test.ts` | -| A4 | Field mapping catches budget-unit confusion, requires/deps confusion and silently dropped keys | proven | mutants `budget-inner-alias-is-dropped`, `over-maximum-budget-is-accepted`, `assumption-may-carry-a-dependency` | -| A5 | Every legal event interleaving of at most four units is enumerated, and every publication at every prefix is checked for its obligations, its inputs and its source | proven | `tests/integration/ooo-publication-invariants.test.ts` (18 cases) over `src/integration/task-semantics-interleavings.ts`. The enumeration is a merge of per-unit scripts, capped at four units and six events, and the checks read the derived view rather than a second model: a dispatch must rest on closed, accepted inputs, a completion (the accepted set closed over the same predicate) must have its verdict bound to the bytes, its bytes present, no cancellation, and inputs that are present and accepted. Both halves are pinned: each checker condition is deleted by a mutant and caught by the case that names it - `the-completion-does-not-bind-the-verdict-to-the-bytes`, `the-input-is-not-required-to-be-accepted`, `the-input-may-be-cancelled`, `the-completion-ignores-a-cancelled-unit`, `the-dispatch-does-not-require-a-closed-input`, `the-merge-enumerates-one-order` - and the clean run over the design's script sets reports nothing. The design's caveat stands and is quoted in the module: passing says this finite model satisfies the listed properties, not that any Agent program is correct | +| node | obligation | state | evidence | +| ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| A1 | The compiler returns a legal plan or a refusal naming task, field and obligation | proven | `tests/integration/task-semantics.test.ts` | +| A2 | The offline model dispatches without publishing or storing run state | proven | `tests/integration/task-semantics-model.test.ts` | +| A3 | The design's six discriminating cases are executable | proven | `tests/integration/task-semantics-cases.test.ts` | +| A4 | Field mapping catches budget-unit confusion, requires/deps confusion and silently dropped keys | proven | mutants `budget-inner-alias-is-dropped`, `over-maximum-budget-is-accepted`, `assumption-may-carry-a-dependency` | +| A5 | Every legal event interleaving of at most four units is enumerated, and every publication at every prefix is checked for its obligations, its inputs and its source | proven | `tests/integration/ooo-publication-invariants.test.ts` (18 cases) over `src/integration/task-semantics-interleavings.ts`. The enumeration is a merge of per-unit scripts, capped at four units and six events, and the checks read the derived view rather than a second model: a dispatch must rest on closed, accepted inputs, a completion (the accepted set closed over the same predicate) must have its verdict bound to the bytes, its bytes present, no cancellation, and inputs that are present and accepted. Both halves are pinned: each checker condition is deleted by a mutant and caught by the case that names it - `the-completion-does-not-bind-the-verdict-to-the-bytes`, `the-input-is-not-required-to-be-accepted`, `the-input-may-be-cancelled`, `the-completion-ignores-a-cancelled-unit`, `the-dispatch-does-not-require-a-closed-input` - and the clean run over the design's script sets reports nothing. The merge's completeness rests on the enumeration's own count assertion rather than on a mutant: the case `the merge enumerates every legal order, not one of them` asserts the multinomial count, which a merge returning one order fails, and the tooth that showed it (`the-merge-enumerates-one-order`) was retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). The design's caveat stands and is quoted in the module: passing says this finite model satisfies the listed properties, not that any Agent program is correct | ## B. Persistence (design: "进入持久化接入时另须证明") @@ -142,7 +144,7 @@ separate claims. | B2 | A retry does not deliver twice | proven | `tests/core/task-board-deliverable.test.ts`; mutants `stale-claim-may-deliver-again`, `deliverer-may-judge-its-own-work` | | B3 | A transaction failure cannot write a verdict without its association | proven | `tests/integration/ooo-transition-atomicity.test.ts`; mutants `a-method-opens-its-own-transaction`, `nested-write-transaction-is-allowed`, `swallowed-failure-still-commits` | | B4 | A retained board entry is not cleared by TTL | proven | `tests/core/task-board-retention.test.ts`; mutants `prune-ignores-retention`, `bounded-pin-never-expires` | -| B5 | The status query and dependency release call one predicate | proven | `tests/integration/ooo-acceptance-one-predicate.test.ts`; mutants `the-board-read-path-stops-calling-the-predicate`, `the-board-decides-acceptance-on-its-own` | +| B5 | The status query and dependency release call one predicate | proven | `tests/integration/ooo-acceptance-one-predicate.test.ts`: the case `acceptance has one home, and both readers reach it` counts the call sites itself - one definition of the predicate in `src/`, one `acceptedFact({` in each reader, zero `verdict ===` comparisons in `ooo-board.ts` - so a rename that stops the call and a second decision inside it each fail a count rather than a named mutant. The two teeth that showed both (`the-board-read-path-stops-calling-the-predicate`, `the-board-decides-acceptance-on-its-own`) were retired on 2026-09-24, once the counts stated the rule directly | | B6 | A generic board write cannot move a managed round's own state, cannot take the live claim it holds, and cannot move an entry a run has adopted outside that run's transition | proven | the state half is structural and asserted (`tests/integration/ooo-managed-fence.test.ts`: the round's owner, attempt, acceptance and terminal reason stay its own facts, and a generic claim on its entry changes none of them). The claim it holds is not structural and is tooth-backed: `a-live-claim-can-be-taken-by-another-agent` - the store's CAS stops requiring the holder to be the claimant, and the same test's direct second reader then succeeds. The design's owed sentence - a generic write on a managed entry is applied _inside_ the coordinating transaction, and `judge`/`resolve` cannot go around the run's fence - is implemented and pinned by `tests/integration/ooo-managed-write.test.ts` (6 cases): a direct verb on an adopted entry is refused and the entry does not move, while an entry no run adopts takes the path it always did; a coordinated write lands the board transition and the run's fact in one transaction and a failure after the board write leaves neither; a cancelled run takes no further lifecycle writes; a write for the wrong run, or for an entry no run adopted, is refused before anything is written; and the daemon's own `claim` verb routes an adopted entry through the transition, writing the fact with its own store. Mutants: `a-managed-entry-ignores-the-coordinated-scope` (the store's guard), `the-daemon-verb-skips-the-coordinated-path` (the routing), `a-coordinated-write-skips-its-run-fact`, `a-cancelled-run-still-accepts-writes`, `a-coordinated-write-skips-the-binding-recheck`. What is still the design's step 3 rather than this row: nothing adopts entries into a run yet, so the research drivers' direct writes are not refused today - they will be, and are meant to be, once the runner and thin adapters work through the coordinator | | B7 | JSONL export failure does not change the terminal state | not applicable | no JSONL export exists on this branch to fail | | B8 | The field-mapping checks catch the three confusions | proven | A4: the offline compiler owns this check | @@ -172,7 +174,7 @@ separate claims. | D9 | A post-commit notification failure is recorded, not thrown | proven | `tests/integration/ooo-post-commit-notification.test.ts`: a unit whose post-commit `afterCommit` throws still submits as `accepted`, its artifact is the board's, and `lastNotificationFailure()` reports the message, while the same unit with a reachable notification reports nothing. Verified by mutation (hand-run, 2026-09-18): deleting the `try`/`catch` in `submit()` makes the notification escape and the case fails. The implementer was always the board (`src/integration/ooo-board.ts`, `lastNotificationFailure()`); the evals suite that used to pin it drove a real round only because the artifact envelope is the board's business, and the round was retired ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The seam took three tries to find - `publishReady` is private and is called without a port at the commit, and `afterCommit` also runs on non-commit paths, so the failure is gated on a commit having landed | | D10 | `runCycle`'s early-cancel branch is reachable, or it is dead code | not applicable | **Not applicable: the branch went with the round.** It was reachable and pinned twice (`evals/ooo-execution/early-cancel.test.ts` failed with "database is not open" when the round closed the store it borrowed; verified by hand, restored byte-identically), and the retirement deleted both the branch and its subject ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The property it stood behind - nobody who borrows a store closes it - now holds by construction, because no client opens one: the daemon is the only writer and drivers borrow a served store (D14) | | D11 | The run records live in the store's schema, and their typed writes join the store's transaction | proven | `task_run_manifest`, `task_run_tasks` and `task_run_facts` are created by `migrate()` like every other table (`tests/core/store/schema.test.ts` names them), and `NmgStoreBase` writes them through `registerTaskRun`, `freezeTaskRunTask` and `appendTaskRunFact`, each taking the same optional port the board writes take: outside a transition it opens one, inside it joins the caller's. What the cases prove (`tests/core/store/task-runs.test.ts`, 8/8): a board write and a run fact written through one port land together, and when the fact write fails the board row does not survive it; a retried fact is recorded once and keeps its first sequence; two runs in one store keep separate task and fact namespaces; a run's plan cannot be replaced and a frozen task cannot be redefined; an entry bound by two runs is refused rather than answered with one of them; the reads register nothing. Mutants: `the-run-fact-opens-its-own-transaction`, `a-retried-run-fact-is-appended-twice`, `a-second-plan-overwrites-the-frozen-one`, `a-frozen-task-is-replaced-by-a-different-definition`, `a-second-run-adopts-a-bound-entry`. One layer deliberately has no mutant of its own: the `UNIQUE (run_id, kind, task_id, attempt)` constraint is a storage-level backstop behind the explicit identity check, so a mutant that dropped only the constraint would survive - the check is the deciding layer and is the one mutated above | -| D12 | The binding of a logical task to its board entry is a run fact, and one routing rule decides what a binding makes managed | proven | `bindRunEntry` in `src/integration/task-coordinator.ts` records the binding as the run fact `entry-bound` (the name the D11 store tests already used), and it is the binding that the store's fence reads, so a bound entry's lifecycle write is refused unless it runs inside the run's coordinated transition. The op joins a caller's transaction when given a port, which is what lets a host create the entry and bind it in one transition: `tests/integration/ooo-managed-adopt.test.ts` (5 cases) checks that the board never holds an entry the run does not (a failure after both writes leaves neither), that a retry on the same task and attempt is not a second binding while another entry for that task and attempt is refused rather than silently dropped, and that every refusal names a stored fact (unregistered run, cancelled run, a task the run never froze, an entry that is not on the named channel, an entry another run already carries). `coordinatedEntryWrite` moves to the same module as the single routing rule, so the daemon's board verbs and any in-process writer decide "managed" in one place instead of each caller's belief about the entry; `src/cli/service.ts` no longer keeps its own copy. Mutants: `an-adopted-entry-takes-the-direct-path`, `a-binding-ignores-whether-the-task-was-frozen`, `a-binding-does-not-check-the-entry-exists`, `one-entry-is-bound-to-two-tasks`, `a-second-entry-rebinds-the-task`, `the-daemon-verb-skips-the-coordinated-path` (re-anchored to the daemon's claim handler, since the routing it mutates moved out of that file) | +| D12 | The binding of a logical task to its board entry is a run fact, and one routing rule decides what a binding makes managed | proven | `bindRunEntry` in `src/integration/task-coordinator.ts` records the binding as the run fact `entry-bound` (the name the D11 store tests already used), and it is the binding that the store's fence reads, so a bound entry's lifecycle write is refused unless it runs inside the run's coordinated transition. The op joins a caller's transaction when given a port, which is what lets a host create the entry and bind it in one transition: `tests/integration/ooo-managed-adopt.test.ts` (5 cases) checks that the board never holds an entry the run does not (a failure after both writes leaves neither), that a retry on the same task and attempt is not a second binding while another entry for that task and attempt is refused rather than silently dropped, and that every refusal names a stored fact (unregistered run, cancelled run, a task the run never froze, an entry that is not on the named channel, an entry another run already carries). The idempotence half rests on that case rather than on a mutant - `a-second-entry-rebinds-the-task` was retired on 2026-09-24 - because the case asserts the retry's shape and the refusal by name. `coordinatedEntryWrite` moves to the same module as the single routing rule, so the daemon's board verbs and any in-process writer decide "managed" in one place instead of each caller's belief about the entry; `src/cli/service.ts` no longer keeps its own copy. Mutants: `an-adopted-entry-takes-the-direct-path`, `a-binding-ignores-whether-the-task-was-frozen`, `a-binding-does-not-check-the-entry-exists`, `one-entry-is-bound-to-two-tasks`, `the-daemon-verb-skips-the-coordinated-path` (re-anchored to the daemon's claim handler, since the routing it mutates moved out of that file) | | D13 | The run surface another process reaches is the daemon's: register, freeze, bind, cancel and status, and nothing else writes run facts | proven | `taskRun` in `src/cli/protocol.ts` over `registerRun` / `freezeRunPlan` / `bindRunEntry` / `cancelRun` / `taskRunStatus` / `createBoundEntry` in `src/integration/task-coordinator.ts`; `tests/cli/task-run-surface.test.ts` (11 cases) drives all of it through `service.invoke`, so what it exercises is the wire a second process uses, and it asserts `hello.methods` carries `taskRun` - the advertisement a client gates the `adopt` field on. Registration is the only transition that does not need a run to exist already (`coordinateRunWrite` refuses an unregistered run, which is what every later transition and every managed write rests on). A plan freeze is one transition: the request's array order becomes the positions, a dangling dependency and a self-dependency are refused by name while the plan is still only a proposal, and a batch the store refuses leaves the earlier tasks unfrozen - `the-plan-freezes-one-task-per-transaction`, `every-task-is-frozen-at-position-zero`, `the-plan-may-freeze-a-dangling-dependency`, `a-task-may-depend-on-itself` - and a cancelled run takes no further plan (`a-cancelled-run-takes-a-new-plan`). Adoption rides the transition that creates the entry (`createBoundEntry` takes the port), so a refusal leaves no entry behind (`the-entry-is-created-before-its-binding-is-checked`) and the wire cannot drop the request silently (`the-wire-drops-an-adoption-request`), which is the one failure mode the epoch rule in `design.md` is about. The binding fact records the channel that names its entry - that is what lets `status` resolve a binding without searching the board (`a-binding-does-not-record-its-channel`). `cancel` is the only writer of `run-cancelled`: before this row `src/` had none, so the state the fence and the dispatch derivation both read was reachable only from a test (the gap G4 named); a run-level cancellation carries the schema's empty task id (`a-run-cancellation-names-a-task`), a task-level one requires the task to have been frozen (`a-cancellation-ignores-whether-the-task-was-frozen`, `an-unknown-run-can-be-cancelled`), and cancelling twice is one fact because the fact's own identity is the duplicate key. `status` is a read and derives nothing: a run the store does not know stays unknown (`status-registers-the-run-it-cannot-find`), and ready/blocked/accepted stay with the shared pure function rather than with this view. The CLI exposes the operator half (`nmg run status`, `nmg run cancel`), which is also what the "every RPC method is a CLI command or intentionally RPC-only" gate asks for. One thing this pass corrected about itself: the first version's parser also computed a `position` per task, which `freezeRunPlan` immediately overwrote - the mutant that changed it survived, which is how the dead field was found, and it was deleted rather than kept | | D14 | The evidence drivers reach the board through the daemon and open no database of their own | proven | `evals/ooo-execution/round-client.ts` resolves the daemon from the store's own lease and refuses by name when nothing serves it; `round-host.ts` serves an existing round store the way the product daemon does (`NmgService` + `serveHttp` + the store's lease) and is that store's only writer while it runs; `board-worker.ts`, `board-deliver.ts` and `board-judge.ts` now use the protocol's board verbs (`read`, `claim`, `release`, `deliver`, `judge`) and no longer import the store. `tests/integration/ooo-evidence-drivers.test.ts`: both end-to-end cases host the store and spawn the drivers as separate processes against the endpoint (spawned, not `spawnSync`, because the test process is what answers them); the managed case registers a run and freezes its plan through `taskRun`, adopts the entry in the same `put` that creates it, then runs the worker and the judge as separate processes - the run's log afterwards is exactly `entry-bound`, `board-claim`, `board-deliver`, `board-judge`, with the binding resolved to the entry the worker claimed and its verdict accepted. Two cases hold the boundary itself: a driver with no daemon refuses by name instead of falling back to the file, and no `board-*.ts` driver may mention the store or omit the round client. Mutants: `a-driver-falls-back-to-opening-the-store`, `the-round-client-does-not-require-a-daemon`, plus the drivers' two earlier ones (`worker-reads-only-one-reporter-shape`, `judge-may-judge-its-own-delivery`), re-pointed at the renamed case. Two further rules came out of building this, and both are about the same defect the first attempt hit: a process that serves an endpoint cannot answer a call to it while it is blocked, so "the answer will come" is not an assumption a client may make. `round-client.ts` therefore bounds every call (`ROUND_CALL_TIMEOUT_MS`, an opt-in bound on the thin client's `httpCall`, which by default still leaves the platform's own) and reports a timeout by naming the bound, the endpoint and the pid that serves it rather than the transport's; and it refuses when the lease's pid is the caller's own, directing an in-process host to `host.call(...)` (`round-host.ts`), which is the design's shape for an offline host. The test is now a client too: it starts the host as its own process (`[host] serving `, idle timeout as the backstop for a host whose owner died) instead of hosting in-process, and asks it to shut down rather than killing it, so what runs on the way out is the release path - a lease held by a process that is gone is a store nothing can serve. Cases: "a client refuses to call the endpoint its own process serves", "a call to a host that never answers gives up in seconds and names the reason" (raced against its own 5s deadline, so the check of the bound is itself bounded), "a host releases its lease when it stops, so the next host can take the store". Mutants: `the-round-client-calls-the-endpoint-it-serves`, `the-round-client-has-no-limit-on-how-long-it-waits`, and - new target, `src/cli/http-server.ts` - `the-serving-process-never-releases-its-lease`. Not covered by this row: `live-continuation.ts` still creates its own store and constructs `BoardAdmission` - the runner half of the same design sentence and the next obligation. `round-runner.ts` did the same and was retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)) | diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 6a668344..3602b142 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -300,22 +300,6 @@ const TARGETS: readonly Target[] = [ to: "const db = new BoardAdmission(databasePath) as unknown as DatabaseSync;", expect: "the query port reads a round without migrating, publishing or exposing a write", }, - { - // Located inside the reader: a rename of the call proves the check counts call sites. - name: "the-board-read-path-stops-calling-the-predicate", - ast: { within: "readAccepted" }, - from: "acceptedFact({", - to: "locallyAccepted({", - expect: "acceptance has one home, and both readers reach it", - }, - { - // A bypass that still type-checks: the suite must notice the second decision. - name: "the-board-decides-acceptance-on-its-own", - ast: { within: "readAccepted" }, - from: "!acceptedFact({", - to: '!(recorded?.verdict === "accepted" ? false : true) && !acceptedFact({', - expect: "acceptance has one home, and both readers reach it", - }, { // The claim is the write that would corrupt a neighbour run: the same task id exists in // every run, so a claim that is not scoped by run claims somebody else's row too. @@ -597,16 +581,6 @@ const TARGETS: readonly Target[] = [ }, expect: "the round's own answer is the shared rule's answer, not an ordering's", }, - { - // The answer a caller asks for carries the rule's own reasoning, so a refused unit that has - // no cause of its own is still named by the plan-level gate that held it back. This mutant - // silences every refusal at once: an empty legal set with no reasons is exactly the answer - // the design says a caller cannot be given. - name: "a-refused-unit-is-silent", - derive: { within: "selection", operator: "replace-argument", call: "causes.set", arg: 1 }, - to: "[]", - expect: "no refusal in the answer is silent, over the flags a plan's facts can carry", - }, { // A cause names the gate; it never restates the condition. Dropping one leaves a refusal // whose reason the rule can no longer give, even when another gate would still be true. @@ -773,15 +747,6 @@ const TARGETS: readonly Target[] = [ to: " void checkInputs;", expect: "the checker reports a dispatch whose input is not accepted", }, - { - // The enumeration is the other half of the claim: a merge that stops at the first order - // checks one interleaving and reports it as all of them. - name: "the-merge-enumerates-one-order", - ast: { within: "interleavings" }, - from: " return out;", - to: " return out.slice(0, 1);", - expect: "the merge enumerates every legal order, not one of them", - }, ], }, { @@ -1120,14 +1085,6 @@ const TARGETS: readonly Target[] = [ to: " if (false && bound && (bound.runId !== request.runId || bound.taskId !== request.taskId))", expect: "a binding refuses what the store does not hold", }, - { - // A second entry for the same task and attempt is a disagreement. The stored fact is keyed - // by task and attempt, so accepting it would keep the first binding and report the second. - name: "a-second-entry-rebinds-the-task", - from: " if (existing && existing.entryId !== request.entryId)", - to: " if (false && existing && existing.entryId !== request.entryId)", - expect: "a binding is idempotent for its task and attempt, and refuses a second entry", - }, { // The transition is the run's record of what happened to its entry; without it the board // moved and the run has nothing to read. From 4f50675664320616d0a7f8bf3d7df42e85351d91 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:26:03 +0800 Subject: [PATCH 26/32] Six more teeth retire where one case states their rules Second group of the retirement pass, same criterion as the first: a tooth may go when a named case's assertion fails for exactly the violation the tooth introduces, and that case's wording states the rule. - The five rules of a declared budget - the licence is the budget's part and not its head, a budget above one must name each handoff's target, every startable task gets a handoff, a startable handoff is not retired by a republish, and a multi-slot run directs its handoffs - are all stated by one case, evals/ooo-execution/board-slots.test.ts's "a declared budget holds two claims at once, and the store is why each handoff is directed". Its assertions name each rule: the refusal message, one handoff per startable task, serialState null for each, the same ids across a republish, and the non-head claimed first. So the-licence-is-the-head-whatever-the-budget, a-second-slot-is-declared-without-a-target, only-the-heads-handoff-is-published, a-startable-handoff-is-retired-as-unselected and a-multi-slot-handoff-is-published-un-directed are retired. The F2b-slot row keeps seven teeth, and the rules themselves keep a witness that says them. - claim-is-not-scoped-to-its-run (row B1) is the tooth the namespace experiment first watched survive: tests/integration/ooo-run-namespace.test.ts was strengthened until it caught it, and that record names the assertion that did it (a raw-row read of the other run's owner). A case whose history is exactly this tooth is a better witness than the tooth's own name. Docs and code in one commit: rows B1 and F2b-slot now name the checks that carry their rules and say when each tooth was retired, the obligations ledger gains these dated readings, and the record's fourth plan item is marked with this group, the two candidates judged not retirable so far, and the two stale expects still outstanding. Readings at this revision: mutation:anchors -> anchors: 138 of 138 resolve, over 23 targets, exit 0; --targets=src/integration/ooo-board.ts -> mutants: 15 of 15 caught by the named test, restored byte-identically; test:product -> 1562 pass, 0 fail, exit 0; verify:static -> exit 0. One reading to carry, not a code failure: the first test:product run of this pass failed tests/core/graph-cycles.test.ts's first case on EPERM from rmSync of its %TEMP% scratch directory, and the file then passed 8 of 8 alone and the suite 1562 of 1562. That is the Windows temp-tree removal class the ledger already records as undiagnosed, and the ledger now says so with this instance. --- ...-09-24-mutants-are-derived-not-anchored.md | 20 +++++- ...-mutants-are-derived-not-anchored.zh-CN.md | 4 +- .../design/task-unit-semantics-obligations.md | 7 ++- tools/mutation-teeth.ts | 62 ------------------- 4 files changed, 26 insertions(+), 67 deletions(-) diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index aaa104fe..295706ca 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -137,7 +137,25 @@ severalWaits`, and the prediction made before running it - one dead anchor, seve and the four targets the pass touched sweep `70 of 70 caught`, restored byte-identically 5 of 5. One finding to carry forward: `a-refused-unit-is-silent` was named by **no** ledger row, and the case that catches it is unowned too - an orphan tooth whose retirement removed the orphan rather than a - row's pin. **Remaining:** the rest of the ~18-20 candidates, and the I/E teeth. + row's pin. + **Second group landed (six teeth, 2026-09-24):** the five rules of the declared budget - the licence + being the budget's part rather than the head, a budget above one requiring a target, one handoff per + startable task, a startable handoff surviving a republish, and a multi-slot handoff being directed - + are all stated by one case, `evals/ooo-execution/board-slots.test.ts`'s `a declared budget holds two +claims at once, and the store is why each handoff is directed`, whose assertions name each rule (the + refusal message, the handoff count, the null `serialState`, the identity across a republish, and the + non-head claimed first). They were retired from the F2b-slot row, which keeps seven teeth. The sixth, + `claim-is-not-scoped-to-its-run` (row B1), is the tooth the namespace experiment first saw survive - + the case was strengthened until it caught it, and the record documents the raw-row read that was + needed, which makes that case a better witness than the tooth's name. Measured after: + `anchors: 138 of 138 resolve`, and `--targets=src/integration/ooo-board.ts` sweeps `15 of 15 caught`, + restored byte-identically. + **Remaining:** the I/E teeth, the store-target teeth (`a-retried-run-fact-is-appended-twice`, + `a-frozen-task-is-replaced-by-a-different-definition`), and the two candidates judged **not** + retirable so far - `every-task-is-frozen-at-position-zero` and `a-binding-does-not-record-its-channel`, + whose named case is a register/freeze/adopt/read-back round-trip, weaker than the rule it is asked to + carry. The stale `expect` on `next-is-not-the-head-of-the-ordered-candidates` and on + `fusion-continues-from-an-unverified-answer` is still outstanding. ## Acceptance criteria diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index 53d76430..459f671a 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -46,7 +46,9 @@ - `the-merge-enumerates-one-order`(A5 行):用例断言多重组合计数(10),而只返回一种顺序的合并会失败。 - `a-second-entry-rebinds-the-task`(D12 行):用例断言重试是同一个绑定、且第二条 entry 被按名拒绝。 退役后实测:`anchors: 144 of 144 resolve`(登记处是 144 颗牙,不是 149),这趟碰到的四个目标 sweep `70 of 70 caught`,5 of 5 逐字节还原。 - 一个要带下去的发见:`a-refused-unit-is-silent` **没有任何**台账行点名它,抓住它的那条用例同样无人认领——一颗孤儿牙,退役它移除的是孤儿,而不是某一行的钉子。**尚未完成:**其余约 18-20 个候选,以及 I/E 牙。 + 一个要带下去的发见:`a-refused-unit-is-silent` **没有任何**台账行点名它,抓住它的那条用例同样无人认领——一颗孤儿牙,退役它移除的是孤儿,而不是某一行的钉子。 + **第二组已落地(六颗,2026-09-24):**声明的预算那五条规则——许可的是预算的那一部分而不是表头、预算大于一必须点名目标、每个可开始任务都有一个 handoff、可开始的 handoff 会在重发布中存活、多槽位的 handoff 是被定向的——都由同一条用例陈述:`evals/ooo-execution/board-slots.test.ts` 的 `a declared budget holds two claims at once, and the store is why each handoff is directed`,其断言逐条点名了这些规则(拒绝消息、handoff 计数、`serialState` 为 null、重发布后 id 不变、先认领非表头那个)。这五颗从 F2b-slot 行退役,该行保留七颗。第六颗 `claim-is-not-scoped-to-its-run`(B1 行)是命名空间实验里最初**没被抓住**的那颗牙——用例被加强到能抓住它,实验记录写下了当时必需的"裸行读取",所以那条用例比这颗牙的名字是更好的证人。退役后实测:`anchors: 138 of 138 resolve`,`--targets=src/integration/ooo-board.ts` sweep `15 of 15 caught`,逐字节还原。 + **尚未完成:**I/E 牙、store 目标那两颗(`a-retried-run-fact-is-appended-twice`、`a-frozen-task-is-replaced-by-a-different-definition`),以及目前判定为**不可退役**的两颗——`every-task-is-frozen-at-position-zero` 与 `a-binding-does-not-record-its-channel`,它们的点名用例是"注册/冻结/采纳/读回"的往返,弱于要它承担的规则。`next-is-not-the-head-of-the-ordered-candidates` 与 `fusion-continues-from-an-unverified-answer` 的过期 `expect` 仍未处理。 ## 验收标准 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index 93f26648..2c747293 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -22,6 +22,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `node --experimental-strip-types --test tests/tools/mutation-anchor.test.ts` -> 16 pass, 0 fail, exit 0 (2026-09-24: the resolver's six operators and its refusals, over source strings, no filesystem) - `npm run mutation:teeth -- --anchors-only` -> `anchors: 149 of 149 resolve, over 23 targets`, exit 0 in 0.6 s - `npm run mutation:teeth -- --anchors-only` -> `anchors: 144 of 144 resolve, over 23 targets`, exit 0 in 0.7 s (2026-09-24, after five teeth were retired in favour of the checks that already state their rules) +- `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts` -> `mutants: 15 of 15 caught by the named test`, restored byte-identically (2026-09-24, after six more teeth were retired there: the five budget/publication rules and the run-scoping claim) - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) (2026-09-24: every tooth still applies where it is aimed; one anchor is re-taken through whitespace normalization, `the-completion-does-not-bind-the-verdict-to-the-bytes`) - `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) - `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. @@ -140,7 +141,7 @@ separate claims. | node | obligation | state | evidence | | ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| B1 | Two rounds with the same taskId do not collide | proven | `tests/integration/ooo-run-namespace.test.ts`; mutant `claim-is-not-scoped-to-its-run` | +| B1 | Two rounds with the same taskId do not collide | proven | `tests/integration/ooo-run-namespace.test.ts`: the case `two runs in one store do not collide, do not see each other, and cancel separately` asserts each half directly - a claim in one run leaves the other run's row for the same task id alone (read as a raw row, which is the assertion the namespace experiment record shows was needed), a cancellation is per run, and a reopen by name sees the right one. The tooth `claim-is-not-scoped-to-its-run` was retired on 2026-09-24, once that case stated the rule ([record](../experiments/execution/ooo-run-namespace-2026-09-13.md), [decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)) | | B2 | A retry does not deliver twice | proven | `tests/core/task-board-deliverable.test.ts`; mutants `stale-claim-may-deliver-again`, `deliverer-may-judge-its-own-work` | | B3 | A transaction failure cannot write a verdict without its association | proven | `tests/integration/ooo-transition-atomicity.test.ts`; mutants `a-method-opens-its-own-transaction`, `nested-write-transaction-is-allowed`, `swallowed-failure-still-commits` | | B4 | A retained board entry is not cleared by TTL | proven | `tests/core/task-board-retention.test.ts`; mutants `prune-ignores-retention`, `bounded-pin-never-expires` | @@ -251,7 +252,7 @@ and the paid stage then ran on a family held out of it. | F2a: the plan has one home, and the round's log names it | landed | the plan is one value with one home: the spec the driver runs (`PlanDriverSpec.plan`) and the run manifest the store freezes (D12). F2a's round-side carriers (`DEFAULT_ROUND_PLAN`, `CycleOptions.plan`, `openRoundStore(path, plan)`, `round-plan.test.ts` with its 6 cases) were retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); [the arms' driver decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) is what made the driver the home in the first place | | F2b: a research-side driver for arbitrary legal plans | landed, one arm | `evals/ooo-execution/plan-driver.ts` (`runPlan`, `comparePlanSlots`, `verifyParent` path, refusal-naming CLI) + `plan-driver.test.ts` (15 cases) + 12 named mutants (`tools/mutation-teeth.ts`, target `evals/ooo-execution/plan-driver.ts`) + `BoardAdmission.candidates()` (the ordered legal set; `next()` is its head, with its own mutant) + [the decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) | | F2c: pick the parent task from the sweep's turning point | landed | Two families of the shape the design's parent family needs - a frozen interface, three independent builders, one summary that depends on all three - as a spec pair each: `evals/ooo-execution/fixtures/report/` (the instrument's own) and `fixtures/pipeline/` (**held out**: written after the driver, and not the family anything was tuned against). The coarse spec is one unit over the whole task checked by every frozen test; the fine spec is four units with per-unit checks and no `join`, so the parent check composes all four. Sibling units are code-independent by construction (the summary takes the derived values as parameters), which is what lets a unit be verified before its siblings exist. Offline, with the instrument's own answers and a wrong one: `evals/ooo-execution/families.test.ts` (8 cases over both families) - both plans accept, a wrong answer is rejected by its own check and takes the composition with it, and a unit that declares no checks and has no file-wide list is refused by name. Driver support this needed: per-unit `checks` with the file-wide list as a fallback, a `canned` worker (the instrument's answer, so the family's acceptance is shown before any model is paid), and the parent composition fix recorded in the pilot's experiment record | -| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected`, `a-multi-slot-handoff-is-published-un-directed`, `the-driver-declares-one-slot-whatever-the-spec-says`, `the-driver-awaits-each-unit-instead-of-the-batch`. [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | +| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-driver-declares-one-slot-whatever-the-spec-says`, `the-driver-awaits-each-unit-instead-of-the-batch`. The budget's licence and publication rules are pinned by `evals/ooo-execution/board-slots.test.ts`'s own case instead: `a declared budget holds two claims at once, and the store is why each handoff is directed` refuses a budget above one with no target by that message, asserts one handoff per startable task with `serialState` null for each, keeps every startable handoff across a republish by identity, and claims the non-head first - so `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected` and `a-multi-slot-handoff-is-published-un-directed` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | | F3: real-model pilot (A 3 / B 3 / C 2, current pi model, directional only) | landed | `evals/ooo-execution/pilot.ts` (`--live` required, `--report` to re-aggregate recorded runs with no model call, refuses a merge of two instruments, seeded arm order) + [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md). Fixed: `deepseek/deepseek-v4-flash`, envelope limits `turns 6 / reads 3 / 120 s` for every arm, A 3 / B 3 / C 2, arms drawn from a seeded shuffle, the held-out `pipeline` family. Measured (8 runs, 23 model calls, every run complete, all 8 parent checks accepted): per run A 8.6 s / 11.2 k tokens, B 21.5 s / 31.5 k tokens, C 16.2 s / 29.9 k tokens, host 1.6 s / 3.4 s / 3.6 s. All three of F1's expectations held: B is 2.5× A in wall time and 2.8× in tokens (one slot buys nothing), C recovers part of it (0.76× B) and not the 2× a pure model-call overlap would give, and the host cost grows with candidates rather than slots. Wasted cost 0, human intervention 0. **Directional only**: n = 8, one model, one held-out family, and no quality difference was available to measure - every arm accepted everything | | F4: execution fusion - legality, accounting and the driver policy | landed, live arm measured | Rules: `sharedSessionLegal`/`fusionSuccessors`/`fusionCandidates` in `src/integration/ooo-execution.ts` (the design's five conditions, one line each, composed with the board's candidate answer) + `tests/integration/ooo-fusion.test.ts` (12 cases) + 8 named mutants (target `src/integration/ooo-execution.ts`). Accounting: `fusionAccounting`/`fusionVerdict` + a fusion block in `cost-model.ts --sweep` - two lines kept apart, `unmeasured` until a run prices the session startup + `cost-model.test.ts` (12 cases) + 4 named mutants. Policy: `PlanDriverSpec.fusion` + `PlanRun.sessions` in `evals/ooo-execution/plan-driver.ts`, with the board's candidate set as the authority on staleness/cancellation/delivery/waits + `plan-driver.test.ts` (15 cases) + 11 named mutants. Scoped sweeps on this revision: `src/integration/ooo-execution.ts` 17 of 17 caught, `evals/ooo-execution/cost-model.ts` 4 of 4, `src/core/store/clock.ts` 4 of 4, `evals/ooo-execution/plan-driver.ts` 11 of 11, each restored byte-identically. A driver mutant that only restated `sharedSessionLegal`'s own rule was deleted rather than kept: the suite could not distinguish it from the shared predicate, which is what one home for that rule means. Both functions this slice pushed above the complexity limit (`sharedSessionLegal` 18, `runOneUnit` 16) were brought under it by extracting helpers, not by raising the threshold; `npm run lint` and `npm run check` are clean on this revision. **The live half landed.** `createPiSessionRunner` holds one Pi session and its tool surface across units (each unit re-points one mutable `UnitState` box; `patchSessionInput` is the one place a unit's prompt, snapshot and bounds are built, and `executePiPatch`/`executePiSnapshot` are thin callers of it), `PiRun` separates a unit's own `tokens`/`cacheRead`/`cacheWrite` from the session's `sessionTokens`, and `piSessionWorker` + `--session-runner` hold one runner per driver session. First paid D arm (2026-09-19, `deepseek-v4-flash`, 2 units, `--slots 1`, `turns: 6`, 3 reps per bound, only `fusion.unitsPerSession` differing): fused ran one session of two units and the control two sessions of one, quality parity in all six runs (every unit accepted), median 22 533 against 22 498 tokens and 11 048 against 12 948 ms - so no token saving yet (~1.9 s per run, ~15 % of the unfused wall, which prices the session-startup term at ~1.9 s instead of leaving it assumed) and per unit the second one cost ~8 % less while the first cost more: a chain's tool surface is the union of its units' capabilities because a session's surface is fixed at creation, so a unit can spend a turn on a tool that refuses by name. Also fixed here: `specFrom` had silently dropped a spec file's `fusion` block, so a spec asking for fusion ran as the control arm. | | F5: speculation lifecycle - one declared fact, three outcomes | landed (offline); the paid E arm ran once, no gain claimed | `SpeculationAssumption` / `ResolvedPredicate` / `SpeculationCandidate` / `isBoundedSpeculation` / `speculationOutcome` in `src/integration/ooo-execution.ts`, beside the fusion conditions: the assumption is a declaration the summary binds to (it never discovers for itself that the guess was false), the first experiment's bounds are a predicate (exactly one pending fact, nothing prepared from the guess - a speculative successor or an irreversible operation each refuse it by name), and the outcome has three states rather than two - **true** publishes, **false** discards the candidate and returns `sessionReusable: false`, which is what makes "失效会话不能复用到真实路径" a rule the caller must honour instead of a note, and **unknown** (no reading, an unattested reading, or evidence about another version) waits without publishing. Asking for the outcome of a candidate that is not the bounded shape throws rather than folding a fourth state into the three. Cases: `tests/integration/ooo-speculation.test.ts` (9), the last of which joins this half to fusion's condition 5 - an invalidated branch is not a legal predecessor for the real path. Mutants: 5 (`speculation-guesses-several-facts-at-once`, `a-guess-with-no-evidence-publishes`, `an-unattested-reading-counts-as-evidence`, `evidence-about-another-version-is-the-same-fact`, `a-contradicted-guess-keeps-its-session`); the target's sweep is 22 of 22 caught. The E arm's instrument now exists (`evals/ooo-execution/speculation-pilot.ts`, registered) and ran once (2026-09-19, 6 paid units, ~43 k tokens): it decides the guessed fact, prepares the candidate, applies `speculationOutcome`, and verifies a published candidate with the unit's own frozen check. **No result is claimed**: the quality term was false in all four verified candidates, so by the design's own rule the latency and cost shape may not be reported as a gain. The search behind those failures is now closed and its first reading was wrong: `artifactEnvelope` builds two legitimate shapes (a patch, and a conclusion with `kind, conclusion, summary, evidence, citations`), and the instrument had fed every artifact to the patch reader. Eight of nine attempts answered with a conclusion, which this unit's check cannot pass and the board would refuse; the one patch attempt failed on a real mistake (`rows` for `lines`). The instrument now reads by kind, keeps every artifact, candidate tree and check output, and the run is archived. Measured outcome of the arm at this shape: the post-fact cost drops from ~6.2 s of work to 175 ms of verification when the fact holds, the false-fact case wastes 20 332 tokens, and the prepared candidate was publishable in 0 of 3 holding reps - so the cost is real, the gain is not, and the binding constraint is the candidate's admissibility | @@ -365,7 +366,7 @@ The harness note from that pass is closed: a failed clean run now reports the ob (`observedFailures` in `tools/mutation-teeth.ts`, main's #63 - this branch takes it by merging main, not by a change of its own), so the verdict is diagnosable instead of repeated. The one-in-four flake the note also recorded was never diagnosed here; the Windows temp-tree removal and shadow-lock flakes -main repaired in #63/#64 are the nearest known causes, and they arrive in the same merge. +main repaired in #63/#64 are the nearest known causes, and they arrive in the same merge. It surfaced once more on 2026-09-24, in the retirement pass: `tests/core/graph-cycles.test.ts`'s first case failed on `EPERM` from `rmSync` of its scratch directory under `%TEMP%`, and the file passed 8 of 8 when re-run alone - so the reading to carry is that this class is still live and still not a failure of the code under test. **Closed since (same branch, later pass).** Everything the bullet list above left open is now built rather than described, and two of its statements were corrected by doing it: the store gained diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 3602b142..9c1557bc 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -300,16 +300,6 @@ const TARGETS: readonly Target[] = [ to: "const db = new BoardAdmission(databasePath) as unknown as DatabaseSync;", expect: "the query port reads a round without migrating, publishing or exposing a write", }, - { - // The claim is the write that would corrupt a neighbour run: the same task id exists in - // every run, so a claim that is not scoped by run claims somebody else's row too. - name: "claim-is-not-scoped-to-its-run", - ast: { within: "claim" }, - from: ' "UPDATE ooo_probe_facts SET attempt=?, owner=?, claim_time=? WHERE run_id=? AND id=?",', - to: ' "UPDATE ooo_probe_facts SET attempt=?, owner=?, claim_time=? WHERE ? IS NOT NULL AND id=?",', - expect: - "two runs in one store do not collide, do not see each other, and cancel separately", - }, { // The composed write must join the transition it is called in: a publication that opens its // own boundary commits even when the transition around it fails. @@ -429,58 +419,6 @@ const TARGETS: readonly Target[] = [ expect: "contract: a verified patch candidate is what dependents bind to, and only acceptance releases them", }, - { - // The licence is the budget's part of the ordered set, not its head: with two declared slots a - // second task is claimable, and with one it is not. Narrowing it back to the head is exactly - // the rule the C arm could not cross. - name: "the-licence-is-the-head-whatever-the-budget", - ast: { within: "claimableRow" }, - from: ' if (!this.startable().includes(id)) throw new Error("task not selected by narrow dispatch");', - to: ' if (this.next() !== id) throw new Error("task not selected by narrow dispatch");', - expect: - "a declared budget holds two claims at once, and the store is why each handoff is directed", - }, - { - // A budget above one is unusable without a target, because the store queues a second - // un-directed actionable entry behind the first. Accepting it silently would report two slots - // and deliver one. - name: "a-second-slot-is-declared-without-a-target", - ast: { within: "admissionSlots" }, - from: " if (slots > 1 && !options.handoffTarget)", - to: " if (false && slots > 1 && !options.handoffTarget)", - expect: - "a declared budget holds two claims at once, and the store is why each handoff is directed", - }, - { - // Every startable task gets its handoff, not only the head: publishing one is what makes the - // other slots claims rather than a promise. - name: "only-the-heads-handoff-is-published", - ast: { within: "publishReady" }, - from: " if (!startable.includes(row.id)) continue;", - to: " if (row.id !== startable[0]) continue;", - expect: - "a declared budget holds two claims at once, and the store is why each handoff is directed", - }, - { - // A startable handoff is not retired just because it is not the head: retiring it would - // withdraw the second slot's offer right after publishing it. - name: "a-startable-handoff-is-retired-as-unselected", - ast: { within: "publishReady" }, - from: " if (startable.includes(row.id) || this.live(row)) continue;", - to: " if (row.id === startable[0] || this.live(row)) continue;", - expect: - "a declared budget holds two claims at once, and the store is why each handoff is directed", - }, - { - // The store's serialization is why a multi-slot run directs its handoffs. Publishing them - // un-directed leaves the second one queued as `pending`, and its claim is refused. - name: "a-multi-slot-handoff-is-published-un-directed", - ast: { within: "publishReady" }, - from: " this.slots > 1 ? this.handoffTarget!(row.id) : undefined,", - to: " undefined,", - expect: - "a declared budget holds two claims at once, and the store is why each handoff is directed", - }, ], }, { From fff284ebc913b5b67350c84a6f9a5753834e9b57 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 16:27:21 +0800 Subject: [PATCH 27/32] Third group: two teeth retire, and four links were wrong Two retirements, same criterion as the first two groups: a-retried-run-fact-is-appended-twice and a-frozen-task-is-replaced-by-a-different-definition (row D11) go, because that row's prose already rested on the two cases for those rules and the cases state them outright ("appending the same fact twice records it once and keeps the first sequence", "freezing a task twice is a no-op, and a different definition for it is refused"). The group's larger finding is what the first whole-register sweep said: 136 of 136 caught, but only 134 caught by the case each expect named. Four links were wrong, each in its own way, and all four are repaired: - fusion-continues-from-an-unverified-answer: its named case refused the pair for a second reason as well (the successor declared the predecessor as a dependency, so the dependency rule could refuse it alone), and a case that passes for another reason pins nothing. The pair now declares no dependency, which leaves the predecessor's verdict as the only condition that can refuse it, and the tooth's own expect case now fails under the mutant. - next-is-not-the-head-of-the-ordered-candidates: its expect named a case in another target. The board's budget case now asserts next() is the head of the ordered set - the rule in its own words - and that is the name the tooth carries. - the-caller-rebuilds-the-shared-floor: its expect was a paraphrase that named no case at all. Re-pointed to "a declining route's narrow run verifies on its own tests and nothing else", which fails under it. - the-grace-is-zero: its named case derived its fixture from the constant under test (half = CLOCK_GRACE_MS / 2000), so zeroing the grace moved the stamp onto now and the case passed while the bug was live. The fixture is a literal now and the case fails under the mutant again. Re-run after the repairs: mutants: 136 of 136 caught by the named test, all 136 by the case their expect names, 23 of 23 targets restored byte-identically, exit 0 in 217 s. The class is the one this arc opened with - a link that goes stale while the artifact still looks right - so the reading worth keeping is not the count of teeth but the count of teeth whose named case is the one that fails. One behaviour recorded rather than changed: a named case that does not finish inside its bound counts as caught, with the reason printed (the-pass-asks-a-unit-it-already-failed-again, 30 s). The tool's own comment says a run that never ends proves nothing; the code counts it as caught and says why. Resolving that tension changes what the ledger's proven means, so it is left for a decision rather than a fix. Docs and code in one commit: row D11 names the cases for the two retired rules and says when the teeth went; the obligations ledger gains the whole-register reading with all four repairs named; the record's third group notes the same, the two candidates still judged not retirable, and the timeout tension. Readings at this revision: whole register -> 136 of 136 caught, 136 by name, 23 of 23 restored, 217 s; mutation:anchors -> anchors: 136 of 136 resolve; test:product -> 1562 pass, 0 fail, exit 0; verify:static -> exit 0; lint -> 0 findings. --- ...-09-24-mutants-are-derived-not-anchored.md | 32 +++++++++++++++---- ...-mutants-are-derived-not-anchored.zh-CN.md | 9 +++++- .../design/task-unit-semantics-obligations.md | 7 ++-- evals/ooo-execution/board-slots.test.ts | 3 ++ tests/core/store/current-value-window.test.ts | 5 ++- tests/integration/ooo-fusion.test.ts | 5 ++- tools/mutation-teeth.ts | 22 ++----------- 7 files changed, 51 insertions(+), 32 deletions(-) diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index 295706ca..056d9794 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -150,12 +150,32 @@ claims at once, and the store is why each handoff is directed`, whose assertions needed, which makes that case a better witness than the tooth's name. Measured after: `anchors: 138 of 138 resolve`, and `--targets=src/integration/ooo-board.ts` sweeps `15 of 15 caught`, restored byte-identically. - **Remaining:** the I/E teeth, the store-target teeth (`a-retried-run-fact-is-appended-twice`, - `a-frozen-task-is-replaced-by-a-different-definition`), and the two candidates judged **not** - retirable so far - `every-task-is-frozen-at-position-zero` and `a-binding-does-not-record-its-channel`, - whose named case is a register/freeze/adopt/read-back round-trip, weaker than the rule it is asked to - carry. The stale `expect` on `next-is-not-the-head-of-the-ordered-candidates` and on - `fusion-continues-from-an-unverified-answer` is still outstanding. + **Third group landed (two teeth retired, four links repaired, 2026-09-24):** + `a-retried-run-fact-is-appended-twice` and `a-frozen-task-is-replaced-by-a-different-definition` (row + D11) go, because the row's own prose already rested on the cases for those two rules and the cases say + them outright. The group's larger finding is what the first **whole-register** sweep said: 136 of 136 + caught, but only **134 by the case each `expect` named**. Four links were wrong, each in its own way: + - `fusion-continues-from-an-unverified-answer`'s case refused the pair for a second reason as well, so + it passed under the mutant. The pair now carries no dependency between the two units, which leaves + the predecessor's verdict as the only condition that can refuse it. + - `next-is-not-the-head-of-the-ordered-candidates`'s `expect` named a case in another target. The + board's budget case now asserts `next()` is the head of the ordered set, and that is the name it + carries. + - `the-caller-rebuilds-the-shared-floor`'s `expect` was a paraphrase that named no case at all. + - `the-grace-is-zero`'s named case derived its fixture from the constant under test + (`half = CLOCK_GRACE_MS / 2000`), so zeroing the grace moved the stamp onto `now` and the case passed + while the bug was live. The fixture is a literal now. + The re-run after those repairs: **136 of 136 caught, all 136 by the case its `expect` names**, 23 of 23 + targets restored byte-identically, 217 s. The class is the same one this record opened with - a link + that goes stale while the artifact still looks right - so the reading to keep is not the count of teeth + but the count of teeth whose named case is the one that fails. + One behaviour recorded rather than changed: a named case that does not finish inside its bound counts as + caught, with the reason printed (`the-pass-asks-a-unit-it-already-failed-again`, 30 s). The tool's own + comment says a run that never ends proves nothing; the code counts it as caught and says why. That + tension is left alone, because resolving it changes what the ledger's `proven` means. + **Remaining:** the I/E teeth, and the two candidates judged **not** retirable so far - + `every-task-is-frozen-at-position-zero` and `a-binding-does-not-record-its-channel`, whose named case is + a register/freeze/adopt/read-back round-trip, weaker than the rule it is asked to carry. ## Acceptance criteria diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index 459f671a..00a5f992 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -48,7 +48,14 @@ 退役后实测:`anchors: 144 of 144 resolve`(登记处是 144 颗牙,不是 149),这趟碰到的四个目标 sweep `70 of 70 caught`,5 of 5 逐字节还原。 一个要带下去的发见:`a-refused-unit-is-silent` **没有任何**台账行点名它,抓住它的那条用例同样无人认领——一颗孤儿牙,退役它移除的是孤儿,而不是某一行的钉子。 **第二组已落地(六颗,2026-09-24):**声明的预算那五条规则——许可的是预算的那一部分而不是表头、预算大于一必须点名目标、每个可开始任务都有一个 handoff、可开始的 handoff 会在重发布中存活、多槽位的 handoff 是被定向的——都由同一条用例陈述:`evals/ooo-execution/board-slots.test.ts` 的 `a declared budget holds two claims at once, and the store is why each handoff is directed`,其断言逐条点名了这些规则(拒绝消息、handoff 计数、`serialState` 为 null、重发布后 id 不变、先认领非表头那个)。这五颗从 F2b-slot 行退役,该行保留七颗。第六颗 `claim-is-not-scoped-to-its-run`(B1 行)是命名空间实验里最初**没被抓住**的那颗牙——用例被加强到能抓住它,实验记录写下了当时必需的"裸行读取",所以那条用例比这颗牙的名字是更好的证人。退役后实测:`anchors: 138 of 138 resolve`,`--targets=src/integration/ooo-board.ts` sweep `15 of 15 caught`,逐字节还原。 - **尚未完成:**I/E 牙、store 目标那两颗(`a-retried-run-fact-is-appended-twice`、`a-frozen-task-is-replaced-by-a-different-definition`),以及目前判定为**不可退役**的两颗——`every-task-is-frozen-at-position-zero` 与 `a-binding-does-not-record-its-channel`,它们的点名用例是"注册/冻结/采纳/读回"的往返,弱于要它承担的规则。`next-is-not-the-head-of-the-ordered-candidates` 与 `fusion-continues-from-an-unverified-answer` 的过期 `expect` 仍未处理。 + **第三组已落地(退役两颗、修好四条链接,2026-09-24):**`a-retried-run-fact-is-appended-twice` 与 `a-frozen-task-is-replaced-by-a-different-definition`(D11 行)退役——该行的正文本来就靠这两条用例承载那两条规则,而用例把规则直说了。这一组更大的发现是**第一次全登记处 sweep**的读数:136 颗全部被抓,但只有 **134 颗是被各自 `expect` 点名的那条用例抓住的**。四条链接都错了,而且各错各的: + - `fusion-continues-from-an-unverified-answer` 的点名用例还因为另一个原因拒绝这一对,于是在 mutant 下照样通过。现在这一对的两个单元之间没有依赖,前驱的裁决成为唯一能拒绝它的条件。 + - `next-is-not-the-head-of-the-ordered-candidates` 的 `expect` 指的是另一个目标里的用例。董事会(board)的预算用例现在断言 `next()` 就是有序集合的表头,它的名字随之更正。 + - `the-caller-rebuilds-the-shared-floor` 的 `expect` 是一句转述,不对应任何用例。 + - `the-grace-is-zero` 的点名用例从**被测常量本身**推导夹具(`half = CLOCK_GRACE_MS / 2000`),于是把 grace 归零会把时间戳挪到 `now` 上,缺陷活着而用例照过。夹具现在是字面量。 + 修完重跑:**136 of 136 caught,且 136 颗全部由各自 `expect` 点名的用例抓住**,23 of 23 目标逐字节还原,217 秒。这一类正是本记录开头写的那一类——**凭据看着没问题,链接悄悄过期**——所以要记住的读数不是牙的数量,而是"点名用例就是失败那个"的牙的数量。 + 一条记录而未改动的行为:点名用例在自己的时限内没跑完,也被算作抓住,并把原因打印出来(`the-pass-asks-a-unit-it-already-failed-again`,30 秒)。工具自己的注释说"永远跑不完的一轮什么也证明不了",代码却把它算作抓住并说明了理由。这个矛盾留在原地,因为解决它会改变台账里 `proven` 的含义。 + **尚未完成:**I/E 牙,以及目前判定为**不可退役**的两颗——`every-task-is-frozen-at-position-zero` 与 `a-binding-does-not-record-its-channel`,它们的点名用例是"注册/冻结/采纳/读回"的往返,弱于要它承担的规则。 ## 验收标准 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index 2c747293..60917484 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -20,10 +20,11 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth` -> 136 of 136 caught by the named test, 20 of 20 targets restored byte-identically, exit 0. Run on 2026-09-19 as four lanes, one sweep per tree, using `git worktree add --detach` on the same commit for three of them: a sweep is sequential _within_ a tree because its mutants substitute into the same file, and parallel across trees, where each lane also gets the isolation property that no lane's suites can read another lane's mutant. Lanes: 42 of 42 (`ooo-board`, `task-coordinator`), 41 of 41 (`base`, `ooo-execution`'s 17, `task-semantics-interleavings`), 42 of 42 (thirteen small targets) and 11 of 11 (`plan-driver`) - the last serialised into its own lane because its suite has a 25 s case and does real candidate verification (~92 s per run, against ~2 s for the cheap suites). **That lane's cost has since changed**: the arms' checks are data checks now, so at `c01d3fe` its clean run is 5.5 s and its five mutants are caught by the case each names in 0.8-1.2 s on the same target; the numbers in this paragraph describe the 2026-09-19 instrument. Three things the run itself taught, all fixed and pinned afterwards: the lock's `live` flag was never written on substitution (a multi-hunk edit failed as a whole and only the restore half was reapplied), so the field lied about a running sweep; `NODE_TEST_CONTEXT` inherited when a sweep is started from inside a `node --test` process made the nested runner exit 0 having run no test at all, which the harness reported as "the suite passed" and which turned every mutant of that target into a false "not caught"; and the refusal in `agent:verify` fired on `--dry-run` too, which made two of the verifier's own tests fail while a sweep held the tree (a dry run reads the plan and the route config, not the mutated file, so it is exempt now). The lane that reported a clean-run failure (`tools/agent-verify.ts`) had found the last of those three. A sweep is also refused while any lock is present, including one whose owner died, because a killed sweep leaves its mutant in the target (post-mortem 0003). Interruption note: two lane processes were killed by the console that launched them and were relaunched; the JSON each run writes at its end survived even when the buffered stdout summary was lost, so the lane results above were read from those files rather than from stdout. (the retirement pass added the driver's interleaving mutant; was 111 of 111 before this pass: `src/integration/task-semantics-interleavings.ts` gained three budget mutants and `evals/ooo-execution/plan-driver.ts` three for the per-unit checks, the canned worker and the parent composition). How these runs are scheduled (scoped during a change, full before a push, detached with a collected result) is a standing rule of the repository now, in [`skills/repo-development/SKILL.md`](../../skills/repo-development/SKILL.md) with its measured costs in [the decision](../decisions/implemented/2026-09-18-detached-long-checks.md) - `npm run complexity:gate` -> exit 0, 18 methods above 15 unchanged from baseline. It caught the E pass's first version (adding the advisers option pushed `BoardAdmission`'s constructor to 17, so the options check moved into `admissionAdvice()`) rather than the threshold being raised; the slot pass added its option check the same way (`admissionSlots`), and the ordered-set/`publishReady` reads stayed under it. - `node --experimental-strip-types --test tests/tools/mutation-anchor.test.ts` -> 16 pass, 0 fail, exit 0 (2026-09-24: the resolver's six operators and its refusals, over source strings, no filesystem) -- `npm run mutation:teeth -- --anchors-only` -> `anchors: 149 of 149 resolve, over 23 targets`, exit 0 in 0.6 s +- `npm run mutation:teeth -- --anchors-only` -> `anchors: 149 of 149 resolve, over 23 targets`, exit 0 in 0.6 s (2026-09-24: every tooth still applies where it is aimed; one anchor is re-taken through whitespace normalization, `the-completion-does-not-bind-the-verdict-to-the-bytes`) - `npm run mutation:teeth -- --anchors-only` -> `anchors: 144 of 144 resolve, over 23 targets`, exit 0 in 0.7 s (2026-09-24, after five teeth were retired in favour of the checks that already state their rules) - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts` -> `mutants: 15 of 15 caught by the named test`, restored byte-identically (2026-09-24, after six more teeth were retired there: the five budget/publication rules and the run-scoping claim) -- `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) (2026-09-24: every tooth still applies where it is aimed; one anchor is re-taken through whitespace normalization, `the-completion-does-not-bind-the-verdict-to-the-bytes`) +- `npm run mutation:teeth` (the whole register) -> `mutants: 136 of 136 caught by the named test`, 23 of 23 targets restored byte-identically, exit 0 in 217 s (2026-09-24), **every one of them caught by the case its `expect` names**. The first whole-register reading (231 s, 134 of 136 by name) is what surfaced three wrong links, all repaired in the same pass: `the-caller-rebuilds-the-shared-floor` (its `expect` was a paraphrase of no case, re-pointed), `the-grace-is-zero` (its named case derived its fixture from the constant under test, so it passed while the bug was live - the fixture is a literal now), and `next-is-not-the-head-of-the-ordered-candidates` (re-pointed to `at the default budget the licence is still the head of the ordered set`, which now asserts the head). A fourth was found by the pilot sweep and repaired here: `fusion-continues-from-an-unverified-answer`'s case refused the pair for a second reason as well, so it passed under the mutant - with no dependency between the two units it is refused by that one condition only. One row still carries a note rather than an assertion: `the-pass-asks-a-unit-it-already-failed-again` is caught by a named case that does not finish inside its 30 s bound, which this tool counts as caught and prints with that reason. +- `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) - `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) - `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. - `node --experimental-strip-types --test --test-concurrency=4 tests/integration/ooo-ordinary-failure.test.ts tests/integration/ooo-managed-fence.test.ts tests/integration/ooo-read-paths-agree.test.ts tests/integration/ooo-round-query.test.ts tests/integration/ooo-task-tables.test.ts` -> 5, 3, 1, 2 and 4 pass, 0 fail, exit 0 @@ -174,7 +175,7 @@ separate claims. | D8 | Shutdown stops new work, lets the work in flight finish, then closes once | proven | `src/cli/service.ts`: `close()` refuses while calls are in flight, `drain()` is the awaited middle step, and `inFlight` counts what the daemon accepted. `tests/cli/service-drain.test.ts` holds an accepted call in flight (search opens the store, so the counter sees work a synchronous close could not), requires `drain(0)` to fail rather than return quietly, then closes once; `tests/cli/service.test.ts` 51/0 and `archive-shutdown` 3/0 are the regression evidence | | D9 | A post-commit notification failure is recorded, not thrown | proven | `tests/integration/ooo-post-commit-notification.test.ts`: a unit whose post-commit `afterCommit` throws still submits as `accepted`, its artifact is the board's, and `lastNotificationFailure()` reports the message, while the same unit with a reachable notification reports nothing. Verified by mutation (hand-run, 2026-09-18): deleting the `try`/`catch` in `submit()` makes the notification escape and the case fails. The implementer was always the board (`src/integration/ooo-board.ts`, `lastNotificationFailure()`); the evals suite that used to pin it drove a real round only because the artifact envelope is the board's business, and the round was retired ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The seam took three tries to find - `publishReady` is private and is called without a port at the commit, and `afterCommit` also runs on non-commit paths, so the failure is gated on a commit having landed | | D10 | `runCycle`'s early-cancel branch is reachable, or it is dead code | not applicable | **Not applicable: the branch went with the round.** It was reachable and pinned twice (`evals/ooo-execution/early-cancel.test.ts` failed with "database is not open" when the round closed the store it borrowed; verified by hand, restored byte-identically), and the retirement deleted both the branch and its subject ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The property it stood behind - nobody who borrows a store closes it - now holds by construction, because no client opens one: the daemon is the only writer and drivers borrow a served store (D14) | -| D11 | The run records live in the store's schema, and their typed writes join the store's transaction | proven | `task_run_manifest`, `task_run_tasks` and `task_run_facts` are created by `migrate()` like every other table (`tests/core/store/schema.test.ts` names them), and `NmgStoreBase` writes them through `registerTaskRun`, `freezeTaskRunTask` and `appendTaskRunFact`, each taking the same optional port the board writes take: outside a transition it opens one, inside it joins the caller's. What the cases prove (`tests/core/store/task-runs.test.ts`, 8/8): a board write and a run fact written through one port land together, and when the fact write fails the board row does not survive it; a retried fact is recorded once and keeps its first sequence; two runs in one store keep separate task and fact namespaces; a run's plan cannot be replaced and a frozen task cannot be redefined; an entry bound by two runs is refused rather than answered with one of them; the reads register nothing. Mutants: `the-run-fact-opens-its-own-transaction`, `a-retried-run-fact-is-appended-twice`, `a-second-plan-overwrites-the-frozen-one`, `a-frozen-task-is-replaced-by-a-different-definition`, `a-second-run-adopts-a-bound-entry`. One layer deliberately has no mutant of its own: the `UNIQUE (run_id, kind, task_id, attempt)` constraint is a storage-level backstop behind the explicit identity check, so a mutant that dropped only the constraint would survive - the check is the deciding layer and is the one mutated above | +| D11 | The run records live in the store's schema, and their typed writes join the store's transaction | proven | `task_run_manifest`, `task_run_tasks` and `task_run_facts` are created by `migrate()` like every other table (`tests/core/store/schema.test.ts` names them), and `NmgStoreBase` writes them through `registerTaskRun`, `freezeTaskRunTask` and `appendTaskRunFact`, each taking the same optional port the board writes take: outside a transition it opens one, inside it joins the caller's. What the cases prove (`tests/core/store/task-runs.test.ts`, 8/8): a board write and a run fact written through one port land together, and when the fact write fails the board row does not survive it; a retried fact is recorded once and keeps its first sequence; two runs in one store keep separate task and fact namespaces; a run's plan cannot be replaced and a frozen task cannot be redefined; an entry bound by two runs is refused rather than answered with one of them; the reads register nothing. Mutants: `the-run-fact-opens-its-own-transaction`, `a-second-plan-overwrites-the-frozen-one`, `a-second-run-adopts-a-bound-entry`. Two rules this row states are carried by the cases rather than by a mutant: `appending the same fact twice records it once and keeps the first sequence` and `freezing a task twice is a no-op, and a different definition for it is refused` say the retry identity and the frozen definition in their own assertions, so `a-retried-run-fact-is-appended-twice` and `a-frozen-task-is-replaced-by-a-different-definition` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). One layer deliberately has no mutant of its own: the `UNIQUE (run_id, kind, task_id, attempt)` constraint is a storage-level backstop behind the explicit identity check, so a mutant that dropped only the constraint would survive - the check is the deciding layer and is the one mutated above | | D12 | The binding of a logical task to its board entry is a run fact, and one routing rule decides what a binding makes managed | proven | `bindRunEntry` in `src/integration/task-coordinator.ts` records the binding as the run fact `entry-bound` (the name the D11 store tests already used), and it is the binding that the store's fence reads, so a bound entry's lifecycle write is refused unless it runs inside the run's coordinated transition. The op joins a caller's transaction when given a port, which is what lets a host create the entry and bind it in one transition: `tests/integration/ooo-managed-adopt.test.ts` (5 cases) checks that the board never holds an entry the run does not (a failure after both writes leaves neither), that a retry on the same task and attempt is not a second binding while another entry for that task and attempt is refused rather than silently dropped, and that every refusal names a stored fact (unregistered run, cancelled run, a task the run never froze, an entry that is not on the named channel, an entry another run already carries). The idempotence half rests on that case rather than on a mutant - `a-second-entry-rebinds-the-task` was retired on 2026-09-24 - because the case asserts the retry's shape and the refusal by name. `coordinatedEntryWrite` moves to the same module as the single routing rule, so the daemon's board verbs and any in-process writer decide "managed" in one place instead of each caller's belief about the entry; `src/cli/service.ts` no longer keeps its own copy. Mutants: `an-adopted-entry-takes-the-direct-path`, `a-binding-ignores-whether-the-task-was-frozen`, `a-binding-does-not-check-the-entry-exists`, `one-entry-is-bound-to-two-tasks`, `the-daemon-verb-skips-the-coordinated-path` (re-anchored to the daemon's claim handler, since the routing it mutates moved out of that file) | | D13 | The run surface another process reaches is the daemon's: register, freeze, bind, cancel and status, and nothing else writes run facts | proven | `taskRun` in `src/cli/protocol.ts` over `registerRun` / `freezeRunPlan` / `bindRunEntry` / `cancelRun` / `taskRunStatus` / `createBoundEntry` in `src/integration/task-coordinator.ts`; `tests/cli/task-run-surface.test.ts` (11 cases) drives all of it through `service.invoke`, so what it exercises is the wire a second process uses, and it asserts `hello.methods` carries `taskRun` - the advertisement a client gates the `adopt` field on. Registration is the only transition that does not need a run to exist already (`coordinateRunWrite` refuses an unregistered run, which is what every later transition and every managed write rests on). A plan freeze is one transition: the request's array order becomes the positions, a dangling dependency and a self-dependency are refused by name while the plan is still only a proposal, and a batch the store refuses leaves the earlier tasks unfrozen - `the-plan-freezes-one-task-per-transaction`, `every-task-is-frozen-at-position-zero`, `the-plan-may-freeze-a-dangling-dependency`, `a-task-may-depend-on-itself` - and a cancelled run takes no further plan (`a-cancelled-run-takes-a-new-plan`). Adoption rides the transition that creates the entry (`createBoundEntry` takes the port), so a refusal leaves no entry behind (`the-entry-is-created-before-its-binding-is-checked`) and the wire cannot drop the request silently (`the-wire-drops-an-adoption-request`), which is the one failure mode the epoch rule in `design.md` is about. The binding fact records the channel that names its entry - that is what lets `status` resolve a binding without searching the board (`a-binding-does-not-record-its-channel`). `cancel` is the only writer of `run-cancelled`: before this row `src/` had none, so the state the fence and the dispatch derivation both read was reachable only from a test (the gap G4 named); a run-level cancellation carries the schema's empty task id (`a-run-cancellation-names-a-task`), a task-level one requires the task to have been frozen (`a-cancellation-ignores-whether-the-task-was-frozen`, `an-unknown-run-can-be-cancelled`), and cancelling twice is one fact because the fact's own identity is the duplicate key. `status` is a read and derives nothing: a run the store does not know stays unknown (`status-registers-the-run-it-cannot-find`), and ready/blocked/accepted stay with the shared pure function rather than with this view. The CLI exposes the operator half (`nmg run status`, `nmg run cancel`), which is also what the "every RPC method is a CLI command or intentionally RPC-only" gate asks for. One thing this pass corrected about itself: the first version's parser also computed a `position` per task, which `freezeRunPlan` immediately overwrote - the mutant that changed it survived, which is how the dead field was found, and it was deleted rather than kept | | D14 | The evidence drivers reach the board through the daemon and open no database of their own | proven | `evals/ooo-execution/round-client.ts` resolves the daemon from the store's own lease and refuses by name when nothing serves it; `round-host.ts` serves an existing round store the way the product daemon does (`NmgService` + `serveHttp` + the store's lease) and is that store's only writer while it runs; `board-worker.ts`, `board-deliver.ts` and `board-judge.ts` now use the protocol's board verbs (`read`, `claim`, `release`, `deliver`, `judge`) and no longer import the store. `tests/integration/ooo-evidence-drivers.test.ts`: both end-to-end cases host the store and spawn the drivers as separate processes against the endpoint (spawned, not `spawnSync`, because the test process is what answers them); the managed case registers a run and freezes its plan through `taskRun`, adopts the entry in the same `put` that creates it, then runs the worker and the judge as separate processes - the run's log afterwards is exactly `entry-bound`, `board-claim`, `board-deliver`, `board-judge`, with the binding resolved to the entry the worker claimed and its verdict accepted. Two cases hold the boundary itself: a driver with no daemon refuses by name instead of falling back to the file, and no `board-*.ts` driver may mention the store or omit the round client. Mutants: `a-driver-falls-back-to-opening-the-store`, `the-round-client-does-not-require-a-daemon`, plus the drivers' two earlier ones (`worker-reads-only-one-reporter-shape`, `judge-may-judge-its-own-delivery`), re-pointed at the renamed case. Two further rules came out of building this, and both are about the same defect the first attempt hit: a process that serves an endpoint cannot answer a call to it while it is blocked, so "the answer will come" is not an assumption a client may make. `round-client.ts` therefore bounds every call (`ROUND_CALL_TIMEOUT_MS`, an opt-in bound on the thin client's `httpCall`, which by default still leaves the platform's own) and reports a timeout by naming the bound, the endpoint and the pid that serves it rather than the transport's; and it refuses when the lease's pid is the caller's own, directing an in-process host to `host.call(...)` (`round-host.ts`), which is the design's shape for an offline host. The test is now a client too: it starts the host as its own process (`[host] serving `, idle timeout as the backstop for a host whose owner died) instead of hosting in-process, and asks it to shut down rather than killing it, so what runs on the way out is the release path - a lease held by a process that is gone is a store nothing can serve. Cases: "a client refuses to call the endpoint its own process serves", "a call to a host that never answers gives up in seconds and names the reason" (raced against its own 5s deadline, so the check of the bound is itself bounded), "a host releases its lease when it stops, so the next host can take the store". Mutants: `the-round-client-calls-the-endpoint-it-serves`, `the-round-client-has-no-limit-on-how-long-it-waits`, and - new target, `src/cli/http-server.ts` - `the-serving-process-never-releases-its-lease`. Not covered by this row: `live-continuation.ts` still creates its own store and constructs `BoardAdmission` - the runner half of the same design sentence and the next obligation. `round-runner.ts` did the same and was retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)) | diff --git a/evals/ooo-execution/board-slots.test.ts b/evals/ooo-execution/board-slots.test.ts index 1df5db85..486b25bf 100644 --- a/evals/ooo-execution/board-slots.test.ts +++ b/evals/ooo-execution/board-slots.test.ts @@ -141,6 +141,9 @@ test("at the default budget the licence is still the head of the ordered set", ( "the legal set is what a caller may report or rank, whatever the budget", ); assert.deepEqual(gate.startable(), ["P"], "the licence is the budget's part of that set"); + // The ordered set has a head, and it is one task: a caller that asks "what next" gets the same + // answer the licence names, rather than the second candidate or the whole set. + assert.equal(gate.next(), "P", "the head of the ordered set is the next task"); const published = handoffs(gate); assert.equal(published.length, 1, "one slot publishes one handoff"); assert.equal(published[0]!.to, null, "the default budget publishes the broadcast handoff"); diff --git a/tests/core/store/current-value-window.test.ts b/tests/core/store/current-value-window.test.ts index 9964ad4e..e13c5e23 100644 --- a/tests/core/store/current-value-window.test.ts +++ b/tests/core/store/current-value-window.test.ts @@ -33,7 +33,10 @@ import { } from "../../../src/core/store/clock.ts"; import { NmgStore } from "../../../src/core/store.ts"; -const half = (CLOCK_GRACE_MS / 2000).toFixed(3); +// Deliberately a literal, not `CLOCK_GRACE_MS / 2000`: a fixture derived from the constant under test +// moves with it, so zeroing the grace would move the stamp onto `now` and the case would pass while the +// bug it exists for was live. The case below asserts the relation to the constant instead. +const half = "0.025"; /** A memory whose validity boundaries are stamped in SQL, then reopened through the store: the read * then compares SQLite's clock against SQLite's clock. */ diff --git a/tests/integration/ooo-fusion.test.ts b/tests/integration/ooo-fusion.test.ts index f403cc31..2dba0750 100644 --- a/tests/integration/ooo-fusion.test.ts +++ b/tests/integration/ooo-fusion.test.ts @@ -72,7 +72,10 @@ test("fusion refuses a successor whose visibility the session would widen", () = }); test("fusion refuses to continue from a unit whose verdict is not accepted", () => { - const tasks = [task("before"), task("after", { dependencies: ["before"] })]; + // No dependency between them, so this pair is refused by exactly one condition: the predecessor's + // verdict. With a dependency the same assertion would pass for the other reason, and the case would + // claim a rule it does not pin. + const tasks = [task("before"), task("after")]; assert.equal(legal({ tasks }), false); }); diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 9c1557bc..31235f62 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -230,23 +230,6 @@ const TARGETS: readonly Target[] = [ to: " if (\n false &&\n String(existing.plan_digest) !== input.planDigest &&\n String(existing.policy) !== input.policy\n )", expect: "a run registers once, and a second plan for the same run is refused", }, - { - // Frozen means frozen: the same task id with a different definition is a different plan, - // and replacing it in place would rewrite the input a decision was already read against. - name: "a-frozen-task-is-replaced-by-a-different-definition", - ast: { within: "insertTaskRunTask" }, - from: " if (!same)", - to: " if (!same && false)", - expect: "freezing a task twice is a no-op, and a different definition for it is refused", - }, - { - // The fact's own identity is what makes a retry after a lost response append once. - name: "a-retried-run-fact-is-appended-twice", - ast: { within: "insertTaskRunFact" }, - from: " if (known) return { sequence: Number(known.sequence), recorded: false };", - to: " if (known && false) return { sequence: Number(known.sequence), recorded: false };", - expect: "appending the same fact twice records it once and keeps the first sequence", - }, { // The fact write has to join the transition the caller opened, not open a second one; the // board write and the run fact of one transition stand or fall together. @@ -416,8 +399,7 @@ const TARGETS: readonly Target[] = [ ast: { within: "next" }, from: " return this.candidates()[0] ?? null;", to: " return this.candidates()[1] ?? null;", - expect: - "contract: a verified patch candidate is what dependents bind to, and only acceptance releases them", + expect: "at the default budget the licence is still the head of the ordered set", }, ], }, @@ -1220,7 +1202,7 @@ const TARGETS: readonly Target[] = [ name: "the-caller-rebuilds-the-shared-floor", from: "? [...narrowPlan.shared, ...(route.tests.length ? [nodeTestCheckName(route.id)] : [])]", to: '? ["check", "docs:check", "format:check", "glossary:check", "lint", "package:check", "rtm:check", ...(route.tests.length ? [nodeTestCheckName(route.id)] : [])]', - expect: "a route that declines the shared checks plans only its own tests", + expect: "a declining route's narrow run verifies on its own tests and nothing else", }, ], }, From a8103248be6652979f7c204b8dfce5b915e89f9b Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 16:48:33 +0800 Subject: [PATCH 28/32] The insertion-shaped residue retires where a check states its rule The 27 teeth whose anchor is a place where code must not appear are the ones a derived selector cannot express, so each of them was either going to stay hand-written forever or give way to a check that says the same thing. 26 give way, and one is kept with its reason. The method, applied per tooth rather than in bulk: the register-wide sweep first says which case fails under the tooth; that case's assertion is then read to see whether it fails for exactly this violation and states the rule in its own words. What qualified: - a count over the store's own source - "the store runs its transaction boundary in exactly one place" counts one BEGIN, one COMMIT and one ROLLBACK, all three inside writeTransaction (a-method-opens-its- own-transaction); - an enumeration of a set that must not change - the accepted prefix, the fallbacks by source id, the outcome of a scoring order, the orders the ordered mode allows, a run's tasks after a refused freeze (reopening-one-task-clears-every-acceptance, a-failing-source-takes-the-decision-with-it, the-ordering-adds-a-task-to-the-set, ordered-mode-becomes-any-topological-order, the-plan-freezes-one-task-per-transaction); - a refusal by name plus a raw read back - the second reader of a ready task is refused and the owner is then read out of the table, the status of an unknown run stays unknown, a driver without a daemon refuses and no driver imports the store (a-live-claim-can-be-taken-by-another-agent, status-registers-the-run-it- cannot-find, a-driver-falls-back-to-opening-the-store); - the file's own hash and a file that must not be created (the-read-only-factory-opens-a-writable-handle). One is kept: every-task-is-frozen-at-position-zero. Its named case is a register/freeze/adopt/read-back round-trip, weaker than "the position comes from the array order, not from the request" - the same reason a-binding-does-not-record-its-channel stays, as recorded when the earlier groups were judged. Seven of the 26 turned out to be orphans: no row in the obligations ledger names round-publication-opens-its-own-transaction, round-never-releases-its-pin, ordered-mode-becomes-any-topological-order, the-loop-awaits-each-unit-instead-of-the-batch, a-unit-is-dispatched-twice-in-one-batch, a-unit-nothing-checks-is-still-a-unit or the-parent-check-ignores-its-own-verdict, so each of those rules has a check and no row. The checks are named in the record. Whether those seven rules deserve rows is a question about the ledger, and it is left as one rather than answered by inventing rows here. The target tools/agent-verify.ts had a single tooth and goes with it, so the register is 110 teeth over 22 entries covering 21 files. Docs and code in one commit: the obligations ledger's rows stop naming the retired teeth as evidence and name the case that states each rule instead (B3, B6, C3, D3, D6, D11, D13, D14, E1, E3, F2b-slot, F5, the G rows and the prose that listed them); its readings gain the whole-register run for this pass and its "how a row earns proven" paragraph gains the retirement criterion, including the two ways a case can look like it states a rule without doing so (a fixture derived from the constant under test, a case that refuses for a second reason as well) - both of which this arc found by reading assertions; the record's item 3 is marked landed with the numbers, the kept tooth, the seven orphans and the next question. Readings at this revision: whole register -> mutants: 110 of 110 caught by the named test, all 110 by the case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 162 s; mutation:anchors -> anchors: 110 of 110 resolve, over 22 targets; test:product -> 1562 pass, 0 fail, exit 0; verify:static -> exit 0; lint -> 0 findings. --- ...-09-24-mutants-are-derived-not-anchored.md | 37 ++- ...-mutants-are-derived-not-anchored.zh-CN.md | 5 +- .../design/task-unit-semantics-obligations.md | 110 ++++---- tools/mutation-teeth.ts | 234 ++---------------- 4 files changed, 114 insertions(+), 272 deletions(-) diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index 056d9794..9e062e9a 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -82,6 +82,35 @@ and the replacement bytes are computed from the syntax tree on every run.** appear_, so they are the last to move and the first to be reconsidered: where the rule already has a check whose form is relational or enumerative, the tooth is retired and the check is named in its place. + + **Landed (2026-09-24): 26 of the 27 insertion-shaped teeth are gone.** Each went the same way: the + sweep first showed that the tooth's named case was the one that failed under it, and reading that case + showed the rule stated outright - a count over the store's own source (`the store runs its transaction +boundary in exactly one place`: one BEGIN, one COMMIT and one ROLLBACK, all three inside + `writeTransaction`), an enumeration of the accepted set, of the fallbacks or of what a freeze left + behind, a refusal by name plus a raw read back (`one reader of the same ready task is given the claim, +the second is refused` reads the owner out of the table), or the file's own hash (a read-only open + neither creates, migrates nor writes). One was kept in this set: `every-task-is-frozen-at-position-zero`, + whose named case is a register/freeze/adopt/read-back round-trip and so is weaker than "the position + comes from the array order, not from the request" - the same reason `a-binding-does-not-record-its-channel` + stays. Seven of the 26 were **orphans** - no row named them - so the rule each pinned has a check and no + row: `round-publication-opens-its-own-transaction` (`the round's own publication rolls back with the +transition that made it`), `round-never-releases-its-pin` (`cancelling a round releases the pins it held, +so nothing it referenced leaks`), `ordered-mode-becomes-any-topological-order` (`the ordered mode is the +declared plan order and nothing else`), `the-loop-awaits-each-unit-instead-of-the-batch` and + `a-unit-is-dispatched-twice-in-one-batch` (`a unit's check is outstanding while an independent unit's +worker runs`, which also asserts "each unit is dispatched once"), `a-unit-nothing-checks-is-still-a-unit` + (`evals/ooo-execution/families.test.ts`'s per-family "a unit nothing checks is refused rather than + accepted on nothing") and `the-parent-check-ignores-its-own-verdict` (`the parent check is the composed +acceptance, and a failing check is reported as such`). Whether those seven rules deserve rows of their own + is the next question this reading raises, and it is a question about the ledger rather than about the + register. + + The reading after the retirement: **110 of 110 caught by the named test, all 110 by the case their + `expect` names**, 22 of 22 targets restored byte-identically, 162 s; `anchors: 110 of 110 resolve, over 22 +targets`. The target `tools/agent-verify.ts` went with its single tooth, so the register is 22 entries + over 21 files. + 4. **A check that a tooth still applies becomes part of the static contract**: an anchors-only pass that resolves every mutant, without running a suite, in single-digit seconds, failing when a site cannot be resolved, resolves more than once, or when a target claims more teeth than it can apply. The @@ -173,9 +202,11 @@ claims at once, and the store is why each handoff is directed`, whose assertions caught, with the reason printed (`the-pass-asks-a-unit-it-already-failed-again`, 30 s). The tool's own comment says a run that never ends proves nothing; the code counts it as caught and says why. That tension is left alone, because resolving it changes what the ledger's `proven` means. - **Remaining:** the I/E teeth, and the two candidates judged **not** retirable so far - - `every-task-is-frozen-at-position-zero` and `a-binding-does-not-record-its-channel`, whose named case is - a register/freeze/adopt/read-back round-trip, weaker than the rule it is asked to carry. + **Remaining:** the mass conversion of the swap-shaped teeth - the six operators can express them, and that + is the deferred half of the vocabulary work - and the one insertion-shaped tooth left standing, + `every-task-is-frozen-at-position-zero`, together with `a-binding-does-not-record-its-channel`: both keep + a named case that is a register/freeze/adopt/read-back round-trip, weaker than the rule each is asked to + carry. ## Acceptance criteria diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index 00a5f992..f2e44b8f 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -55,7 +55,10 @@ - `the-grace-is-zero` 的点名用例从**被测常量本身**推导夹具(`half = CLOCK_GRACE_MS / 2000`),于是把 grace 归零会把时间戳挪到 `now` 上,缺陷活着而用例照过。夹具现在是字面量。 修完重跑:**136 of 136 caught,且 136 颗全部由各自 `expect` 点名的用例抓住**,23 of 23 目标逐字节还原,217 秒。这一类正是本记录开头写的那一类——**凭据看着没问题,链接悄悄过期**——所以要记住的读数不是牙的数量,而是"点名用例就是失败那个"的牙的数量。 一条记录而未改动的行为:点名用例在自己的时限内没跑完,也被算作抓住,并把原因打印出来(`the-pass-asks-a-unit-it-already-failed-again`,30 秒)。工具自己的注释说"永远跑不完的一轮什么也证明不了",代码却把它算作抓住并说明了理由。这个矛盾留在原地,因为解决它会改变台账里 `proven` 的含义。 - **尚未完成:**I/E 牙,以及目前判定为**不可退役**的两颗——`every-task-is-frozen-at-position-zero` 与 `a-binding-does-not-record-its-channel`,它们的点名用例是"注册/冻结/采纳/读回"的往返,弱于要它承担的规则。 + **已落地(2026-09-24):27 颗"插入型"牙里,26 颗退掉了。**每一颗都走同一条路:先由 sweep 证明它的点名用例就是失败的那条,再读那条用例,确认规则被直说——对 store 自身源码计数(`the store runs its transaction boundary in exactly one place`:一个 BEGIN、一个 COMMIT、一个 ROLLBACK,且都在 `writeTransaction` 里)、对已接受集合/回退项/冻结残留做枚举、按名字拒绝后再把原行读回来(`one reader of the same ready task is given the claim, the second is refused` 直接从表里读 owner),或者用文件自身的哈希(`a read-only open neither creates, migrates nor writes`)。这一组里留下了一颗:`every-task-is-frozen-at-position-zero`,它的点名用例是"注册/冻结/采纳/读回"的往返,弱于"位置来自数组顺序而不是请求",与 `a-binding-does-not-record-its-channel` 留下的理由相同。26 颗里有 7 颗是**孤儿**——没有任何一行点名它们——它们各自钉住的规则有检查、却没有行:`round-publication-opens-its-own-transaction`(`the round's own publication rolls back with the transition that made it`)、`round-never-releases-its-pin`(`cancelling a round releases the pins it held, so nothing it referenced leaks`)、`ordered-mode-becomes-any-topological-order`(`the ordered mode is the declared plan order and nothing else`)、`the-loop-awaits-each-unit-instead-of-the-batch` 与 `a-unit-is-dispatched-twice-in-one-batch`(`a unit's check is outstanding while an independent unit's worker runs`,它同时断言"each unit is dispatched once")、`a-unit-nothing-checks-is-still-a-unit`(`evals/ooo-execution/families.test.ts` 里每个 family 的"a unit nothing checks is refused rather than accepted on nothing")以及 `the-parent-check-ignores-its-own-verdict`(`the parent check is the composed acceptance, and a failing check is reported as such`)。这七条规则是否该各自有一行,是这次读数提出的下一个问题——它是台账的问题,不是登记处的问题。 + + 退役后的读数:**110 of 110 caught by the named test,且 110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 目标逐字节还原,162 秒;`anchors: 110 of 110 resolve, over 22 targets`。目标 `tools/agent-verify.ts` 随它唯一那颗牙一起消失,所以登记处现在是 21 个文件上的 22 个条目。 + **尚未完成:**剩余的"替换型"牙的整体迁移——六个算子能表达它们,那是词表工作中被推迟的一半;以及仍立着的插入型 `every-task-is-frozen-at-position-zero` 与 `a-binding-does-not-record-its-channel`:两者保留的点名用例都是"注册/冻结/采纳/读回"的往返,弱于各自要承担的规则。 ## 验收标准 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index 60917484..97325fb4 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -24,6 +24,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth -- --anchors-only` -> `anchors: 144 of 144 resolve, over 23 targets`, exit 0 in 0.7 s (2026-09-24, after five teeth were retired in favour of the checks that already state their rules) - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts` -> `mutants: 15 of 15 caught by the named test`, restored byte-identically (2026-09-24, after six more teeth were retired there: the five budget/publication rules and the run-scoping claim) - `npm run mutation:teeth` (the whole register) -> `mutants: 136 of 136 caught by the named test`, 23 of 23 targets restored byte-identically, exit 0 in 217 s (2026-09-24), **every one of them caught by the case its `expect` names**. The first whole-register reading (231 s, 134 of 136 by name) is what surfaced three wrong links, all repaired in the same pass: `the-caller-rebuilds-the-shared-floor` (its `expect` was a paraphrase of no case, re-pointed), `the-grace-is-zero` (its named case derived its fixture from the constant under test, so it passed while the bug was live - the fixture is a literal now), and `next-is-not-the-head-of-the-ordered-candidates` (re-pointed to `at the default budget the licence is still the head of the ordered set`, which now asserts the head). A fourth was found by the pilot sweep and repaired here: `fusion-continues-from-an-unverified-answer`'s case refused the pair for a second reason as well, so it passed under the mutant - with no dependency between the two units it is refused by that one condition only. One row still carries a note rather than an assertion: `the-pass-asks-a-unit-it-already-failed-again` is caught by a named case that does not finish inside its 30 s bound, which this tool counts as caught and prints with that reason. +- `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were **orphans**: no row named `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` or `the-parent-check-ignores-its-own-verdict`, so their rules have a case but no row; the check each case makes is in the [record](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) - `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) - `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. @@ -41,6 +42,17 @@ landed so far on the pilot file). `npm run mutation:teeth -- --anchors-only` ans anchor still applies - under a second, no suite run, nothing written - because a tooth that stopped matching its rule is otherwise only visible as a check that stopped counting. +A tooth may also be **retired**, and the criterion is the same one that asks whether it is worth keeping: +the check that catches the violation has to state the rule itself, in a form that fails for that +violation and no other - a count, an enumeration, a relation read back as a raw row, a refusal by name. +Behavioural differentials do not qualify (B5 needed call-site counts, not a comparison of outcomes), and +neither does a case whose fixture moves with the constant under test or that refuses for a second reason +as well; both were found by reading the assertion rather than by trusting the sweep. A retirement is only +made after the sweep says the named case is the one that fails under that tooth, and the row that used to +name the tooth names the check instead, in the same commit. 26 teeth were retired this way on 2026-09-24, +including every insertion-shaped tooth whose rule had such a check; the ones that stay are named, with +their reasons, in [the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md). + ## Where a proof lives: two evidence bases, and neither stands for the other The rows cite test files, and those files split into two bases that **do not share storage** and whose @@ -140,45 +152,45 @@ separate claims. ## B. Persistence (design: "进入持久化接入时另须证明") -| node | obligation | state | evidence | -| ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| B1 | Two rounds with the same taskId do not collide | proven | `tests/integration/ooo-run-namespace.test.ts`: the case `two runs in one store do not collide, do not see each other, and cancel separately` asserts each half directly - a claim in one run leaves the other run's row for the same task id alone (read as a raw row, which is the assertion the namespace experiment record shows was needed), a cancellation is per run, and a reopen by name sees the right one. The tooth `claim-is-not-scoped-to-its-run` was retired on 2026-09-24, once that case stated the rule ([record](../experiments/execution/ooo-run-namespace-2026-09-13.md), [decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)) | -| B2 | A retry does not deliver twice | proven | `tests/core/task-board-deliverable.test.ts`; mutants `stale-claim-may-deliver-again`, `deliverer-may-judge-its-own-work` | -| B3 | A transaction failure cannot write a verdict without its association | proven | `tests/integration/ooo-transition-atomicity.test.ts`; mutants `a-method-opens-its-own-transaction`, `nested-write-transaction-is-allowed`, `swallowed-failure-still-commits` | -| B4 | A retained board entry is not cleared by TTL | proven | `tests/core/task-board-retention.test.ts`; mutants `prune-ignores-retention`, `bounded-pin-never-expires` | -| B5 | The status query and dependency release call one predicate | proven | `tests/integration/ooo-acceptance-one-predicate.test.ts`: the case `acceptance has one home, and both readers reach it` counts the call sites itself - one definition of the predicate in `src/`, one `acceptedFact({` in each reader, zero `verdict ===` comparisons in `ooo-board.ts` - so a rename that stops the call and a second decision inside it each fail a count rather than a named mutant. The two teeth that showed both (`the-board-read-path-stops-calling-the-predicate`, `the-board-decides-acceptance-on-its-own`) were retired on 2026-09-24, once the counts stated the rule directly | -| B6 | A generic board write cannot move a managed round's own state, cannot take the live claim it holds, and cannot move an entry a run has adopted outside that run's transition | proven | the state half is structural and asserted (`tests/integration/ooo-managed-fence.test.ts`: the round's owner, attempt, acceptance and terminal reason stay its own facts, and a generic claim on its entry changes none of them). The claim it holds is not structural and is tooth-backed: `a-live-claim-can-be-taken-by-another-agent` - the store's CAS stops requiring the holder to be the claimant, and the same test's direct second reader then succeeds. The design's owed sentence - a generic write on a managed entry is applied _inside_ the coordinating transaction, and `judge`/`resolve` cannot go around the run's fence - is implemented and pinned by `tests/integration/ooo-managed-write.test.ts` (6 cases): a direct verb on an adopted entry is refused and the entry does not move, while an entry no run adopts takes the path it always did; a coordinated write lands the board transition and the run's fact in one transaction and a failure after the board write leaves neither; a cancelled run takes no further lifecycle writes; a write for the wrong run, or for an entry no run adopted, is refused before anything is written; and the daemon's own `claim` verb routes an adopted entry through the transition, writing the fact with its own store. Mutants: `a-managed-entry-ignores-the-coordinated-scope` (the store's guard), `the-daemon-verb-skips-the-coordinated-path` (the routing), `a-coordinated-write-skips-its-run-fact`, `a-cancelled-run-still-accepts-writes`, `a-coordinated-write-skips-the-binding-recheck`. What is still the design's step 3 rather than this row: nothing adopts entries into a run yet, so the research drivers' direct writes are not refused today - they will be, and are meant to be, once the runner and thin adapters work through the coordinator | -| B7 | JSONL export failure does not change the terminal state | not applicable | no JSONL export exists on this branch to fail | -| B8 | The field-mapping checks catch the three confusions | proven | A4: the offline compiler owns this check | -| B9 | A cancelled task is neither dispatched nor read as a closed input | proven | found by A5's enumeration rather than by reading: acceptance already refused a cancelled task (the one predicate reads `cancellations`), while the eligibility rule could not see the cancellation at all - `DispatchTask` carried no such fact - so a cancelled task stayed selectable as the next dispatch. `DispatchTask.cancelled` now carries the same recorded fact the predicate reads, and `nextTask` refuses it in `ready()` and in `valid()`; pinned by the cases "a cancelled task is not dispatched, and nothing reads one as a closed input" and "nextTask refuses a task marked cancelled, whatever else the caller set", and by mutants `a-cancelled-task-is-still-dispatched` and `the-dispatch-does-not-require-a-cancelled-input-to-be-closed`. Scope: the round has no per-task cancellation source yet (its cancel is run-level, and `recordedFacts` fills no `cancellations`), so nothing changes for it today - this closes the shared derivation's half of the design's 取消后的有效性 | +| node | obligation | state | evidence | +| ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| B1 | Two rounds with the same taskId do not collide | proven | `tests/integration/ooo-run-namespace.test.ts`: the case `two runs in one store do not collide, do not see each other, and cancel separately` asserts each half directly - a claim in one run leaves the other run's row for the same task id alone (read as a raw row, which is the assertion the namespace experiment record shows was needed), a cancellation is per run, and a reopen by name sees the right one. The tooth `claim-is-not-scoped-to-its-run` was retired on 2026-09-24, once that case stated the rule ([record](../experiments/execution/ooo-run-namespace-2026-09-13.md), [decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)) | +| B2 | A retry does not deliver twice | proven | `tests/core/task-board-deliverable.test.ts`; mutants `stale-claim-may-deliver-again`, `deliverer-may-judge-its-own-work` | +| B3 | A transaction failure cannot write a verdict without its association | proven | `tests/integration/ooo-transition-atomicity.test.ts`; mutant `swallowed-failure-still-commits`; the two teeth that pinned the boundary itself (`a-method-opens-its-own-transaction`, `nested-write-transaction-is-allowed`) were retired on 2026-09-24, because `a write entry reached inside a transition without a port is refused, not nested` states the refusal and `the store runs its transaction boundary in exactly one place` counts one BEGIN, one COMMIT and one ROLLBACK in the store's source, all three inside `writeTransaction` | +| B4 | A retained board entry is not cleared by TTL | proven | `tests/core/task-board-retention.test.ts`; mutants `prune-ignores-retention`, `bounded-pin-never-expires` | +| B5 | The status query and dependency release call one predicate | proven | `tests/integration/ooo-acceptance-one-predicate.test.ts`: the case `acceptance has one home, and both readers reach it` counts the call sites itself - one definition of the predicate in `src/`, one `acceptedFact({` in each reader, zero `verdict ===` comparisons in `ooo-board.ts` - so a rename that stops the call and a second decision inside it each fail a count rather than a named mutant. The two teeth that showed both (`the-board-read-path-stops-calling-the-predicate`, `the-board-decides-acceptance-on-its-own`) were retired on 2026-09-24, once the counts stated the rule directly | +| B6 | A generic board write cannot move a managed round's own state, cannot take the live claim it holds, and cannot move an entry a run has adopted outside that run's transition | proven | the state half is structural and asserted (`tests/integration/ooo-managed-fence.test.ts`: the round's owner, attempt, acceptance and terminal reason stay its own facts, and a generic claim on its entry changes none of them). The claim it holds is not structural: it was carried by `a-live-claim-can-be-taken-by-another-agent`, retired on 2026-09-24 - the store's CAS stops requiring the holder to be the claimant, and the same test's direct second reader then succeeds, so the case states the rule itself. The design's owed sentence - a generic write on a managed entry is applied _inside_ the coordinating transaction, and `judge`/`resolve` cannot go around the run's fence - is implemented and pinned by `tests/integration/ooo-managed-write.test.ts` (6 cases): a direct verb on an adopted entry is refused and the entry does not move, while an entry no run adopts takes the path it always did; a coordinated write lands the board transition and the run's fact in one transaction and a failure after the board write leaves neither; a cancelled run takes no further lifecycle writes; a write for the wrong run, or for an entry no run adopted, is refused before anything is written; and the daemon's own `claim` verb routes an adopted entry through the transition, writing the fact with its own store. Mutants: `a-managed-entry-ignores-the-coordinated-scope` (the store's guard), `the-daemon-verb-skips-the-coordinated-path` (the routing), `a-coordinated-write-skips-its-run-fact`, `a-cancelled-run-still-accepts-writes`, `a-coordinated-write-skips-the-binding-recheck`. What is still the design's step 3 rather than this row: nothing adopts entries into a run yet, so the research drivers' direct writes are not refused today - they will be, and are meant to be, once the runner and thin adapters work through the coordinator | +| B7 | JSONL export failure does not change the terminal state | not applicable | no JSONL export exists on this branch to fail | +| B8 | The field-mapping checks catch the three confusions | proven | A4: the offline compiler owns this check | +| B9 | A cancelled task is neither dispatched nor read as a closed input | proven | found by A5's enumeration rather than by reading: acceptance already refused a cancelled task (the one predicate reads `cancellations`), while the eligibility rule could not see the cancellation at all - `DispatchTask` carried no such fact - so a cancelled task stayed selectable as the next dispatch. `DispatchTask.cancelled` now carries the same recorded fact the predicate reads, and `nextTask` refuses it in `ready()` and in `valid()`; pinned by the cases "a cancelled task is not dispatched, and nothing reads one as a closed input" and "nextTask refuses a task marked cancelled, whatever else the caller set", and by mutants `a-cancelled-task-is-still-dispatched` and `the-dispatch-does-not-require-a-cancelled-input-to-be-closed`. Scope: the round has no per-task cancellation source yet (its cancel is run-level, and `recordedFacts` fills no `cancellations`), so nothing changes for it today - this closes the shared derivation's half of the design's 取消后的有效性 | ## C. Recomputation (design: "重算检查须证明") -| node | obligation | state | evidence | -| ---- | ----------------------------------------------------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| C1 | Deleting every derived cache yields the same view | proven | `tests/integration/ooo-task-tables.test.ts`; mutant `derived-rebuild-is-a-no-op` | -| C2 | A lease crossing its boundary invalidates the old view | proven | `evals/ooo-execution/recovery.test.ts`; mutant `stale-claim-may-deliver-again` | -| C3 | Concurrent readers of one ready task: one legal claim | proven | `tests/integration/ooo-managed-fence.test.ts`. The test now reaches the board directly as well, because the round's own `live()` check refuses first and a mutant that broke only the store would have survived: a second reader on its own connection calls `claimTaskBoardEntry` and is refused, which is what makes "one legal claim" a property of the store rather than of one caller's discipline. Mutant `a-live-claim-can-be-taken-by-another-agent` (caught) | -| C4 | An unknown external result in a crash window is not guessed | proven | `tests/integration/ooo-external-window.test.ts`. `externalReady` refuses to make an event ready while a check for that task is outstanding ("managed check requires bound terminal evidence"), so the window cannot be closed by announcing it; after a restart the stored fact is still `external_ready = 0` with no artifact, and an invented event name is refused | +| node | obligation | state | evidence | +| ---- | ----------------------------------------------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| C1 | Deleting every derived cache yields the same view | proven | `tests/integration/ooo-task-tables.test.ts`; mutant `derived-rebuild-is-a-no-op` | +| C2 | A lease crossing its boundary invalidates the old view | proven | `evals/ooo-execution/recovery.test.ts`; mutant `stale-claim-may-deliver-again` | +| C3 | Concurrent readers of one ready task: one legal claim | proven | `tests/integration/ooo-managed-fence.test.ts`. The test now reaches the board directly as well, because the round's own `live()` check refuses first and a mutant that broke only the store would have survived: a second reader on its own connection calls `claimTaskBoardEntry` and is refused, which is what makes "one legal claim" a property of the store rather than of one caller's discipline. The tooth `a-live-claim-can-be-taken-by-another-agent` (caught) was retired on 2026-09-24: `one reader of the same ready task is given the claim, the second is refused` asserts the refusal and then reads the owner back as a raw row | +| C4 | An unknown external result in a crash window is not guessed | proven | `tests/integration/ooo-external-window.test.ts`. `externalReady` refuses to make an event ready while a check for that task is outstanding ("managed check requires bound terminal evidence"), so the window cannot be closed by announcing it; after a restart the stored fact is still `external_ready = 0` with no artifact, and an invented event name is refused | ## D. Lifecycle and integration pre-conditions (design: 事务参与与连接生命周期, 当前实现与接入前置条件) -| node | obligation | state | evidence | -| ---- | ------------------------------------------------------------------------------------------------------------------------------------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| D1 | Only the owner opens, migrates, checkpoints and closes a Store | proven | the round borrows its store and never closes it (D7), the read-only open is its own capability, and the daemon's close now refuses to drop accepted calls and is idempotent (D8). The narrow port and the borrowed status view carry the rest | -| D2 | A port exposes no raw connection, SQL, transaction control or `close()` | proven | `RoundQueryPort` and `TransactionPort` in `src/integration/ooo-board.ts`; `tests/integration/ooo-round-query.test.ts` | -| D3 | A write inside an open transition without a port is refused | proven | `tests/core/store-transaction-port.test.ts`; mutant `nested-write-transaction-is-allowed` | -| D4 | status's borrowed view migrates nothing, publishes nothing, initialises nothing | proven | `tests/integration/ooo-round-query.test.ts`; mutant `the-status-read-path-opens-the-rounds-store` (the mutant opens the writer's path, and the suite catches it) | -| D5 | The two read paths agree on the same facts and the same evaluation time | proven | `tests/integration/ooo-read-paths-agree.test.ts`. The comparison is no longer vacuous: the case accepts a task through the round's own host-verified path first, then requires both paths to report the same acceptance _and the same artifact bytes_. Mutant `the-offline-reader-decides-acceptance-on-its-own` (the borrowed port answers `{}` from its own rule; caught) | -| D6 | A read-only open is protected by the handle, not by `query_only` | proven | `tests/core/store-readonly-open.test.ts`; mutant `the-read-only-factory-opens-a-writable-handle` | -| D7 | `round` consumes an operations port and never calls `close()`; the outer host owns the Store | not applicable | **Retired with its subject.** The rule was proven on the round's ports and is now the shape of every host instead: the daemon owns the Store, a driver borrows a served store and closes nothing of its own, and the coordinated-write fence is the store's own state rather than a caller's discipline (D14). The round-side carriers (`OooRoundOperations`, `openRoundStore()`, `CycleOptions.operations`, the constructor-installed specs) were deleted with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)) | -| D8 | Shutdown stops new work, lets the work in flight finish, then closes once | proven | `src/cli/service.ts`: `close()` refuses while calls are in flight, `drain()` is the awaited middle step, and `inFlight` counts what the daemon accepted. `tests/cli/service-drain.test.ts` holds an accepted call in flight (search opens the store, so the counter sees work a synchronous close could not), requires `drain(0)` to fail rather than return quietly, then closes once; `tests/cli/service.test.ts` 51/0 and `archive-shutdown` 3/0 are the regression evidence | -| D9 | A post-commit notification failure is recorded, not thrown | proven | `tests/integration/ooo-post-commit-notification.test.ts`: a unit whose post-commit `afterCommit` throws still submits as `accepted`, its artifact is the board's, and `lastNotificationFailure()` reports the message, while the same unit with a reachable notification reports nothing. Verified by mutation (hand-run, 2026-09-18): deleting the `try`/`catch` in `submit()` makes the notification escape and the case fails. The implementer was always the board (`src/integration/ooo-board.ts`, `lastNotificationFailure()`); the evals suite that used to pin it drove a real round only because the artifact envelope is the board's business, and the round was retired ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The seam took three tries to find - `publishReady` is private and is called without a port at the commit, and `afterCommit` also runs on non-commit paths, so the failure is gated on a commit having landed | -| D10 | `runCycle`'s early-cancel branch is reachable, or it is dead code | not applicable | **Not applicable: the branch went with the round.** It was reachable and pinned twice (`evals/ooo-execution/early-cancel.test.ts` failed with "database is not open" when the round closed the store it borrowed; verified by hand, restored byte-identically), and the retirement deleted both the branch and its subject ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The property it stood behind - nobody who borrows a store closes it - now holds by construction, because no client opens one: the daemon is the only writer and drivers borrow a served store (D14) | -| D11 | The run records live in the store's schema, and their typed writes join the store's transaction | proven | `task_run_manifest`, `task_run_tasks` and `task_run_facts` are created by `migrate()` like every other table (`tests/core/store/schema.test.ts` names them), and `NmgStoreBase` writes them through `registerTaskRun`, `freezeTaskRunTask` and `appendTaskRunFact`, each taking the same optional port the board writes take: outside a transition it opens one, inside it joins the caller's. What the cases prove (`tests/core/store/task-runs.test.ts`, 8/8): a board write and a run fact written through one port land together, and when the fact write fails the board row does not survive it; a retried fact is recorded once and keeps its first sequence; two runs in one store keep separate task and fact namespaces; a run's plan cannot be replaced and a frozen task cannot be redefined; an entry bound by two runs is refused rather than answered with one of them; the reads register nothing. Mutants: `the-run-fact-opens-its-own-transaction`, `a-second-plan-overwrites-the-frozen-one`, `a-second-run-adopts-a-bound-entry`. Two rules this row states are carried by the cases rather than by a mutant: `appending the same fact twice records it once and keeps the first sequence` and `freezing a task twice is a no-op, and a different definition for it is refused` say the retry identity and the frozen definition in their own assertions, so `a-retried-run-fact-is-appended-twice` and `a-frozen-task-is-replaced-by-a-different-definition` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). One layer deliberately has no mutant of its own: the `UNIQUE (run_id, kind, task_id, attempt)` constraint is a storage-level backstop behind the explicit identity check, so a mutant that dropped only the constraint would survive - the check is the deciding layer and is the one mutated above | -| D12 | The binding of a logical task to its board entry is a run fact, and one routing rule decides what a binding makes managed | proven | `bindRunEntry` in `src/integration/task-coordinator.ts` records the binding as the run fact `entry-bound` (the name the D11 store tests already used), and it is the binding that the store's fence reads, so a bound entry's lifecycle write is refused unless it runs inside the run's coordinated transition. The op joins a caller's transaction when given a port, which is what lets a host create the entry and bind it in one transition: `tests/integration/ooo-managed-adopt.test.ts` (5 cases) checks that the board never holds an entry the run does not (a failure after both writes leaves neither), that a retry on the same task and attempt is not a second binding while another entry for that task and attempt is refused rather than silently dropped, and that every refusal names a stored fact (unregistered run, cancelled run, a task the run never froze, an entry that is not on the named channel, an entry another run already carries). The idempotence half rests on that case rather than on a mutant - `a-second-entry-rebinds-the-task` was retired on 2026-09-24 - because the case asserts the retry's shape and the refusal by name. `coordinatedEntryWrite` moves to the same module as the single routing rule, so the daemon's board verbs and any in-process writer decide "managed" in one place instead of each caller's belief about the entry; `src/cli/service.ts` no longer keeps its own copy. Mutants: `an-adopted-entry-takes-the-direct-path`, `a-binding-ignores-whether-the-task-was-frozen`, `a-binding-does-not-check-the-entry-exists`, `one-entry-is-bound-to-two-tasks`, `the-daemon-verb-skips-the-coordinated-path` (re-anchored to the daemon's claim handler, since the routing it mutates moved out of that file) | -| D13 | The run surface another process reaches is the daemon's: register, freeze, bind, cancel and status, and nothing else writes run facts | proven | `taskRun` in `src/cli/protocol.ts` over `registerRun` / `freezeRunPlan` / `bindRunEntry` / `cancelRun` / `taskRunStatus` / `createBoundEntry` in `src/integration/task-coordinator.ts`; `tests/cli/task-run-surface.test.ts` (11 cases) drives all of it through `service.invoke`, so what it exercises is the wire a second process uses, and it asserts `hello.methods` carries `taskRun` - the advertisement a client gates the `adopt` field on. Registration is the only transition that does not need a run to exist already (`coordinateRunWrite` refuses an unregistered run, which is what every later transition and every managed write rests on). A plan freeze is one transition: the request's array order becomes the positions, a dangling dependency and a self-dependency are refused by name while the plan is still only a proposal, and a batch the store refuses leaves the earlier tasks unfrozen - `the-plan-freezes-one-task-per-transaction`, `every-task-is-frozen-at-position-zero`, `the-plan-may-freeze-a-dangling-dependency`, `a-task-may-depend-on-itself` - and a cancelled run takes no further plan (`a-cancelled-run-takes-a-new-plan`). Adoption rides the transition that creates the entry (`createBoundEntry` takes the port), so a refusal leaves no entry behind (`the-entry-is-created-before-its-binding-is-checked`) and the wire cannot drop the request silently (`the-wire-drops-an-adoption-request`), which is the one failure mode the epoch rule in `design.md` is about. The binding fact records the channel that names its entry - that is what lets `status` resolve a binding without searching the board (`a-binding-does-not-record-its-channel`). `cancel` is the only writer of `run-cancelled`: before this row `src/` had none, so the state the fence and the dispatch derivation both read was reachable only from a test (the gap G4 named); a run-level cancellation carries the schema's empty task id (`a-run-cancellation-names-a-task`), a task-level one requires the task to have been frozen (`a-cancellation-ignores-whether-the-task-was-frozen`, `an-unknown-run-can-be-cancelled`), and cancelling twice is one fact because the fact's own identity is the duplicate key. `status` is a read and derives nothing: a run the store does not know stays unknown (`status-registers-the-run-it-cannot-find`), and ready/blocked/accepted stay with the shared pure function rather than with this view. The CLI exposes the operator half (`nmg run status`, `nmg run cancel`), which is also what the "every RPC method is a CLI command or intentionally RPC-only" gate asks for. One thing this pass corrected about itself: the first version's parser also computed a `position` per task, which `freezeRunPlan` immediately overwrote - the mutant that changed it survived, which is how the dead field was found, and it was deleted rather than kept | -| D14 | The evidence drivers reach the board through the daemon and open no database of their own | proven | `evals/ooo-execution/round-client.ts` resolves the daemon from the store's own lease and refuses by name when nothing serves it; `round-host.ts` serves an existing round store the way the product daemon does (`NmgService` + `serveHttp` + the store's lease) and is that store's only writer while it runs; `board-worker.ts`, `board-deliver.ts` and `board-judge.ts` now use the protocol's board verbs (`read`, `claim`, `release`, `deliver`, `judge`) and no longer import the store. `tests/integration/ooo-evidence-drivers.test.ts`: both end-to-end cases host the store and spawn the drivers as separate processes against the endpoint (spawned, not `spawnSync`, because the test process is what answers them); the managed case registers a run and freezes its plan through `taskRun`, adopts the entry in the same `put` that creates it, then runs the worker and the judge as separate processes - the run's log afterwards is exactly `entry-bound`, `board-claim`, `board-deliver`, `board-judge`, with the binding resolved to the entry the worker claimed and its verdict accepted. Two cases hold the boundary itself: a driver with no daemon refuses by name instead of falling back to the file, and no `board-*.ts` driver may mention the store or omit the round client. Mutants: `a-driver-falls-back-to-opening-the-store`, `the-round-client-does-not-require-a-daemon`, plus the drivers' two earlier ones (`worker-reads-only-one-reporter-shape`, `judge-may-judge-its-own-delivery`), re-pointed at the renamed case. Two further rules came out of building this, and both are about the same defect the first attempt hit: a process that serves an endpoint cannot answer a call to it while it is blocked, so "the answer will come" is not an assumption a client may make. `round-client.ts` therefore bounds every call (`ROUND_CALL_TIMEOUT_MS`, an opt-in bound on the thin client's `httpCall`, which by default still leaves the platform's own) and reports a timeout by naming the bound, the endpoint and the pid that serves it rather than the transport's; and it refuses when the lease's pid is the caller's own, directing an in-process host to `host.call(...)` (`round-host.ts`), which is the design's shape for an offline host. The test is now a client too: it starts the host as its own process (`[host] serving `, idle timeout as the backstop for a host whose owner died) instead of hosting in-process, and asks it to shut down rather than killing it, so what runs on the way out is the release path - a lease held by a process that is gone is a store nothing can serve. Cases: "a client refuses to call the endpoint its own process serves", "a call to a host that never answers gives up in seconds and names the reason" (raced against its own 5s deadline, so the check of the bound is itself bounded), "a host releases its lease when it stops, so the next host can take the store". Mutants: `the-round-client-calls-the-endpoint-it-serves`, `the-round-client-has-no-limit-on-how-long-it-waits`, and - new target, `src/cli/http-server.ts` - `the-serving-process-never-releases-its-lease`. Not covered by this row: `live-continuation.ts` still creates its own store and constructs `BoardAdmission` - the runner half of the same design sentence and the next obligation. `round-runner.ts` did the same and was retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)) | +| node | obligation | state | evidence | +| ---- | ------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| D1 | Only the owner opens, migrates, checkpoints and closes a Store | proven | the round borrows its store and never closes it (D7), the read-only open is its own capability, and the daemon's close now refuses to drop accepted calls and is idempotent (D8). The narrow port and the borrowed status view carry the rest | +| D2 | A port exposes no raw connection, SQL, transaction control or `close()` | proven | `RoundQueryPort` and `TransactionPort` in `src/integration/ooo-board.ts`; `tests/integration/ooo-round-query.test.ts` | +| D3 | A write inside an open transition without a port is refused | proven | `tests/core/store-transaction-port.test.ts`; the case `a write entry reached inside a transition without a port is refused, not nested` states it, and it also asserts that the refused write left nothing behind; the tooth `nested-write-transaction-is-allowed` was retired on 2026-09-24 | +| D4 | status's borrowed view migrates nothing, publishes nothing, initialises nothing | proven | `tests/integration/ooo-round-query.test.ts`; mutant `the-status-read-path-opens-the-rounds-store` (the mutant opens the writer's path, and the suite catches it) | +| D5 | The two read paths agree on the same facts and the same evaluation time | proven | `tests/integration/ooo-read-paths-agree.test.ts`. The comparison is no longer vacuous: the case accepts a task through the round's own host-verified path first, then requires both paths to report the same acceptance _and the same artifact bytes_. Mutant `the-offline-reader-decides-acceptance-on-its-own` (the borrowed port answers `{}` from its own rule; caught) | +| D6 | A read-only open is protected by the handle, not by `query_only` | proven | `tests/core/store-readonly-open.test.ts`; the case `a read-only open neither creates, migrates nor writes` states it - the refusal, the missing file that was not created, and the unchanged hash of the store; the tooth `the-read-only-factory-opens-a-writable-handle` was retired on 2026-09-24 | +| D7 | `round` consumes an operations port and never calls `close()`; the outer host owns the Store | not applicable | **Retired with its subject.** The rule was proven on the round's ports and is now the shape of every host instead: the daemon owns the Store, a driver borrows a served store and closes nothing of its own, and the coordinated-write fence is the store's own state rather than a caller's discipline (D14). The round-side carriers (`OooRoundOperations`, `openRoundStore()`, `CycleOptions.operations`, the constructor-installed specs) were deleted with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)) | +| D8 | Shutdown stops new work, lets the work in flight finish, then closes once | proven | `src/cli/service.ts`: `close()` refuses while calls are in flight, `drain()` is the awaited middle step, and `inFlight` counts what the daemon accepted. `tests/cli/service-drain.test.ts` holds an accepted call in flight (search opens the store, so the counter sees work a synchronous close could not), requires `drain(0)` to fail rather than return quietly, then closes once; `tests/cli/service.test.ts` 51/0 and `archive-shutdown` 3/0 are the regression evidence | +| D9 | A post-commit notification failure is recorded, not thrown | proven | `tests/integration/ooo-post-commit-notification.test.ts`: a unit whose post-commit `afterCommit` throws still submits as `accepted`, its artifact is the board's, and `lastNotificationFailure()` reports the message, while the same unit with a reachable notification reports nothing. Verified by mutation (hand-run, 2026-09-18): deleting the `try`/`catch` in `submit()` makes the notification escape and the case fails. The implementer was always the board (`src/integration/ooo-board.ts`, `lastNotificationFailure()`); the evals suite that used to pin it drove a real round only because the artifact envelope is the board's business, and the round was retired ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The seam took three tries to find - `publishReady` is private and is called without a port at the commit, and `afterCommit` also runs on non-commit paths, so the failure is gated on a commit having landed | +| D10 | `runCycle`'s early-cancel branch is reachable, or it is dead code | not applicable | **Not applicable: the branch went with the round.** It was reachable and pinned twice (`evals/ooo-execution/early-cancel.test.ts` failed with "database is not open" when the round closed the store it borrowed; verified by hand, restored byte-identically), and the retirement deleted both the branch and its subject ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The property it stood behind - nobody who borrows a store closes it - now holds by construction, because no client opens one: the daemon is the only writer and drivers borrow a served store (D14) | +| D11 | The run records live in the store's schema, and their typed writes join the store's transaction | proven | `task_run_manifest`, `task_run_tasks` and `task_run_facts` are created by `migrate()` like every other table (`tests/core/store/schema.test.ts` names them), and `NmgStoreBase` writes them through `registerTaskRun`, `freezeTaskRunTask` and `appendTaskRunFact`, each taking the same optional port the board writes take: outside a transition it opens one, inside it joins the caller's. What the cases prove (`tests/core/store/task-runs.test.ts`, 8/8): a board write and a run fact written through one port land together, and when the fact write fails the board row does not survive it; a retried fact is recorded once and keeps its first sequence; two runs in one store keep separate task and fact namespaces; a run's plan cannot be replaced and a frozen task cannot be redefined; an entry bound by two runs is refused rather than answered with one of them; the reads register nothing. Mutant: `a-second-run-adopts-a-bound-entry`. Two further teeth were retired here on 2026-09-24, because the cases state the rules: `the-run-fact-opens-its-own-transaction` (`a board write and a run fact land together, and neither lands alone` opens one transition through the port, asserts both halves, then asserts neither survives a failure) and `a-second-plan-overwrites-the-frozen-one` (`a run registers once, and a second plan for the same run is refused` asserts the refusal by name and that the manifest still reads the first plan). Two rules this row states are carried by the cases rather than by a mutant: `appending the same fact twice records it once and keeps the first sequence` and `freezing a task twice is a no-op, and a different definition for it is refused` say the retry identity and the frozen definition in their own assertions, so `a-retried-run-fact-is-appended-twice` and `a-frozen-task-is-replaced-by-a-different-definition` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). One layer deliberately has no mutant of its own: the `UNIQUE (run_id, kind, task_id, attempt)` constraint is a storage-level backstop behind the explicit identity check, so a mutant that dropped only the constraint would survive - the check is the deciding layer and is the one mutated above | +| D12 | The binding of a logical task to its board entry is a run fact, and one routing rule decides what a binding makes managed | proven | `bindRunEntry` in `src/integration/task-coordinator.ts` records the binding as the run fact `entry-bound` (the name the D11 store tests already used), and it is the binding that the store's fence reads, so a bound entry's lifecycle write is refused unless it runs inside the run's coordinated transition. The op joins a caller's transaction when given a port, which is what lets a host create the entry and bind it in one transition: `tests/integration/ooo-managed-adopt.test.ts` (5 cases) checks that the board never holds an entry the run does not (a failure after both writes leaves neither), that a retry on the same task and attempt is not a second binding while another entry for that task and attempt is refused rather than silently dropped, and that every refusal names a stored fact (unregistered run, cancelled run, a task the run never froze, an entry that is not on the named channel, an entry another run already carries). The idempotence half rests on that case rather than on a mutant - `a-second-entry-rebinds-the-task` was retired on 2026-09-24 - because the case asserts the retry's shape and the refusal by name. `coordinatedEntryWrite` moves to the same module as the single routing rule, so the daemon's board verbs and any in-process writer decide "managed" in one place instead of each caller's belief about the entry; `src/cli/service.ts` no longer keeps its own copy. Mutants: `an-adopted-entry-takes-the-direct-path`, `a-binding-ignores-whether-the-task-was-frozen`, `a-binding-does-not-check-the-entry-exists`, `one-entry-is-bound-to-two-tasks`, `the-daemon-verb-skips-the-coordinated-path` (re-anchored to the daemon's claim handler, since the routing it mutates moved out of that file) | +| D13 | The run surface another process reaches is the daemon's: register, freeze, bind, cancel and status, and nothing else writes run facts | proven | `taskRun` in `src/cli/protocol.ts` over `registerRun` / `freezeRunPlan` / `bindRunEntry` / `cancelRun` / `taskRunStatus` / `createBoundEntry` in `src/integration/task-coordinator.ts`; `tests/cli/task-run-surface.test.ts` (11 cases) drives all of it through `service.invoke`, so what it exercises is the wire a second process uses, and it asserts `hello.methods` carries `taskRun` - the advertisement a client gates the `adopt` field on. Registration is the only transition that does not need a run to exist already (`coordinateRunWrite` refuses an unregistered run, which is what every later transition and every managed write rests on). A plan freeze is one transition: the request's array order becomes the positions, a dangling dependency and a self-dependency are refused by name while the plan is still only a proposal, and a batch the store refuses leaves the earlier tasks unfrozen - `the-plan-freezes-one-task-per-transaction`, `every-task-is-frozen-at-position-zero`, `the-plan-may-freeze-a-dangling-dependency`, `a-task-may-depend-on-itself` - and a cancelled run takes no further plan (`a-cancelled-run-takes-a-new-plan`). Four teeth named in this row were retired on 2026-09-24, each because `tests/cli/task-run-surface.test.ts` states its rule through the daemon's own surface: `the-plan-freezes-one-task-per-transaction` (`a refused freeze leaves the plan exactly as it was`), `a-cancelled-run-takes-a-new-plan` (`a cancelled run takes no further plan`), `the-entry-is-created-before-its-binding-is-checked` (`adoption is part of the transition that creates the entry, so a refusal leaves no entry`) and `status-registers-the-run-it-cannot-find` (`status is a read: an unknown run has no manifest and is not registered by being asked`). Adoption rides the transition that creates the entry (`createBoundEntry` takes the port), so a refusal leaves no entry behind (`the-entry-is-created-before-its-binding-is-checked`) and the wire cannot drop the request silently (`the-wire-drops-an-adoption-request`), which is the one failure mode the epoch rule in `design.md` is about. The binding fact records the channel that names its entry - that is what lets `status` resolve a binding without searching the board (`a-binding-does-not-record-its-channel`). `cancel` is the only writer of `run-cancelled`: before this row `src/` had none, so the state the fence and the dispatch derivation both read was reachable only from a test (the gap G4 named); a run-level cancellation carries the schema's empty task id (`a-run-cancellation-names-a-task`), a task-level one requires the task to have been frozen (`a-cancellation-ignores-whether-the-task-was-frozen`, `an-unknown-run-can-be-cancelled`), and cancelling twice is one fact because the fact's own identity is the duplicate key. `status` is a read and derives nothing: a run the store does not know stays unknown (`status-registers-the-run-it-cannot-find`), and ready/blocked/accepted stay with the shared pure function rather than with this view. The CLI exposes the operator half (`nmg run status`, `nmg run cancel`), which is also what the "every RPC method is a CLI command or intentionally RPC-only" gate asks for. One thing this pass corrected about itself: the first version's parser also computed a `position` per task, which `freezeRunPlan` immediately overwrote - the mutant that changed it survived, which is how the dead field was found, and it was deleted rather than kept | +| D14 | The evidence drivers reach the board through the daemon and open no database of their own | proven | `evals/ooo-execution/round-client.ts` resolves the daemon from the store's own lease and refuses by name when nothing serves it; `round-host.ts` serves an existing round store the way the product daemon does (`NmgService` + `serveHttp` + the store's lease) and is that store's only writer while it runs; `board-worker.ts`, `board-deliver.ts` and `board-judge.ts` now use the protocol's board verbs (`read`, `claim`, `release`, `deliver`, `judge`) and no longer import the store. `tests/integration/ooo-evidence-drivers.test.ts`: both end-to-end cases host the store and spawn the drivers as separate processes against the endpoint (spawned, not `spawnSync`, because the test process is what answers them); the managed case registers a run and freezes its plan through `taskRun`, adopts the entry in the same `put` that creates it, then runs the worker and the judge as separate processes - the run's log afterwards is exactly `entry-bound`, `board-claim`, `board-deliver`, `board-judge`, with the binding resolved to the entry the worker claimed and its verdict accepted. Two cases hold the boundary itself: a driver with no daemon refuses by name instead of falling back to the file, and no `board-*.ts` driver may mention the store or omit the round client. Mutants: `the-round-client-does-not-require-a-daemon`, plus the drivers' two earlier ones (`worker-reads-only-one-reporter-shape`, `judge-may-judge-its-own-delivery`), re-pointed at the renamed case. Two further rules came out of building this, and both are about the same defect the first attempt hit: a process that serves an endpoint cannot answer a call to it while it is blocked, so "the answer will come" is not an assumption a client may make. `round-client.ts` therefore bounds every call (`ROUND_CALL_TIMEOUT_MS`, an opt-in bound on the thin client's `httpCall`, which by default still leaves the platform's own) and reports a timeout by naming the bound, the endpoint and the pid that serves it rather than the transport's; and it refuses when the lease's pid is the caller's own, directing an in-process host to `host.call(...)` (`round-host.ts`), which is the design's shape for an offline host. The test is now a client too: it starts the host as its own process (`[host] serving `, idle timeout as the backstop for a host whose owner died) instead of hosting in-process, and asks it to shut down rather than killing it, so what runs on the way out is the release path - a lease held by a process that is gone is a store nothing can serve. Cases: "a client refuses to call the endpoint its own process serves", "a call to a host that never answers gives up in seconds and names the reason" (raced against its own 5s deadline, so the check of the bound is itself bounded), "a host releases its lease when it stops, so the next host can take the store". Mutants: `the-round-client-calls-the-endpoint-it-serves`, `the-round-client-has-no-limit-on-how-long-it-waits`, and - new target, `src/cli/http-server.ts` - `the-serving-process-never-releases-its-lease`. Not covered by this row: `live-continuation.ts` still creates its own store and constructs `BoardAdmission` - the runner half of the same design sentence and the next obligation. `round-runner.ts` did the same and was retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). `a-driver-falls-back-to-opening-the-store` was retired on 2026-09-24: the case above asserts both halves of the boundary (the refusal by name, and that no driver mentions the store), which is the check the tooth used to stand in for. | ## E. Optional HA/MGR integration (design: 复用 autodiff、HA 与 MGR) @@ -207,9 +219,9 @@ and never read for legality, so a soft premise cannot make anything legal. | node | obligation | state | evidence | | ---- | ------------------------------------------------------------------------------------------------ | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| E1 | An out-of-set action is refused even when it scores highest | proven | `tests/integration/ooo-advisers.test.ts` "a suggestion outside the legal set is refused, however high it scores" (the 1e9 score buys order inside the set, never a member); mutants `a-suggestion-outside-the-legal-set-is-scored`, `the-ordering-adds-a-task-to-the-set` (the ordering returns `[...legal, ...best.keys()]`) | +| E1 | An out-of-set action is refused even when it scores highest | proven | `tests/integration/ooo-advisers.test.ts` "a suggestion outside the legal set is refused, however high it scores" (the 1e9 score buys order inside the set, never a member); mutant `a-suggestion-outside-the-legal-set-is-scored`; `the-ordering-adds-a-task-to-the-set` was retired on 2026-09-24, because `a suggestion outside the legal set is refused, however high it scores` enumerates the outcome - one refusal, the order sorted equal to the input set, one adopted entry - which a member added by the ordering fails | | E2 | A soft premise or `requires`-style gate cannot unlock a real dependency | proven | the same file, "a soft premise cannot unlock a dependency the shared rules refused": the fixture plan's `A` depends on `B`, and the suggestion for `A` carries `assumptions: ["requires: B", "hypothesis: B is satisfied"]` with a 1e9 score while `selectableTasks` returns `["B"]`; it is refused as out-of-set, and the same suggestion is adopted in the control case where `B`'s artifact is really accepted, so the refusal is about legality and not about the fixture. No line of the seam reads `assumptions`; the out-of-set refusal is the layer that enforces this, which is why E1's mutant is what fails this case too | -| E3 | A disabled or failing source falls back to the rule policy, and the run is told why | proven | "a disabled or failing source falls back to the rule policy, and says why": a `enabled: false` source and one that throws leave the order equal to the rule order, with two named fallbacks (`disabled`, `failed: no trained state is available`); mutants `a-disabled-source-is-asked-anyway`, `a-failing-source-takes-the-decision-with-it` | +| E3 | A disabled or failing source falls back to the rule policy, and the run is told why | proven | "a disabled or failing source falls back to the rule policy, and says why": a `enabled: false` source and one that throws leave the order equal to the rule order, with two named fallbacks (`disabled`, `failed: no trained state is available`); mutant `a-disabled-source-is-asked-anyway`; `a-failing-source-takes-the-decision-with-it` was retired on 2026-09-24, because `a disabled or failing source falls back to the rule policy, and says why` enumerates the fallbacks by source id and their reasons | | E4 | A score does not cross a session or a branch | proven | "a score from another session or branch is not reused": provenance from another session, and from another branch, is refused by name - the projection carries both identities and the module holds no state to remember one anyway; mutant `a-score-from-another-scope-is-reused` | | E5 | A changed parameter or projection version does not reuse an old suggestion | proven | "a changed parameter or projection version makes an old score a new one": both mismatches are reported as re-scored with the differing input named, and neither orders the set; mutant `a-version-mismatch-still-counts-as-the-same-reading` | | E6 | A score whose history is missing is never reported as a reproduction | proven | "a score that cannot name its own history is re-scored, never reported as a reproduction": a provenance without `observationOrder`/`initialState`, and one with a _different_ observation order, are both re-scored with the missing inputs named; mutant `a-score-with-no-recorded-history-counts-as-a-reproduction` | @@ -247,16 +259,16 @@ orders the offline layer first ("离线模型先覆盖不同粒度、依赖密 发现逻辑错误和成本转折点,不能预测真实模型质量"), and the repository had none: this pass built it, and the paid stage then ran on a family held out of it. -| Step | State | Evidence | -| -------------------------------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| F1: advisory offline cost model | landed | `evals/ooo-execution/cost-model.ts` (`--sweep`, derived graphs, self-checks) + `cost-model.test.ts` (12 cases) + [the sweep record](../experiments/execution/ooo-cost-model-2026-09-17.md) | -| F2a: the plan has one home, and the round's log names it | landed | the plan is one value with one home: the spec the driver runs (`PlanDriverSpec.plan`) and the run manifest the store freezes (D12). F2a's round-side carriers (`DEFAULT_ROUND_PLAN`, `CycleOptions.plan`, `openRoundStore(path, plan)`, `round-plan.test.ts` with its 6 cases) were retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); [the arms' driver decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) is what made the driver the home in the first place | -| F2b: a research-side driver for arbitrary legal plans | landed, one arm | `evals/ooo-execution/plan-driver.ts` (`runPlan`, `comparePlanSlots`, `verifyParent` path, refusal-naming CLI) + `plan-driver.test.ts` (15 cases) + 12 named mutants (`tools/mutation-teeth.ts`, target `evals/ooo-execution/plan-driver.ts`) + `BoardAdmission.candidates()` (the ordered legal set; `next()` is its head, with its own mutant) + [the decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) | -| F2c: pick the parent task from the sweep's turning point | landed | Two families of the shape the design's parent family needs - a frozen interface, three independent builders, one summary that depends on all three - as a spec pair each: `evals/ooo-execution/fixtures/report/` (the instrument's own) and `fixtures/pipeline/` (**held out**: written after the driver, and not the family anything was tuned against). The coarse spec is one unit over the whole task checked by every frozen test; the fine spec is four units with per-unit checks and no `join`, so the parent check composes all four. Sibling units are code-independent by construction (the summary takes the derived values as parameters), which is what lets a unit be verified before its siblings exist. Offline, with the instrument's own answers and a wrong one: `evals/ooo-execution/families.test.ts` (8 cases over both families) - both plans accept, a wrong answer is rejected by its own check and takes the composition with it, and a unit that declares no checks and has no file-wide list is refused by name. Driver support this needed: per-unit `checks` with the file-wide list as a fallback, a `canned` worker (the instrument's answer, so the family's acceptance is shown before any model is paid), and the parent composition fix recorded in the pilot's experiment record | -| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-driver-declares-one-slot-whatever-the-spec-says`, `the-driver-awaits-each-unit-instead-of-the-batch`. The budget's licence and publication rules are pinned by `evals/ooo-execution/board-slots.test.ts`'s own case instead: `a declared budget holds two claims at once, and the store is why each handoff is directed` refuses a budget above one with no target by that message, asserts one handoff per startable task with `serialState` null for each, keeps every startable handoff across a republish by identity, and claims the non-head first - so `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected` and `a-multi-slot-handoff-is-published-un-directed` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | -| F3: real-model pilot (A 3 / B 3 / C 2, current pi model, directional only) | landed | `evals/ooo-execution/pilot.ts` (`--live` required, `--report` to re-aggregate recorded runs with no model call, refuses a merge of two instruments, seeded arm order) + [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md). Fixed: `deepseek/deepseek-v4-flash`, envelope limits `turns 6 / reads 3 / 120 s` for every arm, A 3 / B 3 / C 2, arms drawn from a seeded shuffle, the held-out `pipeline` family. Measured (8 runs, 23 model calls, every run complete, all 8 parent checks accepted): per run A 8.6 s / 11.2 k tokens, B 21.5 s / 31.5 k tokens, C 16.2 s / 29.9 k tokens, host 1.6 s / 3.4 s / 3.6 s. All three of F1's expectations held: B is 2.5× A in wall time and 2.8× in tokens (one slot buys nothing), C recovers part of it (0.76× B) and not the 2× a pure model-call overlap would give, and the host cost grows with candidates rather than slots. Wasted cost 0, human intervention 0. **Directional only**: n = 8, one model, one held-out family, and no quality difference was available to measure - every arm accepted everything | -| F4: execution fusion - legality, accounting and the driver policy | landed, live arm measured | Rules: `sharedSessionLegal`/`fusionSuccessors`/`fusionCandidates` in `src/integration/ooo-execution.ts` (the design's five conditions, one line each, composed with the board's candidate answer) + `tests/integration/ooo-fusion.test.ts` (12 cases) + 8 named mutants (target `src/integration/ooo-execution.ts`). Accounting: `fusionAccounting`/`fusionVerdict` + a fusion block in `cost-model.ts --sweep` - two lines kept apart, `unmeasured` until a run prices the session startup + `cost-model.test.ts` (12 cases) + 4 named mutants. Policy: `PlanDriverSpec.fusion` + `PlanRun.sessions` in `evals/ooo-execution/plan-driver.ts`, with the board's candidate set as the authority on staleness/cancellation/delivery/waits + `plan-driver.test.ts` (15 cases) + 11 named mutants. Scoped sweeps on this revision: `src/integration/ooo-execution.ts` 17 of 17 caught, `evals/ooo-execution/cost-model.ts` 4 of 4, `src/core/store/clock.ts` 4 of 4, `evals/ooo-execution/plan-driver.ts` 11 of 11, each restored byte-identically. A driver mutant that only restated `sharedSessionLegal`'s own rule was deleted rather than kept: the suite could not distinguish it from the shared predicate, which is what one home for that rule means. Both functions this slice pushed above the complexity limit (`sharedSessionLegal` 18, `runOneUnit` 16) were brought under it by extracting helpers, not by raising the threshold; `npm run lint` and `npm run check` are clean on this revision. **The live half landed.** `createPiSessionRunner` holds one Pi session and its tool surface across units (each unit re-points one mutable `UnitState` box; `patchSessionInput` is the one place a unit's prompt, snapshot and bounds are built, and `executePiPatch`/`executePiSnapshot` are thin callers of it), `PiRun` separates a unit's own `tokens`/`cacheRead`/`cacheWrite` from the session's `sessionTokens`, and `piSessionWorker` + `--session-runner` hold one runner per driver session. First paid D arm (2026-09-19, `deepseek-v4-flash`, 2 units, `--slots 1`, `turns: 6`, 3 reps per bound, only `fusion.unitsPerSession` differing): fused ran one session of two units and the control two sessions of one, quality parity in all six runs (every unit accepted), median 22 533 against 22 498 tokens and 11 048 against 12 948 ms - so no token saving yet (~1.9 s per run, ~15 % of the unfused wall, which prices the session-startup term at ~1.9 s instead of leaving it assumed) and per unit the second one cost ~8 % less while the first cost more: a chain's tool surface is the union of its units' capabilities because a session's surface is fixed at creation, so a unit can spend a turn on a tool that refuses by name. Also fixed here: `specFrom` had silently dropped a spec file's `fusion` block, so a spec asking for fusion ran as the control arm. | -| F5: speculation lifecycle - one declared fact, three outcomes | landed (offline); the paid E arm ran once, no gain claimed | `SpeculationAssumption` / `ResolvedPredicate` / `SpeculationCandidate` / `isBoundedSpeculation` / `speculationOutcome` in `src/integration/ooo-execution.ts`, beside the fusion conditions: the assumption is a declaration the summary binds to (it never discovers for itself that the guess was false), the first experiment's bounds are a predicate (exactly one pending fact, nothing prepared from the guess - a speculative successor or an irreversible operation each refuse it by name), and the outcome has three states rather than two - **true** publishes, **false** discards the candidate and returns `sessionReusable: false`, which is what makes "失效会话不能复用到真实路径" a rule the caller must honour instead of a note, and **unknown** (no reading, an unattested reading, or evidence about another version) waits without publishing. Asking for the outcome of a candidate that is not the bounded shape throws rather than folding a fourth state into the three. Cases: `tests/integration/ooo-speculation.test.ts` (9), the last of which joins this half to fusion's condition 5 - an invalidated branch is not a legal predecessor for the real path. Mutants: 5 (`speculation-guesses-several-facts-at-once`, `a-guess-with-no-evidence-publishes`, `an-unattested-reading-counts-as-evidence`, `evidence-about-another-version-is-the-same-fact`, `a-contradicted-guess-keeps-its-session`); the target's sweep is 22 of 22 caught. The E arm's instrument now exists (`evals/ooo-execution/speculation-pilot.ts`, registered) and ran once (2026-09-19, 6 paid units, ~43 k tokens): it decides the guessed fact, prepares the candidate, applies `speculationOutcome`, and verifies a published candidate with the unit's own frozen check. **No result is claimed**: the quality term was false in all four verified candidates, so by the design's own rule the latency and cost shape may not be reported as a gain. The search behind those failures is now closed and its first reading was wrong: `artifactEnvelope` builds two legitimate shapes (a patch, and a conclusion with `kind, conclusion, summary, evidence, citations`), and the instrument had fed every artifact to the patch reader. Eight of nine attempts answered with a conclusion, which this unit's check cannot pass and the board would refuse; the one patch attempt failed on a real mistake (`rows` for `lines`). The instrument now reads by kind, keeps every artifact, candidate tree and check output, and the run is archived. Measured outcome of the arm at this shape: the post-fact cost drops from ~6.2 s of work to 175 ms of verification when the fact holds, the false-fact case wastes 20 332 tokens, and the prepared candidate was publishable in 0 of 3 holding reps - so the cost is real, the gain is not, and the binding constraint is the candidate's admissibility | +| Step | State | Evidence | +| -------------------------------------------------------------------------- | ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| F1: advisory offline cost model | landed | `evals/ooo-execution/cost-model.ts` (`--sweep`, derived graphs, self-checks) + `cost-model.test.ts` (12 cases) + [the sweep record](../experiments/execution/ooo-cost-model-2026-09-17.md) | +| F2a: the plan has one home, and the round's log names it | landed | the plan is one value with one home: the spec the driver runs (`PlanDriverSpec.plan`) and the run manifest the store freezes (D12). F2a's round-side carriers (`DEFAULT_ROUND_PLAN`, `CycleOptions.plan`, `openRoundStore(path, plan)`, `round-plan.test.ts` with its 6 cases) were retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); [the arms' driver decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) is what made the driver the home in the first place | +| F2b: a research-side driver for arbitrary legal plans | landed, one arm | `evals/ooo-execution/plan-driver.ts` (`runPlan`, `comparePlanSlots`, `verifyParent` path, refusal-naming CLI) + `plan-driver.test.ts` (15 cases) + 12 named mutants (`tools/mutation-teeth.ts`, target `evals/ooo-execution/plan-driver.ts`) + `BoardAdmission.candidates()` (the ordered legal set; `next()` is its head, with its own mutant) + [the decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) | +| F2c: pick the parent task from the sweep's turning point | landed | Two families of the shape the design's parent family needs - a frozen interface, three independent builders, one summary that depends on all three - as a spec pair each: `evals/ooo-execution/fixtures/report/` (the instrument's own) and `fixtures/pipeline/` (**held out**: written after the driver, and not the family anything was tuned against). The coarse spec is one unit over the whole task checked by every frozen test; the fine spec is four units with per-unit checks and no `join`, so the parent check composes all four. Sibling units are code-independent by construction (the summary takes the derived values as parameters), which is what lets a unit be verified before its siblings exist. Offline, with the instrument's own answers and a wrong one: `evals/ooo-execution/families.test.ts` (8 cases over both families) - both plans accept, a wrong answer is rejected by its own check and takes the composition with it, and a unit that declares no checks and has no file-wide list is refused by name. Driver support this needed: per-unit `checks` with the file-wide list as a fallback, a `canned` worker (the instrument's answer, so the family's acceptance is shown before any model is paid), and the parent composition fix recorded in the pilot's experiment record | +| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-driver-awaits-each-unit-instead-of-the-batch`. `the-driver-declares-one-slot-whatever-the-spec-says` was retired on 2026-09-24, because `a declared slot count is reached, and the claims overlap in time` asserts the requested count, the used count and the overlap. The budget's licence and publication rules are pinned by `evals/ooo-execution/board-slots.test.ts`'s own case instead: `a declared budget holds two claims at once, and the store is why each handoff is directed` refuses a budget above one with no target by that message, asserts one handoff per startable task with `serialState` null for each, keeps every startable handoff across a republish by identity, and claims the non-head first - so `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected` and `a-multi-slot-handoff-is-published-un-directed` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | +| F3: real-model pilot (A 3 / B 3 / C 2, current pi model, directional only) | landed | `evals/ooo-execution/pilot.ts` (`--live` required, `--report` to re-aggregate recorded runs with no model call, refuses a merge of two instruments, seeded arm order) + [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md). Fixed: `deepseek/deepseek-v4-flash`, envelope limits `turns 6 / reads 3 / 120 s` for every arm, A 3 / B 3 / C 2, arms drawn from a seeded shuffle, the held-out `pipeline` family. Measured (8 runs, 23 model calls, every run complete, all 8 parent checks accepted): per run A 8.6 s / 11.2 k tokens, B 21.5 s / 31.5 k tokens, C 16.2 s / 29.9 k tokens, host 1.6 s / 3.4 s / 3.6 s. All three of F1's expectations held: B is 2.5× A in wall time and 2.8× in tokens (one slot buys nothing), C recovers part of it (0.76× B) and not the 2× a pure model-call overlap would give, and the host cost grows with candidates rather than slots. Wasted cost 0, human intervention 0. **Directional only**: n = 8, one model, one held-out family, and no quality difference was available to measure - every arm accepted everything | +| F4: execution fusion - legality, accounting and the driver policy | landed, live arm measured | Rules: `sharedSessionLegal`/`fusionSuccessors`/`fusionCandidates` in `src/integration/ooo-execution.ts` (the design's five conditions, one line each, composed with the board's candidate answer) + `tests/integration/ooo-fusion.test.ts` (12 cases) + 8 named mutants (target `src/integration/ooo-execution.ts`). Accounting: `fusionAccounting`/`fusionVerdict` + a fusion block in `cost-model.ts --sweep` - two lines kept apart, `unmeasured` until a run prices the session startup + `cost-model.test.ts` (12 cases) + 4 named mutants. Policy: `PlanDriverSpec.fusion` + `PlanRun.sessions` in `evals/ooo-execution/plan-driver.ts`, with the board's candidate set as the authority on staleness/cancellation/delivery/waits + `plan-driver.test.ts` (15 cases) + 11 named mutants. Scoped sweeps on this revision: `src/integration/ooo-execution.ts` 17 of 17 caught, `evals/ooo-execution/cost-model.ts` 4 of 4, `src/core/store/clock.ts` 4 of 4, `evals/ooo-execution/plan-driver.ts` 11 of 11, each restored byte-identically. A driver mutant that only restated `sharedSessionLegal`'s own rule was deleted rather than kept: the suite could not distinguish it from the shared predicate, which is what one home for that rule means. Both functions this slice pushed above the complexity limit (`sharedSessionLegal` 18, `runOneUnit` 16) were brought under it by extracting helpers, not by raising the threshold; `npm run lint` and `npm run check` are clean on this revision. **The live half landed.** `createPiSessionRunner` holds one Pi session and its tool surface across units (each unit re-points one mutable `UnitState` box; `patchSessionInput` is the one place a unit's prompt, snapshot and bounds are built, and `executePiPatch`/`executePiSnapshot` are thin callers of it), `PiRun` separates a unit's own `tokens`/`cacheRead`/`cacheWrite` from the session's `sessionTokens`, and `piSessionWorker` + `--session-runner` hold one runner per driver session. First paid D arm (2026-09-19, `deepseek-v4-flash`, 2 units, `--slots 1`, `turns: 6`, 3 reps per bound, only `fusion.unitsPerSession` differing): fused ran one session of two units and the control two sessions of one, quality parity in all six runs (every unit accepted), median 22 533 against 22 498 tokens and 11 048 against 12 948 ms - so no token saving yet (~1.9 s per run, ~15 % of the unfused wall, which prices the session-startup term at ~1.9 s instead of leaving it assumed) and per unit the second one cost ~8 % less while the first cost more: a chain's tool surface is the union of its units' capabilities because a session's surface is fixed at creation, so a unit can spend a turn on a tool that refuses by name. Also fixed here: `specFrom` had silently dropped a spec file's `fusion` block, so a spec asking for fusion ran as the control arm. | +| F5: speculation lifecycle - one declared fact, three outcomes | landed (offline); the paid E arm ran once, no gain claimed | `SpeculationAssumption` / `ResolvedPredicate` / `SpeculationCandidate` / `isBoundedSpeculation` / `speculationOutcome` in `src/integration/ooo-execution.ts`, beside the fusion conditions: the assumption is a declaration the summary binds to (it never discovers for itself that the guess was false), the first experiment's bounds are a predicate (exactly one pending fact, nothing prepared from the guess - a speculative successor or an irreversible operation each refuse it by name), and the outcome has three states rather than two - **true** publishes, **false** discards the candidate and returns `sessionReusable: false`, which is what makes "失效会话不能复用到真实路径" a rule the caller must honour instead of a note, and **unknown** (no reading, an unattested reading, or evidence about another version) waits without publishing. Asking for the outcome of a candidate that is not the bounded shape throws rather than folding a fourth state into the three. Cases: `tests/integration/ooo-speculation.test.ts` (9), the last of which joins this half to fusion's condition 5 - an invalidated branch is not a legal predecessor for the real path. Mutants: 4 (`speculation-guesses-several-facts-at-once`, `an-unattested-reading-counts-as-evidence`, `evidence-about-another-version-is-the-same-fact`, `a-contradicted-guess-keeps-its-session`); the target's sweep is 22 of 22 caught. The E arm's instrument now exists (`evals/ooo-execution/speculation-pilot.ts`, registered) and ran once (2026-09-19, 6 paid units, ~43 k tokens): it decides the guessed fact, prepares the candidate, applies `speculationOutcome`, and verifies a published candidate with the unit's own frozen check. **No result is claimed**: the quality term was false in all four verified candidates, so by the design's own rule the latency and cost shape may not be reported as a gain. The search behind those failures is now closed and its first reading was wrong: `artifactEnvelope` builds two legitimate shapes (a patch, and a conclusion with `kind, conclusion, summary, evidence, citations`), and the instrument had fed every artifact to the patch reader. Eight of nine attempts answered with a conclusion, which this unit's check cannot pass and the board would refuse; the one patch attempt failed on a real mistake (`rows` for `lines`). The instrument now reads by kind, keeps every artifact, candidate tree and check output, and the run is archived. Measured outcome of the arm at this shape: the post-fact cost drops from ~6.2 s of work to 175 ms of verification when the fact holds, the false-fact case wastes 20 332 tokens, and the prepared candidate was publishable in 0 of 3 holding reps - so the cost is real, the gain is not, and the binding constraint is the candidate's admissibility | ### What F2b measured: a run can hold exactly one claim @@ -349,7 +361,7 @@ guard but a slice of the integration still outstanding: board verbs is the shared-store integration the design's 当前实现 section already names as the target shape that is not in place, not a check that can be added to one method. - The harmful direction is closed without it, which is why B6 can still carry teeth: a generic claim - cannot take a live claim (`a-live-claim-can-be-taken-by-another-agent`), acceptance is the board's + cannot take a live claim (`a-live-claim-can-be-taken-by-another-agent`, retired 2026-09-24), acceptance is the board's verdict bound to _this_ attempt's artifact digest (`verdict-lookup-not-bound-to-the-artifact`), a withdrawn acceptance withdraws the release of the dependent (`round-releases-dependents-on-delivered-bytes`), and an entry retired while verification was @@ -395,10 +407,10 @@ enters through an existing operation, with a stated contract and execution autho | ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | G1 | An ordinary board handoff reaches semantics, execution and acceptance with no `ooo_round` | proven | `BoardAdmission.next()` now answers with `nextTask(dispatchTasks(compileTaskUnits({plan, specs}).units, facts))`, so the coordination path consults the compiler it used to ignore (`src/integration/ooo-board.ts:812` and `:822`, and the hand-rolled candidate mapping is gone from `src/`: `rg -c schedulable src/` is 0 (the config's own anchor for that rule is what used to match, and it moved with the rule)). `tests/integration/ooo-ordinary-handoff.test.ts` walks claim -> deliver -> independent verdict -> accepted with the board's own verbs; the only mention of `ooo_round` in it is the comment saying it is untouched. Verified by mutation by hand: dropping `task.accepted` from `nextTask`'s validity check fails it | | G2 | A real artifact is accepted on that path, and accepting it releases the dependent | proven | the same test: B's artifact is accepted by the board's own verdict, and `deriveStatus(units, facts).ready` goes from `["B"]` to `["A"]` over the recorded facts while `next()` gives the same two answers, so the two views are held together rather than assumed to agree. The same mutation (validity no longer requiring acceptance) fails the test at its first assertion: `assert.equal(gate.next(), "B")` at tests/integration/ooo-ordinary-handoff.test.ts:68 returns null (AssertionError: null !== 'B'), so the `ready` assertion later in the test is never evaluated. Measured, not recalled: the mutation was run and restored byte-identically while correcting this sentence. The rule this row's release depends on - an artifact delivered without an accepted verdict is not selectable - moved with the wiring, and its tooth moved with it: `selection-ignores-a-withdrawn-acceptance` is now anchored on `nextTask`'s `selectable` line in `src/integration/ooo-execution.ts` and is a configured target (full harness: 7 of 7 targets, 31 of 31 caught, 7 of 7 restored byte-identically, exit 0). | -| G3 | Failure, cancellation, the ordered fallback, and the accepted prefix | proven | `tests/integration/ooo-ordinary-failure.test.ts` (5/5): a refused deliverable neither accepts nor releases and is not re-selected on its own until the coordinator reopens it; a later refusal does not withdraw the prefix that was already accepted; a cancellation names the lease it revoked and every claim behind it is refused; with no fusion point the plan falls back to its declared order; a split that maps no parent obligation is refused by name with its location. Selection is tooth-backed (`selection-ignores-a-withdrawn-acceptance`, `a-live-claim-does-not-block-selection`, `cancellation-does-not-stop-a-claim`) and so is the prefix: mutant `reopening-one-task-clears-every-acceptance` (reopen fences every accepted task instead of the dependents of the one reopened; caught by the fifth case, which then finds only A accepted where it requires A and B). The design's fusion clause ("融合中途失败只保留已接受前缀") is **not applicable on this path**: fusion is declared unmodelled and refused by name (`UNMODELLED_ACTIONS`), so nothing on the ordinary path fuses and nothing can fail midway; the prefix property itself is what the fifth case asserts. The offline model covers fused-midway failure for its own enumerated plans under row A3 | -| G4 | The old entry's query and cancel are reachable from existing facilities | proven | query: `tests/integration/ooo-round-query.test.ts` "the query port reads a round without migrating, publishing or exposing a write", tooth-backed by D4's `the-status-read-path-opens-the-rounds-store`. cancel: the product's own lifecycle operation, tooth-backed at integration level by the durability case in the same file ("the terminal decision outlives the host that made it, and still refuses new work": a second host that opens the store reads the same terminal reason, a new claim is refused, the revoked claim is gone and the attempt is fenced). Mutant `cancelling-a-round-forgets-its-reason` (cancel stops persisting the reason; caught by that case - the cross-process suite that used to name it too (`evals/ooo-execution/cancellation.test.ts`) was retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)), which leaves the integration case and the mutant as this row's evidence). Neither file mentions `ooo_round` or the extension: 0 hits | +| G3 | Failure, cancellation, the ordered fallback, and the accepted prefix | proven | `tests/integration/ooo-ordinary-failure.test.ts` (5/5): a refused deliverable neither accepts nor releases and is not re-selected on its own until the coordinator reopens it; a later refusal does not withdraw the prefix that was already accepted; a cancellation names the lease it revoked and every claim behind it is refused; with no fusion point the plan falls back to its declared order; a split that maps no parent obligation is refused by name with its location. Selection is tooth-backed (`selection-ignores-a-withdrawn-acceptance`, `a-live-claim-does-not-block-selection`, `cancellation-does-not-stop-a-claim`) and so is the prefix: the tooth `reopening-one-task-clears-every-acceptance` was retired on 2026-09-24 (reopen fences every accepted task instead of the dependents of the one reopened; the fifth case states the rule, which then finds only A accepted where it requires A and B). The design's fusion clause ("融合中途失败只保留已接受前缀") is **not applicable on this path**: fusion is declared unmodelled and refused by name (`UNMODELLED_ACTIONS`), so nothing on the ordinary path fuses and nothing can fail midway; the prefix property itself is what the fifth case asserts. The offline model covers fused-midway failure for its own enumerated plans under row A3 | +| G4 | The old entry's query and cancel are reachable from existing facilities | proven | query: `tests/integration/ooo-round-query.test.ts` "the query port reads a round without migrating, publishing or exposing a write", tooth-backed by D4's `the-status-read-path-opens-the-rounds-store`. cancel: the product's own lifecycle operation, tooth-backed at integration level by the durability case in the same file ("the terminal decision outlives the host that made it, and still refuses new work": a second host that opens the store reads the same terminal reason, a new claim is refused, the revoked claim is gone and the attempt is fenced). The tooth `cancelling-a-round-forgets-its-reason` was retired on 2026-09-24 (cancel stops persisting the reason; the case above states it - the cross-process suite that used to name it too (`evals/ooo-execution/cancellation.test.ts`) was retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)), which leaves the integration case and the mutant as this row's evidence). Neither file mentions `ooo_round` or the extension: 0 hits | | G5 | `ooo_round` leaves the product tool directory, with the tool directory, adapter, docs and hidden-features registry updated together | proven | Removed: `.pi/extensions/nmg/ooo-round.ts`, its import and its `registerTool` block in `.pi/extensions/nmg/index.ts` (0 remaining mentions of `ooo_round` in the tool directory), and `tests/extensions/nmg/ooo-round.test.ts`; the three tool lists in `tests/extensions/nmg/index.test.ts` no longer name it. What remains is the ordinary path: `nmg_board` for handoffs, the shared layer for selection. `ooo-execution.ts` stays in the extension because it is not the tool - `evals/ooo-execution/live-{cycle,patch,pi}.ts` import it as the eval-side Pi execution adapter, and G7 runs through them. `npm run ooo:round` also stays: it is the research CLI under `evals/`, not a product surface. Docs record the removal in the bootstrap design's S3 row and in the design's own `ooo_round` sentence. The hidden-features registry is deliberately unchanged: its OoO rows describe the research CLI and the live eval entries, whose gates did not change, and the Pi tool was registered by default, so it never had a hidden-feature row to remove. | -| G6 | The task view is a projection of existing facts: no second editable task truth, and a model cannot confirm `accepted`, change a lease or overwrite a cancellation by patching state | proven | the state side is covered (B5, B6, D3, D4). The consequences the design's information table has are now asserted rather than described: the frozen plan has one owner and a second, different plan is refused (`tests/integration/ooo-task-tables.test.ts`, mutant `a-second-plan-silently-adopts-the-run`), a terminal decision cannot be overwritten by a later patch (mutant `a-second-cancellation-overwrites-the-first-decision`), and acceptance is only ever the host-verified artifact plus the board's verdict about that digest (`verdict-lookup-not-bound-to-the-artifact`, `round-releases-dependents-on-delivered-bytes`). What stays prose is the table's owner assignment itself - which layer owns which information class is a design statement, not something a test can read off one column | +| G6 | The task view is a projection of existing facts: no second editable task truth, and a model cannot confirm `accepted`, change a lease or overwrite a cancellation by patching state | proven | the state side is covered (B5, B6, D3, D4). The consequences the design's information table has are now asserted rather than described: the frozen plan has one owner and a second, different plan is refused (`tests/integration/ooo-task-tables.test.ts`, mutant `a-second-plan-silently-adopts-the-run`), a terminal decision cannot be overwritten by a later patch (the tooth `a-second-cancellation-overwrites-the-first-decision` was retired on 2026-09-24, because `a cancellation names the lease it revokes, and the batch behind it is refused` states it), and acceptance is only ever the host-verified artifact plus the board's verdict about that digest (`verdict-lookup-not-bound-to-the-artifact`, `round-releases-dependents-on-delivered-bytes`). What stays prose is the table's owner assignment itself - which layer owns which information class is a design statement, not something a test can read off one column | | G7 | One real handoff at a legal boundary, with another Agent session continuing the same parent task from the view plus retrievable evidence, judged by the fixed parent check | proven | `docs/experiments/execution/ooo-real-continuation-2026-09-14.md`, with the run's raw records beside it (`ooo-real-continuation-2026-09-14-g7-run.jsonl`, `-merge-retries.jsonl`) and the executable check `evals/ooo-execution/live-continuation.ts`. Five structurally identical, content-distinct tasks; every role a separate process; the parent is the channel and the continuation a second handoff in it; part 2 retrieves the part-1 artifact from the view and verifies it against the digest the store recorded before it may continue; the fixed parent check (6/8/7/6/7 frozen cases) decides the boundary and the parent. Model `deepseek/deepseek-v4-flash`; one clean pass over the five tasks = 9 model calls, 80 081 tokens, 84 395 ms, plus a four-attempt sample of `merge`. Result: four of five tasks accepted end to end; `merge`'s continuation failed 2 of 4 attempts with an unparseable patch artifact, and its parent stage then recorded **no verdict** rather than judging the part-1 artifact. The document states the cost and the limits, including that failed attempts' token counts are not recorded A follow-up comparison (`docs/experiments/execution/ooo-real-continuation-comparison-2026-09-14.md`) re-reads every stage record, classifies each failure by cause, and finds that 13 of the 20 continuation failures were the harness's own defects rather than the model's, while the stable first stage failed once in 35 samples. | Two consequences for the nodes above it: diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 31235f62..c22b6354 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -149,22 +149,6 @@ const TARGETS: readonly Target[] = [ "tests/integration/ooo-managed-write.test.ts", ], mutants: [ - { - // The handle, not a PRAGMA, is what makes this read-only: restoring writability must fail the test. - name: "the-read-only-factory-opens-a-writable-handle", - ast: { call: "DatabaseSync", argCount: 2 }, - to: "new DatabaseSync(databasePath)", - expect: "a read-only open neither creates, migrates nor writes", - }, - { - // The store owns the boundary: a write reached inside a transition without its port must be - // refused rather than become a second BEGIN. - name: "nested-write-transaction-is-allowed", - ast: { within: "writeTransaction" }, - from: ' if (this.openTransaction)\n throw new Error("a write transaction is already open: join it with the port it issued");', - to: ' if (this.openTransaction && false)\n throw new Error("a write transaction is already open: join it with the port it issued");', - expect: "a write entry reached inside a transition without a port is refused, not nested", - }, { // A failure the caller swallows still forbids the commit: nothing may be written up to the // failure and then kept by a normal return value. @@ -174,25 +158,7 @@ const TARGETS: readonly Target[] = [ to: " void state.rollbackOnly;", expect: "a failure the caller swallows still forbids the commit", }, - { - // The mechanical invariant: a hand-rolled BEGIN anywhere in the store makes a second - // boundary possible, and behaviour tests would not notice a path that still works. - name: "a-method-opens-its-own-transaction", - ast: { within: "removeMemoryFromChain" }, - from: " return this.writeTransaction(() => {", - to: ' this.db.exec("BEGIN IMMEDIATE");\n return this.writeTransaction(() => {', - expect: "the store runs its transaction boundary in exactly one place", - }, - { - // The claim CAS is the whole of the fence: a reader who cannot take a live claim must - // lose it, and the condition that says so is the only thing between two readers and the - // same work. - name: "a-live-claim-can-be-taken-by-another-agent", - ast: { within: "claimTaskBoardEntry" }, - from: " AND (\n (claimed_by IS NULL OR claim_expires_at IS NULL OR claim_expires_at <= ?)\n OR claimed_by = ?\n )`,", - to: " AND (\n (claimed_by IS NULL OR claim_expires_at IS NULL OR claim_expires_at <= ?)\n OR ? IS NOT NULL\n )`,", - expect: "one reader of the same ready task is given the claim, the second is refused", - }, + { name: "stale-claim-may-deliver-again", ast: { within: "claimTaskBoardEntry" }, @@ -221,24 +187,7 @@ const TARGETS: readonly Target[] = [ to: ' "DELETE FROM task_board_retentions WHERE 0",', expect: "a bounded pin stops pinning when its bound passes", }, - { - // A run's plan is what every later decision is read against, so a second registration - // must not be able to replace it. - name: "a-second-plan-overwrites-the-frozen-one", - ast: { within: "insertTaskRunManifest" }, - from: " if (\n String(existing.plan_digest) !== input.planDigest ||\n String(existing.policy) !== input.policy\n )", - to: " if (\n false &&\n String(existing.plan_digest) !== input.planDigest &&\n String(existing.policy) !== input.policy\n )", - expect: "a run registers once, and a second plan for the same run is refused", - }, - { - // The fact write has to join the transition the caller opened, not open a second one; the - // board write and the run fact of one transition stand or fall together. - name: "the-run-fact-opens-its-own-transaction", - ast: { within: "appendTaskRunFact" }, - from: " return port\n ? this.withPort(port, () => this.insertTaskRunFact(input))\n : this.writeTransaction(() => this.insertTaskRunFact(input));", - to: " void port;\n return this.writeTransaction(() => this.insertTaskRunFact(input));", - expect: "a board write and a run fact land together, and neither lands alone", - }, + { // One managed entry belongs to one run: answering with the first binding would hand a // second run's facts to whoever asked. @@ -283,15 +232,7 @@ const TARGETS: readonly Target[] = [ to: "const db = new BoardAdmission(databasePath) as unknown as DatabaseSync;", expect: "the query port reads a round without migrating, publishing or exposing a write", }, - { - // The composed write must join the transition it is called in: a publication that opens its - // own boundary commits even when the transition around it fails. - name: "round-publication-opens-its-own-transaction", - ast: { within: "publish" }, - from: " },\n port,\n ).id;", - to: " },\n ).id;", - expect: "the round's own publication rolls back with the transition that made it", - }, + { // The cache exists to be recomputable. A rebuild that returns without writing is the // difference between "the sources decide" and "the schema says so". @@ -331,13 +272,7 @@ const TARGETS: readonly Target[] = [ expect: "acceptance survives the entry's own TTL, because the round retains what it references", }, - { - name: "round-never-releases-its-pin", - ast: { within: "fenceRow" }, - from: " // The artifact is being cleared, so the round no longer relies on this entry's\n // verdict: the pins go with the value they protected.\n this.releaseRowRetention(row);", - to: " // The artifact is being cleared, so the round no longer relies on this entry's\n // verdict: the pins go with the value they protected.", - expect: "cancelling a round releases the pins it held, so nothing it referenced leaks", - }, + { // The borrowed view is one implementation serving two paths. Letting the offline port // answer from its own rule is exactly the divergence this target exists to catch. @@ -347,24 +282,7 @@ const TARGETS: readonly Target[] = [ to: " accepted: () => ({}),", expect: "the owner's view and the offline reader report the same facts", }, - { - // Reopening withdraws what was built from the value that no longer exists. Clearing every - // accepted task instead takes back work the host already accepted. - name: "reopening-one-task-clears-every-acceptance", - ast: { within: "reopen" }, - from: " const affected = new Set([id]);", - to: " const affected = new Set(rows.map((row) => row.id));", - expect: "a later refusal does not withdraw the prefix that was already accepted", - }, - { - // The first decision is the one that took effect. A second cancellation that rewrites it - // is a state patch overwriting a terminal fact. - name: "a-second-cancellation-overwrites-the-first-decision", - ast: { within: "cancel" }, - from: " const already = this.cancelled();\n if (already !== null) return [];", - to: " const already = this.cancelled();\n if (false && already !== null) return [];", - expect: "a cancellation names the lease it revokes, and the batch behind it is refused", - }, + { // The frozen plan is the run's input, and installing a second one over it is the second // editable task truth the design forbids. The constructor refuses by policy digest. @@ -382,15 +300,7 @@ const TARGETS: readonly Target[] = [ to: ' if (false && !this.live(row)) return "stale";', expect: "a claim the board retires inside the verification window cannot be committed", }, - { - // The decision is a fact in the store. Keeping it only in the process that made it is - // what would let a restart resume a stopped round. - name: "cancelling-a-round-forgets-its-reason", - ast: { within: "cancel" }, - from: ' .prepare("UPDATE ooo_probe_runs SET cancel_reason=?, cancelled_at=? WHERE run_id=?")\n .run(reason.slice(0, 1_000), new Date(this.now).toISOString(), this.runId);', - to: ' .prepare("UPDATE ooo_probe_runs SET cancel_reason=NULL, cancelled_at=? WHERE run_id=?")\n .run(reason.slice(0, 1_000), new Date(this.now).toISOString(), this.runId);', - expect: "the terminal decision outlives the host that made it, and still refuses new work", - }, + { // Selection and ranking must not be able to disagree with each other. `next()` is the head of // `candidates()`, so a caller that starts one unit and a caller that starts several read the @@ -531,17 +441,7 @@ const TARGETS: readonly Target[] = [ to: " candidate.assumptions.length >= 1 &&", expect: "the first experiment allows one pending fact, and a second is refused by name", }, - { - // A missing reading is not permission to publish: the design's "不确定就等待" is the whole - // reason the outcome has three states instead of two. - name: "a-guess-with-no-evidence-publishes", - ast: { within: "speculationOutcome" }, - from: - ' outcome: "wait",\n sessionReusable: true,\n' + - " reason: `no evidence for ${assumption.predicateId}`,", - to: ' outcome: "publish",\n sessionReusable: true,\n reason: "assumed",', - expect: "no evidence waits: an unknown fact never becomes a silent publish", - }, + { name: "an-unattested-reading-counts-as-evidence", derive: { @@ -673,12 +573,6 @@ const TARGETS: readonly Target[] = [ target: "src/integration/task-semantics-model.ts", suites: ["tests/integration/task-semantics-model.test.ts"], mutants: [ - { - name: "ordered-mode-becomes-any-topological-order", - from: ' if (mode === "ordered") return [[...ids]];', - to: ' if (mode === "ordered" && !ids.length) return [[...ids]];', - expect: "the ordered mode is the declared plan order and nothing else", - }, { name: "fusion-rollback-not-counted", from: " rollbacks += 1;\n retries += 1;", @@ -722,15 +616,6 @@ const TARGETS: readonly Target[] = [ expect: "the board drivers run the protocol end to end through the daemon that serves the store", }, - { - // The whole boundary is that a driver reaches the board through the daemon. A convenience - // import of the store is how that boundary rots, and the structural check is what catches it - // rather than a later round discovering a second writer. - name: "a-driver-falls-back-to-opening-the-store", - from: "const state = roundDaemon(resolve(values.daemon));", - to: 'const state = roundDaemon(resolve(values.daemon));\nconst store = (await import("../../src/core/store/base.ts")).NmgStoreBase;', - expect: "a driver refuses without a daemon, and no driver opens a database of its own", - }, ], }, { @@ -781,15 +666,7 @@ const TARGETS: readonly Target[] = [ to: " const batch = legal.slice(0, 1);", expect: "a declared slot count is reached, and the claims overlap in time", }, - { - // The batch is the unit of overlap, and the overlap that matters is a unit's *check* beside - // another unit's work: awaiting each unit in turn keeps a batch's claims from ever running - // beside each other, which is the property the C arm buys. - name: "the-loop-awaits-each-unit-instead-of-the-batch", - from: ' const held = await Promise.all(\n batch.map((id) => (chainPath ? runChain(id) : dispatch(id).then((a) => "unit" in a))),\n );', - to: ' const held: boolean[] = [];\n for (const id of batch)\n held.push(await (chainPath ? runChain(id) : dispatch(id).then((a) => "unit" in a)));', - expect: "a unit's check is outstanding while an independent unit's worker runs", - }, + { // What may run is the board's answer, not the plan's order: dispatching the declared plan // instead would run a unit whose dependencies are not accepted yet. @@ -798,12 +675,7 @@ const TARGETS: readonly Target[] = [ to: " const legal = input.plan.filter((id) => !attempted.has(id));", expect: "the loop runs what the board offers, in the order the board offers it", }, - { - name: "a-unit-is-dispatched-twice-in-one-batch", - from: " const batch = legal.slice(0, input.slots);", - to: " const batch = [...legal, ...legal].slice(0, input.slots);", - expect: "a unit's check is outstanding while an independent unit's worker runs", - }, + { // The bound is what keeps a fused run from swallowing the plan. Without it one session would // run every legal successor in turn. The comparison lives in `nextSessionMove` (the shared @@ -845,33 +717,13 @@ const TARGETS: readonly Target[] = [ target: "evals/ooo-execution/plan-driver.ts", suites: ["evals/ooo-execution/plan-driver.test.ts", "evals/ooo-execution/families.test.ts"], mutants: [ - { - // The slot count is declared to the board, not only promised to the loop: a driver that asks - // the loop for one slot while telling the board a different count cannot overlap claims. - name: "the-driver-declares-one-slot-whatever-the-spec-says", - from: " board: gate,\n plan: planIds,\n slots: spec.slots,", - to: " board: gate,\n plan: planIds,\n slots: 1,", - expect: "a declared slot count is reached, and the claims overlap in time", - }, { name: "a-unit-ignores-the-checks-it-declares", from: " const checks = unit.checks ? checkList(unit.checks) : fallback;", to: " const checks = fallback;", expect: "report: both plans accept the instrument's answers, and the same composed ones", }, - { - name: "a-unit-nothing-checks-is-still-a-unit", - from: " if (!checks)\n throw new Error(\n `${id}: no checks", - to: " if (!checks && false)\n throw new Error(\n `${id}: no checks", - expect: "a unit nothing checks is refused rather than accepted on nothing", - }, - { - name: "the-parent-check-ignores-its-own-verdict", - from: " return {\n verdict: verified.verdict,", - to: ' return {\n verdict: "accept",', - expect: - "the parent check is the composed acceptance, and a failing check is reported as such", - }, + { // A submitted patch carries the unit's whole frozen view, so composing by overwriting the // candidate with each accepted submission puts the *last* unit's untouched copies of its @@ -914,12 +766,7 @@ const TARGETS: readonly Target[] = [ to: " if (false && !legalSet.has(suggestion.taskId)) {", expect: "a suggestion outside the legal set is refused, however high it scores", }, - { - name: "the-ordering-adds-a-task-to-the-set", - from: " return [...legal].sort((left, right) => {", - to: " return [...legal, ...best.keys()].sort((left, right) => {", - expect: "a suggestion outside the legal set is refused, however high it scores", - }, + { name: "an-unmodelled-action-is-scored", from: ' if (suggestion.action !== "next") {', @@ -932,12 +779,7 @@ const TARGETS: readonly Target[] = [ to: " if (false && source.enabled === false) {", expect: "a disabled or failing source falls back to the rule policy, and says why", }, - { - name: "a-failing-source-takes-the-decision-with-it", - from: " } catch (error) {", - to: " } catch (error) {\n throw error;", - expect: "a disabled or failing source falls back to the rule policy, and says why", - }, + { name: "a-score-from-another-scope-is-reused", from: " if (\n provenance.sessionId !== projection.sessionId ||\n provenance.branchId !== projection.branchId\n ) {", @@ -1045,14 +887,7 @@ const TARGETS: readonly Target[] = [ to: " if (false && dependency === task.taskId)", expect: "a freeze cannot dangle, repeat a task, or lean on itself", }, - { - // Freezing is one transition: a batch where the store refuses one task must not leave the - // earlier ones frozen, or a plan exists that no caller ever proposed. - name: "the-plan-freezes-one-task-per-transaction", - from: " return store.coordinateRunWrite(request.runId, (port) => {\n // The array order is the plan order: the position comes from here, not from the request.\n request.tasks.forEach((task, position) =>\n store.freezeTaskRunTask({ ...task, runId: request.runId, position }, port),\n );\n return { runId: request.runId, frozen: request.tasks.length };\n });", - to: " request.tasks.forEach((task, position) =>\n store.freezeTaskRunTask({ ...task, runId: request.runId, position }),\n );\n return { runId: request.runId, frozen: request.tasks.length };", - expect: "a refused freeze leaves the plan exactly as it was", - }, + { // The plan order is the array order: the position comes from that loop, so freezing every // task at zero would leave the stored plan's order to the task ids. @@ -1061,15 +896,7 @@ const TARGETS: readonly Target[] = [ to: " request.tasks.forEach((task) =>\n store.freezeTaskRunTask({ ...task, runId: request.runId, position: 0 }, port),\n );", expect: "a run registers, freezes a plan, adopts entries, and reads it all back", }, - { - // A cancelled run is closed: its plan is not extended behind the cancellation that every - // other rule in this file already honours. - name: "a-cancelled-run-takes-a-new-plan", - from: " const refusal = managedWriteRefusal(store, request.runId);\n if (refusal) throw new Error(refusal);", - to: " const refusal: string | null = null;\n if (refusal) throw new Error(refusal);", - ast: { within: "freezeRunPlan" }, - expect: "a cancelled run takes no further plan", - }, + { // The binding records which channel carries the entry, which is what lets a status reader // resolve it without searching every channel. @@ -1078,15 +905,7 @@ const TARGETS: readonly Target[] = [ to: " payload: null,", expect: "a run registers, freezes a plan, adopts entries, and reads it all back", }, - { - // Creating the entry and adopting it are one transition. Two calls would leave an unmanaged - // entry behind when the binding is refused - the hole the run fence exists to close. - name: "the-entry-is-created-before-its-binding-is-checked", - from: " return store.writeTransaction((port) => {\n const entry = store.putTaskBoardEntry(request.entry, port);", - to: " return store.writeTransaction(() => {\n const entry = store.putTaskBoardEntry(request.entry);", - expect: - "adoption is part of the transition that creates the entry, so a refusal leaves no entry", - }, + { // A run-level cancellation is the run's fact, not a task's: the schema's empty task id is // what keeps it from colliding with a task that has no name. @@ -1112,15 +931,6 @@ const TARGETS: readonly Target[] = [ expect: "status is a read: an unknown run has no manifest and is not registered by being asked", }, - { - // A status read registers and appends nothing: a view that repaired what it could not find - // would make its own answer true. - name: "status-registers-the-run-it-cannot-find", - from: " manifest: store.taskRunManifest(runId),", - to: ' manifest: (store.registerTaskRun({ runId, planDigest: "", policy: "", revision: "", retention: "" }), store.taskRunManifest(runId)),', - expect: - "status is a read: an unknown run has no manifest and is not registered by being asked", - }, ], }, { @@ -1192,20 +1002,6 @@ const TARGETS: readonly Target[] = [ }, ], }, - { - // One home for the check list: the plan decides it, including a route's decline. A caller that - // rebuilt the floor from the constant would execute checks the plan said not to. - target: "tools/agent-verify.ts", - suites: ["tests/tools/agent-verify.test.ts"], - mutants: [ - { - name: "the-caller-rebuilds-the-shared-floor", - from: "? [...narrowPlan.shared, ...(route.tests.length ? [nodeTestCheckName(route.id)] : [])]", - to: '? ["check", "docs:check", "format:check", "glossary:check", "lint", "package:check", "rtm:check", ...(route.tests.length ? [nodeTestCheckName(route.id)] : [])]', - expect: "a declining route's narrow run verifies on its own tests and nothing else", - }, - ], - }, { // Fusion legality gets its own entry for the same file: the harness reads only the first failures // of a suite run, so a second suite in the existing target pushed that target's own named failures From 337cc00cab52e8f86be8302346cf9fdedab37ff5 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 21:51:49 +0800 Subject: [PATCH 29/32] 59 teeth become operator-and-selector, and three gaps in the vocabulary close The register-wide conversion the pilot started, taken target by target: 59 more hand-written teeth are now a name, an operator and a selector, so a rename or a moved line cannot retire them. Per operator: condition-never 44 (a guard or a variable initializer), neutralize-term 9, replace-property 9, replace-argument 6, drop-statement 6, condition-holds 3. The register is 77 derived of 110. Reading after the conversion: mutants: 110 of 110 caught by the named test, all 110 by the case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 171 s; anchors: 110 of 110 resolve. The conversion does not make a tooth bite - the operator realises the same violation the hand anchor did - so the reading that matters is the second one: every tooth is still caught by the case it names. Three sites the operators could not name, each closed with a case in tests/tools/mutation-anchor.test.ts (16 cases to 19) rather than argued about: - A class constructor is a member. BoardAdmission's constructor refuses a second plan while it opens the store, and uniqueMember knew only methods and function declarations, so that site had no `within` to write. The name is `constructor`; matching it is what let a-second-plan-silently-adopts-the-run convert. - A condition written across lines is one condition. The candidate filter compared raw text while the whole-condition check compared whitespace-normalized text, so a fragment of a wrapped `if` found no candidate and then matched nothing. Both normalize now, and the filter no longer demands the fragment be unique inside the candidate: `b` appears three times in `a && (b || !b)` and that is still the condition a selector naming `b` means. - A guard clause is a statement. `if (...) throw ...;` was not among the statements drop-statement would remove, because an `if` is an IfStatement rather than an expression, a declaration, a `return` or a `throw`. It is now, and statements resolve to the innermost one containing the fragment - the rule conditions already followed, and what keeps a fragment from matching both a guard and the statement inside it. One wrong `within` was shipped in this pass and the sweep is what caught it, which is the argument for sweeping after a conversion instead of trusting the anchors pass: the-completion-ignores-a-cancelled-unit was aimed at checkDispatch rather than checkCompletion (the fragment occurs in both members), so the anchors pass resolved, the mutant was still caught - and it was caught by the suite, not by the case that names a completion of a cancelled unit. Re-aimed, that case fails again. What is left hand-written is 33 teeth, and they cluster by the operator they would need: a comparison rewritten (3), a fragment inside a template or SQL string (5), an iterable emptied or a filter dropped (4), a call or `new` replaced or unwrapped (6), an index moved (2), an initializer replaced by a different expression (3), a condition negated (2), a statement rewritten into another statement (3), two statements sharing one line (2), a literal swapped (1), a site at module top level where there is no member to name (1), and one tooth that changes two things at once. Four clusters look worth an operator (replace-comparison, replace-fragment, empty-iterable, replace-callee); the record says so and leaves the decision open rather than adding operators nobody has asked for yet. Docs and code in one commit: the record's item 2 carries the numbers, the three gaps and the wrong `within`; its remaining list is the 33-tooth classification; the obligations ledger says 77 of 110 teeth are derived, names the four candidate operators, and gains this pass's whole-register reading beside the one it already had. The tools' own headers say a constructor's name and which statements drop-statement covers. Readings at this revision: whole register -> 110 of 110 caught, all by name, 22 of 22 restored, 171 s; mutation:anchors -> 110 of 110 resolve, over 22 targets; mutation-anchor cases -> 19 pass, 0 fail; test:product -> 1565 pass, 0 fail, exit 0; verify:static -> exit 0; lint -> 0 findings. --- ...-09-24-mutants-are-derived-not-anchored.md | 36 ++ ...-mutants-are-derived-not-anchored.zh-CN.md | 8 + .../design/task-unit-semantics-obligations.md | 7 +- tests/tools/mutation-anchor.test.ts | 47 ++ tools/mutation-anchor.ts | 54 ++- tools/mutation-teeth.ts | 441 ++++++++++++------ 6 files changed, 441 insertions(+), 152 deletions(-) diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index 9e062e9a..8d00756d 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -129,6 +129,42 @@ targets`. The target `tools/agent-verify.ts` went with its single tooth, so the strings rather than by a filesystem. **Landed:** the resolver is `tools/mutation-anchor.ts` (it held no state, so it moved out of the sweep script and the tests can call it), with 16 cases over source strings and no filesystem access. + **Extended to the whole register (2026-09-24): 59 more teeth converted, and the register is now 77 + derived of 110** - every class of site the six operators can express, taken target by target and + swept after each: `condition-never` on a guard or an initializer (44 in total), `neutralize-term` (9), + `replace-property` (9), `replace-argument` (6), `drop-statement` (6), `condition-holds` (3). The + reading after it: `110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, + 22 of 22 restored byte-identically, 171 s. + Three vocabulary gaps showed up as _sites the operators could not name_, and each was closed with a + case in `tests/tools/mutation-anchor.test.ts` (now 19 cases) rather than argued about: + - **A class constructor is a member.** `BoardAdmission`'s constructor refuses a second plan while it + opens the store, and `uniqueMember` only knew method and function declarations, so that site had no + `within` at all. Its name is `constructor`; matching it is what let + `a-second-plan-silently-adopts-the-run` convert. + - **A condition written across lines is one condition.** The candidate filter compared raw text while + the whole-condition check compared whitespace-normalized text, so a fragment of a wrapped `if` found + no candidate and then matched nothing. Both now normalize, and a filter does not require the + fragment to be unique _inside_ the candidate - `b` appears three times in `a && (b || !b)`, and that + is still the condition a selector naming `b` means. + - **A guard clause is a statement.** `if (...) throw ...;` was not among the statements + `drop-statement` would remove, because an `if` is an `IfStatement` rather than an expression, + declaration, `return` or `throw`. It is now, and statements resolve to the **innermost** one + containing the fragment, the rule conditions already followed - which is also what keeps a fragment + from matching both a guard and the statement inside it. + One wrong `within` was shipped and the sweep caught it, which is the argument for sweeping after a + conversion rather than trusting the anchors pass: `the-completion-ignores-a-cancelled-unit` was aimed + at `checkDispatch` instead of `checkCompletion` (the fragment occurs in both members), so the anchors + pass resolved, the mutant was still caught - and it was caught by the _suite_, not by the case that + names a completion of a cancelled unit. Re-aimed, that case fails again. + **What is left hand-written is 33 teeth, and they cluster by the operator they would need:** a + comparison rewritten (`=== "none"` to `!== "always"`, 3), a fragment inside a template or a SQL string + (5), an iterable emptied or a filter dropped (4), a call or `new` replaced or unwrapped (6), an index + moved (`[0]` to `[1]`, 2), an initializer replaced by a different expression (3), a condition negated + (2), a statement rewritten into another statement (3), two statements sharing one line (2), a literal + swapped (1), a site at module top level, where there is no member to name (1), and one tooth that + changes two things at once (`every-task-is-frozen-at-position-zero`, kept deliberately). Four of these + clusters look worth an operator (`replace-comparison`, `replace-fragment`, `empty-iterable`, + `replace-callee`); the rest are single sites or shapes a general operator would make ambiguous. 2. The pilot target converted and swept: same names, same cases, same catches; then a refactor inside it that demonstrates a derived tooth surviving what a byte anchor did not. **Landed**, with one correction: the demo first refactor was a rename plus a condition lifted into `const closed = spent || diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index f2e44b8f..5121478b 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -29,6 +29,14 @@ **一颗牙的耐久部分是它的名字、它的算子、它的选择器,以及必须失败的那个用例。位置与替换字节每次从语法树算出来。** 1. **六个作用在具名选择器上的算子**,作用域限定在一个成员内:`condition-never`(选择器指到的那条条件永不成立)、`condition-holds`(它始终成立——镜像的那个,因为写成 `return a && b` 的规则说的是“这条条件为真”,把它换成 `false` 是把规则反过来,而不是去掉)、`neutralize-term`(该条件里的一项变成它的恒等元——`&&` 下为 `true`,`||` 下为 `false`)、`replace-argument`(某个调用的第 _n_ 个实参换成声明的片段)、`replace-property`(具名对象属性的值换成声明的片段)、`drop-statement`(选择器指到的那条语句连行带缩进被删掉)。两条规则让选择器保持诚实:匹配到多于一处、或者片段同时命中两个候选的,一律拒绝而不是猜;两个“整条件”算子拒绝只命名条件一部分的片段,因为整段替换会把 mutant 悄悄放大成“这条规则的每一条理由”——一项有它自己的算子。规则用不上这些时,仍可用“成员作用域 + 最小片段”的旧形式;而**“最小”是要点**:不要消息文本、不要同级实参、一项能表达的不要写成整条语句。 + **已推广到整个登记处(2026-09-24):又转了 59 颗,登记处现在是 110 颗里 77 颗 derived。**六个算子能表达的每一类点位都按目标逐个转完,每转完一个目标就 sweep:`condition-never`(守卫或初始化器,共 44)、`neutralize-term`(9)、`replace-property`(9)、`replace-argument`(6)、`drop-statement`(6)、`condition-holds`(3)。转完的读数:`110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐字节还原,171 秒。 + 过程中暴露了三个"算子点不到这个点位"的缺口,每个都用 `tests/tools/mutation-anchor.test.ts` 里的一条用例补上(现在 19 条),而不是靠讨论: + - **类构造函数也是成员。**`BoardAdmission` 的构造函数在打开 store 时拒绝第二个 plan,而 `uniqueMember` 只认方法声明和函数声明,那个点位当时连 `within` 都写不出来。它的名字是 `constructor`;认它之后 `a-second-plan-silently-adopts-the-run` 才转得动。 + - **跨行写的条件仍然是一个条件。**候选过滤器比的是原始文本,而整条件检查比的是空白归一化后的文本,于是被折行 `if` 的片段先找不到候选、再匹配不上。现在两边都归一化;并且过滤器不要求片段在候选内部唯一——`b` 在 `a && (b || !b)` 里出现三次,而它仍然是指名 `b` 的选择器所指的那个条件。 + - **守卫语句也是一种语句。**`if (...) throw ...;` 原本不在 `drop-statement` 会删的语句里,因为 `if` 是 `IfStatement` 而非表达式、声明、`return` 或 `throw`。现在它算,并且语句解析取**最内层**那个包含片段的语句——条件早就是这个规则,这也正是不让片段同时命中守卫和守卫里那条语句的原因。 + 有一处 `within` 我写错了,是 sweep 抓出来的,这也正是"转完必须 sweep、不能只信 anchors pass"的理由:`the-completion-ignores-a-cancelled-unit` 被指向 `checkDispatch` 而不是 `checkCompletion`(那个片段在两个成员里都出现),于是 anchors pass 照样通过、mutant 照样被抓——但抓它的是**整个 suite**,不是那条"报告一个被取消单元的完成"的用例。重新瞄准后,那条用例又失败了。 + **剩下 33 颗仍是手写,且按"需要什么算子"成簇:**比较运算符被改写(`=== "none"` 变 `!== "always"`,3)、模板或 SQL 字符串里的片段(5)、可迭代对象被清空或过滤被删(4)、调用或 `new` 被替换/脱壳(6)、下标移动(`[0]`→`[1]`,2)、初始化器被换成另一个表达式(3)、条件被取反(2)、语句被改写成另一条语句(3)、两条语句挤在一行(2)、字面量被替换(1)、点位在模块顶层因而没有成员可命名(1),以及一颗一次改两处的牙(`every-task-is-frozen-at-position-zero`,故意留着)。这四簇看起来值得加算子(`replace-comparison`、`replace-fragment`、`empty-iterable`、`replace-callee`);其余是单点位,或者用通用算子表达反而会变得有歧义的形状。 + 2. **把 133 颗 G/V/R 牙逐目标改成派生形式**,从 `src/integration/ooo-execution.ts` 开始试点:该文件在两条目标条目里共 24 颗(16 颗融合合法性推演、8 颗会话融合);同样的名字、同样的 `expect` 用例,转换后同样 24/24 被抓——其中 19 颗派生、5 颗保留手写(下面具名)——然后在该文件里做一次真重构,证明派生出来的位置都还在,而字节锚点会在这次重构里退役。 试点文件里那 5 颗残余,每一颗都是单个表达式或单个值,而不是整条语句或消息:`selection-ignores-a-withdrawn-acceptance`(整个 `const` 初始化式被替换)、`the-budget-is-not-cut-from-the-startable-set`(被返回表达式里的调用被删)、`next-task-is-not-the-head-of-the-legal-set`(下标 `[0]` 变 `[1]`)、`speculation-guesses-several-facts-at-once`(`length === 1` 变 `>= 1`)、`a-guess-with-no-evidence-publishes`(插入一个分支并重写消息)。 3. **派生形式表达不了的,保留手写,并且说明为什么。** 16 颗 I/E 牙——插入、包装、第二份规则——它们的锚点是“某段代码不该出现的地方”,所以它们最后才搬,也最先被重新考虑:凡是规则已经有一个关系式或穷举式检查的,就退役这颗牙,并在它原来的位置上点名那个检查。 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index 97325fb4..cd9c75df 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -24,6 +24,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth -- --anchors-only` -> `anchors: 144 of 144 resolve, over 23 targets`, exit 0 in 0.7 s (2026-09-24, after five teeth were retired in favour of the checks that already state their rules) - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts` -> `mutants: 15 of 15 caught by the named test`, restored byte-identically (2026-09-24, after six more teeth were retired there: the five budget/publication rules and the run-scoping claim) - `npm run mutation:teeth` (the whole register) -> `mutants: 136 of 136 caught by the named test`, 23 of 23 targets restored byte-identically, exit 0 in 217 s (2026-09-24), **every one of them caught by the case its `expect` names**. The first whole-register reading (231 s, 134 of 136 by name) is what surfaced three wrong links, all repaired in the same pass: `the-caller-rebuilds-the-shared-floor` (its `expect` was a paraphrase of no case, re-pointed), `the-grace-is-zero` (its named case derived its fixture from the constant under test, so it passed while the bug was live - the fixture is a literal now), and `next-is-not-the-head-of-the-ordered-candidates` (re-pointed to `at the default budget the licence is still the head of the ordered set`, which now asserts the head). A fourth was found by the pilot sweep and repaired here: `fusion-continues-from-an-unverified-answer`'s case refused the pair for a second reason as well, so it passed under the mutant - with no dependency between the two units it is refused by that one condition only. One row still carries a note rather than an assertion: `the-pass-asks-a-unit-it-already-failed-again` is caught by a named case that does not finish inside its 30 s bound, which this tool counts as caught and prints with that reason. +- `npm run mutation:teeth` (the whole register, after 59 teeth were converted to derived selectors) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 171 s (2026-09-24). The conversion is not what makes them bite; it is what keeps them aimed: one wrong `within` was shipped in this pass and the sweep is what caught it, as a tooth caught by the suite rather than by the case that names its rule. - `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were **orphans**: no row named `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` or `the-parent-check-ignores-its-own-verdict`, so their rules have a case but no row; the check each case makes is in the [record](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) - `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) @@ -37,8 +38,10 @@ description rather than a pin. Every mutant name below was read from A tooth is a **name, an operator and a selector** - not a copy of a line: the site and the replacement bytes are resolved from the syntax tree on every run, so a rename or a reflow does not retire the pin -while the rule still stands ([the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md), -landed so far on the pilot file). `npm run mutation:teeth -- --anchors-only` answers whether every +while the rule still stands ([the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). +77 of the register's 110 teeth are in that form now; the 33 that are not are the shapes the six operators +cannot express, clustered and counted in the same record, with the four clusters that look worth an +operator named there. `npm run mutation:teeth -- --anchors-only` answers whether every anchor still applies - under a second, no suite run, nothing written - because a tooth that stopped matching its rule is otherwise only visible as a check that stopped counting. diff --git a/tests/tools/mutation-anchor.test.ts b/tests/tools/mutation-anchor.test.ts index eb392c91..d1516b0d 100644 --- a/tests/tools/mutation-anchor.test.ts +++ b/tests/tools/mutation-anchor.test.ts @@ -217,6 +217,53 @@ test("a member the selector cannot tell apart is refused: none, several, or a fr ); }); +test("a class constructor is a member, and its name is `constructor`", () => { + const classSource = `class Gate {\n constructor(runs, wanted) {\n const recorded = runs.find((run) => run.id === wanted.id);\n if (recorded && recorded.policy !== wanted.policy) refuse("policy changed");\n }\n}\n`; + const inConstructor = named("x", { + within: "constructor", + operator: "condition-never", + condition: "recorded && recorded.policy !== wanted.policy", + }); + assert.equal( + selected(classSource, inConstructor), + "recorded && recorded.policy !== wanted.policy", + ); + assert.match( + refusal( + classSource, + named("x", { within: "Gate", operator: "condition-never", condition: "recorded" }), + ), + /member Gate matched 0 members/, + ); +}); + +test("a condition written across lines is one condition, and a reflowed fragment still names it", () => { + const wrapped = `function refuseSuggestion(suggestion, projection) {\n const provenance = suggestion.provenance;\n if (\n provenance.sessionId !== projection.sessionId ||\n provenance.branchId !== projection.branchId\n ) {\n return "other scope";\n }\n return null;\n}\n`; + const scope = named("x", { + within: "refuseSuggestion", + operator: "condition-never", + condition: + "provenance.sessionId !== projection.sessionId || provenance.branchId !== projection.branchId", + }); + assert.match( + selected(wrapped, scope), + /^provenance\.sessionId !== projection\.sessionId \|\|\n\s+provenance\.branchId !== projection\.branchId$/u, + ); +}); + +test("a guard clause is a statement that can be dropped, listed among the statements that can", () => { + const guarded = `function claim(id, agentId, cancelled) {\n if (!id || !agentId) throw new Error("task and agent required");\n if (cancelled !== null) throw new Error("round cancelled");\n return id;\n}\n`; + const guard = named("x", { + within: "claim", + operator: "drop-statement", + statement: 'if (cancelled !== null) throw new Error("round cancelled");', + }); + assert.equal( + selected(guarded, guard), + ' if (cancelled !== null) throw new Error("round cancelled");\n', + ); +}); + test("when a fragment sits inside two decision positions, the innermost one is the site", () => { const nested = named("x", { within: "selection", diff --git a/tools/mutation-anchor.ts b/tools/mutation-anchor.ts index ada74812..d907a8a6 100644 --- a/tools/mutation-anchor.ts +++ b/tools/mutation-anchor.ts @@ -51,7 +51,8 @@ export interface Derive { | "replace-argument" /** The value of the object property named `property` becomes the declared `to` fragment. */ | "replace-property" - /** The statement the selector identifies is removed, with its line and its indentation. */ + /** The statement the selector identifies is removed, with its line and its indentation: an + * expression, a declaration, a `return`, a `throw`, or a guard clause. */ | "drop-statement"; /** Which guard: a fragment its condition's own text contains, e.g. `existing.deliveredBy`. */ readonly condition?: string; @@ -81,6 +82,14 @@ export interface Site { readonly retaken: boolean; } +/** Does this text contain the fragment, ignoring the line breaks the formatter chose? A filter that + * asks "which candidate mentions this" must not require the fragment to be unique inside the + * candidate: `b` appears three times in `a && (b || !b)`, and that condition is still the one a + * selector naming `b` is talking about. */ +function containsText(haystack: string, fragment: string): boolean { + return haystack.replace(/\s+/gu, " ").includes(fragment.replace(/\s+/gu, " ").trim()); +} + /** Whitespace-normalized text search: exact bytes first, then reflowed form. * * The commit hook runs prettier, so a reflowed anchor must not retire a tooth. More than one match @@ -203,11 +212,14 @@ function collect(root: ts.Node, isWanted: (node: ts.Node) => return found; } -/** The one member with this name in the file, or why it is not one site. */ +/** The one member with this name in the file, or why it is not one site. A class constructor is a + * member too: its name is written `constructor`, and code that refuses a second plan while it opens + * the store lives there and nowhere else. */ function uniqueMember(source: ts.SourceFile, name: string): ts.Node | { reason: string } { const named = (node: ts.Node): boolean => - (ts.isMethodDeclaration(node) || ts.isFunctionDeclaration(node)) && - node.name?.getText(source) === name; + ((ts.isMethodDeclaration(node) || ts.isFunctionDeclaration(node)) && + node.name?.getText(source) === name) || + (ts.isConstructorDeclaration(node) && name === "constructor"); const members = collect(source, named); if (members.length !== 1) return { @@ -288,18 +300,32 @@ function statementSite( if (derive.statement === undefined) return { reason: "drop-statement needs a `statement` fragment" }; const fragment = derive.statement; + // A guard clause is a statement like any other: `if (...) throw ...;` is dropped whole, and the + // statement matcher normalizes whitespace so a reflow cannot make it unfindable. const containing = (node: ts.Node): boolean => (ts.isExpressionStatement(node) || ts.isVariableStatement(node) || ts.isReturnStatement(node) || - ts.isThrowStatement(node)) && - node.getText(source).includes(fragment); + ts.isThrowStatement(node) || + ts.isIfStatement(node)) && + containsText(node.getText(source), fragment); const statements = collect(member, containing); - if (statements.length !== 1) + // The innermost statement wins, the same rule conditions follow: a guard whose body is itself a + // statement is the guard, not the two of them. + const innermost = statements.filter( + (statement) => + !statements.some( + (other) => + other !== statement && + other.getStart(source) >= statement.getStart(source) && + other.getEnd() <= statement.getEnd(), + ), + ); + if (innermost.length !== 1) return { - reason: `${statements.length} statements in ${derive.within} contain the fragment, refusing to claim a check`, + reason: `${innermost.length} statements in ${derive.within} contain the fragment, refusing to claim a check`, }; - const statement = statements[0]!; + const statement = innermost[0]!; // The whole line goes: the indentation to its left and the newline to its right. A trailing comment // on that line would be dropped with it, and a statement that shares its line with anything else // would leave half of that line behind, so both are refused rather than silently damaged. @@ -355,10 +381,12 @@ function conditionSite( derive: Derive, ): ts.Expression | { reason: string } { const conditions = decisionExpressions(member); - const matching = - derive.condition === undefined - ? conditions - : conditions.filter((condition) => condition.getText(source).includes(derive.condition!)); + // Matched with the same whitespace normalization as everything else: a condition written across + // lines is still one condition, and a fragment that names it must not have to reproduce the line + // breaks prettier chose. + const mentions = (condition: ts.Expression): boolean => + derive.condition === undefined || containsText(condition.getText(source), derive.condition); + const matching = conditions.filter(mentions); const innermost = matching.filter( (condition) => !matching.some( diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index c22b6354..1ed71582 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -32,7 +32,8 @@ * * Where a mutant says where it applies: * - * - `derive: { within, operator, ... }` is a **selector plus an operator**: the member is named, a + * - `derive: { within, operator, ... }` is a **selector plus an operator**: the member is named + * (`constructor` when the site is in a class constructor), a * fragment inside it is matched (the same matcher as below, so a reflow cannot break it), and the * bytes to write are computed on every run. This is the form a tooth should have. Its identity is * the rule and the site it names, never the text that happens to be there today, and it is what @@ -91,8 +92,11 @@ const TARGETS: readonly Target[] = [ }, { name: "read-or-write-path-may-leave-the-frozen-files", - from: " if (!usable.has(path)) {", - to: " if (false && !usable.has(path)) {", + derive: { + within: "refusePermissionExpansion", + operator: "condition-never", + condition: "!usable.has(path)", + }, expect: "refuses a read or write path outside the frozen files", }, { @@ -109,8 +113,12 @@ const TARGETS: readonly Target[] = [ }, { name: "over-maximum-budget-is-accepted", - from: " if (spec.budget && !unknownBudget.length && !within(spec.budget, MAX_PATCH_BUDGET)) {", - to: " if (false) {", + derive: { + within: "refuseOutOfRange", + operator: "condition-never", + condition: + "spec.budget && !unknownBudget.length && !within(spec.budget, MAX_PATCH_BUDGET)", + }, expect: "refuses an unknown key inside budget or limits, and a range above the maximum", }, { @@ -121,8 +129,11 @@ const TARGETS: readonly Target[] = [ }, { name: "assumption-may-carry-a-dependency", - from: ' } else if (typeof task === "string" && !dependencies.includes(task)) {', - to: " } else if (false) {", + derive: { + within: "refuseRequirements", + operator: "condition-never", + condition: 'typeof task === "string" && !dependencies.includes(task)', + }, expect: "case 2: a lost obligation or a widened permission is refused, and an assumption cannot stand in for a dependency", }, @@ -130,9 +141,13 @@ const TARGETS: readonly Target[] = [ // The status query is the same rule read back: if it kept asking with a one-task budget it // would report a ready set narrower than what the run declared, and the two answers would differ. name: "the-status-query-ignores-the-declared-budget", - ast: { within: "deriveStatus" }, - from: " const ready = startableTasks(tasks, slots);", - to: " const ready = startableTasks(tasks, 1);", + derive: { + within: "deriveStatus", + operator: "replace-argument", + call: "startableTasks", + arg: 1, + }, + to: "1", expect: "a run that declares more slots reports the tasks it may start, not just the head", }, ], @@ -153,24 +168,30 @@ const TARGETS: readonly Target[] = [ // A failure the caller swallows still forbids the commit: nothing may be written up to the // failure and then kept by a normal return value. name: "swallowed-failure-still-commits", - ast: { within: "withPort" }, - from: " state.rollbackOnly = true;", - to: " void state.rollbackOnly;", + derive: { + within: "withPort", + operator: "drop-statement", + statement: "state.rollbackOnly = true;", + }, expect: "a failure the caller swallows still forbids the commit", }, { name: "stale-claim-may-deliver-again", - ast: { within: "claimTaskBoardEntry" }, - from: " if (!renewed) {", - to: " if (false && !renewed) {", + derive: { + within: "claimTaskBoardEntry", + operator: "condition-never", + condition: "!renewed", + }, expect: "renewing your own live claim does not start a new attempt", }, { name: "deliverer-may-judge-its-own-work", - ast: { within: "judgeTaskBoardEntry" }, - from: " if (existing.deliveredBy === input.agentId) {", - to: " if (false && existing.deliveredBy === input.agentId) {", + derive: { + within: "judgeTaskBoardEntry", + operator: "condition-never", + condition: "existing.deliveredBy === input.agentId", + }, expect: "the deliverer cannot judge its own deliverable", }, { @@ -192,18 +213,22 @@ const TARGETS: readonly Target[] = [ // One managed entry belongs to one run: answering with the first binding would hand a // second run's facts to whoever asked. name: "a-second-run-adopts-a-bound-entry", - ast: { within: "taskRunForEntry" }, - from: " if (runs.size > 1)", - to: " if (runs.size > 1 && false)", + derive: { + within: "taskRunForEntry", + operator: "condition-never", + condition: "runs.size > 1", + }, expect: "an entry bound by two runs is refused rather than answered with one of them", }, { // A managed entry's write belongs to its run's scope: the store is the only thing that can // tell a coordinated write from a verb reached around it. name: "a-managed-entry-ignores-the-coordinated-scope", - ast: { within: "requireManagedWriteScope" }, - from: " if (this.coordinatedRun === binding.runId) return;", - to: " if (true) return;", + derive: { + within: "requireManagedWriteScope", + operator: "condition-holds", + condition: "this.coordinatedRun === binding.runId", + }, expect: "a direct board verb cannot move an entry a run has adopted", }, ], @@ -237,38 +262,51 @@ const TARGETS: readonly Target[] = [ // The cache exists to be recomputable. A rebuild that returns without writing is the // difference between "the sources decide" and "the schema says so". name: "derived-rebuild-is-a-no-op", - ast: { within: "refreshDerived" }, - from: " this.putInputDigest(\n row.id,\n row.attempt >= 1 ? (frozen?.digest ?? this.inputDigest(row)) : null,\n );", - to: " void row.id;", + derive: { + within: "refreshDerived", + operator: "drop-statement", + statement: "this.putInputDigest(", + }, expect: "deleting the derived cache and rebuilding it yields the same view", }, { name: "round-releases-dependents-on-delivered-bytes", - ast: { within: "readAccepted" }, - from: " verdict: recorded?.verdict ?? null,", - to: ' verdict: "accepted",', + derive: { + within: "readAccepted", + operator: "replace-property", + property: "verdict", + }, + to: '"accepted"', expect: "an outside rejection withdraws the release of a dependent, and the round fails closed", }, { name: "verdict-lookup-not-bound-to-the-artifact", - ast: { within: "readAccepted" }, - from: " const recorded = verdictOf.get(roundChannel(runId), digest) as unknown as", - to: ' const recorded = verdictOf.get(roundChannel(runId), "%") as unknown as', + derive: { + within: "readAccepted", + operator: "replace-argument", + call: "verdictOf.get", + arg: 1, + }, + to: '"%"', expect: "the board verdict is what accepts an artifact, not the round's own column", }, { name: "cancellation-does-not-stop-a-claim", - ast: { within: "claim" }, - from: ' if (!id || !agentId) throw new Error("task and agent required");\n if (this.cancelled() !== null) throw new Error("round cancelled");', - to: ' if (!id || !agentId) throw new Error("task and agent required");', + derive: { + within: "claim", + operator: "drop-statement", + statement: 'if (this.cancelled() !== null) throw new Error("round cancelled");', + }, expect: "a cancellation names the lease it revokes, and the batch behind it is refused", }, { name: "round-does-not-pin-what-it-references", - ast: { within: "publishReady" }, - from: ' this.retainTaskBoardEntry({\n taskId: this.channel,\n entryId,\n owner: RETENTION_OWNER,\n reason: `round ${this.runId ?? "initial"} handoff for ${row.id}`,\n now: new Date(this.now).toISOString(),\n });', - to: " void entryId;", + derive: { + within: "publishReady", + operator: "drop-statement", + statement: "this.retainTaskBoardEntry({", + }, expect: "acceptance survives the entry's own TTL, because the round retains what it references", }, @@ -277,9 +315,12 @@ const TARGETS: readonly Target[] = [ // The borrowed view is one implementation serving two paths. Letting the offline port // answer from its own rule is exactly the divergence this target exists to catch. name: "the-offline-reader-decides-acceptance-on-its-own", - ast: { within: "openRoundQuery" }, - from: " accepted: () => readAccepted(db, resolved),", - to: " accepted: () => ({}),", + derive: { + within: "openRoundQuery", + operator: "replace-property", + property: "accepted", + }, + to: "() => ({})", expect: "the owner's view and the offline reader report the same facts", }, @@ -287,17 +328,22 @@ const TARGETS: readonly Target[] = [ // The frozen plan is the run's input, and installing a second one over it is the second // editable task truth the design forbids. The constructor refuses by policy digest. name: "a-second-plan-silently-adopts-the-run", - from: " if (recorded && recorded.policy !== wanted)", - to: " if (false && recorded && recorded.policy !== wanted)", + derive: { + within: "constructor", + operator: "condition-never", + condition: "recorded && recorded.policy !== wanted", + }, expect: "the frozen plan has one owner, and a second, different plan is refused", }, { // Verification is await-capable, so a ticket can be retired while it runs. Dropping the // re-check at the commit is how a decision made before the wait is applied after it. name: "the-commit-trusts-a-claim-the-board-retired", - ast: { within: "commitArtifact" }, - from: ' if (!this.live(row)) return "stale";', - to: ' if (false && !this.live(row)) return "stale";', + derive: { + within: "commitArtifact", + operator: "condition-never", + condition: "!this.live(row)", + }, expect: "a claim the board retires inside the verification window cannot be committed", }, @@ -484,9 +530,11 @@ const TARGETS: readonly Target[] = [ // Dropping the declaration restores a preference no plan enabled, which is what made a plan's // meaning depend on which planner read it. name: "the-continuation-is-not-declared", - ast: { within: "nextSessionMove" }, - from: ' if (!enabledConstraints(input.plan).includes("repair-first"))', - to: " if (false)", + derive: { + within: "nextSessionMove", + operator: "condition-never", + condition: '!enabledConstraints(input.plan).includes("repair-first")', + }, expect: "the continuation is a declared constraint, not the planner's default", }, { @@ -516,16 +564,22 @@ const TARGETS: readonly Target[] = [ { // A task someone is working is not offered to a second worker, at any budget. name: "the-budget-offers-a-claimed-task", - from: " if (input.claimed.includes(unit))", - to: " if (false && input.claimed.includes(unit))", + derive: { + within: "checkBudget", + operator: "condition-never", + condition: "input.claimed.includes(unit)", + }, expect: "the budget properties fire on a hand-built view, so deleting them cannot pass quietly", }, { // A bigger budget adds candidates; it may not drop one a smaller budget offered. name: "a-bigger-budget-may-drop-a-candidate", - from: " if (!input.ready.includes(unit))", - to: " if (false && !input.ready.includes(unit))", + derive: { + within: "checkBudget", + operator: "condition-never", + condition: "!input.ready.includes(unit)", + }, expect: "the budget properties fire on a hand-built view, so deleting them cannot pass quietly", }, @@ -533,38 +587,49 @@ const TARGETS: readonly Target[] = [ // Every condition in the checker is deleted once, and the case that names it has to fail: // a condition no case can reach is a comment, not a check. name: "the-completion-does-not-bind-the-verdict-to-the-bytes", - from: " else if (verdict.digest !== artifact)", - to: " else if (false && verdict.digest !== artifact)", + derive: { + within: "checkCompletion", + operator: "condition-never", + condition: "verdict.digest !== artifact", + }, expect: "the checker reports a completion whose verdict judged other bytes", }, { // Whether an input is current, and whether its bytes are the bytes its verdict judged, is the // one acceptance predicate's answer; the checker asks it instead of comparing by hand. name: "the-input-is-not-required-to-be-accepted", - ast: { within: "checkInputs" }, - from: " if (dependencyUnit && !isAccepted(dependencyUnit, context.facts))", - to: " if (false && dependencyUnit && !isAccepted(dependencyUnit, context.facts))", + derive: { + within: "checkInputs", + operator: "condition-never", + condition: "dependencyUnit && !isAccepted(dependencyUnit, context.facts)", + }, expect: "the checker reports a completion resting on an input that drifted", }, { name: "the-input-may-be-cancelled", - ast: { within: "checkInputs" }, - from: " if (cancelled(context.facts, dependency))", - to: " if (false && cancelled(context.facts, dependency))", + derive: { + within: "checkInputs", + operator: "condition-never", + condition: "cancelled(context.facts, dependency)", + }, expect: "the checker reports a completion resting on a cancelled input", }, { name: "the-completion-ignores-a-cancelled-unit", - ast: { within: "checkCompletion" }, - from: " if (cancelled(context.facts, unit.id))", - to: " if (false)", + derive: { + within: "checkCompletion", + operator: "condition-never", + condition: "cancelled(context.facts, unit.id)", + }, expect: "the checker reports a completion of a cancelled unit", }, { name: "the-dispatch-does-not-require-a-closed-input", - ast: { within: "checkDispatch" }, - from: ' checkInputs(context, unit, "dispatch");', - to: " void checkInputs;", + derive: { + within: "checkDispatch", + operator: "drop-statement", + statement: 'checkInputs(context, unit, "dispatch");', + }, expect: "the checker reports a dispatch whose input is not accepted", }, ], @@ -583,8 +648,13 @@ const TARGETS: readonly Target[] = [ // The closure takes the unit-level predicate, so the marker names the two conditions it // joins: a unit's own acceptance and every dependency being in the set. name: "acceptance-closure-dropped", - from: " own(unit) && unit.inputs.dependencies.every((dependency) => accepted.has(dependency));", - to: " own(unit);", + derive: { + within: "acceptedClosure", + operator: "neutralize-term", + condition: + "own(unit) && unit.inputs.dependencies.every((dependency) => accepted.has(dependency))", + term: "unit.inputs.dependencies.every((dependency) => accepted.has(dependency))", + }, expect: "a dependency that is not accepted makes fusion pay a rollback and a retry", }, ], @@ -626,24 +696,31 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "the-round-client-does-not-require-a-daemon", - from: ' if (!state || state.transport !== "http" || !state.host || !state.port || !state.token) {', - to: ' if (false && (!state || state.transport !== "http" || !state.host || !state.port || !state.token)) {', + derive: { + within: "roundDaemon", + operator: "condition-never", + condition: + '!state || state.transport !== "http" || !state.host || !state.port || !state.token', + }, expect: "a driver refuses without a daemon, and no driver opens a database of its own", }, { // Whether a call to an endpoint the caller itself serves can be answered depends on the // caller not blocking - the assumption that turned this failure into a 305-second wait. name: "the-round-client-calls-the-endpoint-it-serves", - from: " if (state.pid === process.pid) {", - to: " if (false && state.pid === process.pid) {", + derive: { + within: "roundDaemon", + operator: "condition-never", + condition: "state.pid === process.pid", + }, expect: "a client refuses to call the endpoint its own process serves", }, { // Without a bound, a blocked host is indistinguishable from a slow one, and the caller waits // out the transport's own timeout instead of being told what to look at. name: "the-round-client-has-no-limit-on-how-long-it-waits", - from: " return await httpCall(state, method, params, { timeoutMs });", - to: " return await httpCall(state, method, params, {});", + derive: { within: "call", operator: "replace-argument", call: "httpCall", arg: 3 }, + to: "{}", expect: "a call to a host that never answers gives up in seconds and names the reason", }, ], @@ -662,8 +739,13 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "the-loop-ignores-the-slot-count", - from: " const batch = legal.slice(0, input.slots);", - to: " const batch = legal.slice(0, 1);", + derive: { + within: "dispatchPlan", + operator: "replace-argument", + call: "legal.slice", + arg: 1, + }, + to: "1", expect: "a declared slot count is reached, and the claims overlap in time", }, @@ -682,16 +764,19 @@ const TARGETS: readonly Target[] = [ // layer, not a mutation target), so what this mutant breaks is the loop's hand-off of the // declared bound: the invariant is unchanged, its anchor follows the code carrying it. name: "fusion-ignores-the-declared-bound", - from: " bound: input.sessions?.bound ?? 1,", - to: " bound: Number.MAX_SAFE_INTEGER,", + derive: { within: "dispatchPlan", operator: "replace-property", property: "bound" }, + to: "Number.MAX_SAFE_INTEGER", expect: "a fused chain stops at the declared bound and does not swallow the plan", }, { // The evidence of fusion is the session the worker reported, not the one the loop asked for: // a worker that quietly starts its own session must not be reported as fused. name: "fusion-counts-a-session-the-worker-did-not-use", - from: " if (unit.sessionId !== session.id) {", - to: " if (false && unit.sessionId !== session.id) {", + derive: { + within: "dispatchPlan", + operator: "condition-never", + condition: "unit.sessionId !== session.id", + }, expect: "a worker that starts its own session is not reported as fusion", }, { @@ -704,8 +789,11 @@ const TARGETS: readonly Target[] = [ }, { name: "a-failed-worker-is-reported-as-a-run-that-finished", - from: " if (result.failure !== undefined || result.artifact === undefined)", - to: " if (false && (result.failure !== undefined || result.artifact === undefined))", + derive: { + within: "dispatchUnit", + operator: "condition-never", + condition: "result.failure !== undefined || result.artifact === undefined", + }, expect: "a unit whose worker failed is asked once in a pass", }, ], @@ -744,8 +832,7 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "the-serving-process-never-releases-its-lease", - from: " lease.release();", - to: " // lease.release();", + derive: { within: "serveHttp", operator: "drop-statement", statement: "lease.release();" }, expect: "a host releases its lease when it stops, so the next host can take the store", }, ], @@ -762,34 +849,50 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "a-suggestion-outside-the-legal-set-is-scored", - from: " if (!legalSet.has(suggestion.taskId)) {", - to: " if (false && !legalSet.has(suggestion.taskId)) {", + derive: { + within: "refuseSuggestion", + operator: "condition-never", + condition: "!legalSet.has(suggestion.taskId)", + }, expect: "a suggestion outside the legal set is refused, however high it scores", }, { name: "an-unmodelled-action-is-scored", - from: ' if (suggestion.action !== "next") {', - to: " if (false) {", + derive: { + within: "refuseSuggestion", + operator: "condition-never", + condition: 'suggestion.action !== "next"', + }, expect: "an unmodelled action is refused rather than scored", }, { name: "a-disabled-source-is-asked-anyway", - from: " if (source.enabled === false) {", - to: " if (false && source.enabled === false) {", + derive: { + within: "orderCandidates", + operator: "condition-never", + condition: "source.enabled === false", + }, expect: "a disabled or failing source falls back to the rule policy, and says why", }, { name: "a-score-from-another-scope-is-reused", - from: " if (\n provenance.sessionId !== projection.sessionId ||\n provenance.branchId !== projection.branchId\n ) {", - to: " if (false) {", + derive: { + within: "refuseSuggestion", + operator: "condition-never", + condition: + "provenance.sessionId !== projection.sessionId || provenance.branchId !== projection.branchId", + }, expect: "a score from another session or branch is not reused", }, { name: "a-version-mismatch-still-counts-as-the-same-reading", - from: " if (provenance.parametersVersion !== projection.parametersVersion)\n missing.push(`parametersVersion=${provenance.parametersVersion}`);", - to: " if (false) missing.push(`parametersVersion=${provenance.parametersVersion}`);", + derive: { + within: "missingValidityInputs", + operator: "condition-never", + condition: "provenance.parametersVersion !== projection.parametersVersion", + }, expect: "a changed parameter or projection version makes an old score a new one", }, { @@ -801,8 +904,11 @@ const TARGETS: readonly Target[] = [ }, { name: "a-ranking-survives-into-the-claim", - from: " return legalNow.includes(adopted.taskId)", - to: " return true", + derive: { + within: "revalidateSuggestion", + operator: "condition-holds", + condition: "legalNow.includes(adopted.taskId)", + }, expect: "an adopted ranking is re-checked where the write happens", }, ], @@ -828,23 +934,32 @@ const TARGETS: readonly Target[] = [ // A run cannot adopt an entry for work it never froze: otherwise the binding names a task // no decision was ever read against. name: "a-binding-ignores-whether-the-task-was-frozen", - from: " if (!isFrozen(store, request.runId, request.taskId))", - to: " if (false && !isFrozen(store, request.runId, request.taskId))", + derive: { + within: "bindRunEntry", + operator: "condition-never", + condition: "!isFrozen(store, request.runId, request.taskId)", + }, expect: "a binding refuses what the store does not hold", }, { // The binding names an entry the board really holds, on the channel the caller names. name: "a-binding-does-not-check-the-entry-exists", - from: " if (!store.getTaskBoardEntryById(request.boardTaskId, request.entryId))", - to: " if (false && !store.getTaskBoardEntryById(request.boardTaskId, request.entryId))", + derive: { + within: "bindRunEntry", + operator: "condition-never", + condition: "!store.getTaskBoardEntryById(request.boardTaskId, request.entryId)", + }, expect: "a binding refuses what the store does not hold", }, { // One entry carries one task: without this a second run would fence an entry it does not // own, and the fence would refuse the first run's own writes. name: "one-entry-is-bound-to-two-tasks", - from: " if (bound && (bound.runId !== request.runId || bound.taskId !== request.taskId))", - to: " if (false && bound && (bound.runId !== request.runId || bound.taskId !== request.taskId))", + derive: { + within: "bindRunEntry", + operator: "condition-never", + condition: "bound && (bound.runId !== request.runId || bound.taskId !== request.taskId)", + }, expect: "a binding refuses what the store does not hold", }, { @@ -867,24 +982,33 @@ const TARGETS: readonly Target[] = [ { // The binding is re-read where the write happens, not where the caller decided to make it. name: "a-coordinated-write-skips-the-binding-recheck", - from: " if (binding.runId !== request.runId)", - to: " if (false && binding.runId !== request.runId)", + derive: { + within: "coordinatedBoardWrite", + operator: "condition-never", + condition: "binding.runId !== request.runId", + }, expect: "a coordinated write refuses an entry that is not this run's", }, { // The run surface's transitions: a plan the run cannot satisfy is refused while it is still // a proposal rather than frozen into a task that can never be ready. name: "the-plan-may-freeze-a-dangling-dependency", - from: " if (!known.has(dependency))", - to: " if (false && !known.has(dependency))", + derive: { + within: "freezeRunPlan", + operator: "condition-never", + condition: "!known.has(dependency)", + }, expect: "a freeze cannot dangle, repeat a task, or lean on itself", }, { // A task that waits for itself is a task that is never ready, and the freeze is the last // point at which that is still only a proposal. name: "a-task-may-depend-on-itself", - from: " if (dependency === task.taskId)", - to: " if (false && dependency === task.taskId)", + derive: { + within: "freezeRunPlan", + operator: "condition-never", + condition: "dependency === task.taskId", + }, expect: "a freeze cannot dangle, repeat a task, or lean on itself", }, @@ -901,8 +1025,8 @@ const TARGETS: readonly Target[] = [ // The binding records which channel carries the entry, which is what lets a status reader // resolve it without searching every channel. name: "a-binding-does-not-record-its-channel", - from: " payload: JSON.stringify({ boardTaskId: request.boardTaskId }),", - to: " payload: null,", + derive: { within: "bindRunEntry", operator: "replace-property", property: "payload" }, + to: "null", expect: "a run registers, freezes a plan, adopts entries, and reads it all back", }, @@ -910,24 +1034,36 @@ const TARGETS: readonly Target[] = [ // A run-level cancellation is the run's fact, not a task's: the schema's empty task id is // what keeps it from colliding with a task that has no name. name: "a-run-cancellation-names-a-task", - from: ' taskId: request.taskId ?? "",', - to: ' taskId: request.taskId ?? "-",', + derive: { + within: "cancelRun", + operator: "replace-property", + property: "taskId", + in: "taskId: request.taskId", + }, + to: 'request.taskId ?? "-"', expect: "cancelling a run is recorded once, stops its managed writes, and is readable", }, { // Cancelling a task the plan never froze would name nothing while reading as a fact about // the run. name: "a-cancellation-ignores-whether-the-task-was-frozen", - from: " if (request.taskId !== undefined && !isFrozen(store, request.runId, request.taskId))", - to: " if (false && request.taskId !== undefined && !isFrozen(store, request.runId, request.taskId))", + derive: { + within: "cancelRun", + operator: "condition-never", + condition: + "request.taskId !== undefined && !isFrozen(store, request.runId, request.taskId)", + }, expect: "cancelling one task names it, and a task the plan never froze cannot be cancelled", }, { // There is nothing to cancel in a run this store cannot name, and the refusal says so // rather than leaving it to the transaction's own message. name: "an-unknown-run-can-be-cancelled", - from: " if (!store.taskRunManifest(request.runId))", - to: " if (false && !store.taskRunManifest(request.runId))", + derive: { + within: "cancelRun", + operator: "condition-never", + condition: "!store.taskRunManifest(request.runId)", + }, expect: "status is a read: an unknown run has no manifest and is not registered by being asked", }, @@ -951,8 +1087,8 @@ const TARGETS: readonly Target[] = [ // The wire drops an adoption request: the entry is created, the caller is told the put // succeeded, and no run manages it - the silent divergence the epoch rule exists for. name: "the-wire-drops-an-adoption-request", - from: " adopt: optionalAdoption(params.adopt),", - to: " adopt: undefined,", + derive: { within: "parseTaskBoardParams", operator: "replace-property", property: "adopt" }, + to: "undefined", expect: "adoption is part of the transition that creates the entry, so a refusal leaves no entry", }, @@ -966,8 +1102,11 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "the-declined-shared-checks-still-run", - from: ' const declined = route.verify.sharedChecks === "none";', - to: " const declined = false;", + derive: { + within: "planNarrowVerify", + operator: "condition-never", + condition: 'route.verify.sharedChecks === "none"', + }, expect: "a route that declares the shared checks not applicable narrows to its own tests", }, { @@ -988,15 +1127,22 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "a-declined-floor-may-have-no-tests", - from: 'if (sharedChecks === "none" && !route.tests.length) {', - to: 'if (false && sharedChecks === "none" && !route.tests.length) {', + derive: { + within: "validateRouteVerify", + operator: "condition-never", + condition: 'sharedChecks === "none" && !route.tests.length', + }, expect: "verify.sharedChecks must be a known declaration, and declining needs its own tests", }, { name: "an-unknown-shared-checks-value-is-accepted", - from: 'if (sharedChecks !== undefined && sharedChecks !== "always" && sharedChecks !== "none") {', - to: 'if (false && sharedChecks !== undefined && sharedChecks !== "always" && sharedChecks !== "none") {', + derive: { + within: "validateRouteVerify", + operator: "condition-never", + condition: + 'sharedChecks !== undefined && sharedChecks !== "always" && sharedChecks !== "none"', + }, expect: "verify.sharedChecks must be a known declaration, and declining needs its own tests", }, @@ -1096,8 +1242,12 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "fusion-books-the-shared-startup-per-unit", - from: " sharedStartupMs: sessions * params.sessionStartMs,", - to: " sharedStartupMs: shape.units * params.sessionStartMs,", + derive: { + within: "fusionAccounting", + operator: "replace-property", + property: "sharedStartupMs", + }, + to: "shape.units * params.sessionStartMs", expect: "fusion books the shared startup once per session, not once per unit", }, { @@ -1108,14 +1258,21 @@ const TARGETS: readonly Target[] = [ }, { name: "fusion-removes-boundaries-that-are-not-there", - from: " boundarySavedMs: (shape.units - sessions) * params.contextMsPerUnit,", - to: " boundarySavedMs: shape.units * params.contextMsPerUnit,", + derive: { + within: "fusionAccounting", + operator: "replace-property", + property: "boundarySavedMs", + }, + to: "shape.units * params.contextMsPerUnit", expect: "a fusion bound of one unit removes no boundary and still pays the startup", }, { name: "fusion-reads-a-gain-out-of-an-assumed-term", - from: ' if (!params.sessionStartMeasured) return "unmeasured";', - to: ' if (false) return "unmeasured";', + derive: { + within: "fusionVerdict", + operator: "condition-never", + condition: "!params.sessionStartMeasured", + }, expect: "an assumed session startup never reads as a gain", }, ], @@ -1163,14 +1320,24 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "the-identity-is-not-hex", - from: ' return createHash("sha256").update(bytes).digest("hex");', - to: ' return createHash("sha256").update(bytes).digest("base64url");', + derive: { + within: "workDigest", + operator: "replace-argument", + call: 'createHash("sha256").update(bytes).digest', + arg: 0, + }, + to: '"base64url"', expect: "the identity of work bytes is sha256 in hex, at full length", }, { name: "the-json-variant-does-not-digest-json", - from: " return workDigest(JSON.stringify(value));", - to: " return workDigest(String(value));", + derive: { + within: "workDigestOf", + operator: "replace-argument", + call: "workDigest", + arg: 0, + }, + to: "String(value)", expect: "the JSON variant digests JSON text, so key order is the caller's", }, ], From b515f94ddf324836e44ddac40eaf8c97c2488e9c Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 22:21:25 +0800 Subject: [PATCH 30/32] Seven operators the catalogues name, and 18 more teeth converted The 33 teeth still hand-written after the register-wide conversion were classified by what they would need, and the classification was checked against what mutation tools actually name (Stryker's supported mutators, pitest's mutator list, cargo-mutants' patterns, Cosmic Ray's operator concept). Thirteen of those sites fall into mutation classes every catalogue carries, so the vocabulary grew instead of the class being declared impossible: negate-condition pitest NEGATE_CONDITIONALS / Stryker boolean literals 2 teeth negate-comparison pitest NEGATE_CONDITIONALS / Stryker EqualityOperator 2 teeth remove-conditionals pitest REMOVE_CONDITIONALS 1 tooth remove-call Stryker filter/slice/sort removals / pitest VOID_METHOD_CALLS 2 teeth replace-call Stryker MethodExpression / pitest CONSTRUCTOR_CALLS 1 tooth replace-initializer pitest PRIMITIVE_RETURNS, INLINE_CONSTS / literal mutators 5 teeth replace-iterable Stryker ArrayDeclaration / pitest EMPTY_RETURNS 3 teeth Each computes its own bytes where the mutation determines them (false, true, the negated operator, the call's receiver, the guard's body) and takes the mutant's `to` only where the new value is a choice - the rule the earlier operators already followed. Operators are now a table of resolvers over one context, so adding one is an entry plus its name in the union, and the dispatch's complexity stopped growing with the vocabulary. Two widenings, each shown by a real tooth that could not convert: - A function bound to a name is a member. `const count = (label) => {...}` is the whole of board-worker.ts's logic and it is not a function declaration, so its only tooth had no scope to name; the selector now accepts a variable whose initializer is an arrow function or a function expression. - Two identical calls need a holder to tell them apart. dispatchPlan calls board.candidates() on offer and again filtered; `in` now names a fragment of the statement the call sits in, as it already named the holding object literal for replace-property and the call's own text for replace-argument. A wrong `within` was shipped again, and this time the anchors pass caught it loudly: the guard is in runParentCheck and I named instrumentCommit, so it refused with "0 guards in instrumentCommit match the selector". The first pass's wrong member was the quiet kind - checkDispatch instead of checkCompletion - where the pass resolved, the mutant was still caught, and only the sweep showed it was caught by the suite rather than by the case that names the rule. Wrong member is loud when it finds nothing, silent when it finds the same shape twice: the sweep after a conversion is not optional. Reading: anchors 110 of 110 resolve over 22 targets; mutants 110 of 110 caught by the named test, all 110 by the case their expect names, 22 of 22 targets restored byte-identically, exit 0 in 176 s; the anchor cases 30 pass; test:product 1576 pass; verify:static exit 0; lint 0 findings; complexity gate ok. 95 of the 110 teeth are derived now. The residue is 15 teeth and its classification is measured, not guessed: three fragments inside a string (two SQL clauses, one SQLite unit - catalogues mutate whole literals, and SQL has its own catalogue); two at module top level where there is no member to name; two moving an array index, which corrects this record's own earlier claim that the catalogues cover that class (FirstToLast is a method named first); two writing two statements on one line; one rewriting a comparison into a range test, where the catalogue's negation would not realise the violation the name states; four bespoke expression rewrites; and one that changes two things at once. Three clusters would repay an operator and are recorded as open (clause-level SQL mutation, a file-scope selector, an index operator); none is a widening, and the measured cost of not having them is seven teeth. The catalogue check also supplies the rationale this record was missing: pitest's "Less is more" names subsumption - a mutant subsumed by others adds runtime, not confidence - which is what the retirement criterion measures rather than guesses at; cargo-mutants' unviable mutants and its documented refusals to generate some mutations are the same argument used here for leaving rare shapes alone. And one thing is deliberately not borrowed: those tools count a mutant killed when any test fails, while this register counts it only when the case it names is the one that fails, which is why "caught by the suite, not by the case that names the rule" is a defect it reports and they cannot. Docs and code in one commit: the record carries the operator table, the two widenings, both wrong-member episodes, the measured residue and the two catalogue lessons; the ledger says 95 of 110 are derived and gains this pass's readings; the teeth header says which names a member can have and that the operators are the catalogues' mutation classes. --- ...-09-24-mutants-are-derived-not-anchored.md | 138 +++++-- ...-mutants-are-derived-not-anchored.zh-CN.md | 41 +- .../design/task-unit-semantics-obligations.md | 7 +- tests/tools/mutation-anchor.test.ts | 262 +++++++++++++ tools/mutation-anchor.ts | 353 +++++++++++++++++- tools/mutation-teeth.ts | 159 ++++---- 6 files changed, 831 insertions(+), 129 deletions(-) diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index 8d00756d..42a46ece 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -129,42 +129,108 @@ targets`. The target `tools/agent-verify.ts` went with its single tooth, so the strings rather than by a filesystem. **Landed:** the resolver is `tools/mutation-anchor.ts` (it held no state, so it moved out of the sweep script and the tests can call it), with 16 cases over source strings and no filesystem access. - **Extended to the whole register (2026-09-24): 59 more teeth converted, and the register is now 77 - derived of 110** - every class of site the six operators can express, taken target by target and - swept after each: `condition-never` on a guard or an initializer (44 in total), `neutralize-term` (9), - `replace-property` (9), `replace-argument` (6), `drop-statement` (6), `condition-holds` (3). The - reading after it: `110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, - 22 of 22 restored byte-identically, 171 s. - Three vocabulary gaps showed up as _sites the operators could not name_, and each was closed with a - case in `tests/tools/mutation-anchor.test.ts` (now 19 cases) rather than argued about: - - **A class constructor is a member.** `BoardAdmission`'s constructor refuses a second plan while it - opens the store, and `uniqueMember` only knew method and function declarations, so that site had no - `within` at all. Its name is `constructor`; matching it is what let - `a-second-plan-silently-adopts-the-run` convert. - - **A condition written across lines is one condition.** The candidate filter compared raw text while - the whole-condition check compared whitespace-normalized text, so a fragment of a wrapped `if` found - no candidate and then matched nothing. Both now normalize, and a filter does not require the - fragment to be unique _inside_ the candidate - `b` appears three times in `a && (b || !b)`, and that - is still the condition a selector naming `b` means. - - **A guard clause is a statement.** `if (...) throw ...;` was not among the statements - `drop-statement` would remove, because an `if` is an `IfStatement` rather than an expression, - declaration, `return` or `throw`. It is now, and statements resolve to the **innermost** one - containing the fragment, the rule conditions already followed - which is also what keeps a fragment - from matching both a guard and the statement inside it. - One wrong `within` was shipped and the sweep caught it, which is the argument for sweeping after a - conversion rather than trusting the anchors pass: `the-completion-ignores-a-cancelled-unit` was aimed - at `checkDispatch` instead of `checkCompletion` (the fragment occurs in both members), so the anchors - pass resolved, the mutant was still caught - and it was caught by the _suite_, not by the case that - names a completion of a cancelled unit. Re-aimed, that case fails again. - **What is left hand-written is 33 teeth, and they cluster by the operator they would need:** a - comparison rewritten (`=== "none"` to `!== "always"`, 3), a fragment inside a template or a SQL string - (5), an iterable emptied or a filter dropped (4), a call or `new` replaced or unwrapped (6), an index - moved (`[0]` to `[1]`, 2), an initializer replaced by a different expression (3), a condition negated - (2), a statement rewritten into another statement (3), two statements sharing one line (2), a literal - swapped (1), a site at module top level, where there is no member to name (1), and one tooth that - changes two things at once (`every-task-is-frozen-at-position-zero`, kept deliberately). Four of these - clusters look worth an operator (`replace-comparison`, `replace-fragment`, `empty-iterable`, - `replace-callee`); the rest are single sites or shapes a general operator would make ambiguous. + **Extended to the whole register (2026-09-24): 59 more teeth converted, then the operators the + catalogues name, and the register is now 95 derived of 110.** First pass: every class of site the + six operators already had, taken target by target and swept after each - `condition-never` on a + guard or an initializer (44 in total), `neutralize-term` (9), `replace-property` (9), + `replace-argument` (6), `drop-statement` (6), `condition-holds` (3). + Second pass: the 33 that were left were classified by what they would need, and the classification + was checked against what mutation tools actually name (Stryker's supported mutators, pitest's + mutator list, cargo-mutants' patterns, Cosmic Ray's operator concept). Thirteen of those sites fell + into mutation classes every catalogue carries, so the vocabulary grew by seven operators rather than + the class being declared impossible: + + | operator | what the catalogues call it | teeth | + | --------------------- | ----------------------------------------------------------------------- | ----- | + | `negate-condition` | pitest `NEGATE_CONDITIONALS`; Stryker boolean literals | 2 | + | `negate-comparison` | pitest `NEGATE_CONDITIONALS`; Stryker `EqualityOperator` | 2 | + | `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | + | `remove-call` | Stryker's filter/slice/sort removals; pitest `VOID_METHOD_CALLS` | 2 | + | `replace-call` | Stryker `MethodExpression`; pitest `CONSTRUCTOR_CALLS` | 1 | + | `replace-initializer` | pitest `PRIMITIVE_RETURNS`, `INLINE_CONSTS`; Stryker's literal mutators | 5 | + | `replace-iterable` | Stryker `ArrayDeclaration`; pitest `EMPTY_RETURNS` | 3 | + + Each one computes its own bytes where the mutation determines them (`false`, `true`, the negated + operator, the call's receiver, the guard's body) and takes the mutant's `to` only where the new value + is a choice - the same rule the earlier operators follow. `tests/tools/mutation-anchor.test.ts` has a + case per operator, including each refusal (an `else`, a block body, a call that is not a method call, + a missing initializer), and is now 30 cases. + Two widenings came with them, both of the same kind as the ones before - a site the vocabulary could + not name, shown by a real tooth failing to convert: + + - **A function bound to a name is a member.** `const count = (label) => {...}` in + `evals/ooo-execution/board-worker.ts` is the whole of that file's logic, and it is not a function + declaration, so its only tooth had no scope to name. The selector now accepts a variable whose + initializer is an arrow function or a function expression. + - **Two identical calls need a holder to tell them apart.** `dispatchPlan` calls + `board.candidates()` on offer and again filtered; the callee text is the same in both. `in` now + names a fragment of the statement the call sits in (it already named the holding object literal for + `replace-property` and the call's own text for `replace-argument`). + + One wrong `within` was shipped in the first pass and the **anchors pass**, not the sweep, is what + caught it: `the-completion-ignores-a-cancelled-unit` was aimed at `checkDispatch` instead of + `checkCompletion` (the fragment occurs in both members). That is the quieter failure: the anchors pass + resolves, the mutant is caught - by the _suite_, not by the case that names a completion of a + cancelled unit. In the second pass the same mistake was loud instead, because the guard I aimed at is + in `runParentCheck` and I named `instrumentCommit`: the pass refused with "0 guards in + instrumentCommit match the selector". A wrong member is loud when it finds nothing and silent when it + finds the same shape twice, which is why the sweep after a conversion is not optional. + **Reading after both passes:** `anchors: 110 of 110 resolve, over 22 targets`; `mutants: 110 of 110 +caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored + byte-identically, exit 0 in 176 s. + + **What is left hand-written is 15 teeth, and the classification is now measured rather than + guessed:** + + - three are the same shape the tools also cannot express: a fragment **inside a string** - two SQL + clauses (`prune-ignores-retention`, `bounded-pin-never-expires`) and one SQLite unit + (`the-grace-uses-a-unit-sqlite-does-not-know`). Catalogues mutate the whole literal + (`StringLiteral` to `""`), and SQL has its own catalogue (SQLMutation: clause deletion, `AND` to + `OR`, NULL mutations), so the honest next step for these is clause-level SQL operators, not byte + surgery; + - two are at **module top level**, where there is no member to name: a guard in a script + (`judge-may-judge-its-own-delivery`) and `export const CLOCK_GRACE_MS = 50`. Giving `derive` a + file-scope mode is a design decision, not a widening, and it is not taken here; + - two move an **index** (`[0]` to `[1]`). This one corrects this record's own earlier list: it said + "index moved (2)" was among the classes the catalogues cover, and that is wrong. `FirstToLast` is a + method _named_ `first`; no catalogue checked here mutates an array index; + - two write **two statements on one line** (`fusion-rollback-not-counted`, + `a-score-with-no-recorded-history-counts-as-a-reproduction`) - the line discipline of + `drop-statement`, not a mutation class; + - one rewrites a comparison into a range test (`=== 1` to `>= 1`), where the catalogue's negation + (`!== 1`) would not realise the violation the name states; + - four are bespoke expression rewrites with no operator behind them: a constructor swapped for a + crafted substitute of the same interface (`the-status-read-path-opens-the-rounds-store`), a + multi-line call replaced by a stub fact (`a-coordinated-write-skips-its-run-fact`), a wrapper + unwrapped to its inner call (`the-daemon-verb-skips-the-coordinated-path`), and + `view-adapter-accepts-bytes`, where two statements become one `return`; + - one, `every-task-is-frozen-at-position-zero`, changes two things at once and is kept deliberately. + + Three of those clusters would repay an operator and are recorded as open: clause-level SQL mutation, + a file-scope selector, and an index operator. None is added here, because none is a widening - each + is a new scoping rule or a new domain - and the measured cost of not having them is seven teeth. + Two things the catalogues say about the shape of this work are worth carrying: + + - **Subsumption.** pitest's _Less is more_ names the effect this record's retirement criterion + measures: a mutant that the combination of others already subsumes adds no confidence, only + runtime. pitest avoids them with a fixed default set; this register measures it instead - the sweep + says which case fails, and the case is then read to see whether it states the rule. That is why 37 + teeth were retired rather than kept, and why the answer to "should every conceivable mutation be + here" is no. + - **Unviable mutants.** cargo-mutants counts a mutant that does not compile separately and prints it + only on request; it excludes test functions and `unsafe`, and it documents outright why it does not + generate some mutations (`==` to `<` "too prone to generate false positives", `-a` to `+a` "too + prone to generate unviable cases"). This register has no unviable mutants by construction - every + tooth has a named case - but the argument is the same one used above for leaving the residual + clusters alone: a general operator over a shape that is rare here produces mostly noise. + + The one thing this design does _not_ borrow from those tools is the definition of a kill. StrykerJS, + pitest and cargo-mutants count a mutant as killed when _any_ test fails; several can report which + test did it (`fullMutationMatrix`), but the purpose is to find the assertion to strengthen, not to + gate. Here a tooth is counted only when **the case it names** is the one that fails, which is why + "caught by the suite, not by the case that names the rule" is a defect this register reports and none + of those tools would. + 2. The pilot target converted and swept: same names, same cases, same catches; then a refactor inside it that demonstrates a derived tooth surviving what a byte anchor did not. **Landed**, with one correction: the demo first refactor was a rename plus a condition lifted into `const closed = spent || diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index 5121478b..fe2c4b3a 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -29,13 +29,40 @@ **一颗牙的耐久部分是它的名字、它的算子、它的选择器,以及必须失败的那个用例。位置与替换字节每次从语法树算出来。** 1. **六个作用在具名选择器上的算子**,作用域限定在一个成员内:`condition-never`(选择器指到的那条条件永不成立)、`condition-holds`(它始终成立——镜像的那个,因为写成 `return a && b` 的规则说的是“这条条件为真”,把它换成 `false` 是把规则反过来,而不是去掉)、`neutralize-term`(该条件里的一项变成它的恒等元——`&&` 下为 `true`,`||` 下为 `false`)、`replace-argument`(某个调用的第 _n_ 个实参换成声明的片段)、`replace-property`(具名对象属性的值换成声明的片段)、`drop-statement`(选择器指到的那条语句连行带缩进被删掉)。两条规则让选择器保持诚实:匹配到多于一处、或者片段同时命中两个候选的,一律拒绝而不是猜;两个“整条件”算子拒绝只命名条件一部分的片段,因为整段替换会把 mutant 悄悄放大成“这条规则的每一条理由”——一项有它自己的算子。规则用不上这些时,仍可用“成员作用域 + 最小片段”的旧形式;而**“最小”是要点**:不要消息文本、不要同级实参、一项能表达的不要写成整条语句。 - **已推广到整个登记处(2026-09-24):又转了 59 颗,登记处现在是 110 颗里 77 颗 derived。**六个算子能表达的每一类点位都按目标逐个转完,每转完一个目标就 sweep:`condition-never`(守卫或初始化器,共 44)、`neutralize-term`(9)、`replace-property`(9)、`replace-argument`(6)、`drop-statement`(6)、`condition-holds`(3)。转完的读数:`110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐字节还原,171 秒。 - 过程中暴露了三个"算子点不到这个点位"的缺口,每个都用 `tests/tools/mutation-anchor.test.ts` 里的一条用例补上(现在 19 条),而不是靠讨论: - - **类构造函数也是成员。**`BoardAdmission` 的构造函数在打开 store 时拒绝第二个 plan,而 `uniqueMember` 只认方法声明和函数声明,那个点位当时连 `within` 都写不出来。它的名字是 `constructor`;认它之后 `a-second-plan-silently-adopts-the-run` 才转得动。 - - **跨行写的条件仍然是一个条件。**候选过滤器比的是原始文本,而整条件检查比的是空白归一化后的文本,于是被折行 `if` 的片段先找不到候选、再匹配不上。现在两边都归一化;并且过滤器不要求片段在候选内部唯一——`b` 在 `a && (b || !b)` 里出现三次,而它仍然是指名 `b` 的选择器所指的那个条件。 - - **守卫语句也是一种语句。**`if (...) throw ...;` 原本不在 `drop-statement` 会删的语句里,因为 `if` 是 `IfStatement` 而非表达式、声明、`return` 或 `throw`。现在它算,并且语句解析取**最内层**那个包含片段的语句——条件早就是这个规则,这也正是不让片段同时命中守卫和守卫里那条语句的原因。 - 有一处 `within` 我写错了,是 sweep 抓出来的,这也正是"转完必须 sweep、不能只信 anchors pass"的理由:`the-completion-ignores-a-cancelled-unit` 被指向 `checkDispatch` 而不是 `checkCompletion`(那个片段在两个成员里都出现),于是 anchors pass 照样通过、mutant 照样被抓——但抓它的是**整个 suite**,不是那条"报告一个被取消单元的完成"的用例。重新瞄准后,那条用例又失败了。 - **剩下 33 颗仍是手写,且按"需要什么算子"成簇:**比较运算符被改写(`=== "none"` 变 `!== "always"`,3)、模板或 SQL 字符串里的片段(5)、可迭代对象被清空或过滤被删(4)、调用或 `new` 被替换/脱壳(6)、下标移动(`[0]`→`[1]`,2)、初始化器被换成另一个表达式(3)、条件被取反(2)、语句被改写成另一条语句(3)、两条语句挤在一行(2)、字面量被替换(1)、点位在模块顶层因而没有成员可命名(1),以及一颗一次改两处的牙(`every-task-is-frozen-at-position-zero`,故意留着)。这四簇看起来值得加算子(`replace-comparison`、`replace-fragment`、`empty-iterable`、`replace-callee`);其余是单点位,或者用通用算子表达反而会变得有歧义的形状。 + **已推广到整个登记处(2026-09-24):又转了 59 颗,随后补上"现成目录点名的算子",登记处现在是 110 颗里 95 颗 derived。**第一遍:六个算子已经能表达的每一类点位,按目标逐个转完、每转完一个就 sweep——`condition-never`(守卫或初始化器,共 44)、`neutralize-term`(9)、`replace-property`(9)、`replace-argument`(6)、`drop-statement`(6)、`condition-holds`(3)。 + 第二遍:把剩下的 33 颗按"缺什么"分类,再拿现成变异工具真正点名的东西核对(Stryker 的 supported mutators、pitest 的 mutator 列表、cargo-mutants 的 patterns、Cosmic Ray 的 operator 概念)。其中十三颗落在"每个目录都有的变异类"里,于是词汇增加了七个算子,而不是宣布这一类做不到: + + | 算子 | 现成目录里的名字 | 颗数 | + | --------------------- | ------------------------------------------------------------- | ---- | + | `negate-condition` | pitest `NEGATE_CONDITIONALS`;Stryker 布尔字面量 | 2 | + | `negate-comparison` | pitest `NEGATE_CONDITIONALS`;Stryker `EqualityOperator` | 2 | + | `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | + | `remove-call` | Stryker 的 filter/slice/sort 删除;pitest `VOID_METHOD_CALLS` | 2 | + | `replace-call` | Stryker `MethodExpression`;pitest `CONSTRUCTOR_CALLS` | 1 | + | `replace-initializer` | pitest `PRIMITIVE_RETURNS`、`INLINE_CONSTS`;Stryker 字面量类 | 5 | + | `replace-iterable` | Stryker `ArrayDeclaration`;pitest `EMPTY_RETURNS` | 3 | + + 凡是"变异本身决定字节"的(`false`、`true`、取反后的算子、调用的接收者、守卫的 body)都由算子算出来;只有当新值是一个**选择**时才取 mutant 的 `to`——与先前算子同一条规矩。`tests/tools/mutation-anchor.test.ts` 每个算子一条用例,并包含它各自的拒绝情形(有 `else`、body 是块、调用的不是方法、没有初始化器),现在是 30 条。 + 随之而来两处"扩宽",与此前的性质相同——都是**真实的一颗牙转不过去**才暴露出来的点位: + - **绑定到名字的函数也是成员。**`evals/ooo-execution/board-worker.ts` 里 `const count = (label) => {...}` 是该文件逻辑的全部,它不是函数声明,于是那颗牙连作用域都写不出来。选择器现在接受"初始化器是箭头函数或函数表达式的变量"。 + - **两个一模一样的调用要靠"容器"区分。**`dispatchPlan` 里 `board.candidates()` 出现两次(一次 onOffer、一次接 filter),callee 文本完全相同。`in` 现在可以指"该调用所在语句里的一段文本"(它此前已经能指 `replace-property` 的宿主对象字面量、`replace-argument` 的调用自身文本)。 + 第一遍我又写错了一处 `within`,这次是**anchors pass**(不是 sweep)抓出来的:`the-completion-ignores-a-cancelled-unit` 被指向 `checkDispatch` 而不是 `checkCompletion`(同一片段在两个成员里都有)。这是更安静的那种失败:anchors pass 照样通过、mutant 照样被抓——但抓它的是**整个 suite**,不是那条"报告一个被取消单元的完成"的用例。第二遍同类错误反而是响的:我瞄的那个守卫在 `runParentCheck`,我却写了 `instrumentCommit`,pass 直接拒绝——"0 guards in instrumentCommit match the selector"。**成员名写错,找不到时是响的、找到两处同形状时是哑的**,这就是"转完必须 sweep"不可省的理由。 + **两遍之后的读数:**`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,176 秒。 + + **剩下手写的是 15 颗,分类这一次是量出来的、不是猜的:** + - 三颗是**连现成工具也表达不了**的形状:**字符串里的片段**——两条 SQL 子句(`prune-ignores-retention`、`bounded-pin-never-expires`)和一处 SQLite 单位(`the-grace-uses-a-unit-sqlite-does-not-know`)。目录只会替换整个字面量(`StringLiteral` → `""`),而 SQL 有自己的目录(SQLMutation:删子句、`AND` 换 `OR`、NULL 变异),所以这三颗诚实的方向是"子句级 SQL 算子",不是对字节动手术; + - 两颗在**模块顶层**,没有成员可命名:脚本里的一个守卫(`judge-may-judge-its-own-delivery`)和 `export const CLOCK_GRACE_MS = 50`。给 `derive` 加"文件级作用域"是设计决定而不是扩宽,这里不做; + - 两颗动**下标**(`[0]` → `[1]`)。这一颗同时纠正本记录自己此前的说法:之前把"下标移动(2)"列为目录覆盖的类,这是错的——`FirstToLast` 是**名为 `first` 的方法**,我查过的目录里没有哪个变异数组下标; + - 两颗是**两条语句挤在同一行**(`fusion-rollback-not-counted`、`a-score-with-no-recorded-history-counts-as-a-reproduction`)——这是 `drop-statement` 的行纪律问题,不是变异类; + - 一颗把比较改写成范围判断(`=== 1` → `>= 1`),而目录里的取反(`!== 1`)**不会**实现名字所陈述的那个违规; + - 四颗是无算子可依的定制表达式改写:构造调用被换成同接口的手写替身(`the-status-read-path-opens-the-rounds-store`)、多行调用被换成桩事实(`a-coordinated-write-skips-its-run-fact`)、包装被拆到内层调用(`the-daemon-verb-skips-the-coordinated-path`),以及 `view-adapter-accepts-bytes`(两条语句变成一个 `return`); + - 一颗 `every-task-is-frozen-at-position-zero` 一次改两处,故意留着。 + + 其中三簇"值得加算子",作为待决写下:子句级 SQL 变异、文件级选择器、下标算子。这里都不加,因为它们都不是**扩宽**——各自是一条新的作用域规则或一个新领域——而不加它们的代价是量出来的七颗。 + 目录里还有两件事值得带过来: + - **subsumption(被覆盖)。**pitest 的《Less is more》给这件事起了名:一个 mutant 若已被其他 mutant 的组合覆盖,它带来的只是运行时间,不是信心。pitest 靠"固定的默认集合"回避;本登记处是**量**它——sweep 先说出是哪条用例失败,再读那条用例确认它确实陈述了规则。这正是 37 颗牙被退役而不是留着的原因,也是"是不是每种能想到的变异都该进来"的答案是"不"的原因。 + - **unviable mutant(编译不过的变异体)。**cargo-mutants 把这类单独计数、默认不打印;它排除测试函数与 `unsafe`,并且明确写出**为什么不生成某些变异**(`==` 换成 `<`"太容易产生假阳性"、`-a` 换成 `+a`"太容易不编译")。本登记处**按构造就没有** unviable(每颗都有点名的用例),但理由是同一条:为一个在本仓库很少出现的形状加通用算子,产出的多半是噪音。 + 这个设计与那些工具**唯一没有借用**的一点,是"消灭"的定义。StrykerJS、pitest、cargo-mutants 把"任意一条测试失败"就算 killed;其中几个还能报告是哪条测试(`fullMutationMatrix`),但用途是找出该加强哪条断言,不是当门禁。这里只有**点名的那条**用例失败才算数,所以"被整个 suite 抓到、却没被点名用例抓到"是本登记处会报、而那些工具一个都不会报的缺陷。 2. **把 133 颗 G/V/R 牙逐目标改成派生形式**,从 `src/integration/ooo-execution.ts` 开始试点:该文件在两条目标条目里共 24 颗(16 颗融合合法性推演、8 颗会话融合);同样的名字、同样的 `expect` 用例,转换后同样 24/24 被抓——其中 19 颗派生、5 颗保留手写(下面具名)——然后在该文件里做一次真重构,证明派生出来的位置都还在,而字节锚点会在这次重构里退役。 试点文件里那 5 颗残余,每一颗都是单个表达式或单个值,而不是整条语句或消息:`selection-ignores-a-withdrawn-acceptance`(整个 `const` 初始化式被替换)、`the-budget-is-not-cut-from-the-startable-set`(被返回表达式里的调用被删)、`next-task-is-not-the-head-of-the-legal-set`(下标 `[0]` 变 `[1]`)、`speculation-guesses-several-facts-at-once`(`length === 1` 变 `>= 1`)、`a-guess-with-no-evidence-publishes`(插入一个分支并重写消息)。 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index cd9c75df..d94b411e 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -24,6 +24,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth -- --anchors-only` -> `anchors: 144 of 144 resolve, over 23 targets`, exit 0 in 0.7 s (2026-09-24, after five teeth were retired in favour of the checks that already state their rules) - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts` -> `mutants: 15 of 15 caught by the named test`, restored byte-identically (2026-09-24, after six more teeth were retired there: the five budget/publication rules and the run-scoping claim) - `npm run mutation:teeth` (the whole register) -> `mutants: 136 of 136 caught by the named test`, 23 of 23 targets restored byte-identically, exit 0 in 217 s (2026-09-24), **every one of them caught by the case its `expect` names**. The first whole-register reading (231 s, 134 of 136 by name) is what surfaced three wrong links, all repaired in the same pass: `the-caller-rebuilds-the-shared-floor` (its `expect` was a paraphrase of no case, re-pointed), `the-grace-is-zero` (its named case derived its fixture from the constant under test, so it passed while the bug was live - the fixture is a literal now), and `next-is-not-the-head-of-the-ordered-candidates` (re-pointed to `at the default budget the licence is still the head of the ordered set`, which now asserts the head). A fourth was found by the pilot sweep and repaired here: `fusion-continues-from-an-unverified-answer`'s case refused the pair for a second reason as well, so it passed under the mutant - with no dependency between the two units it is refused by that one condition only. One row still carries a note rather than an assertion: `the-pass-asks-a-unit-it-already-failed-again` is caught by a named case that does not finish inside its 30 s bound, which this tool counts as caught and prints with that reason. +- `npm run mutation:teeth` (the whole register, after the operators the catalogues name were added and 18 more teeth converted) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 176 s (2026-09-24); `npm run mutation:anchors` -> `anchors: 110 of 110 resolve, over 22 targets`; `tests/tools/mutation-anchor.test.ts` -> 30 pass. - `npm run mutation:teeth` (the whole register, after 59 teeth were converted to derived selectors) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 171 s (2026-09-24). The conversion is not what makes them bite; it is what keeps them aimed: one wrong `within` was shipped in this pass and the sweep is what caught it, as a tooth caught by the suite rather than by the case that names its rule. - `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were **orphans**: no row named `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` or `the-parent-check-ignores-its-own-verdict`, so their rules have a case but no row; the check each case makes is in the [record](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) @@ -39,9 +40,9 @@ description rather than a pin. Every mutant name below was read from A tooth is a **name, an operator and a selector** - not a copy of a line: the site and the replacement bytes are resolved from the syntax tree on every run, so a rename or a reflow does not retire the pin while the rule still stands ([the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). -77 of the register's 110 teeth are in that form now; the 33 that are not are the shapes the six operators -cannot express, clustered and counted in the same record, with the four clusters that look worth an -operator named there. `npm run mutation:teeth -- --anchors-only` answers whether every +95 of the register's 110 teeth are in that form now; the 15 that are not are the shapes the operators +cannot express, clustered and counted in the same record, with the three clusters that would repay an +operator named there (clause-level SQL mutation, a file-scope selector, an index operator). `npm run mutation:teeth -- --anchors-only` answers whether every anchor still applies - under a second, no suite run, nothing written - because a tooth that stopped matching its rule is otherwise only visible as a check that stopped counting. diff --git a/tests/tools/mutation-anchor.test.ts b/tests/tools/mutation-anchor.test.ts index d1516b0d..7bfe91b0 100644 --- a/tests/tools/mutation-anchor.test.ts +++ b/tests/tools/mutation-anchor.test.ts @@ -320,3 +320,265 @@ test("matchText still refuses a marker that occurs twice, and re-takes a reflowe const reflowed = matchText("a &&\n b", "a && b"); assert.ok(!("reason" in reflowed) && reflowed.retaken); }); + +test("a negated condition is the negation, and the `!` on a whole condition is the one that goes", () => { + const above = `function selection(binding, cancelled) {\n if (!binding) return apply();\n if (binding && cancelled) return null;\n return binding;\n}\n`; + assert.equal( + selected( + above, + named("x", { + within: "selection", + operator: "negate-condition", + condition: "!binding", + }), + ), + "!", + ); + assert.equal( + selected( + above, + named("x", { + within: "selection", + operator: "negate-condition", + condition: "binding && cancelled", + }), + ), + "binding && cancelled", + ); +}); + +test("the negated condition is wrapped, not unwrapped, when the `!` binds one operand only", () => { + const above = `function selection(a, b) {\n if (!a || b) return 1;\n return 0;\n}\n`; + const site = locate( + above, + named("x", { within: "selection", operator: "negate-condition", condition: "!a || b" }), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "!a || b"); + // The replacement is the whole condition wrapped: dropping the `!` would negate `a` alone. + assert.equal(site.replacement, "!(!a || b)"); +}); + +test("a comparison is negated in place, keeping both sides and the spacing around them", () => { + const above = `function selection(verdict) {\n if (verdict !== "accepted") return false;\n return true;\n}\n`; + const site = locate( + above, + named("x", { + within: "selection", + operator: "negate-comparison", + condition: 'verdict !== "accepted"', + }), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "!=="); + assert.equal(site.replacement, "==="); + assert.match( + refusal( + above, + named("x", { + within: "selection", + operator: "negate-comparison", + condition: 'verdict === "absent"', + }), + ), + /0 comparisons in selection match the selector/, + ); +}); + +test("a dropped guard keeps its body, and refuses an else or a block body rather than rewriting it", () => { + const above = `function selection(spec, path, content, files) {\n if (spec.baseline[path] !== content) files[path] = content;\n return files;\n}\n`; + assert.equal( + selected( + above, + named("x", { + within: "selection", + operator: "remove-conditionals", + condition: "spec.baseline[path] !== content", + }), + ), + "if (spec.baseline[path] !== content) files[path] = content;", + ); + const withElse = `function selection(a, b) {\n if (a) b();\n else c();\n return a;\n}\n`; + assert.match( + refusal( + withElse, + named("x", { within: "selection", operator: "remove-conditionals", condition: "a" }), + ), + /has an else/, + ); + const withBlock = `function selection(a, b) {\n if (a) {\n b();\n }\n return a;\n}\n`; + assert.match( + refusal( + withBlock, + named("x", { within: "selection", operator: "remove-conditionals", condition: "a" }), + ), + /body in selection is a block/, + ); +}); + +test("removing a call leaves its receiver, and a call that is not a method is refused", () => { + const above = `function selection(board, attempted, values) {\n const legal = board.candidates().filter((id) => !attempted.has(id));\n const plain = Number(values);\n return [legal, plain];\n}\n`; + assert.equal( + selected( + above, + named("x", { + within: "selection", + operator: "remove-call", + call: "board.candidates().filter", + }), + ), + "board.candidates().filter((id) => !attempted.has(id))", + ); + assert.match( + refusal(above, named("x", { within: "selection", operator: "remove-call", call: "Number" })), + /needs a method call/, + ); +}); + +test("a call is replaced by the declared source of the same shape", () => { + const above = `function selection(board, input) {\n const legal = board.candidates().filter(Boolean);\n return legal;\n}\n`; + const site = locate( + above, + named( + "x", + { within: "selection", operator: "replace-call", call: "board.candidates" }, + { to: "input.plan" }, + ), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "board.candidates()"); + assert.equal(site.replacement, "input.plan"); +}); + +test("two calls to the same callee are told apart by the statement each sits in", () => { + const above = `function selection(board, attempted) { + const onOffer = board.candidates(); + const legal = board.candidates().filter((id) => !attempted.has(id)); + return [onOffer, legal]; +} +`; + assert.match( + refusal( + above, + named( + "x", + { within: "selection", operator: "replace-call", call: "board.candidates" }, + { to: "input.plan" }, + ), + ), + /2 calls to board.candidates in selection match the selector/, + ); + const site = locate( + above, + named( + "x", + { + within: "selection", + operator: "replace-call", + call: "board.candidates", + in: "board.candidates().filter", + }, + { to: "input.plan" }, + ), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "board.candidates()"); +}); + +test("a variable's value is replaced by the declared one, and an absent one is refused", () => { + const above = `function selection(shape) {\n const sessions = Math.ceil(shape.units / per);\n return sessions;\n}\n`; + const site = locate( + above, + named( + "x", + { within: "selection", operator: "replace-initializer", variable: "sessions" }, + { to: "shape.units" }, + ), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "Math.ceil(shape.units / per)"); + const absent = `function selection(shape) {\n const sessions = shape.units;\n return sessions;\n}\n`; + assert.match( + refusal( + absent, + named( + "x", + { within: "selection", operator: "replace-initializer", variable: "count" }, + { to: "0" }, + ), + ), + /0 declarations of count/, + ); +}); + +test("a loop's iterable is replaced by the declared one, and the loop is named when there are several", () => { + const above = `function selection(budgets, other) {\n const seen = [];\n for (const budget of budgets) {\n seen.push(budget);\n }\n for (const item of other) {\n seen.push(item);\n }\n return seen;\n}\n`; + assert.equal( + selected( + above, + named( + "x", + { within: "selection", operator: "replace-iterable", iterable: "of budgets" }, + { to: "[]" }, + ), + ), + "budgets", + ); + assert.match( + refusal( + above, + named( + "x", + { within: "selection", operator: "replace-iterable", iterable: "of" }, + { to: "[]" }, + ), + ), + /2 loops in selection match the selector/, + ); +}); + +test("a call is told from another by a fragment of its own text", () => { + const above = `function selection() {\n const a = clockNow("later") + clockNow("earlier");\n return a;\n}\n`; + assert.equal( + selected( + above, + named( + "x", + { + within: "selection", + operator: "replace-argument", + call: "clockNow", + arg: 0, + in: '"later"', + }, + { to: '"earlier"' }, + ), + ), + '"later"', + ); + assert.match( + refusal( + above, + named( + "x", + { within: "selection", operator: "replace-argument", call: "clockNow", arg: 0 }, + { to: '"earlier"' }, + ), + ), + /matched 2 sites/, + ); +}); + +test("a function bound to a name is a member, and the selectors search its body", () => { + const above = `const selection = (values) => {\n const spec = new RegExp("x").exec(values);\n return spec;\n};\n`; + const site = locate( + above, + named( + "x", + { within: "selection", operator: "replace-initializer", variable: "spec" }, + { to: "null" }, + ), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), 'new RegExp("x").exec(values)'); +}); diff --git a/tools/mutation-anchor.ts b/tools/mutation-anchor.ts index d907a8a6..afbcc0f0 100644 --- a/tools/mutation-anchor.ts +++ b/tools/mutation-anchor.ts @@ -45,27 +45,56 @@ export interface Derive { * `condition-never`, because a rule written as `return a && b` says "this holds" - replacing it * with `false` would reverse the rule instead of removing it. */ | "condition-holds" + /** The condition the selector identifies becomes its negation. The operator pitest calls Negate + * Conditionals: a rule that refuses when the condition holds now refuses when it does not. */ + | "negate-condition" + /** The comparison the selector identifies gets the negated comparison operator (`===` becomes + * `!==`, `<` becomes `>=`). The narrower sibling of `negate-condition`, for a rule whose point is + * the comparison itself. */ + | "negate-comparison" + /** The guard the selector identifies is dropped and its body is kept, so the guarded code runs + * whatever the condition said. pitest's Remove Conditionals. */ + | "remove-conditionals" + /** The call the selector identifies becomes its receiver, so the method's work is skipped: a + * `filter` that no longer filters, a `slice` that no longer cuts. pitest's Void Method Calls and + * Stryker's filter/slice/sort removals. */ + | "remove-call" + /** The call the selector identifies becomes the declared `to` fragment, for a call replaced by a + * different source of the same shape. */ + | "replace-call" /** One term of that condition becomes its identity: `true` under `&&`, `false` under `||`. */ | "neutralize-term" /** Argument `arg` of the call to `call` becomes the declared `to` fragment. */ | "replace-argument" /** The value of the object property named `property` becomes the declared `to` fragment. */ | "replace-property" + /** The initializer of the variable named `variable` becomes the declared `to` fragment: the value + * bound here is a different one, which is what `primitive returns` and `inline constant` mutate. */ + | "replace-initializer" + /** The iterable of the loop the selector identifies becomes the declared `to` fragment, so the + * loop runs over an empty or a shortened collection. */ + | "replace-iterable" /** The statement the selector identifies is removed, with its line and its indentation: an * expression, a declaration, a `return`, a `throw`, or a guard clause. */ | "drop-statement"; - /** Which guard: a fragment its condition's own text contains, e.g. `existing.deliveredBy`. */ + /** Which guard or which comparison: a fragment its own text contains, e.g. `existing.deliveredBy`. */ readonly condition?: string; /** Which term of it, for `neutralize-term`: the term's own text, e.g. `!task.cancelled`. */ readonly term?: string; - /** Which call, for `replace-argument`: its callee as written, e.g. `startableTasks`. */ + /** Which call, for `replace-argument`, `remove-call` and `replace-call`: its callee as written, + * e.g. `startableTasks`, or `board.candidates().filter` when the receiver is what tells it apart. */ readonly call?: string; /** Which argument, for `replace-argument`, counting from zero. */ readonly arg?: number; /** Which property, for `replace-property`: its name as written, e.g. `sessionReusable`. */ readonly property?: string; - /** For `replace-property`: a fragment of the object literal that holds it, when the same property - * name is written in several literals of one member and only one of them is the site. */ + /** Which variable, for `replace-initializer`: its name as written. */ + readonly variable?: string; + /** Which loop, for `replace-iterable`: a fragment of the loop's own text, e.g. `of unknownBudget`. */ + readonly iterable?: string; + /** The holder that tells two same-named sites apart: for `replace-property`, a fragment of the + * object literal that holds it; for `replace-argument`, a fragment of the call's own text; for + * `remove-call` and `replace-call`, a fragment of the statement the call sits in. */ readonly in?: string; /** Which statement, for `drop-statement`: a fragment of the statement's own text. */ readonly statement?: string; @@ -214,18 +243,30 @@ function collect(root: ts.Node, isWanted: (node: ts.Node) => /** The one member with this name in the file, or why it is not one site. A class constructor is a * member too: its name is written `constructor`, and code that refuses a second plan while it opens - * the store lives there and nowhere else. */ + * the store lives there and nowhere else. So is a function bound to a name - `const count = (label) + * => ...` names one function, and whether the formatter wrote it as a declaration is not a fact about + * which rules live inside it. */ function uniqueMember(source: ts.SourceFile, name: string): ts.Node | { reason: string } { + const boundFunction = (node: ts.Node): boolean => + ts.isVariableDeclaration(node) && + node.name.getText(source) === name && + node.initializer !== undefined && + (ts.isArrowFunction(node.initializer) || ts.isFunctionExpression(node.initializer)); const named = (node: ts.Node): boolean => ((ts.isMethodDeclaration(node) || ts.isFunctionDeclaration(node)) && node.name?.getText(source) === name) || + boundFunction(node) || (ts.isConstructorDeclaration(node) && name === "constructor"); const members = collect(source, named); if (members.length !== 1) return { reason: `member ${name} matched ${members.length} members, refusing to claim a check`, }; - return members[0]!; + const member = members[0]!; + // A name bound to a function is the function's body for every purpose here, so the selectors search + // the body rather than the declaration that wraps it. + if (ts.isVariableDeclaration(member) && member.initializer) return member.initializer; + return member; } /** Argument `arg` of the one call to `call` inside the member. */ @@ -240,7 +281,10 @@ function argumentSite( if (to === undefined) return { reason: "replace-argument needs the mutant's `to` fragment" }; const calls = collect( member, - (node) => ts.isCallExpression(node) && node.expression.getText(source) === derive.call, + (node) => + ts.isCallExpression(node) && + node.expression.getText(source) === derive.call && + (derive.in === undefined || containsText(node.getText(source), derive.in)), ); if (calls.length !== 1) return { @@ -480,6 +524,292 @@ function neutralizedTerm( }; } +/** What every resolver is given: the parsed file, the one member, the selector, and the bytes a + * value-substituting operator declares. */ +interface Context { + readonly source: ts.SourceFile; + readonly text: string; + readonly member: ts.Node; + readonly derive: Derive; + readonly to: string | undefined; +} + +type Resolver = (context: Context) => Site | { reason: string }; + +/** The inner sites of one collection, with the nested ones dropped: a guard whose body is a statement + * is the guard and not the two of them, and the same rule serves conditions, statements, comparisons + * and guards. */ +function innermost(source: ts.SourceFile, nodes: readonly T[]): T[] { + return nodes.filter( + (node) => + !nodes.some( + (other) => + other !== node && + other.getStart(source) >= node.getStart(source) && + other.getEnd() <= node.getEnd(), + ), + ); +} + +/** The one collection site the fragment identifies, or why there is not one. */ +function oneSite( + source: ts.SourceFile, + derive: Derive, + nodes: readonly T[], + what: string, +): T | { reason: string } { + const inner = innermost(source, nodes); + if (inner.length !== 1) + return { + reason: `${inner.length} ${what} in ${derive.within} match the selector, refusing to claim a check`, + }; + return inner[0]!; +} + +/** Every comparison operator, and the one that reverses it. */ +const NEGATED_COMPARISON: Readonly> = { + "===": "!==", + "!==": "===", + "==": "!=", + "!=": "==", + "<": ">=", + "<=": ">", + ">": "<=", + ">=": "<", +}; + +/** The condition the selector identifies becomes its negation. A condition that is already a whole + * negation drops its `!` - which is exactly the negation - and anything else is wrapped, because + * removing a `!` that binds only the first operand (`!a || b`) would not negate the whole. */ +function negateCondition(context: Context): Site | { reason: string } { + const condition = conditionSite(context.source, context.member, context.derive); + if ("reason" in condition) return condition; + if ( + ts.isPrefixUnaryExpression(condition) && + condition.operator === ts.SyntaxKind.ExclamationToken + ) + return { + start: condition.getStart(context.source), + end: condition.operand.getStart(context.source), + replacement: "", + retaken: false, + }; + return { + start: condition.getStart(context.source), + end: condition.getEnd(), + replacement: `!(${condition.getText(context.source)})`, + retaken: false, + }; +} + +/** The comparison the selector identifies gets the operator that reverses it. */ +function negateComparison(context: Context): Site | { reason: string } { + const { source, derive } = context; + if (derive.condition === undefined) + return { reason: "negate-comparison needs a `condition` fragment naming the comparison" }; + const fragment = derive.condition; + const comparing = (node: ts.Node): boolean => + ts.isBinaryExpression(node) && + NEGATED_COMPARISON[node.operatorToken.getText(source)] !== undefined && + containsText(node.getText(source), fragment); + const found = oneSite( + source, + derive, + collect(context.member, comparing), + "comparisons", + ); + if ("reason" in found) return found; + const written = found.operatorToken.getText(source); + const after = found.left.getEnd(); + const gap = context.text.slice(after, found.right.getStart()); + const at = gap.indexOf(written); + if (at < 0) + return { reason: `cannot read ${written} in ${derive.within}, refusing to claim a check` }; + return { + start: after + at, + end: after + at + written.length, + replacement: NEGATED_COMPARISON[written]!, + retaken: false, + }; +} + +/** The guard the selector identifies goes, and its body stays: the guarded code runs whatever the + * condition said. A guard with an `else` would have to choose a branch, and a block body would have + * to be unindented, so both are refused rather than rewritten. */ +function removeConditionals(context: Context): Site | { reason: string } { + const { source, derive } = context; + if (derive.condition === undefined) + return { reason: "remove-conditionals needs a `condition` fragment naming the guard" }; + const fragment = derive.condition; + const found = oneSite( + source, + derive, + collect( + context.member, + (node) => ts.isIfStatement(node) && containsText(node.expression.getText(source), fragment), + ), + "guards", + ); + if ("reason" in found) return found; + if (found.elseStatement) + return { reason: `the guard in ${derive.within} has an else, refusing to choose a branch` }; + if (ts.isBlock(found.thenStatement)) + return { + reason: `the guard's body in ${derive.within} is a block, refusing to rewrite its indentation`, + }; + return { + start: found.getStart(source), + end: found.getEnd(), + replacement: found.thenStatement.getText(source), + retaken: false, + }; +} + +/** The text of the nearest statement a node sits in. Two identical expressions can be written in one + * member - `board.candidates()` on offer and the same call filtered - and the statement that holds + * each is what tells them apart without naming a line. */ +function holderStatementText(source: ts.SourceFile, node: ts.Node): string { + let holder: ts.Node = node; + while (holder.parent && !ts.isStatement(holder)) holder = holder.parent; + return ts.isStatement(holder) ? holder.getText(source) : node.getText(source); +} + +/** The one call the selector names, by the callee as written. */ +function callSite(context: Context): ts.CallExpression | { reason: string } { + const { source, derive } = context; + if (derive.call === undefined) return { reason: `${derive.operator} needs a \`call\` fragment` }; + const found = oneSite( + source, + derive, + collect( + context.member, + (node) => + ts.isCallExpression(node) && + node.expression.getText(source) === derive.call && + (derive.in === undefined || containsText(holderStatementText(source, node), derive.in)), + ), + `calls to ${derive.call}`, + ); + if ("reason" in found) return found; + return found; +} + +/** The call becomes its receiver: the method's work is skipped, and what it was called on is what is + * left. Only a method call has a receiver to fall back to, so anything else is refused. */ +function removeCall(context: Context): Site | { reason: string } { + const call = callSite(context); + if ("reason" in call) return call; + if (!ts.isPropertyAccessExpression(call.expression)) + return { reason: `remove-call needs a method call, and ${context.derive.call} is not one` }; + return { + start: call.getStart(context.source), + end: call.getEnd(), + replacement: call.expression.expression.getText(context.source), + retaken: false, + }; +} + +/** The call becomes the declared fragment: a different source of the same shape. */ +function replaceCall(context: Context): Site | { reason: string } { + if (context.to === undefined) return { reason: "replace-call needs the mutant's `to` fragment" }; + const call = callSite(context); + if ("reason" in call) return call; + return { + start: call.getStart(context.source), + end: call.getEnd(), + replacement: context.to, + retaken: false, + }; +} + +/** The initializer of the one variable the selector names becomes the declared fragment. */ +function replaceInitializer(context: Context): Site | { reason: string } { + const { source, derive } = context; + if (derive.variable === undefined) + return { reason: "replace-initializer needs a `variable` name" }; + if (context.to === undefined) + return { reason: "replace-initializer needs the mutant's `to` fragment" }; + const found = oneSite( + source, + derive, + collect( + context.member, + (node) => ts.isVariableDeclaration(node) && node.name.getText(source) === derive.variable, + ), + `declarations of ${derive.variable}`, + ); + if ("reason" in found) return found; + const initializer = found.initializer; + if (!initializer) + return { reason: `${derive.variable} has no initializer, refusing to claim a check` }; + return { + start: initializer.getStart(source), + end: initializer.getEnd(), + replacement: context.to, + retaken: false, + }; +} + +/** The iterable of the one loop the selector identifies becomes the declared fragment: the loop body + * runs over a collection that is empty, or shorter than the one the code meant. */ +function replaceIterable(context: Context): Site | { reason: string } { + const { source, derive } = context; + if (derive.iterable === undefined) + return { reason: "replace-iterable needs an `iterable` fragment naming the loop" }; + if (context.to === undefined) + return { reason: "replace-iterable needs the mutant's `to` fragment" }; + const fragment = derive.iterable; + const found = oneSite( + source, + derive, + collect( + context.member, + (node) => ts.isForOfStatement(node) && containsText(node.getText(source), fragment), + ), + "loops", + ); + if ("reason" in found) return found; + return { + start: found.expression.getStart(source), + end: found.expression.getEnd(), + replacement: context.to, + retaken: false, + }; +} + +/** The operators, each one a function of the selector and the declared bytes. Adding an operator means + * adding an entry here and its name to `Derive["operator"]`: nothing else in the register changes. */ +const RESOLVERS: Readonly> = { + "condition-never": (context) => { + const condition = conditionSite(context.source, context.member, context.derive); + if ("reason" in condition) return condition; + return wholeCondition(context.source, condition, context.derive); + }, + "condition-holds": (context) => { + const condition = conditionSite(context.source, context.member, context.derive); + if ("reason" in condition) return condition; + return wholeCondition(context.source, condition, context.derive); + }, + "negate-condition": negateCondition, + "negate-comparison": negateComparison, + "remove-conditionals": removeConditionals, + "remove-call": removeCall, + "replace-call": replaceCall, + "replace-initializer": replaceInitializer, + "replace-iterable": replaceIterable, + "neutralize-term": (context) => { + const condition = conditionSite(context.source, context.member, context.derive); + if ("reason" in condition) return condition; + return neutralizedTerm(context.source, condition, context.derive); + }, + "replace-argument": (context) => + argumentSite(context.source, context.member, context.derive, context.to), + "replace-property": (context) => + propertySite(context.source, context.member, context.derive, context.to), + "drop-statement": (context) => + statementSite(context.source, context.member, context.derive, context.text), +}; + /** * Resolve a derived mutant: the selector picks the code, the operator says what to do to it, and the * bytes are computed here rather than stored. Every selector is scoped to one member, and every @@ -494,12 +824,5 @@ function deriveSite( const source = ts.createSourceFile("mutant.ts", text, ts.ScriptTarget.Latest, true); const member = uniqueMember(source, derive.within); if ("reason" in member) return member; - if (derive.operator === "replace-argument") return argumentSite(source, member, derive, to); - if (derive.operator === "replace-property") return propertySite(source, member, derive, to); - if (derive.operator === "drop-statement") return statementSite(source, member, derive, text); - const condition = conditionSite(source, member, derive); - if ("reason" in condition) return condition; - return derive.operator === "neutralize-term" - ? neutralizedTerm(source, condition, derive) - : wholeCondition(source, condition, derive); + return RESOLVERS[derive.operator]({ source, text, member, derive, to }); } diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 1ed71582..5d01fd04 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -32,14 +32,17 @@ * * Where a mutant says where it applies: * - * - `derive: { within, operator, ... }` is a **selector plus an operator**: the member is named - * (`constructor` when the site is in a class constructor), a - * fragment inside it is matched (the same matcher as below, so a reflow cannot break it), and the - * bytes to write are computed on every run. This is the form a tooth should have. Its identity is - * the rule and the site it names, never the text that happens to be there today, and it is what - * lets a rename, a reordered condition or a lifted line leave the tooth aimed at the same rule. - * Selectors that match more than one site - or a fragment that fits two conditions - are refused, - * never guessed at. + * - `derive: { within, operator, ... }` is a **selector plus an operator**: the member is named (a + * method, a function declaration, a class constructor under the name `constructor`, or a function + * bound to a variable), a fragment inside it is matched (the same matcher as below, so a reflow + * cannot break it), and the bytes to write are computed on every run. This is the form a tooth + * should have. Its identity is the rule and the site it names, never the text that happens to be + * there today, and it is what lets a rename, a reordered condition or a lifted line leave the tooth + * aimed at the same rule. Selectors that match more than one site - or a fragment that fits two + * conditions - are refused, never guessed at. The operators are the mutation classes the tools and + * the literature name (conditionals-to-false/true, negate conditionals, remove conditionals, method + * call removal, constant and collection substitution), kept deliberately few: see + * `tools/mutation-anchor.ts` for each one's selector and the catalogue it comes from. * - `ast: { within: "" }` (or `ast: { call, argCount }`) locates the site through the syntax * tree. Use this in any file that is still being edited. It survives reformatting, and it refuses * when the code it guards has moved out of the member it belongs to - a move that a byte anchor @@ -86,8 +89,11 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "any-verdict-counts-as-accepted", - from: ' if (fact.verdict !== "accepted") return false;', - to: " if (fact.verdict === null) return false;", + derive: { + within: "acceptedFact", + operator: "negate-comparison", + condition: 'fact.verdict !== "accepted"', + }, expect: "acceptedFact is the single rule, and the view adapter does not grow its own", }, { @@ -101,14 +107,22 @@ const TARGETS: readonly Target[] = [ }, { name: "spec-level-alias-is-dropped", - from: " for (const key of unknownKeys(spec, SPEC_FIELDS)) {", - to: " for (const key of [] as string[]) {", + derive: { + within: "refuseSpecFields", + operator: "replace-iterable", + iterable: "of unknownKeys(spec, SPEC_FIELDS)", + }, + to: "[] as string[]", expect: "refuses the aliases that would turn a byte budget into a token claim", }, { name: "budget-inner-alias-is-dropped", - from: " for (const key of unknownBudget) {", - to: " for (const key of [] as string[]) {", + derive: { + within: "refuseOutOfRange", + operator: "replace-iterable", + iterable: "of unknownBudget", + }, + to: "[] as string[]", expect: "refuses an unknown key inside budget or limits, and a range above the maximum", }, { @@ -394,9 +408,8 @@ const TARGETS: readonly Target[] = [ }, { name: "selection-ignores-a-withdrawn-acceptance", - ast: { within: "selection" }, - from: " const selectable = plan.filter((task) => task.accepted || !task.delivered);", - to: " const selectable = plan;", + derive: { within: "selection", operator: "replace-initializer", variable: "selectable" }, + to: "plan", expect: "an outside rejection withdraws the release of a dependent, and the round fails closed", }, @@ -426,12 +439,8 @@ const TARGETS: readonly Target[] = [ expect: "a declared budget is spent by claims in flight, not by the next task's rank", }, { - // The cut is the whole point of declaring slots: without it the budget is a comment, and a - // run that asked for two would start the whole legal set. name: "the-budget-is-not-cut-from-the-startable-set", - ast: { within: "startableTasks" }, - from: " return legal.slice(0, room);", - to: " return legal;", + derive: { within: "startableTasks", operator: "remove-call", call: "legal.slice" }, expect: "a declared budget is spent by claims in flight, not by the next task's rank", }, { @@ -538,12 +547,13 @@ const TARGETS: readonly Target[] = [ expect: "the continuation is a declared constraint, not the planner's default", }, { - // A name the protocol does not define is refused by name; reading it as "nothing was asked" - // is what would make the declaration decoration. name: "an-unknown-constraint-is-ignored", - ast: { within: "enabledConstraints" }, - from: " const unknown = named.filter((name) => !(PLAN_CONSTRAINTS as readonly string[]).includes(name));", - to: " const unknown: readonly string[] = [];", + derive: { + within: "enabledConstraints", + operator: "replace-initializer", + variable: "unknown", + }, + to: "[] as readonly string[]", expect: "a constraint the protocol does not define is refused by name, never ignored", }, ], @@ -553,12 +563,9 @@ const TARGETS: readonly Target[] = [ suites: ["tests/integration/ooo-publication-invariants.test.ts"], mutants: [ { - // Every declared budget publishes its own set, and the ones a one-slot run cannot offer are - // the whole point of declaring more. Deriving every prefix at one slot hides them. name: "every-budget-is-walked-at-one-slot", - ast: { within: "budgetViews" }, - from: " for (const budget of budgets) {", - to: " for (const budget of budgets.slice(0, 1)) {", + derive: { within: "budgetViews", operator: "replace-iterable", iterable: "of budgets" }, + to: "budgets.slice(0, 1)", expect: "a declared budget publishes more than one slot can, and never a claimed task", }, { @@ -669,8 +676,8 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "worker-reads-only-one-reporter-shape", - from: ' const spec = new RegExp(`^\u2139 ${label} (\\\\d+)$`, "m").exec(body);', - to: " const spec = null;", + derive: { within: "count", operator: "replace-initializer", variable: "spec" }, + to: "null", expect: "the worker claims, runs the named suite and delivers its digest", }, ], @@ -750,11 +757,14 @@ const TARGETS: readonly Target[] = [ }, { - // What may run is the board's answer, not the plan's order: dispatching the declared plan - // instead would run a unit whose dependencies are not accepted yet. name: "the-loop-dispatches-the-plan-instead-of-what-the-board-offers", - from: " const legal = board.candidates().filter((id) => !attempted.has(id));", - to: " const legal = input.plan.filter((id) => !attempted.has(id));", + derive: { + within: "dispatchPlan", + operator: "replace-call", + call: "board.candidates", + in: "board.candidates().filter", + }, + to: "input.plan", expect: "the loop runs what the board offers, in the order the board offers it", }, @@ -780,11 +790,12 @@ const TARGETS: readonly Target[] = [ expect: "a worker that starts its own session is not reported as fusion", }, { - // One pass takes each unit at most once: without the record of what was attempted, a unit the - // board offers again after a failed worker is asked for again and again in the same pass. name: "the-pass-asks-a-unit-it-already-failed-again", - from: " const legal = board.candidates().filter((id) => !attempted.has(id));", - to: " const legal = board.candidates();", + derive: { + within: "dispatchPlan", + operator: "remove-call", + call: "board.candidates().filter", + }, expect: "a unit whose worker failed is asked once in a pass", }, { @@ -807,18 +818,18 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "a-unit-ignores-the-checks-it-declares", - from: " const checks = unit.checks ? checkList(unit.checks) : fallback;", - to: " const checks = fallback;", + derive: { within: "specFrom", operator: "replace-initializer", variable: "checks" }, + to: "fallback", expect: "report: both plans accept the instrument's answers, and the same composed ones", }, { - // A submitted patch carries the unit's whole frozen view, so composing by overwriting the - // candidate with each accepted submission puts the *last* unit's untouched copies of its - // siblings over the work they did. Only the files a unit changed are its work. name: "the-parent-takes-the-last-units-whole-view", - from: " if (spec.baseline[path] !== content) files[path] = content;", - to: " files[path] = content;", + derive: { + within: "runParentCheck", + operator: "remove-conditionals", + condition: "spec.baseline[path] !== content", + }, expect: "report: both plans accept the instrument's answers, and the same composed ones", }, ], @@ -922,11 +933,12 @@ const TARGETS: readonly Target[] = [ ], mutants: [ { - // Binding is what makes an entry managed, so the refusal has to read the stored fact rather - // than whatever the caller believes about the entry. name: "an-adopted-entry-takes-the-direct-path", - from: " if (!binding) return request.apply();", - to: " if (binding) return request.apply();", + derive: { + within: "coordinatedEntryWrite", + operator: "negate-condition", + condition: "!binding", + }, expect: "the routing rule sends a managed entry to its run and leaves an unmanaged one alone", }, @@ -971,12 +983,12 @@ const TARGETS: readonly Target[] = [ expect: "a coordinated write lands the board transition and the run's fact together", }, { - // A cancelled run is the end of its managed entries' lifecycle, and the fence is the only - // thing that says so. name: "a-cancelled-run-still-accepts-writes", - ast: { within: "managedWriteRefusal" }, - from: " if (!cancelled) return null;", - to: " if (cancelled) return null;", + derive: { + within: "managedWriteRefusal", + operator: "negate-condition", + condition: "!cancelled", + }, expect: "a cancelled run takes no further lifecycle writes on what it adopted", }, { @@ -1110,11 +1122,12 @@ const TARGETS: readonly Target[] = [ expect: "a route that declares the shared checks not applicable narrows to its own tests", }, { - // An absent declaration means "always". Reading anything that is not the explicit - // "always" as a decline would silently drop the floor for every route that never asked. name: "an-undeclared-route-is-read-as-declining", - from: 'route.verify.sharedChecks === "none"', - to: 'route.verify.sharedChecks !== "always"', + derive: { + within: "planNarrowVerify", + operator: "negate-comparison", + condition: 'route.verify.sharedChecks === "none"', + }, expect: "a change cleanly owned by one leaf route narrows to its own tests", }, ], @@ -1252,8 +1265,12 @@ const TARGETS: readonly Target[] = [ }, { name: "fusion-counts-one-session-per-unit", - from: " const sessions = Math.ceil(shape.units / per);", - to: " const sessions = shape.units;", + derive: { + within: "fusionAccounting", + operator: "replace-initializer", + variable: "sessions", + }, + to: "shape.units", expect: "fusion books the shared startup once per session, not once per unit", }, { @@ -1286,14 +1303,20 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "the-window-does-not-grace-valid-from", - from: ' `((${alias}.valid_from IS NULL OR ${alias}.valid_from <= ${clockNow("later")})` +', - to: ' `((${alias}.valid_from IS NULL OR ${alias}.valid_from <= ${clockNow("earlier")})` +', + derive: { + within: "currentlyValid", + operator: "replace-argument", + call: "clockNow", + arg: 0, + in: '"later"', + }, + to: '"earlier"', expect: "a value stamped a moment in the future is current, not missing", }, { name: "the-window-does-not-grace-expiry", - from: ' return `(${alias}.expires_at IS NULL OR ${alias}.expires_at > ${clockNow("earlier")})`;', - to: ' return `(${alias}.expires_at IS NULL OR ${alias}.expires_at > ${clockNow("later")})`;', + derive: { within: "notExpired", operator: "replace-argument", call: "clockNow", arg: 0 }, + to: '"later"', expect: "the freshness half widens one grace into the past and no further", }, { From 76912913c2936c2361a8613d9861d9fc6ff41164 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 22:58:44 +0800 Subject: [PATCH 31/32] Every tooth is derived now: the last 15 convert, and the register has no hand-written anchor left Asked why the remaining 15 were not converted, the honest answer turned out to be that they could be - and that three of the reasons recorded for leaving them were wrong. Measured, one at a time, each conversion keeping its name and its named case: - the file scope: `within` is optional, and without it the whole file is the scope with the selector required to be unique in it (a top-level `CLOCK_GRACE_MS`, a guard in a module-level script); - `replace-index`: the element access the selector identifies gets the declared index (2 teeth); - `replace-literal-fragment`: a piece of text written inside a literal becomes the declared fragment, with the same holder-statement disambiguator the call selectors have, because one SQL predicate is written seven times in one member (3 teeth: two clauses and one SQLite time unit); - the call selectors accept `new X(...)` as a call site (1 tooth, the round store opened as a constructor call); - `uniqueMember` accepts a function bound to an object property, `claim: (store, parsed) => ...` (1 tooth, the daemon's handler table); - `replace-property` writes a shorthand property out, `position` to `position: 0`, because a value cannot go where the name was (1 tooth); - five more converted with operators that already existed: a call replaced by a boolean, a value replaced by a stub fact, a wrapper unwrapped to its inner call, an iterable shortened, a comparison negated. Three recorded reasons were wrong, and the sweep is what said so: - "the catalogue's negation would not realise the violation" (speculation-guesses-several-facts-at-once): it does, and the named case fails under `!== 1`; - "it changes two things at once" (every-task-is-frozen-at-position-zero): only one of the two is load-bearing, because `noUnusedParameters` is not set, so the callback keeps the parameter it no longer reads; - "two statements sharing one line" (two teeth): the hand anchors had to span two statements, but the mutations are single-statement deletions and both convert to `drop-statement`. Two mistakes made and caught inside this pass, both worth keeping: - a hand-typed `expect` reads as a broken mutant. Retyping two case names from memory instead of copying them made both teeth report "survived" - the tool filters the suite by the named case, so a name that matches no case is indistinguishable from a mutant nothing catches. Loud, not a false pass, and the conversion scripts now copy the string rather than retyping it; - rerunning a conversion script after the vocabulary had grown rewrote two teeth back to their earlier form; the anchors pass refused both at once, which is the case for the pass being in the static contract seen from the other side - it catches edits to the register, not only drift in the code. Reading: anchors 110 of 110 resolve over 22 targets; mutants 110 of 110 caught by the named test, all 110 by the case their expect names, 22 of 22 restored byte-identically, exit 0 in 160 s; anchor cases 34 pass; test:product 1580 pass; verify:static exit 0; lint 0 findings; complexity gate ok; format check clean. The register is 110 teeth over 22 targets covering 21 files, every one a name plus an operator plus a selector, with 26 operators - each a mutation class the catalogues name or a slot they name. What is not claimed: a derived tooth is not stronger evidence than the byte anchor it replaced. It is the same violation caught by the same case; what it buys is that the tooth stays aimed at its rule across renames and reflows, which is the failure this record exists for. The operators are also not the whole catalogue, only the classes this register's rules need. Docs and code in one commit: the record's residue section is replaced by the conversion table, the three corrected reasons and the two mistakes; the ledger says all 110 teeth are derived and gains this reading; the teeth header says which names a scope can have, module level included. --- ...-09-24-mutants-are-derived-not-anchored.md | 88 ++++++---- ...-mutants-are-derived-not-anchored.zh-CN.md | 35 ++-- .../design/task-unit-semantics-obligations.md | 10 +- tests/tools/mutation-anchor.test.ts | 94 ++++++++++ tools/mutation-anchor.ts | 160 +++++++++++++++--- tools/mutation-teeth.ts | 122 ++++++------- 6 files changed, 379 insertions(+), 130 deletions(-) diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md index 42a46ece..273719bb 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md @@ -129,8 +129,8 @@ targets`. The target `tools/agent-verify.ts` went with its single tooth, so the strings rather than by a filesystem. **Landed:** the resolver is `tools/mutation-anchor.ts` (it held no state, so it moved out of the sweep script and the tests can call it), with 16 cases over source strings and no filesystem access. - **Extended to the whole register (2026-09-24): 59 more teeth converted, then the operators the - catalogues name, and the register is now 95 derived of 110.** First pass: every class of site the + **Extended to the whole register (2026-09-24), in three waves: 59 teeth, then 18, then the last 15 - + and the register is now 110 derived of 110.** First pass: every class of site the six operators already had, taken target by target and swept after each - `condition-never` on a guard or an initializer (44 in total), `neutralize-term` (9), `replace-property` (9), `replace-argument` (6), `drop-statement` (6), `condition-holds` (3). @@ -179,36 +179,60 @@ targets`. The target `tools/agent-verify.ts` went with its single tooth, so the caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 176 s. - **What is left hand-written is 15 teeth, and the classification is now measured rather than - guessed:** - - - three are the same shape the tools also cannot express: a fragment **inside a string** - two SQL - clauses (`prune-ignores-retention`, `bounded-pin-never-expires`) and one SQLite unit - (`the-grace-uses-a-unit-sqlite-does-not-know`). Catalogues mutate the whole literal - (`StringLiteral` to `""`), and SQL has its own catalogue (SQLMutation: clause deletion, `AND` to - `OR`, NULL mutations), so the honest next step for these is clause-level SQL operators, not byte - surgery; - - two are at **module top level**, where there is no member to name: a guard in a script - (`judge-may-judge-its-own-delivery`) and `export const CLOCK_GRACE_MS = 50`. Giving `derive` a - file-scope mode is a design decision, not a widening, and it is not taken here; - - two move an **index** (`[0]` to `[1]`). This one corrects this record's own earlier list: it said - "index moved (2)" was among the classes the catalogues cover, and that is wrong. `FirstToLast` is a - method _named_ `first`; no catalogue checked here mutates an array index; - - two write **two statements on one line** (`fusion-rollback-not-counted`, - `a-score-with-no-recorded-history-counts-as-a-reproduction`) - the line discipline of - `drop-statement`, not a mutation class; - - one rewrites a comparison into a range test (`=== 1` to `>= 1`), where the catalogue's negation - (`!== 1`) would not realise the violation the name states; - - four are bespoke expression rewrites with no operator behind them: a constructor swapped for a - crafted substitute of the same interface (`the-status-read-path-opens-the-rounds-store`), a - multi-line call replaced by a stub fact (`a-coordinated-write-skips-its-run-fact`), a wrapper - unwrapped to its inner call (`the-daemon-verb-skips-the-coordinated-path`), and - `view-adapter-accepts-bytes`, where two statements become one `return`; - - one, `every-task-is-frozen-at-position-zero`, changes two things at once and is kept deliberately. - - Three of those clusters would repay an operator and are recorded as open: clause-level SQL mutation, - a file-scope selector, and an index operator. None is added here, because none is a widening - each - is a new scoping rule or a new domain - and the measured cost of not having them is seven teeth. + **Then every one of the 33 converted, and the register is 110 derived of 110 with no hand-written + tooth left (2026-09-24).** The three that were declared open above as "worth an operator" were taken + instead of argued about, and each turned out to be a small, nameable widening rather than a new + domain: + + | what the residue needed | what was actually added | teeth | + | ------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----- | + | a file-scope selector | `within` is optional: without it the whole file is the scope, and the selector has to be unique in it | 2 | + | an index operator | `replace-index`: the element access the selector identifies gets the declared index | 2 | + | clause-level SQL mutation | `replace-literal-fragment`: a piece of text written inside a literal becomes the declared fragment, with the same holder disambiguator the call selectors have | 3 | + | a function bound to a property | `uniqueMember` accepts a function bound to a variable _or an object property_ (`claim: (store, parsed) => ...`) | 1 | + | a shorthand property | `replace-property` writes a shorthand out (`position` becomes `position: 0`), because a value cannot go where the name was | 1 | + | a guard inside a module-level script | the file scope above, with `condition-never` | 1 | + | a constructor call | the call selectors accept `new X(...)` as a call site | 1 | + | a call replaced by a boolean, a value replaced by a stub fact, a wrapper unwrapped, an iterable shortened, a comparison negated | the operators that already existed, used with a declared `to` | 5 | + + Three things this pass settled by measurement rather than by argument, each of which had been + written down here as a reason **not** to convert: + + - `speculation-guesses-several-facts-at-once` was kept because the catalogue's negation (`=== 1` to + `!== 1`) "would not realise the violation the name states". It does: the named case fails under the + negated comparison. The reasoning was wrong and the sweep said so. + - `every-task-is-frozen-at-position-zero` was kept because it "changes two things at once". Only one + of the two changes is load-bearing: `noUnusedParameters` is not set, so the callback can keep the + parameter it no longer reads, and the tooth converts to a single `replace-property`. + - The two teeth that write two statements on one line were never about the line discipline at all. + Their hand anchors had to span two statements; the mutations are single-statement deletions, and + both converted to `drop-statement`. + + Two mistakes were made and caught inside this pass, and both are worth keeping: + + - **A hand-typed `expect` reads as a broken mutant.** Retyping two case names from memory instead of + copying them from the register made both teeth report **"survived"** - the tool filters the suite by + the named case, so a name that matches no case looks exactly like a mutant nothing catches. It is a + loud failure (an alarm, not a false pass), and it is the second time this register has been bitten + by the `expect` field being an assertion of its own: the conversion script now copies the string, + and the copy is why the same conversions then read `caught by the named test`. + - **A conversion script rerun against a converted register.** Re-running the first wave's script after + the vocabulary had grown rewrote two teeth back to their earlier form (`in` dropped, `within` back + to the guessed member), and the anchors pass refused both immediately. That is the argument for the + anchors pass being in the static contract, seen from the other side: it also catches _edits to the + register_, not only drift in the code it points at. + + **The register is now 110 teeth, every one a name plus an operator plus a selector, over 22 targets + covering 21 files; 26 operators, all of them a mutation class the catalogues name or a slot they name + (a value, an argument, a property, an initializer, an iterable, an index, a callee, a literal + fragment).** Reading: `anchors: 110 of 110 resolve, over 22 targets`; `mutants: 110 of 110 caught by +the named test`, **all 110 by the case their `expect` names**, 22 of 22 restored byte-identically, + exit 0 in 160 s; `tests/tools/mutation-anchor.test.ts` 34 pass. + What is _not_ claimed: a derived tooth is not stronger evidence than the byte anchor it replaced - it + is the same violation, caught by the same case. What it buys is that the tooth stays aimed at its rule + across renames and reflows, which is the failure this record exists for. The operators are also not + the whole catalogue: they are the classes this register's rules happen to need, and adding one is now + an entry in a table plus its name in the union. Two things the catalogues say about the shape of this work are worth carrying: - **Subsumption.** pitest's _Less is more_ names the effect this record's retirement criterion diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index fe2c4b3a..73b110dc 100644 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -29,7 +29,7 @@ **一颗牙的耐久部分是它的名字、它的算子、它的选择器,以及必须失败的那个用例。位置与替换字节每次从语法树算出来。** 1. **六个作用在具名选择器上的算子**,作用域限定在一个成员内:`condition-never`(选择器指到的那条条件永不成立)、`condition-holds`(它始终成立——镜像的那个,因为写成 `return a && b` 的规则说的是“这条条件为真”,把它换成 `false` 是把规则反过来,而不是去掉)、`neutralize-term`(该条件里的一项变成它的恒等元——`&&` 下为 `true`,`||` 下为 `false`)、`replace-argument`(某个调用的第 _n_ 个实参换成声明的片段)、`replace-property`(具名对象属性的值换成声明的片段)、`drop-statement`(选择器指到的那条语句连行带缩进被删掉)。两条规则让选择器保持诚实:匹配到多于一处、或者片段同时命中两个候选的,一律拒绝而不是猜;两个“整条件”算子拒绝只命名条件一部分的片段,因为整段替换会把 mutant 悄悄放大成“这条规则的每一条理由”——一项有它自己的算子。规则用不上这些时,仍可用“成员作用域 + 最小片段”的旧形式;而**“最小”是要点**:不要消息文本、不要同级实参、一项能表达的不要写成整条语句。 - **已推广到整个登记处(2026-09-24):又转了 59 颗,随后补上"现成目录点名的算子",登记处现在是 110 颗里 95 颗 derived。**第一遍:六个算子已经能表达的每一类点位,按目标逐个转完、每转完一个就 sweep——`condition-never`(守卫或初始化器,共 44)、`neutralize-term`(9)、`replace-property`(9)、`replace-argument`(6)、`drop-statement`(6)、`condition-holds`(3)。 + **已推广到整个登记处(2026-09-24),分三波:59 颗、18 颗,最后 15 颗——登记处现在是 110 颗里 110 颗 derived。**第一遍:六个算子已经能表达的每一类点位,按目标逐个转完、每转完一个就 sweep——`condition-never`(守卫或初始化器,共 44)、`neutralize-term`(9)、`replace-property`(9)、`replace-argument`(6)、`drop-statement`(6)、`condition-holds`(3)。 第二遍:把剩下的 33 颗按"缺什么"分类,再拿现成变异工具真正点名的东西核对(Stryker 的 supported mutators、pitest 的 mutator 列表、cargo-mutants 的 patterns、Cosmic Ray 的 operator 概念)。其中十三颗落在"每个目录都有的变异类"里,于是词汇增加了七个算子,而不是宣布这一类做不到: | 算子 | 现成目录里的名字 | 颗数 | @@ -49,16 +49,31 @@ 第一遍我又写错了一处 `within`,这次是**anchors pass**(不是 sweep)抓出来的:`the-completion-ignores-a-cancelled-unit` 被指向 `checkDispatch` 而不是 `checkCompletion`(同一片段在两个成员里都有)。这是更安静的那种失败:anchors pass 照样通过、mutant 照样被抓——但抓它的是**整个 suite**,不是那条"报告一个被取消单元的完成"的用例。第二遍同类错误反而是响的:我瞄的那个守卫在 `runParentCheck`,我却写了 `instrumentCommit`,pass 直接拒绝——"0 guards in instrumentCommit match the selector"。**成员名写错,找不到时是响的、找到两处同形状时是哑的**,这就是"转完必须 sweep"不可省的理由。 **两遍之后的读数:**`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,176 秒。 - **剩下手写的是 15 颗,分类这一次是量出来的、不是猜的:** - - 三颗是**连现成工具也表达不了**的形状:**字符串里的片段**——两条 SQL 子句(`prune-ignores-retention`、`bounded-pin-never-expires`)和一处 SQLite 单位(`the-grace-uses-a-unit-sqlite-does-not-know`)。目录只会替换整个字面量(`StringLiteral` → `""`),而 SQL 有自己的目录(SQLMutation:删子句、`AND` 换 `OR`、NULL 变异),所以这三颗诚实的方向是"子句级 SQL 算子",不是对字节动手术; - - 两颗在**模块顶层**,没有成员可命名:脚本里的一个守卫(`judge-may-judge-its-own-delivery`)和 `export const CLOCK_GRACE_MS = 50`。给 `derive` 加"文件级作用域"是设计决定而不是扩宽,这里不做; - - 两颗动**下标**(`[0]` → `[1]`)。这一颗同时纠正本记录自己此前的说法:之前把"下标移动(2)"列为目录覆盖的类,这是错的——`FirstToLast` 是**名为 `first` 的方法**,我查过的目录里没有哪个变异数组下标; - - 两颗是**两条语句挤在同一行**(`fusion-rollback-not-counted`、`a-score-with-no-recorded-history-counts-as-a-reproduction`)——这是 `drop-statement` 的行纪律问题,不是变异类; - - 一颗把比较改写成范围判断(`=== 1` → `>= 1`),而目录里的取反(`!== 1`)**不会**实现名字所陈述的那个违规; - - 四颗是无算子可依的定制表达式改写:构造调用被换成同接口的手写替身(`the-status-read-path-opens-the-rounds-store`)、多行调用被换成桩事实(`a-coordinated-write-skips-its-run-fact`)、包装被拆到内层调用(`the-daemon-verb-skips-the-coordinated-path`),以及 `view-adapter-accepts-bytes`(两条语句变成一个 `return`); - - 一颗 `every-task-is-frozen-at-position-zero` 一次改两处,故意留着。 + **随后这 33 颗全部转完,登记处现在是 110 颗里 110 颗 derived,没有一颗手写(2026-09-24)。**上面写成"值得加算子"的三簇,是拿实现去验而不是继续争论:每一簇都不是新领域,而是一处小且可命名的扩宽—— - 其中三簇"值得加算子",作为待决写下:子句级 SQL 变异、文件级选择器、下标算子。这里都不加,因为它们都不是**扩宽**——各自是一条新的作用域规则或一个新领域——而不加它们的代价是量出来的七颗。 + | 残留需要什么 | 实际加了什么 | 颗数 | + | ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | ---- | + | 文件级选择器 | `within` 变可选:不写就用整个文件作作用域,选择器必须在文件里唯一 | 2 | + | 下标算子 | `replace-index`:选择器指到的元素访问换成声明的下标 | 2 | + | 子句级 SQL 变异 | `replace-literal-fragment`:字面量内部的一段文本换成声明的片段,并带上调用选择器那套"容器"消歧 | 3 | + | 绑定到对象属性的函数 | `uniqueMember` 接受绑定到变量**或对象属性**的函数(`claim: (store, parsed) => ...`) | 1 | + | 简写属性 | `replace-property` 把简写写开(`position` 变 `position: 0`),因为值不能写在名字的位置上 | 1 | + | 模块级脚本里的守卫 | 上面的文件级作用域,配 `condition-never` | 1 | + | 构造函数调用 | 调用选择器接受 `new X(...)` 作为调用点 | 1 | + | 调用换成布尔、值换成桩事实、包装拆开、可迭代缩短、比较取反 | 已有的算子,配声明的 `to` | 5 | + + 这一遍里有三件事是**量出来**而不是讲道理定的,而它们此前都被写在本记录里当作"不转"的理由: + + - `speculation-guesses-several-facts-at-once` 当初留着,理由是"目录里的取反(`=== 1` → `!== 1`)不会实现名字所说的违规"。事实上会:换成取反比较后,点名的那条用例照样失败。理由是错的,是 sweep 说了算。 + - `every-task-is-frozen-at-position-zero` 当初留着,理由是"一次改两处"。两处里只有一处是承重的:`noUnusedParameters` 没开,回调可以继续保留它不再读的参数,于是这颗牙变成单条 `replace-property`。 + - 那两颗"两条语句挤一行"从来不是行纪律的事。它们的手写锚点**不得不**横跨两条语句,而变异本身是删掉一条语句,两颗都转成了 `drop-statement`。 + + 这一遍里犯的两个错都被当场抓住,都值得记下来: + - **`expect` 是手打的,读起来就像一颗坏掉的牙。**有两颗牙我凭记忆重打了用例名(没有从登记处抄),结果两颗都报 **"survived"**——工具按点名的用例过滤 suite,名字对不上任何用例时,看起来就跟"没任何用例抓到"一模一样。它是响的失败(警报,不是误判通过),但这是这个登记处第二次被 `expect` 字段咬到:转换脚本现在**抄**这个字符串,抄过之后同样两颗读数是 `caught by the named test`。 + - **转换脚本对着已转换的登记处又跑了一遍。**词汇长大之后重跑第一遍的脚本,把两颗牙改回了早期形态(`in` 掉了、`within` 回到猜的那个成员),anchors pass 立刻拒绝了两个。这正是"把 anchors pass 放进静态契约"的另一个方向的理由:它也抓**对登记处本身的编辑**,不只是抓它指向的代码漂移。 + + **登记处现在是 110 颗牙,每一颗都是"名字 + 算子 + 选择器",覆盖 22 个目标、21 个文件;26 个算子,每一个要么是目录点名的变异类,要么是目录点名的槽位**(值、实参、属性、初始化器、可迭代、下标、被调、字面量片段)。读数:`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,160 秒;`tests/tools/mutation-anchor.test.ts` 34 条全过。 + 没有声称的东西:derived 形态**不比**它替代的字节锚点更强——违规是同一个、抓它的用例是同一条。它换来的是:改名、重排之后这颗牙仍瞄准原来那条规则,而这正是本记录要解决的那个失败。算子也不等于整个目录:它们只是本登记处的规则恰好需要的那些类,加一个算子现在就是表里加一行、类型里加一个名字。 目录里还有两件事值得带过来: - **subsumption(被覆盖)。**pitest 的《Less is more》给这件事起了名:一个 mutant 若已被其他 mutant 的组合覆盖,它带来的只是运行时间,不是信心。pitest 靠"固定的默认集合"回避;本登记处是**量**它——sweep 先说出是哪条用例失败,再读那条用例确认它确实陈述了规则。这正是 37 颗牙被退役而不是留着的原因,也是"是不是每种能想到的变异都该进来"的答案是"不"的原因。 - **unviable mutant(编译不过的变异体)。**cargo-mutants 把这类单独计数、默认不打印;它排除测试函数与 `unsafe`,并且明确写出**为什么不生成某些变异**(`==` 换成 `<`"太容易产生假阳性"、`-a` 换成 `+a`"太容易不编译")。本登记处**按构造就没有** unviable(每颗都有点名的用例),但理由是同一条:为一个在本仓库很少出现的形状加通用算子,产出的多半是噪音。 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index d94b411e..ec2607ae 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -24,6 +24,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth -- --anchors-only` -> `anchors: 144 of 144 resolve, over 23 targets`, exit 0 in 0.7 s (2026-09-24, after five teeth were retired in favour of the checks that already state their rules) - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts` -> `mutants: 15 of 15 caught by the named test`, restored byte-identically (2026-09-24, after six more teeth were retired there: the five budget/publication rules and the run-scoping claim) - `npm run mutation:teeth` (the whole register) -> `mutants: 136 of 136 caught by the named test`, 23 of 23 targets restored byte-identically, exit 0 in 217 s (2026-09-24), **every one of them caught by the case its `expect` names**. The first whole-register reading (231 s, 134 of 136 by name) is what surfaced three wrong links, all repaired in the same pass: `the-caller-rebuilds-the-shared-floor` (its `expect` was a paraphrase of no case, re-pointed), `the-grace-is-zero` (its named case derived its fixture from the constant under test, so it passed while the bug was live - the fixture is a literal now), and `next-is-not-the-head-of-the-ordered-candidates` (re-pointed to `at the default budget the licence is still the head of the ordered set`, which now asserts the head). A fourth was found by the pilot sweep and repaired here: `fusion-continues-from-an-unverified-answer`'s case refused the pair for a second reason as well, so it passed under the mutant - with no dependency between the two units it is refused by that one condition only. One row still carries a note rather than an assertion: `the-pass-asks-a-unit-it-already-failed-again` is caught by a named case that does not finish inside its 30 s bound, which this tool counts as caught and prints with that reason. +- `npm run mutation:teeth` (the whole register, every tooth derived) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 160 s (2026-09-24); `npm run mutation:anchors` -> `anchors: 110 of 110 resolve, over 22 targets`; `tests/tools/mutation-anchor.test.ts` -> 34 pass; `test:product` -> 1580 pass; `verify:static` exit 0; lint 0 findings; complexity gate ok. - `npm run mutation:teeth` (the whole register, after the operators the catalogues name were added and 18 more teeth converted) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 176 s (2026-09-24); `npm run mutation:anchors` -> `anchors: 110 of 110 resolve, over 22 targets`; `tests/tools/mutation-anchor.test.ts` -> 30 pass. - `npm run mutation:teeth` (the whole register, after 59 teeth were converted to derived selectors) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 171 s (2026-09-24). The conversion is not what makes them bite; it is what keeps them aimed: one wrong `within` was shipped in this pass and the sweep is what caught it, as a tooth caught by the suite rather than by the case that names its rule. - `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were **orphans**: no row named `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` or `the-parent-check-ignores-its-own-verdict`, so their rules have a case but no row; the check each case makes is in the [record](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. @@ -40,9 +41,12 @@ description rather than a pin. Every mutant name below was read from A tooth is a **name, an operator and a selector** - not a copy of a line: the site and the replacement bytes are resolved from the syntax tree on every run, so a rename or a reflow does not retire the pin while the rule still stands ([the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). -95 of the register's 110 teeth are in that form now; the 15 that are not are the shapes the operators -cannot express, clustered and counted in the same record, with the three clusters that would repay an -operator named there (clause-level SQL mutation, a file-scope selector, an index operator). `npm run mutation:teeth -- --anchors-only` answers whether every +**All 110 of the register's teeth are in that form now**, over 22 targets covering 21 files, with the +26 operators the record lists - each one a mutation class the catalogues name or a slot they name. A +tooth with no operator to name it is no longer residue but a missing operator: the last 15 converted +after `within` became optional (the file is a scope), the call selectors took `new X(...)`, and three +operators were added for the shapes that had none (an index, a fragment of a literal, a shorthand +property written out). `npm run mutation:teeth -- --anchors-only` answers whether every anchor still applies - under a second, no suite run, nothing written - because a tooth that stopped matching its rule is otherwise only visible as a check that stopped counting. diff --git a/tests/tools/mutation-anchor.test.ts b/tests/tools/mutation-anchor.test.ts index 7bfe91b0..978a0d13 100644 --- a/tests/tools/mutation-anchor.test.ts +++ b/tests/tools/mutation-anchor.test.ts @@ -582,3 +582,97 @@ test("a function bound to a name is a member, and the selectors search its body" assert.ok(!("reason" in site)); assert.equal(above.slice(site.start, site.end), 'new RegExp("x").exec(values)'); }); + +test("the index of an element access is replaced, and the access is named when there are several", () => { + const above = `function selection(board) {\n const head = board.candidates()[0] ?? null;\n const tail = board.candidates()[1] ?? null;\n return [head, tail];\n}\n`; + assert.equal( + selected( + above, + named( + "x", + { within: "selection", operator: "replace-index", access: "candidates()[0]" }, + { to: "1" }, + ), + ), + "0", + ); + assert.match( + refusal( + above, + named( + "x", + { within: "selection", operator: "replace-index", access: "candidates()" }, + { to: "1" }, + ), + ), + /2 element accesses in selection match the selector/, + ); +}); + +test("a piece of text inside a literal is replaced where it is written, and a repeated one is refused", () => { + const above = `function prune() {\n return db.prepare(\n \`DELETE FROM entries WHERE expires_at <= ?\n AND id NOT IN (SELECT entry_id FROM retentions)\`,\n );\n}\n`; + assert.equal( + selected( + above, + named( + "x", + { + within: "prune", + operator: "replace-literal-fragment", + text: "AND id NOT IN (SELECT entry_id FROM retentions)", + }, + { to: "" }, + ), + ), + "AND id NOT IN (SELECT entry_id FROM retentions)", + ); + const quoted = `function prune() {\n const a = "WHERE 0";\n const b = "WHERE 0";\n return [a, b];\n}\n`; + assert.match( + refusal( + quoted, + named( + "x", + { within: "prune", operator: "replace-literal-fragment", text: "WHERE 0" }, + { to: "" }, + ), + ), + /2 literals in prune match the selector/, + ); +}); + +test("a constructor call is a call site, and the module is a scope when no member is named", () => { + const above = `function openRoundQuery(path) {\n const db = new DatabaseSync(path, { readOnly: true });\n return db;\n}\n`; + const site = locate( + above, + named( + "x", + { within: "openRoundQuery", operator: "replace-call", call: "DatabaseSync" }, + { to: "new BoardAdmission(path) as unknown as DatabaseSync" }, + ), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "new DatabaseSync(path, { readOnly: true })"); + const module = `if (entry.deliveredBy === agentId) {\n throw new Error("the deliverer may not judge its own work");\n}\n`; + const guard = locate( + module, + named("x", { operator: "condition-never", condition: "entry.deliveredBy === agentId" }), + ); + assert.ok(!("reason" in guard)); + assert.equal(module.slice(guard.start, guard.end), "entry.deliveredBy === agentId"); + assert.equal(guard.replacement, "false"); +}); + +test("a shorthand property is written out, because the value cannot go where the name was", () => { + const above = `function selection(task, runId) {\n return { ...task, runId, position };\n}\n`; + const site = locate( + above, + named( + "x", + { within: "selection", operator: "replace-property", property: "position" }, + { to: "0" }, + ), + ); + assert.ok(!("reason" in site)); + assert.equal(above.slice(site.start, site.end), "position"); + assert.equal(site.replacement, "position: 0"); +}); diff --git a/tools/mutation-anchor.ts b/tools/mutation-anchor.ts index afbcc0f0..63345a5e 100644 --- a/tools/mutation-anchor.ts +++ b/tools/mutation-anchor.ts @@ -36,8 +36,10 @@ export interface Mutant { /** The operators a derived mutant may use, and the selector each one needs. */ export interface Derive { - /** The member whose body holds the site. Exact, and required: the scope is what makes it unique. */ - readonly within: string; + /** The member whose body holds the site. The scope is what makes a fragment unique, so it is named + * wherever there is a member to name; omitted at module level, where the whole file is the scope and + * the selector has to be unique in it. */ + readonly within?: string; readonly operator: /** The condition the selector identifies never holds: it becomes `false`. */ | "condition-never" @@ -74,6 +76,12 @@ export interface Derive { /** The iterable of the loop the selector identifies becomes the declared `to` fragment, so the * loop runs over an empty or a shortened collection. */ | "replace-iterable" + /** The index of the element access the selector identifies becomes the declared `to` fragment: + * the code reads the next element instead of the head of the collection. */ + | "replace-index" + /** A piece of text written inside a literal becomes the declared `to` fragment, for the data a + * catalogue only mutates whole: a SQL clause, a format unit. */ + | "replace-literal-fragment" /** The statement the selector identifies is removed, with its line and its indentation: an * expression, a declaration, a `return`, a `throw`, or a guard clause. */ | "drop-statement"; @@ -92,6 +100,11 @@ export interface Derive { readonly variable?: string; /** Which loop, for `replace-iterable`: a fragment of the loop's own text, e.g. `of unknownBudget`. */ readonly iterable?: string; + /** Which element access, for `replace-index`: a fragment of its own text, e.g. `candidates()[0]`. */ + readonly access?: string; + /** Which piece of text, for `replace-literal-fragment`: the text itself, as it is written inside + * the literal, e.g. `retained_until IS NOT NULL`. */ + readonly text?: string; /** The holder that tells two same-named sites apart: for `replace-property`, a fragment of the * object literal that holds it; for `replace-argument`, a fragment of the call's own text; for * `remove-call` and `replace-call`, a fragment of the statement the call sits in. */ @@ -244,14 +257,19 @@ function collect(root: ts.Node, isWanted: (node: ts.Node) => /** The one member with this name in the file, or why it is not one site. A class constructor is a * member too: its name is written `constructor`, and code that refuses a second plan while it opens * the store lives there and nowhere else. So is a function bound to a name - `const count = (label) - * => ...` names one function, and whether the formatter wrote it as a declaration is not a fact about - * which rules live inside it. */ + * => ...` or `claim: (store, parsed) => ...` names one function, and whether the formatter wrote it as + * a declaration is not a fact about which rules live inside it. */ function uniqueMember(source: ts.SourceFile, name: string): ts.Node | { reason: string } { + const bound = (node: ts.Node, initializer: ts.Expression | undefined): boolean => + initializer !== undefined && + (ts.isArrowFunction(initializer) || ts.isFunctionExpression(initializer)); const boundFunction = (node: ts.Node): boolean => - ts.isVariableDeclaration(node) && - node.name.getText(source) === name && - node.initializer !== undefined && - (ts.isArrowFunction(node.initializer) || ts.isFunctionExpression(node.initializer)); + (ts.isVariableDeclaration(node) && + node.name.getText(source) === name && + bound(node, node.initializer)) || + (ts.isPropertyAssignment(node) && + node.name.getText(source) === name && + bound(node, node.initializer)); const named = (node: ts.Node): boolean => ((ts.isMethodDeclaration(node) || ts.isFunctionDeclaration(node)) && node.name?.getText(source) === name) || @@ -263,17 +281,20 @@ function uniqueMember(source: ts.SourceFile, name: string): ts.Node | { reason: reason: `member ${name} matched ${members.length} members, refusing to claim a check`, }; const member = members[0]!; + const body = + ts.isVariableDeclaration(member) || ts.isPropertyAssignment(member) + ? member.initializer + : undefined; // A name bound to a function is the function's body for every purpose here, so the selectors search // the body rather than the declaration that wraps it. - if (ts.isVariableDeclaration(member) && member.initializer) return member.initializer; - return member; + return body ?? member; } /** Argument `arg` of the one call to `call` inside the member. */ function argumentSite( source: ts.SourceFile, member: ts.Node, - derive: Derive, + derive: Scoped, to: string | undefined, ): Site | { reason: string } { if (derive.call === undefined || derive.arg === undefined) @@ -314,14 +335,18 @@ function holderText(source: ts.SourceFile, node: ts.Node): string { function propertySite( source: ts.SourceFile, member: ts.Node, - derive: Derive, + derive: Scoped, to: string | undefined, ): Site | { reason: string } { if (derive.property === undefined) return { reason: "replace-property needs a `property` name" }; if (to === undefined) return { reason: "replace-property needs the mutant's `to` fragment" }; + // A shorthand is a property too: `{ position }` writes the same property as `{ position: position }`, + // and it is replaced whole (`position: 0`), because writing the value where the name was would leave + // a spread element rather than a property. const named = (node: ts.Node): boolean => - ts.isPropertyAssignment(node) && node.name.getText(source) === derive.property; - const properties = collect(member, named); + (ts.isPropertyAssignment(node) || ts.isShorthandPropertyAssignment(node)) && + node.name.getText(source) === derive.property; + const properties = collect(member, named); const held = derive.in === undefined ? properties @@ -330,7 +355,15 @@ function propertySite( return { reason: `property ${derive.property} matched ${held.length} sites in ${derive.within}, refusing to claim a check`, }; - const value = held[0]!.initializer; + const property = held[0]!; + if (ts.isShorthandPropertyAssignment(property)) + return { + start: property.getStart(source), + end: property.getEnd(), + replacement: `${derive.property}: ${to}`, + retaken: false, + }; + const value = property.initializer; return { start: value.getStart(source), end: value.getEnd(), replacement: to, retaken: false }; } @@ -338,7 +371,7 @@ function propertySite( function statementSite( source: ts.SourceFile, member: ts.Node, - derive: Derive, + derive: Scoped, text: string, ): Site | { reason: string } { if (derive.statement === undefined) @@ -422,7 +455,7 @@ function decisionExpressions(member: ts.Node): ts.Expression[] { function conditionSite( source: ts.SourceFile, member: ts.Node, - derive: Derive, + derive: Scoped, ): ts.Expression | { reason: string } { const conditions = decisionExpressions(member); // Matched with the same whitespace normalization as everything else: a condition written across @@ -454,7 +487,7 @@ function conditionSite( function wholeCondition( source: ts.SourceFile, condition: ts.Expression, - derive: Derive, + derive: Scoped, ): Site | { reason: string } { if (derive.condition === undefined) return { reason: `${derive.operator} needs a \`condition\` fragment` }; @@ -504,7 +537,7 @@ function identityLiteral(before: string, after: string): string | null { function neutralizedTerm( source: ts.SourceFile, condition: ts.Expression, - derive: Derive, + derive: Scoped, ): Site | { reason: string } { if (derive.term === undefined) return { reason: "neutralize-term needs a `term` fragment" }; const text = condition.getText(source); @@ -524,13 +557,19 @@ function neutralizedTerm( }; } -/** What every resolver is given: the parsed file, the one member, the selector, and the bytes a +/** The selector as the resolvers see it: the scope is always named, because `deriveSite` fills in + * "the module" when the mutant did not name one. */ +interface Scoped extends Omit { + readonly within: string; +} + +/** What every resolver is given: the parsed file, the one scope, the selector, and the bytes a * value-substituting operator declares. */ interface Context { readonly source: ts.SourceFile; readonly text: string; readonly member: ts.Node; - readonly derive: Derive; + readonly derive: Scoped; readonly to: string | undefined; } @@ -554,7 +593,7 @@ function innermost(source: ts.SourceFile, nodes: readonly T[] /** The one collection site the fragment identifies, or why there is not one. */ function oneSite( source: ts.SourceFile, - derive: Derive, + derive: Scoped, nodes: readonly T[], what: string, ): T | { reason: string } { @@ -675,16 +714,16 @@ function holderStatementText(source: ts.SourceFile, node: ts.Node): string { } /** The one call the selector names, by the callee as written. */ -function callSite(context: Context): ts.CallExpression | { reason: string } { +function callSite(context: Context): ts.CallExpression | ts.NewExpression | { reason: string } { const { source, derive } = context; if (derive.call === undefined) return { reason: `${derive.operator} needs a \`call\` fragment` }; const found = oneSite( source, derive, - collect( + collect( context.member, (node) => - ts.isCallExpression(node) && + (ts.isCallExpression(node) || ts.isNewExpression(node)) && node.expression.getText(source) === derive.call && (derive.in === undefined || containsText(holderStatementText(source, node), derive.in)), ), @@ -777,6 +816,66 @@ function replaceIterable(context: Context): Site | { reason: string } { }; } +/** The index of the one element access the selector identifies becomes the declared fragment: the + * code reads the next element instead of the head of the collection. */ +function replaceIndex(context: Context): Site | { reason: string } { + const { source, derive } = context; + if (derive.access === undefined) + return { reason: "replace-index needs an `access` fragment naming the element access" }; + if (context.to === undefined) return { reason: "replace-index needs the mutant's `to` fragment" }; + const fragment = derive.access; + const found = oneSite( + source, + derive, + collect( + context.member, + (node) => ts.isElementAccessExpression(node) && containsText(node.getText(source), fragment), + ), + "element accesses", + ); + if ("reason" in found) return found; + return { + start: found.argumentExpression.getStart(source), + end: found.argumentExpression.getEnd(), + replacement: context.to, + retaken: false, + }; +} + +/** A piece of text written inside a literal becomes the declared fragment. Catalogues mutate whole + * literals; the data a query is made of keeps its own rules inside one, and this is the narrowest + * operator that names that: the text is matched where it is written, and it must be written once. */ +function replaceLiteralFragment(context: Context): Site | { reason: string } { + const { source, derive } = context; + if (derive.text === undefined) + return { reason: "replace-literal-fragment needs a `text` fragment" }; + if (context.to === undefined) + return { reason: "replace-literal-fragment needs the mutant's `to` fragment" }; + const fragment = derive.text; + const holders = collect( + context.member, + (node) => + (ts.isStringLiteralLike(node) || + ts.isNoSubstitutionTemplateLiteral(node) || + ts.isTemplateExpression(node)) && + containsText(node.getText(source), fragment) && + (derive.in === undefined || containsText(holderStatementText(source, node), derive.in)), + ); + const found = oneSite(source, derive, holders, "literals"); + if ("reason" in found) return found; + const at = found.getText(source).indexOf(fragment); + if (at < 0) + return { + reason: `cannot read ${fragment} in the literal it was found in, refusing to claim a check`, + }; + return { + start: found.getStart(source) + at, + end: found.getStart(source) + at + fragment.length, + replacement: context.to, + retaken: false, + }; +} + /** The operators, each one a function of the selector and the declared bytes. Adding an operator means * adding an entry here and its name to `Derive["operator"]`: nothing else in the register changes. */ const RESOLVERS: Readonly> = { @@ -797,6 +896,8 @@ const RESOLVERS: Readonly> = { "replace-call": replaceCall, "replace-initializer": replaceInitializer, "replace-iterable": replaceIterable, + "replace-index": replaceIndex, + "replace-literal-fragment": replaceLiteralFragment, "neutralize-term": (context) => { const condition = conditionSite(context.source, context.member, context.derive); if ("reason" in condition) return condition; @@ -818,11 +919,14 @@ const RESOLVERS: Readonly> = { */ function deriveSite( text: string, - derive: Derive, + derive: Scoped, to: string | undefined, ): Site | { reason: string } { const source = ts.createSourceFile("mutant.ts", text, ts.ScriptTarget.Latest, true); - const member = uniqueMember(source, derive.within); + // Module level is a scope like any other, and the only one that has no member to name: a mutant + // written there says so by naming none, and its selector then has to be unique in the whole file. + const scoped: Scoped = { ...derive, within: derive.within ?? "the module" }; + const member = derive.within === undefined ? source : uniqueMember(source, derive.within); if ("reason" in member) return member; - return RESOLVERS[derive.operator]({ source, text, member, derive, to }); + return RESOLVERS[scoped.operator]({ source, text, member, derive: scoped, to }); } diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index 5d01fd04..b5a0fa0e 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -32,10 +32,11 @@ * * Where a mutant says where it applies: * - * - `derive: { within, operator, ... }` is a **selector plus an operator**: the member is named (a - * method, a function declaration, a class constructor under the name `constructor`, or a function - * bound to a variable), a fragment inside it is matched (the same matcher as below, so a reflow - * cannot break it), and the bytes to write are computed on every run. This is the form a tooth + * - `derive: { within, operator, ... }` is a **selector plus an operator**: the scope is named (a + * method, a function declaration, a class constructor under the name `constructor`, a function + * bound to a variable or to an object property - or omitted at module level, where the file is the + * scope), a fragment inside it is matched (the same matcher as below, so a reflow cannot break it), + * and the bytes to write are computed on every run. This is the form a tooth * should have. Its identity is the rule and the site it names, never the text that happens to be * there today, and it is what lets a rename, a reordered condition or a lifted line leave the tooth * aimed at the same rule. Selectors that match more than one site - or a fragment that fits two @@ -137,8 +138,8 @@ const TARGETS: readonly Target[] = [ }, { name: "view-adapter-accepts-bytes", - from: " const recorded = facts.verdicts?.[unit.id];", - to: " return Boolean(facts.artifacts?.[unit.id]);", + derive: { within: "isAccepted", operator: "replace-call", call: "acceptedFact" }, + to: "Boolean(artifact)", expect: "acceptedFact is the single rule, and the view adapter does not grow its own", }, { @@ -210,16 +211,23 @@ const TARGETS: readonly Target[] = [ }, { name: "prune-ignores-retention", - ast: { within: "pruneExpiredTaskBoardEntries" }, - from: " `DELETE FROM task_board_entries WHERE task_id = ? AND expires_at <= ?\n AND id NOT IN (SELECT entry_id FROM task_board_retentions)`,", - to: " `DELETE FROM task_board_entries WHERE task_id = ? AND expires_at <= ?`,", + derive: { + within: "pruneExpiredTaskBoardEntries", + operator: "replace-literal-fragment", + text: "AND id NOT IN (SELECT entry_id FROM task_board_retentions)", + in: "DELETE FROM task_board_entries WHERE task_id", + }, + to: "", expect: "a retained entry, its delivery and its acknowledgement survive the prune", }, { name: "bounded-pin-never-expires", - ast: { within: "expireStaleRetentions" }, - from: ' "DELETE FROM task_board_retentions WHERE retained_until IS NOT NULL AND retained_until <= ?",', - to: ' "DELETE FROM task_board_retentions WHERE 0",', + derive: { + within: "expireStaleRetentions", + operator: "replace-literal-fragment", + text: "retained_until IS NOT NULL AND retained_until <= ?", + }, + to: "0", expect: "a bounded pin stops pinning when its bound passes", }, @@ -263,12 +271,9 @@ const TARGETS: readonly Target[] = [ ], mutants: [ { - // The borrowed view must not reopen the writer's path: this mutant makes the read path - // migrate and publish the store it was asked only to read. name: "the-status-read-path-opens-the-rounds-store", - ast: { within: "openRoundQuery" }, - from: "const db = new DatabaseSync(databasePath, { readOnly: true });", - to: "const db = new BoardAdmission(databasePath) as unknown as DatabaseSync;", + derive: { within: "openRoundQuery", operator: "replace-call", call: "DatabaseSync" }, + to: "new BoardAdmission(databasePath) as unknown as DatabaseSync", expect: "the query port reads a round without migrating, publishing or exposing a write", }, @@ -362,13 +367,9 @@ const TARGETS: readonly Target[] = [ }, { - // Selection and ranking must not be able to disagree with each other. `next()` is the head of - // `candidates()`, so a caller that starts one unit and a caller that starts several read the - // same order - and a mutant that moves the head by one breaks the round's dispatch. name: "next-is-not-the-head-of-the-ordered-candidates", - ast: { within: "next" }, - from: " return this.candidates()[0] ?? null;", - to: " return this.candidates()[1] ?? null;", + derive: { within: "next", operator: "replace-index", access: "this.candidates()[0]" }, + to: "1", expect: "at the default budget the licence is still the head of the ordered set", }, ], @@ -479,21 +480,22 @@ const TARGETS: readonly Target[] = [ "the answer names the gate a unit was refused through, and the gates are the rule's own", }, { - // `nextTask` returns the head of what it decided, so an ordering cannot disagree with the - // shared rule about what may be selected: dropping the head moves both. name: "next-task-is-not-the-head-of-the-legal-set", - ast: { within: "nextTask" }, - from: " return selectableTasks(plan, slots)[0] ?? null;", - to: " return selectableTasks(plan, slots)[1] ?? null;", + derive: { + within: "nextTask", + operator: "replace-index", + access: "selectableTasks(plan, slots)[0]", + }, + to: "1", expect: "the round's own answer is the shared rule's answer, not an ordering's", }, { - // The design's first experiment allows exactly one pending fact. This lets a candidate guess - // two at once, which is the boundary the design says to prove before widening. name: "speculation-guesses-several-facts-at-once", - ast: { within: "isBoundedSpeculation" }, - from: " candidate.assumptions.length === 1 &&", - to: " candidate.assumptions.length >= 1 &&", + derive: { + within: "isBoundedSpeculation", + operator: "negate-comparison", + condition: "candidate.assumptions.length === 1", + }, expect: "the first experiment allows one pending fact, and a second is refused by name", }, @@ -647,8 +649,7 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "fusion-rollback-not-counted", - from: " rollbacks += 1;\n retries += 1;", - to: " retries += 1;", + derive: { within: "run", operator: "drop-statement", statement: "rollbacks += 1;" }, expect: "a dependency that is not accepted makes fusion pay a rollback and a retry", }, { @@ -688,8 +689,7 @@ const TARGETS: readonly Target[] = [ mutants: [ { name: "judge-may-judge-its-own-delivery", - from: " if (entry.deliveredBy === agentId) {", - to: " if (false && entry.deliveredBy === agentId) {", + derive: { operator: "condition-never", condition: "entry.deliveredBy === agentId" }, expect: "the board drivers run the protocol end to end through the daemon that serves the store", }, @@ -908,8 +908,11 @@ const TARGETS: readonly Target[] = [ }, { name: "a-score-with-no-recorded-history-counts-as-a-reproduction", - from: ' if (!provenance.observationOrder?.length) missing.push("observationOrder");\n else if (!sameOrder(provenance.observationOrder, projection.observationOrder))\n missing.push(`observationOrder=${provenance.observationOrder.join(",")}`);', - to: "", + derive: { + within: "missingValidityInputs", + operator: "drop-statement", + statement: "!provenance.observationOrder?.length", + }, expect: "a score that cannot name its own history is re-scored, never reported as a reproduction", }, @@ -975,11 +978,13 @@ const TARGETS: readonly Target[] = [ expect: "a binding refuses what the store does not hold", }, { - // The transition is the run's record of what happened to its entry; without it the board - // moved and the run has nothing to read. name: "a-coordinated-write-skips-its-run-fact", - from: " const fact = store.appendTaskRunFact(\n {\n runId: request.runId,\n kind: managedTransitionKind(request.verb),\n taskId: binding.taskId,\n attempt: binding.attempt,\n entryId: request.entryId,\n payload: JSON.stringify({ actorId: request.actorId, status }),\n },\n port,\n );", - to: " const fact = { sequence: 0, recorded: true };", + derive: { + within: "coordinatedBoardWrite", + operator: "replace-initializer", + variable: "fact", + }, + to: "{ sequence: 0, recorded: true }", expect: "a coordinated write lands the board transition and the run's fact together", }, { @@ -1025,11 +1030,9 @@ const TARGETS: readonly Target[] = [ }, { - // The plan order is the array order: the position comes from that loop, so freezing every - // task at zero would leave the stored plan's order to the task ids. name: "every-task-is-frozen-at-position-zero", - from: " request.tasks.forEach((task, position) =>\n store.freezeTaskRunTask({ ...task, runId: request.runId, position }, port),\n );", - to: " request.tasks.forEach((task) =>\n store.freezeTaskRunTask({ ...task, runId: request.runId, position: 0 }, port),\n );", + derive: { within: "freezeRunPlan", operator: "replace-property", property: "position" }, + to: "0", expect: "a run registers, freezes a plan, adopts entries, and reads it all back", }, @@ -1086,13 +1089,14 @@ const TARGETS: readonly Target[] = [ suites: ["tests/integration/ooo-managed-write.test.ts", "tests/cli/task-run-surface.test.ts"], mutants: [ { - // The routing is what keeps the daemon's verbs out of the store's refusal: a managed entry - // reached directly from a handler cannot be moved at all. The rule itself lives in the - // coordinator (one home for it), so this mutant pins that the daemon's claim still goes - // through it rather than at the store. name: "the-daemon-verb-skips-the-coordinated-path", - from: ' entry: coordinatedEntryWrite(store, {\n verb: "claim",\n entryId: p.entryId,\n actorId: p.agentId,\n apply: () => store.claimTaskBoardEntry(p),\n }),', - to: " entry: store.claimTaskBoardEntry(p),", + derive: { + within: "claim", + operator: "replace-call", + call: "coordinatedEntryWrite", + in: 'verb: "claim"', + }, + to: "store.claimTaskBoardEntry(p)", expect: "a daemon board verb routes a managed entry through the run's transition", }, { @@ -1321,14 +1325,18 @@ const TARGETS: readonly Target[] = [ }, { name: "the-grace-is-zero", - from: "export const CLOCK_GRACE_MS = 50;", - to: "export const CLOCK_GRACE_MS = 0;", + derive: { operator: "replace-initializer", variable: "CLOCK_GRACE_MS" }, + to: "0", expect: "a value stamped a moment in the future is current, not missing", }, { name: "the-grace-uses-a-unit-sqlite-does-not-know", - from: " return `strftime('%Y-%m-%dT%H:%M:%fZ', 'now', '${modifier}${seconds} seconds')`;", - to: " return `strftime('%Y-%m-%dT%H:%M:%fZ', 'now', '${modifier}${CLOCK_GRACE_MS} milliseconds')`;", + derive: { + within: "clockNow", + operator: "replace-literal-fragment", + text: "${seconds} seconds", + }, + to: "${CLOCK_GRACE_MS} milliseconds", expect: "a just-written memory is never read as not active", }, ], From 41768c83db5da5e2460288e17c8da1631689b99a Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 23:12:01 +0800 Subject: [PATCH 32/32] docs(decisions): the derived-mutants record is implemented, not proposed The four Plan items landed, and the acceptance criterion is met more strongly than it asked: the hand-written residue is zero, not small. The record proposed a doctrine; the doctrine is now the register's only form, so it moves to docs/decisions/implemented/ and takes the implemented format. - Problem keeps the motivation and the measurement that removed an option (the three-way split does not hold: every caught mutant is externally observable). - Proposal and Plan become Decision - the durable part is a name, an operator, a selector and the case that must fail - and Implementation state: the three waves, the widenings the residue forced, the corrections a run caught rather than a review, the readings, and what is not claimed. - Acceptance criteria is gone (the implemented format forbids it); its content is where the readings are. The old Risks become Consequences, together with the costs of the new form and the two questions left open on purpose. - Both languages in this commit, and the seven inbound links in docs/design/task-unit-semantics-obligations.md follow the move. - docs/design/ci-cd-and-quality.md: the anchors pass resolves 110 teeth now, not the 149 the sentence was written about, and it takes about a second. Readings at this revision: anchors 110 of 110 resolve over 22 targets; the whole register 110 of 110 caught, all 110 by the case each expect names, 22 of 22 targets restored byte-identically, exit 0 in 160 s; verify:static exit 0. --- ...-09-24-mutants-are-derived-not-anchored.md | 248 +++++++++++ ...-mutants-are-derived-not-anchored.zh-CN.md | 106 +++++ ...-09-24-mutants-are-derived-not-anchored.md | 401 ------------------ ...-mutants-are-derived-not-anchored.zh-CN.md | 138 ------ docs/design/ci-cd-and-quality.md | 2 +- .../design/task-unit-semantics-obligations.md | 26 +- 6 files changed, 368 insertions(+), 553 deletions(-) create mode 100644 docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md create mode 100644 docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md delete mode 100644 docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md delete mode 100644 docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md diff --git a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md new file mode 100644 index 00000000..23995016 --- /dev/null +++ b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md @@ -0,0 +1,248 @@ +# A tooth is an operator over a symbol, not a copy of a line + +[中文](2026-09-24-mutants-are-derived-not-anchored.zh-CN.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [Tests do not need a filesystem](2026-09-20-tests-need-no-filesystem.md), [The checks read a live mutant](../../postmortem/0003-checks-read-a-live-mutant.md), [Bound agent verification as one run](2026-09-23-verification-whole-run-deadline.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md), [Mechanism, not policy](../proposed/2026-09-21-mechanism-not-policy.md) + +## Problem + +Every rule this repository has decided to protect is protected the same way: a named mutant in +`tools/mutation-teeth.ts` that replaces a byte range with a hand-written replacement, plus the name of +the case that must fail when it does. That evidence is honest - a mutant is the only form that speaks +about the implementation's freedom to be wrong - and it is also the most fragile artifact in the tree, +because **the anchor is the mutant's identity**. It is a copy of a line, so the line moving, being +inlined, or having its message text extracted retires the tooth. + +Measured on 2026-09-24, at 149 mutants over 22 targets: + +| What the anchor points at | Mutants | Share | Already scoped to a member | +| ----------------------------------------------------- | ------- | ----- | -------------------------- | +| A guard or predicate term neutralised (G) | 64 | 43% | 30 | +| A value, argument, index or callee substituted (V) | 58 | 39% | 26 | +| A statement or call removed (R) | 11 | 7% | 6 | +| Code inserted: a statement, branch or second copy (I) | 8 | 5% | 4 | +| An expression replaced wholesale (E) | 8 | 5% | 5 | + +Two readings decided this record. **89% of the teeth (133 of 149) were a site plus one of a handful of +operators** - disable a condition, drop a term, substitute an argument, remove a statement - so the +operator is the intent and the site is incidental; storing the site as text was a choice, not a +requirement. And **78 of 149 (52%) had no structural scope at all**: a byte fragment matched anywhere in +the file, which is exactly the shape that dies first. Both failures observed while planning were of that +shape: a statement whose inner call was inlined, and a refusal whose message text was extracted into a +`subject` expression. A third form showed up beside them: two teeth whose named case no longer caught +them (`fusion-continues-from-an-unverified-answer`, `next-is-not-the-head-of-the-ordered-candidates` read +as "caught by the suite, not the named case"), which is the same coupling one layer out - the anchor +held, the sentence "this case is the one that fails" went stale. + +A measurement taken while planning also had to be reported, because it removed an option: the three-way +split "derivable / expressible as input / needs an internal perturbation" **does not hold**. A targeted +mutant is by definition _observable from outside_ - that is what "caught" means - so +"input-expressible" was true of all 149. The three cases predicted to need an internal perturbation all +had an external observation: the status read path opening a writable handle is observed by the store +file not changing, a second copy of the acceptance rule by a verdict whose digest does not match the +artifact, and a transaction wrapper by a plan whose third task is refused while the first two stay +frozen. So the replacement direction was not "input checks instead of mutants" but **derive the mutant +instead of storing it**. + +## Decision + +**A mutant's durable part is its name, its operator, its selector and the case that must fail. The site +and the replacement bytes are computed from the syntax tree on every run.** + +- **The operator is the mutation class; the selector is the site.** A tooth says + `condition-never(within: judgeTaskBoardEntry, condition: existing.deliveredBy)`, not two lines of + source. The tool resolves the range and writes the bytes, so a rename, a reflow or a lifted statement + leaves the tooth aimed at the same rule, and every replacement is computed rather than stored - `to` + exists only where the new value is a _choice_ (a different argument, property, initializer, iterable, + index, callee or literal fragment). +- **The vocabulary is the one the catalogues name**, not one invented here: pitest's + `NEGATE_CONDITIONALS`, `REMOVE_CONDITIONALS`, `VOID_METHOD_CALLS`, `PRIMITIVE_RETURNS`; + Stryker's `ConditionalExpression`, `EqualityOperator`, `ArrayDeclaration`, `MethodExpression`, + `BlockRemoval`; cargo-mutants' patterns. A site with no operator behind it is a **missing operator**, + not a permanent hand-written anchor - that is what turned the last residue of 15 into the last three + operators plus three widenings. +- **A selector that is ambiguous is refused, never guessed at**: more than one matching member, call, + property, statement, comparison, literal or element access is a failure, and so is a fragment that + names only part of a condition the operator replaces whole (a term has its own operator). +- **Scope is named; module level is a scope too.** A selector is normally scoped to one member - a + method, a function declaration, a class constructor under the name `constructor`, or a function bound + to a variable or to an object property. Where there is no member, `within` is omitted and the whole + file is the scope, with the selector required to be unique in it. +- **The anchors pass is part of the static contract.** `npm run mutation:anchors` (`--anchors-only`) + resolves every tooth without running a suite, and is one of `verify:static`'s checks and one of the + `ci-and-tests` route's atomic checks. The full sweep stays out of the gate and remains the standing + rule before a push, because the pass proves a site still resolves, not that the mutant is still + caught. +- **The evidence standard is clarified, not loosened.** "How a row earns `proven`" already requires a + test that fails when the rule is broken; the clarification is that the demonstration may be a code + perturbation (derived or hand-written) or an input-side case whose assertion contrasts two inputs or + enumerates a bounded space. A row that used to name a tooth may name a contrast check instead, in the + same commit that retires the tooth. +- **A tooth may be retired when a check states its rule.** The criterion: the sweep first shows the + tooth's named case is the one that fails under it, **and** reading that case's assertion shows it fails + for exactly this violation and states the rule - a count, an enumeration, a relation read back as a + raw row, or a refusal by name. A behavioural differential does not qualify. Two things that look like + witnesses and are not: a fixture derived from the constant under test, and a case that refuses for a + second reason as well. + +## Implementation state + +**The register is 110 teeth over 22 target entries covering 21 files, every one of them a name plus an +operator plus a selector, with no hand-written anchor left. 26 operators. 37 teeth have been retired in +favour of checks that state their rule.** + +`tools/mutation-anchor.ts` holds the resolver: `Mutant`, `Derive`, `Site`, `matchText`, `locate`, and a +table of resolvers, one per operator, so adding an operator is an entry in that table plus its name in +the type union. It holds no state and reads no file, so the sweep and the anchors-only pass ask the same +function the same question and `tests/tools/mutation-anchor.test.ts` proves it over source strings - 34 +cases, one per operator plus every refusal, with no filesystem access. Where the mutation determines its +own bytes (`false`, `true`, the negated operator, the call's receiver, the guard's body, the empty +collection) nothing is stored. + +| operator | what the catalogues call it | teeth | +| ------------------------------------------------------------- | --------------------------------------------------------------------------- | ---------- | +| `condition-never`, `condition-holds` | pitest `FALSE_RETURNS` / `TRUE_RETURNS`; Stryker `ConditionalExpression` | 45 + 3 | +| `neutralize-term` | pitest `NEGATE_CONDITIONALS`; Stryker logical/boolean literals | 9 | +| `negate-condition` | pitest `NEGATE_CONDITIONALS`; Stryker boolean literals | 2 | +| `negate-comparison` | pitest `NEGATE_CONDITIONALS`; Stryker `EqualityOperator` | 3 | +| `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | +| `drop-statement` | Stryker `BlockRemoval`; pitest `VOID_METHOD_CALLS` | 8 | +| `remove-call` | Stryker filter/slice/sort removals; pitest `VOID_METHOD_CALLS` | 2 | +| `replace-call` | Stryker `MethodExpression`; pitest `CONSTRUCTOR_CALLS` | 4 | +| `replace-argument`, `replace-property`, `replace-initializer` | pitest `PRIMITIVE_RETURNS`, `INLINE_CONSTS`; Stryker's literal mutators | 8 + 10 + 7 | +| `replace-iterable` | Stryker `ArrayDeclaration`; pitest `EMPTY_RETURNS` | 3 | +| `replace-index` | (no catalogue entry; the nearest, `FirstToLast`, is a method named `first`) | 2 | +| `replace-literal-fragment` | SQLMutation's clause-level operators, at fragment granularity | 3 | + +Three waves converted the register, each swept after it: 59 teeth with the six operators the resolver +started with, 18 with the seven the catalogues name, and the last 15 with three more operators and three +widenings. The widenings were all found by real teeth failing to convert, never by argument: + +| what the residue needed | what was added | teeth | +| ---------------------------- | ----------------------------------------------------------------------------------------------------------------- | ----- | +| a function bound to a name | `uniqueMember` accepts a function bound to a variable or to an object property (`claim: (store) => ...`) | 2 | +| two identical call sites | `in` names a fragment of the statement a call or a literal sits in | 3 | +| module level | `within` is optional: the file is the scope, the selector must be unique in it | 2 | +| a constructor call | the call selectors accept `new X(...)` as a call site | 1 | +| a shorthand property | `replace-property` writes it out (`position` becomes `position: 0`), because a value cannot go where the name was | 1 | +| an index, a literal fragment | `replace-index`, `replace-literal-fragment` | 5 | + +**Readings, at this revision:** `anchors: 110 of 110 resolve, over 22 targets`; `mutants: 110 of 110 +caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored +byte-identically, exit 0 in 160 s; `tests/tools/mutation-anchor.test.ts` 34 pass; `test:product` 1580 +pass; `verify:static` exit 0; lint 0 findings; complexity gate ok. An earlier full-run reading on the +same arc: 136 teeth, 136 of 136 caught, all 136 by name, 217 s. + +**Corrections this record carries, each caught by a run rather than by review:** + +- The pilot's refactor prediction was wrong: a condition lifted into `const closed = spent || +severalWaits` retired a derived tooth as well as the hand anchor, because a condition bound to a name + was not a decision position for the selector. The selector was widened to include a variable's + initializer, and the corrected prediction then held exactly. +- The first whole-register sweep said 136 of 136 caught but only **134 by the case each `expect` named**. + Four links were wrong: a case that refused the pair for a second reason as well; an `expect` naming a + case in another target; an `expect` that was a paraphrase naming no case; and a named case whose + fixture was derived from the constant under test (`half = CLOCK_GRACE_MS / 2000`), so zeroing the + constant moved the stamp and the case passed while the bug was live. +- A wrong `within` was shipped twice. The quiet one: `the-completion-ignores-a-cancelled-unit` aimed at + `checkDispatch` instead of `checkCompletion` - the anchors pass resolved, the mutant was caught, but by + the _suite_ rather than by the case that names a completion of a cancelled unit. The loud one: a guard + aimed at `instrumentCommit` instead of `runParentCheck`, refused immediately as "0 guards in + instrumentCommit match the selector". **A wrong member is loud when it finds nothing and silent when it + finds the same shape twice**, which is why the sweep after a conversion is not optional. +- Three reasons recorded for _not_ converting the last 15 were measured wrong: the catalogue's negation + (`=== 1` to `!== 1`) does realise the violation the name states; `every-task-is-frozen-at-position-zero` + changes only one load-bearing thing, because `noUnusedParameters` is not set; and the two teeth that + "shared a line" were only ever hand anchors that had to span two statements - the mutations are single + statement removals. The lesson kept: attempt every site with the operators available and measure, + before writing down a reason not to. +- A hand-typed `expect` reads as a broken mutant. Retyping two case names from memory instead of copying + them made both teeth report **"survived"**: the tool filters the suite by the named case, so a name + that matches no case is indistinguishable from a mutant nothing catches. It is a loud failure, never a + false pass, and it is the second time this register has been bitten by `expect` being an assertion of + its own. Conversion scripts now copy the string from the register. +- Rerunning a conversion script after the vocabulary had grown rewrote two teeth back to their earlier + form; the anchors pass refused both at once. That is the case for the pass being in the static + contract seen from the other side: it also catches edits to the register, not only drift in the code it + points at. + +**Retirements.** 37 teeth went, in four groups, each only after the sweep showed the named case was the +one that failed under it and the case's assertion was read. Eleven in the first two groups (the two +call-site counts behind row B5, the 32-combination enumeration, the multinomial count of 10 behind row +A5, a rebinding refused by name, and the five rules of the declared budget stated by one case). Two in +the third (row D11, whose prose already rested on the cases). 26 in the fourth: the insertion-shaped +teeth, whose anchor is a place where code must not appear, and which a derived selector cannot express - +the only honest alternatives were a permanent fragile anchor or a check that states the rule. One of that +set was kept and then converted, and seven were **orphans** - no ledger row named them, so the rule each +pinned has a check and no row. + +**What is not claimed.** A derived tooth is not stronger evidence than the byte anchor it replaced: it is +the same violation, caught by the same case. What it buys is that the tooth stays aimed at its rule +across renames and reflows, which is the failure this record exists for. The operators are not the whole +catalogue either - they are the classes this register's rules happen to need. And a conversion is not a +licence to stop sweeping: `mutation:anchors` proves resolution, not a live catch. + +**What the catalogues say, kept deliberately.** pitest's _Less is more_ names the effect the retirement +criterion measures: a mutant the combination of others already subsumes adds runtime, not confidence. +pitest avoids them with a fixed default set; this register measures it. cargo-mutants counts a mutant +that does not compile separately and documents why it refuses some mutations (`==` to `<` "too prone to +generate false positives", `-a` to `+a` "too prone to generate unviable cases") - the same argument used +here for not adding operators for shapes that are rare. The one thing this design does not borrow is the +definition of a kill: those tools count a mutant killed when _any_ test fails; here a tooth is counted +only when **the case it names** is the one that fails, which is why "caught by the suite, not by the case +that names the rule" is a defect this register reports and none of those tools would. + +## Alternatives considered + +- **Keep hand-written anchors and make them smaller.** Rejected as the whole answer: it is a discipline, + not a mechanism, and the discipline is what 52% of the teeth violated. It survives as a rule for the + replacement bytes, which stay minimal on purpose: no message text, no sibling argument, no whole + statement where a term does. +- **Anchor by AST node only, keeping byte replacements** (the older `ast.within` plus a text pair). + Rejected as insufficient: measured, 71 of 149 teeth already had that, and two teeth died inside it, + because the _replacement_ and the _fragment_ were still copies of lines. +- **Replace mutants with input-side checks wholesale.** Rejected on the measurement above: it does not + discriminate (every caught mutant is externally observable), and it would drop the only evidence that + speaks about an unforeseen implementation slip rather than a declared violation. +- **Insert the mutation into the code behind a runtime switch** (mutant schemata, `mutation_active("...")` + guards). Rejected for this repository: the guard is real code in `src/`, and a switched-off mutation + path is a policy word in the mechanism layer, which [mechanism, not + policy](../proposed/2026-09-21-mechanism-not-policy.md) forbids. +- **Auto-generate mutants from operators over the whole tree** (what pitest, StrykerJS and cargo-mutants + do). Rejected as the form here: it would replace named evidence with a score, and the ledger needs a + named tooth per row. The derived form keeps the operator idea and the naming. +- **Leave the sweep out of every gate and rely on the standing rule.** Rejected: the measured case for + this record is that nobody noticed two teeth had stopped biting, and a rule that depends on being + remembered is the shape this repository already replaced elsewhere. + +## Consequences + +- **The register's cost is now a vocabulary rather than a maintenance duty.** A reader of a `from`/`to` + pair sees the change; a reader of `condition-never(within: judgeTaskBoardEntry, condition: +existing.deliveredBy)` has to know the operator. Mitigation: the operators are few and fixed, each one + is documented beside its type with the catalogue it comes from and the selector it needs, and the tool + prints the bytes it wrote when asked. +- **A derived selector is still scoped to a decision position, so a rule rewritten into a different shape + can still retire its tooth.** The difference is that it now dies visibly: the anchors pass names the + tooth and the fragment it could not resolve, which is exactly the failure that went unnoticed before. +- **A wrong selector mutates the wrong node, and a mutant that does not compile is not a caught tooth.** + Mitigation: exactly one match or refusal, plus the two readings the sweep already takes (a clean run + must pass, and the named case must be the one that fails - a mutant that only breaks the build is + reported as caught by the suite, not by the named case). +- **Retiring a tooth can lose a named case.** A relational check may catch a class where the tooth caught + a specific wrong behaviour, so a row could claim more than it shows. Mitigation: a retirement names the + check that replaces it, in the ledger row and in the commit that retires the tooth. +- **A conversion is evidence surgery.** Every conversion touches the artifact that says the code is + protected, so a mistake reduces coverage quietly. Mitigation: one target at a time, the same names and + cases, the catch results recorded before and after, and the post-mortem rule about a live mutant + (status the file under test before any run). +- **The pass can be trusted instead of run.** An anchors-only pass proves the site still resolves, not + that the mutant is still caught, and a tooth whose `expect` went stale still passes it. Mitigation: the + full sweep remains the standing rule before a push, and the sweep is the reading that counts - not the + number of teeth, but **the number of teeth whose named case is the one that fails**. +- **Open, and deliberately not decided here:** the seven orphan teeth whose rules have a check and no + ledger row (whether those rules deserve rows is a question about the ledger, not about the register); + and the `expect` field's own failure mode, which is loud but only on the slow path - a report that says + whether the named case exists in the target's suites would make it immediate. diff --git a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md new file mode 100644 index 00000000..ccfa4345 --- /dev/null +++ b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -0,0 +1,106 @@ +# 牙是"作用于符号的算子",不是"某一行的一次拷贝" + +[English](2026-09-24-mutants-are-derived-not-anchored.md) + +**Status:** implemented +**Approved:** explicit +**Relates to:** [测试不需要文件系统](2026-09-20-tests-need-no-filesystem.md)、[检查读到的是一个还活着的 mutant](../../postmortem/0003-checks-read-a-live-mutant.md)、[把 agent 验证限制为一次运行](2026-09-23-verification-whole-run-deadline.md)、[契约的义务](../../design/task-unit-semantics-obligations.md)、[机制,而非策略](../proposed/2026-09-21-mechanism-not-policy.md) + +## 问题 + +本仓库决定保护的每一条规则,都用同一种方式保护:`tools/mutation-teeth.ts` 里一个具名 mutant,把一段字节换成手写的替换内容,外加"当它存在时哪条用例必须失败"的用例名。这种证据是诚实的——只有 mutant 这种形式能谈"实现本来还可以怎么写错"——同时也是整棵树里最脆的东西,因为**锚点就是这颗 mutant 的身份**。它是某一行的一次拷贝,所以那一行被移动、被内联、消息文本被抽出去,这颗牙就退役了。 + +2026-09-24 实测,当时 149 颗 mutant、22 个目标: + +| 锚点指向什么 | 颗数 | 占比 | 已经限定到成员 | +| ---------------------------------------------- | ---- | ---- | -------------- | +| 守卫或谓词里的一项被中和(G) | 64 | 43% | 30 | +| 值、实参、下标或被调被替换(V) | 58 | 39% | 26 | +| 语句或调用被删除(R) | 11 | 7% | 6 | +| 插入代码:一条语句、一个分支或第二份拷贝(I) | 8 | 5% | 4 | +| 整个表达式被替换(E) | 8 | 5% | 5 | + +两个读数决定了本记录。**89% 的牙(149 之 133)是"一个点位 + 少数几个算子之一"**——让条件失效、中和一项、替换实参、删除语句——也就是算子才是意图、点位是附带,把点位存成文本只是一个选择而不是必需。而且 **149 之 78(52%)完全没有结构性作用域**:一个在文件里任意匹配的字节片段,正是最先死掉的形状。规划期间观察到的两处失效都是这个形状:一条语句里层调用被内联、一处拒绝的理由消息被抽成 `subject` 表达式。旁边还出现了第三种:两颗牙点名的用例已经抓不住它们(`fusion-continues-from-an-unverified-answer`、`next-is-not-the-head-of-the-ordered-candidates` 读数变成"被整个 suite 抓到,不是点名用例"),这是同一层耦合往外挪了一格——锚点还在,而"这条用例是唯一会失败的那条"这句话旧了。 + +规划期间做的另一个测量也必须写下,因为它砍掉了一个选项:三分法"可推导 / 可用输入表达 / 需要内部扰动"**不成立**。定向 mutant 按定义就是**从外部可观察**的——"抓到"就是这个意思——所以"可用输入表达"对 149 颗全部成立。被预测为"需要内部扰动"的三处都有外部观察:状态读路径打开可写句柄,可以由 store 文件没有变化观察到;第二份验收规则,可以由一个 digest 与产物不匹配的裁决观察到;事务包装,可以由"第三个任务被拒绝而前两个仍处于冻结"观察到。所以替换方向不是"用输入检查替代 mutant",而是**把 mutant 推导出来而不是存下来**。 + +## 决策 + +**一颗 mutant 的耐久部分是它的名字、它的算子、它的选择器,以及必须失败的那条用例。点位和替换字节每次运行时从语法树算出来。** + +- **算子是变异类,选择器是点位。**一颗牙写的是 + `condition-never(within: judgeTaskBoardEntry, condition: existing.deliveredBy)`,而不是两行源码。工具解析范围并写入字节,所以改名、重排、语句被上提之后,这颗牙仍然瞄着同一条规则;每一处替换都是算出来的而不是存下来的——只有当新值是一个**选择**时才需要 `to`(换哪个实参、属性、初始化器、可迭代、下标、被调或字面量片段)。 +- **词汇就是现成目录点名的那些**,不是这里发明的:pitest 的 `NEGATE_CONDITIONALS`、`REMOVE_CONDITIONALS`、`VOID_METHOD_CALLS`、`PRIMITIVE_RETURNS`;Stryker 的 `ConditionalExpression`、`EqualityOperator`、`ArrayDeclaration`、`MethodExpression`、`BlockRemoval`;cargo-mutants 的 patterns。**没有算子能点名的点位,是"缺了一个算子",而不是永久的手写锚点**——正是这条把最后的 15 颗残留变成了最后三个算子加三处扩宽。 +- **选择器有歧义就拒绝,绝不猜**:成员、调用、属性、语句、比较、字面量或元素访问匹配到多于一个就是失败;被整条件算子替换、而片段只指名条件一部分的,也拒绝(一项有它自己的算子)。 +- **作用域要命名;模块级也是一种作用域。**选择器通常限定在一个成员内——方法、函数声明、名字为 `constructor` 的类构造函数,或绑定到变量、绑定到对象属性的函数。没有成员可命名时省略 `within`,整个文件就是作用域,并要求选择器在文件里唯一。 +- **anchors pass 是静态契约的一部分。**`npm run mutation:anchors`(`--anchors-only`)不跑任何用例就解析全部牙,是 `verify:static` 的检查之一,也是 `ci-and-tests` 路由的原子检查之一。全量 sweep 不进闸门,仍是推送前的常设规则——因为这个 pass 证明的是"点位还解析得出来",不是"mutant 还抓得住"。 +- **证据标准是被澄清,不是被放松。**"一行怎么才算 `proven`"本来就要求"规则被破坏时有用例失败";澄清之处在于:这个演示可以是代码扰动(推导的或手写的),也可以是输入侧用例,只要它的断言对着两个输入做对比、或枚举了一个有界空间。原先指向某颗牙的行,可以在退役那颗牙的同一个提交里改为指向那条对比检查。 +- **当某条检查已经陈述了规则,那颗牙可以退役。**标准是:先由 sweep 说明"在这颗牙之下失败的就是它点名的用例",**并且**读那条用例的断言,确认它失败的原因恰好是这个违规、并且把规则说了出来——一个计数、一次枚举、把原始行读回来做关系判断、或按名字拒绝。行为差分不算数。有两样东西看着像见证人其实不是:**从被测常量推出来的 fixture**,以及**同时因为第二个理由而拒绝的用例**。 + +## 落地情况 + +**登记处现在是 110 颗牙、22 个目标条目、覆盖 21 个文件,每一颗都是"名字 + 算子 + 选择器",没有一颗手写锚点。26 个算子。已有 37 颗牙退役,改为由陈述了该规则的检查承担。** + +`tools/mutation-anchor.ts` 是解析器:`Mutant`、`Derive`、`Site`、`matchText`、`locate`,以及一张"一个算子一个解析函数"的表,所以加一个算子就是表里加一行、类型里加一个名字。它不持有状态、不读文件,因此 sweep 与 anchors-only pass 问的是同一个函数的同一个问题,而 `tests/tools/mutation-anchor.test.ts` 用源码字符串证明它——34 条用例,每个算子一条并覆盖各自的拒绝情形,不碰文件系统。凡是变异自己决定字节的(`false`、`true`、取反后的算子、调用的接收者、守卫的 body、空集合),什么都不存。 + +| 算子 | 现成目录里的名字 | 颗数 | +| ------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------- | ---------- | +| `condition-never`、`condition-holds` | pitest `FALSE_RETURNS` / `TRUE_RETURNS`;Stryker `ConditionalExpression` | 45 + 3 | +| `neutralize-term` | pitest `NEGATE_CONDITIONALS`;Stryker 逻辑/布尔字面量 | 9 | +| `negate-condition` | pitest `NEGATE_CONDITIONALS`;Stryker 布尔字面量 | 2 | +| `negate-comparison` | pitest `NEGATE_CONDITIONALS`;Stryker `EqualityOperator` | 3 | +| `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | +| `drop-statement` | Stryker `BlockRemoval`;pitest `VOID_METHOD_CALLS` | 8 | +| `remove-call` | Stryker filter/slice/sort 删除;pitest `VOID_METHOD_CALLS` | 2 | +| `replace-call` | Stryker `MethodExpression`;pitest `CONSTRUCTOR_CALLS` | 4 | +| `replace-argument`、`replace-property`、`replace-initializer` | pitest `PRIMITIVE_RETURNS`、`INLINE_CONSTS`;Stryker 字面量类 | 8 + 10 + 7 | +| `replace-iterable` | Stryker `ArrayDeclaration`;pitest `EMPTY_RETURNS` | 3 | +| `replace-index` | (目录里没有;最接近的 `FirstToLast` 指的是名为 `first` 的方法) | 2 | +| `replace-literal-fragment` | SQLMutation 的子句级算子,降到片段粒度 | 3 | + +登记处分三波转换,每一波之后都 sweep:用解析器最初那六个算子转了 59 颗,用目录点名的七个算子转了 18 颗,最后 15 颗用另外三个算子加三处扩宽转完。三处扩宽全都是**真实的牙转不过去**才发现的,没有一处是靠讨论定下来的: + +| 残留需要什么 | 实际加了什么 | 颗数 | +| -------------------- | ------------------------------------------------------------------------------------------------------------------- | ---- | +| 绑定到名字的函数 | `uniqueMember` 接受绑定到变量或对象属性的函数(`claim: (store) => ...`) | 2 | +| 两个一模一样的调用点 | `in` 指"调用或字面量所在语句里的一段文本" | 3 | +| 模块级 | `within` 变可选:文件就是作用域,选择器必须在文件里唯一 | 2 | +| 构造函数调用 | 调用选择器接受 `new X(...)` 作为调用点 | 1 | +| 简写属性 | `replace-property` 把它写开(`position` 变 `position: 0`),因为值不能写在名字的位置上 | 1 | +| 下标、字面量片段 | `replace-index`、`replace-literal-fragment` | 5 | + +**本版读数:**`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,160 秒;`tests/tools/mutation-anchor.test.ts` 34 条全过;`test:product` 1580 全过;`verify:static` exit 0;lint 0 findings;复杂度门禁通过。同一条弧上更早的一次全量读数是:136 颗牙,136 of 136 caught,136 全部按名字,217 秒。 + +**本记录携带的更正,每一条都是被运行结果抓出来的、不是被评审看出来的:** + +- 试点的重构预测是错的:把一个条件上提成 `const closed = spent || severalWaits` 之后,除了手写锚点,还顺带退役了一颗 derived 牙,因为"绑定到名字的条件"当时不算决策位置。选择器因此扩宽到包含变量初始化器,改后的预测随后精确成立。 +- 第一次全量 sweep 说 136 of 136 caught,但只有 **134 颗是按各自 `expect` 点名的用例抓到的**。四处链接是错的:一条用例同时因为第二个理由拒绝那对单元;一处 `expect` 点名的是另一个目标里的用例;一处 `expect` 是句改写、根本没点名任何用例;还有一条点名用例的 fixture 是从被测常量推出来的(`half = CLOCK_GRACE_MS / 2000`),于是把常量归零后时间戳被挪走、bug 还活着用例却通过了。 +- `within` 写错发生过两次。安静的那次:`the-completion-ignores-a-cancelled-unit` 指向了 `checkDispatch` 而不是 `checkCompletion`——anchors pass 照样通过、mutant 照样被抓,但抓它的是**整个 suite**,不是那条"报告一个被取消单元的完成"的用例。响的那次:守卫在 `runParentCheck` 我写了 `instrumentCommit`,立刻被拒——"0 guards in instrumentCommit match the selector"。**成员名写错,找不到时是响的、找到两处同形状时是哑的**,这就是"转完必须 sweep"不可省的理由。 +- 最后 15 颗当初"不转"的三条理由,实测都是错的:目录里的取反(`=== 1` → `!== 1`)**确实**实现了名字所说的违规;`every-task-is-frozen-at-position-zero` 只有一处承重,因为 `noUnusedParameters` 没开;而两颗"两条语句挤一行"从来只是手写锚点不得不横跨两条语句——变异本身是删掉一条语句。留下的教训:**把理由写下来之前,先拿现有算子去试、去量**。 +- `expect` 手打,读起来就像一颗坏掉的牙。两颗牙凭记忆重打用例名(而不是从登记处抄),结果都报 **"survived"**:工具按点名的用例过滤 suite,名字对不上任何用例时,与"没有任何用例抓到"无法区分。它是响的失败、从不是误判通过,但这是这个登记处第二次被"`expect` 自己也是一条断言"咬到。转换脚本现在从登记处**抄**这个字符串。 +- 词汇长大之后重跑旧的转换脚本,把两颗牙改回了早期形态;anchors pass 一次性把两个都拒了。这正是"把这个 pass 放进静态契约"从另一个方向看到的理由:它也抓**对登记处本身的编辑**,不只抓它指向的代码漂移。 + +**退役。** 37 颗牙走了,分四组,每组都先由 sweep 说明"失败的就是它点名的用例"、再读那条用例的断言。前两组十一颗(B5 行背后的两处调用点计数、32 种组合的枚举、A5 行背后的多项式计数 10、按名字拒绝的重新绑定,以及由一条用例陈述的预算五规则)。第三组两颗(D11 行,其正文本来就建立在对应用例上)。第四组 26 颗:插入形状的牙——它们的锚点是"代码不得出现的地方",derived 选择器表达不了,诚实的选择只有"永久脆弱的锚点"或"一条陈述了规则的检查"。这一组里有一颗后来也转掉了,另有七颗是**孤儿**——没有任何账本行点名它们,于是它们钉的规则有检查、没有行。 + +**没有声称的东西。** derived 牙**不比**它替代的字节锚点更强:违规是同一个、抓它的用例是同一条。它换来的是改名、重排之后仍瞄准原来那条规则,而这正是本记录要解决的那个失败。算子也不等于整个目录——它们只是本登记处的规则恰好需要的那些类。转换也不是"可以不再 sweep"的许可:`mutation:anchors` 证明的是解析成功,不是"还抓得住"。 + +**有意带过来的目录结论。** pitest 的《Less is more》给退役标准所量的效应起了名:一个已被其他 mutant 组合覆盖的 mutant,带来的是运行时间而不是信心。pitest 用固定默认集合回避它,本登记处是**量**它。cargo-mutants 把编译不过的 mutant 单独计数,并写明为什么某些变异不生成(`==` 换成 `<`"太容易产生假阳性"、`-a` 换成 `+a`"太容易不编译")——与这里"不为罕见形状加算子"是同一条论据。唯一**不借用**的是"消灭"的定义:那些工具是"任意一条测试失败"就算 killed;这里只有**点名的那条**用例失败才算数,所以"被整个 suite 抓到、却没被点名用例抓到"是本登记处会报、而那些工具一个都不会报的缺陷。 + +## 考虑过的替代方案 + +- **继续手写锚点、只是写得更小。**作为整体答案被否:那是纪律而不是机制,而纪律正是 52% 的牙当时违反的东西。它作为"替换字节"的规则保留下来,并且是刻意最小:不要消息文本、不要兄弟实参、能用一项时不要整条语句。 +- **只按 AST 节点锚定,替换字节照旧**(早先的 `ast.within` 加一对文本)。被否,理由是实测:149 颗里已有 71 颗是这个形态,而本轮仍有两颗死在它上面——因为**替换内容**和**片段**仍是对行的拷贝。 +- **整体改用输入侧检查替代 mutant。**按上面的测量被否:它没有区分力(每一颗被抓到的 mutant 从外部都可见),而且会丢掉唯一一种能谈"预料之外的实现疏忽"而非"已声明违规"的证据。 +- **把变异插进代码、放在运行时开关之后**(mutant schemata、`mutation_active("...")` 守卫)。本仓库否掉:那个守卫是 `src/` 里的真实代码,而"被关掉的变异路径"是机制层里的策略词,[机制,而非策略](../proposed/2026-09-21-mechanism-not-policy.md)禁止这样做。 +- **对整个语法树按算子自动生成 mutant**(pitest、StrykerJS、cargo-mutants 的做法)。作为本仓库的形式被否:它会把具名证据换成一个分数,而账本每一行都需要一颗具名的牙。derived 形态保留了算子这个想法和"具名"。 +- **干脆不进任何闸门、只靠常设规则。**被否:本记录的实测就是"没人注意到两颗牙已经停摆",而"依赖被记住"的规则正是本仓库在别处已经替换掉的形状。 + +## 后果 + +- **登记处的成本,从"维护义务"变成了"一套词汇"。**读 `from`/`to` 对的人直接看到改动;读 `condition-never(within: judgeTaskBoardEntry, condition: existing.deliveredBy)` 的人得知道那个算子。缓解:算子少而固定,每一个都在其类型旁写明它来自哪个目录、需要什么选择器;需要时工具会打印它写入的字节。 +- **derived 选择器仍然限定在"决策位置",所以规则被改写成另一种形状时,牙仍可能退役。**不同之处是它现在**死得看得见**:anchors pass 会点名那颗牙和它解析不出的片段——而这正是此前无人察觉的那种失效。 +- **选择器写错会变异错的节点;而编译不过的 mutant 不算"抓到的牙"。**缓解:匹配必须恰好一个,否则拒绝;再加 sweep 本来就取的两个读数(干净树必须通过;点名用例必须是失败的那条——只把构建弄坏的 mutant 会被报成"被 suite 抓到,不是点名用例")。 +- **退役会丢掉一条具名用例。**关系型检查抓的是一类,而牙抓的可能是某个具体错误行为,于是行可能声称得比展示的多。缓解:退役时在账本行里、在同一个提交里点名替代它的那条检查。 +- **转换是"对证据动手"。**每一次转换都触碰"这份代码被保护着"的那件证物,因此错误会安静地降低覆盖。缓解:一次一个目标、同名同用例、前后都记下抓取结果,以及那条关于活 mutant 的 postmortem 规则(任何运行前先看被测文件的状态)。 +- **这个 pass 可能被信任而不是被运行。**anchors-only pass 证明点位还解析得出来,不证明 mutant 还抓得住;一颗 `expect` 已经陈旧的牙照样能通过它。缓解:全量 sweep 仍是推送前的常设规则,而真正要看的读数不是牙的数量,而是**"失败的就是它点名那条用例"的牙的数量**。 +- **开放、且刻意不在这里决定:**那七颗孤儿牙——它们钉的规则有检查、没有账本行(这些规则该不该有自己的行,是账本的问题而不是登记处的问题);以及 `expect` 字段自身的失效模式——它是响的,但只在慢路径上;一份"点名用例在目标套件里存在与否"的报告能让它立刻可见。 diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md deleted file mode 100644 index 273719bb..00000000 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md +++ /dev/null @@ -1,401 +0,0 @@ -# A tooth is an operator over a symbol, not a copy of a line - -[中文](2026-09-24-mutants-are-derived-not-anchored.zh-CN.md) - -**Status:** proposed -**Approved:** explicit -**Relates to:** [Tests do not need a filesystem](../implemented/2026-09-20-tests-need-no-filesystem.md), [The checks read a live mutant](../../postmortem/0003-checks-read-a-live-mutant.md), [Bound agent verification as one run](../implemented/2026-09-23-verification-whole-run-deadline.md), [The contract's obligations](../../design/task-unit-semantics-obligations.md), [Mechanism, not policy](2026-09-21-mechanism-not-policy.md) - -## Problem - -Every rule this repository has decided to protect is protected the same way: a named mutant in -`tools/mutation-teeth.ts` that replaces a byte range with a hand-written replacement, plus the name of -the case that must fail when it does. That evidence is honest - a mutant is the only form that speaks -about the implementation's freedom to be wrong - and it is also the most fragile artifact in the tree, -because **the anchor is the mutant's identity**. It is a copy of a line, so the line moving, being -inlined, or having its message text extracted retires the tooth. - -Measured on 2026-09-24, at 149 mutants over 22 targets: - -| What the anchor points at | Mutants | Share | Already scoped to a member | -| ----------------------------------------------------- | ------- | ----- | -------------------------- | -| A guard or predicate term neutralised (G) | 64 | 43% | 30 | -| A value, argument, index or callee substituted (V) | 58 | 39% | 26 | -| A statement or call removed (R) | 11 | 7% | 6 | -| Code inserted: a statement, branch or second copy (I) | 8 | 5% | 4 | -| An expression replaced wholesale (E) | 8 | 5% | 5 | - -Two readings decide this record. **89% of the teeth (133 of 149) are a site plus one of a handful of -operators** - disable a condition, drop a term, substitute an argument, remove a statement - so the -operator is the intent and the site is incidental; storing the site as text is a choice, not a -requirement. And **78 of 149 (52%) have no structural scope at all**: they are a byte fragment matched -anywhere in the file, which is exactly the shape that dies first. Both failures this session were of -that shape: a statement whose inner call was inlined, and a refusal whose message text was extracted -into a `subject` expression. A third form showed up beside them: two teeth whose named case no longer -catches them (`fusion-continues-from-an-unverified-answer`, `next-is-not-the-head-of-the-ordered-candidates` -read as "caught by the suite, not the named case"), which is the same coupling one layer out - the -anchor held, the sentence "this case is the one that fails" went stale. - -A measurement taken while planning this record also has to be reported, because it removed an option: -the three-way split "derivable / expressible as input / needs an internal perturbation" **does not -hold**. A targeted mutant is by definition _observable from outside_ - that is what "caught" means - so -"input-expressible" is true of all 149. The three cases predicted to need an internal perturbation all -had an external observation: the status read path opening a writable handle is observed by the store -file not changing, a second copy of the acceptance rule by a verdict whose digest does not match the -artifact, and a transaction wrapper by a plan whose third task is refused while the first two stay -frozen. So the replacement direction is not "input checks instead of mutants"; it is **derive the -mutant instead of storing it**, with input-side checks as the form for the rules where a contrast or an -enumeration already says it better. - -## Proposal - -**A mutant's durable part is its name, its operator, its selector and the case that must fail. The site -and the replacement bytes are computed from the syntax tree on every run.** - -1. **Six operators over a named selector**, resolved inside a member: `condition-never` (the condition - the selector identifies never holds), `condition-holds` (it always holds - the mirror, because a rule - written as `return a && b` says the condition is true and replacing it with `false` would reverse the - rule instead of removing it), `neutralize-term` (one term of that condition becomes its identity - - `true` under `&&`, `false` under `||`), `replace-argument` (argument _n_ of a call becomes the - declared fragment), `replace-property` (the value of a named object property becomes it), and - `drop-statement` (the statement the selector identifies is removed with its line). Two rules keep - selectors honest: a selector that matches more than one site, or a fragment that fits two candidates, - is refused rather than guessed at; and the two whole-condition operators refuse a fragment that - names only part of a condition, because replacing all of it would silently widen the mutant into - "every reason this rule has" - a term has its own operator. Where a rule needs none of them, a - `within`-scoped minimal fragment stays available, and "minimal" is the point: no message text, no - sibling argument, no whole statement where a term does. -2. **Move the 133 G/V/R teeth to the derived form**, target by target, beginning with - `src/integration/ooo-execution.ts` as the pilot: it holds 24 teeth over two target entries (16 fusion- - legality and speculation, 8 session fusion), the same mutant names, the same `expect` cases, and the - same 24 of 24 caught after conversion - 19 of them derived, 5 staying hand-written and named below - - plus a real refactor in that file that leaves every derived site applying where a byte anchor - retires. - The five residues in the pilot file, each a single expression or value rather than a statement or a - message: `selection-ignores-a-withdrawn-acceptance` (a whole `const` initializer replaced), - `the-budget-is-not-cut-from-the-startable-set` (a returned expression's call removed), - `next-task-is-not-the-head-of-the-legal-set` (an index `[0]` becomes `[1]`), - `speculation-guesses-several-facts-at-once` (`length === 1` becomes `>= 1`), and - `a-guess-with-no-evidence-publishes` (an inserted branch plus a rewritten message). -3. **What a derived form cannot express stays hand-written and says so.** The 16 I/E teeth - insertions, - wrappers, a second copy of a rule - are the ones whose anchor is a _place where code must not - appear_, so they are the last to move and the first to be reconsidered: where the rule already has a - check whose form is relational or enumerative, the tooth is retired and the check is named in its - place. - - **Landed (2026-09-24): 26 of the 27 insertion-shaped teeth are gone.** Each went the same way: the - sweep first showed that the tooth's named case was the one that failed under it, and reading that case - showed the rule stated outright - a count over the store's own source (`the store runs its transaction -boundary in exactly one place`: one BEGIN, one COMMIT and one ROLLBACK, all three inside - `writeTransaction`), an enumeration of the accepted set, of the fallbacks or of what a freeze left - behind, a refusal by name plus a raw read back (`one reader of the same ready task is given the claim, -the second is refused` reads the owner out of the table), or the file's own hash (a read-only open - neither creates, migrates nor writes). One was kept in this set: `every-task-is-frozen-at-position-zero`, - whose named case is a register/freeze/adopt/read-back round-trip and so is weaker than "the position - comes from the array order, not from the request" - the same reason `a-binding-does-not-record-its-channel` - stays. Seven of the 26 were **orphans** - no row named them - so the rule each pinned has a check and no - row: `round-publication-opens-its-own-transaction` (`the round's own publication rolls back with the -transition that made it`), `round-never-releases-its-pin` (`cancelling a round releases the pins it held, -so nothing it referenced leaks`), `ordered-mode-becomes-any-topological-order` (`the ordered mode is the -declared plan order and nothing else`), `the-loop-awaits-each-unit-instead-of-the-batch` and - `a-unit-is-dispatched-twice-in-one-batch` (`a unit's check is outstanding while an independent unit's -worker runs`, which also asserts "each unit is dispatched once"), `a-unit-nothing-checks-is-still-a-unit` - (`evals/ooo-execution/families.test.ts`'s per-family "a unit nothing checks is refused rather than - accepted on nothing") and `the-parent-check-ignores-its-own-verdict` (`the parent check is the composed -acceptance, and a failing check is reported as such`). Whether those seven rules deserve rows of their own - is the next question this reading raises, and it is a question about the ledger rather than about the - register. - - The reading after the retirement: **110 of 110 caught by the named test, all 110 by the case their - `expect` names**, 22 of 22 targets restored byte-identically, 162 s; `anchors: 110 of 110 resolve, over 22 -targets`. The target `tools/agent-verify.ts` went with its single tooth, so the register is 22 entries - over 21 files. - -4. **A check that a tooth still applies becomes part of the static contract**: an anchors-only pass that - resolves every mutant, without running a suite, in single-digit seconds, failing when a site cannot be - resolved, resolves more than once, or when a target claims more teeth than it can apply. The - `ci-and-tests` route then lists it beside the other atomic checks, and the whole-run deadline - ([bound agent verification](../implemented/2026-09-23-verification-whole-run-deadline.md)) is a - constraint on it: the pass is cheap, the full sweep stays out of the gate. -5. **The evidence standard is clarified, not loosened.** "How a row earns `proven`" already says a test - that fails when the rule is broken; the clarification is that the demonstration may be a code - perturbation (derived or hand-written) or an input-side case whose assertion contrasts two inputs or - enumerates a bounded space. A row that used to name a tooth may name a contrast check instead, in the - same commit that retires the tooth. - -## Plan - -1. This record, plus the tool's operator vocabulary and its resolver, proven by tests over source - strings rather than by a filesystem. **Landed:** the resolver is `tools/mutation-anchor.ts` (it held - no state, so it moved out of the sweep script and the tests can call it), with 16 cases over source - strings and no filesystem access. - **Extended to the whole register (2026-09-24), in three waves: 59 teeth, then 18, then the last 15 - - and the register is now 110 derived of 110.** First pass: every class of site the - six operators already had, taken target by target and swept after each - `condition-never` on a - guard or an initializer (44 in total), `neutralize-term` (9), `replace-property` (9), - `replace-argument` (6), `drop-statement` (6), `condition-holds` (3). - Second pass: the 33 that were left were classified by what they would need, and the classification - was checked against what mutation tools actually name (Stryker's supported mutators, pitest's - mutator list, cargo-mutants' patterns, Cosmic Ray's operator concept). Thirteen of those sites fell - into mutation classes every catalogue carries, so the vocabulary grew by seven operators rather than - the class being declared impossible: - - | operator | what the catalogues call it | teeth | - | --------------------- | ----------------------------------------------------------------------- | ----- | - | `negate-condition` | pitest `NEGATE_CONDITIONALS`; Stryker boolean literals | 2 | - | `negate-comparison` | pitest `NEGATE_CONDITIONALS`; Stryker `EqualityOperator` | 2 | - | `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | - | `remove-call` | Stryker's filter/slice/sort removals; pitest `VOID_METHOD_CALLS` | 2 | - | `replace-call` | Stryker `MethodExpression`; pitest `CONSTRUCTOR_CALLS` | 1 | - | `replace-initializer` | pitest `PRIMITIVE_RETURNS`, `INLINE_CONSTS`; Stryker's literal mutators | 5 | - | `replace-iterable` | Stryker `ArrayDeclaration`; pitest `EMPTY_RETURNS` | 3 | - - Each one computes its own bytes where the mutation determines them (`false`, `true`, the negated - operator, the call's receiver, the guard's body) and takes the mutant's `to` only where the new value - is a choice - the same rule the earlier operators follow. `tests/tools/mutation-anchor.test.ts` has a - case per operator, including each refusal (an `else`, a block body, a call that is not a method call, - a missing initializer), and is now 30 cases. - Two widenings came with them, both of the same kind as the ones before - a site the vocabulary could - not name, shown by a real tooth failing to convert: - - - **A function bound to a name is a member.** `const count = (label) => {...}` in - `evals/ooo-execution/board-worker.ts` is the whole of that file's logic, and it is not a function - declaration, so its only tooth had no scope to name. The selector now accepts a variable whose - initializer is an arrow function or a function expression. - - **Two identical calls need a holder to tell them apart.** `dispatchPlan` calls - `board.candidates()` on offer and again filtered; the callee text is the same in both. `in` now - names a fragment of the statement the call sits in (it already named the holding object literal for - `replace-property` and the call's own text for `replace-argument`). - - One wrong `within` was shipped in the first pass and the **anchors pass**, not the sweep, is what - caught it: `the-completion-ignores-a-cancelled-unit` was aimed at `checkDispatch` instead of - `checkCompletion` (the fragment occurs in both members). That is the quieter failure: the anchors pass - resolves, the mutant is caught - by the _suite_, not by the case that names a completion of a - cancelled unit. In the second pass the same mistake was loud instead, because the guard I aimed at is - in `runParentCheck` and I named `instrumentCommit`: the pass refused with "0 guards in - instrumentCommit match the selector". A wrong member is loud when it finds nothing and silent when it - finds the same shape twice, which is why the sweep after a conversion is not optional. - **Reading after both passes:** `anchors: 110 of 110 resolve, over 22 targets`; `mutants: 110 of 110 -caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored - byte-identically, exit 0 in 176 s. - - **Then every one of the 33 converted, and the register is 110 derived of 110 with no hand-written - tooth left (2026-09-24).** The three that were declared open above as "worth an operator" were taken - instead of argued about, and each turned out to be a small, nameable widening rather than a new - domain: - - | what the residue needed | what was actually added | teeth | - | ------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----- | - | a file-scope selector | `within` is optional: without it the whole file is the scope, and the selector has to be unique in it | 2 | - | an index operator | `replace-index`: the element access the selector identifies gets the declared index | 2 | - | clause-level SQL mutation | `replace-literal-fragment`: a piece of text written inside a literal becomes the declared fragment, with the same holder disambiguator the call selectors have | 3 | - | a function bound to a property | `uniqueMember` accepts a function bound to a variable _or an object property_ (`claim: (store, parsed) => ...`) | 1 | - | a shorthand property | `replace-property` writes a shorthand out (`position` becomes `position: 0`), because a value cannot go where the name was | 1 | - | a guard inside a module-level script | the file scope above, with `condition-never` | 1 | - | a constructor call | the call selectors accept `new X(...)` as a call site | 1 | - | a call replaced by a boolean, a value replaced by a stub fact, a wrapper unwrapped, an iterable shortened, a comparison negated | the operators that already existed, used with a declared `to` | 5 | - - Three things this pass settled by measurement rather than by argument, each of which had been - written down here as a reason **not** to convert: - - - `speculation-guesses-several-facts-at-once` was kept because the catalogue's negation (`=== 1` to - `!== 1`) "would not realise the violation the name states". It does: the named case fails under the - negated comparison. The reasoning was wrong and the sweep said so. - - `every-task-is-frozen-at-position-zero` was kept because it "changes two things at once". Only one - of the two changes is load-bearing: `noUnusedParameters` is not set, so the callback can keep the - parameter it no longer reads, and the tooth converts to a single `replace-property`. - - The two teeth that write two statements on one line were never about the line discipline at all. - Their hand anchors had to span two statements; the mutations are single-statement deletions, and - both converted to `drop-statement`. - - Two mistakes were made and caught inside this pass, and both are worth keeping: - - - **A hand-typed `expect` reads as a broken mutant.** Retyping two case names from memory instead of - copying them from the register made both teeth report **"survived"** - the tool filters the suite by - the named case, so a name that matches no case looks exactly like a mutant nothing catches. It is a - loud failure (an alarm, not a false pass), and it is the second time this register has been bitten - by the `expect` field being an assertion of its own: the conversion script now copies the string, - and the copy is why the same conversions then read `caught by the named test`. - - **A conversion script rerun against a converted register.** Re-running the first wave's script after - the vocabulary had grown rewrote two teeth back to their earlier form (`in` dropped, `within` back - to the guessed member), and the anchors pass refused both immediately. That is the argument for the - anchors pass being in the static contract, seen from the other side: it also catches _edits to the - register_, not only drift in the code it points at. - - **The register is now 110 teeth, every one a name plus an operator plus a selector, over 22 targets - covering 21 files; 26 operators, all of them a mutation class the catalogues name or a slot they name - (a value, an argument, a property, an initializer, an iterable, an index, a callee, a literal - fragment).** Reading: `anchors: 110 of 110 resolve, over 22 targets`; `mutants: 110 of 110 caught by -the named test`, **all 110 by the case their `expect` names**, 22 of 22 restored byte-identically, - exit 0 in 160 s; `tests/tools/mutation-anchor.test.ts` 34 pass. - What is _not_ claimed: a derived tooth is not stronger evidence than the byte anchor it replaced - it - is the same violation, caught by the same case. What it buys is that the tooth stays aimed at its rule - across renames and reflows, which is the failure this record exists for. The operators are also not - the whole catalogue: they are the classes this register's rules happen to need, and adding one is now - an entry in a table plus its name in the union. - Two things the catalogues say about the shape of this work are worth carrying: - - - **Subsumption.** pitest's _Less is more_ names the effect this record's retirement criterion - measures: a mutant that the combination of others already subsumes adds no confidence, only - runtime. pitest avoids them with a fixed default set; this register measures it instead - the sweep - says which case fails, and the case is then read to see whether it states the rule. That is why 37 - teeth were retired rather than kept, and why the answer to "should every conceivable mutation be - here" is no. - - **Unviable mutants.** cargo-mutants counts a mutant that does not compile separately and prints it - only on request; it excludes test functions and `unsafe`, and it documents outright why it does not - generate some mutations (`==` to `<` "too prone to generate false positives", `-a` to `+a` "too - prone to generate unviable cases"). This register has no unviable mutants by construction - every - tooth has a named case - but the argument is the same one used above for leaving the residual - clusters alone: a general operator over a shape that is rare here produces mostly noise. - - The one thing this design does _not_ borrow from those tools is the definition of a kill. StrykerJS, - pitest and cargo-mutants count a mutant as killed when _any_ test fails; several can report which - test did it (`fullMutationMatrix`), but the purpose is to find the assertion to strengthen, not to - gate. Here a tooth is counted only when **the case it names** is the one that fails, which is why - "caught by the suite, not by the case that names the rule" is a defect this register reports and none - of those tools would. - -2. The pilot target converted and swept: same names, same cases, same catches; then a refactor inside it - that demonstrates a derived tooth surviving what a byte anchor did not. **Landed**, with one - correction: the demo first refactor was a rename plus a condition lifted into `const closed = spent || -severalWaits`, and the prediction made before running it - one dead anchor, seven derived teeth - surviving - was **wrong**: two sites died, the hand anchor and the derived - `a-live-claim-does-not-block-selection`, because a condition bound to a name was not a decision - position for the selector. Lifting a condition into a local is a refactor a maintainer makes, so the - selector was widened to include a variable's initializer; the corrected prediction (only the hand - anchor dies) then held exactly: `anchors: 148 of 149 resolve`, one failure, and the file restored - byte-identically. -3. The anchors-only pass, wired into the static contract with its route and design updates. **Landed:** - `npm run mutation:anchors` (`--anchors-only`) is one of `verify:static`'s checks and one of the - `ci-and-tests` route's atomic checks, in the same order the route-contract test enforces. Measured - on this revision: all 149 anchors resolve in 0.98 s inside a full `npm run agent:verify`, which is - 112 s end to end - inside its 150-second budget - and a standalone `verify:static` is 41 s. The full - sweep stays out of the gate. -4. Retirement pass over the ~18-20 teeth whose rule already has a relational or enumerative check, and - over the I/E teeth that can be replaced; the ledger's `proven` sentence updated in the same commit. - **First group landed (five teeth, 2026-09-24), and the criterion is the point of it:** a tooth may go - when a named case's assertion _fails for exactly the violation the tooth introduces_ and that case's - wording states the rule - so the row can name the check instead of the mutant. Applied by reading the - assertion and then re-running the target's sweep with the tooth gone. - - `the-board-read-path-stops-calling-the-predicate` and `the-board-decides-acceptance-on-its-own` - (row B5): the case counts the call sites itself - one definition of the predicate in `src/`, one - `acceptedFact({` in each reader, zero `verdict ===` comparisons in `ooo-board.ts` - so both - violations fail a count. This is the pair that made the criterion worth writing down: a - behavioural differential would not have caught either, a count does. - - `a-refused-unit-is-silent`: the case enumerates the 32 flag combinations a plan's facts can carry - and asserts every refused unit has a reason. - - `the-merge-enumerates-one-order` (row A5): the case asserts the multinomial count (10), which a - merge returning one order fails. - - `a-second-entry-rebinds-the-task` (row D12): the case asserts a retry is the same binding and a - second entry is refused by name. - Measured after the retirements: `anchors: 144 of 144 resolve` (the register is 144 teeth, not 149), - and the four targets the pass touched sweep `70 of 70 caught`, restored byte-identically 5 of 5. - One finding to carry forward: `a-refused-unit-is-silent` was named by **no** ledger row, and the case - that catches it is unowned too - an orphan tooth whose retirement removed the orphan rather than a - row's pin. - **Second group landed (six teeth, 2026-09-24):** the five rules of the declared budget - the licence - being the budget's part rather than the head, a budget above one requiring a target, one handoff per - startable task, a startable handoff surviving a republish, and a multi-slot handoff being directed - - are all stated by one case, `evals/ooo-execution/board-slots.test.ts`'s `a declared budget holds two -claims at once, and the store is why each handoff is directed`, whose assertions name each rule (the - refusal message, the handoff count, the null `serialState`, the identity across a republish, and the - non-head claimed first). They were retired from the F2b-slot row, which keeps seven teeth. The sixth, - `claim-is-not-scoped-to-its-run` (row B1), is the tooth the namespace experiment first saw survive - - the case was strengthened until it caught it, and the record documents the raw-row read that was - needed, which makes that case a better witness than the tooth's name. Measured after: - `anchors: 138 of 138 resolve`, and `--targets=src/integration/ooo-board.ts` sweeps `15 of 15 caught`, - restored byte-identically. - **Third group landed (two teeth retired, four links repaired, 2026-09-24):** - `a-retried-run-fact-is-appended-twice` and `a-frozen-task-is-replaced-by-a-different-definition` (row - D11) go, because the row's own prose already rested on the cases for those two rules and the cases say - them outright. The group's larger finding is what the first **whole-register** sweep said: 136 of 136 - caught, but only **134 by the case each `expect` named**. Four links were wrong, each in its own way: - - `fusion-continues-from-an-unverified-answer`'s case refused the pair for a second reason as well, so - it passed under the mutant. The pair now carries no dependency between the two units, which leaves - the predecessor's verdict as the only condition that can refuse it. - - `next-is-not-the-head-of-the-ordered-candidates`'s `expect` named a case in another target. The - board's budget case now asserts `next()` is the head of the ordered set, and that is the name it - carries. - - `the-caller-rebuilds-the-shared-floor`'s `expect` was a paraphrase that named no case at all. - - `the-grace-is-zero`'s named case derived its fixture from the constant under test - (`half = CLOCK_GRACE_MS / 2000`), so zeroing the grace moved the stamp onto `now` and the case passed - while the bug was live. The fixture is a literal now. - The re-run after those repairs: **136 of 136 caught, all 136 by the case its `expect` names**, 23 of 23 - targets restored byte-identically, 217 s. The class is the same one this record opened with - a link - that goes stale while the artifact still looks right - so the reading to keep is not the count of teeth - but the count of teeth whose named case is the one that fails. - One behaviour recorded rather than changed: a named case that does not finish inside its bound counts as - caught, with the reason printed (`the-pass-asks-a-unit-it-already-failed-again`, 30 s). The tool's own - comment says a run that never ends proves nothing; the code counts it as caught and says why. That - tension is left alone, because resolving it changes what the ledger's `proven` means. - **Remaining:** the mass conversion of the swap-shaped teeth - the six operators can express them, and that - is the deferred half of the vocabulary work - and the one insertion-shaped tooth left standing, - `every-task-is-frozen-at-position-zero`, together with `a-binding-does-not-record-its-channel`: both keep - a named case that is a register/freeze/adopt/read-back round-trip, weaker than the rule each is asked to - carry. - -## Acceptance criteria - -- A mutant declared as an operator over a selector stores no byte range, and the tool resolves the site - and the replacement from the syntax tree. -- A refactor that inlines, renames, reflows or extracts code around a derived site leaves the tooth - applying and still caught by its named case; a byte-anchored tooth in the same place does not survive - the same refactor (the pilot names both). -- The anchors-only pass runs without suites in under ten seconds on this tree, fails on an unresolvable - or ambiguous site, reports the claimed-versus-applicable gap, and leaves the tree byte-identical. -- Adding it to the static contract keeps a full `agent:verify` inside its 150-second budget. -- Teeth retired in favour of a contrast or enumeration are named in the commit that retires them, and - the row they proved names that check afterwards. -- The residue of hand-written fragments is enumerated by name, and every one of them is a single - expression or term rather than a statement or a message. - -## Risks - -- **A derived selector is still scoped to a decision position.** Measured in the pilot: a condition lifted - out of its `if` into `const closed = spent || severalWaits` retired the tooth - the selector looked for - a condition in a decision position, and a name bound for a decision made a line later was not one - until it was added. A tooth can therefore still die when its rule is rewritten into a _different_ - shape. The difference is that it now dies visibly: the anchors-only pass names the tooth and the - fragment it could not resolve, which is exactly the failure that went unnoticed before. -- **A wrong selector mutates the wrong node.** A derived site is resolved by the tool, so a selector that - matches another call can produce a mutant nobody intended. Mitigation: exactly one match or refusal, - which the tool already does for `ast`, plus the existing two readings (a clean run must pass, and the - named case must be the one that fails - a mutant that does not compile is reported as "by the suite, - not the named case" rather than as a caught tooth). -- **Derivation hides the perturbation from the reader.** A reader of a `from`/`to` pair sees the change; - a reader of `condition-never(within: judgeTaskBoardEntry, condition: existing.deliveredBy)` has to - know the operator. Mitigation: the operators are few and fixed, they are documented beside the type, - and the tool prints the bytes it wrote when asked. -- **Retiring a tooth can lose a named case.** A relational check may catch a class where the tooth caught - a specific wrong behaviour, so the row would claim more than it shows. Mitigation: a retirement names - the check that replaces it, in the ledger row and in the commit. -- **The pass could be trusted instead of run.** An anchors-only pass proves the site still resolves, not - that the mutant is still caught; a tooth whose named case went stale still passes it. Mitigation: the - pass reports the stale-`expect` reading the sweep already produces, and the full sweep remains the - standing rule before a push. -- **Conversion is evidence surgery.** Every conversion touches the artifact that says the code is - protected, so a mistake reduces coverage silently. Mitigation: one target per commit, the same names - and cases, the catch results recorded before and after, and the post-mortem rule about a live mutant - (status the file under test before any run). - -## Alternatives considered - -- **Keep hand-written anchors and make them smaller.** Rejected as the whole answer: it is a discipline, - not a mechanism, and the discipline is what 52% of the teeth currently violate. Kept as a rule for the - residue. -- **Anchor by AST node only, keeping byte replacements** (today's `ast.within` plus a text pair). - Rejected as insufficient: measured, 71 of 149 teeth already have that, and two teeth died inside it - this session, because the _replacement_ and the _fragment_ are still copies of lines. -- **Replace mutants with input-side checks wholesale.** Rejected on the measurement above: it does not - discriminate (every caught mutant is externally observable), and it would drop the only evidence that - speaks about an unforeseen implementation slip rather than a declared violation. -- **Insert the mutation into the code behind a runtime switch** (mutant schemata, `mutation_active("...")` - guards). Rejected for this repository: the guard is real code in `src/`, and a switched-off mutation - path is a policy word in the mechanism layer, which [mechanism, not - policy](2026-09-21-mechanism-not-policy.md) forbids. -- **Auto-generate mutants from operators over the whole tree** (what PIT, StrykerJS and cargo-mutants - do). Rejected as the form here: it would replace named evidence with a score, and the ledger needs a - named tooth per row. The derived form keeps the operator idea and the naming. -- **Leave the sweep out of every gate and rely on the standing rule.** Rejected: the measured case for - this record is that nobody noticed two teeth had stopped biting, and a rule that depends on being - remembered is the shape this repository already replaced elsewhere. diff --git a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md deleted file mode 100644 index 73b110dc..00000000 --- a/docs/decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ /dev/null @@ -1,138 +0,0 @@ -# 牙是算子加符号,不是一行文本的副本 - -[English](2026-09-24-mutants-are-derived-not-anchored.md) - -**Status:** proposed -**Approved:** explicit -**Relates to:** [测试不需要文件系统](../implemented/2026-09-20-tests-need-no-filesystem.md)、[检查读到的是一个活着的 mutant](../../postmortem/0003-checks-read-a-live-mutant.md)、[把 agent 验证限制成一次运行](../implemented/2026-09-23-verification-whole-run-deadline.md)、[契约的义务](../../design/task-unit-semantics-obligations.md)、[机制,不是策略](2026-09-21-mechanism-not-policy.md) - -## 问题 - -这个仓库决定要保护的每一条规则,都用同一种方式保护:`tools/mutation-teeth.ts` 里一颗具名 mutant——把一段字节替换成手写的替换文本,外加“哪一个用例必须因此失败”的名字。这条证据是诚实的(只有 mutant 这种形式能谈论“实现自己有可能错”),同时它也是树里最脆的产物,因为**锚点就是 mutant 的身份**:它是某一行的副本,于是那一行搬家、被内联、或者消息文本被抽走,这颗牙就退役了。 - -2026-09-24 在 149 颗牙、22 个目标上实测: - -| 锚点指向什么 | 颗数 | 占比 | 其中已有成员级作用域 | -| ------------------------------------- | ---- | ---- | -------------------- | -| 守卫或谓词项被中和(G) | 64 | 43% | 30 | -| 值、实参、下标或被调者被替换(V) | 58 | 39% | 26 | -| 语句或调用被删(R) | 11 | 7% | 6 | -| 插入代码:语句、分支或第二份规则(I) | 8 | 5% | 4 | -| 整段表达式被替换(E) | 8 | 5% | 5 | - -两个读数为这份记录定性。**89% 的牙(149 里的 133)就是“一个位置 + 少数几种算子”**——让条件永不成立、去掉一项、替换一个实参、删掉一条语句——也就是说算子是意图,位置只是陪衬;把位置存成文本是一种选择,而不是必需。而 **149 里有 78 颗(52%)根本没有结构作用域**:一个在全文件里匹配的字节片段,正是死得最快的那种形状。本轮两次失效都是这个形状:一条语句里被内联掉的内部调用,以及一段拒绝消息被抽成 `subject` 表达式。旁边还出现了第三种形状:两颗牙的点名用例已经不再是抓住它们的那个用例(`fusion-continues-from-an-unverified-answer`、`next-is-not-the-head-of-the-ordered-candidates` 被读成“caught by the suite, not the named case”)——同一类耦合往外一层:锚点还在,“这个用例就是失败的那个”这句话过期了。 - -规划这份记录时做的另一次测量也必须写下来,因为它排除了一个选项:三分法“可派生 / 可用输入式反例 / 必须扰动内部”**不成立**。一颗被抓的 mutant 按定义就是**从外部可观察**的(这就是“被抓”的含义),所以“可用输入式反例”对 149 颗全都成立。三颗预测“必须扰动内部”的牙都找得到外部观察:状态读路径开了可写句柄,用“store 文件没有变化”观察;acceptance 规则的第二份副本,用一个 digest 与产物不匹配的裁决观察;事务包装,用“计划里第三个任务被拒时前两个仍冻结”观察。所以替换方向不是“用输入检查替代 mutant”,而是**派生 mutant,而不是把它存下来**;输入式检查留给那些“对比或穷举本来就说得更清楚”的规则。 - -## 提案 - -**一颗牙的耐久部分是它的名字、它的算子、它的选择器,以及必须失败的那个用例。位置与替换字节每次从语法树算出来。** - -1. **六个作用在具名选择器上的算子**,作用域限定在一个成员内:`condition-never`(选择器指到的那条条件永不成立)、`condition-holds`(它始终成立——镜像的那个,因为写成 `return a && b` 的规则说的是“这条条件为真”,把它换成 `false` 是把规则反过来,而不是去掉)、`neutralize-term`(该条件里的一项变成它的恒等元——`&&` 下为 `true`,`||` 下为 `false`)、`replace-argument`(某个调用的第 _n_ 个实参换成声明的片段)、`replace-property`(具名对象属性的值换成声明的片段)、`drop-statement`(选择器指到的那条语句连行带缩进被删掉)。两条规则让选择器保持诚实:匹配到多于一处、或者片段同时命中两个候选的,一律拒绝而不是猜;两个“整条件”算子拒绝只命名条件一部分的片段,因为整段替换会把 mutant 悄悄放大成“这条规则的每一条理由”——一项有它自己的算子。规则用不上这些时,仍可用“成员作用域 + 最小片段”的旧形式;而**“最小”是要点**:不要消息文本、不要同级实参、一项能表达的不要写成整条语句。 - **已推广到整个登记处(2026-09-24),分三波:59 颗、18 颗,最后 15 颗——登记处现在是 110 颗里 110 颗 derived。**第一遍:六个算子已经能表达的每一类点位,按目标逐个转完、每转完一个就 sweep——`condition-never`(守卫或初始化器,共 44)、`neutralize-term`(9)、`replace-property`(9)、`replace-argument`(6)、`drop-statement`(6)、`condition-holds`(3)。 - 第二遍:把剩下的 33 颗按"缺什么"分类,再拿现成变异工具真正点名的东西核对(Stryker 的 supported mutators、pitest 的 mutator 列表、cargo-mutants 的 patterns、Cosmic Ray 的 operator 概念)。其中十三颗落在"每个目录都有的变异类"里,于是词汇增加了七个算子,而不是宣布这一类做不到: - - | 算子 | 现成目录里的名字 | 颗数 | - | --------------------- | ------------------------------------------------------------- | ---- | - | `negate-condition` | pitest `NEGATE_CONDITIONALS`;Stryker 布尔字面量 | 2 | - | `negate-comparison` | pitest `NEGATE_CONDITIONALS`;Stryker `EqualityOperator` | 2 | - | `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | - | `remove-call` | Stryker 的 filter/slice/sort 删除;pitest `VOID_METHOD_CALLS` | 2 | - | `replace-call` | Stryker `MethodExpression`;pitest `CONSTRUCTOR_CALLS` | 1 | - | `replace-initializer` | pitest `PRIMITIVE_RETURNS`、`INLINE_CONSTS`;Stryker 字面量类 | 5 | - | `replace-iterable` | Stryker `ArrayDeclaration`;pitest `EMPTY_RETURNS` | 3 | - - 凡是"变异本身决定字节"的(`false`、`true`、取反后的算子、调用的接收者、守卫的 body)都由算子算出来;只有当新值是一个**选择**时才取 mutant 的 `to`——与先前算子同一条规矩。`tests/tools/mutation-anchor.test.ts` 每个算子一条用例,并包含它各自的拒绝情形(有 `else`、body 是块、调用的不是方法、没有初始化器),现在是 30 条。 - 随之而来两处"扩宽",与此前的性质相同——都是**真实的一颗牙转不过去**才暴露出来的点位: - - **绑定到名字的函数也是成员。**`evals/ooo-execution/board-worker.ts` 里 `const count = (label) => {...}` 是该文件逻辑的全部,它不是函数声明,于是那颗牙连作用域都写不出来。选择器现在接受"初始化器是箭头函数或函数表达式的变量"。 - - **两个一模一样的调用要靠"容器"区分。**`dispatchPlan` 里 `board.candidates()` 出现两次(一次 onOffer、一次接 filter),callee 文本完全相同。`in` 现在可以指"该调用所在语句里的一段文本"(它此前已经能指 `replace-property` 的宿主对象字面量、`replace-argument` 的调用自身文本)。 - 第一遍我又写错了一处 `within`,这次是**anchors pass**(不是 sweep)抓出来的:`the-completion-ignores-a-cancelled-unit` 被指向 `checkDispatch` 而不是 `checkCompletion`(同一片段在两个成员里都有)。这是更安静的那种失败:anchors pass 照样通过、mutant 照样被抓——但抓它的是**整个 suite**,不是那条"报告一个被取消单元的完成"的用例。第二遍同类错误反而是响的:我瞄的那个守卫在 `runParentCheck`,我却写了 `instrumentCommit`,pass 直接拒绝——"0 guards in instrumentCommit match the selector"。**成员名写错,找不到时是响的、找到两处同形状时是哑的**,这就是"转完必须 sweep"不可省的理由。 - **两遍之后的读数:**`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,176 秒。 - - **随后这 33 颗全部转完,登记处现在是 110 颗里 110 颗 derived,没有一颗手写(2026-09-24)。**上面写成"值得加算子"的三簇,是拿实现去验而不是继续争论:每一簇都不是新领域,而是一处小且可命名的扩宽—— - - | 残留需要什么 | 实际加了什么 | 颗数 | - | ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | ---- | - | 文件级选择器 | `within` 变可选:不写就用整个文件作作用域,选择器必须在文件里唯一 | 2 | - | 下标算子 | `replace-index`:选择器指到的元素访问换成声明的下标 | 2 | - | 子句级 SQL 变异 | `replace-literal-fragment`:字面量内部的一段文本换成声明的片段,并带上调用选择器那套"容器"消歧 | 3 | - | 绑定到对象属性的函数 | `uniqueMember` 接受绑定到变量**或对象属性**的函数(`claim: (store, parsed) => ...`) | 1 | - | 简写属性 | `replace-property` 把简写写开(`position` 变 `position: 0`),因为值不能写在名字的位置上 | 1 | - | 模块级脚本里的守卫 | 上面的文件级作用域,配 `condition-never` | 1 | - | 构造函数调用 | 调用选择器接受 `new X(...)` 作为调用点 | 1 | - | 调用换成布尔、值换成桩事实、包装拆开、可迭代缩短、比较取反 | 已有的算子,配声明的 `to` | 5 | - - 这一遍里有三件事是**量出来**而不是讲道理定的,而它们此前都被写在本记录里当作"不转"的理由: - - - `speculation-guesses-several-facts-at-once` 当初留着,理由是"目录里的取反(`=== 1` → `!== 1`)不会实现名字所说的违规"。事实上会:换成取反比较后,点名的那条用例照样失败。理由是错的,是 sweep 说了算。 - - `every-task-is-frozen-at-position-zero` 当初留着,理由是"一次改两处"。两处里只有一处是承重的:`noUnusedParameters` 没开,回调可以继续保留它不再读的参数,于是这颗牙变成单条 `replace-property`。 - - 那两颗"两条语句挤一行"从来不是行纪律的事。它们的手写锚点**不得不**横跨两条语句,而变异本身是删掉一条语句,两颗都转成了 `drop-statement`。 - - 这一遍里犯的两个错都被当场抓住,都值得记下来: - - **`expect` 是手打的,读起来就像一颗坏掉的牙。**有两颗牙我凭记忆重打了用例名(没有从登记处抄),结果两颗都报 **"survived"**——工具按点名的用例过滤 suite,名字对不上任何用例时,看起来就跟"没任何用例抓到"一模一样。它是响的失败(警报,不是误判通过),但这是这个登记处第二次被 `expect` 字段咬到:转换脚本现在**抄**这个字符串,抄过之后同样两颗读数是 `caught by the named test`。 - - **转换脚本对着已转换的登记处又跑了一遍。**词汇长大之后重跑第一遍的脚本,把两颗牙改回了早期形态(`in` 掉了、`within` 回到猜的那个成员),anchors pass 立刻拒绝了两个。这正是"把 anchors pass 放进静态契约"的另一个方向的理由:它也抓**对登记处本身的编辑**,不只是抓它指向的代码漂移。 - - **登记处现在是 110 颗牙,每一颗都是"名字 + 算子 + 选择器",覆盖 22 个目标、21 个文件;26 个算子,每一个要么是目录点名的变异类,要么是目录点名的槽位**(值、实参、属性、初始化器、可迭代、下标、被调、字面量片段)。读数:`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,160 秒;`tests/tools/mutation-anchor.test.ts` 34 条全过。 - 没有声称的东西:derived 形态**不比**它替代的字节锚点更强——违规是同一个、抓它的用例是同一条。它换来的是:改名、重排之后这颗牙仍瞄准原来那条规则,而这正是本记录要解决的那个失败。算子也不等于整个目录:它们只是本登记处的规则恰好需要的那些类,加一个算子现在就是表里加一行、类型里加一个名字。 - 目录里还有两件事值得带过来: - - **subsumption(被覆盖)。**pitest 的《Less is more》给这件事起了名:一个 mutant 若已被其他 mutant 的组合覆盖,它带来的只是运行时间,不是信心。pitest 靠"固定的默认集合"回避;本登记处是**量**它——sweep 先说出是哪条用例失败,再读那条用例确认它确实陈述了规则。这正是 37 颗牙被退役而不是留着的原因,也是"是不是每种能想到的变异都该进来"的答案是"不"的原因。 - - **unviable mutant(编译不过的变异体)。**cargo-mutants 把这类单独计数、默认不打印;它排除测试函数与 `unsafe`,并且明确写出**为什么不生成某些变异**(`==` 换成 `<`"太容易产生假阳性"、`-a` 换成 `+a`"太容易不编译")。本登记处**按构造就没有** unviable(每颗都有点名的用例),但理由是同一条:为一个在本仓库很少出现的形状加通用算子,产出的多半是噪音。 - 这个设计与那些工具**唯一没有借用**的一点,是"消灭"的定义。StrykerJS、pitest、cargo-mutants 把"任意一条测试失败"就算 killed;其中几个还能报告是哪条测试(`fullMutationMatrix`),但用途是找出该加强哪条断言,不是当门禁。这里只有**点名的那条**用例失败才算数,所以"被整个 suite 抓到、却没被点名用例抓到"是本登记处会报、而那些工具一个都不会报的缺陷。 - -2. **把 133 颗 G/V/R 牙逐目标改成派生形式**,从 `src/integration/ooo-execution.ts` 开始试点:该文件在两条目标条目里共 24 颗(16 颗融合合法性推演、8 颗会话融合);同样的名字、同样的 `expect` 用例,转换后同样 24/24 被抓——其中 19 颗派生、5 颗保留手写(下面具名)——然后在该文件里做一次真重构,证明派生出来的位置都还在,而字节锚点会在这次重构里退役。 - 试点文件里那 5 颗残余,每一颗都是单个表达式或单个值,而不是整条语句或消息:`selection-ignores-a-withdrawn-acceptance`(整个 `const` 初始化式被替换)、`the-budget-is-not-cut-from-the-startable-set`(被返回表达式里的调用被删)、`next-task-is-not-the-head-of-the-legal-set`(下标 `[0]` 变 `[1]`)、`speculation-guesses-several-facts-at-once`(`length === 1` 变 `>= 1`)、`a-guess-with-no-evidence-publishes`(插入一个分支并重写消息)。 -3. **派生形式表达不了的,保留手写,并且说明为什么。** 16 颗 I/E 牙——插入、包装、第二份规则——它们的锚点是“某段代码不该出现的地方”,所以它们最后才搬,也最先被重新考虑:凡是规则已经有一个关系式或穷举式检查的,就退役这颗牙,并在它原来的位置上点名那个检查。 -4. **“牙是否还咬得住”成为静态契约的一部分**:一趟只做解析、不跑用例的 anchors-only 检查,个位秒数内完成,在位置无法解析、解析出多于一处、或者某个目标声称的颗数多于能应用的颗数时失败。`ci-and-tests` 路由随后把它与其它原子检查并列;而[整轮预算](../implemented/2026-09-23-verification-whole-run-deadline.md)是它的约束:这趟检查必须便宜,全量 sweep 仍然不进闸门。 -5. **证据标准是被澄清,不是被放宽。** “一行如何挣得 `proven`”原文已经写的是“一个在规则被破坏时会失败的检查”;澄清之处在于:这个证明可以是一次**代码扰动**(派生的或手写的),也可以是一个**输入侧用例**,其断言对比两组输入、或穷举一个有界空间。某一行以前点名一颗牙,现在可以改为点名一个对比检查——在同一提交里退役那颗牙。 - -## 计划 - -1. 本文,加上工具的算子词汇表与解析器;用“对源码字符串”而不是对文件系统的测试来证明它。**已落地:**解析器是 `tools/mutation-anchor.ts`(它不持有任何状态,所以从 sweep 脚本里搬出来,测试可以直接调用),16 个用例全部作用在源码字符串上,不碰文件系统。 -2. 试点目标转换并重跑:同样的名字、同样的用例、同样的抓法;随后在该文件内做一次重构,证明派生出来的牙活过了字节锚点活不过去的那次改动。**已落地,带一处更正:**演示用的那次重构是“改一个局部变量名 + 把条件抬成 `const closed = spent || severalWaits`”,而运行前写下的预测——死一颗锚点、七个派生位置存活——**是错的**:死了两处,手写锚点和派生的 `a-live-claim-does-not-block-selection`,因为“绑定到名字上的条件”当时不算选择器眼里的一次决策位置。把条件抬成局部常量是维护者真会做的重构,于是选择器扩到了变量初始化式;更正后的预测(只有手写锚点会死)随后精确成立:`anchors: 148 of 149 resolve`,一处失败,文件逐字节还原。 -3. anchors-only 这趟检查接入静态契约,并同步路由与设计文档。**已落地:**`npm run mutation:anchors`(`--anchors-only`)现在是 `verify:static` 的一条检查,也是 `ci-and-tests` 路由的一条原子检查,顺序与路由契约测试强制的一致。本版本实测:一次完整的 `npm run agent:verify` 里,149 颗锚点全部解析,用时 0.98 秒,整轮 112 秒——留在 150 秒预算之内;单独跑 `verify:static` 是 41 秒。全量 sweep 仍不进闸门。 -4. 对“规则已有关系式/穷举式检查”的那约 18-20 颗,以及能被替代的 I/E 牙,做退役;台账里 `proven` 那句话在同一提交里更新。**第一组已落地(五颗,2026-09-24),而准则本身才是这件事的要点:**当一条具名用例的断言*恰好因为这颗牙引入的违规*而失败、且该用例的措辞陈述了那条规则时,这颗牙才可以退役——于是台账行点名的变成那条检查,而不是 mutant。判定方式是先读断言,再把牙删掉后重跑该目标的 sweep。 - - `the-board-read-path-stops-calling-the-predicate` 与 `the-board-decides-acceptance-on-its-own`(B5 行):用例自己数调用点——`src/` 里谓词只定义一次、两个读取方各有一个 `acceptedFact({`、`ooo-board.ts` 里 `verdict ===` 出现零次——所以两种违规都让某个计数失败。正是这两颗让这条准则值得写下来:一个行为差异对比两颗都抓不住,计数能。 - - `a-refused-unit-is-silent`:用例穷举了计划事实可能携带的 32 种 flag 组合,并断言每一个被拒单元都有理由。 - - `the-merge-enumerates-one-order`(A5 行):用例断言多重组合计数(10),而只返回一种顺序的合并会失败。 - - `a-second-entry-rebinds-the-task`(D12 行):用例断言重试是同一个绑定、且第二条 entry 被按名拒绝。 - 退役后实测:`anchors: 144 of 144 resolve`(登记处是 144 颗牙,不是 149),这趟碰到的四个目标 sweep `70 of 70 caught`,5 of 5 逐字节还原。 - 一个要带下去的发见:`a-refused-unit-is-silent` **没有任何**台账行点名它,抓住它的那条用例同样无人认领——一颗孤儿牙,退役它移除的是孤儿,而不是某一行的钉子。 - **第二组已落地(六颗,2026-09-24):**声明的预算那五条规则——许可的是预算的那一部分而不是表头、预算大于一必须点名目标、每个可开始任务都有一个 handoff、可开始的 handoff 会在重发布中存活、多槽位的 handoff 是被定向的——都由同一条用例陈述:`evals/ooo-execution/board-slots.test.ts` 的 `a declared budget holds two claims at once, and the store is why each handoff is directed`,其断言逐条点名了这些规则(拒绝消息、handoff 计数、`serialState` 为 null、重发布后 id 不变、先认领非表头那个)。这五颗从 F2b-slot 行退役,该行保留七颗。第六颗 `claim-is-not-scoped-to-its-run`(B1 行)是命名空间实验里最初**没被抓住**的那颗牙——用例被加强到能抓住它,实验记录写下了当时必需的"裸行读取",所以那条用例比这颗牙的名字是更好的证人。退役后实测:`anchors: 138 of 138 resolve`,`--targets=src/integration/ooo-board.ts` sweep `15 of 15 caught`,逐字节还原。 - **第三组已落地(退役两颗、修好四条链接,2026-09-24):**`a-retried-run-fact-is-appended-twice` 与 `a-frozen-task-is-replaced-by-a-different-definition`(D11 行)退役——该行的正文本来就靠这两条用例承载那两条规则,而用例把规则直说了。这一组更大的发现是**第一次全登记处 sweep**的读数:136 颗全部被抓,但只有 **134 颗是被各自 `expect` 点名的那条用例抓住的**。四条链接都错了,而且各错各的: - - `fusion-continues-from-an-unverified-answer` 的点名用例还因为另一个原因拒绝这一对,于是在 mutant 下照样通过。现在这一对的两个单元之间没有依赖,前驱的裁决成为唯一能拒绝它的条件。 - - `next-is-not-the-head-of-the-ordered-candidates` 的 `expect` 指的是另一个目标里的用例。董事会(board)的预算用例现在断言 `next()` 就是有序集合的表头,它的名字随之更正。 - - `the-caller-rebuilds-the-shared-floor` 的 `expect` 是一句转述,不对应任何用例。 - - `the-grace-is-zero` 的点名用例从**被测常量本身**推导夹具(`half = CLOCK_GRACE_MS / 2000`),于是把 grace 归零会把时间戳挪到 `now` 上,缺陷活着而用例照过。夹具现在是字面量。 - 修完重跑:**136 of 136 caught,且 136 颗全部由各自 `expect` 点名的用例抓住**,23 of 23 目标逐字节还原,217 秒。这一类正是本记录开头写的那一类——**凭据看着没问题,链接悄悄过期**——所以要记住的读数不是牙的数量,而是"点名用例就是失败那个"的牙的数量。 - 一条记录而未改动的行为:点名用例在自己的时限内没跑完,也被算作抓住,并把原因打印出来(`the-pass-asks-a-unit-it-already-failed-again`,30 秒)。工具自己的注释说"永远跑不完的一轮什么也证明不了",代码却把它算作抓住并说明了理由。这个矛盾留在原地,因为解决它会改变台账里 `proven` 的含义。 - **已落地(2026-09-24):27 颗"插入型"牙里,26 颗退掉了。**每一颗都走同一条路:先由 sweep 证明它的点名用例就是失败的那条,再读那条用例,确认规则被直说——对 store 自身源码计数(`the store runs its transaction boundary in exactly one place`:一个 BEGIN、一个 COMMIT、一个 ROLLBACK,且都在 `writeTransaction` 里)、对已接受集合/回退项/冻结残留做枚举、按名字拒绝后再把原行读回来(`one reader of the same ready task is given the claim, the second is refused` 直接从表里读 owner),或者用文件自身的哈希(`a read-only open neither creates, migrates nor writes`)。这一组里留下了一颗:`every-task-is-frozen-at-position-zero`,它的点名用例是"注册/冻结/采纳/读回"的往返,弱于"位置来自数组顺序而不是请求",与 `a-binding-does-not-record-its-channel` 留下的理由相同。26 颗里有 7 颗是**孤儿**——没有任何一行点名它们——它们各自钉住的规则有检查、却没有行:`round-publication-opens-its-own-transaction`(`the round's own publication rolls back with the transition that made it`)、`round-never-releases-its-pin`(`cancelling a round releases the pins it held, so nothing it referenced leaks`)、`ordered-mode-becomes-any-topological-order`(`the ordered mode is the declared plan order and nothing else`)、`the-loop-awaits-each-unit-instead-of-the-batch` 与 `a-unit-is-dispatched-twice-in-one-batch`(`a unit's check is outstanding while an independent unit's worker runs`,它同时断言"each unit is dispatched once")、`a-unit-nothing-checks-is-still-a-unit`(`evals/ooo-execution/families.test.ts` 里每个 family 的"a unit nothing checks is refused rather than accepted on nothing")以及 `the-parent-check-ignores-its-own-verdict`(`the parent check is the composed acceptance, and a failing check is reported as such`)。这七条规则是否该各自有一行,是这次读数提出的下一个问题——它是台账的问题,不是登记处的问题。 - - 退役后的读数:**110 of 110 caught by the named test,且 110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 目标逐字节还原,162 秒;`anchors: 110 of 110 resolve, over 22 targets`。目标 `tools/agent-verify.ts` 随它唯一那颗牙一起消失,所以登记处现在是 21 个文件上的 22 个条目。 - **尚未完成:**剩余的"替换型"牙的整体迁移——六个算子能表达它们,那是词表工作中被推迟的一半;以及仍立着的插入型 `every-task-is-frozen-at-position-zero` 与 `a-binding-does-not-record-its-channel`:两者保留的点名用例都是"注册/冻结/采纳/读回"的往返,弱于各自要承担的规则。 - -## 验收标准 - -- 以“算子 + 选择器”声明的 mutant 不存任何字节区间,位置与替换文本由工具从语法树解析。 -- 一次把周围代码内联、改名、重排或抽出的重构之后,派生位置仍能应用、且仍被点名用例抓住;同一处的字节锚点则不然(试点要把两者都点出来)。 -- anchors-only 检查在本地树上不跑用例、十秒内完成;无法解析或解析出多于一处时失败;报告“声称 vs 可应用”的差额;并让工作树逐字节不变。 -- 把它加进静态契约之后,一次完整的 `agent:verify` 仍留在 150 秒预算之内。 -- 因对比或穷举而退役的牙,在退役它的那次提交里被点名,它原来保护的那一行之后点名那个检查。 -- 手写片段的残余按名字列出,且每一颗都是单个表达式或单项,而不是一条语句或一段消息。 - -## 风险 - -- **派生选择器仍然被限定在“一次决策位置”上。** 试点实测:把条件从它的 `if` 里抬成 `const closed = spent || severalWaits` 会直接让那颗牙退役——选择器找的是决策位置上的条件,而“为下一行的决策先绑个名字”在扩展之前不算一处。所以当规则被改写成*另一种形状*时,牙仍可能死掉。区别在于它现在是**可见地**死:anchors-only 那趟检查会点出牙的名字和它解析不到的片段,而这恰恰是以前没人注意到的那种失效。 -- **选择器写错会改到错误的节点。** 派生位置由工具解析,所以匹配到另一个调用的选择器会产出一颗没人想要的 mutant。缓解:恰好一处否则拒绝(工具对 `ast` 已经如此),加上原有的两重读数(clean run 必须通过;必须是点名用例失败——编译不过的 mutant 会被读成“by the suite, not the named case”,而不是被算作抓到的牙)。 -- **派生把扰动藏起来了。** 读 `from`/`to` 的人看得见改动;读 `condition-never(within: judgeTaskBoardEntry, condition: existing.deliveredBy)` 的人得知道这个算子。缓解:算子少而固定、写在类型旁边;工具在被要求时打印它实际写入的字节。 -- **退役一颗牙可能丢掉一个点名用例。** 关系式检查抓的可能是一类,而那颗牙抓的是一个具体的错行为,于是那一行会宣称超过它所展示的。缓解:退役时在台账行与提交里点名替代它的检查。 -- **这趟检查可能被信任而不是被运行。** anchors-only 只证明位置还能解析,不证明牙还被抓住;点名用例过期的那类牙照样通过。缓解:这趟检查报告 sweep 已经会产出的“点名用例过期”读数,而全量 sweep 仍是推送前的常设规则。 -- **转换是在证据上动手术。** 每次转换都碰到“这份代码被保护着”的凭据本身,所以一个错误会静默地减少覆盖。缓解:一个提交一个目标;名字与用例不变;转换前后都记录抓取结果;并遵守“活着的 mutant”那条验尸规则(任何一次运行之前先看被测文件的状态)。 - -## 考虑过的替代方案 - -- **保留手写锚点,只把它们写小一点。** 作为全部答案被拒绝:那是一条纪律,不是一种机制,而 52% 的牙今天违反的正是这条纪律。作为残余部分的规则保留。 -- **只用 AST 定位,替换仍是字节**(即今天的 `ast.within` 加一对文本)。作为不够被拒绝:实测已有 71/149 是这样,本轮仍有两颗牙死在它里面,因为**替换文本**和**片段**依旧是行的副本。 -- **整体用输入侧检查替换 mutant。** 按上面的测量拒绝:它不具区分度(每颗被抓的 mutant 都是外部可观察的),而且会丢掉唯一一种谈论“未预料到的实现偏差”而不是“声明的违反”的证据。 -- **把变异插进代码、用运行时开关激活**(mutant schemata,`mutation_active("...")` 这类守卫)。本仓拒绝:守卫是 `src/` 里的真代码,而一条被关掉的变异路径就是机制层里的策略词,[机制,不是策略](2026-09-21-mechanism-not-policy.md) 已经禁止。 -- **全树按算子自动生成 mutant**(PIT、StrykerJS、cargo-mutants 的做法)。作为本仓的形式被拒绝:它会把具名证据换成一个分数,而台账每一行需要一个具名的牙。派生形式保留了算子这个主意,也保留了具名。 -- **把这趟检查留在闸门之外,只靠常设规则。** 拒绝:这份记录的实测理由就是没人发现两颗牙早已不再咬人,而一条依赖“有人记得”的规则,是仓库在别处已经替换掉的形状。 diff --git a/docs/design/ci-cd-and-quality.md b/docs/design/ci-cd-and-quality.md index de1ac9d5..fad73648 100644 --- a/docs/design/ci-cd-and-quality.md +++ b/docs/design/ci-cd-and-quality.md @@ -59,7 +59,7 @@ exit_criteria: Replace with a stable contract test or remove after the redesign `npm run check:tests`(`tsc -p tsconfig.tests.json --noUnusedLocals --noUnusedParameters`,`tests/` 面第一次被类型检查)刻意不进入任何阻塞契约,只在 CI 以 advisory 步骤运行。 -`verify:static` 中的 `mutation:anchors`(`tools/mutation-teeth.ts --anchors-only`)只做一件事:把 149 颗具名 mutant 的位置全部解析一遍,不跑任何用例、不写任何字节,在 0.6 秒内回答“每一颗牙是否还瞄着东西”。它进入静态契约是因为**一颗锚点失效时没有别的检查会注意到**:全量 sweep 不跑(`mutation:teeth` 不在任何 CI 作业里),而一颗匹配不到位置的牙在 sweep 报告里只是“不可应用”并被排除出分母——本轮修掉的两颗牙就是这样悄无声息地停摆的。全量 sweep 仍然不进闸门:它是分钟级、要跑用例,属于推送前的常设规则([决策](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md))。 +`verify:static` 中的 `mutation:anchors`(`tools/mutation-teeth.ts --anchors-only`)只做一件事:把 110 颗具名 mutant 的位置全部解析一遍,不跑任何用例、不写任何字节,在 1 秒内回答“每一颗牙是否还瞄着东西”。它进入静态契约是因为**一颗锚点失效时没有别的检查会注意到**:全量 sweep 不跑(`mutation:teeth` 不在任何 CI 作业里),而一颗匹配不到位置的牙在 sweep 报告里只是“不可应用”并被排除出分母——本轮修掉的两颗牙就是这样悄无声息地停摆的。全量 sweep 仍然不进闸门:它是分钟级、要跑用例,属于推送前的常设规则([决策](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md))。 `verify:static` 中的 `complexity:gate` 默认以 `git merge-base HEAD origin/main` 为基线(可用 `--base ` 显式覆盖)。基线必须是 merge base 而不是 `HEAD`:后者只比较未提交的工作树,于是已提交到分支的改动完全不可见 —— 在 CI 的干净检出上它永远报“无改动”,等于每个 PR 都没有被这条 gate 检查过。因此每次运行都会**陈述自己用了哪个基线**,并**点名它未能测量的改动文件**(ESLint 拒绝某路径、或文件根本无法解析,都会产出“零发现”,与“量过且干净”无法区分)。 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index ec2607ae..e49493ee 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -27,7 +27,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth` (the whole register, every tooth derived) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 160 s (2026-09-24); `npm run mutation:anchors` -> `anchors: 110 of 110 resolve, over 22 targets`; `tests/tools/mutation-anchor.test.ts` -> 34 pass; `test:product` -> 1580 pass; `verify:static` exit 0; lint 0 findings; complexity gate ok. - `npm run mutation:teeth` (the whole register, after the operators the catalogues name were added and 18 more teeth converted) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 176 s (2026-09-24); `npm run mutation:anchors` -> `anchors: 110 of 110 resolve, over 22 targets`; `tests/tools/mutation-anchor.test.ts` -> 30 pass. - `npm run mutation:teeth` (the whole register, after 59 teeth were converted to derived selectors) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 171 s (2026-09-24). The conversion is not what makes them bite; it is what keeps them aimed: one wrong `within` was shipped in this pass and the sweep is what caught it, as a tooth caught by the suite rather than by the case that names its rule. -- `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were **orphans**: no row named `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` or `the-parent-check-ignores-its-own-verdict`, so their rules have a case but no row; the check each case makes is in the [record](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. +- `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were **orphans**: no row named `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` or `the-parent-check-ignores-its-own-verdict`, so their rules have a case but no row; the check each case makes is in the [record](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) - `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) - `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. @@ -40,7 +40,7 @@ description rather than a pin. Every mutant name below was read from A tooth is a **name, an operator and a selector** - not a copy of a line: the site and the replacement bytes are resolved from the syntax tree on every run, so a rename or a reflow does not retire the pin -while the rule still stands ([the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). +while the rule still stands ([the decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)). **All 110 of the register's teeth are in that form now**, over 22 targets covering 21 files, with the 26 operators the record lists - each one a mutation class the catalogues name or a slot they name. A tooth with no operator to name it is no longer residue but a missing operator: the last 15 converted @@ -59,7 +59,7 @@ as well; both were found by reading the assertion rather than by trusting the sw made after the sweep says the named case is the one that fails under that tooth, and the row that used to name the tooth names the check instead, in the same commit. 26 teeth were retired this way on 2026-09-24, including every insertion-shaped tooth whose rule had such a check; the ones that stay are named, with -their reasons, in [the decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md). +their reasons, in [the decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md). ## Where a proof lives: two evidence bases, and neither stands for the other @@ -150,19 +150,19 @@ separate claims. ## A. Offline semantics (the design's first slice, already landed) -| node | obligation | state | evidence | -| ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| A1 | The compiler returns a legal plan or a refusal naming task, field and obligation | proven | `tests/integration/task-semantics.test.ts` | -| A2 | The offline model dispatches without publishing or storing run state | proven | `tests/integration/task-semantics-model.test.ts` | -| A3 | The design's six discriminating cases are executable | proven | `tests/integration/task-semantics-cases.test.ts` | -| A4 | Field mapping catches budget-unit confusion, requires/deps confusion and silently dropped keys | proven | mutants `budget-inner-alias-is-dropped`, `over-maximum-budget-is-accepted`, `assumption-may-carry-a-dependency` | -| A5 | Every legal event interleaving of at most four units is enumerated, and every publication at every prefix is checked for its obligations, its inputs and its source | proven | `tests/integration/ooo-publication-invariants.test.ts` (18 cases) over `src/integration/task-semantics-interleavings.ts`. The enumeration is a merge of per-unit scripts, capped at four units and six events, and the checks read the derived view rather than a second model: a dispatch must rest on closed, accepted inputs, a completion (the accepted set closed over the same predicate) must have its verdict bound to the bytes, its bytes present, no cancellation, and inputs that are present and accepted. Both halves are pinned: each checker condition is deleted by a mutant and caught by the case that names it - `the-completion-does-not-bind-the-verdict-to-the-bytes`, `the-input-is-not-required-to-be-accepted`, `the-input-may-be-cancelled`, `the-completion-ignores-a-cancelled-unit`, `the-dispatch-does-not-require-a-closed-input` - and the clean run over the design's script sets reports nothing. The merge's completeness rests on the enumeration's own count assertion rather than on a mutant: the case `the merge enumerates every legal order, not one of them` asserts the multinomial count, which a merge returning one order fails, and the tooth that showed it (`the-merge-enumerates-one-order`) was retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). The design's caveat stands and is quoted in the module: passing says this finite model satisfies the listed properties, not that any Agent program is correct | +| node | obligation | state | evidence | +| ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| A1 | The compiler returns a legal plan or a refusal naming task, field and obligation | proven | `tests/integration/task-semantics.test.ts` | +| A2 | The offline model dispatches without publishing or storing run state | proven | `tests/integration/task-semantics-model.test.ts` | +| A3 | The design's six discriminating cases are executable | proven | `tests/integration/task-semantics-cases.test.ts` | +| A4 | Field mapping catches budget-unit confusion, requires/deps confusion and silently dropped keys | proven | mutants `budget-inner-alias-is-dropped`, `over-maximum-budget-is-accepted`, `assumption-may-carry-a-dependency` | +| A5 | Every legal event interleaving of at most four units is enumerated, and every publication at every prefix is checked for its obligations, its inputs and its source | proven | `tests/integration/ooo-publication-invariants.test.ts` (18 cases) over `src/integration/task-semantics-interleavings.ts`. The enumeration is a merge of per-unit scripts, capped at four units and six events, and the checks read the derived view rather than a second model: a dispatch must rest on closed, accepted inputs, a completion (the accepted set closed over the same predicate) must have its verdict bound to the bytes, its bytes present, no cancellation, and inputs that are present and accepted. Both halves are pinned: each checker condition is deleted by a mutant and caught by the case that names it - `the-completion-does-not-bind-the-verdict-to-the-bytes`, `the-input-is-not-required-to-be-accepted`, `the-input-may-be-cancelled`, `the-completion-ignores-a-cancelled-unit`, `the-dispatch-does-not-require-a-closed-input` - and the clean run over the design's script sets reports nothing. The merge's completeness rests on the enumeration's own count assertion rather than on a mutant: the case `the merge enumerates every legal order, not one of them` asserts the multinomial count, which a merge returning one order fails, and the tooth that showed it (`the-merge-enumerates-one-order`) was retired on 2026-09-24 ([decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)). The design's caveat stands and is quoted in the module: passing says this finite model satisfies the listed properties, not that any Agent program is correct | ## B. Persistence (design: "进入持久化接入时另须证明") | node | obligation | state | evidence | | ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| B1 | Two rounds with the same taskId do not collide | proven | `tests/integration/ooo-run-namespace.test.ts`: the case `two runs in one store do not collide, do not see each other, and cancel separately` asserts each half directly - a claim in one run leaves the other run's row for the same task id alone (read as a raw row, which is the assertion the namespace experiment record shows was needed), a cancellation is per run, and a reopen by name sees the right one. The tooth `claim-is-not-scoped-to-its-run` was retired on 2026-09-24, once that case stated the rule ([record](../experiments/execution/ooo-run-namespace-2026-09-13.md), [decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)) | +| B1 | Two rounds with the same taskId do not collide | proven | `tests/integration/ooo-run-namespace.test.ts`: the case `two runs in one store do not collide, do not see each other, and cancel separately` asserts each half directly - a claim in one run leaves the other run's row for the same task id alone (read as a raw row, which is the assertion the namespace experiment record shows was needed), a cancellation is per run, and a reopen by name sees the right one. The tooth `claim-is-not-scoped-to-its-run` was retired on 2026-09-24, once that case stated the rule ([record](../experiments/execution/ooo-run-namespace-2026-09-13.md), [decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)) | | B2 | A retry does not deliver twice | proven | `tests/core/task-board-deliverable.test.ts`; mutants `stale-claim-may-deliver-again`, `deliverer-may-judge-its-own-work` | | B3 | A transaction failure cannot write a verdict without its association | proven | `tests/integration/ooo-transition-atomicity.test.ts`; mutant `swallowed-failure-still-commits`; the two teeth that pinned the boundary itself (`a-method-opens-its-own-transaction`, `nested-write-transaction-is-allowed`) were retired on 2026-09-24, because `a write entry reached inside a transition without a port is refused, not nested` states the refusal and `the store runs its transaction boundary in exactly one place` counts one BEGIN, one COMMIT and one ROLLBACK in the store's source, all three inside `writeTransaction` | | B4 | A retained board entry is not cleared by TTL | proven | `tests/core/task-board-retention.test.ts`; mutants `prune-ignores-retention`, `bounded-pin-never-expires` | @@ -195,7 +195,7 @@ separate claims. | D8 | Shutdown stops new work, lets the work in flight finish, then closes once | proven | `src/cli/service.ts`: `close()` refuses while calls are in flight, `drain()` is the awaited middle step, and `inFlight` counts what the daemon accepted. `tests/cli/service-drain.test.ts` holds an accepted call in flight (search opens the store, so the counter sees work a synchronous close could not), requires `drain(0)` to fail rather than return quietly, then closes once; `tests/cli/service.test.ts` 51/0 and `archive-shutdown` 3/0 are the regression evidence | | D9 | A post-commit notification failure is recorded, not thrown | proven | `tests/integration/ooo-post-commit-notification.test.ts`: a unit whose post-commit `afterCommit` throws still submits as `accepted`, its artifact is the board's, and `lastNotificationFailure()` reports the message, while the same unit with a reachable notification reports nothing. Verified by mutation (hand-run, 2026-09-18): deleting the `try`/`catch` in `submit()` makes the notification escape and the case fails. The implementer was always the board (`src/integration/ooo-board.ts`, `lastNotificationFailure()`); the evals suite that used to pin it drove a real round only because the artifact envelope is the board's business, and the round was retired ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The seam took three tries to find - `publishReady` is private and is called without a port at the commit, and `afterCommit` also runs on non-commit paths, so the failure is gated on a commit having landed | | D10 | `runCycle`'s early-cancel branch is reachable, or it is dead code | not applicable | **Not applicable: the branch went with the round.** It was reachable and pinned twice (`evals/ooo-execution/early-cancel.test.ts` failed with "database is not open" when the round closed the store it borrowed; verified by hand, restored byte-identically), and the retirement deleted both the branch and its subject ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). The property it stood behind - nobody who borrows a store closes it - now holds by construction, because no client opens one: the daemon is the only writer and drivers borrow a served store (D14) | -| D11 | The run records live in the store's schema, and their typed writes join the store's transaction | proven | `task_run_manifest`, `task_run_tasks` and `task_run_facts` are created by `migrate()` like every other table (`tests/core/store/schema.test.ts` names them), and `NmgStoreBase` writes them through `registerTaskRun`, `freezeTaskRunTask` and `appendTaskRunFact`, each taking the same optional port the board writes take: outside a transition it opens one, inside it joins the caller's. What the cases prove (`tests/core/store/task-runs.test.ts`, 8/8): a board write and a run fact written through one port land together, and when the fact write fails the board row does not survive it; a retried fact is recorded once and keeps its first sequence; two runs in one store keep separate task and fact namespaces; a run's plan cannot be replaced and a frozen task cannot be redefined; an entry bound by two runs is refused rather than answered with one of them; the reads register nothing. Mutant: `a-second-run-adopts-a-bound-entry`. Two further teeth were retired here on 2026-09-24, because the cases state the rules: `the-run-fact-opens-its-own-transaction` (`a board write and a run fact land together, and neither lands alone` opens one transition through the port, asserts both halves, then asserts neither survives a failure) and `a-second-plan-overwrites-the-frozen-one` (`a run registers once, and a second plan for the same run is refused` asserts the refusal by name and that the manifest still reads the first plan). Two rules this row states are carried by the cases rather than by a mutant: `appending the same fact twice records it once and keeps the first sequence` and `freezing a task twice is a no-op, and a different definition for it is refused` say the retry identity and the frozen definition in their own assertions, so `a-retried-run-fact-is-appended-twice` and `a-frozen-task-is-replaced-by-a-different-definition` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). One layer deliberately has no mutant of its own: the `UNIQUE (run_id, kind, task_id, attempt)` constraint is a storage-level backstop behind the explicit identity check, so a mutant that dropped only the constraint would survive - the check is the deciding layer and is the one mutated above | +| D11 | The run records live in the store's schema, and their typed writes join the store's transaction | proven | `task_run_manifest`, `task_run_tasks` and `task_run_facts` are created by `migrate()` like every other table (`tests/core/store/schema.test.ts` names them), and `NmgStoreBase` writes them through `registerTaskRun`, `freezeTaskRunTask` and `appendTaskRunFact`, each taking the same optional port the board writes take: outside a transition it opens one, inside it joins the caller's. What the cases prove (`tests/core/store/task-runs.test.ts`, 8/8): a board write and a run fact written through one port land together, and when the fact write fails the board row does not survive it; a retried fact is recorded once and keeps its first sequence; two runs in one store keep separate task and fact namespaces; a run's plan cannot be replaced and a frozen task cannot be redefined; an entry bound by two runs is refused rather than answered with one of them; the reads register nothing. Mutant: `a-second-run-adopts-a-bound-entry`. Two further teeth were retired here on 2026-09-24, because the cases state the rules: `the-run-fact-opens-its-own-transaction` (`a board write and a run fact land together, and neither lands alone` opens one transition through the port, asserts both halves, then asserts neither survives a failure) and `a-second-plan-overwrites-the-frozen-one` (`a run registers once, and a second plan for the same run is refused` asserts the refusal by name and that the manifest still reads the first plan). Two rules this row states are carried by the cases rather than by a mutant: `appending the same fact twice records it once and keeps the first sequence` and `freezing a task twice is a no-op, and a different definition for it is refused` say the retry identity and the frozen definition in their own assertions, so `a-retried-run-fact-is-appended-twice` and `a-frozen-task-is-replaced-by-a-different-definition` were retired on 2026-09-24 ([decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)). One layer deliberately has no mutant of its own: the `UNIQUE (run_id, kind, task_id, attempt)` constraint is a storage-level backstop behind the explicit identity check, so a mutant that dropped only the constraint would survive - the check is the deciding layer and is the one mutated above | | D12 | The binding of a logical task to its board entry is a run fact, and one routing rule decides what a binding makes managed | proven | `bindRunEntry` in `src/integration/task-coordinator.ts` records the binding as the run fact `entry-bound` (the name the D11 store tests already used), and it is the binding that the store's fence reads, so a bound entry's lifecycle write is refused unless it runs inside the run's coordinated transition. The op joins a caller's transaction when given a port, which is what lets a host create the entry and bind it in one transition: `tests/integration/ooo-managed-adopt.test.ts` (5 cases) checks that the board never holds an entry the run does not (a failure after both writes leaves neither), that a retry on the same task and attempt is not a second binding while another entry for that task and attempt is refused rather than silently dropped, and that every refusal names a stored fact (unregistered run, cancelled run, a task the run never froze, an entry that is not on the named channel, an entry another run already carries). The idempotence half rests on that case rather than on a mutant - `a-second-entry-rebinds-the-task` was retired on 2026-09-24 - because the case asserts the retry's shape and the refusal by name. `coordinatedEntryWrite` moves to the same module as the single routing rule, so the daemon's board verbs and any in-process writer decide "managed" in one place instead of each caller's belief about the entry; `src/cli/service.ts` no longer keeps its own copy. Mutants: `an-adopted-entry-takes-the-direct-path`, `a-binding-ignores-whether-the-task-was-frozen`, `a-binding-does-not-check-the-entry-exists`, `one-entry-is-bound-to-two-tasks`, `the-daemon-verb-skips-the-coordinated-path` (re-anchored to the daemon's claim handler, since the routing it mutates moved out of that file) | | D13 | The run surface another process reaches is the daemon's: register, freeze, bind, cancel and status, and nothing else writes run facts | proven | `taskRun` in `src/cli/protocol.ts` over `registerRun` / `freezeRunPlan` / `bindRunEntry` / `cancelRun` / `taskRunStatus` / `createBoundEntry` in `src/integration/task-coordinator.ts`; `tests/cli/task-run-surface.test.ts` (11 cases) drives all of it through `service.invoke`, so what it exercises is the wire a second process uses, and it asserts `hello.methods` carries `taskRun` - the advertisement a client gates the `adopt` field on. Registration is the only transition that does not need a run to exist already (`coordinateRunWrite` refuses an unregistered run, which is what every later transition and every managed write rests on). A plan freeze is one transition: the request's array order becomes the positions, a dangling dependency and a self-dependency are refused by name while the plan is still only a proposal, and a batch the store refuses leaves the earlier tasks unfrozen - `the-plan-freezes-one-task-per-transaction`, `every-task-is-frozen-at-position-zero`, `the-plan-may-freeze-a-dangling-dependency`, `a-task-may-depend-on-itself` - and a cancelled run takes no further plan (`a-cancelled-run-takes-a-new-plan`). Four teeth named in this row were retired on 2026-09-24, each because `tests/cli/task-run-surface.test.ts` states its rule through the daemon's own surface: `the-plan-freezes-one-task-per-transaction` (`a refused freeze leaves the plan exactly as it was`), `a-cancelled-run-takes-a-new-plan` (`a cancelled run takes no further plan`), `the-entry-is-created-before-its-binding-is-checked` (`adoption is part of the transition that creates the entry, so a refusal leaves no entry`) and `status-registers-the-run-it-cannot-find` (`status is a read: an unknown run has no manifest and is not registered by being asked`). Adoption rides the transition that creates the entry (`createBoundEntry` takes the port), so a refusal leaves no entry behind (`the-entry-is-created-before-its-binding-is-checked`) and the wire cannot drop the request silently (`the-wire-drops-an-adoption-request`), which is the one failure mode the epoch rule in `design.md` is about. The binding fact records the channel that names its entry - that is what lets `status` resolve a binding without searching the board (`a-binding-does-not-record-its-channel`). `cancel` is the only writer of `run-cancelled`: before this row `src/` had none, so the state the fence and the dispatch derivation both read was reachable only from a test (the gap G4 named); a run-level cancellation carries the schema's empty task id (`a-run-cancellation-names-a-task`), a task-level one requires the task to have been frozen (`a-cancellation-ignores-whether-the-task-was-frozen`, `an-unknown-run-can-be-cancelled`), and cancelling twice is one fact because the fact's own identity is the duplicate key. `status` is a read and derives nothing: a run the store does not know stays unknown (`status-registers-the-run-it-cannot-find`), and ready/blocked/accepted stay with the shared pure function rather than with this view. The CLI exposes the operator half (`nmg run status`, `nmg run cancel`), which is also what the "every RPC method is a CLI command or intentionally RPC-only" gate asks for. One thing this pass corrected about itself: the first version's parser also computed a `position` per task, which `freezeRunPlan` immediately overwrote - the mutant that changed it survived, which is how the dead field was found, and it was deleted rather than kept | | D14 | The evidence drivers reach the board through the daemon and open no database of their own | proven | `evals/ooo-execution/round-client.ts` resolves the daemon from the store's own lease and refuses by name when nothing serves it; `round-host.ts` serves an existing round store the way the product daemon does (`NmgService` + `serveHttp` + the store's lease) and is that store's only writer while it runs; `board-worker.ts`, `board-deliver.ts` and `board-judge.ts` now use the protocol's board verbs (`read`, `claim`, `release`, `deliver`, `judge`) and no longer import the store. `tests/integration/ooo-evidence-drivers.test.ts`: both end-to-end cases host the store and spawn the drivers as separate processes against the endpoint (spawned, not `spawnSync`, because the test process is what answers them); the managed case registers a run and freezes its plan through `taskRun`, adopts the entry in the same `put` that creates it, then runs the worker and the judge as separate processes - the run's log afterwards is exactly `entry-bound`, `board-claim`, `board-deliver`, `board-judge`, with the binding resolved to the entry the worker claimed and its verdict accepted. Two cases hold the boundary itself: a driver with no daemon refuses by name instead of falling back to the file, and no `board-*.ts` driver may mention the store or omit the round client. Mutants: `the-round-client-does-not-require-a-daemon`, plus the drivers' two earlier ones (`worker-reads-only-one-reporter-shape`, `judge-may-judge-its-own-delivery`), re-pointed at the renamed case. Two further rules came out of building this, and both are about the same defect the first attempt hit: a process that serves an endpoint cannot answer a call to it while it is blocked, so "the answer will come" is not an assumption a client may make. `round-client.ts` therefore bounds every call (`ROUND_CALL_TIMEOUT_MS`, an opt-in bound on the thin client's `httpCall`, which by default still leaves the platform's own) and reports a timeout by naming the bound, the endpoint and the pid that serves it rather than the transport's; and it refuses when the lease's pid is the caller's own, directing an in-process host to `host.call(...)` (`round-host.ts`), which is the design's shape for an offline host. The test is now a client too: it starts the host as its own process (`[host] serving `, idle timeout as the backstop for a host whose owner died) instead of hosting in-process, and asks it to shut down rather than killing it, so what runs on the way out is the release path - a lease held by a process that is gone is a store nothing can serve. Cases: "a client refuses to call the endpoint its own process serves", "a call to a host that never answers gives up in seconds and names the reason" (raced against its own 5s deadline, so the check of the bound is itself bounded), "a host releases its lease when it stops, so the next host can take the store". Mutants: `the-round-client-calls-the-endpoint-it-serves`, `the-round-client-has-no-limit-on-how-long-it-waits`, and - new target, `src/cli/http-server.ts` - `the-serving-process-never-releases-its-lease`. Not covered by this row: `live-continuation.ts` still creates its own store and constructs `BoardAdmission` - the runner half of the same design sentence and the next obligation. `round-runner.ts` did the same and was retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)). `a-driver-falls-back-to-opening-the-store` was retired on 2026-09-24: the case above asserts both halves of the boundary (the refusal by name, and that no driver mentions the store), which is the check the tooth used to stand in for. | @@ -273,7 +273,7 @@ and the paid stage then ran on a family held out of it. | F2a: the plan has one home, and the round's log names it | landed | the plan is one value with one home: the spec the driver runs (`PlanDriverSpec.plan`) and the run manifest the store freezes (D12). F2a's round-side carriers (`DEFAULT_ROUND_PLAN`, `CycleOptions.plan`, `openRoundStore(path, plan)`, `round-plan.test.ts` with its 6 cases) were retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); [the arms' driver decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) is what made the driver the home in the first place | | F2b: a research-side driver for arbitrary legal plans | landed, one arm | `evals/ooo-execution/plan-driver.ts` (`runPlan`, `comparePlanSlots`, `verifyParent` path, refusal-naming CLI) + `plan-driver.test.ts` (15 cases) + 12 named mutants (`tools/mutation-teeth.ts`, target `evals/ooo-execution/plan-driver.ts`) + `BoardAdmission.candidates()` (the ordered legal set; `next()` is its head, with its own mutant) + [the decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) | | F2c: pick the parent task from the sweep's turning point | landed | Two families of the shape the design's parent family needs - a frozen interface, three independent builders, one summary that depends on all three - as a spec pair each: `evals/ooo-execution/fixtures/report/` (the instrument's own) and `fixtures/pipeline/` (**held out**: written after the driver, and not the family anything was tuned against). The coarse spec is one unit over the whole task checked by every frozen test; the fine spec is four units with per-unit checks and no `join`, so the parent check composes all four. Sibling units are code-independent by construction (the summary takes the derived values as parameters), which is what lets a unit be verified before its siblings exist. Offline, with the instrument's own answers and a wrong one: `evals/ooo-execution/families.test.ts` (8 cases over both families) - both plans accept, a wrong answer is rejected by its own check and takes the composition with it, and a unit that declares no checks and has no file-wide list is refused by name. Driver support this needed: per-unit `checks` with the file-wide list as a fallback, a `canned` worker (the instrument's answer, so the family's acceptance is shown before any model is paid), and the parent composition fix recorded in the pilot's experiment record | -| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-driver-awaits-each-unit-instead-of-the-batch`. `the-driver-declares-one-slot-whatever-the-spec-says` was retired on 2026-09-24, because `a declared slot count is reached, and the claims overlap in time` asserts the requested count, the used count and the overlap. The budget's licence and publication rules are pinned by `evals/ooo-execution/board-slots.test.ts`'s own case instead: `a declared budget holds two claims at once, and the store is why each handoff is directed` refuses a budget above one with no target by that message, asserts one handoff per startable task with `serialState` null for each, keeps every startable handoff across a republish by identity, and claims the non-head first - so `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected` and `a-multi-slot-handoff-is-published-un-directed` were retired on 2026-09-24 ([decision](../decisions/proposed/2026-09-24-mutants-are-derived-not-anchored.md)). [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | +| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-driver-awaits-each-unit-instead-of-the-batch`. `the-driver-declares-one-slot-whatever-the-spec-says` was retired on 2026-09-24, because `a declared slot count is reached, and the claims overlap in time` asserts the requested count, the used count and the overlap. The budget's licence and publication rules are pinned by `evals/ooo-execution/board-slots.test.ts`'s own case instead: `a declared budget holds two claims at once, and the store is why each handoff is directed` refuses a budget above one with no target by that message, asserts one handoff per startable task with `serialState` null for each, keeps every startable handoff across a republish by identity, and claims the non-head first - so `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected` and `a-multi-slot-handoff-is-published-un-directed` were retired on 2026-09-24 ([decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)). [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | | F3: real-model pilot (A 3 / B 3 / C 2, current pi model, directional only) | landed | `evals/ooo-execution/pilot.ts` (`--live` required, `--report` to re-aggregate recorded runs with no model call, refuses a merge of two instruments, seeded arm order) + [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md). Fixed: `deepseek/deepseek-v4-flash`, envelope limits `turns 6 / reads 3 / 120 s` for every arm, A 3 / B 3 / C 2, arms drawn from a seeded shuffle, the held-out `pipeline` family. Measured (8 runs, 23 model calls, every run complete, all 8 parent checks accepted): per run A 8.6 s / 11.2 k tokens, B 21.5 s / 31.5 k tokens, C 16.2 s / 29.9 k tokens, host 1.6 s / 3.4 s / 3.6 s. All three of F1's expectations held: B is 2.5× A in wall time and 2.8× in tokens (one slot buys nothing), C recovers part of it (0.76× B) and not the 2× a pure model-call overlap would give, and the host cost grows with candidates rather than slots. Wasted cost 0, human intervention 0. **Directional only**: n = 8, one model, one held-out family, and no quality difference was available to measure - every arm accepted everything | | F4: execution fusion - legality, accounting and the driver policy | landed, live arm measured | Rules: `sharedSessionLegal`/`fusionSuccessors`/`fusionCandidates` in `src/integration/ooo-execution.ts` (the design's five conditions, one line each, composed with the board's candidate answer) + `tests/integration/ooo-fusion.test.ts` (12 cases) + 8 named mutants (target `src/integration/ooo-execution.ts`). Accounting: `fusionAccounting`/`fusionVerdict` + a fusion block in `cost-model.ts --sweep` - two lines kept apart, `unmeasured` until a run prices the session startup + `cost-model.test.ts` (12 cases) + 4 named mutants. Policy: `PlanDriverSpec.fusion` + `PlanRun.sessions` in `evals/ooo-execution/plan-driver.ts`, with the board's candidate set as the authority on staleness/cancellation/delivery/waits + `plan-driver.test.ts` (15 cases) + 11 named mutants. Scoped sweeps on this revision: `src/integration/ooo-execution.ts` 17 of 17 caught, `evals/ooo-execution/cost-model.ts` 4 of 4, `src/core/store/clock.ts` 4 of 4, `evals/ooo-execution/plan-driver.ts` 11 of 11, each restored byte-identically. A driver mutant that only restated `sharedSessionLegal`'s own rule was deleted rather than kept: the suite could not distinguish it from the shared predicate, which is what one home for that rule means. Both functions this slice pushed above the complexity limit (`sharedSessionLegal` 18, `runOneUnit` 16) were brought under it by extracting helpers, not by raising the threshold; `npm run lint` and `npm run check` are clean on this revision. **The live half landed.** `createPiSessionRunner` holds one Pi session and its tool surface across units (each unit re-points one mutable `UnitState` box; `patchSessionInput` is the one place a unit's prompt, snapshot and bounds are built, and `executePiPatch`/`executePiSnapshot` are thin callers of it), `PiRun` separates a unit's own `tokens`/`cacheRead`/`cacheWrite` from the session's `sessionTokens`, and `piSessionWorker` + `--session-runner` hold one runner per driver session. First paid D arm (2026-09-19, `deepseek-v4-flash`, 2 units, `--slots 1`, `turns: 6`, 3 reps per bound, only `fusion.unitsPerSession` differing): fused ran one session of two units and the control two sessions of one, quality parity in all six runs (every unit accepted), median 22 533 against 22 498 tokens and 11 048 against 12 948 ms - so no token saving yet (~1.9 s per run, ~15 % of the unfused wall, which prices the session-startup term at ~1.9 s instead of leaving it assumed) and per unit the second one cost ~8 % less while the first cost more: a chain's tool surface is the union of its units' capabilities because a session's surface is fixed at creation, so a unit can spend a turn on a tool that refuses by name. Also fixed here: `specFrom` had silently dropped a spec file's `fusion` block, so a spec asking for fusion ran as the control arm. | | F5: speculation lifecycle - one declared fact, three outcomes | landed (offline); the paid E arm ran once, no gain claimed | `SpeculationAssumption` / `ResolvedPredicate` / `SpeculationCandidate` / `isBoundedSpeculation` / `speculationOutcome` in `src/integration/ooo-execution.ts`, beside the fusion conditions: the assumption is a declaration the summary binds to (it never discovers for itself that the guess was false), the first experiment's bounds are a predicate (exactly one pending fact, nothing prepared from the guess - a speculative successor or an irreversible operation each refuse it by name), and the outcome has three states rather than two - **true** publishes, **false** discards the candidate and returns `sessionReusable: false`, which is what makes "失效会话不能复用到真实路径" a rule the caller must honour instead of a note, and **unknown** (no reading, an unattested reading, or evidence about another version) waits without publishing. Asking for the outcome of a candidate that is not the bounded shape throws rather than folding a fourth state into the three. Cases: `tests/integration/ooo-speculation.test.ts` (9), the last of which joins this half to fusion's condition 5 - an invalidated branch is not a legal predecessor for the real path. Mutants: 4 (`speculation-guesses-several-facts-at-once`, `an-unattested-reading-counts-as-evidence`, `evidence-about-another-version-is-the-same-fact`, `a-contradicted-guess-keeps-its-session`); the target's sweep is 22 of 22 caught. The E arm's instrument now exists (`evals/ooo-execution/speculation-pilot.ts`, registered) and ran once (2026-09-19, 6 paid units, ~43 k tokens): it decides the guessed fact, prepares the candidate, applies `speculationOutcome`, and verifies a published candidate with the unit's own frozen check. **No result is claimed**: the quality term was false in all four verified candidates, so by the design's own rule the latency and cost shape may not be reported as a gain. The search behind those failures is now closed and its first reading was wrong: `artifactEnvelope` builds two legitimate shapes (a patch, and a conclusion with `kind, conclusion, summary, evidence, citations`), and the instrument had fed every artifact to the patch reader. Eight of nine attempts answered with a conclusion, which this unit's check cannot pass and the board would refuse; the one patch attempt failed on a real mistake (`rows` for `lines`). The instrument now reads by kind, keeps every artifact, candidate tree and check output, and the run is archived. Measured outcome of the arm at this shape: the post-fact cost drops from ~6.2 s of work to 175 ms of verification when the fact holds, the false-fact case wastes 20 332 tokens, and the prepared candidate was publishable in 0 of 3 holding reps - so the cost is real, the gain is not, and the binding constraint is the candidate's admissibility |