From 45d6767c9367bc9e23a5506ab9f3d3c3f0e3e5c8 Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 23:34:50 +0800 Subject: [PATCH 1/2] tools(mutation-teeth): an expect that names no case says so, instead of surviving A tooth names the case that must fail, and a sweep filters a suite by that name so one tooth costs one case instead of a whole suite. When the name matches no case, the filtered run passes with no case in it - and that pass was reported as a surviving mutant. The two are different defects: one is a rule nothing checks, the other is evidence whose name went stale. This is how two teeth were read as broken while they were only misnamed, and the tool had no way to say which had happened. - `ranACase` (tools/mutation-anchor.ts, pure, no filesystem) reads the run's own report: a marked line that is not the suite file is a case, a suite line alone means the filter matched nothing, and an unrecognized reporter answers unknown - a tooth is not accused on evidence the harness cannot parse. The suite file's own line is what makes the answer possible: it runs whatever the filter did. - The runner asks it before it believes a pass, reports the tooth as `misnamed` with the name it wanted, and counts it as a problem (non-zero exit). The whole-suite fallback is not run for it: the defect is the name, not the coverage. - `tests/tools/mutation-anchor.test.ts`: one case covering the file line alone, a matched case, a failed case, two suites where nothing matched, and the unknown reporter - 35 cases in the file now. - The record's own failure-mode note moves from "open" to answered, in the same commit. Measured: the whole register 110 of 110 caught by the named test, all 110 by the case each expect names, 22 of 22 targets restored byte-identically, exit 0 in 162 s - so the new check raises no false alarm anywhere in the register. A deliberately misnamed tooth reads `misnamed (no case in this target's suites is named "...")`, the file is restored byte-identically, and the run exits non-zero. --- ...-09-24-mutants-are-derived-not-anchored.md | 25 ++++--- ...-mutants-are-derived-not-anchored.zh-CN.md | 67 ++++++++++--------- tests/tools/mutation-anchor.test.ts | 24 ++++++- tools/mutation-anchor.ts | 25 +++++++ tools/mutation-teeth.ts | 56 ++++++++++++---- 5 files changed, 141 insertions(+), 56 deletions(-) diff --git a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md index 23995016..0e01c9b8 100644 --- a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md @@ -96,7 +96,7 @@ favour of checks that state their rule.** `tools/mutation-anchor.ts` holds the resolver: `Mutant`, `Derive`, `Site`, `matchText`, `locate`, and a table of resolvers, one per operator, so adding an operator is an entry in that table plus its name in the type union. It holds no state and reads no file, so the sweep and the anchors-only pass ask the same -function the same question and `tests/tools/mutation-anchor.test.ts` proves it over source strings - 34 +function the same question and `tests/tools/mutation-anchor.test.ts` proves it over source strings - 35 cases, one per operator plus every refusal, with no filesystem access. Where the mutation determines its own bytes (`false`, `true`, the negated operator, the call's receiver, the guard's body, the empty collection) nothing is stored. @@ -131,9 +131,11 @@ widenings. The widenings were all found by real teeth failing to convert, never **Readings, at this revision:** `anchors: 110 of 110 resolve, over 22 targets`; `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored -byte-identically, exit 0 in 160 s; `tests/tools/mutation-anchor.test.ts` 34 pass; `test:product` 1580 -pass; `verify:static` exit 0; lint 0 findings; complexity gate ok. An earlier full-run reading on the -same arc: 136 teeth, 136 of 136 caught, all 136 by name, 217 s. +byte-identically, exit 0 in 162 s; `tests/tools/mutation-anchor.test.ts` 35 pass; `test:product` 1580 +pass; `verify:static` exit 0; lint 0 findings; complexity gate ok. The run that added the misnamed check +reads exactly the same, which is also what says the check raises no false alarm across the whole +register. An earlier full-run reading on the same arc: 136 teeth, 136 of 136 caught, all 136 by name, +217 s. **Corrections this record carries, each caught by a run rather than by review:** @@ -242,7 +244,14 @@ existing.deliveredBy)` has to know the operator. Mitigation: the operators are f that the mutant is still caught, and a tooth whose `expect` went stale still passes it. Mitigation: the full sweep remains the standing rule before a push, and the sweep is the reading that counts - not the number of teeth, but **the number of teeth whose named case is the one that fails**. -- **Open, and deliberately not decided here:** the seven orphan teeth whose rules have a check and no - ledger row (whether those rules deserve rows is a question about the ledger, not about the register); - and the `expect` field's own failure mode, which is loud but only on the slow path - a report that says - whether the named case exists in the target's suites would make it immediate. +- **The `expect` field's own failure mode is answered.** It was loud but ambiguous: a name no case + carries made the filtered run pass with no case in it, and that pass was reported as a surviving + mutant - which is how two teeth were read as broken while they were only misnamed. `ranACase` reads the + run's own report before believing a pass: a marked line that is not the suite file is a case, a suite + line alone means the filter matched nothing, and an unrecognized reporter answers unknown rather than + accusing a tooth. A misnamed tooth is reported as `misnamed` with the name it wanted, it is a problem + with a non-zero exit, and the whole-suite fallback is not run for it - the defect is the name, not the + coverage. +- **Still open, and deliberately not decided here:** the seven orphan teeth whose rules have a check and + no ledger row (whether those rules deserve rows is a question about the ledger, not about the + register). diff --git a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index ccfa4345..11f4a1eb 100644 --- a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -12,13 +12,13 @@ 2026-09-24 实测,当时 149 颗 mutant、22 个目标: -| 锚点指向什么 | 颗数 | 占比 | 已经限定到成员 | -| ---------------------------------------------- | ---- | ---- | -------------- | -| 守卫或谓词里的一项被中和(G) | 64 | 43% | 30 | -| 值、实参、下标或被调被替换(V) | 58 | 39% | 26 | -| 语句或调用被删除(R) | 11 | 7% | 6 | -| 插入代码:一条语句、一个分支或第二份拷贝(I) | 8 | 5% | 4 | -| 整个表达式被替换(E) | 8 | 5% | 5 | +| 锚点指向什么 | 颗数 | 占比 | 已经限定到成员 | +| --------------------------------------------- | ---- | ---- | -------------- | +| 守卫或谓词里的一项被中和(G) | 64 | 43% | 30 | +| 值、实参、下标或被调被替换(V) | 58 | 39% | 26 | +| 语句或调用被删除(R) | 11 | 7% | 6 | +| 插入代码:一条语句、一个分支或第二份拷贝(I) | 8 | 5% | 4 | +| 整个表达式被替换(E) | 8 | 5% | 5 | 两个读数决定了本记录。**89% 的牙(149 之 133)是"一个点位 + 少数几个算子之一"**——让条件失效、中和一项、替换实参、删除语句——也就是算子才是意图、点位是附带,把点位存成文本只是一个选择而不是必需。而且 **149 之 78(52%)完全没有结构性作用域**:一个在文件里任意匹配的字节片段,正是最先死掉的形状。规划期间观察到的两处失效都是这个形状:一条语句里层调用被内联、一处拒绝的理由消息被抽成 `subject` 表达式。旁边还出现了第三种:两颗牙点名的用例已经抓不住它们(`fusion-continues-from-an-unverified-answer`、`next-is-not-the-head-of-the-ordered-candidates` 读数变成"被整个 suite 抓到,不是点名用例"),这是同一层耦合往外挪了一格——锚点还在,而"这条用例是唯一会失败的那条"这句话旧了。 @@ -41,35 +41,35 @@ **登记处现在是 110 颗牙、22 个目标条目、覆盖 21 个文件,每一颗都是"名字 + 算子 + 选择器",没有一颗手写锚点。26 个算子。已有 37 颗牙退役,改为由陈述了该规则的检查承担。** -`tools/mutation-anchor.ts` 是解析器:`Mutant`、`Derive`、`Site`、`matchText`、`locate`,以及一张"一个算子一个解析函数"的表,所以加一个算子就是表里加一行、类型里加一个名字。它不持有状态、不读文件,因此 sweep 与 anchors-only pass 问的是同一个函数的同一个问题,而 `tests/tools/mutation-anchor.test.ts` 用源码字符串证明它——34 条用例,每个算子一条并覆盖各自的拒绝情形,不碰文件系统。凡是变异自己决定字节的(`false`、`true`、取反后的算子、调用的接收者、守卫的 body、空集合),什么都不存。 - -| 算子 | 现成目录里的名字 | 颗数 | -| ------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------- | ---------- | -| `condition-never`、`condition-holds` | pitest `FALSE_RETURNS` / `TRUE_RETURNS`;Stryker `ConditionalExpression` | 45 + 3 | -| `neutralize-term` | pitest `NEGATE_CONDITIONALS`;Stryker 逻辑/布尔字面量 | 9 | -| `negate-condition` | pitest `NEGATE_CONDITIONALS`;Stryker 布尔字面量 | 2 | -| `negate-comparison` | pitest `NEGATE_CONDITIONALS`;Stryker `EqualityOperator` | 3 | -| `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | -| `drop-statement` | Stryker `BlockRemoval`;pitest `VOID_METHOD_CALLS` | 8 | -| `remove-call` | Stryker filter/slice/sort 删除;pitest `VOID_METHOD_CALLS` | 2 | -| `replace-call` | Stryker `MethodExpression`;pitest `CONSTRUCTOR_CALLS` | 4 | -| `replace-argument`、`replace-property`、`replace-initializer` | pitest `PRIMITIVE_RETURNS`、`INLINE_CONSTS`;Stryker 字面量类 | 8 + 10 + 7 | -| `replace-iterable` | Stryker `ArrayDeclaration`;pitest `EMPTY_RETURNS` | 3 | -| `replace-index` | (目录里没有;最接近的 `FirstToLast` 指的是名为 `first` 的方法) | 2 | -| `replace-literal-fragment` | SQLMutation 的子句级算子,降到片段粒度 | 3 | +`tools/mutation-anchor.ts` 是解析器:`Mutant`、`Derive`、`Site`、`matchText`、`locate`,以及一张"一个算子一个解析函数"的表,所以加一个算子就是表里加一行、类型里加一个名字。它不持有状态、不读文件,因此 sweep 与 anchors-only pass 问的是同一个函数的同一个问题,而 `tests/tools/mutation-anchor.test.ts` 用源码字符串证明它——35 条用例,每个算子一条并覆盖各自的拒绝情形,不碰文件系统。凡是变异自己决定字节的(`false`、`true`、取反后的算子、调用的接收者、守卫的 body、空集合),什么都不存。 + +| 算子 | 现成目录里的名字 | 颗数 | +| ------------------------------------------------------------- | ------------------------------------------------------------------------ | ---------- | +| `condition-never`、`condition-holds` | pitest `FALSE_RETURNS` / `TRUE_RETURNS`;Stryker `ConditionalExpression` | 45 + 3 | +| `neutralize-term` | pitest `NEGATE_CONDITIONALS`;Stryker 逻辑/布尔字面量 | 9 | +| `negate-condition` | pitest `NEGATE_CONDITIONALS`;Stryker 布尔字面量 | 2 | +| `negate-comparison` | pitest `NEGATE_CONDITIONALS`;Stryker `EqualityOperator` | 3 | +| `remove-conditionals` | pitest `REMOVE_CONDITIONALS` | 1 | +| `drop-statement` | Stryker `BlockRemoval`;pitest `VOID_METHOD_CALLS` | 8 | +| `remove-call` | Stryker filter/slice/sort 删除;pitest `VOID_METHOD_CALLS` | 2 | +| `replace-call` | Stryker `MethodExpression`;pitest `CONSTRUCTOR_CALLS` | 4 | +| `replace-argument`、`replace-property`、`replace-initializer` | pitest `PRIMITIVE_RETURNS`、`INLINE_CONSTS`;Stryker 字面量类 | 8 + 10 + 7 | +| `replace-iterable` | Stryker `ArrayDeclaration`;pitest `EMPTY_RETURNS` | 3 | +| `replace-index` | (目录里没有;最接近的 `FirstToLast` 指的是名为 `first` 的方法) | 2 | +| `replace-literal-fragment` | SQLMutation 的子句级算子,降到片段粒度 | 3 | 登记处分三波转换,每一波之后都 sweep:用解析器最初那六个算子转了 59 颗,用目录点名的七个算子转了 18 颗,最后 15 颗用另外三个算子加三处扩宽转完。三处扩宽全都是**真实的牙转不过去**才发现的,没有一处是靠讨论定下来的: -| 残留需要什么 | 实际加了什么 | 颗数 | -| -------------------- | ------------------------------------------------------------------------------------------------------------------- | ---- | -| 绑定到名字的函数 | `uniqueMember` 接受绑定到变量或对象属性的函数(`claim: (store) => ...`) | 2 | -| 两个一模一样的调用点 | `in` 指"调用或字面量所在语句里的一段文本" | 3 | -| 模块级 | `within` 变可选:文件就是作用域,选择器必须在文件里唯一 | 2 | -| 构造函数调用 | 调用选择器接受 `new X(...)` 作为调用点 | 1 | -| 简写属性 | `replace-property` 把它写开(`position` 变 `position: 0`),因为值不能写在名字的位置上 | 1 | -| 下标、字面量片段 | `replace-index`、`replace-literal-fragment` | 5 | +| 残留需要什么 | 实际加了什么 | 颗数 | +| -------------------- | -------------------------------------------------------------------------------------- | ---- | +| 绑定到名字的函数 | `uniqueMember` 接受绑定到变量或对象属性的函数(`claim: (store) => ...`) | 2 | +| 两个一模一样的调用点 | `in` 指"调用或字面量所在语句里的一段文本" | 3 | +| 模块级 | `within` 变可选:文件就是作用域,选择器必须在文件里唯一 | 2 | +| 构造函数调用 | 调用选择器接受 `new X(...)` 作为调用点 | 1 | +| 简写属性 | `replace-property` 把它写开(`position` 变 `position: 0`),因为值不能写在名字的位置上 | 1 | +| 下标、字面量片段 | `replace-index`、`replace-literal-fragment` | 5 | -**本版读数:**`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,160 秒;`tests/tools/mutation-anchor.test.ts` 34 条全过;`test:product` 1580 全过;`verify:static` exit 0;lint 0 findings;复杂度门禁通过。同一条弧上更早的一次全量读数是:136 颗牙,136 of 136 caught,136 全部按名字,217 秒。 +**本版读数:**`anchors: 110 of 110 resolve, over 22 targets`;`mutants: 110 of 110 caught by the named test`,**110 颗全部由各自 `expect` 点名的用例抓住**,22 of 22 逐位元组还原,exit 0,162 秒;`tests/tools/mutation-anchor.test.ts` 35 条全过;`test:product` 1580 全过;`verify:static` exit 0;lint 0 findings;复杂度门禁通过。加上 misnamed 检查的那一次运行读数完全相同,这也就是"这条检查在整张登记处上没有假报"的证据。同一条弧上更早的一次全量读数是:136 颗牙,136 of 136 caught,136 全部按名字,217 秒。 **本记录携带的更正,每一条都是被运行结果抓出来的、不是被评审看出来的:** @@ -103,4 +103,5 @@ - **退役会丢掉一条具名用例。**关系型检查抓的是一类,而牙抓的可能是某个具体错误行为,于是行可能声称得比展示的多。缓解:退役时在账本行里、在同一个提交里点名替代它的那条检查。 - **转换是"对证据动手"。**每一次转换都触碰"这份代码被保护着"的那件证物,因此错误会安静地降低覆盖。缓解:一次一个目标、同名同用例、前后都记下抓取结果,以及那条关于活 mutant 的 postmortem 规则(任何运行前先看被测文件的状态)。 - **这个 pass 可能被信任而不是被运行。**anchors-only pass 证明点位还解析得出来,不证明 mutant 还抓得住;一颗 `expect` 已经陈旧的牙照样能通过它。缓解:全量 sweep 仍是推送前的常设规则,而真正要看的读数不是牙的数量,而是**"失败的就是它点名那条用例"的牙的数量**。 -- **开放、且刻意不在这里决定:**那七颗孤儿牙——它们钉的规则有检查、没有账本行(这些规则该不该有自己的行,是账本的问题而不是登记处的问题);以及 `expect` 字段自身的失效模式——它是响的,但只在慢路径上;一份"点名用例在目标套件里存在与否"的报告能让它立刻可见。 +- **`expect` 字段自身的失效模式已经答了。**它本是响的、却有歧义:一个没有任何用例的名字会让过滤后的运行"零用例通过",而这个通过被报成"mutant 存活"——那两颗牙就是这样被读成坏掉的,而它们只是名字写错了。`ranACase` 在相信一次通过之前先读运行自己的报告:带标记的行里若不是套件文件,就是一条用例;只有套件文件的行意味着过滤没匹配到任何东西;不认识的 reporter 则答"未知",而不是去指控一颗牙。名字写错的牙现在报 `misnamed` 并附上它要的那个名字,算 problem、退出码非零,而且不再为它跑整套 fallback——缺陷是名字,不是覆盖。 +- **仍然开放、刻意不在这里决定:**那七颗孤儿牙——它们钉的规则有检查、没有账本行(这些规则该不该有自己的行,是账本的问题而不是登记处的问题)。 diff --git a/tests/tools/mutation-anchor.test.ts b/tests/tools/mutation-anchor.test.ts index 978a0d13..71bd197e 100644 --- a/tests/tools/mutation-anchor.test.ts +++ b/tests/tools/mutation-anchor.test.ts @@ -11,7 +11,7 @@ import assert from "node:assert/strict"; import test from "node:test"; -import { locate, matchText, type Mutant } from "../../tools/mutation-anchor.ts"; +import { locate, matchText, ranACase, type Mutant } from "../../tools/mutation-anchor.ts"; const SOURCE = `function selection(tasks, slots) { const spent = slots < 1; @@ -676,3 +676,25 @@ test("a shorthand property is written out, because the value cannot go where the assert.equal(above.slice(site.start, site.end), "position"); assert.equal(site.replacement, "position: 0"); }); + +// A tooth names the case that must fail, and a sweep filters the suite by that name so one tooth costs one +// case instead of a whole suite. A name no case carries therefore makes the filtered run pass with nothing +// in it, and that pass reads exactly like a surviving mutant - the misreading this asks about. The suite +// file's own line is what makes the answer possible: it runs whatever the filter did. +test("a filtered run that matched no case is not a run that found the mutant innocent", () => { + const nothingMatched = + "\u2714 tests\\tools\\mutation-anchor.test.ts (277.9136ms)\n\u2139 tests 1\n\u2139 pass 1\n"; + assert.equal(ranACase(nothingMatched), false); + const oneMatched = + "\u2714 a shorthand property is written out (6.8911ms)\n\u2139 tests 1\n\u2139 pass 1\n"; + assert.equal(ranACase(oneMatched), true); + assert.equal(ranACase("\u2714 a case that passed\n\u2716 a case that failed\n"), true); + // A second suite in the same run, where nothing matched: the file line alone is still not a case. + assert.equal( + ranACase("\u2714 tests\\tools\\a.test.ts (12ms)\n\u2714 tests\\tools\\b.test.ts (9ms)\n"), + false, + ); + // A reporter this harness does not recognize is not evidence against a tooth: unknown, not absent. + assert.equal(ranACase("ok 1 - tests/tools/mutation-anchor.test.ts\n"), undefined); + assert.equal(ranACase(""), undefined); +}); diff --git a/tools/mutation-anchor.ts b/tools/mutation-anchor.ts index 63345a5e..bb0a2d24 100644 --- a/tools/mutation-anchor.ts +++ b/tools/mutation-anchor.ts @@ -163,6 +163,31 @@ export function matchText( return { start: match.index, end: match.index + match[0].length, retaken: true }; } +/** Did any case name run in this suite output? `true` when a case line is there, `false` when the only + * line the reporter marked is the suite file itself, and `undefined` when the output carries no marked + * line at all. + * + * A mutant names the case that must fail, and a sweep filters a suite by that name so one tooth costs + * one case instead of a whole suite. A name that matches no case is not a mutant nothing catches: the + * filtered run passes with no case in it, and that pass reads exactly like a surviving mutant. The two + * are different defects - one is a rule nothing checks, the other is evidence whose name went stale - + * so the runner asks this before it believes a pass, in the same spirit as `matchText` refusing to + * claim a site it cannot find. + * + * The suite file line is what makes the answer possible: a file's own test runs whether or not any + * case inside it matched, so a marked line whose name ends in `.ts` says nothing about the filter. + * An unrecognized reporter answers `undefined`, never `false`: a tooth is not accused on evidence the + * harness cannot read. */ +export function ranACase(report: string): boolean | undefined { + const marked = report + .split("\n") + .map((line) => /^\s*[\u2714\u2716]\s+(.*)$/u.exec(line)?.[1]) + .filter((name): name is string => name !== undefined) + .map((name) => name.replace(/\s+\([\d.]+ms\)\s*$/u, "").trim()); + if (marked.length === 0) return undefined; + return marked.some((name) => !/\.tsx?$/u.test(name)); +} + /** Where a mutant applies: by selector when it derives one, by syntax tree when it says so, by bytes * otherwise. * diff --git a/tools/mutation-teeth.ts b/tools/mutation-teeth.ts index b5a0fa0e..42fea4f1 100644 --- a/tools/mutation-teeth.ts +++ b/tools/mutation-teeth.ts @@ -6,9 +6,11 @@ * load-bearing line with a plausible-but-wrong version; the suite must then fail, and the *expected * * Two rules the config learned the hard way. `expect` is the NAME of the test that must fail - - * prose there makes a real catch read as "NOT caught". The `ast` locator is indentation-sensitive, - * so an anchor whose leading spaces no longer match the file is a stale anchor to be fixed, not a - * cosmetic difference; and a site that cannot be located is a failure, never a claimed check. + * prose there makes a real catch read as "NOT caught", and a name no case carries makes a filtered run + * pass with nothing in it, which is reported as `misnamed` rather than as a survivor. The `ast` locator + * is indentation-sensitive, so an anchor whose leading spaces no longer match the file is a stale anchor + * to be fixed, not a cosmetic difference; and a site that cannot be located is a failure, never a + * claimed check. * test* must be the one that fails. Three outcomes per target, not two: * * 1. the clean tree passes the target's suites; @@ -72,7 +74,7 @@ import { type MutationLock, } from "./mutation-lock.ts"; import { writeJsonAtomic } from "./parts/fs.ts"; -import { locate, type Mutant } from "./mutation-anchor.ts"; +import { locate, ranACase, type Mutant } from "./mutation-anchor.ts"; interface Target { readonly target: string; @@ -1386,6 +1388,9 @@ interface MutantOutcome { readonly ms?: number; /** True when the named case was what failed, rather than the whole suite catching it. */ readonly caughtByName?: boolean; + /** True when no case in the target's suites carries the name this tooth requires - the tooth cannot be + * evaluated at all, which is a defect of the evidence rather than a surviving mutant. */ + readonly misnamed?: boolean; } interface Outcome { @@ -1649,9 +1654,15 @@ for (const { target, suites, mutants: declared } of selected) { const mutantStartedAt = Date.now(); const named = runSuites(present, mutant.expect); const caughtByName = !named.ok && named.out.includes(mutant.expect); + // A filtered run passes when the name matches no case in these suites, and that pass reads exactly + // like a surviving mutant. The two are different defects, and only one of them is about the code, so + // the run's own report is asked which it is - the failure mode that made two teeth read as broken + // when they were only misnamed. A tooth with no case to name cannot be rescued by the whole suite, + // so the fallback is skipped for it: the defect is the name, not the coverage. + const namesNoCase = named.ok && ranACase(named.out) === false; // A case that never finishes is this mutant's own answer - a loop broken enough to spin - and // running the whole suite after it would only spend the same bound again to learn nothing. - const result = caughtByName || named.timedOut ? named : runSuites(present); + const result = caughtByName || named.timedOut || namesNoCase ? named : runSuites(present); const ms = Date.now() - mutantStartedAt; const caught = !result.ok && result.out.includes(mutant.expect); // A mutant that makes the case spin is not a passing mutant, but it is not an assertion failure @@ -1660,20 +1671,35 @@ for (const { target, suites, mutants: declared } of selected) { // whole suite. const didNotFinish = named.timedOut === true; mutantOutcomes.push( - caught || didNotFinish + namesNoCase ? { name: mutant.name, applicable: true, - caught: true, - caughtByName: caughtByName || didNotFinish, + caught: false, + misnamed: true, ms, - ...(didNotFinish - ? { note: `the case did not finish within ${patternTimeoutMs() / 1000}s` } - : {}), + note: `no case in this target's suites is named "${mutant.expect}"`, } - : { name: mutant.name, applicable: true, caught, caughtByName, ms, note: "survived" }, + : caught || didNotFinish + ? { + name: mutant.name, + applicable: true, + caught: true, + caughtByName: caughtByName || didNotFinish, + ms, + ...(didNotFinish + ? { note: `the case did not finish within ${patternTimeoutMs() / 1000}s` } + : {}), + } + : { name: mutant.name, applicable: true, caught, caughtByName, ms, note: "survived" }, ); - if (!caught && !didNotFinish) { + if (namesNoCase) { + problems.push( + ` mutant ${mutant.name}: the case it names is not in this target's suites, so a passing "${ + mutant.expect + }" run proves nothing about it`, + ); + } else if (!caught && !didNotFinish) { const observed = observedFailures(result.out); problems.push( ` mutant ${mutant.name}: NOT caught by "${mutant.expect}" (suite passed: ${result.ok}; observed: ${ @@ -1719,7 +1745,9 @@ for (const outcome of outcomes) outcome.mutants .map( (mutant) => - `${mutant.name} ${mutant.caught ? "caught" : "survived"}${ + `${mutant.name} ${ + mutant.misnamed ? "misnamed" : mutant.caught ? "caught" : "survived" + }${ mutant.caughtByName === false ? " (by the suite, not the named case)" : mutant.note From d88e23c04938678909760ed902a4cceb6888330c Mon Sep 17 00:00:00 2001 From: wefio <48851810+wefio@users.noreply.github.com> Date: Fri, 25 Sep 2026 23:38:46 +0800 Subject: [PATCH 2/2] docs(ledger): the seven rules no row claimed get one home each The retirement pass left seven teeth whose rules no row named. The reading was "a case but no row"; the rules themselves were recovered from the register before the retirement, each replacement case was located and its assertion read, and the answer is not seven new rows - most of these rules were already carried by a row that simply did not name the case: - `round-publication-opens-its-own-transaction` is B3's rule. Its replacement case lives in B3's own suite and asserts the transition's rollback for the publication half, saying in the assertion what the alternative would mean. Named in B3. - `round-never-releases-its-pin` is the other side of B4's retention rule. The case asserts the pins held while a round runs, then cancels and asserts the retention list is empty, because both kinds of pin go with the values they protected. Added to B4. - `ordered-mode-becomes-any-topological-order` had no row at all: the design's sentence about publication order had never been given one. It is A6 now, evidenced by the case that asserts the exact declared order, one ordered plan, and that the baseline never gains order freedom. Row counts: A 5 -> 6. - The four judge-loop rules are F2b-slot's and F2c's: the batch rule and the duplicated dispatch share one case, `a unit's check is outstanding while an independent unit's worker runs`, and the unit with no checks and the parent check are the two cases F2c's row now names. So the readings line no longer calls them orphans, and the record no longer leaves the question open - what remains open is only whether "a check but no row" is a defect at all, and here the defect was in the ledger rather than in the register. Both languages in this commit. Readings: `docs:check` 304 files, 0 errors, 0 warnings; `rtm:check` exit 0; `test:product` 1581 pass, 0 fail, exit 0. --- ...26-09-24-mutants-are-derived-not-anchored.md | 17 ++++++++++++----- ...24-mutants-are-derived-not-anchored.zh-CN.md | 4 ++-- docs/design/task-unit-semantics-obligations.md | 14 ++++++++------ 3 files changed, 22 insertions(+), 13 deletions(-) diff --git a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md index 0e01c9b8..d71c508b 100644 --- a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md +++ b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md @@ -177,8 +177,9 @@ A5, a rebinding refused by name, and the five rules of the declared budget state the third (row D11, whose prose already rested on the cases). 26 in the fourth: the insertion-shaped teeth, whose anchor is a place where code must not appear, and which a derived selector cannot express - the only honest alternatives were a permanent fragile anchor or a check that states the rule. One of that -set was kept and then converted, and seven were **orphans** - no ledger row named them, so the rule each -pinned has a check and no row. +set was kept and then converted, and seven named no ledger row at all: the rules were real, but no row +mentioned them. Each of the seven was given a home in the row that already carries its rule rather than a +row invented for it - two rows extended, one row added ([the ledger](../../design/task-unit-semantics-obligations.md)). **What is not claimed.** A derived tooth is not stronger evidence than the byte anchor it replaced: it is the same violation, caught by the same case. What it buys is that the tooth stays aimed at its rule @@ -252,6 +253,12 @@ existing.deliveredBy)` has to know the operator. Mitigation: the operators are f accusing a tooth. A misnamed tooth is reported as `misnamed` with the name it wanted, it is a problem with a non-zero exit, and the whole-suite fallback is not run for it - the defect is the name, not the coverage. -- **Still open, and deliberately not decided here:** the seven orphan teeth whose rules have a check and - no ledger row (whether those rules deserve rows is a question about the ledger, not about the - register). +- **The seven rules no row named are placed, not left dangling.** When the insertion-shaped teeth were + retired, seven of them named a rule no row claimed: a publication joining the caller's transition, a + round releasing the pins it held, the ordered mode being the declared plan order, the batch's units in + flight together, a unit with no checks refused, the parent check reporting its composed verdict. Each was + given a home in the row that already carries its rule - B3 gained the publication case (its own suite), + B4 the pin release, the ordered mode became its own row (A6, the design's own sentence about publication + order, which had no row), and the four judge-loop rules are F2b-slot's and F2c's, whose texts now name + the cases. What remains open is only whether a rule with a check but no row is a defect at all: it was + one here, but the defect was in the ledger, not in the register. diff --git a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md index 11f4a1eb..43761a32 100644 --- a/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md +++ b/docs/decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.zh-CN.md @@ -80,7 +80,7 @@ - `expect` 手打,读起来就像一颗坏掉的牙。两颗牙凭记忆重打用例名(而不是从登记处抄),结果都报 **"survived"**:工具按点名的用例过滤 suite,名字对不上任何用例时,与"没有任何用例抓到"无法区分。它是响的失败、从不是误判通过,但这是这个登记处第二次被"`expect` 自己也是一条断言"咬到。转换脚本现在从登记处**抄**这个字符串。 - 词汇长大之后重跑旧的转换脚本,把两颗牙改回了早期形态;anchors pass 一次性把两个都拒了。这正是"把这个 pass 放进静态契约"从另一个方向看到的理由:它也抓**对登记处本身的编辑**,不只抓它指向的代码漂移。 -**退役。** 37 颗牙走了,分四组,每组都先由 sweep 说明"失败的就是它点名的用例"、再读那条用例的断言。前两组十一颗(B5 行背后的两处调用点计数、32 种组合的枚举、A5 行背后的多项式计数 10、按名字拒绝的重新绑定,以及由一条用例陈述的预算五规则)。第三组两颗(D11 行,其正文本来就建立在对应用例上)。第四组 26 颗:插入形状的牙——它们的锚点是"代码不得出现的地方",derived 选择器表达不了,诚实的选择只有"永久脆弱的锚点"或"一条陈述了规则的检查"。这一组里有一颗后来也转掉了,另有七颗是**孤儿**——没有任何账本行点名它们,于是它们钉的规则有检查、没有行。 +**退役。** 37 颗牙走了,分四组,每组都先由 sweep 说明"失败的就是它点名的用例"、再读那条用例的断言。前两组十一颗(B5 行背后的两处调用点计数、32 种组合的枚举、A5 行背后的多项式计数 10、按名字拒绝的重新绑定,以及由一条用例陈述的预算五规则)。第三组两颗(D11 行,其正文本来就建立在对应用例上)。第四组 26 颗:插入形状的牙——它们的锚点是"代码不得出现的地方",derived 选择器表达不了,诚实的选择只有"永久脆弱的锚点"或"一条陈述了规则的检查"。这一组里有一颗后来也转掉了,另有七颗点名的规则没有任何账本行声称(规则是真的,只是行里没写);这七条后来各自被安顿到本来就承载它的行,或是补了一行([账本](../../design/task-unit-semantics-obligations.md))。 **没有声称的东西。** derived 牙**不比**它替代的字节锚点更强:违规是同一个、抓它的用例是同一条。它换来的是改名、重排之后仍瞄准原来那条规则,而这正是本记录要解决的那个失败。算子也不等于整个目录——它们只是本登记处的规则恰好需要的那些类。转换也不是"可以不再 sweep"的许可:`mutation:anchors` 证明的是解析成功,不是"还抓得住"。 @@ -104,4 +104,4 @@ - **转换是"对证据动手"。**每一次转换都触碰"这份代码被保护着"的那件证物,因此错误会安静地降低覆盖。缓解:一次一个目标、同名同用例、前后都记下抓取结果,以及那条关于活 mutant 的 postmortem 规则(任何运行前先看被测文件的状态)。 - **这个 pass 可能被信任而不是被运行。**anchors-only pass 证明点位还解析得出来,不证明 mutant 还抓得住;一颗 `expect` 已经陈旧的牙照样能通过它。缓解:全量 sweep 仍是推送前的常设规则,而真正要看的读数不是牙的数量,而是**"失败的就是它点名那条用例"的牙的数量**。 - **`expect` 字段自身的失效模式已经答了。**它本是响的、却有歧义:一个没有任何用例的名字会让过滤后的运行"零用例通过",而这个通过被报成"mutant 存活"——那两颗牙就是这样被读成坏掉的,而它们只是名字写错了。`ranACase` 在相信一次通过之前先读运行自己的报告:带标记的行里若不是套件文件,就是一条用例;只有套件文件的行意味着过滤没匹配到任何东西;不认识的 reporter 则答"未知",而不是去指控一颗牙。名字写错的牙现在报 `misnamed` 并附上它要的那个名字,算 problem、退出码非零,而且不再为它跑整套 fallback——缺陷是名字,不是覆盖。 -- **仍然开放、刻意不在这里决定:**那七颗孤儿牙——它们钉的规则有检查、没有账本行(这些规则该不该有自己的行,是账本的问题而不是登记处的问题)。 +- **那七条没有行点名的规则是"安顿了",不是"挂着"。**退役插入形状的牙时,其中七颗点名的规则没有任何行声称:发布加入调用方的事务、一轮在自己持有的 pin 被清除时释放它们、ordered 模式就是声明顺序、一个批次的各单元同时在途、没有检查的单元被拒、父检查报告它合成的裁决。每一条都被安顿到本来就承载它的那一行——B3 拿到发布那条用例(同一套件)、B4 拿到 pin 释放、ordered 模式自成一行为 A6(设计里关于发布顺序的那一句原本没有行),四条判官循环的规则则属于 F2b-slot 与 F2c,它们的正文现在点名了这些用例。仍然开着的只剩一个问题:"有检查、没有行"的规则到底算不算缺陷——这里算,但缺陷在账本,不在登记处。 diff --git a/docs/design/task-unit-semantics-obligations.md b/docs/design/task-unit-semantics-obligations.md index e49493ee..881903e7 100644 --- a/docs/design/task-unit-semantics-obligations.md +++ b/docs/design/task-unit-semantics-obligations.md @@ -3,7 +3,7 @@ **Authority:** living ledger for `docs/design/task-unit-semantics.md` — each row is one obligation from that design; progress is counted in rows moved to `proven`, not in edits made. -Counts at this revision, computed from the rows below rather than from memory: **A** 5 proven, 0 partly, 0 owed (5 rows); **B** 9 proven, 0 partly, 0 owed (9 rows, B7 nothing to fail); **C** 4 proven, 0 partly, 0 owed (4 rows); **D** 13 proven, 0 partly, 0 owed, 1 not applicable (14 rows; D10's subject, `runCycle`, was retired by [the retirement decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); **E** 9 proven as a seam with offline proofs and no source wired (9 rows); **G** 7 proven, 0 partly, 0 owed (7 rows). **F** complete except the fusion slice: its offline layer, the plan's single home, the arms' driver and its slot budget, the two task families and the real-model pilot have all landed and are recorded below, the last one in [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md) with its own sample-size caveat; **F4** (execution fusion) has landed its offline half (legality, accounting, the driver's session policy) **and** the live mechanism it needed - `createPiSessionRunner` in the extension, `piSessionWorker` + `--session-runner` in the driver - and the paid D arm is measured: fused `unitsPerSession: 2` against the `1` control, 3 reps each, one session of two units against two sessions of one, quality parity, ~1.9 s faster per run and token-neutral on that plan shape - a two-rep trend whose own spreads (1.2 s each) exceeded the 1.9 s gap, so it does not price the session-startup term the cost model had left `unmeasured`; on the four-unit fine plan the same cap experiment measures 3.2-4.2 s of wall clock per avoided session with the saving five times its within-cell spread, while its token columns settle nothing at three reps ([the D arm](../experiments/execution/ooo-arms-pilot-2026-09-18.md#d-arm-fusion-added-2026-09-19)). The A, B and C cells were then run on the plain path with the same fixture, worker, envelope and parent check (coarse one rep: 9 726 ms / 9 018 tokens; fine at one slot two reps: 21 885 and 23 273 ms, median 22 579; the same spec at two slots two reps: 18 565 and 19 336 ms, median 18 950, and cheaper in money than the one-slot pair), which also made the chain surface's own price visible: declaring fusion at a bound of one costs 3 607 ms and 15 412 tokens more than the plain path for the same plan ([the A-D cells](../experiments/execution/archive/ooo-arms-2026-09-19/README.md)). Their reports record `inputTokens`/`outputTokens`/`cacheRead`/`cost` per unit and the commit they ran from, so a comparison can name its instrument instead of arguing about it. **F5** (speculation lifecycle) landed offline in the same shared module: the design's `assumptions=[{predicateId, version, expected}]` as a declaration, the first experiment's bounds (exactly one pending fact, no speculative successor, no irreversible operation) refused by name, and the three outcomes read from authoritative evidence - true publishes, false discards the candidate and closes its branch session, unknown waits, because a missing reading is not permission to publish - with 9 cases and 5 named mutants. The **E arm** itself (budgeted speculation against a real model) has since run twice, on 2026-09-19, through the instrument it needed - one that decides the guessed fact and re-runs the real path under a new ticket - with no gain claimed: [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md#e-arm-bounded-speculation-first-run-2026-09-19) and [the archive README](../experiments/execution/archive/ooo-arms-2026-09-19/README.md). This sentence is what this paragraph used to over-claim the other way: it read as though the arm programme were one run short, while the measurement phase had closed. The counts sentence above is what this line used to over-claim: it read "**F** complete" while the F4 row said otherwise. F1 built the advisory cost model the design orders before any paid call (`evals/ooo-execution/cost-model.ts`, derived plan graphs, self-checks, no quality term) and recorded its sweep in [the experiment record](../experiments/execution/ooo-cost-model-2026-09-17.md); the sweep says one slot buys nothing (so B is predicted worse than A on cost), a chain or a two-unit refinement loses at every granularity, and the host check queue is what caps fine granularity. F2a then made the plan a value with one home and made the round's log name the plan it was given, which the runner's byte-identical copy of the default made worth doing on its own; the decision that the arms get their own research-side driver, with the 42 couplings measured behind it, is [recorded here](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md). F2b then built that driver (`evals/ooo-execution/plan-driver.ts`, one slot at a time, with the ordered legal set taken from `BoardAdmission.candidates()` rather than re-derived), and **building it measured that the C arm has no mechanism**: a run can hold exactly one claim, so `slots: 4` runs with `slotsUsed: 1` and names the refusal instead of reporting a time. The operator's decision on that finding is [recorded here](../decisions/implemented/2026-09-18-declared-slot-budget.md) — a run declares its **slot budget** (`slots`, default 1), the ordered legal set is cut to it after ordering, and the status query reports the same cut — and it has landed (row `F2b-slot` below, with the measurement kept in "What F2b measured"). Before it, E landed: `src/integration/task-advisers.ts` is the port a source may speak to and the rules it cannot break, `selectableTasks` is the legal set the shared rules already decided (with `nextTask` as its head), and `BoardAdmission` takes optional advisers whose absence is the rule policy - nine obligations proven offline with nine registered mutants, no gate changed and no HA or MGR implementation wired. Before it, D14 gave the drivers a daemon to be clients of and made that client boundary pessimistic (a bounded, named call, and no client calling the endpoint it serves). +Counts at this revision, computed from the rows below rather than from memory: **A** 6 proven, 0 partly, 0 owed (6 rows); **B** 9 proven, 0 partly, 0 owed (9 rows, B7 nothing to fail); **C** 4 proven, 0 partly, 0 owed (4 rows); **D** 13 proven, 0 partly, 0 owed, 1 not applicable (14 rows; D10's subject, `runCycle`, was retired by [the retirement decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); **E** 9 proven as a seam with offline proofs and no source wired (9 rows); **G** 7 proven, 0 partly, 0 owed (7 rows). **F** complete except the fusion slice: its offline layer, the plan's single home, the arms' driver and its slot budget, the two task families and the real-model pilot have all landed and are recorded below, the last one in [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md) with its own sample-size caveat; **F4** (execution fusion) has landed its offline half (legality, accounting, the driver's session policy) **and** the live mechanism it needed - `createPiSessionRunner` in the extension, `piSessionWorker` + `--session-runner` in the driver - and the paid D arm is measured: fused `unitsPerSession: 2` against the `1` control, 3 reps each, one session of two units against two sessions of one, quality parity, ~1.9 s faster per run and token-neutral on that plan shape - a two-rep trend whose own spreads (1.2 s each) exceeded the 1.9 s gap, so it does not price the session-startup term the cost model had left `unmeasured`; on the four-unit fine plan the same cap experiment measures 3.2-4.2 s of wall clock per avoided session with the saving five times its within-cell spread, while its token columns settle nothing at three reps ([the D arm](../experiments/execution/ooo-arms-pilot-2026-09-18.md#d-arm-fusion-added-2026-09-19)). The A, B and C cells were then run on the plain path with the same fixture, worker, envelope and parent check (coarse one rep: 9 726 ms / 9 018 tokens; fine at one slot two reps: 21 885 and 23 273 ms, median 22 579; the same spec at two slots two reps: 18 565 and 19 336 ms, median 18 950, and cheaper in money than the one-slot pair), which also made the chain surface's own price visible: declaring fusion at a bound of one costs 3 607 ms and 15 412 tokens more than the plain path for the same plan ([the A-D cells](../experiments/execution/archive/ooo-arms-2026-09-19/README.md)). Their reports record `inputTokens`/`outputTokens`/`cacheRead`/`cost` per unit and the commit they ran from, so a comparison can name its instrument instead of arguing about it. **F5** (speculation lifecycle) landed offline in the same shared module: the design's `assumptions=[{predicateId, version, expected}]` as a declaration, the first experiment's bounds (exactly one pending fact, no speculative successor, no irreversible operation) refused by name, and the three outcomes read from authoritative evidence - true publishes, false discards the candidate and closes its branch session, unknown waits, because a missing reading is not permission to publish - with 9 cases and 5 named mutants. The **E arm** itself (budgeted speculation against a real model) has since run twice, on 2026-09-19, through the instrument it needed - one that decides the guessed fact and re-runs the real path under a new ticket - with no gain claimed: [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md#e-arm-bounded-speculation-first-run-2026-09-19) and [the archive README](../experiments/execution/archive/ooo-arms-2026-09-19/README.md). This sentence is what this paragraph used to over-claim the other way: it read as though the arm programme were one run short, while the measurement phase had closed. The counts sentence above is what this line used to over-claim: it read "**F** complete" while the F4 row said otherwise. F1 built the advisory cost model the design orders before any paid call (`evals/ooo-execution/cost-model.ts`, derived plan graphs, self-checks, no quality term) and recorded its sweep in [the experiment record](../experiments/execution/ooo-cost-model-2026-09-17.md); the sweep says one slot buys nothing (so B is predicted worse than A on cost), a chain or a two-unit refinement loses at every granularity, and the host check queue is what caps fine granularity. F2a then made the plan a value with one home and made the round's log name the plan it was given, which the runner's byte-identical copy of the default made worth doing on its own; the decision that the arms get their own research-side driver, with the 42 couplings measured behind it, is [recorded here](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md). F2b then built that driver (`evals/ooo-execution/plan-driver.ts`, one slot at a time, with the ordered legal set taken from `BoardAdmission.candidates()` rather than re-derived), and **building it measured that the C arm has no mechanism**: a run can hold exactly one claim, so `slots: 4` runs with `slotsUsed: 1` and names the refusal instead of reporting a time. The operator's decision on that finding is [recorded here](../decisions/implemented/2026-09-18-declared-slot-budget.md) — a run declares its **slot budget** (`slots`, default 1), the ordered legal set is cut to it after ordering, and the status query reports the same cut — and it has landed (row `F2b-slot` below, with the measurement kept in "What F2b measured"). Before it, E landed: `src/integration/task-advisers.ts` is the port a source may speak to and the rules it cannot break, `selectableTasks` is the legal set the shared rules already decided (with `nextTask` as its head), and `BoardAdmission` takes optional advisers whose absence is the rule policy - nine obligations proven offline with nine registered mutants, no gate changed and no HA or MGR implementation wired. Before it, D14 gave the drivers a daemon to be clients of and made that client boundary pessimistic (a bounded, named call, and no client calling the endpoint it serves). Verification commands, run in the worktree that holds this branch, with the values they returned at this revision (re-run them rather than trusting the numbers; the harness writes no log file). Readings taken at @@ -27,7 +27,7 @@ an earlier revision say so, because a reading belongs to the instrument that pro - `npm run mutation:teeth` (the whole register, every tooth derived) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 160 s (2026-09-24); `npm run mutation:anchors` -> `anchors: 110 of 110 resolve, over 22 targets`; `tests/tools/mutation-anchor.test.ts` -> 34 pass; `test:product` -> 1580 pass; `verify:static` exit 0; lint 0 findings; complexity gate ok. - `npm run mutation:teeth` (the whole register, after the operators the catalogues name were added and 18 more teeth converted) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 176 s (2026-09-24); `npm run mutation:anchors` -> `anchors: 110 of 110 resolve, over 22 targets`; `tests/tools/mutation-anchor.test.ts` -> 30 pass. - `npm run mutation:teeth` (the whole register, after 59 teeth were converted to derived selectors) -> `mutants: 110 of 110 caught by the named test`, **all 110 by the case their `expect` names**, 22 of 22 targets restored byte-identically, exit 0 in 171 s (2026-09-24). The conversion is not what makes them bite; it is what keeps them aimed: one wrong `within` was shipped in this pass and the sweep is what caught it, as a tooth caught by the suite rather than by the case that names its rule. -- `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were **orphans**: no row named `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` or `the-parent-check-ignores-its-own-verdict`, so their rules have a case but no row; the check each case makes is in the [record](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. +- `npm run mutation:teeth` (the whole register) -> `mutants: 110 of 110 caught by the named test`, 22 of 22 targets restored byte-identically, exit 0 in 162 s (2026-09-24, after **26 more teeth were retired** in favour of the checks that state their rules - every one of the 27 insertion-shaped teeth this pass examined except `every-task-is-frozen-at-position-zero`, whose named case is a register/freeze/adopt/read-back round-trip and therefore weaker than the rule). All 110 are caught by the case their `expect` names. Seven of the 26 were first read as **orphans**, because no row named the tooth: `round-publication-opens-its-own-transaction`, `round-never-releases-its-pin`, `ordered-mode-becomes-any-topological-order`, `the-loop-awaits-each-unit-instead-of-the-batch`, `a-unit-is-dispatched-twice-in-one-batch`, `a-unit-nothing-checks-is-still-a-unit` and `the-parent-check-ignores-its-own-verdict`. Each rule was then given a home rather than a new home invented for it: the publication rule is B3's (the case that replaced it is in B3's suite and states the same transition), the pin release is B4's other side, the ordered mode is the **A6** row added for it (the design sentence it implements had no row), and the four judge-loop rules are F2b-slot's and F2c's - the batch rule and the duplicated dispatch share one case, `a unit's check is outstanding while an independent unit's worker runs`, and the unit with no checks and the parent check are the two cases F2c now names. The check each case makes is in the [record](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md). The target `tools/agent-verify.ts` went with its single tooth. - `npm run mutation:teeth -- --targets=src/integration/ooo-board.ts,src/integration/ooo-execution.ts,src/integration/task-semantics-interleavings.ts,src/integration/task-coordinator.ts` -> `mutants: 70 of 70 caught by the named test`, restored byte-identically 5 of 5 (2026-09-24, the four targets the retirement touched) - `npm run mutation:teeth -- --targets=src/integration/ooo-execution.ts` -> `mutants: 24 of 24 caught by the named test`, restored byte-identically 2 of 2 (2026-09-24: the pilot file after conversion, 19 of its 24 teeth derived; 23 are caught by the case their `expect` names and one by the suite alone, the stale name `fusion-continues-from-an-unverified-answer`, which the conversion neither caused nor fixed) - `npm run lint` (now over `src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/`) -> 0 findings, exit 0; `npm run check` -> exit 0. `npm run agent:verify` on a path under `evals/` still fails on the evaluation route's TAP rule for the skipped LoCoMo bridge - by decision, the rule stays and the reason names the skip. **A caveat found with the LSP, and the state of it now**: `evals/**` is in no `tsconfig`, so `npm run check`/`check:tests` never type-check it and only the LSP sees those files. The diagnostics this pass recorded were repaired on 2026-09-23 (the two evidence drivers' flag narrowing, `live-continuation.ts`'s declared result shape, and `tests/integration/ooo-evidence-drivers.test.ts`'s `{ pid: 0 }` fallback, commit `e31fe776` plus the test fix beside it), so every file this arc touched reports none: the LSP reading for those - the two drivers, `live-continuation.ts`, the store board's test, and `ooo-evidence-drivers.test.ts` - is zero. `src/` is also clean, which `npm run check` (exit 0) covers. What the LSP still reports is 31 diagnostics in seven `evals/` files this arc does not own: `benchmarks/run.ts` 7, `longmemeval/run.ts` 13, `controller/run.ts` 3, `natural-maintenance/audit.ts` 3, `hierarchy-scale/run.ts` 2, `longmemeval/score.ts` 2, `omnimemeval/bridge.ts` 1 - a slice of its own, and the reading a reader should expect in the meantime is that number rather than zero. @@ -158,14 +158,16 @@ separate claims. | A4 | Field mapping catches budget-unit confusion, requires/deps confusion and silently dropped keys | proven | mutants `budget-inner-alias-is-dropped`, `over-maximum-budget-is-accepted`, `assumption-may-carry-a-dependency` | | A5 | Every legal event interleaving of at most four units is enumerated, and every publication at every prefix is checked for its obligations, its inputs and its source | proven | `tests/integration/ooo-publication-invariants.test.ts` (18 cases) over `src/integration/task-semantics-interleavings.ts`. The enumeration is a merge of per-unit scripts, capped at four units and six events, and the checks read the derived view rather than a second model: a dispatch must rest on closed, accepted inputs, a completion (the accepted set closed over the same predicate) must have its verdict bound to the bytes, its bytes present, no cancellation, and inputs that are present and accepted. Both halves are pinned: each checker condition is deleted by a mutant and caught by the case that names it - `the-completion-does-not-bind-the-verdict-to-the-bytes`, `the-input-is-not-required-to-be-accepted`, `the-input-may-be-cancelled`, `the-completion-ignores-a-cancelled-unit`, `the-dispatch-does-not-require-a-closed-input` - and the clean run over the design's script sets reports nothing. The merge's completeness rests on the enumeration's own count assertion rather than on a mutant: the case `the merge enumerates every legal order, not one of them` asserts the multinomial count, which a merge returning one order fails, and the tooth that showed it (`the-merge-enumerates-one-order`) was retired on 2026-09-24 ([decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)). The design's caveat stands and is quoted in the module: passing says this finite model satisfies the listed properties, not that any Agent program is correct | +| A6 | The publication order is the frozen plan's stable topological order, and the model's `ordered` mode is that order rather than any topological one | proven | the design sentence is "最初使用冻结计划的稳定拓扑顺序作为发布顺序;生成和验证可以越序". `tests/integration/task-semantics-model.test.ts`: the case `the ordered mode is the declared plan order and nothing else` asserts exactly one ordered plan, that its order is the declared plan order, that the mode reports one order, and that the baseline never gains order freedom while the out-of-order mode gains two - so a mode returning any topological order fails a count rather than a wording. The tooth `ordered-mode-becomes-any-topological-order` was retired on 2026-09-24, once that case stated the rule ([decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)) | + ## B. Persistence (design: "进入持久化接入时另须证明") | node | obligation | state | evidence | | ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | B1 | Two rounds with the same taskId do not collide | proven | `tests/integration/ooo-run-namespace.test.ts`: the case `two runs in one store do not collide, do not see each other, and cancel separately` asserts each half directly - a claim in one run leaves the other run's row for the same task id alone (read as a raw row, which is the assertion the namespace experiment record shows was needed), a cancellation is per run, and a reopen by name sees the right one. The tooth `claim-is-not-scoped-to-its-run` was retired on 2026-09-24, once that case stated the rule ([record](../experiments/execution/ooo-run-namespace-2026-09-13.md), [decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)) | | B2 | A retry does not deliver twice | proven | `tests/core/task-board-deliverable.test.ts`; mutants `stale-claim-may-deliver-again`, `deliverer-may-judge-its-own-work` | -| B3 | A transaction failure cannot write a verdict without its association | proven | `tests/integration/ooo-transition-atomicity.test.ts`; mutant `swallowed-failure-still-commits`; the two teeth that pinned the boundary itself (`a-method-opens-its-own-transaction`, `nested-write-transaction-is-allowed`) were retired on 2026-09-24, because `a write entry reached inside a transition without a port is refused, not nested` states the refusal and `the store runs its transaction boundary in exactly one place` counts one BEGIN, one COMMIT and one ROLLBACK in the store's source, all three inside `writeTransaction` | -| B4 | A retained board entry is not cleared by TTL | proven | `tests/core/task-board-retention.test.ts`; mutants `prune-ignores-retention`, `bounded-pin-never-expires` | +| B3 | A transaction failure cannot write a verdict without its association | proven | `tests/integration/ooo-transition-atomicity.test.ts`; mutant `swallowed-failure-still-commits`; the two teeth that pinned the boundary itself (`a-method-opens-its-own-transaction`, `nested-write-transaction-is-allowed`) were retired on 2026-09-24, because `a write entry reached inside a transition without a port is refused, not nested` states the refusal and `the store runs its transaction boundary in exactly one place` counts one BEGIN, one COMMIT and one ROLLBACK in the store's source, all three inside `writeTransaction`. The case `the round's own publication rolls back with the transition that made it` states the publication half itself: the publication takes the port the caller's transition issued, so a failing transition takes the publication with it, and the assertion says what the alternative would mean - a publication that opened its own transaction would have survived it. The tooth `round-publication-opens-its-own-transaction` was retired on 2026-09-24, once that case stated the rule. | +| B4 | A retained board entry is not cleared by TTL | proven | `tests/core/task-board-retention.test.ts`; mutants `prune-ignores-retention`, `bounded-pin-never-expires`. The other side of retention - what happens to the pins when the value they protected is cleared - is pinned by `evals/ooo-execution/patch-cycle.test.ts`'s case `cancelling a round releases the pins it held, so nothing it referenced leaks`: it asserts the round holds pins while it runs, then cancels and asserts the retention list is empty, because both kinds of pin (the handoff published for the successor and the entry whose verdict accepted the patch) go with the values they protected, and a leaked pin would keep an entry alive forever in a round nobody is running. The tooth `round-never-releases-its-pin` was retired on 2026-09-24, once that case stated the rule. | | B5 | The status query and dependency release call one predicate | proven | `tests/integration/ooo-acceptance-one-predicate.test.ts`: the case `acceptance has one home, and both readers reach it` counts the call sites itself - one definition of the predicate in `src/`, one `acceptedFact({` in each reader, zero `verdict ===` comparisons in `ooo-board.ts` - so a rename that stops the call and a second decision inside it each fail a count rather than a named mutant. The two teeth that showed both (`the-board-read-path-stops-calling-the-predicate`, `the-board-decides-acceptance-on-its-own`) were retired on 2026-09-24, once the counts stated the rule directly | | B6 | A generic board write cannot move a managed round's own state, cannot take the live claim it holds, and cannot move an entry a run has adopted outside that run's transition | proven | the state half is structural and asserted (`tests/integration/ooo-managed-fence.test.ts`: the round's owner, attempt, acceptance and terminal reason stay its own facts, and a generic claim on its entry changes none of them). The claim it holds is not structural: it was carried by `a-live-claim-can-be-taken-by-another-agent`, retired on 2026-09-24 - the store's CAS stops requiring the holder to be the claimant, and the same test's direct second reader then succeeds, so the case states the rule itself. The design's owed sentence - a generic write on a managed entry is applied _inside_ the coordinating transaction, and `judge`/`resolve` cannot go around the run's fence - is implemented and pinned by `tests/integration/ooo-managed-write.test.ts` (6 cases): a direct verb on an adopted entry is refused and the entry does not move, while an entry no run adopts takes the path it always did; a coordinated write lands the board transition and the run's fact in one transaction and a failure after the board write leaves neither; a cancelled run takes no further lifecycle writes; a write for the wrong run, or for an entry no run adopted, is refused before anything is written; and the daemon's own `claim` verb routes an adopted entry through the transition, writing the fact with its own store. Mutants: `a-managed-entry-ignores-the-coordinated-scope` (the store's guard), `the-daemon-verb-skips-the-coordinated-path` (the routing), `a-coordinated-write-skips-its-run-fact`, `a-cancelled-run-still-accepts-writes`, `a-coordinated-write-skips-the-binding-recheck`. What is still the design's step 3 rather than this row: nothing adopts entries into a run yet, so the research drivers' direct writes are not refused today - they will be, and are meant to be, once the runner and thin adapters work through the coordinator | | B7 | JSONL export failure does not change the terminal state | not applicable | no JSONL export exists on this branch to fail | @@ -272,8 +274,8 @@ and the paid stage then ran on a family held out of it. | F1: advisory offline cost model | landed | `evals/ooo-execution/cost-model.ts` (`--sweep`, derived graphs, self-checks) + `cost-model.test.ts` (12 cases) + [the sweep record](../experiments/execution/ooo-cost-model-2026-09-17.md) | | F2a: the plan has one home, and the round's log names it | landed | the plan is one value with one home: the spec the driver runs (`PlanDriverSpec.plan`) and the run manifest the store freezes (D12). F2a's round-side carriers (`DEFAULT_ROUND_PLAN`, `CycleOptions.plan`, `openRoundStore(path, plan)`, `round-plan.test.ts` with its 6 cases) were retired with the round on 2026-09-18 ([decision](../decisions/implemented/2026-09-18-retire-the-round-instrument.md)); [the arms' driver decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) is what made the driver the home in the first place | | F2b: a research-side driver for arbitrary legal plans | landed, one arm | `evals/ooo-execution/plan-driver.ts` (`runPlan`, `comparePlanSlots`, `verifyParent` path, refusal-naming CLI) + `plan-driver.test.ts` (15 cases) + 12 named mutants (`tools/mutation-teeth.ts`, target `evals/ooo-execution/plan-driver.ts`) + `BoardAdmission.candidates()` (the ordered legal set; `next()` is its head, with its own mutant) + [the decision](../decisions/implemented/2026-09-17-arms-get-their-own-driver.md) | -| F2c: pick the parent task from the sweep's turning point | landed | Two families of the shape the design's parent family needs - a frozen interface, three independent builders, one summary that depends on all three - as a spec pair each: `evals/ooo-execution/fixtures/report/` (the instrument's own) and `fixtures/pipeline/` (**held out**: written after the driver, and not the family anything was tuned against). The coarse spec is one unit over the whole task checked by every frozen test; the fine spec is four units with per-unit checks and no `join`, so the parent check composes all four. Sibling units are code-independent by construction (the summary takes the derived values as parameters), which is what lets a unit be verified before its siblings exist. Offline, with the instrument's own answers and a wrong one: `evals/ooo-execution/families.test.ts` (8 cases over both families) - both plans accept, a wrong answer is rejected by its own check and takes the composition with it, and a unit that declares no checks and has no file-wide list is refused by name. Driver support this needed: per-unit `checks` with the file-wide list as a fallback, a `canned` worker (the instrument's answer, so the family's acceptance is shown before any model is paid), and the parent composition fix recorded in the pilot's experiment record | -| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-driver-awaits-each-unit-instead-of-the-batch`. `the-driver-declares-one-slot-whatever-the-spec-says` was retired on 2026-09-24, because `a declared slot count is reached, and the claims overlap in time` asserts the requested count, the used count and the overlap. The budget's licence and publication rules are pinned by `evals/ooo-execution/board-slots.test.ts`'s own case instead: `a declared budget holds two claims at once, and the store is why each handoff is directed` refuses a budget above one with no target by that message, asserts one handoff per startable task with `serialState` null for each, keeps every startable handoff across a republish by identity, and claims the non-head first - so `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected` and `a-multi-slot-handoff-is-published-un-directed` were retired on 2026-09-24 ([decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)). [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | +| F2c: pick the parent task from the sweep's turning point | landed | Two families of the shape the design's parent family needs - a frozen interface, three independent builders, one summary that depends on all three - as a spec pair each: `evals/ooo-execution/fixtures/report/` (the instrument's own) and `fixtures/pipeline/` (**held out**: written after the driver, and not the family anything was tuned against). The coarse spec is one unit over the whole task checked by every frozen test; the fine spec is four units with per-unit checks and no `join`, so the parent check composes all four. Sibling units are code-independent by construction (the summary takes the derived values as parameters), which is what lets a unit be verified before its siblings exist. Offline, with the instrument's own answers and a wrong one: `evals/ooo-execution/families.test.ts` (8 cases over both families) - both plans accept, a wrong answer is rejected by its own check and takes the composition with it (the case `the parent check is the composed acceptance, and a failing check is reported as such` - the tooth `the-parent-check-ignores-its-own-verdict` was retired on 2026-09-24 once that case stated the rule), and a unit that declares no checks and has no file-wide list is refused by name rather than accepted on nothing (the case `a unit nothing checks is refused rather than accepted on nothing`, which retired the tooth `a-unit-nothing-checks-is-still-a-unit` the same day). Driver support this needed: per-unit `checks` with the file-wide list as a fallback, a `canned` worker (the instrument's answer, so the family's acceptance is shown before any model is paid), and the parent composition fix recorded in the pilot's experiment record | +| F2b-slot: the C arm's mechanism (a run declares its slot budget) | landed | Rules: `selectableTasks(plan, slots)` / `startableTasks(plan, slots)` / `nextTask(plan, slots)` / `remainingSlots(plan, slots)` / `deriveStatus(units, facts, slots)` in `src/integration/ooo-execution.ts` + `src/integration/task-semantics.ts`; the ordered legal set is cut to `slots - claimed` **after** ordering, and the cut is what a claim licence may name. Admission: `BoardAdmissionOptions.slots` (default 1) + `handoffTarget` (required above 1, because the store queues a second un-directed actionable), `publishReady` offers every startable task one directed handoff and keeps it across a republish, and `claimableRow` checks `startable()`. Driver: `plan-driver.ts` declares the spec's count and names each claimant with one function. Cases: `board-slots.test.ts` (3), `narrow-dispatch.test.ts` (6, two new), `plan-driver.test.ts` (8, one replaced by the overlap case and one added by the retirement pass - the added one is `a unit's check is outstanding while an independent unit's worker runs`, which measures the windows a worker and a check occupy rather than the two durations, and it is this row's own statement of the batch rule: the tooth `the-loop-awaits-each-unit-instead-of-the-batch` awaited the batch one unit at a time, and `a-unit-is-dispatched-twice-in-one-batch` dispatched one twice, both pinned to that same case and both retired on 2026-09-24 once it stated the rule), `tests/integration/task-semantics.test.ts`. Mutants: `a-live-claim-does-not-block-selection` (re-anchored), `a-claimed-task-stays-on-offer`, `the-budget-is-not-cut-from-the-startable-set`, `half-a-slot-is-a-smaller-budget`, `the-status-query-ignores-the-declared-budget`, `the-driver-awaits-each-unit-instead-of-the-batch`. `the-driver-declares-one-slot-whatever-the-spec-says` was retired on 2026-09-24, because `a declared slot count is reached, and the claims overlap in time` asserts the requested count, the used count and the overlap. The budget's licence and publication rules are pinned by `evals/ooo-execution/board-slots.test.ts`'s own case instead: `a declared budget holds two claims at once, and the store is why each handoff is directed` refuses a budget above one with no target by that message, asserts one handoff per startable task with `serialState` null for each, keeps every startable handoff across a republish by identity, and claims the non-head first - so `the-licence-is-the-head-whatever-the-budget`, `a-second-slot-is-declared-without-a-target`, `only-the-heads-handoff-is-published`, `a-startable-handoff-is-retired-as-unselected` and `a-multi-slot-handoff-is-published-un-directed` were retired on 2026-09-24 ([decision](../decisions/implemented/2026-09-24-mutants-are-derived-not-anchored.md)). [Decision](../decisions/implemented/2026-09-18-declared-slot-budget.md) | | F3: real-model pilot (A 3 / B 3 / C 2, current pi model, directional only) | landed | `evals/ooo-execution/pilot.ts` (`--live` required, `--report` to re-aggregate recorded runs with no model call, refuses a merge of two instruments, seeded arm order) + [the pilot record](../experiments/execution/ooo-arms-pilot-2026-09-18.md). Fixed: `deepseek/deepseek-v4-flash`, envelope limits `turns 6 / reads 3 / 120 s` for every arm, A 3 / B 3 / C 2, arms drawn from a seeded shuffle, the held-out `pipeline` family. Measured (8 runs, 23 model calls, every run complete, all 8 parent checks accepted): per run A 8.6 s / 11.2 k tokens, B 21.5 s / 31.5 k tokens, C 16.2 s / 29.9 k tokens, host 1.6 s / 3.4 s / 3.6 s. All three of F1's expectations held: B is 2.5× A in wall time and 2.8× in tokens (one slot buys nothing), C recovers part of it (0.76× B) and not the 2× a pure model-call overlap would give, and the host cost grows with candidates rather than slots. Wasted cost 0, human intervention 0. **Directional only**: n = 8, one model, one held-out family, and no quality difference was available to measure - every arm accepted everything | | F4: execution fusion - legality, accounting and the driver policy | landed, live arm measured | Rules: `sharedSessionLegal`/`fusionSuccessors`/`fusionCandidates` in `src/integration/ooo-execution.ts` (the design's five conditions, one line each, composed with the board's candidate answer) + `tests/integration/ooo-fusion.test.ts` (12 cases) + 8 named mutants (target `src/integration/ooo-execution.ts`). Accounting: `fusionAccounting`/`fusionVerdict` + a fusion block in `cost-model.ts --sweep` - two lines kept apart, `unmeasured` until a run prices the session startup + `cost-model.test.ts` (12 cases) + 4 named mutants. Policy: `PlanDriverSpec.fusion` + `PlanRun.sessions` in `evals/ooo-execution/plan-driver.ts`, with the board's candidate set as the authority on staleness/cancellation/delivery/waits + `plan-driver.test.ts` (15 cases) + 11 named mutants. Scoped sweeps on this revision: `src/integration/ooo-execution.ts` 17 of 17 caught, `evals/ooo-execution/cost-model.ts` 4 of 4, `src/core/store/clock.ts` 4 of 4, `evals/ooo-execution/plan-driver.ts` 11 of 11, each restored byte-identically. A driver mutant that only restated `sharedSessionLegal`'s own rule was deleted rather than kept: the suite could not distinguish it from the shared predicate, which is what one home for that rule means. Both functions this slice pushed above the complexity limit (`sharedSessionLegal` 18, `runOneUnit` 16) were brought under it by extracting helpers, not by raising the threshold; `npm run lint` and `npm run check` are clean on this revision. **The live half landed.** `createPiSessionRunner` holds one Pi session and its tool surface across units (each unit re-points one mutable `UnitState` box; `patchSessionInput` is the one place a unit's prompt, snapshot and bounds are built, and `executePiPatch`/`executePiSnapshot` are thin callers of it), `PiRun` separates a unit's own `tokens`/`cacheRead`/`cacheWrite` from the session's `sessionTokens`, and `piSessionWorker` + `--session-runner` hold one runner per driver session. First paid D arm (2026-09-19, `deepseek-v4-flash`, 2 units, `--slots 1`, `turns: 6`, 3 reps per bound, only `fusion.unitsPerSession` differing): fused ran one session of two units and the control two sessions of one, quality parity in all six runs (every unit accepted), median 22 533 against 22 498 tokens and 11 048 against 12 948 ms - so no token saving yet (~1.9 s per run, ~15 % of the unfused wall, which prices the session-startup term at ~1.9 s instead of leaving it assumed) and per unit the second one cost ~8 % less while the first cost more: a chain's tool surface is the union of its units' capabilities because a session's surface is fixed at creation, so a unit can spend a turn on a tool that refuses by name. Also fixed here: `specFrom` had silently dropped a spec file's `fusion` block, so a spec asking for fusion ran as the control arm. | | F5: speculation lifecycle - one declared fact, three outcomes | landed (offline); the paid E arm ran once, no gain claimed | `SpeculationAssumption` / `ResolvedPredicate` / `SpeculationCandidate` / `isBoundedSpeculation` / `speculationOutcome` in `src/integration/ooo-execution.ts`, beside the fusion conditions: the assumption is a declaration the summary binds to (it never discovers for itself that the guess was false), the first experiment's bounds are a predicate (exactly one pending fact, nothing prepared from the guess - a speculative successor or an irreversible operation each refuse it by name), and the outcome has three states rather than two - **true** publishes, **false** discards the candidate and returns `sessionReusable: false`, which is what makes "失效会话不能复用到真实路径" a rule the caller must honour instead of a note, and **unknown** (no reading, an unattested reading, or evidence about another version) waits without publishing. Asking for the outcome of a candidate that is not the bounded shape throws rather than folding a fourth state into the three. Cases: `tests/integration/ooo-speculation.test.ts` (9), the last of which joins this half to fusion's condition 5 - an invalidated branch is not a legal predecessor for the real path. Mutants: 4 (`speculation-guesses-several-facts-at-once`, `an-unattested-reading-counts-as-evidence`, `evidence-about-another-version-is-the-same-fact`, `a-contradicted-guess-keeps-its-session`); the target's sweep is 22 of 22 caught. The E arm's instrument now exists (`evals/ooo-execution/speculation-pilot.ts`, registered) and ran once (2026-09-19, 6 paid units, ~43 k tokens): it decides the guessed fact, prepares the candidate, applies `speculationOutcome`, and verifies a published candidate with the unit's own frozen check. **No result is claimed**: the quality term was false in all four verified candidates, so by the design's own rule the latency and cost shape may not be reported as a gain. The search behind those failures is now closed and its first reading was wrong: `artifactEnvelope` builds two legitimate shapes (a patch, and a conclusion with `kind, conclusion, summary, evidence, citations`), and the instrument had fed every artifact to the patch reader. Eight of nine attempts answered with a conclusion, which this unit's check cannot pass and the board would refuse; the one patch attempt failed on a real mistake (`rows` for `lines`). The instrument now reads by kind, keeps every artifact, candidate tree and check output, and the run is archived. Measured outcome of the arm at this shape: the post-fact cost drops from ~6.2 s of work to 175 ms of verification when the fact holds, the false-fact case wastes 20 332 tokens, and the prepared candidate was publishable in 0 of 3 holding reps - so the cost is real, the gain is not, and the binding constraint is the candidate's admissibility |