Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 16 additions & 1 deletion agent-context.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,22 @@ routes:
- skills/repo-development/SKILL.md
tests: [tests/**]
verify:
blocking: [verify:static, test:product]
# Agent verification lists the static contract's atomic checks so a full
# plan runs each shared check once. CI still invokes verify:static.
blocking:
- build
- package:check
- check
- check:lock
- lint
- format:check
- docs:check
- agent:context:check
- complexity:gate
- verify:packages
- glossary:check
- rtm:check
- test:product
advisory: [verify:research, verify:chaos]

- id: packaging
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# Require nonzero evidence from the product coverage run

[中文](2026-09-23-coverage-evidence-liveness.zh-CN.md)

**Status:** implemented
**Approved:** explicit

## Problem

A successful test process can leave c8 with no covered source in its configured
include set. That run produces a 0% report even though all tests passed, so the
coverage track needs an explicit evidence check.

## Decision

The existing `test:coverage` command uses c8's coverage check with a 0.01% line
floor. This is a liveness floor for the existing included source set, not a
quality target. It makes a passing full coverage track require at least some
measured covered lines. The test list and concurrency remain the same.

## Alternatives considered

- Treat test exit status alone as coverage evidence. Rejected because c8 can
produce a zero-coverage report after a passing out-of-scope test subset.
- Set a high total or per-file threshold now. Rejected because that is a
separate coverage-quality policy requiring a full baseline and ownership.
- Write a custom report validator. Rejected while c8's built-in check covers
the observed empty-evidence failure.

## Consequences

- A zero-coverage run fails instead of silently producing an apparently valid
coverage artifact.
- This floor does not certify that every expected file was exercised; test
inventory and a future coverage policy remain separate decisions.
- Remove this floor if c8's built-in check produces a false pass or false fail
for this configured source set, and replace it with an evidenced check.
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# 要求产品覆盖率运行产生非零证据

[English](2026-09-23-coverage-evidence-liveness.md)

**Status:** implemented
**Approved:** explicit

## 问题

测试进程成功时,c8 仍可能没有测到配置包含的任何源码。这样即使测试全部通过,
也会生成 0% 报告,因此覆盖率轨需要明确检查证据是否存在。

## 决策

现有 `test:coverage` 命令使用 c8 自带的覆盖率检查,要求行覆盖率至少 0.01%。
这是现有源码集合的活性下限,不是质量目标。正式覆盖率轨只有测到一些源码行
才算通过。测试清单与并发度保持原样。

## 考虑过的替代方案

- 仅凭测试退出码认定覆盖率证据有效:范围外测试即使通过,c8 也可能报告零覆盖率。
- 现在设置较高的总体或逐文件门槛:那是另一项覆盖率质量政策,需要完整基线与归属。
- 编写自定义报告校验器:c8 内建检查已覆盖这次观察到的空证据失效。

## 后果

- 零覆盖率运行会失败,不再留下看似有效的报告。
- 这个下限不能证明每个预期文件都被执行;测试清单和未来的覆盖率政策是独立问题。
- 若 c8 内建检查对此源码集合误通过或误失败,应移除此下限并换成有证据的检查。
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# Route-only verifier tests omit unrelated shared checks

[中文](2026-09-23-route-only-verifier-fixtures.zh-CN.md)

**Status:** implemented
**Approved:** auto
**Relates to:** [Bound agent verification as one run](2026-09-23-verification-whole-run-deadline.md)

## Problem

Four narrow verifier tests assert only route-test failures or their TAP evidence. Their fixture also ran seven successful shared npm checks before reaching the asserted failure, spending time on processes irrelevant to those assertions.

## Decision

Those fixture cases use the existing `verify.sharedChecks: none` declaration. They retain the real verifier CLI, Git change discovery, Node route test, receipt, and output assertions. A separate default narrow test continues to assert that all seven shared checks ran, and the full-mode test continues to assert the full declared blocking set.

The resource selection is local to these test fixtures. It does not change any repository route or the product verifier's default behavior.

## Alternatives considered

- Keep every shared check in each route-failure test. It repeats seven child processes without exercising a distinct shared-check assertion.
- Replace the CLI with an in-process mock. This would lose the route-test execution and evidence boundary that the tests protect.
- Drop shared checks from every narrow test. This would remove direct coverage of the default shared-check floor.

## Consequences

All 28 verifier tests passed before and after the change. The targeted file took 32.672 seconds before and 22.244 seconds after in one local pair. The four affected cases together fell from about 17.55 to 5.53 seconds; the whole-file difference remains subject to machine load. See the [resource trial](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial).

If any route-only case starts asserting shared-check execution, remove its opt-out. Revert the fixture change if the complete blocking verifier route fails or its default narrow representative stops proving shared checks.
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# 只断言路由结果的验证器测试省去无关共享检查

[English](2026-09-23-route-only-verifier-fixtures.md)

**Status:** implemented
**Approved:** auto
**Relates to:** [整轮有界的 Agent 验证](2026-09-23-verification-whole-run-deadline.zh-CN.md)

## 问题

四项 narrow 验证器测试只断言路由测试失败及其 TAP 证据,fixture 却在断言之前执行七个成功的共享 npm 检查,为无关性质启动了多个子进程。

## 决策

这些 fixture 使用已有的 `verify.sharedChecks: none` 声明,保留真实 verifier CLI、Git 变更发现、Node 路由测试、receipt 与输出断言。另一个默认 narrow 测试继续断言七项共享检查均已运行,full-mode 测试继续断言完整的声明阻塞集。

该资源选择只作用于测试 fixture,不改变仓库路由或产品验证器的默认行为。

## 考虑过的替代方案

- 每项路由失败测试都运行共享检查:重复启动七个子进程,却没有增加共享检查的独立断言。
- 用进程内 mock 替换 CLI:会失去路由测试执行和证据链边界。
- 所有 narrow 测试都省去共享检查:会失去对默认共享检查下限的直接覆盖。

## 后果

修改前后 28 项验证器测试均通过。一次本机对照中,目标文件由 32.672 秒降到 22.244 秒;四项受影响测试合计由约 17.55 秒降到 5.53 秒。整文件时间仍受机器负载影响。方法见[资源实验](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial)。

若某个路由测试开始断言共享检查执行,就撤销它的 opt-out。若完整阻塞验证失败,或默认 narrow 代表测试不再证明共享检查,则撤销这项 fixture 改动。
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# Use memory SQLite for single-connection store tests

[中文](2026-09-23-single-connection-store-tests-in-memory.zh-CN.md)

**Status:** implemented
**Approved:** auto
**Relates to:** [Tests do not need a filesystem](2026-09-20-tests-need-no-filesystem.md)

## Problem

`tests/core/store.test.ts` and `tests/core/store/maintenance.test.ts` created and removed a separate file database for tests whose assertions use only one connection. File setup and writes spent time on a persistence boundary those tests did not observe.

## Decision

The single-connection helper opens a fresh `NmgStore(":memory:")` for each test and closes it afterward. Tests that assert restart persistence, migration from an existing file, or an independent reader continue to create separate file databases. The test's Safety or Contract role and blocking product coverage do not change. Resource choice follows the property being asserted; it is independent of narrow/full verification scope.

The shared integration `testDatabase()` remains file-backed because the daemon and service use its path. CLI, Git, lease, and cross-process tests keep real workspaces and processes.

## Alternatives considered

- Keep every store test file-backed. This exercises disk I/O repeatedly but adds no new persistence assertion to tests that never reopen or share the database.
- Use one shared file database. This removes setup but lets records and schema state leak between tests and makes parallel runs order-dependent.
- Move every test to memory. This would erase the independent connection, restart, and migration evidence.

## Consequences

The 52 store tests passed with the original file helper in 12.153 seconds, and with the single-connection helper in memory in 9.363 and 8.706 seconds in two local targeted runs. The 55 maintenance tests passed in 2.494 seconds with a file store and 1.470 seconds with a memory store. A 40-run setup probe measured mean 0.258 ms for directory create/remove, 17.419 ms for memory store open/close, and 32.874 ms for file store open/close plus directory cleanup. These local timings are comparative evidence, not a worst-case guarantee; the [experiment](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial) records the method.

If a test begins asserting a file, WAL, restart, or cross-connection property, move it to a unique file-backed fixture. Revert this choice if the full blocking test route fails or if file-backed behavior is left without direct coverage.
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# 单连接 Store 测试使用内存 SQLite

[English](2026-09-23-single-connection-store-tests-in-memory.md)

**Status:** implemented
**Approved:** auto
**Relates to:** [测试不需要文件系统](2026-09-20-tests-need-no-filesystem.zh-CN.md)

## 问题

`tests/core/store.test.ts` 与 `tests/core/store/maintenance.test.ts` 原先为只通过一个连接断言行为的测试创建和删除独立文件数据库。这些测试没有观察文件持久化边界,却承担了文件数据库的成本。

## 决策

单连接 helper 为每项测试打开独立的 `NmgStore(":memory:")`,结束后关闭。验证重启后持久化、已有文件的迁移或独立连接读取的测试继续使用各自的文件数据库。测试的 Safety/Contract 类型与产品阻塞覆盖不变。资源由断言需要观察的性质决定,与 narrow/full 范围无关。

集成用的 `testDatabase()` 仍使用文件路径,以供 daemon 和 service 共享。CLI、Git、lease 与跨进程测试继续使用真实工作区和进程。

## 考虑过的替代方案

- 所有 store 测试都保留文件数据库:重复运行磁盘 I/O,却不会让从不重开的测试增加持久化断言。
- 共用一个文件数据库:减少准备成本,但测试间记录和 schema 状态会串扰,并行结果依赖顺序。
- 所有测试都改用内存:会失去独立连接、重启与迁移证据。

## 后果

52 项 store 测试使用原文件 helper 时通过,用时 12.153 秒;单连接 helper 改为内存后两次均通过,用时 9.363 和 8.706 秒。55 项 maintenance 测试使用文件库时通过,用时 2.494 秒;改为内存库后通过,用时 1.470 秒。40 次准备成本探针测得建删目录平均 0.258 毫秒、内存 store 开关平均 17.419 毫秒、文件 store 开关并清理目录平均 32.874 毫秒。这些是本机对照数据,不是最坏时间保证;方法见[实验记录](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial)。

若测试开始断言文件、WAL、重启或跨连接性质,应改为独立文件 fixture。若完整阻塞测试失败,或文件持久化行为失去直接测试覆盖,则撤销本决策。
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Bound agent verification as one run

[中文](2026-09-23-verification-whole-run-deadline.zh-CN.md)

**Status:** implemented
**Approved:** explicit

## Problem

`agent:verify --timeout-ms` previously gave every serial check a fresh timeout. A
150-second setting could therefore take many times 150 seconds and did not meet
the requested time to a result.

## Decision

Use 150 seconds as the default budget for the entire `npm run agent:verify`
invocation. `--timeout-ms` adjusts that total budget. An outer watchdog returns
an incomplete failure and writes failure evidence if setup or other synchronous
work exceeds it. Direct, RCP full, and RCP narrow checks receive only the
remaining time. Once the budget is exhausted, pending checks are recorded as
failed without being started, and the final verdict fails. A partial set of
passed checks cannot be promoted to a passing verification.

The Agent verifier's `ci-and-tests` route declares the atomic checks of the
`verify:static` package contract and `test:product`. The full plan deduplicates
identical check names across routes, executes each atomic check once, and keeps
its own result and route attribution. A test checks that this route stays equal
to the static package contract. CI retains `verify:static` as its reproducible
named entry point.

The non-RCP full Agent plan runs independent static checks in bounded groups of
three after the build and package barriers. The approved group does not rewrite
root source or `dist`; subpackage building and complexity probes use separate
paths. `test:product` remains after those
groups. The plan keeps declaration-order results, per-check failures, and one
shared deadline. RCP and narrow verification retain their existing serial
execution.

## Alternatives considered

- Keep a timeout for each command. Rejected because serial checks can exceed
the requested result time by a large multiple.
- Stop without recording pending checks. Rejected because a missing check is easy
to mistake for an irrelevant check.
- Run `verify:static` and its constituent checks again as independent route
checks. Rejected because repeated work spends the same run budget without
adding independent coverage. Skipping the independent results instead would
lose per-check attribution.
- Launch every blocking check together. Rejected because build and packaging
rewrite generated files that other checks read, and an unrestricted process
burst competes with the product suite.

## Consequences

- A timeout is an incomplete verification requiring remediation and a new run.
- The outer process bounds the official CLI's result time. It kills the direct
verifier process on timeout, but cannot guarantee that npm's descendant
processes have exited; a later run must inspect a live mutation lock before
trusting the tree.
- The Agent plan has more individually attributed check results, but does not
repeat static checks already shared by other routes. CI's contract is unchanged.
- Bounded concurrency shortens the static phase on the measured worktree, but
individual CPU-heavy checks take longer under contention. The 150-second
deadline still fails closed if a run cannot finish.
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
# 将 Agent 验证限制为一整轮预算

[English](2026-09-23-verification-whole-run-deadline.md)

**Status:** implemented
**Approved:** explicit

## 问题

过去 `agent:verify --timeout-ms` 给串行的每项检查重新计时。设为 150 秒时,
整轮仍可能花费数倍于 150 秒,无法满足出结果的时限。

## 决策

官方 `npm run agent:verify` 入口默认使用 150 秒整轮预算,`--timeout-ms` 调整
的是这个总预算。外层看门进程在同步准备等阶段越界时返回 `incomplete` 失败并写入
失败证据。普通、RCP full 和 RCP narrow 检查都只获得剩余时间。预算耗尽后,不再
启动待运行的检查,而是逐项记为失败;整轮结论也失败。局部通过不能冒充整轮通过。

Agent 验证器的 `ci-and-tests` route 列出 `verify:static` package contract 的原子检查和
`test:product`。完整计划按同名检查跨 route 去重,每项原子检查只执行一次,同时保留独立
结果及 route 归因。测试约束该 route 与静态 package contract 保持一致。CI 继续以
`verify:static` 作为可复现的命名入口。

非 RCP 的 Agent full 计划在构建与打包屏障之后,以最多三个并发执行相互独立的静态检查;
这些检查不改写根源码或 `dist`,子包构建和复杂度探针使用独立路径。
`test:product` 在这些检查完成后运行。结果仍按声明顺序排列,保留逐项失败归因,且共用
同一整轮预算。RCP 与 narrow 验证保持原有串行执行。

## 考虑过的替代方案

- 保留逐命令超时:串行检查会使总时长远超用户要求,因此拒绝。
- 停止运行但不记录待检查项:缺失的检查容易被误认为不适用,因此拒绝。
- 同时运行 `verify:static` 和它包含的独立 route 检查:重复工作消耗同一整轮预算,
却不增加独立覆盖;直接跳过独立结果又会失去逐项归因,因此拒绝。
- 将所有阻塞检查同时启动:构建与打包会改写其他检查读取的生成文件;无限制的进程并发
还会与产品测试争用资源,因此拒绝。

## 后果

- 超时属于未完成验证,需要整改并重新运行。
- 外层进程限制官方 CLI 的出结果时间,超时时终止直接的验证器进程;但不能保证
npm 的后代进程均已退出。重新运行前仍须检查活动的 mutation lock。
- Agent 计划保留更多逐项归因的结果,但不重复执行其他 route 共享的静态检查。
CI contract 不变。
- 有界并发在实测工作树上缩短静态阶段,但 CPU 密集检查会因资源争用而各自变慢。
未能按时完成时,150 秒整轮预算仍失败关闭。
Loading
Loading