diff --git a/agent-context.yaml b/agent-context.yaml index 43182b95..540b81b9 100644 --- a/agent-context.yaml +++ b/agent-context.yaml @@ -137,7 +137,22 @@ routes: - skills/repo-development/SKILL.md tests: [tests/**] verify: - blocking: [verify:static, test:product] + # Agent verification lists the static contract's atomic checks so a full + # plan runs each shared check once. CI still invokes verify:static. + blocking: + - build + - package:check + - check + - check:lock + - lint + - format:check + - docs:check + - agent:context:check + - complexity:gate + - verify:packages + - glossary:check + - rtm:check + - test:product advisory: [verify:research, verify:chaos] - id: packaging diff --git a/docs/decisions/implemented/2026-09-23-coverage-evidence-liveness.md b/docs/decisions/implemented/2026-09-23-coverage-evidence-liveness.md new file mode 100644 index 00000000..b91293cd --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-coverage-evidence-liveness.md @@ -0,0 +1,37 @@ +# Require nonzero evidence from the product coverage run + +[中文](2026-09-23-coverage-evidence-liveness.zh-CN.md) + +**Status:** implemented +**Approved:** explicit + +## Problem + +A successful test process can leave c8 with no covered source in its configured +include set. That run produces a 0% report even though all tests passed, so the +coverage track needs an explicit evidence check. + +## Decision + +The existing `test:coverage` command uses c8's coverage check with a 0.01% line +floor. This is a liveness floor for the existing included source set, not a +quality target. It makes a passing full coverage track require at least some +measured covered lines. The test list and concurrency remain the same. + +## Alternatives considered + +- Treat test exit status alone as coverage evidence. Rejected because c8 can + produce a zero-coverage report after a passing out-of-scope test subset. +- Set a high total or per-file threshold now. Rejected because that is a + separate coverage-quality policy requiring a full baseline and ownership. +- Write a custom report validator. Rejected while c8's built-in check covers + the observed empty-evidence failure. + +## Consequences + +- A zero-coverage run fails instead of silently producing an apparently valid + coverage artifact. +- This floor does not certify that every expected file was exercised; test + inventory and a future coverage policy remain separate decisions. +- Remove this floor if c8's built-in check produces a false pass or false fail + for this configured source set, and replace it with an evidenced check. diff --git a/docs/decisions/implemented/2026-09-23-coverage-evidence-liveness.zh-CN.md b/docs/decisions/implemented/2026-09-23-coverage-evidence-liveness.zh-CN.md new file mode 100644 index 00000000..f1e268a7 --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-coverage-evidence-liveness.zh-CN.md @@ -0,0 +1,29 @@ +# 要求产品覆盖率运行产生非零证据 + +[English](2026-09-23-coverage-evidence-liveness.md) + +**Status:** implemented +**Approved:** explicit + +## 问题 + +测试进程成功时,c8 仍可能没有测到配置包含的任何源码。这样即使测试全部通过, +也会生成 0% 报告,因此覆盖率轨需要明确检查证据是否存在。 + +## 决策 + +现有 `test:coverage` 命令使用 c8 自带的覆盖率检查,要求行覆盖率至少 0.01%。 +这是现有源码集合的活性下限,不是质量目标。正式覆盖率轨只有测到一些源码行 +才算通过。测试清单与并发度保持原样。 + +## 考虑过的替代方案 + +- 仅凭测试退出码认定覆盖率证据有效:范围外测试即使通过,c8 也可能报告零覆盖率。 +- 现在设置较高的总体或逐文件门槛:那是另一项覆盖率质量政策,需要完整基线与归属。 +- 编写自定义报告校验器:c8 内建检查已覆盖这次观察到的空证据失效。 + +## 后果 + +- 零覆盖率运行会失败,不再留下看似有效的报告。 +- 这个下限不能证明每个预期文件都被执行;测试清单和未来的覆盖率政策是独立问题。 +- 若 c8 内建检查对此源码集合误通过或误失败,应移除此下限并换成有证据的检查。 diff --git a/docs/decisions/implemented/2026-09-23-route-only-verifier-fixtures.md b/docs/decisions/implemented/2026-09-23-route-only-verifier-fixtures.md new file mode 100644 index 00000000..a8d55d19 --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-route-only-verifier-fixtures.md @@ -0,0 +1,29 @@ +# Route-only verifier tests omit unrelated shared checks + +[中文](2026-09-23-route-only-verifier-fixtures.zh-CN.md) + +**Status:** implemented +**Approved:** auto +**Relates to:** [Bound agent verification as one run](2026-09-23-verification-whole-run-deadline.md) + +## Problem + +Four narrow verifier tests assert only route-test failures or their TAP evidence. Their fixture also ran seven successful shared npm checks before reaching the asserted failure, spending time on processes irrelevant to those assertions. + +## Decision + +Those fixture cases use the existing `verify.sharedChecks: none` declaration. They retain the real verifier CLI, Git change discovery, Node route test, receipt, and output assertions. A separate default narrow test continues to assert that all seven shared checks ran, and the full-mode test continues to assert the full declared blocking set. + +The resource selection is local to these test fixtures. It does not change any repository route or the product verifier's default behavior. + +## Alternatives considered + +- Keep every shared check in each route-failure test. It repeats seven child processes without exercising a distinct shared-check assertion. +- Replace the CLI with an in-process mock. This would lose the route-test execution and evidence boundary that the tests protect. +- Drop shared checks from every narrow test. This would remove direct coverage of the default shared-check floor. + +## Consequences + +All 28 verifier tests passed before and after the change. The targeted file took 32.672 seconds before and 22.244 seconds after in one local pair. The four affected cases together fell from about 17.55 to 5.53 seconds; the whole-file difference remains subject to machine load. See the [resource trial](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial). + +If any route-only case starts asserting shared-check execution, remove its opt-out. Revert the fixture change if the complete blocking verifier route fails or its default narrow representative stops proving shared checks. diff --git a/docs/decisions/implemented/2026-09-23-route-only-verifier-fixtures.zh-CN.md b/docs/decisions/implemented/2026-09-23-route-only-verifier-fixtures.zh-CN.md new file mode 100644 index 00000000..c3046734 --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-route-only-verifier-fixtures.zh-CN.md @@ -0,0 +1,29 @@ +# 只断言路由结果的验证器测试省去无关共享检查 + +[English](2026-09-23-route-only-verifier-fixtures.md) + +**Status:** implemented +**Approved:** auto +**Relates to:** [整轮有界的 Agent 验证](2026-09-23-verification-whole-run-deadline.zh-CN.md) + +## 问题 + +四项 narrow 验证器测试只断言路由测试失败及其 TAP 证据,fixture 却在断言之前执行七个成功的共享 npm 检查,为无关性质启动了多个子进程。 + +## 决策 + +这些 fixture 使用已有的 `verify.sharedChecks: none` 声明,保留真实 verifier CLI、Git 变更发现、Node 路由测试、receipt 与输出断言。另一个默认 narrow 测试继续断言七项共享检查均已运行,full-mode 测试继续断言完整的声明阻塞集。 + +该资源选择只作用于测试 fixture,不改变仓库路由或产品验证器的默认行为。 + +## 考虑过的替代方案 + +- 每项路由失败测试都运行共享检查:重复启动七个子进程,却没有增加共享检查的独立断言。 +- 用进程内 mock 替换 CLI:会失去路由测试执行和证据链边界。 +- 所有 narrow 测试都省去共享检查:会失去对默认共享检查下限的直接覆盖。 + +## 后果 + +修改前后 28 项验证器测试均通过。一次本机对照中,目标文件由 32.672 秒降到 22.244 秒;四项受影响测试合计由约 17.55 秒降到 5.53 秒。整文件时间仍受机器负载影响。方法见[资源实验](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial)。 + +若某个路由测试开始断言共享检查执行,就撤销它的 opt-out。若完整阻塞验证失败,或默认 narrow 代表测试不再证明共享检查,则撤销这项 fixture 改动。 diff --git a/docs/decisions/implemented/2026-09-23-single-connection-store-tests-in-memory.md b/docs/decisions/implemented/2026-09-23-single-connection-store-tests-in-memory.md new file mode 100644 index 00000000..7912642a --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-single-connection-store-tests-in-memory.md @@ -0,0 +1,29 @@ +# Use memory SQLite for single-connection store tests + +[中文](2026-09-23-single-connection-store-tests-in-memory.zh-CN.md) + +**Status:** implemented +**Approved:** auto +**Relates to:** [Tests do not need a filesystem](2026-09-20-tests-need-no-filesystem.md) + +## Problem + +`tests/core/store.test.ts` and `tests/core/store/maintenance.test.ts` created and removed a separate file database for tests whose assertions use only one connection. File setup and writes spent time on a persistence boundary those tests did not observe. + +## Decision + +The single-connection helper opens a fresh `NmgStore(":memory:")` for each test and closes it afterward. Tests that assert restart persistence, migration from an existing file, or an independent reader continue to create separate file databases. The test's Safety or Contract role and blocking product coverage do not change. Resource choice follows the property being asserted; it is independent of narrow/full verification scope. + +The shared integration `testDatabase()` remains file-backed because the daemon and service use its path. CLI, Git, lease, and cross-process tests keep real workspaces and processes. + +## Alternatives considered + +- Keep every store test file-backed. This exercises disk I/O repeatedly but adds no new persistence assertion to tests that never reopen or share the database. +- Use one shared file database. This removes setup but lets records and schema state leak between tests and makes parallel runs order-dependent. +- Move every test to memory. This would erase the independent connection, restart, and migration evidence. + +## Consequences + +The 52 store tests passed with the original file helper in 12.153 seconds, and with the single-connection helper in memory in 9.363 and 8.706 seconds in two local targeted runs. The 55 maintenance tests passed in 2.494 seconds with a file store and 1.470 seconds with a memory store. A 40-run setup probe measured mean 0.258 ms for directory create/remove, 17.419 ms for memory store open/close, and 32.874 ms for file store open/close plus directory cleanup. These local timings are comparative evidence, not a worst-case guarantee; the [experiment](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial) records the method. + +If a test begins asserting a file, WAL, restart, or cross-connection property, move it to a unique file-backed fixture. Revert this choice if the full blocking test route fails or if file-backed behavior is left without direct coverage. diff --git a/docs/decisions/implemented/2026-09-23-single-connection-store-tests-in-memory.zh-CN.md b/docs/decisions/implemented/2026-09-23-single-connection-store-tests-in-memory.zh-CN.md new file mode 100644 index 00000000..a871f774 --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-single-connection-store-tests-in-memory.zh-CN.md @@ -0,0 +1,29 @@ +# 单连接 Store 测试使用内存 SQLite + +[English](2026-09-23-single-connection-store-tests-in-memory.md) + +**Status:** implemented +**Approved:** auto +**Relates to:** [测试不需要文件系统](2026-09-20-tests-need-no-filesystem.zh-CN.md) + +## 问题 + +`tests/core/store.test.ts` 与 `tests/core/store/maintenance.test.ts` 原先为只通过一个连接断言行为的测试创建和删除独立文件数据库。这些测试没有观察文件持久化边界,却承担了文件数据库的成本。 + +## 决策 + +单连接 helper 为每项测试打开独立的 `NmgStore(":memory:")`,结束后关闭。验证重启后持久化、已有文件的迁移或独立连接读取的测试继续使用各自的文件数据库。测试的 Safety/Contract 类型与产品阻塞覆盖不变。资源由断言需要观察的性质决定,与 narrow/full 范围无关。 + +集成用的 `testDatabase()` 仍使用文件路径,以供 daemon 和 service 共享。CLI、Git、lease 与跨进程测试继续使用真实工作区和进程。 + +## 考虑过的替代方案 + +- 所有 store 测试都保留文件数据库:重复运行磁盘 I/O,却不会让从不重开的测试增加持久化断言。 +- 共用一个文件数据库:减少准备成本,但测试间记录和 schema 状态会串扰,并行结果依赖顺序。 +- 所有测试都改用内存:会失去独立连接、重启与迁移证据。 + +## 后果 + +52 项 store 测试使用原文件 helper 时通过,用时 12.153 秒;单连接 helper 改为内存后两次均通过,用时 9.363 和 8.706 秒。55 项 maintenance 测试使用文件库时通过,用时 2.494 秒;改为内存库后通过,用时 1.470 秒。40 次准备成本探针测得建删目录平均 0.258 毫秒、内存 store 开关平均 17.419 毫秒、文件 store 开关并清理目录平均 32.874 毫秒。这些是本机对照数据,不是最坏时间保证;方法见[实验记录](../../experiments/verification/product-file-partitions-2026-09-23.md#resource-boundary-trial)。 + +若测试开始断言文件、WAL、重启或跨连接性质,应改为独立文件 fixture。若完整阻塞测试失败,或文件持久化行为失去直接测试覆盖,则撤销本决策。 diff --git a/docs/decisions/implemented/2026-09-23-verification-whole-run-deadline.md b/docs/decisions/implemented/2026-09-23-verification-whole-run-deadline.md new file mode 100644 index 00000000..084ed700 --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-verification-whole-run-deadline.md @@ -0,0 +1,64 @@ +# Bound agent verification as one run + +[中文](2026-09-23-verification-whole-run-deadline.zh-CN.md) + +**Status:** implemented +**Approved:** explicit + +## Problem + +`agent:verify --timeout-ms` previously gave every serial check a fresh timeout. A +150-second setting could therefore take many times 150 seconds and did not meet +the requested time to a result. + +## Decision + +Use 150 seconds as the default budget for the entire `npm run agent:verify` +invocation. `--timeout-ms` adjusts that total budget. An outer watchdog returns +an incomplete failure and writes failure evidence if setup or other synchronous +work exceeds it. Direct, RCP full, and RCP narrow checks receive only the +remaining time. Once the budget is exhausted, pending checks are recorded as +failed without being started, and the final verdict fails. A partial set of +passed checks cannot be promoted to a passing verification. + +The Agent verifier's `ci-and-tests` route declares the atomic checks of the +`verify:static` package contract and `test:product`. The full plan deduplicates +identical check names across routes, executes each atomic check once, and keeps +its own result and route attribution. A test checks that this route stays equal +to the static package contract. CI retains `verify:static` as its reproducible +named entry point. + +The non-RCP full Agent plan runs independent static checks in bounded groups of +three after the build and package barriers. The approved group does not rewrite +root source or `dist`; subpackage building and complexity probes use separate +paths. `test:product` remains after those +groups. The plan keeps declaration-order results, per-check failures, and one +shared deadline. RCP and narrow verification retain their existing serial +execution. + +## Alternatives considered + +- Keep a timeout for each command. Rejected because serial checks can exceed + the requested result time by a large multiple. +- Stop without recording pending checks. Rejected because a missing check is easy + to mistake for an irrelevant check. +- Run `verify:static` and its constituent checks again as independent route + checks. Rejected because repeated work spends the same run budget without + adding independent coverage. Skipping the independent results instead would + lose per-check attribution. +- Launch every blocking check together. Rejected because build and packaging + rewrite generated files that other checks read, and an unrestricted process + burst competes with the product suite. + +## Consequences + +- A timeout is an incomplete verification requiring remediation and a new run. +- The outer process bounds the official CLI's result time. It kills the direct + verifier process on timeout, but cannot guarantee that npm's descendant + processes have exited; a later run must inspect a live mutation lock before + trusting the tree. +- The Agent plan has more individually attributed check results, but does not + repeat static checks already shared by other routes. CI's contract is unchanged. +- Bounded concurrency shortens the static phase on the measured worktree, but + individual CPU-heavy checks take longer under contention. The 150-second + deadline still fails closed if a run cannot finish. diff --git a/docs/decisions/implemented/2026-09-23-verification-whole-run-deadline.zh-CN.md b/docs/decisions/implemented/2026-09-23-verification-whole-run-deadline.zh-CN.md new file mode 100644 index 00000000..a85bf345 --- /dev/null +++ b/docs/decisions/implemented/2026-09-23-verification-whole-run-deadline.zh-CN.md @@ -0,0 +1,47 @@ +# 将 Agent 验证限制为一整轮预算 + +[English](2026-09-23-verification-whole-run-deadline.md) + +**Status:** implemented +**Approved:** explicit + +## 问题 + +过去 `agent:verify --timeout-ms` 给串行的每项检查重新计时。设为 150 秒时, +整轮仍可能花费数倍于 150 秒,无法满足出结果的时限。 + +## 决策 + +官方 `npm run agent:verify` 入口默认使用 150 秒整轮预算,`--timeout-ms` 调整 +的是这个总预算。外层看门进程在同步准备等阶段越界时返回 `incomplete` 失败并写入 +失败证据。普通、RCP full 和 RCP narrow 检查都只获得剩余时间。预算耗尽后,不再 +启动待运行的检查,而是逐项记为失败;整轮结论也失败。局部通过不能冒充整轮通过。 + +Agent 验证器的 `ci-and-tests` route 列出 `verify:static` package contract 的原子检查和 +`test:product`。完整计划按同名检查跨 route 去重,每项原子检查只执行一次,同时保留独立 +结果及 route 归因。测试约束该 route 与静态 package contract 保持一致。CI 继续以 +`verify:static` 作为可复现的命名入口。 + +非 RCP 的 Agent full 计划在构建与打包屏障之后,以最多三个并发执行相互独立的静态检查; +这些检查不改写根源码或 `dist`,子包构建和复杂度探针使用独立路径。 +`test:product` 在这些检查完成后运行。结果仍按声明顺序排列,保留逐项失败归因,且共用 +同一整轮预算。RCP 与 narrow 验证保持原有串行执行。 + +## 考虑过的替代方案 + +- 保留逐命令超时:串行检查会使总时长远超用户要求,因此拒绝。 +- 停止运行但不记录待检查项:缺失的检查容易被误认为不适用,因此拒绝。 +- 同时运行 `verify:static` 和它包含的独立 route 检查:重复工作消耗同一整轮预算, + 却不增加独立覆盖;直接跳过独立结果又会失去逐项归因,因此拒绝。 +- 将所有阻塞检查同时启动:构建与打包会改写其他检查读取的生成文件;无限制的进程并发 + 还会与产品测试争用资源,因此拒绝。 + +## 后果 + +- 超时属于未完成验证,需要整改并重新运行。 +- 外层进程限制官方 CLI 的出结果时间,超时时终止直接的验证器进程;但不能保证 + npm 的后代进程均已退出。重新运行前仍须检查活动的 mutation lock。 +- Agent 计划保留更多逐项归因的结果,但不重复执行其他 route 共享的静态检查。 + CI contract 不变。 +- 有界并发在实测工作树上缩短静态阶段,但 CPU 密集检查会因资源争用而各自变慢。 + 未能按时完成时,150 秒整轮预算仍失败关闭。 diff --git a/docs/decisions/implemented/2026-09-24-resource-matched-test-fixtures.md b/docs/decisions/implemented/2026-09-24-resource-matched-test-fixtures.md new file mode 100644 index 00000000..05c6527e --- /dev/null +++ b/docs/decisions/implemented/2026-09-24-resource-matched-test-fixtures.md @@ -0,0 +1,51 @@ +# Match test fixtures to the resource an assertion observes + +[中文](2026-09-24-resource-matched-test-fixtures.zh-CN.md) + +**Status:** implemented +**Approved:** auto + +## Problem + +The four resource needs in product tests were confused with four fixed runner +partitions. A fixed schedule was slower on the current worktree, but that result +does not answer whether each resource class has avoidable setup cost. + +## Decision + +Choose fixtures by the assertion's observable boundary: in-process data for pure +decisions, an independent in-memory SQLite store for one-connection behavior, a +unique file database or workspace for persistence and path behavior, and real +Git/CLI/HTTP/process resources for external behavior. This choice is orthogonal +to Safety/Contract/Guardrail and narrow/full verification membership. + +RCP orchestration tests that do not observe Git discovery, dirty scope, or commit +provenance inject a fixed `RepositoryProvider` observation and create no Git +repository. The tests that assert those boundaries keep the real provider and +repository. That real-Git fixture still runs `init`, `add`, and `commit`, but +supplies identity through `git -c` on the commit call instead of launching two +separate configuration processes for every fixture. + +## Alternatives considered + +- Treat the classes as fixed Node worker partitions. Rejected: the measured + partition and file-order candidates were slower than the existing route. +- Use a fixed observation for every RCP test. Rejected because scope, commit, + forge, and Git-failure assertions must exercise real repository behavior. +- Initialize Git for every orchestration test. Rejected because these tests + consume the provider result without asserting how Git produced it. +- Share one mutable database or Git repository between tests. Rejected because + test order and concurrent runs could then affect the result. + +## Consequences + +- Ten reconciliation tests retain real Git observation; eight use the fixed + observation. Each real-Git fixture also starts two fewer Git child processes. + The repeated file-level improvement is measured, but a whole-suite speedup + remains unproven. +- The fixture rule guides setup only; route coverage and the 150-second result + deadline remain owned by the verification contracts. +- Restore real Git for any fixed-observation test whose assertion comes to + depend on Git scope or provenance. Repeal the commit-identity shortcut if a + test needs repository-local Git identity configuration or a supported Git + environment behaves differently. diff --git a/docs/decisions/implemented/2026-09-24-resource-matched-test-fixtures.zh-CN.md b/docs/decisions/implemented/2026-09-24-resource-matched-test-fixtures.zh-CN.md new file mode 100644 index 00000000..a790c334 --- /dev/null +++ b/docs/decisions/implemented/2026-09-24-resource-matched-test-fixtures.zh-CN.md @@ -0,0 +1,42 @@ +# 按断言观察的资源选择测试 fixture + +[English](2026-09-24-resource-matched-test-fixtures.md) + +**Status:** implemented +**Approved:** auto + +## 问题 + +产品测试的四类资源需求被误作四个固定的运行器分区。固定调度在当前工作树更慢, +但这个结果不能回答每类资源的准备工作是否有可省的成本。 + +## 决策 + +按断言能观察到的边界选择 fixture:纯决策使用进程内数据;单连接行为使用各项 +独立的内存 SQLite;持久化和路径行为使用唯一文件数据库或工作区;外部行为 +使用真实 Git、CLI、HTTP 或进程资源。这一选择与 Safety/Contract/Guardrail +以及 narrow/full 验证归属正交。 + +不观察 Git 发现、dirty 范围或提交来源的 RCP 编排测试注入固定的 +`RepositoryProvider` 观察值,不创建 Git 仓库。断言这些边界的测试仍使用真实 +提供者和仓库。真实 Git fixture 保留 `init`、`add` 和 `commit`,但将测试提交 +身份通过 `git -c` 传给 `commit`,不再为每个 fixture 分别启动两个配置进程。 + +## 考虑过的替代方案 + +- 把四类资源当成固定 Node worker 分区:实测分区和文件排序候选都比现有入口慢。 +- 所有 RCP 测试都使用固定观察值:范围、提交、forge 和 Git 故障断言需要 + 真实仓库行为,因此拒绝。 +- 所有编排测试都初始化 Git:这些测试只消费提供者结果,并不断言 Git 如何 + 产生结果,因此拒绝。 +- 多项测试共享一个可变数据库或 Git 仓库:运行顺序及并发可能影响结果。 + +## 后果 + +- 十个 reconciliation 测试保留真实 Git 观察;八个使用固定观察值。每个真实 + Git fixture 也少启动两个 Git 子进程。文件级改善已经复测,整套测试提速 + 仍未得到证明。 +- fixture 规则只指导准备工作;route 覆盖与 150 秒出结果时限仍由验证契约负责。 +- 若固定观察值测试后来需要断言 Git 范围或来源,应恢复真实 Git。若测试需要 + 仓库本地的 Git 身份配置,或受支持的 Git 环境行为不同,应撤销这一提交 + 身份捷径。 diff --git a/docs/design/ci-cd-and-quality.md b/docs/design/ci-cd-and-quality.md index 0fa7e47d..e5ff94d4 100644 --- a/docs/design/ci-cd-and-quality.md +++ b/docs/design/ci-cd-and-quality.md @@ -1,7 +1,7 @@ # 代码质量与 CI/CD **Created:** 2026-07-20 -**Updated:** 2026-09-16 +**Updated:** 2026-09-23 **Authority:** 仓库测试、CI 与 Agent 开发流程契约 NMG 的测试负责阻止可复现错误,不负责冻结尚未验证的设计。产品契约、研究测量与故障注入使用不同执行轨道,避免 benchmark 便利逻辑反向定义产品行为。 @@ -39,6 +39,8 @@ exit_criteria: Replace with a stable contract test or remove after the redesign `agent:context:check` 会拒绝缺少 `reason`、`review_after` 或 `exit_criteria` 的 active guardrail。满足退出条件后必须删除,或明确晋升为长期 Safety/Contract 测试。 +测试资源按断言边界选择,与上述阻塞分类正交:只断言路由测试失败的 narrow verifier fixture 可声明不适用共享检查,同时保留真实 CLI、Git 和路由测试;默认 narrow 的代表测试仍验证共享检查执行([决策](../decisions/implemented/2026-09-23-route-only-verifier-fixtures.md))。 + ## 2. 执行轨道 | 命令 | 内容 | 用途 | @@ -63,7 +65,7 @@ exit_criteria: Replace with a stable contract test or remove after the redesign 任何启动 daemon / MCP server 的测试在模块顶层调用 `tests/helpers/test-env.ts` 的 `stripProviderEnv()`,清空环境中的 `NMG_EMBED*` / `NMG_SUMMARY*` / `NMG_JUDGE*`。该约定由 `tests/support/test-env-guard.test.ts` 机械检查:扫描 `tests/**/*.test.ts`,凡是 spawn daemon(`connectDaemon(`、`StdioClientTransport`、`bin/nmg.mjs`、tutorial 脚本、pi extension harness)却没有模块顶层调用者一律失败。被测 daemon 继承测试进程环境;存在 ambient provider(如 `NMG_EMBED_PROVIDER=gemini`)时,recall 会依赖外部服务而变得缓慢且不确定。需要 provider 的测试必须显式传入环境,不得依赖 ambient 配置。 -覆盖率轨将测试并发限制为 2,避免 c8 插桩与跨进程/SQLite 测试叠加时制造内存峰值和 Windows 文件锁假失败;普通 product 轨仍使用并发 4。 +覆盖率轨将测试并发限制为 2,避免 c8 插桩与跨进程/SQLite 测试叠加时制造内存峰值和 Windows 文件锁假失败;普通 product 轨仍使用并发 4。覆盖率轨还以 c8 的 0.01% 行覆盖率下限检查是否产生非零源码证据;它是活性检查,不是覆盖率质量目标([决策](../decisions/implemented/2026-09-23-coverage-evidence-liveness.md))。 ## 3. CI 契约 @@ -99,6 +101,17 @@ TestRuntime 测试通过 `withTestRuntime(...)` 或显式 `dispose()` 获取 RAII 式回收。插件依赖必须显式:database 要求 workspace,daemon 要求 database;缺少依赖时失败,不偷偷创建隐含全局状态。daemon 在进程内使用真实 HTTP JSON-RPC handler,因此能验证协议,同时不会遗留后台进程。 +测试资源按断言需观察的边界分为四类,按需取得最轻且隔离的资源([决策](../decisions/implemented/2026-09-24-resource-matched-test-fixtures.md)): + +| 断言边界 | 测试资源 | 可省的准备工作 | +| --- | --- | --- | +| 纯数据或决策 | 进程内输入 | 工作区、数据库和外部进程 | +| 单个 SQLite 连接的读写 | 每项独立的 `:memory:` store | 临时数据库文件及其清理 | +| 重开、旧库迁移、WAL、独立连接或路径 | 每项唯一的文件数据库或工作区 | 与断言无关的额外文件和服务 | +| CLI、Git、HTTP、daemon lease 或跨进程共享 | 断言这些边界时使用实际进程、端口及所需工作区 | 不参与断言的重复启动和检查 | + +观察文件或跨进程语义的测试保留真实边界;只断言控制平面策略、receipt 或 harness 编排时,可以注入固定的 `RepositoryProvider` 观察值,另由真实 Git 集成测试覆盖发现、dirty 范围、提交和 forge 绑定。集成 `testDatabase()` 使用真实文件路径供 daemon 共享。真实 Git fixture 保留 `init/add/commit`,提交身份只传给 `commit`,避免每项再启动两次 `git config`。资源类别只决定 fixture 和可安全减少的准备工作,不改变 Safety/Contract/Guardrail 分类、阻塞地位或 narrow/full 验证范围;后两者由既有 route 与验证契约决定。 + 该层精确锁定 `@deepseek-ai/cordis@4.0.1` 作为 **devDependency**,只借用插件 effect/fiber 生命周期。NMG Core、daemon、Pi adapter 和发布包均不依赖 Cordis;不引入其 loader、HMR 或配置系统。Cordis adapter 与 NMG fixture composition 分离,未来若替换框架只需修改生命周期 adapter。 ## 5. Agent 原生仓库上下文 @@ -137,7 +150,7 @@ npm run agent:context -- --scope <目标路径> npm run agent:verify ``` -零参数入口自动读取 Git 变更并选择 route;Git 不可用时失败关闭,不允许空跑后误报成功。共享脏工作树使用 `--scope <本任务路径>` 精确覆盖自动范围。执行器聚合所有命中 route,按首次出现顺序去重,并运行全部 blocking 命令;一个命令退出失败、启动异常、超时或被信号终止,都被归因到该命令且不会阻止后续检查收集证据,最终统一返回失败。默认单命令上限为 30 分钟,可用 `--timeout-ms` 调整。 +零参数入口自动读取 Git 变更并选择 route;Git 不可用时失败关闭,不允许空跑后误报成功。共享脏工作树使用 `--scope <本任务路径>` 精确覆盖自动范围。执行器聚合所有命中 route,按首次出现顺序去重,并运行全部 blocking 命令;`ci-and-tests` route 将 `verify:static` 拆成与 CI contract 一致的原子检查,跨 route 共享的检查在同一计划内只执行一次,同时保留每项的状态和 route 归因。CI 仍运行命名的 `verify:static` contract。非 RCP full 计划在构建、打包等写入屏障后有界并发执行独立静态检查,产品测试等待静态检查完成;结果仍按计划顺序排列。RCP 和 narrow 计划仍串行。一个命令退出失败、启动异常或被信号终止,都被归因到该命令且不会阻止后续检查收集证据,最终统一返回失败。官方 `npm run agent:verify` 入口的整轮验证默认预算为 150 秒,可用 `--timeout-ms` 调整;外层看门进程在超时后返回失败及 `incomplete` 证据。普通执行和 RCP full/narrow 执行也共用从执行器入口开始计算的剩余预算;每个子命令只得到剩余时间。预算用尽时不再启动后续检查,逐项记录为失败,整轮不能报通过。超时意味着验证未完成,需要缩短或重新划分检查后再运行,不能把已通过的局部检查当作整轮通过。 advisory 命令默认只显示、不执行;显式传入 `--include-advisory` 后才运行,且其失败不改变 blocking 结果。`--dry-run` 只生成计划,`--json` 提供机器可读报告,`--require-clean` 为打包/CI 等任务增加干净工作树门槛。每次成功形成报告后会自动覆盖 `.nmg/verification/latest.json`;该文件包含 run ID、起止时间、运行时、Git HEAD、scope、route 和逐命令结果,位于已忽略的 `.nmg/` 下,不形成无限增长的日志历史。 diff --git a/docs/experiments/README.md b/docs/experiments/README.md index 440d281f..66bcee8c 100644 --- a/docs/experiments/README.md +++ b/docs/experiments/README.md @@ -24,5 +24,6 @@ when that prefix repeats the directory name. A rolling summary or note uses | `store/` | storage: consolidation, scale, and the write path | | `runtime/` | embedding backends, autodiff operators, and the controller shadow | | `execution/` | restricted out-of-order execution for agents: what the wait is worth, and which measured failure modes the machinery addresses (the earlier admission run stays at the top level) | +| `verification/` | measured verification planning and coverage experiments | `benchmark-results.md` stays at the top level because it spans topics. diff --git a/docs/experiments/README.zh-CN.md b/docs/experiments/README.zh-CN.md index 9db5e71d..b74b69cf 100644 --- a/docs/experiments/README.zh-CN.md +++ b/docs/experiments/README.zh-CN.md @@ -22,5 +22,6 @@ | `store/` | 存储:consolidation、规模、写路径 | | `runtime/` | 嵌入后端、autodiff 算子、controller shadow | | `execution/` | Agent 的受限乱序执行:等待到底值多少、机制覆盖了哪些实测失败模式(更早的准入实验仍在顶层) | +| `verification/` | 验证规划与覆盖范围的测量实验 | `benchmark-results.md` 留在顶层,因为它跨主题。 diff --git a/docs/experiments/verification/full-gate-atomic-checks-2026-09-23.md b/docs/experiments/verification/full-gate-atomic-checks-2026-09-23.md new file mode 100644 index 00000000..7943ee13 --- /dev/null +++ b/docs/experiments/verification/full-gate-atomic-checks-2026-09-23.md @@ -0,0 +1,35 @@ +# Full verification with atomic static checks — 2026-09-23 + +## Protocol + +On the same Windows working tree and Node v24.19.0, run +`rtk npm run agent:verify -- --full` with the route's original nested +`verify:static` check, then with `ci-and-tests` declaring that contract's +atomic checks. The changed test asserts exact parity between the route's +atomic list and the package script. Both runs use the default 150-second +whole-run budget and execute all blocking checks; research and chaos checks +remain advisory. The working tree also contains the uncommitted whole-run +deadline changes and import-impact probe. + +## Observation + +| Plan | End to end | Product tests | Other checks and overhead | Blocking outcome | +| --- | ---: | ---: | ---: | --- | +| Nested static plus independent route checks | 137.3 s | 85.1 s | 52.2 s | passed | +| Atomic static checks, deduplicated by name | 102.3 s | 72.9 s | 29.4 s | passed | +| Atomic static checks after documentation update | 107.4 s | 72.4 s | 35.0 s | passed | + +The first run reported eight blocking command results. The second reported +13 atomic results, each with its route attribution. The original +`verify:static` script contained six commands that the full plan also ran +independently. Their independent durations totalled 17.5 seconds in the first +run. The observed drop in non-product time was 22.8 seconds; run-to-run +variation and npm startup overhead prevent assigning all of it to the route +change. Product tests varied by 12.3 seconds without a product-test change. + +The atomic run passed 1,533 product tests and every blocking check. This one +pair of runs demonstrates the duplicate execution and the measured reduction; +it does not prove a worst-case bound below 150 seconds for every repository +state or machine. The final run includes the decision, design, and this +experiment document; it passed all 13 blocking checks. The full verifier +remains the source of the final verdict. diff --git a/docs/experiments/verification/full-gate-parallel-static-2026-09-23.md b/docs/experiments/verification/full-gate-parallel-static-2026-09-23.md new file mode 100644 index 00000000..49b2522a --- /dev/null +++ b/docs/experiments/verification/full-gate-parallel-static-2026-09-23.md @@ -0,0 +1,33 @@ +# Bounded parallel static checks — 2026-09-23 + +## Protocol + +The full Agent plan already listed atomic checks and deduplicated identical +commands. This experiment keeps build and package checks as serial barriers, +runs adjacent approved static checks with at most three asynchronous npm +processes, then runs `test:product` after the static group. Each check keeps +its own result and the whole run keeps one 150-second deadline. A scheduler +test checks parallel execution, barriers, declaration-order evidence, and +failure attribution; a real npm test checks asynchronous exit attribution. + +Run `rtk npm run agent:verify -- --full` on Windows with Node v24.19.0. +The serial comparison is the last run reported in +[the atomic-check experiment](full-gate-atomic-checks-2026-09-23.md). + +## Observation + +| Plan | End to end | Product tests | Other checks and overhead | Blocking result | +| --- | ---: | ---: | ---: | --- | +| Serial atomic checks | 100.0 s | 72.2 s | 27.8 s | passed | +| Bounded parallel static checks | 96.0 s | 75.2 s | 20.8 s | passed | + +The static and scheduling portion fell by about seven seconds while product +tests varied by about three seconds. A second parallel full run took 104.5 s, +including 84.2 s of product tests and about 20.3 s of other work; all 13 +blocking checks passed. Parallel `check` and `lint` individually took longer +than in the serial run because processes competed for CPU. These runs support +the static-phase reduction but do not establish a stable end-to-end speedup or +a worst-case bound. The first parallel run passed 1,534 product tests; +the second passed 1,535. Both passed all 13 blocking checks; +research and chaos checks remained advisory and were not run. RCP and narrow +verification paths were not changed by this experiment. diff --git a/docs/experiments/verification/impact-probe-2026-09-23.md b/docs/experiments/verification/impact-probe-2026-09-23.md new file mode 100644 index 00000000..4ba46d7c --- /dev/null +++ b/docs/experiments/verification/impact-probe-2026-09-23.md @@ -0,0 +1,60 @@ +# Import-impact verification probe — 2026-09-23 + +## Question and protocol + +Can a dependency graph select a smaller product-test set for a change to +`src/core/simhash.ts` while preserving an honest account of what remains +unverified? This is a research probe, not an admitted verification route. + +The probe inventories Git-visible TypeScript files, parses static imports and +exports plus literal `import()` and `require()` calls, follows reverse import +edges, and intersects reachable tests with the `test:product` file globs. +It hashes the inspected TypeScript files, `package.json`, `package-lock.json`, +`tsconfig.json`, and the Node version/platform/architecture before and after +candidate execution. It reports candidates with exit code 2, even when they +pass; exit code 1 means the candidate execution or planning failed. + +Commands run from the repository root on Windows: + +```text +rtk node --experimental-strip-types scripts/verification-impact-probe.ts --source src/core/simhash.ts --execute +rtk npm run test:product +rtk npm run agent:verify -- --full +rtk node --experimental-strip-types scripts/verification-impact-probe.ts --source src/core/simhash.ts +``` + +The candidate run and the subsequent full verifier used input digest +`8e861d5c8728603c07ae4aa5ebc04e3519294862b244426ed897440db04a1839`. + +## Observation + +| Run | Result | Elapsed | +| --- | --- | ---: | +| Import graph planning | 618 TypeScript files scanned; 14 candidate files of 175 product-test files | 0.7 s | +| Candidate test execution | Candidate tests passed; probe exited 2 (`candidate-only`) | 14.3 s | +| Full `test:product` within `agent:verify --full` | 1,532 tests passed | 85.1 s | +| Full `agent:verify --full` | All blocking checks passed; advisory research and chaos checks were skipped by policy | 137.3 s end to end | + +The candidate execution was about six times faster than the product-test +portion of this full verifier run. This is a comparison of elapsed test +execution, not a reliability or +whole-verifier speedup claim. The current official gate also performs other +checks and was not replaced in this experiment. + +## Coverage gaps and decision + +The graph found three reachable tests outside `test:product` and four unresolved +imports, including computed module paths and a generated module. It does not +track JavaScript modules, file reads, generated inputs, environment dependencies, +or behavioral requirements. A fixture test demonstrates a test that reads the +changed source file without importing it; the graph omits that test. Other +fixtures check transitive and literal dynamic imports, changing input digests, +missing sources, and the nonpassing CLI exit status. + +Therefore the 14 files are useful candidates for quick feedback, but their pass +cannot certify the SimHash change. The probe always reports +`fullGateRequired: true`. Before a subset can replace any blocking check, its +owner needs a declared input/dependency contract for that check, explicit +handling of unknown dependencies, and counterexample tests showing the selector +escalates when closure cannot be established. The full gate remains the +authoritative result for this experiment. diff --git a/docs/experiments/verification/product-file-partitions-2026-09-23.md b/docs/experiments/verification/product-file-partitions-2026-09-23.md new file mode 100644 index 00000000..9538a36b --- /dev/null +++ b/docs/experiments/verification/product-file-partitions-2026-09-23.md @@ -0,0 +1,54 @@ +# Product test file partition probe (2026-09-23) + +## Question + +Can the product test stage finish sooner by assigning every test file to a bounded set of processes using measured test durations? This measures one stage, not the complete verification deadline or a replacement gate. + +## Reference and difference + +DSH's `test:coverage:partitioned` is a supported coverage route. Its coordinator runs coverage-instrumented tests separately from heavy exempt suites because instrumentation multiplies the runtime of compiler and subprocess fixtures that do not contribute to the per-file coverage threshold. It shares a worker budget between these suites, requires all partitions, and merges coverage before making the coverage decision (`C:/Documents/GitHub/dsh/deepseek-harness/scripts/run-gates.ts`, `scripts/coverage-partitions.ts`). + +NMG's product tests use the Node test runner and have no matching per-file coverage exemption contract. That does not rule out adopting DSH's partition design: first identify which NMG files are slowed by instrumentation, startup, subprocesses, or shared resources, then carry over the mechanisms that address those costs. This probe tests duration-weighted file assignment as one possible mechanism; its separate entry lets us compare it with the current `test:product` route. + +## Method and result + +- Baseline: the existing `test:product` file globs, run once with Node's JUnit reporter and `--test-concurrency=4`. The report contained 175 distinct test files, 1536 passing tests, and an 80,594 ms duration. +- Candidate: a one-off script used the baseline JUnit case durations to assign all 175 files to four partitions. Each partition used one Node test worker. The script compared the actual JUnit file inventory and total test count against the profile before reporting a result. +- Candidate: 43/43/43/46 files and 356/358/396/426 passing tests by partition; all 175 files and 1536 tests were present. Partition durations were 67,852/67,630/65,552/69,213 ms; whole candidate duration was 69,254 ms. This one run was 11,340 ms (14.1%) faster than the baseline stage. +- The candidate emitted `comparison-only` and `fullGateRequired: true`, and exited 2 on successful comparison. It was not wired into `test:product`, `agent:verify`, or CI. + +## Why partitioning may help here + +The baseline JUnit report attributes 218.5 seconds of test-case time to 175 files. The three longest files contribute 35.7 seconds (`tests/tools/agent-verify.test.ts`), 28.3 seconds (`tests/rcp/reconcile.test.ts`), and 25.5 seconds (`tests/cli/process.test.ts`). The first repeatedly launches verifier, Git, npm, and Node subprocesses; the second repeats Git observations and reconciliation; the third launches CLI/daemon processes and contains explicit waits. These operations explain why the file durations are uneven, although the report does not isolate their individual costs. + +A four-worker greedy simulation using those same case durations gives worker loads of 51.7/69.7/50.1/47.0 seconds when files are offered in alphabetical order, and 54.6/54.6/54.6/54.6 seconds with longest-first placement. This is a scheduling model, not a trace of Node's actual worker start times: it omits module loading, process setup, and resource contention. Its 15.1-second reduction in the longest modeled load is consistent with the 11.3-second observed stage reduction. The available evidence points to a long-tail scheduling bubble as the strongest explanation, but does not prove that process separation itself adds no benefit or cost. + +The common test runtime gives each test workspace a unique temporary directory and binds daemon HTTP servers to port 0 (`tests/support/test-runtime.ts`). The complexity gate's probe names include the process ID and live under ignored `.nmg/complexity-probes` (`tools/complexity-gate.ts`), so the fixed `probePath` argument in its test is not itself a shared output file. The current full verifier finishes build/package and static checks before starting `test:product` (`tools/agent-verify.ts`, `src/rcp/verification.ts`), avoiding concurrent writes to generated inputs. These facts support file-level parallelism at a bounded worker count; they do not establish safety for overlapping product tests with build or static checks, or for increasing total concurrency. + +DSH's coverage-exempt lane addresses a different measured cost: coverage instrumentation multiplies heavy fixture runtime without contributing to its per-file threshold. NMG has not measured an analogous instrumentation penalty or established an exempt set, so that part of DSH's design remains an open hypothesis. Its duration-weighted assignment addresses the imbalance visible in this profile. + +## Resource boundary trial + +The existing [no-filesystem decision](../../decisions/implemented/2026-09-20-tests-need-no-filesystem.md) removed candidate worktrees from data-only checks because those checks did not assert workspace behavior. The four resource classes and their fixture boundary are now stated in the [test-runtime design](../../design/ci-cd-and-quality.md#4-可组合测试运行时). The trials below measure particular changes within those classes; they do not change Safety/Contract/Guardrail or narrow/full verification membership. + +On this Windows checkout, 40 repetitions measured directory create/remove at mean 0.258 ms (median 0.228 ms), `NmgStore(":memory:")` open/close at mean 17.419 ms (median 17.106 ms), and file-backed store open/close including a unique directory and removal at mean 32.874 ms (median 32.459 ms). Each store construction ran the schema migration. This isolates startup cost, not a whole test's work; it shows that directory creation alone is a small part of the cost. + +`tests/core/store.test.ts` had 52 passing tests and took 12.153 seconds with its original file-backed `withStore` helper. The 45 single-connection calls were changed to a fresh in-memory store; seven explicit file-backed tests still exercise reopening, external readers, and migration. The same file passed 52/52 tests in 9.363 and 8.706 seconds on two candidate runs. These are one baseline and two candidate measurements on a live machine, so the difference is provisional; it cannot be attributed solely to directory removal. The slow FTS5 case also writes 1,041 records, making database I/O part of its time. The scoped full `agent:verify` run passed all 13 blocking checks in 110.955 seconds under its 150-second budget; `test:product` took 88.695 seconds. That whole-stage time was slower than the earlier standalone 80.594-second baseline, so this run establishes correctness and deadline compliance, not an end-to-end speedup attributable to the fixture change. + +The same single-connection criterion applies to the 55 maintenance tests in `tests/core/store/maintenance.test.ts`: all passed with a file store in 2.494 seconds, then with a fresh memory store per test in 1.470 seconds. That file has no restart or independent-connection assertion. + +Four tests in `tests/tools/agent-verify.test.ts` asserted only route-test failure/TAP evidence but ran seven successful shared npm checks through the fixture. They now declare the existing `verify.sharedChecks: none` on those fixture routes, while retaining the real CLI, Git discovery, route test, receipt, and output; another test still asserts all seven shared checks in default narrow mode. All 28 tests passed before and after. In one paired file run, duration fell from 32.672 to 22.244 seconds; the four affected cases together fell from approximately 17.55 to 5.53 seconds. This measures avoided process work in a targeted run, not a full-gate speedup. + +The RCP repository fixture was another resource trial. Initially, all 18 `tests/rcp/reconcile.test.ts` cases initialized Git with `init`, two `config` commands, `add`, and `commit`; the file took 22.153 seconds. Supplying identity through `git -c` on `commit` removed the two config child processes without changing the real repository boundary; two runs took 21.647/20.494 seconds. Inspection then showed that eight orchestration tests consumed a `RepositoryProvider` observation but did not assert Git discovery, dirty scope, or commit provenance. Those eight now use a fixed provider and skip Git setup; ten scope, failure, and forge tests retain the real provider and repository. All 18 tests passed in two runs of 13.802/13.776 seconds. The exact setup reduction is eight Git initializations and two fewer config processes for each remaining real-Git fixture. This is a repeated file-level result on one Windows checkout, not proof of an end-to-end verification speedup. + +An adjacent whole-run A/B then changed only the two RCP test files between the committed baseline blobs (`fixture.ts` `3da0dd527ad9f7caf3b6892004423c39325478c4`, `reconcile.test.ts` `195f87ef6796431fc109b0869fcbe5704d8c56c4`) and the candidate blobs (`665a76f75ef850eb2070af1c0ed78cae79f33fad`, `a4bfa01a6237c7712922c79a086787736b56c7b0`). The rest of the dirty worktree was held constant and runs were serial. Both `agent:verify` runs passed all 13 blocking checks and 1536 product tests: baseline 77.287 seconds overall with 61.026 seconds in `test:product`; candidate 77.406 seconds overall with 60.869 seconds in `test:product`. The paired differences, +0.119 seconds overall and -0.157 seconds in product tests, are effectively zero at this measurement scale. No worker-start trace was recorded, so the likely explanation that the faster RCP file was not on the suite's critical path remains a hypothesis. This single pair does not establish a whole-run speedup despite the repeatable file-level reduction. + +With the maintenance and route-only fixture changes, a scoped full `agent:verify` run passed all 13 blocking checks in 101.832 seconds under the 150-second budget; `test:product` took 78.844 seconds. The previous full run, before these two changes, took 102.541 seconds with 80.130 seconds in `test:product`. These unpaired whole-run samples are too close and variable to establish a whole-gate speedup. + +A focused c8 run of the 28 verifier tests passed and reported 0% because those tests did not execute any source file in the configured c8 include set. The raw V8 blobs included repository URLs, but none under `src/core`. A separate single-file probe of `tests/core/ann.test.ts` produced 84.42% coverage for `src/core/ann.ts`, so the 0% sample does not show a broken c8 report. With c8's 0.01% line floor, five out-of-scope tool tests passed but the coverage command failed at 0%, while the single ANN test passed at 0.38% overall line coverage. The official `test:coverage` route then passed all 1536 tests with 85.16% line coverage in 126.454 seconds. These runs validate the nonzero-evidence check; they do not isolate an instrumentation-exempt group or establish a coverage-quality target. + +## Limits and next decision + +The earlier one-run speedup did not repeat on the current worktree. A paired rerun after the single-connection and route-only fixture changes used the same 176-file, 1539-test inventory: the ordinary `test:product` route passed in 68.337 seconds (Node reported 67.784 seconds), while the four fixed partitions passed in 87.067 seconds. The profile still described the earlier 175-file, 1536-test run; the candidate assigned the new file a median fallback weight and verified every current file and test from JUnit. Partition durations were 68.856/83.405/80.403/86.999 seconds. An additional single-parent Node run, offering the same 176 files from longest recorded duration to shortest while retaining `--test-concurrency=4`, passed all 1539 tests in 75.453 seconds. Both current candidates were slower than the unchanged route, so the candidate code and profile were removed instead of being promoted. + +These are one paired baseline/partition run and one ordered run on a live Windows checkout, not a distribution of repeat measurements. They establish that the previous 14.1% speedup is not reliable on this tree. The old JUnit weights also precede changes to the longest verifier test file; startup and resource contention remain unisolated. Further scheduling work needs a current profile and repeated paired runs before it can replace `test:product`. The current resource-fixture reductions and the existing 150-second whole-run gate are the validated path for this change. diff --git a/package.json b/package.json index 6258b030..0241f63a 100644 --- a/package.json +++ b/package.json @@ -59,7 +59,7 @@ "lint:fix": "eslint --fix src/ .pi/extensions/ claude-plugins/ workbuddy-plugin/ tests/ evals/ scripts/ tools/", "format": "prettier --write \"src/**/*.ts\" \".pi/**/*.ts\" \"workbuddy-plugin/**/*.ts\"", "format:check": "prettier --check \"src/**/*.ts\" \".pi/**/*.ts\" \"workbuddy-plugin/**/*.ts\"", - "test:coverage": "c8 node --experimental-strip-types --test --test-concurrency=2 \"tests/cli/**/*.test.ts\" \"tests/core/**/*.test.ts\" \"tests/docker/**/*.test.ts\" \"tests/docs/**/*.test.ts\" \"tests/extensions/**/*.test.ts\" \"tests/guardrails/**/*.test.ts\" \"tests/integration/**/*.test.ts\" \"tests/lab/**/*.test.ts\" \"tests/prompts/**/*.test.ts\" \"tests/rcp/**/*.test.ts\" \"tests/scripts/**/*.test.ts\" \"tests/skills/**/*.test.ts\" \"tests/support/**/*.test.ts\" \"tests/tools/**/*.test.ts\"", + "test:coverage": "c8 --check-coverage --lines 0.01 node --experimental-strip-types --test --test-concurrency=2 \"tests/cli/**/*.test.ts\" \"tests/core/**/*.test.ts\" \"tests/docker/**/*.test.ts\" \"tests/docs/**/*.test.ts\" \"tests/extensions/**/*.test.ts\" \"tests/guardrails/**/*.test.ts\" \"tests/integration/**/*.test.ts\" \"tests/lab/**/*.test.ts\" \"tests/prompts/**/*.test.ts\" \"tests/rcp/**/*.test.ts\" \"tests/scripts/**/*.test.ts\" \"tests/skills/**/*.test.ts\" \"tests/support/**/*.test.ts\" \"tests/tools/**/*.test.ts\"", "eval:agents": "node --experimental-strip-types evals/run.ts", "hotspot:modules": "node --experimental-strip-types scripts/hotspot-files.ts", "perf:hotspots": "node --experimental-strip-types scripts/perf-hotspots.ts", @@ -123,7 +123,7 @@ "docs:check": "node --experimental-strip-types scripts/verify-docs.mts", "agent:context": "node --experimental-strip-types tools/repo-context.ts", "agent:context:check": "node --experimental-strip-types tools/repo-context.ts --check", - "agent:verify": "node --experimental-strip-types tools/agent-verify.ts", + "agent:verify": "node --experimental-strip-types tools/agent-verify-watchdog.ts", "ci:uncovered-tests": "node --experimental-strip-types tools/ci-uncovered-tests.ts --check", "complexity:gate": "node --experimental-strip-types tools/complexity-gate.ts", "mutation:teeth": "node --experimental-strip-types tools/mutation-teeth.ts", diff --git a/scripts/verification-impact-probe.ts b/scripts/verification-impact-probe.ts new file mode 100644 index 00000000..5a20c9a5 --- /dev/null +++ b/scripts/verification-impact-probe.ts @@ -0,0 +1,265 @@ +/** Research probe: find test files that can import a changed module. Its output is + * a candidate set, never a verification verdict: imports alone miss file reads, + * computed module paths, and behavioral obligations. */ +import { createHash } from "node:crypto"; +import { spawnSync } from "node:child_process"; +import { existsSync, globSync, readFileSync } from "node:fs"; +import { dirname, isAbsolute, join, relative, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; +import { parseArgs } from "node:util"; + +import ts from "typescript"; + +import { testOutputPassed } from "../src/rcp/trusted.ts"; + +interface ImpactPlan { + source: string; + verdict: "candidate-only"; + fullGateRequired: true; + snapshotDigest: string; + scannedFiles: number; + productTests: number; + candidateTests: string[]; + affectedOutsideProduct: string[]; + unresolvedImports: string[]; + limitations: string[]; +} + +const MODULE_SUFFIXES = ["", ".ts", ".tsx", ".mts", ".cts", "/index.ts"]; + +function repositoryFiles(root: string): string[] { + const git = spawnSync("git", ["ls-files", "-z", "--cached", "--others", "--exclude-standard"], { + cwd: root, + encoding: "utf8", + windowsHide: true, + }); + if (git.error || git.status !== 0) throw new Error(`Git file inventory failed: ${git.stderr}`); + return [ + ...new Set( + git.stdout + .split("\0") + .filter(Boolean) + .map((path) => path.replaceAll("\\", "/")), + ), + ].sort(); +} + +function sourceFiles(files: string[]): string[] { + return files.filter((path) => /\.(?:c|m)?tsx?$/.test(path)); +} + +function productTestFiles(root: string, files: Set): Set { + const packageJson = JSON.parse(readFileSync(join(root, "package.json"), "utf8")) as { + scripts?: Record; + }; + const script = packageJson.scripts?.["test:product"]; + if (!script) throw new Error("test:product script is missing"); + const patterns = [...script.matchAll(/"(tests\/[^"\s]+\.test\.ts)"/g)].map((match) => match[1]!); + if (!patterns.length) throw new Error("test:product has no parseable test globs"); + const tests = new Set(); + for (const pattern of patterns) { + for (const match of globSync(pattern, { cwd: root })) { + const path = match.replaceAll("\\", "/"); + if (files.has(path)) tests.add(path); + } + } + if (!tests.size) throw new Error("test:product matched no tracked tests"); + return tests; +} + +function resolveLocalImport( + from: string, + specifier: string, + files: Set, +): string | undefined { + if (!specifier.startsWith(".")) return undefined; + const base = resolve(dirname(from), specifier); + for (const suffix of MODULE_SUFFIXES) { + const candidate = `${base}${suffix}`; + if (files.has(candidate)) return candidate; + } + if (base.endsWith(".js")) { + const candidate = `${base.slice(0, -3)}.ts`; + if (files.has(candidate)) return candidate; + } + return undefined; +} + +function moduleSpecifiers( + path: string, + source: string, +): { specifiers: string[]; unknown: string[] } { + const tree = ts.createSourceFile(path, source, ts.ScriptTarget.Latest, true); + const specifiers: string[] = []; + const unknown: string[] = []; + const add = (value: ts.Expression | undefined, kind: string) => { + if (value && ts.isStringLiteralLike(value)) specifiers.push(value.text); + else unknown.push(`${path}: ${kind} has a computed path`); + }; + const visit = (node: ts.Node) => { + if (ts.isImportDeclaration(node) || ts.isExportDeclaration(node)) { + if (node.moduleSpecifier) add(node.moduleSpecifier, "module declaration"); + } else if ( + ts.isImportEqualsDeclaration(node) && + ts.isExternalModuleReference(node.moduleReference) + ) { + add(node.moduleReference.expression, "import equals"); + } else if (ts.isCallExpression(node)) { + if (node.expression.kind === ts.SyntaxKind.ImportKeyword) add(node.arguments[0], "import()"); + if (ts.isIdentifier(node.expression) && node.expression.text === "require") + add(node.arguments[0], "require()"); + } + ts.forEachChild(node, visit); + }; + visit(tree); + return { specifiers, unknown }; +} + +function scanImports( + root: string, + sources: string[], + sourceSet: Set, + digest: ReturnType, +): { reverse: Map>; unresolved: string[] } { + const reverse = new Map>(); + const unresolved: string[] = []; + for (const path of sources) { + const absolute = resolve(root, path); + const content = readFileSync(absolute, "utf8"); + digest.update(path).update("\0").update(content).update("\0"); + const imports = moduleSpecifiers(path, content); + unresolved.push(...imports.unknown); + for (const specifier of imports.specifiers) { + if (!specifier.startsWith(".")) continue; + const dependency = resolveLocalImport(absolute, specifier, sourceSet); + if (!dependency) { + unresolved.push(`${path}: unresolved ${specifier}`); + continue; + } + const consumers = reverse.get(dependency) ?? new Set(); + consumers.add(absolute); + reverse.set(dependency, consumers); + } + } + return { reverse, unresolved }; +} + +export function planImpact(root: string, source: string): ImpactPlan { + const absoluteRoot = resolve(root); + const absoluteSource = isAbsolute(source) ? resolve(source) : resolve(absoluteRoot, source); + const inventory = repositoryFiles(absoluteRoot); + const allFiles = new Set(inventory.map((path) => resolve(absoluteRoot, path))); + if (!allFiles.has(absoluteSource) || !existsSync(absoluteSource)) + throw new Error(`source is not a present repository file: ${source}`); + const sources = sourceFiles(inventory); + if (!sources.length) throw new Error("repository has no TypeScript source files"); + const sourceSet = new Set(sources.map((path) => resolve(absoluteRoot, path))); + const productTests = productTestFiles(absoluteRoot, new Set(inventory)); + const digest = createHash("sha256"); + const { reverse, unresolved } = scanImports(absoluteRoot, sources, sourceSet, digest); + for (const path of ["package.json", "package-lock.json", "tsconfig.json"]) { + const absolute = join(absoluteRoot, path); + if (!existsSync(absolute)) throw new Error(`missing verification input: ${path}`); + digest.update(path).update("\0").update(readFileSync(absolute)).update("\0"); + } + digest.update(process.version).update(process.platform).update(process.arch); + + const affected = new Set([absoluteSource]); + const queue = [absoluteSource]; + for (const current of queue) { + for (const consumer of reverse.get(current) ?? []) { + if (affected.has(consumer)) continue; + affected.add(consumer); + queue.push(consumer); + } + } + const relativeAffected = [...affected].map((path) => + relative(absoluteRoot, path).replaceAll("\\", "/"), + ); + const candidateTests = relativeAffected.filter((path) => productTests.has(path)).sort(); + const affectedOutsideProduct = relativeAffected + .filter((path) => /\.test\.[cm]?tsx?$/.test(path) && !productTests.has(path)) + .sort(); + if (!candidateTests.length) throw new Error("impact graph found no product test consumers"); + if (candidateTests.some((path) => !productTests.has(path))) + throw new Error("impact graph selected a test outside test:product"); + return { + source: relative(absoluteRoot, absoluteSource).replaceAll("\\", "/"), + verdict: "candidate-only", + fullGateRequired: true, + snapshotDigest: digest.digest("hex"), + scannedFiles: sources.length, + productTests: productTests.size, + candidateTests, + affectedOutsideProduct, + unresolvedImports: [...new Set(unresolved)].sort(), + limitations: [ + "Import reachability does not prove behavioral test sufficiency.", + "Filesystem reads, generated inputs, and environment dependencies are not mapped.", + "This candidate set is not a passing verification or an authorized replacement for agent:verify.", + ], + }; +} + +const invoked = process.argv[1] ? resolve(process.argv[1]) : ""; +if (invoked === fileURLToPath(import.meta.url)) { + try { + const { values } = parseArgs({ + options: { + root: { type: "string" }, + source: { type: "string" }, + execute: { type: "boolean" }, + }, + }); + if (!values.source) + throw new Error("usage: --source [--root ] [--execute]"); + const root = resolve(values.root ?? process.cwd()); + const started = performance.now(); + const plan = planImpact(root, values.source); + const planningMs = Math.round(performance.now() - started); + let execution: + | { candidateTestsPassed: boolean; durationMs: number; exitCode?: number; reason?: string } + | undefined; + if (values.execute) { + const env = { ...process.env }; + delete env.NODE_TEST_CONTEXT; + const before = plan.snapshotDigest; + const run = spawnSync( + process.execPath, + [ + "--experimental-strip-types", + "--test", + "--test-reporter=tap", + "--test-concurrency=4", + ...plan.candidateTests, + ], + { + cwd: root, + encoding: "utf8", + windowsHide: true, + timeout: 150_000, + maxBuffer: 16 * 1024 * 1024, + env, + }, + ); + const after = planImpact(root, values.source).snapshotDigest; + const candidateTestsPassed = + !run.error && run.status === 0 && testOutputPassed(run.stdout ?? "") && before === after; + execution = { + candidateTestsPassed, + durationMs: Math.round(performance.now() - started - planningMs), + exitCode: run.status ?? undefined, + reason: + before !== after ? "repository inputs changed during execution" : run.error?.message, + }; + } + process.stdout.write( + `${JSON.stringify({ measuredAt: new Date().toISOString(), root, planningMs, ...plan, execution }, null, 2)}\n`, + ); + // Exit 2 means the candidate set was measured, but it is not an admissible gate. + process.exitCode = execution?.candidateTestsPassed === false ? 1 : 2; + } catch (error) { + process.stderr.write(`${error instanceof Error ? error.message : String(error)}\n`); + process.exitCode = 1; + } +} diff --git a/src/rcp/providers.ts b/src/rcp/providers.ts index 333b0237..6c7ae292 100644 --- a/src/rcp/providers.ts +++ b/src/rcp/providers.ts @@ -207,6 +207,15 @@ export class ProcessHarnessProvider implements HarnessProvider { } } +function overallDeadlineFailure(name: string, timeoutMs: number): VerificationCheckResult { + return { + name, + status: "failed", + durationMs: 0, + reason: `verification exceeded ${timeoutMs}ms overall deadline`, + }; +} + export class LocalNpmVerifierProvider implements VerifierProvider { readonly descriptor: ProviderDescriptor = { id: "local-npm-verifier", @@ -218,10 +227,12 @@ export class LocalNpmVerifierProvider implements VerifierProvider { readonly timeoutMs: number; readonly streamOutput: boolean; + readonly remainingMs: () => number; - constructor(timeoutMs = 30 * 60 * 1_000, streamOutput = false) { + constructor(timeoutMs = 30 * 60 * 1_000, streamOutput = false, remainingMs = () => timeoutMs) { this.timeoutMs = timeoutMs; this.streamOutput = streamOutput; + this.remainingMs = remainingMs; } async definitionDigest(request: { @@ -252,6 +263,11 @@ export class LocalNpmVerifierProvider implements VerifierProvider { ); const checks: VerificationCheckResult[] = []; for (const name of request.workOrder.verificationChecks) { + const budget = this.remainingMs(); + if (budget <= 0) { + checks.push(overallDeadlineFailure(name, this.timeoutMs)); + continue; + } const definition = scripts[name]; if (!definition) { checks.push({ @@ -262,7 +278,7 @@ export class LocalNpmVerifierProvider implements VerifierProvider { }); continue; } - checks.push(runNpmScriptCheck(request.root, name, this.timeoutMs, this.streamOutput)); + checks.push(runNpmScriptCheck(request.root, name, budget, this.streamOutput)); } return { provider: this.descriptor, @@ -301,10 +317,12 @@ export class NarrowVerifierProvider implements VerifierProvider { readonly timeoutMs: number; readonly streamOutput: boolean; + readonly remainingMs: () => number; - constructor(timeoutMs = 30 * 60 * 1_000, streamOutput = false) { + constructor(timeoutMs = 30 * 60 * 1_000, streamOutput = false, remainingMs = () => timeoutMs) { this.timeoutMs = timeoutMs; this.streamOutput = streamOutput; + this.remainingMs = remainingMs; } async definitionDigest(request: { @@ -331,10 +349,13 @@ export class NarrowVerifierProvider implements VerifierProvider { const routes = readRouteDeclarations(request.root); const checks: VerificationCheckResult[] = []; for (const name of request.workOrder.verificationChecks) { + const budget = this.remainingMs(); + if (budget <= 0) { + checks.push(overallDeadlineFailure(name, this.timeoutMs)); + continue; + } if (name.startsWith("node-test:")) { - checks.push( - runRouteTestsCheck(request.root, name, routes, this.timeoutMs, this.streamOutput), - ); + checks.push(runRouteTestsCheck(request.root, name, routes, budget, this.streamOutput)); continue; } const definition = scripts[name]; @@ -347,7 +368,7 @@ export class NarrowVerifierProvider implements VerifierProvider { }); continue; } - checks.push(runNpmScriptCheck(request.root, name, this.timeoutMs, this.streamOutput)); + checks.push(runNpmScriptCheck(request.root, name, budget, this.streamOutput)); } return { provider: this.descriptor, diff --git a/src/rcp/verification.ts b/src/rcp/verification.ts index f374641f..b4a30ded 100644 --- a/src/rcp/verification.ts +++ b/src/rcp/verification.ts @@ -1,4 +1,4 @@ -import { spawnSync, type SpawnSyncReturns } from "node:child_process"; +import { execFile, spawnSync, type SpawnSyncReturns } from "node:child_process"; import type { RouteDeclaration, VerificationCheckResult } from "./types.ts"; @@ -51,32 +51,31 @@ export async function executeVerificationPlan( includeAdvisory?: boolean; dryRun?: boolean; run: CommandRunner; + parallel?: { commands: ReadonlySet; concurrency: number }; }, ): Promise { - const results: VerificationCommandResult[] = []; - const execute = async (item: VerificationPlanItem) => { + const execute = async (item: VerificationPlanItem): Promise => { if (options.dryRun) { - results.push({ ...item, status: "skipped", durationMs: 0, reason: "dry run" }); - return; + return { ...item, status: "skipped", durationMs: 0, reason: "dry run" }; } try { const result = await options.run(item.command, item.classification, item.routes); - results.push({ ...result, ...item }); + return { ...result, ...item }; } catch (cause) { const reason = cause instanceof Error ? cause.message : String(cause); - results.push({ + return { ...item, status: "failed", durationMs: 0, errorKind: "runner", reason, output: outputTail(reason), - }); + }; } }; - for (const item of plan.blocking) await execute(item); + const results = await executeBlocking(plan.blocking, execute, options.parallel); if (options.includeAdvisory) { - for (const item of plan.advisory) await execute(item); + for (const item of plan.advisory) results.push(await execute(item)); } else { for (const item of plan.advisory) { results.push({ @@ -95,9 +94,64 @@ export async function executeVerificationPlan( }; } -export function npmCommandRunner(root: string, quiet: boolean, timeoutMs: number): CommandRunner { +async function executeBlocking( + items: VerificationPlanItem[], + execute: (item: VerificationPlanItem) => Promise, + parallel?: { commands: ReadonlySet; concurrency: number }, +): Promise { + const results: VerificationCommandResult[] = []; + for (let index = 0; index < items.length;) { + const item = items[index]!; + if (!parallel?.commands.has(item.command)) { + results.push(await execute(item)); + index += 1; + continue; + } + const group: VerificationPlanItem[] = []; + while (index < items.length && parallel.commands.has(items[index]!.command)) { + group.push(items[index++]!); + } + const slots = new Array(group.length); + let next = 0; + await Promise.all( + Array.from( + { length: Math.min(Math.max(1, parallel.concurrency), group.length) }, + async () => { + while (next < group.length) { + const slot = next++; + slots[slot] = await execute(group[slot]!); + } + }, + ), + ); + results.push(...slots); + } + return results; +} + +export function npmCommandRunner( + root: string, + quiet: boolean, + timeoutMs: number, + remainingMs: () => number = () => timeoutMs, + asyncCommands: ReadonlySet = new Set(), +): CommandRunner { return async (command, classification, routes) => { - const result = runNpmScriptCheck(root, command, timeoutMs, !quiet); + const budget = remainingMs(); + if (budget <= 0) { + return { + command, + classification, + routes, + status: "failed", + durationMs: 0, + reason: `verification exceeded ${timeoutMs}ms overall deadline`, + errorKind: "timeout", + }; + } + const result = asyncCommands.has(command) + ? await runNpmScriptCheckAsync(root, command, budget, !quiet) + : runNpmScriptCheck(root, command, budget, !quiet); return { command, classification, @@ -112,6 +166,48 @@ export function npmCommandRunner(root: string, quiet: boolean, timeoutMs: number }; } +function runNpmScriptCheckAsync( + root: string, + name: string, + timeoutMs: number, + streamOutput: boolean, +): Promise { + const started = performance.now(); + const { executable, args } = npmInvocation(name); + return new Promise((resolveResult) => { + execFile( + executable, + args, + { + cwd: root, + encoding: "utf8", + windowsHide: true, + maxBuffer: 16 * 1024 * 1024, + timeout: timeoutMs, + }, + (error, stdout, stderr) => { + if (streamOutput) streamCapturedOutput(stdout, stderr); + const exitCode = typeof error?.code === "number" ? error.code : error ? undefined : 0; + const reason = error + ? error.killed && performance.now() - started >= timeoutMs - 10 + ? `command exceeded ${timeoutMs}ms timeout` + : error.signal + ? `command terminated by ${error.signal}` + : error.message + : undefined; + resolveResult({ + name, + status: error ? "failed" : "passed", + durationMs: Math.round(performance.now() - started), + exitCode, + reason, + evidence: error ? outputTail(`${stdout}${stderr}${error.message}`) : undefined, + }); + }, + ); + }); +} + export function runNpmScriptCheck( root: string, name: string, diff --git a/tests/core/store.test.ts b/tests/core/store.test.ts index 5039342e..26d40548 100644 --- a/tests/core/store.test.ts +++ b/tests/core/store.test.ts @@ -8,19 +8,17 @@ import test from "node:test"; import { NmgStore } from "../../src/core/store.ts"; import type { MemoryMarker, VectorEmbedder } from "../../src/core/types.ts"; -function withStore(run: (store: NmgStore) => void, embedder?: VectorEmbedder): void { - const directory = mkdtempSync(join(tmpdir(), "nmg-test-")); - const store = new NmgStore(join(directory, "nmg.sqlite"), embedder); +function withMemoryStore(run: (store: NmgStore) => void, embedder?: VectorEmbedder): void { + const store = new NmgStore(":memory:", embedder); try { run(store); } finally { store.close(); - rmSync(directory, { recursive: true, force: true }); } } test("remember persists a memory with traceable evidence", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "NMG uses Pi as its agent harness", nodeName: "NMG architecture", @@ -39,7 +37,7 @@ test("remember persists a memory with traceable evidence", () => { }); test("memory markers persist as extensible control metadata", () => { - withStore((store) => { + withMemoryStore((store) => { const markers: MemoryMarker[] = [ { kind: "forget", attributes: { effect: "revoke" } }, { kind: "future-extension", attributes: { enabled: true, score: 0.8 } }, @@ -60,7 +58,7 @@ test("memory markers persist as extensible control metadata", () => { }); test("memory writes retain accepted and privacy-safe rejected policy decisions", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "The user prefers compact technical explanations", nodeName: "response preference", @@ -93,7 +91,7 @@ test("memory writes retain accepted and privacy-safe rejected policy decisions", }); test("disposable controller probes do not persist retrieval traces", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "Project Atlas stores analytics in DuckDB", nodeName: "Atlas" }); const context = store.searchContext("Atlas analytics", { persistTrace: false }); assert.ok(context.activeGraph); @@ -102,7 +100,7 @@ test("disposable controller probes do not persist retrieval traces", () => { }); test("semantic searchContext preserves Active Graph budgets and tracing", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "Project Atlas stores analytics in DuckDB", nodeName: "Project Atlas storage", @@ -156,7 +154,7 @@ test("memory survives closing and reopening the local database", () => { }); test("search respects local tier and result budgets", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "Docker is only an optional execution backend", nodeName: "sandbox boundary", @@ -169,7 +167,7 @@ test("search respects local tier and result budgets", () => { }); test("the same semantic node is reused without merging evidence", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.remember({ statement: "Cloud sync is optional", nodeName: "deployment", @@ -186,7 +184,7 @@ test("the same semantic node is reused without merging evidence", () => { }); test("scope filters memories without discarding other scopes", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "The production database is PostgreSQL", nodeName: "database", @@ -208,7 +206,7 @@ test("scope filters memories without discarding other scopes", () => { }); test("a newer state supersedes but does not delete the old memory", () => { - withStore((store) => { + withMemoryStore((store) => { const previous = store.remember({ statement: "Project Atlas uses Python 3.11", nodeName: "Project Atlas runtime", @@ -277,7 +275,7 @@ test("session archives checkpoint changes without entering semantic search", () }); test("stateKey automatically supersedes the active state in the same scope", () => { - withStore((store) => { + withMemoryStore((store) => { const previous = store.remember({ statement: "Charity 5K personal best is 27:12", nodeName: "running personal best", @@ -307,7 +305,7 @@ test("stateKey automatically supersedes the active state in the same scope", () }); test("semantically equivalent state keys become aliases before supersession", () => { - withStore((store) => { + withMemoryStore((store) => { const previous = store.remember({ statement: "The charity 5K personal best is 27:12", nodeName: "user-5k-personal-best", @@ -333,7 +331,7 @@ test("semantically equivalent state keys become aliases before supersession", () }); test("state alias repair does not merge a goal with a personal best", () => { - withStore((store) => { + withMemoryStore((store) => { const best = store.remember({ statement: "The charity 5K personal best is 25:50", nodeName: "user-running-5k-personal-best", @@ -354,7 +352,7 @@ test("state alias repair does not merge a goal with a personal best", () => { }); test("events and conversation evidence preserve time, actor, and truth status", () => { - withStore((store) => { + withMemoryStore((store) => { const event = store.remember({ statement: "User visited MoMA", nodeName: "MoMA visit", @@ -376,7 +374,7 @@ test("events and conversation evidence preserve time, actor, and truth status", }); test("claims are persisted and derive the record-level logical rollup", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "The user has used Flask, but has never used Flask-Login.", nodeName: "Flask experience", @@ -417,7 +415,7 @@ test("claims are persisted and derive the record-level logical rollup", () => { }); test("searchContext can restrict evidence to the requested source actor", () => { - withStore((store) => { + withMemoryStore((store) => { const userMemory = store.remember({ statement: "User requested the Sunday rotation", nodeName: "Sunday rotation request", @@ -446,7 +444,7 @@ test("searchContext can restrict evidence to the requested source actor", () => }); test("derived memories retain every source evidence and graph relation", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.remember({ statement: "The blazer must be picked up", nodeName: "blazer pickup", @@ -477,7 +475,7 @@ test("derived memories retain every source evidence and graph relation", () => { }); test("searchContext combines matching memories with typed graph edges", () => { - withStore((store) => { + withMemoryStore((store) => { const preference = store.remember({ statement: "User prefers Adobe Premiere Pro advanced tutorials", nodeName: "video editing preference", @@ -506,7 +504,7 @@ test("searchContext combines matching memories with typed graph edges", () => { }); test("getContext expands selected memory IDs without searching", () => { - withStore((store) => { + withMemoryStore((store) => { const wanted = store.remember({ statement: "The selected memory keeps its exact statement", nodeName: "progressive disclosure", @@ -532,7 +530,7 @@ test("getContext expands selected memory IDs without searching", () => { }); test("aggregation context overfetches and prioritizes countable memories", () => { - withStore((store) => { + withMemoryStore((store) => { for (let index = 0; index < 6; index += 1) { store.remember({ statement: `Assistant gave pickup and return organization tip ${index}`, @@ -557,7 +555,7 @@ test("aggregation context overfetches and prioritizes countable memories", () => }); test("lexical ranking ignores English question words that crowd out countable evidence", () => { - withStore((store) => { + withMemoryStore((store) => { for (let index = 0; index < 12; index += 1) { store.remember({ statement: `How many of the ideas are in the guide and what is it for ${index}?`, @@ -587,7 +585,7 @@ test("lexical ranking ignores English question words that crowd out countable ev }); test("recall cues expose a compressed node directory without memory statements", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "The user's detailed editing preference marker is CERULEAN-OWL", nodeName: "video editing preference", @@ -618,7 +616,7 @@ test("recall cues expose a compressed node directory without memory statements", }); test("resident kernel is query independent and excludes unverified assistant constraints", () => { - withStore((store) => { + withMemoryStore((store) => { const pinned = store.remember({ statement: "Never deploy Project Helix without tests", nodeName: "Project Helix release constraint", @@ -644,7 +642,7 @@ test("resident kernel is query independent and excludes unverified assistant con }); test("node merge preserves memories, evidence, relations, and redirects", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.remember({ statement: "Uses TypeScript", nodeName: "NMG language" }); const second = store.remember({ statement: "Uses SQLite", nodeName: "NMG database" }); const host = store.remember({ statement: "Pi hosts NMG", nodeName: "Pi host" }); @@ -683,7 +681,7 @@ test("node merge preserves memories, evidence, relations, and redirects", () => }); test("node split requires a complete partition and preserves every memory", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.remember({ statement: "Python 2 for ROS", nodeName: "Python environment" }); const second = store.remember({ statement: "Python 3.12 for Windows", @@ -722,7 +720,7 @@ test("vector retrieval finds a semantic synonym without lexical overlap", () => return [0, 0, 1]; }, }; - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "The automobile needs a service", nodeName: "vehicle maintenance", @@ -735,7 +733,7 @@ test("vector retrieval finds a semantic synonym without lexical overlap", () => }); test("learning router changes node ranking from explicit feedback", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.remember({ statement: "One", nodeName: "first project" }); store.remember({ statement: "Two", nodeName: "second project" }); assert.deepEqual(store.routeNodes("alpha signal"), []); @@ -811,7 +809,7 @@ test("embeddings persist as Float32 blobs and a warm node cache accepts appends" }); test("embedding index health tracks pending, success, and retryable failure", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.remember({ statement: "Alpha memory", nodeName: "Alpha node" }); store.beginEmbeddingIndex({ indexId: "provider@index-a", @@ -855,7 +853,7 @@ test("embedding index health tracks pending, success, and retryable failure", () }); test("Huffman-like block rebalance promotes frequent memory only in batches", () => { - withStore((store) => { + withMemoryStore((store) => { const memories = Array.from({ length: 5 }, (_, index) => store.remember({ statement: `Shared topic memory ${index}`, @@ -880,7 +878,7 @@ test("Huffman-like block rebalance promotes frequent memory only in batches", () }); test("appends source messages idempotently within a session", () => { - withStore((store) => { + withMemoryStore((store) => { const input = { content: "I prefer concise answers.", role: "user" as const, @@ -901,7 +899,7 @@ test("appends source messages idempotently within a session", () => { }); test("binds a semantic memory to an existing exact history message", () => { - withStore((store) => { + withMemoryStore((store) => { const history = store.appendHistory({ content: "For ROS Melodic keep Python 2, and do not upgrade this project to Python 3.", role: "user", @@ -922,7 +920,7 @@ test("binds a semantic memory to an existing exact history message", () => { }); test("FTS5 reaches exact cold evidence outside the hot candidate window", () => { - withStore((store) => { + withMemoryStore((store) => { for (let index = 0; index < 520; index += 1) { store.remember({ statement: `Routine note ${index}`, @@ -967,7 +965,7 @@ test("FTS5 reaches exact cold evidence outside the hot candidate window", () => }); test("node vectors route first and node-local search recovers leaf evidence", () => { - withStore((store) => { + withMemoryStore((store) => { const robotics = store.remember({ statement: "Gazebo started after switching to software rendering", nodeName: "ROS rendering", @@ -1006,7 +1004,7 @@ test("node vectors route first and node-local search recovers leaf evidence", () }); test("node creation automatically reuses punctuation and case variants", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.upsertNode({ canonicalName: "ROS Melodic", kind: "project" }); const variant = store.upsertNode({ canonicalName: "ros-melodic", kind: "project" }); assert.equal(variant.id, first.id); @@ -1014,7 +1012,7 @@ test("node creation automatically reuses punctuation and case variants", () => { }); test("node vector routing is deterministic unless hierarchical activation is explicit", () => { - withStore((store) => { + withMemoryStore((store) => { const alpha = store.remember({ statement: "Alpha memory", nodeName: "Alpha node" }); const beta = store.remember({ statement: "Beta memory", nodeName: "Beta node" }); store.upsertExternalNodeEmbeddings("stable-routing", [ @@ -1044,7 +1042,7 @@ test("node vector routing is deterministic unless hierarchical activation is exp }); test("leaf summaries preserve distinctions hidden by one broad node summary", () => { - withStore((store) => { + withMemoryStore((store) => { const rendering = store.remember({ statement: "Gazebo recovered after enabling software rendering", nodeName: "ROS project", @@ -1169,7 +1167,7 @@ test("uncompacted Delta survives restart and participates in hierarchy search", }); test("due leaf rebuild compacts only nodes that cross the Delta threshold", () => { - withStore((store) => { + withMemoryStore((store) => { const first = store.remember({ statement: "Alpha fact one", nodeName: "Alpha" }); store.remember({ statement: "Beta fact one", nodeName: "Beta" }); assert.equal(store.rebuildDueLeafBlocks({ deltaThreshold: 2 }).length, 0); @@ -1187,7 +1185,7 @@ test("due leaf rebuild compacts only nodes that cross the Delta threshold", () = }); test("co-retrieval produces delayed link proposals with evidence and cooldown", () => { - withStore((store) => { + withMemoryStore((store) => { const alpha = store.remember({ statement: "Alpha detail", nodeName: "Alpha" }); const beta = store.remember({ statement: "Beta detail", nodeName: "Beta" }); for (let index = 0; index < 3; index += 1) { @@ -1228,7 +1226,7 @@ test("co-retrieval produces delayed link proposals with evidence and cooldown", }); test("repeated ambiguity proposes an evidence-preserving scoped split", () => { - withStore((store) => { + withMemoryStore((store) => { const python = store.remember({ statement: "ROS uses Python 2", nodeName: "Broad project node", @@ -1323,7 +1321,7 @@ test("topology proposals persist review decisions and ignore weak signals", () = }); test("STG and LTG lifecycle preserves IDs, evidence, expiry, and audit history", () => { - withStore((store) => { + withMemoryStore((store) => { const durable = store.remember({ statement: "Project Helix uses SQLite", nodeName: "Project Helix storage", @@ -1368,7 +1366,7 @@ test("STG and LTG lifecycle preserves IDs, evidence, expiry, and audit history", }); test("Active Graph enforces a shared budget and records verified evidence outcomes", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "ORCHID alpha detail", nodeName: "ORCHID alpha", @@ -1433,7 +1431,7 @@ test("Active Graph enforces a shared budget and records verified evidence outcom }); test("Active Graph records graph expansion paths separately from selected memories", () => { - withStore((store) => { + withMemoryStore((store) => { const seed = store.remember({ statement: "LANTERN seed detail", nodeName: "LANTERN seed" }); const related = store.remember({ statement: "Related implementation uses SQLite", @@ -1465,7 +1463,7 @@ test("Active Graph records graph expansion paths separately from selected memori }); test("edge stability deduplicates tasks and drives auditable reversible consolidation", () => { - withStore((store) => { + withMemoryStore((store) => { const alpha = store.remember({ statement: "Alpha stable evidence", nodeName: "Alpha stable" }); const beta = store.remember({ statement: "Beta stable evidence", nodeName: "Beta stable" }); const observe = (taskId: string, contradicted = false) => { @@ -1561,7 +1559,7 @@ test("P3 schema migrates an existing pre-lifecycle database before creating new }); test("contradictionNotes flags claim pairs with opposite polarity in temporal order", () => { - withStore((store) => { + withMemoryStore((store) => { const earlier = store.remember({ statement: "user: I'm trying to implement the basic homepage route with Flask", nodeName: "beam", @@ -1723,7 +1721,7 @@ test("contradictionNotes flags claim pairs with opposite polarity in temporal or }); test("search defaults to legacy mode: lexical fallback works without matching vectors", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "legacy default mode probe statement", nodeName: "probe node", @@ -1745,7 +1743,7 @@ test("search defaults to legacy mode: lexical fallback works without matching ve }); test("searchContext falls back to lexical when semantic retrieval is empty", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "The retired project codename was silver heron", nodeName: "project codename", diff --git a/tests/core/store/maintenance.test.ts b/tests/core/store/maintenance.test.ts index 40b0534e..fdf143bb 100644 --- a/tests/core/store/maintenance.test.ts +++ b/tests/core/store/maintenance.test.ts @@ -1,24 +1,19 @@ import assert from "node:assert/strict"; -import { mkdtempSync, rmSync } from "node:fs"; -import { tmpdir } from "node:os"; -import { join } from "node:path"; import test from "node:test"; import { NmgStore } from "../../../src/core/store.ts"; -function withStore(run: (store: NmgStore) => void): void { - const directory = mkdtempSync(join(tmpdir(), "nmg-maintenance-")); - const store = new NmgStore(join(directory, "test.sqlite")); +function withMemoryStore(run: (store: NmgStore) => void): void { + const store = new NmgStore(":memory:"); try { run(store); } finally { store.close(); - rmSync(directory, { force: true, recursive: true }); } } test("maintenance proposals enforce attribution, evaluation, and explicit review", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "The deployment region is probably Europe", nodeName: "deployment region", @@ -62,7 +57,7 @@ test("maintenance proposals enforce attribution, evaluation, and explicit review }); test("maintenance proposals do not turn retrieval defects into content mutations", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "The user prefers concise explanations", nodeName: "response preferences", @@ -108,7 +103,7 @@ test("maintenance proposals do not turn retrieval defects into content mutations // ── deleteMemory cascaded deletion ── test("deleteMemory: marks record as deleted and returns pre-deletion snapshot", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "user prefers dark mode", nodeName: "user preferences", @@ -123,7 +118,7 @@ test("deleteMemory: marks record as deleted and returns pre-deletion snapshot", }); test("deleteMemory: cleans up FTS, embedding, and leaf membership", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "Atlas uses SQLite for storage", nodeName: "Atlas storage", @@ -140,13 +135,13 @@ test("deleteMemory: cleans up FTS, embedding, and leaf membership", () => { }); test("deleteMemory: returns null for unknown ids", () => { - withStore((store) => { + withMemoryStore((store) => { assert.equal(store.deleteMemory("does-not-exist"), null); }); }); test("deleteMemory: cascade-deletes derived memories with no remaining sources", () => { - withStore((store) => { + withMemoryStore((store) => { const source = store.remember({ statement: "user needs a standing desk", nodeName: "user equipment", @@ -182,7 +177,7 @@ test("deleteMemory: cascade-deletes derived memories with no remaining sources", }); test("deleteMemory: deleted memories are filtered from getContext", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "user is vegan", nodeName: "user diet", @@ -196,7 +191,7 @@ test("deleteMemory: deleted memories are filtered from getContext", () => { }); test("deleteMemory: scrubs Active Graph references and rejects dependent pending proposals", () => { - withStore((store) => { + withMemoryStore((store) => { const target = store.remember({ statement: "Atlas stores an index in SQLite", nodeName: "Atlas index", @@ -239,7 +234,7 @@ test("deleteMemory: scrubs Active Graph references and rejects dependent pending // ── promoteMemory / demoteMemory tier migration ── test("promoteMemory: promotes STG memory to LTG", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "provisional hypothesis", nodeName: "test hypothesis", @@ -256,7 +251,7 @@ test("promoteMemory: promotes STG memory to LTG", () => { }); test("promoteMemory: already-LTG memory is returned unchanged", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "long-term project name is Atlas", nodeName: "project name", @@ -271,7 +266,7 @@ test("promoteMemory: already-LTG memory is returned unchanged", () => { }); test("demoteMemory: demotes LTG memory to STG", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "permanent project name is Atlas", nodeName: "project name", @@ -286,7 +281,7 @@ test("demoteMemory: demotes LTG memory to STG", () => { }); test("demoteMemory: already-STG memory is returned unchanged", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "temp hypothesis", nodeName: "hypothesis", @@ -303,7 +298,7 @@ test("demoteMemory: already-STG memory is returned unchanged", () => { // ── retentionCandidates ── test("retentionCandidates: reports low-importance event for dormant after aging", () => { - withStore((store) => { + withMemoryStore((store) => { const old = store.remember({ statement: "A disposable historical experiment was attempted", nodeName: "historical experiment", @@ -325,7 +320,7 @@ test("retentionCandidates: reports low-importance event for dormant after aging" }); test("retentionCandidates: constraint type is excluded from candidates", () => { - withStore((store) => { + withMemoryStore((store) => { const constraint = store.remember({ statement: "Never erase the production database", nodeName: "production safety", @@ -346,7 +341,7 @@ test("retentionCandidates: constraint type is excluded from candidates", () => { }); test("retentionCandidates: dormant progresses to quarantine", () => { - withStore((store) => { + withMemoryStore((store) => { const old = store.remember({ statement: "old experiment log entry", nodeName: "experiment log", @@ -366,7 +361,7 @@ test("retentionCandidates: dormant progresses to quarantine", () => { }); test("retentionCandidates: returns empty when no memory qualifies", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "fresh constraint", nodeName: "fresh node", @@ -384,7 +379,7 @@ test("retentionCandidates: returns empty when no memory qualifies", () => { // ── pruneRetrievalTraces ── test("pruneRetrievalTraces: prunes traces beyond maxRows", () => { - withStore((store) => { + withMemoryStore((store) => { for (let i = 0; i < 5; i += 1) { store.recordRetrievalTrace({ query: `query ${i}`, @@ -402,7 +397,7 @@ test("pruneRetrievalTraces: prunes traces beyond maxRows", () => { }); test("pruneRetrievalTraces: no-op when under limits", () => { - withStore((store) => { + withMemoryStore((store) => { store.recordRetrievalTrace({ query: "single query", resultMemoryIds: [], @@ -418,7 +413,7 @@ test("pruneRetrievalTraces: no-op when under limits", () => { // ── retrievalTrace / retrievalTracesCount ── test("retrievalTrace: reads back a recorded trace", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "trace test memory", nodeName: "trace node", @@ -439,7 +434,7 @@ test("retrievalTrace: reads back a recorded trace", () => { }); test("retrievalTrace: returns null for unknown id", () => { - withStore((store) => { + withMemoryStore((store) => { assert.equal(store.retrievalTrace("non-existent-trace-id"), null); }); }); @@ -447,14 +442,14 @@ test("retrievalTrace: returns null for unknown id", () => { // ── perfAggregates ── test("perfAggregates: returns empty array on fresh store", () => { - withStore((store) => { + withMemoryStore((store) => { const aggregates = store.perfAggregates(); assert.deepEqual(aggregates, []); }); }); test("perfAggregates: hot path records search.direct and trace after a searchContext", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "hot path perf probe", nodeName: "hot path", @@ -478,7 +473,7 @@ test("perfAggregates: hot path records search.direct and trace after a searchCon // ── recordActiveGraphAttribution ── test("recordActiveGraphAttribution: records usage on a trace", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "active graph test memory", nodeName: "ag node", @@ -501,7 +496,7 @@ test("recordActiveGraphAttribution: records usage on a trace", () => { }); test("retrieval traces and Active Graph feedback enforce session ownership", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "session-owned active graph memory", nodeName: "owned ag node", @@ -539,7 +534,7 @@ test("retrieval traces and Active Graph feedback enforce session ownership", () }); test("recordActiveGraphAttribution: survives non-existent active graph id gracefully", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "bad graph test", nodeName: "bad graph node", @@ -559,7 +554,7 @@ test("recordActiveGraphAttribution: survives non-existent active graph id gracef // ── memoryWriteEvents ── test("memoryWriteEvents: returns write events for a memory", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "write event test", nodeName: "write event node", @@ -574,7 +569,7 @@ test("memoryWriteEvents: returns write events for a memory", () => { }); test("memoryWriteEvents: returns all events when no id is specified", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "event test 1", nodeName: "event node 1", @@ -595,7 +590,7 @@ test("memoryWriteEvents: returns all events when no id is specified", () => { // ── setMemoryStorageState ── test("setMemoryStorageState: moves to dormant and back to indexed", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "Atlas project uses WAL mode", nodeName: "Atlas WAL", @@ -609,7 +604,7 @@ test("setMemoryStorageState: moves to dormant and back to indexed", () => { }); test("setMemoryStorageState: rejects STG memories", () => { - withStore((store) => { + withMemoryStore((store) => { const provisional = store.remember({ statement: "temp hypothesis for this session", nodeName: "session hypothesis", @@ -625,7 +620,7 @@ test("setMemoryStorageState: rejects STG memories", () => { }); test("setMemoryStorageState: throws for non-existent memory", () => { - withStore((store) => { + withMemoryStore((store) => { assert.throws(() => store.setMemoryStorageState("does-not-exist", "dormant"), /does not exist/); }); }); @@ -633,7 +628,7 @@ test("setMemoryStorageState: throws for non-existent memory", () => { // ── expireShortTermMemories ── test("expireShortTermMemories: expires past-due STG memories", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "expiring hypothesis", nodeName: "expiring node", @@ -650,7 +645,7 @@ test("expireShortTermMemories: expires past-due STG memories", () => { // ── upsertNode ── test("upsertNode: creates a new node and returns existing on duplicate", () => { - withStore((store) => { + withMemoryStore((store) => { const node1 = store.upsertNode({ canonicalName: "unique test node", kind: "concept", @@ -669,7 +664,7 @@ test("upsertNode: creates a new node and returns existing on duplicate", () => { // ── rebuildVectorIndex / rebalanceNode / rebuildLeafBlocks ── test("rebuildVectorIndex: rebuilds embeddings for indexed memories", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "vector index test", nodeName: "vector node", @@ -682,7 +677,7 @@ test("rebuildVectorIndex: rebuilds embeddings for indexed memories", () => { }); test("rebalanceNode: returns result with changedMemoryIds and expectedDepth", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "rebalance test memory", nodeName: "rebalance node", @@ -697,7 +692,7 @@ test("rebalanceNode: returns result with changedMemoryIds and expectedDepth", () }); test("rebalanceDueNodes: returns results for nodes above threshold", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "rebalance due test", nodeName: "rebalance due node", @@ -710,7 +705,7 @@ test("rebalanceDueNodes: returns results for nodes above threshold", () => { }); test("runDueMaintenance compacts write deltas and rebalances access changes within a node budget", () => { - withStore((store) => { + withMemoryStore((store) => { const writeDue = store.remember({ statement: "bounded maintenance write", nodeName: "bounded write node", @@ -752,7 +747,7 @@ test("runDueMaintenance compacts write deltas and rebalances access changes with }); test("answer-overlap attribution is diagnostic and cannot train graph state", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "API answers may repeat retrieved wording", nodeName: "API attribution boundary", @@ -781,7 +776,7 @@ test("answer-overlap attribution is diagnostic and cannot train graph state", () }); test("runDueMaintenance drains distributed write pressure without requiring one hot node", () => { - withStore((store) => { + withMemoryStore((store) => { for (let index = 0; index < 16; index += 1) { store.remember({ statement: `distributed write ${index}`, @@ -808,7 +803,7 @@ test("runDueMaintenance drains distributed write pressure without requiring one }); test("runDueMaintenance drains distributed access pressure without requiring one hot node", () => { - withStore((store) => { + withMemoryStore((store) => { const memories = Array.from({ length: 32 }, (_, index) => store.remember({ statement: `distributed access ${index}`, @@ -836,7 +831,7 @@ test("runDueMaintenance drains distributed access pressure without requiring one }); test("runSemanticMaintenance expires STG records and keeps topology changes proposal-only", () => { - withStore((store) => { + withMemoryStore((store) => { const expired = store.remember({ statement: "expired session-local observation", nodeName: "expired session node", @@ -860,7 +855,7 @@ test("runSemanticMaintenance expires STG records and keeps topology changes prop }); test("rebuildLeafBlocks: creates leaf blocks for a node", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "leaf block test memory", nodeName: "leaf block node", @@ -875,7 +870,7 @@ test("rebuildLeafBlocks: creates leaf blocks for a node", () => { }); test("rebuildLeafBlocks: rebuilds all nodes when no nodeId is given", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "all leaf blocks test 1", nodeName: "leaf all node 1", @@ -896,7 +891,7 @@ test("rebuildLeafBlocks: rebuilds all nodes when no nodeId is given", () => { // ── dirtyLeafNodeIds / pendingIndexDelta / acknowledgeIndexDelta ── test("dirtyLeafNodeIds: returns dirty node ids after delete", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "dirty leaf test", nodeName: "dirty leaf node", @@ -912,7 +907,7 @@ test("dirtyLeafNodeIds: returns dirty node ids after delete", () => { }); test("pendingIndexDelta: returns memory ids with pending changes", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "pending delta test", nodeName: "pending delta node", @@ -925,7 +920,7 @@ test("pendingIndexDelta: returns memory ids with pending changes", () => { }); test("acknowledgeIndexDelta: clears compacted deltas", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "ack delta test", nodeName: "ack delta node", @@ -942,7 +937,7 @@ test("acknowledgeIndexDelta: clears compacted deltas", () => { // ── embedding index lifecycle ── test("beginEmbeddingIndex / completeEmbeddingIndex / embeddingIndexHealth lifecycle", () => { - withStore((store) => { + withMemoryStore((store) => { store.beginEmbeddingIndex({ indexId: "test-index-1", model: "test-model", @@ -959,7 +954,7 @@ test("beginEmbeddingIndex / completeEmbeddingIndex / embeddingIndexHealth lifecy }); test("failEmbeddingIndex: marks index as failed", () => { - withStore((store) => { + withMemoryStore((store) => { store.beginEmbeddingIndex({ indexId: "test-index-fail", model: "test-model", @@ -976,14 +971,14 @@ test("failEmbeddingIndex: marks index as failed", () => { // ── contradictionNotes ── test("contradictionNotes: returns empty map for empty input", () => { - withStore((store) => { + withMemoryStore((store) => { const notes = store.contradictionNotes([]); assert.equal(notes.size, 0); }); }); test("contradictionNotes: returns empty map for memories without claims", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "no claims memory", nodeName: "no claims node", @@ -998,7 +993,7 @@ test("contradictionNotes: returns empty map for memories without claims", () => // ── rebuildDueLeafBlocks ── test("rebuildDueLeafBlocks: rebuilds blocks for nodes with pending deltas", () => { - withStore((store) => { + withMemoryStore((store) => { store.remember({ statement: "due leaf blocks test", nodeName: "due leaf node", @@ -1013,7 +1008,7 @@ test("rebuildDueLeafBlocks: rebuilds blocks for nodes with pending deltas", () = // ── recordConsolidationEvent visibility check ── test("promoteMemory triggers a consolidation event visible via consolidationEvents", () => { - withStore((store) => { + withMemoryStore((store) => { const saved = store.remember({ statement: "consolidation event test", nodeName: "consolidation node", @@ -1033,7 +1028,7 @@ test("promoteMemory triggers a consolidation event visible via consolidationEven // ── deferred retrieval-pair signals (drainPendingTraceSignals) ── test("deferred pair signals materialize on drain and stay idempotent", () => { - withStore((store) => { + withMemoryStore((store) => { const alpha = store.remember({ statement: "deferred alpha", nodeName: "defer alpha" }); const beta = store.remember({ statement: "deferred beta", nodeName: "defer beta" }); for (let index = 0; index < 3; index += 1) { @@ -1067,7 +1062,7 @@ test("deferred pair signals materialize on drain and stay idempotent", () => { }); test("feedback before drain still records observations and drives consolidation", () => { - withStore((store) => { + withMemoryStore((store) => { const alpha = store.remember({ statement: "fb alpha", nodeName: "fb alpha" }); const beta = store.remember({ statement: "fb beta", nodeName: "fb beta" }); const traceId = store.recordRetrievalTrace({ @@ -1092,7 +1087,7 @@ test("feedback before drain still records observations and drives consolidation" }); test("drainPendingTraceSignals is bounded by limit and marks processed traces", () => { - withStore((store) => { + withMemoryStore((store) => { const alpha = store.remember({ statement: "bounded alpha", nodeName: "bounded alpha" }); const beta = store.remember({ statement: "bounded beta", nodeName: "bounded beta" }); for (let index = 0; index < 5; index += 1) { @@ -1111,7 +1106,7 @@ test("drainPendingTraceSignals is bounded by limit and marks processed traces", }); test("runSemanticMaintenance drains deferred pair signals before proposing", () => { - withStore((store) => { + withMemoryStore((store) => { const alpha = store.remember({ statement: "sm alpha", nodeName: "sm alpha" }); const beta = store.remember({ statement: "sm beta", nodeName: "sm beta" }); for (let index = 0; index < 3; index += 1) { diff --git a/tests/rcp/fixture.ts b/tests/rcp/fixture.ts index 3da0dd52..665a76f7 100644 --- a/tests/rcp/fixture.ts +++ b/tests/rcp/fixture.ts @@ -3,7 +3,7 @@ import { mkdirSync, mkdtempSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; import { join } from "node:path"; -export function repositoryFixture(): string { +export function repositoryFixture(options: { git?: boolean } = {}): string { const root = mkdtempSync(join(tmpdir(), "nmg-rcp-")); const files: Record = { "package.json": JSON.stringify({ @@ -36,11 +36,20 @@ export function repositoryFixture(): string { mkdirSync(join(path, ".."), { recursive: true }); writeFileSync(path, content); } - git(root, ["init", "--quiet"]); - git(root, ["config", "user.email", "rcp@example.invalid"]); - git(root, ["config", "user.name", "RCP Test"]); - git(root, ["add", "."]); - git(root, ["commit", "--quiet", "-m", "fixture"]); + if (options.git !== false) { + git(root, ["init", "--quiet"]); + git(root, ["add", "."]); + git(root, [ + "-c", + "user.email=rcp@example.invalid", + "-c", + "user.name=RCP Test", + "commit", + "--quiet", + "-m", + "fixture", + ]); + } return root; } diff --git a/tests/rcp/reconcile.test.ts b/tests/rcp/reconcile.test.ts index 195f87ef..a4bfa01a 100644 --- a/tests/rcp/reconcile.test.ts +++ b/tests/rcp/reconcile.test.ts @@ -23,19 +23,48 @@ import type { RepositoryProvider } from "../../src/rcp/repository.ts"; import type { VerifierProvider } from "../../src/rcp/providers.ts"; import { contractText, repositoryFixture } from "./fixture.ts"; -function setup() { - const root = repositoryFixture(); +function fixedRepository(root: string): RepositoryProvider { + return { + descriptor: { + id: "fixed-test-repository", + version: "1", + capabilities: ["fixed-observation"], + operations: ["observe"], + authority: ["plan", "apply", "continuous"], + }, + observe: async (request) => { + assert.equal(request.root, root); + return { + root: root.replaceAll("\\", "/"), + observedRevision: `sha256:${"0".repeat(64)}`, + git: { available: true, commit: "fixed-test-commit", dirtyFiles: [] }, + files: [], + diagnostics: [], + observedBytes: 0, + }; + }, + }; +} + +function setup(resource: "git" | "observation" = "git") { + const root = repositoryFixture({ git: resource === "git" }); const compiled = compileContract({ text: contractText(), path: join(root, "contract.yaml") }); assert.ok(compiled.contract); - return { root, contract: compiled.contract, routes: readRouteDeclarations(root) }; + return { + root, + contract: compiled.contract, + routes: readRouteDeclarations(root), + repository: resource === "git" ? undefined : fixedRepository(root), + }; } function providers( root: string, harness: HarnessProvider = new ExternalWorkspaceHarnessProvider(), + repository: RepositoryProvider = new LocalRepositoryProvider(), ) { return { - repository: new LocalRepositoryProvider(), + repository, policy: new DefaultPolicyProvider(), harness, verifier: new LocalNpmVerifierProvider(30_000), @@ -44,10 +73,10 @@ function providers( } test("plan mode emits a WorkOrder without executing or recording", async () => { - const value = setup(); + const value = setup("observation"); const result = await reconcileOnce( { ...value, requestedMode: "plan", invocationId: "plan" }, - providers(value.root), + providers(value.root, undefined, value.repository), ); assert.equal(result.status, "planned"); assert.equal(result.receipt, undefined); @@ -85,10 +114,10 @@ test("apply independently verifies, records an immutable receipt, and reuses ide }); test("tampered verified receipt is rejected instead of reused", async () => { - const value = setup(); + const value = setup("observation"); const first = await reconcileOnce( { ...value, requestedMode: "apply", invocationId: "tamper-1", now: fixedClock() }, - providers(value.root), + providers(value.root, undefined, value.repository), ); assert.equal(first.status, "verified"); const tampered = JSON.parse(readFileSync(first.receiptPath!, "utf8")) as Record; @@ -97,15 +126,15 @@ test("tampered verified receipt is rejected instead of reused", async () => { const second = await reconcileOnce( { ...value, requestedMode: "apply", invocationId: "tamper-2" }, - providers(value.root), + providers(value.root, undefined, value.repository), ); assert.equal(second.status, "blocked"); assert.match(second.conditions.at(-1)?.reason ?? "", /invalid receipt/i); }); test("receipt reuse is bound to the current verifier definition", async () => { - const value = setup(); - const base = providers(value.root); + const value = setup("observation"); + const base = providers(value.root, undefined, value.repository); const first = await reconcileOnce( { ...value, requestedMode: "apply", invocationId: "verifier-v1", now: fixedClock() }, base, @@ -235,8 +264,8 @@ test("an implementing harness cannot weaken its verifier during execution", asyn }); test("optional NMG failure degrades without changing authority or verified decision", async () => { - const value = setup(); - const base = providers(value.root); + const value = setup("observation"); + const base = providers(value.root, undefined, value.repository); const result = await reconcileOnce( { ...value, requestedMode: "apply", invocationId: "memory", now: fixedClock() }, { @@ -256,10 +285,10 @@ test("optional NMG failure degrades without changing authority or verified decis }); test("external and process harnesses consume the same WorkOrder contract", async () => { - const value = setup(); + const value = setup("observation"); const external = await reconcileOnce( { ...value, operationKey: "external", requestedMode: "plan" }, - providers(value.root, new ExternalWorkspaceHarnessProvider("codex")), + providers(value.root, new ExternalWorkspaceHarnessProvider("codex"), value.repository), ); const script = join(value.root, "harness.mjs"); writeFileSync( @@ -275,7 +304,7 @@ test("external and process harnesses consume the same WorkOrder contract", async }); test("process harness terminates after its configured timeout", async () => { - const value = setup(); + const value = setup("observation"); const plan = await reconcileOnce( { ...value, @@ -283,7 +312,7 @@ test("process harness terminates after its configured timeout", async () => { requestedMode: "plan", executionTimeoutMs: 50, }, - providers(value.root), + providers(value.root, undefined, value.repository), ); const script = join(value.root, "hanging-harness.mjs"); writeFileSync(script, "process.stdin.resume(); setInterval(() => {}, 1000);\n"); @@ -299,7 +328,7 @@ test("process harness terminates after its configured timeout", async () => { }); test("harness exceptions become failed terminal receipts", async () => { - const value = setup(); + const value = setup("observation"); const harness: HarnessProvider = { descriptor: new ExternalWorkspaceHarnessProvider("throwing-harness").descriptor, execute: async () => { @@ -308,7 +337,7 @@ test("harness exceptions become failed terminal receipts", async () => { }; const result = await reconcileOnce( { ...value, requestedMode: "apply", invocationId: "throwing", now: fixedClock() }, - providers(value.root, harness), + providers(value.root, harness, value.repository), ); assert.equal(result.status, "failed"); assert.equal(result.receipt?.decision, "failed"); @@ -317,7 +346,7 @@ test("harness exceptions become failed terminal receipts", async () => { }); test("failed attempts remain append-only without preventing a later retry", async () => { - const value = setup(); + const value = setup("observation"); const failingHarness: HarnessProvider = { descriptor: new ExternalWorkspaceHarnessProvider("first-attempt").descriptor, execute: async () => ({ @@ -328,12 +357,12 @@ test("failed attempts remain append-only without preventing a later retry", asyn }; const first = await reconcileOnce( { ...value, requestedMode: "apply", invocationId: "attempt-1", now: fixedClock() }, - providers(value.root, failingHarness), + providers(value.root, failingHarness, value.repository), ); assert.equal(first.status, "failed"); const second = await reconcileOnce( { ...value, requestedMode: "apply", invocationId: "attempt-2", now: fixedClock() }, - providers(value.root), + providers(value.root, undefined, value.repository), ); assert.equal(second.status, "verified"); assert.equal(receiptFileCount(value.root), 2); diff --git a/tests/scripts/verification-impact-probe.test.ts b/tests/scripts/verification-impact-probe.test.ts new file mode 100644 index 00000000..bac52486 --- /dev/null +++ b/tests/scripts/verification-impact-probe.test.ts @@ -0,0 +1,102 @@ +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import { mkdirSync, mkdtempSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import test from "node:test"; +import { fileURLToPath } from "node:url"; + +import { planImpact } from "../../scripts/verification-impact-probe.ts"; + +function fixture(): string { + const root = mkdtempSync(join(tmpdir(), "nmg-impact-probe-")); + for (const path of ["src", "tests"]) mkdirSync(join(root, path), { recursive: true }); + writeFileSync( + join(root, "package.json"), + JSON.stringify({ scripts: { "test:product": 'node --test "tests/**/*.test.ts"' } }), + ); + writeFileSync(join(root, "package-lock.json"), "{}\n"); + writeFileSync(join(root, "tsconfig.json"), "{}\n"); + writeFileSync(join(root, "src", "pure.ts"), "export const value = 1;\n"); + writeFileSync( + join(root, "src", "consumer.ts"), + 'import { value } from "./pure.ts";\nexport { value };\n', + ); + writeFileSync(join(root, "tests", "direct.test.ts"), 'import "../src/pure.ts";\n'); + writeFileSync(join(root, "tests", "consumer.test.ts"), 'import "../src/consumer.ts";\n'); + writeFileSync( + join(root, "tests", "unrelated.test.ts"), + 'import test from "node:test";\ntest("unrelated", () => {});\n', + ); + const git = spawnSync("git", ["init", "--quiet"], { + cwd: root, + encoding: "utf8", + windowsHide: true, + }); + assert.equal(git.status, 0, git.stderr); + return root; +} + +test("impact probe follows direct and transitive imports without calling the subset sufficient", () => { + const root = fixture(); + const plan = planImpact(root, "src/pure.ts"); + assert.deepEqual(plan.candidateTests, ["tests/consumer.test.ts", "tests/direct.test.ts"]); + assert.equal(plan.productTests, 3); + assert.equal(plan.fullGateRequired, true); + assert.equal(plan.verdict, "candidate-only"); +}); + +test("impact probe includes literal dynamic imports and reports computed paths", () => { + const root = fixture(); + writeFileSync(join(root, "tests", "dynamic.test.ts"), 'await import("../src/pure.ts");\n'); + writeFileSync( + join(root, "tests", "computed.test.ts"), + 'const path = "../src/pure.ts";\nawait import(path);\n', + ); + const plan = planImpact(root, "src/pure.ts"); + assert.ok(plan.candidateTests.includes("tests/dynamic.test.ts")); + assert.ok(plan.unresolvedImports.some((item) => item.includes("tests/computed.test.ts"))); + assert.equal(plan.fullGateRequired, true); +}); + +test("impact probe exposes a file-read counterexample instead of claiming coverage", () => { + const root = fixture(); + writeFileSync( + join(root, "tests", "text.test.ts"), + 'import { readFileSync } from "node:fs";\nreadFileSync("src/pure.ts", "utf8");\n', + ); + const plan = planImpact(root, "src/pure.ts"); + assert.equal(plan.candidateTests.includes("tests/text.test.ts"), false); + assert.equal(plan.fullGateRequired, true); + assert.ok(plan.limitations.some((item) => item.includes("Filesystem reads"))); +}); + +test("impact probe changes its snapshot key when an input changes", () => { + const root = fixture(); + const before = planImpact(root, "src/pure.ts"); + writeFileSync(join(root, "src", "pure.ts"), "export const value = 2;\n"); + const after = planImpact(root, "src/pure.ts"); + assert.notEqual(before.snapshotDigest, after.snapshotDigest); + assert.deepEqual(before.candidateTests, after.candidateTests); +}); + +test("impact probe refuses an absent source instead of treating zero matches as clean", () => { + const root = fixture(); + assert.throws(() => planImpact(root, "src/missing.ts"), /not a present repository file/); +}); + +test("CLI reports candidates with a nonpassing exit status", () => { + const root = fixture(); + const script = fileURLToPath( + new URL("../../scripts/verification-impact-probe.ts", import.meta.url), + ); + const run = spawnSync( + process.execPath, + ["--experimental-strip-types", script, "--root", root, "--source", "src/pure.ts"], + { cwd: root, encoding: "utf8", windowsHide: true }, + ); + assert.equal(run.status, 2, run.stderr); + const output = JSON.parse(run.stdout) as { verdict: string; fullGateRequired: boolean }; + assert.equal(output.verdict, "candidate-only"); + assert.equal(output.fullGateRequired, true); +}); diff --git a/tests/tools/agent-verify.test.ts b/tests/tools/agent-verify.test.ts index f402f12b..52fe93d1 100644 --- a/tests/tools/agent-verify.test.ts +++ b/tests/tools/agent-verify.test.ts @@ -14,6 +14,7 @@ import { type VerificationCommandResult, } from "../../tools/agent-verify.ts"; import type { AgentContextReport } from "../../tools/repo-context.ts"; +import { npmCommandRunner } from "../../src/rcp/verification.ts"; function report(): AgentContextReport { return { @@ -97,6 +98,78 @@ test("default execution runs every blocking check and leaves advisory work expli assert.equal(results.results.at(-1)?.reason, "advisory checks require --include-advisory"); }); +test("parallel static checks preserve barriers, result order, and individual failures", async () => { + const commands = ["build", "check", "lint", "format:check", "test:product"]; + const plan = { + blocking: commands.map((command) => ({ + command, + classification: "blocking" as const, + routes: [], + })), + advisory: [], + }; + const started: string[] = []; + const events: string[] = []; + let active = 0; + let peak = 0; + const result = await executeVerificationPlan(plan, { + parallel: { commands: new Set(["check", "lint", "format:check"]), concurrency: 2 }, + run: async (command, classification, routes) => { + started.push(command); + events.push(`start:${command}`); + active += 1; + peak = Math.max(peak, active); + await new Promise((resolve) => setTimeout(resolve, command === "check" ? 20 : 5)); + active -= 1; + events.push(`end:${command}`); + return { + command, + classification, + routes, + status: command === "lint" ? "failed" : "passed", + durationMs: 1, + }; + }, + }); + assert.equal(peak, 2); + assert.deepEqual(started.slice(0, 3), ["build", "check", "lint"]); + assert.ok(events.indexOf("end:build") < events.indexOf("start:check")); + assert.ok(events.indexOf("end:format:check") < events.indexOf("start:test:product")); + assert.equal(started.at(-1), "test:product"); + assert.deepEqual( + result.results.map(({ command }) => command), + commands, + ); + assert.equal(result.results.find(({ command }) => command === "lint")?.status, "failed"); + assert.equal(result.ok, false); +}); + +test("asynchronous npm checks attribute their own exit status", async () => { + const root = mkdtempSync(join(tmpdir(), "nmg-agent-verify-async-")); + writeFileSync( + join(root, "package.json"), + JSON.stringify({ scripts: { fail: 'node -e "process.exit(7)"' } }), + ); + const run = npmCommandRunner(root, true, 5_000, () => 5_000, new Set(["fail"])); + const result = await run("fail", "blocking", ["fixture"]); + assert.equal(result.status, "failed"); + assert.equal(result.exitCode, 7); + assert.equal(result.errorKind, "exit"); + assert.deepEqual(result.routes, ["fixture"]); +}); + +test("asynchronous npm checks respect the shared remaining budget", async () => { + const root = mkdtempSync(join(tmpdir(), "nmg-agent-verify-async-timeout-")); + writeFileSync( + join(root, "package.json"), + JSON.stringify({ scripts: { slow: 'node -e "setTimeout(() => {}, 250)"' } }), + ); + const run = npmCommandRunner(root, true, 50, () => 50, new Set(["slow"])); + const result = await run("slow", "blocking", ["fixture"]); + assert.equal(result.status, "failed"); + assert.equal(result.errorKind, "timeout"); +}); + test("advisory failures are reported without failing the blocking result", async () => { const results = await executeVerificationPlan(buildVerificationPlan(report()), { includeAdvisory: true, @@ -635,8 +708,81 @@ test("CLI attributes command timeout and persists the failure", () => { assert.equal(payload.results[0]?.errorKind, "timeout"); }); +test("CLI spends one timeout budget across all blocking checks", () => { + const root = mkdtempSync(join(tmpdir(), "nmg-agent-verify-total-timeout-")); + mkdirSync(join(root, "docs"), { recursive: true }); + writeFileSync(join(root, "docs", "owner.md"), "# Owner\n"); + writeFileSync( + join(root, "package.json"), + JSON.stringify({ + name: "fixture", + version: "1.0.0", + scripts: { + slow: 'node -e "setTimeout(() => {}, 10000)"', + after: 'node -e "process.exit(0)"', + }, + }), + ); + writeFileSync( + join(root, "agent-context.yaml"), + "version: 1\nroutes:\n - id: fixture\n paths: [src/**]\n owners: [docs/owner.md]\n tests: []\n verify:\n blocking: [slow, after]\n advisory: []\n", + ); + + const script = fileURLToPath(new URL("../../tools/agent-verify.ts", import.meta.url)); + const result = spawnSync( + process.execPath, + [ + "--experimental-strip-types", + script, + "--root", + root, + "--scope", + "src/file.ts", + "--timeout-ms", + "50", + "--json", + ], + { encoding: "utf8", windowsHide: true, timeout: 5_000 }, + ); + assert.notEqual(result.status, 0, result.stderr || result.stdout); + const payload = JSON.parse(result.stdout) as { results: VerificationCommandResult[] }; + assert.equal(payload.results.length, 2); + assert.equal(payload.results[1]?.status, "failed"); + assert.equal(payload.results[1]?.durationMs, 0); + assert.match(payload.results[1]?.reason ?? "", /overall deadline/); +}); + +test("official verifier entry returns an incomplete failure when the worker exceeds its deadline", () => { + const root = mkdtempSync(join(tmpdir(), "nmg-agent-verify-watchdog-")); + const script = fileURLToPath(new URL("../../tools/agent-verify-watchdog.ts", import.meta.url)); + const result = spawnSync( + process.execPath, + ["--experimental-strip-types", script, "--root", root, "--timeout-ms", "1", "--json"], + { encoding: "utf8", windowsHide: true, timeout: 5_000 }, + ); + assert.notEqual(result.status, 0, result.stderr || result.stdout); + const payload = JSON.parse(result.stdout) as { + ok: boolean; + incomplete: boolean; + evidencePath: string; + }; + assert.equal(payload.ok, false); + assert.equal(payload.incomplete, true); + const evidence = JSON.parse(readFileSync(payload.evidencePath, "utf8")) as { + result: { ok: boolean }; + incomplete: boolean; + }; + assert.equal(evidence.result.ok, false); + assert.equal(evidence.incomplete, true); +}); + function narrowFixture( - options: { failRouteTest?: boolean; routeTests?: string; skipRouteTest?: boolean } = {}, + options: { + failRouteTest?: boolean; + routeTests?: string; + skipRouteTest?: boolean; + omitSharedChecks?: boolean; + } = {}, ): string { const root = mkdtempSync(join(tmpdir(), "nmg-agent-verify-narrow-")); mkdirSync(join(root, "plugin"), { recursive: true }); @@ -666,7 +812,7 @@ function narrowFixture( ); writeFileSync( join(root, "agent-context.yaml"), - `version: 1\nroutes:\n - id: plugin\n paths: [plugin/**]\n owners: []\n tests: [${options.routeTests ?? "tests/plugin/**"}]\n verify:\n blocking: [check, test:product]\n advisory: []\n`, + `version: 1\nroutes:\n - id: plugin\n paths: [plugin/**]\n owners: []\n tests: [${options.routeTests ?? "tests/plugin/**"}]\n verify:\n${options.omitSharedChecks ? " sharedChecks: none\n" : ""} blocking: [check, test:product]\n advisory: []\n`, ); const git = (args: string[]) => spawnSync("git", args, { cwd: root, encoding: "utf8", windowsHide: true }); @@ -729,7 +875,7 @@ test("--full forces the declared blocking set instead of narrowing", () => { }); test("a failing route test fails the narrow gate instead of passing vacuously", () => { - const result = runVerify(narrowFixture({ failRouteTest: true })); + const result = runVerify(narrowFixture({ failRouteTest: true, omitSharedChecks: true })); assert.notEqual(result.status, 0, result.stderr || result.stdout); const payload = JSON.parse(result.stdout) as { rcp?: { status: string; receiptPath?: string }; @@ -740,11 +886,12 @@ test("a failing route test fails the narrow gate instead of passing vacuously", const receipt = JSON.parse(readFileSync(payload.rcp!.receiptPath!, "utf8")) as { checks: Array<{ name: string; status: string }>; }; + assert.deepEqual(receipt.checks.map((check) => check.name), ["node-test:plugin"]); assert.equal(receipt.checks.find((check) => check.name === "node-test:plugin")?.status, "failed"); }); test("route test patterns that match no files fail closed", () => { - const result = runVerify(narrowFixture({ routeTests: "tests/nope/**" })); + const result = runVerify(narrowFixture({ routeTests: "tests/nope/**", omitSharedChecks: true })); assert.notEqual(result.status, 0, result.stderr || result.stdout); const payload = JSON.parse(result.stdout) as { rcp?: { status: string; receiptPath?: string }; @@ -759,7 +906,7 @@ test("route test patterns that match no files fail closed", () => { }); test("route tests that are only skipped do not pass", () => { - const result = runVerify(narrowFixture({ skipRouteTest: true })); + const result = runVerify(narrowFixture({ skipRouteTest: true, omitSharedChecks: true })); assert.notEqual(result.status, 0, result.stderr || result.stdout); const payload = JSON.parse(result.stdout) as { rcp?: { status: string; receiptPath?: string }; @@ -815,7 +962,7 @@ test("a failed check restates its own last lines and names the evidence file", ( }); test("narrow mode restates the failing check's last lines too", () => { - const root = narrowFixture({ failRouteTest: true }); + const root = narrowFixture({ failRouteTest: true, omitSharedChecks: true }); const script = fileURLToPath(new URL("../../tools/agent-verify.ts", import.meta.url)); const result = spawnSync( process.execPath, diff --git a/tests/tools/test-groups.test.ts b/tests/tools/test-groups.test.ts index 9734434d..35f3c69a 100644 --- a/tests/tools/test-groups.test.ts +++ b/tests/tools/test-groups.test.ts @@ -1,6 +1,7 @@ import assert from "node:assert/strict"; import { readFileSync } from "node:fs"; import test from "node:test"; +import { parse as parseYaml } from "yaml"; const packageJson = JSON.parse( readFileSync(new URL("../../package.json", import.meta.url), "utf8"), @@ -61,3 +62,16 @@ test("local and CI verification groups share named package contracts", () => { assert.match(packageJson.scripts["verify:static"], /agent:context:check/); assert.match(packageJson.scripts["verify:product-ci"], /test:coverage/); }); + +test("agent static checks match the CI contract without nested duplicate execution", () => { + const context = parseYaml( + readFileSync(new URL("../../agent-context.yaml", import.meta.url), "utf8"), + ) as { routes: { id: string; verify: { blocking: string[] } }[] }; + const route = context.routes.find(({ id }) => id === "ci-and-tests"); + assert.ok(route); + const staticChecks = [...packageJson.scripts["verify:static"]!.matchAll(/npm run ([\w:-]+)/gu)].map( + (match) => match[1]!, + ); + assert.deepEqual(route.verify.blocking, [...staticChecks, "test:product"]); + assert.equal(new Set(route.verify.blocking).size, route.verify.blocking.length); +}); diff --git a/tools/agent-verify-watchdog.ts b/tools/agent-verify-watchdog.ts new file mode 100644 index 00000000..d6f6c6fa --- /dev/null +++ b/tools/agent-verify-watchdog.ts @@ -0,0 +1,75 @@ +import { randomUUID } from "node:crypto"; +import { spawnSync } from "node:child_process"; +import { mkdirSync, renameSync, writeFileSync } from "node:fs"; +import { dirname, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; + +const args = process.argv.slice(2); +let root = process.cwd(); +let output: string | undefined; +let timeoutMs = 150_000; +let json = false; +for (let index = 0; index < args.length; index += 1) { + if (args[index] === "--root") root = args[++index] ?? root; + else if (args[index] === "--output") output = args[++index]; + else if (args[index] === "--timeout-ms") { + const value = args[++index]; + if (value && /^\d+$/.test(value) && Number.isSafeInteger(Number(value)) && Number(value) > 0) + timeoutMs = Number(value); + } else if (args[index] === "--json") json = true; +} +root = resolve(root); +const evidencePath = output + ? resolve(root, output) + : resolve(root, ".nmg", "verification", "latest.json"); +const startedAt = new Date().toISOString(); +const worker = fileURLToPath(new URL("./agent-verify.ts", import.meta.url)); +const child = spawnSync(process.execPath, ["--experimental-strip-types", worker, ...args], { + encoding: "utf8", + windowsHide: true, + stdio: json ? "pipe" : "inherit", + maxBuffer: 16 * 1024 * 1024, + timeout: timeoutMs, +}); +const timedOut = (child.error as NodeJS.ErrnoException | undefined)?.code === "ETIMEDOUT"; +if (!timedOut) { + if (child.stdout) process.stdout.write(child.stdout); + if (child.stderr) process.stderr.write(child.stderr); + if (child.error) process.stderr.write(`${child.error.message}\n`); + process.exitCode = child.status ?? 1; +} else { + const reason = `verification exceeded ${timeoutMs}ms overall deadline; result incomplete`; + const result = { + ok: false, + results: [ + { + command: "agent:verify", + classification: "blocking", + routes: [], + status: "failed", + durationMs: timeoutMs, + reason, + errorKind: "timeout", + }, + ], + }; + const evidence = { + schemaVersion: 1, + runId: randomUUID(), + startedAt, + finishedAt: new Date().toISOString(), + runtime: { node: process.version, platform: process.platform, arch: process.arch }, + options: { timeoutMs }, + report: {}, + result, + incomplete: true, + }; + mkdirSync(dirname(evidencePath), { recursive: true }); + const temporary = `${evidencePath}.${process.pid}.tmp`; + writeFileSync(temporary, `${JSON.stringify(evidence, null, 2)}\n`, "utf8"); + renameSync(temporary, evidencePath); + if (json) + process.stdout.write(`${JSON.stringify({ ...result, incomplete: true, evidencePath })}\n`); + else process.stderr.write(`${reason}\nEvidence: ${evidencePath}\n`); + process.exitCode = 1; +} diff --git a/tools/agent-verify.ts b/tools/agent-verify.ts index 065f69f5..901dfb17 100644 --- a/tools/agent-verify.ts +++ b/tools/agent-verify.ts @@ -41,6 +41,23 @@ export { } from "../src/rcp/verification.ts"; import { requirePositiveInteger } from "./parts/numbers.ts"; +// These checks do not rewrite root source or dist after build/package:check. +// verify:packages writes its own subpackage, and complexity:gate uses isolated +// probes. The product suite stays behind this batch to avoid sharing those +// resources or observing generated files mid-write. +const PARALLEL_STATIC_CHECKS = new Set([ + "docs:check", + "glossary:check", + "check", + "check:lock", + "lint", + "format:check", + "agent:context:check", + "complexity:gate", + "verify:packages", + "rtm:check", +]); + export function buildVerificationPlan(report: AgentContextReport) { return buildRouteVerificationPlan(report.routes); } @@ -81,7 +98,7 @@ function parseArgs(args: string[]) { let requireClean = false; let narrow = false; let full = false; - let timeoutMs = 30 * 60 * 1_000; + let timeoutMs = 150_000; let output: string | undefined; let help = false; const scopes: string[] = []; @@ -149,6 +166,7 @@ automatically runs its workspace-ready reconciliation and records a receipt. the change is cleanly owned. --full force the declared whole blocking set (no narrowing) --include-advisory run advisory checks in addition to blocking checks + --timeout-ms whole-run deadline (default: 150000ms) --dry-run print and persist the plan without running checks --require-clean reject a dirty Git worktree --root verify another repository root @@ -225,7 +243,13 @@ interface RcpEvidence { async function executeRcpVerification( report: AgentContextReport, contract: RepositoryContractIr, - options: { root: string; timeoutMs: number; includeAdvisory: boolean; json: boolean }, + options: { + root: string; + timeoutMs: number; + remainingMs: () => number; + includeAdvisory: boolean; + json: boolean; + }, ): Promise<{ result: VerificationRunResult; rcp: RcpEvidence }> { const reconciliation = await reconcileOnce( { @@ -240,7 +264,7 @@ async function executeRcpVerification( repository: new LocalRepositoryProvider(), policy: new DefaultPolicyProvider(), harness: new ExternalWorkspaceHarnessProvider(), - verifier: new LocalNpmVerifierProvider(options.timeoutMs, !options.json), + verifier: new LocalNpmVerifierProvider(options.timeoutMs, !options.json, options.remainingMs), receipts: new FileReceiptSink(join(options.root, ".rcp", "receipts")), }, ); @@ -260,7 +284,7 @@ async function executeRcpVerification( { blocking: [], advisory: plan.advisory }, { includeAdvisory: options.includeAdvisory, - run: npmCommandRunner(options.root, options.json, options.timeoutMs), + run: npmCommandRunner(options.root, options.json, options.timeoutMs, options.remainingMs), }, ); const results = [...blocking, ...advisory.results]; @@ -285,6 +309,7 @@ async function executeNarrowVerification( options: { root: string; timeoutMs: number; + remainingMs: () => number; includeAdvisory: boolean; json: boolean; routeId: string; @@ -306,7 +331,7 @@ async function executeNarrowVerification( repository: new LocalRepositoryProvider(), policy: new DefaultPolicyProvider(), harness: new ExternalWorkspaceHarnessProvider(), - verifier: new NarrowVerifierProvider(options.timeoutMs, !options.json), + verifier: new NarrowVerifierProvider(options.timeoutMs, !options.json, options.remainingMs), receipts: new FileReceiptSink(join(options.root, ".rcp", "receipts")), }, ); @@ -476,6 +501,8 @@ if (invokedPath === fileURLToPath(import.meta.url)) { process.exit(0); } const startedAt = new Date().toISOString(); + const deadlineAt = performance.now() + options.timeoutMs; + const remainingMs = () => Math.max(0, Math.floor(deadlineAt - performance.now())); // A verification is a claim about a tree, and a tree with a live mutant in it is not the tree the // change produced (post-mortem 0003). Refuse rather than report: a passing lane read in that window // is evidence about code that never existed, and that is the reading nobody investigates. @@ -534,6 +561,7 @@ if (invokedPath === fileURLToPath(import.meta.url)) { { root: options.root, timeoutMs: options.timeoutMs, + remainingMs, includeAdvisory: options.includeAdvisory, json: options.json, routeId: route.id, @@ -546,17 +574,40 @@ if (invokedPath === fileURLToPath(import.meta.url)) { report.warnings.push(`narrow escalated to full: ${narrowPlan.escalationReason}`); } execution = contract - ? await executeRcpVerification(report, contract, options) + ? await executeRcpVerification(report, contract, { ...options, remainingMs }) : { result: await executeVerificationPlan(buildVerificationPlan(report), { includeAdvisory: options.includeAdvisory, dryRun: options.dryRun, - run: npmCommandRunner(options.root, options.json, options.timeoutMs), + parallel: report.routes.some((item) => item.id === "ci-and-tests") + ? { commands: PARALLEL_STATIC_CHECKS, concurrency: 3 } + : undefined, + run: npmCommandRunner( + options.root, + options.json, + options.timeoutMs, + remainingMs, + PARALLEL_STATIC_CHECKS, + ), }), rcp: undefined, }; } const { result, rcp } = execution; + if (remainingMs() <= 0) { + result.ok = false; + if (!result.results.some((item) => item.reason?.includes("overall deadline"))) { + result.results.push({ + command: "agent:verify", + classification: "blocking", + routes: report.routes.map((item) => item.id), + status: "failed", + durationMs: 0, + reason: `verification exceeded ${options.timeoutMs}ms overall deadline`, + errorKind: "timeout", + }); + } + } const finishedAt = new Date().toISOString(); const evidence = { schemaVersion: 1,