Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Local default cutover: recovery audit and remaining delivery scopes

- Audited baseline: `157ab7b11`, 2026-09-27, plus this delivery.
- Historical recovery baseline: `157ab7b11`; current inventory: `70b3cca01`, 2026-09-27.
- Owner: overall roadmap #4574 R5/G2; shared authority D2/D3; TS T3/T4.
- Supersedes the **current count**, not historical evidence, in the
[September 24 reconciliation](2026-09-24-default-cutover-reconciliation.md).
Expand All @@ -21,7 +21,7 @@ target before the final digest mismatch stopped recovery. There was also no
independent read-only CLI proof of a restored store's complete retained history
and receipt lookup. These are recovery gaps, not missing capture writers.

## Four scoped deliveries starting with this PR
## Historical recovery delivery allocation

| Delivery | Observable result and remaining boundary |
| --- | --- |
Expand All @@ -30,7 +30,7 @@ and receipt lookup. These are recovery gaps, not missing capture writers.
| **3. Whole-Goal activation and rollback integration** | Reconcile #5054's retained-source inventory, then exercise source drain, saved reviewed cutover, all retained command consumers and fenced recovery/rollback together. Bind a recovered copy through an explicit transition; do not revive a source lease or overwrite later writes. Delete only Python decisions whose callers have actually moved. |
| **4. Default entrypoints and final bounded retirement** | New Goal creation, settings, installation, packaged frontend/Lark/CLI consistently use the qualified local profile. Existing Goals have explicit migration and disable/recovery paths. Remove last legacy business writers after their caller inventory and rollback constraints pass; retain rendering and Host IO. |

This is **four planned new delivery PRs including this one, three afterwards**,
At that recovery checkpoint, this was **four planned new delivery PRs including the recovery PR, three afterwards**,
not a guarantee that no acceptance defect will require another PR. The original
three *architectural packages* are not a decrementing PR counter. This PR closes
one named recovery slice inside package 2; it does not close all of package 2.
Expand Down Expand Up @@ -97,17 +97,73 @@ PostgreSQL 16 server passed, with no skipped checks in these suites. A detached
audit. Recovery of the earlier timed-out SQLite destination passed without
reissuing its committed operations. These checks do not claim active cutover.

## Current delivery inventory and native drain (`70b3cca01`)

The current plan contains **seven delivery slots including this change**: four
already-open PRs and three scoped deliveries. It does not mean seven new PRs,
nor guarantee that seven merges suffice. Earlier counts treated whole-Goal
integration as a single PR before its recovery and execution gaps were bounded;
that was an architectural grouping, not a reliable PR commitment.

| Slot | Existing work / observable completion |
| --- | --- |
| 1 | **#5173**, open: reviewed File↔SQLite selector/fence cutover, durable backup, full-history audit, recovery and retry. Integrate it; do not rebuild it. |
| 2 | **#5144**, open: managed Host execution lifetime/lease supervision. Attached Hosts still need an explicit cancellation boundary. |
| 3 | **#5054**, open: retire legacy Todo event projection/backfill/completion and isolate the experimental supervisor log. |
| 4 | **#4931**, open: SQLite retained-proof encoding/read cost; rerun the applicable formal D2 workload instead of equating an optimization with qualification. |
| 5 | **This change**: move the complete bounded source-outbox drain to TS, remove the Python sequencing/proof/cleanup coordinator and the unused per-entry planning RPC. Keep existing durable formats, full receipt verification and the kernel-lock adapter. |
| 6 | **Whole-Goal integration**: combine the accepted slices with retained consumer parity, interrupted cutover/rollback and post-cutover writes. This drain is one completed subitem, not completion of that scope. |
| 7 | **Default entrypoints and bounded Python retirement**: qualify creation/settings/install/frontend/Lark/CLI, migrate existing Goals explicitly, and delete business writers only after their real callers have moved. |

#5169's content-aware idempotency work is adjacent and must be integrated without
rewriting it; it is not silently counted as another required default-cutover PR.
D1 consumer coverage, D2 capacity/ten-day natural soak and D3 cohort/maintainer
promotion remain **evidence gates**, outside the arithmetic. PostgreSQL service
identity, deployment and operational qualification remain a medium-term scope.

TS now inventories witnessed source files, invokes the existing receipt planner
and transaction owner, checks the monotonic budget after proof, and performs
cursor/cleanup effects under M → primary marker → kernel-lock exclusion. The
Python facade sends one bounded request with no source projection/history and
does not automatically retry a lost response. `shadow_drain_outcome_unknown`
requires a later explicit receipt-based drain; it never asserts no commit.
Unconfigured Goals with no capture state still make no drain RPC.

Real process-death validation found a shared lock defect: a zero-wait acquire
reclaimed a dead owner but returned timeout before trying the now-free path.
It now allows one immediate retry after proven reclamation; live owners still
reject immediately. This uses the shared lock owner rather than a drain-only
sleep or increased timeout. The two crash harnesses now share one scheduling
fixture and kill/reap the actual TS owner at the durable boundary.

The runtime shadow remains a **File candidate**, not a promoted authority or a
new SQLite shadow provider. SQLite remains a supported canonical promotion and
archive-restore target. No provider selector, registry or active Goal is changed
by this delivery. The existing CLI and inline writer drain entrypoints adopt the
same owner; no frontend/Lark configuration contract changes.

Validation for native drain: the new full-batch tests exercise the existing
mixed production-scale Todo fixture; real CLI SIGKILL, filesystem permission,
cursor tampering and File/SQLite reviewed-promotion tests cover the persistence
boundaries. Shared-store regression passes against isolated PostgreSQL 16.
A detached authorized snapshot supplies 1,101 complete Todo records; three
explicitly synthetic source writes produce a fresh four-transaction candidate.
Every original Todo JSON record survives drain and SQLite/File archive restore.
This is not replay of that snapshot's old transaction history or a live migration.

On three local three-entry trials, median drain time is 1.34 s on the audited
base and 0.41 s here, with 11 facade RPCs reduced to one. The kernel-lock process
is lazy and reused within a batch; it releases locks between sections. These
small warm-runtime measurements are not D2 p95/capacity claims. The existing
CLI output-budget suite fails on both base and head at 14,514 versus 14,500
characters for the crowded Turn JSON packet. No ceiling is raised; merge
qualification retains that failure rather than declaring all checks green.

## Shared-runtime latency reconciliation (`96a3b90f4`)

The recovery delivery above is now merged as #5140. #5144 is the open managed
Host process supervision slice; attached hosts still need their declared
cancellation boundary. Whole-Goal activation/rollback and default entrypoint
cutover remain the two planned subsequent implementation PRs. Existing #5054
(retirement) and #4931 (SQLite proof encoding) remain open and are not new work.
Thus the inventory is two planned implementation PRs plus those three existing
PRs, **before this newly reproduced latency repair**. This is an inventory, not
an unconditional completion count: #4224 D2 capacity/soak and D1/D3 evidence
remain gates, and failures may require additional scoped fixes.
The recovery delivery above merged as #5140. The subsequent latency repair is
also on the current audited main. Its observations below remain historical
validation, not another open delivery or a D2 qualification.

The latency repair does not retire another Python owner or close D2. Isolated
fixed File snapshots reproduce 9.9–10.5 second cold history verification,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ checkpoint/delta 格式升级和有界 Python 原型退役已在 main。尤其 #
内容可能先写入隔离目标,直到最终摘要不符才失败。此外缺少独立只读 CLI,证明恢复
目标的完整历史及回执查询仍正确。这是恢复缺口,不是尚未实现事件捕获。

## 从本次开始的四个交付范围
## 当时的恢复交付拆分(历史记录)

| 交付 | 可观察结果及剩余边界 |
| --- | --- |
Expand All @@ -27,7 +27,7 @@ checkpoint/delta 格式升级和有界 Python 原型退役已在 main。尤其 #
| **3. 整 Goal 激活及回退集成** | 对齐 #5054 的保留来源清单,联合验证来源 drain、保存的 reviewed cutover、全部保留命令消费者及 fenced recovery/rollback。通过明确转换绑定恢复副本,不复活旧租约,不覆盖后来写入。仅删除 caller 已迁走的 Python 决策。 |
| **4. 默认入口与最后一批有界退役** | 新建 Goal、settings、安装及打包 frontend/Lark/CLI 一致使用合格本地 profile;存量 Goal 有明确迁移及停用/恢复路径。caller 清单与回退约束通过后,删除最后的旧业务 writer,保留渲染及 Host IO。 |

这是**包含本次在内四个规划新 PR,本次交付后剩三个**;不保证验收不会再发现需要
当时的恢复检查点规划为**包含恢复 PR 在内四个新 PR,交付后剩三个**;不保证验收不会再发现需要
修复的缺陷。原来的三个“架构工作包”不是倒计时 PR 数。本次关闭第 2 包中的一个
具名恢复切片,没有把整个第 2 包标为完成。后续必须指出哪行真正交付,不能再重复
一个不变的“5–8”。
Expand Down Expand Up @@ -76,14 +76,58 @@ PostgreSQL 16 的 4 项跨 provider 检查全部通过,这些套件没有跳
独立真实来源快照通过 File、SQLite 恢复及独立审计;此前超时的 SQLite 目标也成功
恢复,没有重发已提交操作。这些结果不表示已经完成活跃 Goal 切换。

## 当前交付清单与原生 drain(`70b3cca01`)

当前规划有 **7 个交付槽位,包含本次:4 个已有开放 PR,加 3 个交付范围**。
这不意味着还要新开 7 个 PR,也不能保证合并 7 个就足够。此前把整 Goal 整合
当成一个 PR,但恢复与执行边界尚未拆清;那是架构工作包,不能作为准确倒计时。

| 项目 | 已有工作与完成标准 |
| --- | --- |
| 1 | **#5173**:受审 File↔SQLite selector/fence 切换、备份、完整历史核对、恢复和重试。整合已有实现,不重写。 |
| 2 | **#5144**:受管 Host 执行生命周期与租约监督;attached Host 仍需明确取消能力边界。 |
| 3 | **#5054**:退役旧 Todo 事件投影、回填与 completion 分支,独立实验 supervisor 日志。 |
| 4 | **#4931**:SQLite 历史证明编码与读取成本;需正式 D2 工作负载复验,优化代码不等于资格通过。 |
| 5 | **本次**:完整有界 source-outbox drain 归 TS,删除 Python 的顺序、证明、清理编排和无人调用的逐条规划 RPC;保留持久格式、回执证明、内核锁适配。 |
| 6 | **整 Goal 整合**:接通已交付切片,验证保留消费者、中断切换/回退及切换后的新写入。本次 drain 只完成其中一个子项。 |
| 7 | **默认入口与有界 Python 退役**:创建、设置、安装、前端、Lark、CLI 一致;已有 Goal 明确迁移;真实调用者迁走后再删业务 writer。 |

#5169 的内容感知幂等是相邻工作,应避免重写,不把它悄悄加成另一个默认切换必需 PR。
D1 消费者覆盖、D2 容量与至少十天自然 soak、D3 队列与维护者晋升仍是独立证据门。
PostgreSQL 服务身份、部署和运维资格属于中期范围。

TS 统一读取来源文件见证、调用既有回执 planner 与事务 owner,在证明后重新检查
单调时钟预算,并在 M → primary marker → kernel lock 的顺序下更新游标和清理。
Python 只发送一次有界请求,不搬完整 projection/history,也不自动重试响应丢失的
批次。`shadow_drain_outcome_unknown` 表示必须由下次显式 drain 读取回执恢复,
不能解释为“没有提交”。未启用且不存在捕获状态的 Goal 仍不产生 drain RPC。

真正杀进程的测试发现共享锁缺陷:零等待获取已经回收死进程锁,却立即报超时。
现在仅在确认回收后立即再尝试一次;活进程持锁时仍马上退出。这是共享锁修复,
没有为 drain 增加 sleep 或放宽超时。两套崩溃 harness 也共用一个调度 fixture,
在持久化边界杀掉并回收真正执行决策的 TS 进程。

runtime shadow 仍是 **File 候选存储**,不是正式 authority,也没有新增 SQLite
shadow provider。SQLite 是 canonical 晋升和归档恢复目标。本次不切换任何活跃
Goal、provider selector 或 registry。原有 CLI 与 writer 内联 drain 采用同一 owner;
没有前端/Lark 配置合同变化。

本次验证复用混合 production-scale Todo fixture;真实 CLI 的 SIGKILL、文件权限、
游标篡改、File/SQLite 受审晋升覆盖持久边界;共享存储在隔离 PostgreSQL 16 上回归。
授权隔离快照提供 1,101 个完整 Todo,三笔明确标记的合成来源写入形成新的四事务
候选历史。原始 Todo JSON 经 drain 及 SQLite/File 归档恢复仍完整相等;这不是
重放该快照的旧事务历史,也不是活跃 Goal 迁移。

本机三个事务的三次测量中位数由基线 1.34 秒降至本次 0.41 秒,facade RPC 从
11 次降为 1 次。内核锁进程按需创建、批次内复用,各临界区之间释放锁。这是小
负载热运行时测量,不是 D2 p95/容量结论。CLI 输出预算检查在 base/head 都因
拥挤 Turn JSON 为 14,514 字符、超过 14,500 上限而失败;未提高预算,合并资格
保留此失败,不能报告所有检查为绿。

## 共享运行时延迟核对(`96a3b90f4`)

上文恢复交付已作为 #5140 合入。#5144 是待合并的受管 Host 进程监督切片;attached
Host 仍需明确取消能力边界。整 Goal 激活/回退、默认入口切换仍是后续两个规划实现
PR。已有 #5054(退役)、#4931(SQLite 证明编码)仍开放,不重复实现。因此清单是
**两个规划实现 PR,加三个已有 PR,再加本次新复现的延迟修复**。这是工作清单,
不是无条件完成倒计时:#4224 D2 容量/soak、D1/D3 证据仍需验收;失败可产生新的
有界修复,必须指出具体缺陷,不能重新复述一个固定区间。
上文恢复交付 #5140,以及后续共享运行时延迟修复,都已进入当前核对的 main。
下文保留当时的验证事实,不再把它们算作开放工作,也不将其当作 D2 资格。

本次不退役额外 Python owner,也不关闭 D2。隔离固定 File 快照复现了 9.9–10.5 秒
的冷历史校验,期间 ping 在原 10 秒预算内超时。热缓存掩盖问题,交替读取 Goal 又
Expand Down
27 changes: 27 additions & 0 deletions docs/reference/reviewed-coordination-promotion.md
Original file line number Diff line number Diff line change
Expand Up @@ -166,3 +166,30 @@ lost acknowledgement after commit. These are synthetic fault injections, not a
claim of arbitrary process-death or elapsed-soak coverage. Real local rehearsals
must use read-only captured sources and disposable copies; never promote an
active Goal merely to validate this refactor.

## Drain before promotion

`authority-shadow drain` and post-write inline drains use one bounded TS batch.
It verifies the complete retained lineage and exact source byte witnesses before
advancing a cursor or reclaiming outbox files. A cursor is a checkpoint, not
permission to delete; Todo metadata and original transaction receipts remain
part of the existing complete record contract.

```bash
loopx --format json authority-shadow drain --goal-id example-goal \
--max-entries 64 --budget-seconds 30
```

`budget_exhausted=true` with pending entries means another bounded invocation is
needed. The budget controls admission of further effects; it does not cancel an
in-flight durable commit or replace the transport timeout. A live primary writer returns `primary_writer_busy` without waiting for
it. `shadow_drain_outcome_unknown` means the transport lost a trustworthy batch
result: inspect status and invoke drain again explicitly. The next invocation
uses persisted receipts; it must not recreate the business write, delete the
outbox, or assume that no candidate commit happened. Other proof failures require
repairing their reported cause before continuing.

The runtime shadow is still a File candidate. Draining it does not select a
canonical provider, promote a Goal, bypass qualification, or create a SQLite
shadow. File/SQLite promotion uses the reviewed journey above. Feature-off
writers with no prior capture state do not start the drain runtime.
15 changes: 6 additions & 9 deletions examples/shared-goal-authority-e2e/mutants.py
Original file line number Diff line number Diff line change
Expand Up @@ -181,9 +181,9 @@ def command(self) -> list[str]:
"const matched = localAuthorityShadowHeadDigest(request.projection) === localAuthorityShadowHeadDigest(lineage.head.head);",
"const matched = true;")),),
LADDER_ROW + "[s2c2.parity_divergent_detects_foreign_edit]"),
Case("replay_counted_as_delivery", ((COORDINATION + "local_authority_shadow_adapter.py", replacement(
" self._result.replayed += 1\n",
" self._result.delivered += 1\n")),),
Case("replay_counted_as_delivery", ((COORDINATION + "shadow_drain.ts", replacement(
"result.no_op += Number(noOp === true); result.entries.push(summary); result.replayed++; consumed++;",
"result.no_op += Number(noOp === true); result.entries.push(summary); result.delivered++; consumed++;")),),
LADDER_ROW + "[s2c2.sigkill_mid_drain]"),
])

Expand Down Expand Up @@ -252,12 +252,9 @@ def apply(source: str) -> str:
" resolved_source = state_file.resolve(strict=False)",
" return # DELIBERATE MUTANT: allow another goal to bypass source authority.\n resolved_source = state_file.resolve(strict=False)")),),
"tests/control_plane/test_shadow_writer_variant_e2e.py::test_other_goal_cannot_write_a_protected_goal_source_via_state_override[active_capture]"),
Case("cleanup_hides_verified_commit", ((COORDINATION + "local_authority_shadow_adapter.py", replacement(
" if self._result.cursor_before is None:\n"
" self._result.cursor_before = view.get(\"cursor\")\n"
" self._record_view(view)\n",
" if self._result.cursor_before is None:\n"
" self._result.cursor_before = view.get(\"cursor\")\n")),),
Case("cleanup_hides_verified_commit", ((COORDINATION + "shadow_drain.ts", replacement(
" if (plan.view !== null && typeof plan.view === \"object\") observe(plan.view as JsonObject);\n",
"")),),
"tests/control_plane/test_shadow_drain_adversarial.py::test_cleanup_permission_failure_reports_verified_commit_and_recovers[before_commit]"),
Case("native_update_maintenance", ((COORDINATION + "local_authority_runtime.ts",
remove_native_update_maintenance),),
Expand Down
Loading
Loading