Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
# Design follow-up: Goal continuity across restart and replacement

- **Status:** Deferred design note; non-normative, not an accepted new runtime contract.
- **Origin:** Retains useful questions from [Duang777's #5169](https://github.com/loopx-project/loopx/pull/5169), with its implementation and evidence claims narrowed during review.
- **Language mirror:** [中文](goal-immutability-coherence-defense-v0.zh-CN.md).
- **Owning contracts:** [Goal instance identity and orphan recovery](goal-instance-identity-and-orphan-recovery-v0.md), [Goal direction baseline](goal-direction-baseline-v0.md), [governed amendment](shared-goal-alignment-and-governed-amendment-v0.md), [semantic handoff](capable-manager-semantic-handoff-v0.md), and [shared authority](shared-goal-authority-state-provider-v0.md).

This preserves follow-up design value from the original “Goal Immutability as
Coherence Defense” draft. It creates no second roadmap, state owner, acceptance
gate or delivery claim. The owning RFCs decide activation, authority and rollout.
The questions below are proposed qualification scenarios, not evidence of a
remaining defect in every named path.

## What is worth retaining

Durable commitments should outlive a model's working context. A restarted or
replaced Agent should recover the authorized Goal, constraints, accepted work
and unresolved obligations from their existing owners. A same-name replacement
Goal must not inherit an old instance's authority merely because names match.
These are useful long-horizon failure scenarios even when individual storage
and command tests pass.

Three distinctions prevent an overly broad “coherence” guarantee:

| Fact | What it can establish | What it cannot establish |
| --- | --- | --- |
| Exact GoalRef and source-owned instance fence | Which Goal lifetime may admit an action | Whether that action is useful or its output correct |
| Provider revision / CAS | Whether a new write still has its expected storage basis | Current Goal authority if the caller resolves/rebinds the wrong instance; revision tokens are opaque, not ordered counters |
| Operation identity and verified original receipt | Which operation already committed and its original result | Permission to repeat an external effect or attach the result to a replacement Goal |
| Authorized intent / acceptance basis | Which constraints and completion criteria govern this work | Model compliance or outcome correctness without independent evidence |

Immutable **instance identity** does not mean immutable **Goal intent**. Authorized
amendments must remain possible and versioned through their existing owner.
A model can produce a wrong change against a perfectly current CAS revision.
Prompt/context improvements, typed constraints and outcome validation complement
storage fences; none substitutes for all the others.

## Proposed follow-up slices under existing owners

| Slice and owner | Real caller scenario | Decisive acceptance, including recovery |
| --- | --- | --- |
| Instance-qualified continuity — Goal instance RFC; related collaboration/session consumers in [#5106](https://github.com/loopx-project/loopx/pull/5106) and [#5130](https://github.com/loopx-project/loopx/pull/5130) | Retire A through the authorized lifecycle, create same-name B, then deliver A's delayed Todo/result, claim renewal and plan confirmation through their real entrypoints | No mutation or execution authority leaks into B. Typed stale-instance outcomes remain observable; B's legitimate work succeeds. Historical A receipts remain attributable to A where retention/access policy permits. Registry activation and its legacy/off behavior follow the owning RFC. |
| Constraint continuity — direction-baseline and governed-amendment RFCs; roadmap R4 | Resume/rebind an Agent after context loss with stale material/acceptance basis, then repeat with an authorized amendment and refreshed basis | Original constraints and accepted work are recovered from canonical owners; re-evaluation remains Agent-scoped. Unrelated work is not globally blocked. The legitimate amendment can progress; no implicit freeze of all Goal intent. |
| Recoverable late-result disposition — handoff and Effect recovery owners; roadmap R3 | An old request's result arrives after requester/instance replacement or after an external effect has committed but its response was lost | Preserve original request/result lineage and the external effect's durable evidence. Reconcile at the owning ledger; do not silently discard evidence, automatically rebind to B or rerun the effect. An authorized recovery path returns a result or records an explicit terminal disposition. |

Before implementing a slice, inventory current main, related PRs and existing
fixtures. Extend the current owner's missing cases rather than creating a
parallel “semantic certificate” or generic coherence engine. Shared decisions
belong in existing typed TS owners; provider adapters supply physical evidence.
As of this note's 2026-09-27 review, #5106, #5130 and the related App Turn recovery
[#5139](https://github.com/loopx-project/loopx/pull/5139) are open; their merge or
isolated tests alone would not certify the combined journeys above.

## Qualification method and unresolved decisions

Use disposable runtimes, synthetic public-safe Goals and actual supported
backends. Derive expected outcomes from the owning contract before executing:

- Exercise same-instance restart, same-name replacement, authorized amendment,
delayed input, overlapping invalid conditions and valid post-recovery work.
A stale rejection alone is not restored progress.
- Exercise both accidental scope expansion and escape: covered old bindings,
unrelated current work and newly created subjects follow the declared scope.
- For concurrent creation, preserve the registry's declared uniqueness and
linearization contract. Do not assume that every racing request must create
a separate writable Goal; distinct successful lifetimes must never share an ID.
- Inject one fault at a time and demonstrate oracle sensitivity: wrong GoalRef,
missing commit fence, dropped handoff constraint or duplicated effect. Keep
real effect evidence distinct from simulated adapters and model evaluations.

Open decisions belong to the existing owners: whether each producer already
captures sufficient immutable instance/basis evidence; how a stale requester
receives a recoverable outcome; and whether an explicit new producer/schema is
needed. Never derive the original instance from a mutable “current Goal” lookup.
If a format change is necessary, qualify backup/migration and mixed writers;
this note does not predeclare “no migration needed.”

## Delivered boundary and evidence status

[#5169's operation replay](../../reference/authority-operation-replay.md) verifies
matching historical File/SQLite commits without rewinding current state. It does
not implement or qualify the three follow-up journeys. Tests of stale provider
revisions are not tests of Goal replacement, compaction or semantic correctness.

The original draft's quantitative research/experiment tables are not retained as
accepted evidence: this PR does not supply an independently reviewable public
harness, oracle and source provenance for them. Future evidence must name the
exact revision, real entrypoint/backend, fault, independent oracle and recovery
readback. No numeric success rate or claim that “CAS prevents coherence collapse”
is carried forward. This note neither changes File/SQLite defaults nor adds a
new prerequisite to their existing D1–D3 qualification.
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# 后续设计:重启与实例替换中的 Goal 连续性

- **状态:** 延后设计记录;非规范性内容,不是新接受的运行时契约。
- **来源:** 保留 [Duang777 在 #5169 中提出的问题](https://github.com/loopx-project/loopx/pull/5169),评审时收窄其实现与证据宣称。
- **语言镜像:** [English](goal-immutability-coherence-defense-v0.md)。
- **所属契约:** [Goal 实例身份与孤儿恢复](goal-instance-identity-and-orphan-recovery-v0.zh-CN.md)、[Goal 方向基线](goal-direction-baseline-v0.zh-CN.md)、[受治理的修改](shared-goal-alignment-and-governed-amendment-v0.md)、[语义交接](capable-manager-semantic-handoff-v0.zh-CN.md)、[共享权威](shared-goal-authority-state-provider-v0.md)。

这里保留原“Goal 不可变性作为一致性防御”草稿中的后续设计价值,不增加第二份
roadmap、状态 owner、验收门禁或交付宣称。激活、权限和上线由所属 RFC 决定。
以下是建议补充的验收场景,不表示每条路径当前都存在尚未修复的缺陷。

## 值得保留的部分

持久工作承诺应当比模型的工作上下文活得更久。Agent 重启或被替换后,应从已有
owner 恢复获得授权的 Goal、约束、已接受的工作和未完成义务。同名的新 Goal
不能仅因名字相同就继承旧实例的权限。即便单个存储和命令测试已经通过,这些仍是
有价值的长程运行反例。

必须区分以下事实,避免泛化为“语义一致性保证”:

| 事实 | 能证明什么 | 不能证明什么 |
| --- | --- | --- |
| 精确 GoalRef 与源 registry 拥有的实例 fence | 哪次 Goal 生命周期可以接纳动作 | 动作是否有用、输出是否正确 |
| Provider revision / CAS | 新写入是否仍基于预期的存储版本 | 调用方错误解析或重绑定实例时的当前 Goal 权限;revision 是不透明 token,不是可排序计数器 |
| 操作身份与经过验证的原回执 | 哪个操作已经提交及其原结果 | 重复执行外部效果、或把结果挂到替代 Goal 的权限 |
| 已授权的意图/验收基线 | 当前工作应遵守哪些约束和完成条件 | 缺乏独立证据时,模型确实遵守了约束或结果正确 |

**实例身份不可变**不等于 **Goal 意图不可修改**。获得授权的修改必须仍能通过现有
owner 留下版本并生效。模型完全可能基于最新 CAS 版本提交错误改动。
Prompt/上下文优化、类型化约束与结果验证和存储 fence 相互补充,没有一项可以
替代全部其他机制。

## 归入现有 owner 的后续切片

| 切片与 owner | 真实调用场景 | 决定性验收,包含恢复 |
| --- | --- | --- |
| 按实例确认连续性——Goal 实例 RFC;相关 collaboration/session 消费者见 [#5106](https://github.com/loopx-project/loopx/pull/5106)、[#5130](https://github.com/loopx-project/loopx/pull/5130) | 通过获授权的生命周期退役 A,建立同名 B,再经真实入口提交 A 的迟到 Todo/结果、claim 续约、计划确认 | B 不受到错误写入或执行权限污染;过期实例结果可观察,B 的合法工作仍能推进。保留/访问策略允许时,A 的历史回执仍归属于 A。Registry 激活及 legacy/off 行为遵循所属 RFC。 |
| 约束连续性——方向基线、受治理修改 RFC;roadmap R4 | Agent 上下文丢失后以过期材料/验收基线恢复或重绑定,再以已获授权的修改和新基线重复执行 | 从 canonical owner 恢复原约束和已接受工作;重新评估局限于相关 Agent,不把无关工作全部阻塞。合法修改可以推进,不隐式冻结全部 Goal 意图。 |
| 迟到结果的可恢复处置——handoff 与 Effect recovery owner;roadmap R3 | 请求方/实例替换后收到旧请求结果,或外部效果已经提交但响应丢失 | 保留原请求/结果关系和外部效果的持久证据;在所属 ledger 对账,不能静默丢弃证据、自动改挂到 B 或重跑效果。通过获授权的恢复路径返回结果,或明确记录终止处置。 |

实施任一切片前,先核对当前 main、相关 PR 和已有 fixture。补齐当前 owner 的缺口,
不另建“semantic certificate”或通用 coherence 引擎。共享决策放在已有类型化 TS
owner,provider adapter 提供物理存储证据。本记录于 2026-09-27 核对时,#5106、
#5130 及相关 App Turn 恢复 [#5139](https://github.com/loopx-project/loopx/pull/5139)
仍开放;仅合并它们或通过各自孤立测试,不代表上述组合链路已经验收。

## 验证方法与未决事项

使用可丢弃 runtime、公开安全的合成 Goal 和真实受支持 backend;执行前从所属
契约独立推导预期结果:

- 覆盖同实例重启、同名替换、获授权的修改、迟到输入、重叠非法条件,以及恢复后
的合法工作。只有 stale 拒绝不等于恢复推进。
- 同时覆盖范围误扩大与逃逸:旧绑定、无关当前工作和新增对象须遵守声明的范围。
- 并发创建应遵守 registry 的唯一性和线性化契约,不假定每个竞争请求都应建立一份
独立可写 Goal;不同的成功生命周期不能共用实例 ID。
- 每次注入一个故障并验证 oracle 的敏感性:错误 GoalRef、缺失提交 fence、交接约束
丢失或重复效果。区分真实效果证据、模拟 adapter 与模型评测。

未决事项交给已有 owner:各 producer 是否已捕获足够的不可变实例/基线证据;
过期请求方如何收到可恢复结果;是否确需新的 producer/schema。不能通过可变的
“当前 Goal”查询倒推出原实例。如果需要格式修改,必须验收备份/迁移和混合版本
writer;本记录不预先宣称“无需迁移”。

## 本次交付边界与证据状态

[#5169 的操作重放](../../reference/authority-operation-replay.md) 验证历史 File/SQLite
提交的完整意图,且不会回退当前状态;它没有实现或验收以上三组后续链路。
过期 provider revision 的测试不能冒充 Goal 替换、上下文压缩或语义正确性测试。

原草稿的研究百分比和实验成绩表不作为已接受证据保留:本 PR 没有提供可独立审查的
公开 harness、oracle 和来源依据。后续证据须明确精确 revision、真实入口/backend、
故障、独立 oracle 和恢复读回。不沿用数值成功率,也不沿用“CAS 防止语义崩塌”的
结论。本记录不改变 File/SQLite 默认值,也不给现有 D1–D3 验收添加新前置条件。
81 changes: 81 additions & 0 deletions docs/reference/authority-operation-replay.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
# Authority operation replay

File and SQLite accept a direct retry of an already committed operation when
its complete canonical body matches the original transaction. The body is
`next_projection`, `events` and `receipts`; JSON object key order is not intent.
`operation_id` selects that transaction within the opened Goal store.

This supports retry after a lost response. It does not create another state
transition, append events again, restore an old head or authorize another
external effect. If A committed, then B committed, replaying A returns A's
original cursor/provider revision while B remains the current head.

## Commit and recovery contract

| Case | File / SQLite | PostgreSQL / NoKV |
| --- | --- | --- |
| New operation, current CAS basis | Commit atomically | Commit atomically |
| New operation, stale CAS basis | `provider_revision_mismatch` | `provider_revision_mismatch` |
| Existing operation, identical body | `applied` with original cursor/revision | Ordinary commit remains a conflict; recover via `readReceipt` |
| Existing operation, different body | `operation_id_exists` | Conflict; normal revision-check precedence remains |

For a verified historical retry, File/SQLite do not require the caller's CAS
basis to remain current: they are returning a historical fact, not admitting a
new write. Validation, store identity/existing-only admission and the current
store integrity checks still apply. An invalid request fails before replay.
A matching operation with a changed projection, event or receipt is never
acknowledged as the original commit.

File reconstructs the original projection using the retained journal. SQLite
verifies the original checkpoint/delta window in the **same write transaction**
before acknowledging a replay. Comparing the caller to a stored digest alone
is insufficient: a damaged retained receipt/event must fail its own proof.
No full-history audit is added to ordinary retry. Digests detect inconsistent
bytes, not an administrator who rewrites both data and proof.

`CoordinationCommandReceipt` remains the provider-neutral business recovery
owner. It reads and validates the original command receipt even when a provider
returns `applied`, and reconciles conflict or ambiguous responses. Do not remove
that readback: `applied` can describe a historical commit, and the other
providers retain their existing direct-commit behavior. NoKV's ambiguous-write
readback recovery is distinct from its ordinary `commitAuthority` contract.

This changes File/SQLite's previous duplicate-commit rejection behavior.
Concurrent identical lease renewals can now both report `applied`, while only
one renewal/version transition is persisted. Business recovery still validates
request identity; a historical receipt never grants current lease authority.
There is no feature flag, new request field, storage migration or frontend
configuration. Reverting requires reverting the provider behavior, not merely
removing tests. Existing durable receipts keep their original format.

## Scope and verification

The shared-authority RFC owns this storage/recovery boundary. Goal lifetime
identity, lease epochs and provider revisions remain separate contracts.
This change does not implement Goal replacement isolation, semantic correctness
of model output, default provider activation or long-horizon qualification.
The [deferred Goal continuity note](../architecture/rfcs/goal-immutability-coherence-defense-v0.md)
preserves related restart, instance-replacement and constraint-recovery scenarios
under their existing RFC owners; those scenarios are not qualified by this PR.

`authority_operation_replay_conformance.ts` runs on both File and SQLite. It
checks full body drift, canonical key ordering, historical replay and concurrent
same-operation attempts against complete head, receipt and history readback.
SQLite adds retained-row corruption cases that must refuse replay without
changing durable rows. Existing real-process lease renewal and command receipt
suites cover the business entrypoint above these providers.

## 中文说明

File/SQLite 现在可以直接重试已成功提交的操作:操作 ID 相同,而且完整的
projection、events、receipts 相同,才返回原 cursor/revision。A 提交后 B 又提交,
重试 A 只确认 A 的历史结果,不把当前状态退回 A,也不再追加事件。

这不是绕过新写入的 CAS。新操作仍须匹配当前版本;历史重试则须证明原事务。
SQLite 在同一事务内校验对应 checkpoint/delta 窗口,不能只比较数据库中保存的
摘要字段。历史回执或事件损坏时必须拒绝,不能报告成功。

上层 `CoordinationCommandReceipt` 仍须读回并校验业务回执,处理响应丢失及不确定
提交;PostgreSQL/NoKV 的普通重复提交仍返回冲突。并发同意图续约可能从过去的
`applied/recovered` 变成 `applied/applied`,但实际只写入一次。历史成功不授予当前
lease 执行权。此变更没有新配置或存储格式,也不证明 Goal 实例隔离或模型语义正确。
4 changes: 4 additions & 0 deletions docs/reference/sqlite-authority-store.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,10 @@ rotation, corruption repair, or network-filesystem sharing is supported.

## Read integrity

Direct retries follow the [authority operation replay contract](authority-operation-replay.md):
matching full intent returns its verified original position without a new write.


Authority reads share one SQLite snapshot, and writes run the same live proof
inside their transaction before publishing a new commit row. The proof is
layered so that each layer pays only for what it returns:
Expand Down
Loading
Loading