Skip to content

test(native): qualify historical dependency journal archival - #114

Closed
roodboi wants to merge 45 commits into
nextfrom
codex/qualify-retired-rebind
Closed

roodboi wants to merge 45 commits into
nextfrom
codex/qualify-retired-rebind

Conversation

@roodboi

@roodboi roodboi commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

A completed dependency refresh from a prior guest boot can survive absent-publication retirement and block endpoint publication after a later retained restore. This ignored cross-version native regression creates that history with the pinned pre-archive candidate (7dc0811cd0179f40340e48c540699933dc099cf7), an actual refreshed listener and a newer stopped generation. It exercises acknowledged-publisher retirement, interrupted journal archival, resumable restore, endpoint publication, retained marker readback and sibling isolation.

Test-only prior-host-boot setup validates the exact receipt-bound reservation and captured foreground process under the provider lease. After proving the owner dead and every selected socket inactive and unchanged, it projects only the process timestamp and recorded socket device numbers. Real journal, socket inodes, bindings, graph receipt and sibling reservations remain intact. Negative controls use the unchanged production inspector to reject current-boot timestamps, device/inode mismatches and live owners/listeners. The README states this synthetic boundary.

This regression now keeps selected/sibling compose files and private selections/prior-owner evidence under the declared private candidate root on failure. It has no fallback --remove-data cleanup or ephemeral-directory Drop. Four readiness stages are labeled, with bounded foreground failure output saved privately before unwind. Successful graph cleanup remains explicit; afterward, directory disposal verifies the originally created directory identities. Shared fixtures and production guards/assertions are unchanged.

History is preserved: merge b357e034 retains original #114 through 65a7cc82 and incorporates fd284a4d, including next at aa2ae7ee, the socket fixes and Bun 1.4.2 pin. efdcdf6e corrects the synthetic reservation setup; bf8c7f2f67c31f359850f7f09e2931e78f0994fd adds failure preservation. Base remains next.

Validation at bf8c7f2f67c31f359850f7f09e2931e78f0994fd with pinned Bun 1.4.2/Rust 1.97.1, one build job and the preserved worktree-local Cargo cache:

  • Formatting, strict default/all-feature all-target Clippy, staged privacy and diff checks pass.
  • Three focused harness tests pass using actual files and an exiting subprocess: panic/unwind preserves bytes/inodes and diagnostics; explicit successful disposal preserves outside-owner data; replaced directories refuse disposal.
  • Full default Rust suite: 1,016 passed/62 ignored/0 failed, with one test thread. Focused all-feature harness suite: 3 passed/1 ignored/0 failed; the ignored native fixture compiled but did not execute.
  • Exact-head CI 36929769921 passed all eight jobs at this pushed head, including macOS/Ubuntu core, CLI and Docker E2E. This is separate from the open native gate.

CI 36925001671 passed all eight jobs at efdcdf6e; CI 36917274313 passed all eight at b357e034. Earlier full local gates at b357e034 passed frozen Bun install/typecheck/check/test (CLI: 1,737 passed/67 skipped; unchanged database tasks replayed Turbo cache), default Rust (1,013 passed/62 ignored) and all-feature Rust (1,102 passed/85 ignored). An earlier concurrent setup-correction run failed two unchanged socket-cleanup assertions while all three new setup controls passed; that failure remains recorded separately, without attributing its cause to concurrency.

Native qualification remains open after two bounded isolated counterexamples:

  • At b357e034, exit 101 after 34.23s at inspect-absence-with-real-journal, native code host_pin_recovery: production correctly refused synthetic prior-owner evidence paired with a real current-host-boot dependency reservation. The completed journal, selected receipt/reservation and retained data survived matching normal VM shutdown. efdcdf6e fixes that setup inconsistency.
  • At efdcdf6e, using verified signed legacy 7dc0811c and current 90ecf2aa executors, exit 101 after 37.98s at first legacy restore, native code dependency_reservation. A new current-PID reservation with populated socket bindings was retained, but existing test panic cleanup removed both graph volumes and temporary project/private inputs before inspection. The exact rejected pre-failure guard is therefore unproven. Post-panic receipt/history/reservation evidence was preserved with hashes and remained byte-identical across matching normal shutdown; both dedicated VMs are stopped. Original volumes/markers from this second attempt were not preserved. This harness correction addresses that evidence-loss boundary; it has not received another native invocation.

Future qualification requires a fresh dedicated capacity-two pool, verified legacy/current executors accepting that root, a pinned Bun-containing image and relay artifact, one test thread and an external 300s watchdog. The test stops/restarts the entire supplied pool and cannot share application/other-owner resources. The host listener uses an ephemeral loopback port. Neither counterexample qualifies archival, retained restore, application acceptance or physical host-reboot continuity.

Historical CI 36752149904 is preserved: macOS formatting failed on module order, Ubuntu was canceled, and Bun 1.3.9 CLI tests segfaulted/exit 133 without a reported assertion failure. Module order is fixed; carrying the Bun pin alone does not establish why that crash occurred.

Related: HACK-1210. Native gate remains open; no merge or release qualification is implied.

hack-cli-tests added 30 commits September 29, 2026 21:41
Stand-in invocations ran concurrently (a foreground `up` while the driver
polls `ps`) and each did an unlocked read-modify-write of state.json. A
`ps` that loaded the state before the second `up` saved its foreground
token then saved its stale copy over it, so `up` saw its token gone and
exited 0 before readiness; lost call records failed the acceptance test
as well. That failed 26 of 40 local runs, and Runtime state models on CI.

Each invocation now holds an exclusive flock across its read-modify-write
and releases it only before its long waits (the foreground loop and the
ps-hang fault). 40 of 40 local runs pass.
A review of #109 found two startup ownership gaps.

The native HTTPS owner discarded the identity listenPublishedUnixSocket
returned and only recorded one after lstat, chmod and lstat of the path.
A failure after publication then skipped listener and endpoint cleanup,
and a replacement in that window could be chmodded or recorded as the
endpoint. The owner now keeps the published identity at once, compares
its first observation against it, and never changes the endpoint by path.
Failure cleanup always closes the listener, which can no longer unlink
anything but the retired staging name, and then removes the endpoint only
while it is this listener's inode; a replacement is kept.

The helper awaited listen() outside its cleanup, so a listen that bound
the staging name and then failed could leak that socket and listener.
Listening is now inside the cleanup, and a failed start removes only a
staging socket this user created. The helper also sets the endpoint mode
(0600) on the staging inode before publication, so the MCP backend, the
HTTPS owner and the owner challenge no longer chmod the public path.

Controls: bind-then-fail, mode-preparation failure, an endpoint at the
AF_UNIX path limit (and one byte over: refused cleanly on Bun 1.3.9,
bound in full on 1.4.2, never truncated), and an owner fault or
replacement between publication and first observation, on Bun 1.3.9 and
1.4.2.
hack-cli-tests added 13 commits September 30, 2026 14:28
The private-bind fixture hooked chmod on the public mcp.sock to prove the
socket is private before its mode is set and that the creation mask is
restored. Since the endpoint's mode is now set on its staging socket
before publication, the hook never fired and the check failed. It now
observes that staging chmod with the same privacy and mask checks,
requires mode 0600, and fails if the published endpoint is ever chmodded
by path.
…a socket

A second review of the Unix-socket publication found three staging-path
windows where an entry this attempt could not prove was its own could be
changed or removed:

- After an ambiguous partial bind (no recorded identity), cleanup
  unlinked any same-uid socket at the staging name.
- The mode was set by path after an awaited lstat, so a replacement in
  that gap could be chmodded before the postcheck refused it.
- On success the staging name was unlinked unconditionally after the
  awaited link and endpoint lstat.

The socket is now created with exactly its mode (the umask during the
synchronous bind), so nothing is ever chmodded, and its identity is
recorded in the same tick. Publication links only while the staging name
is still that socket, and retirement removes it only while it is (check
and removal back to back). An entry that is not this socket, or cannot be
proven to be, is never removed: on failure it is moved aside while the
server closes, so the runtime's close-time unlink by name cannot reach
it, and then put back with the same inode.

The owner challenge's chmod dependency becomes an afterOwnerSocketPublish
test seam, and the private-bind fixture now requires the socket to be
created private with mode 0600 and no socket name to be chmodded.
…ng entry

The failure close moved an unproven staging entry to one random holding
name: if that name was occupied the close went ahead unprotected, and a
holding entry created between the check and the rename was overwritten.
The entry is now moved with link(2), which never replaces a holding
entry, over a bounded list of fresh holding names, and the staging name
is dropped only while it is still the linked entry. An entry that cannot
be moved aside (every holding name occupied, or not hard-linkable) leaves
the close to proceed, now documented as a residual.

The helper's documentation now states its scope: accidental and
concurrent entries in a private directory, not an adversarial process of
the same user, with the remaining path-based and post-return residuals.
Add explicit selector-bound recovery for a dead legacy HTTPS owner beside an unpublished shared-owner configuration. Preserve original inodes and the CA with resumable exclusive archival, held IPv4/IPv6 guards, and refusal on live or uncertain evidence.
Keep strict receipt identities as the default. Allow only an explicitly selected source-device witness and both current HTTPS inodes, with unchanged inode numbers and matching old/current device mapping. Recheck current owner, host and guest boot, source share and stopped graph scope before archival; preserve historical bytes and report original host-volume continuity as unproven.
@roodboi

roodboi commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by merged #122. This PR’s source was incorporated through #122, squash commit cf4b6e9. Independent acceptance remains in Linear. Closing as superseded; branches and worktrees are retained.

@roodboi roodboi closed this Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant