diff --git a/.ai/skills/hack-repo-verify/SKILL.md b/.ai/skills/hack-repo-verify/SKILL.md index 8acf3367a..21ea9ed6f 100644 --- a/.ai/skills/hack-repo-verify/SKILL.md +++ b/.ai/skills/hack-repo-verify/SKILL.md @@ -53,7 +53,7 @@ untracked candidate additions. Keep private artifacts excluded from the index. | Rust state/recovery | `packages/runtime-core/tests/` plus module tests; malformed/stale ownership, crash window, resume and data preservation | | TLA models | `test:models`, [model mappings](../../../tests/models/tla/README.md), [runner](../../../scripts/check-tla-models.ts) and expected positive/negative controls | | Consumer instructions | [ownership](../../../docs/agent-guidance.md), `bun run generate:agent-plugins`, affected examples and source/render tests | -| Native VM/routing/reclamation | candidate guide (local `_docs/docs/plans/v5/development.md`), relevant ledger gate and isolated host fixture; Linux checks cannot qualify macOS effects | +| Native VM/routing/reclamation | [candidate guide](../../../docs/guides/native-candidate.md), relevant acceptance gate and isolated host fixture; Linux checks cannot qualify macOS effects | | Resource/performance claims | [performance skill](../hack-repo-performance/SKILL.md), matched workload and measured boundaries | Read [CI](../../../.github/workflows/ci.yml) for current hosted gates. A local pass @@ -63,6 +63,14 @@ not default local-product requirements; use them only for explicitly scoped work ## Verify the verifier Treat unavailable prerequisites, ignored tests, cached results and timeouts explicitly. +Before an ignored native graph fixture, read its prerequisites in the +[runtime README](../../../packages/runtime-core/README.md). Some fixtures require +a caller-prepared pool: use the managed provider, engine and network-tool setup, +start the required socket capacity, load the pinned image, and verify runtime +status before invoking the test. Record the current-source test executable and +candidate bundle separately. An unprepared-pool failure does not exercise the +scenario; preserve its evidence and fix setup before rerunning. + TLC must explore the intended states; its negative control must fail the named invariant with the expected trace, not merely return nonzero. A small abstract model needs source mapping and implementation regression evidence; use diff --git a/docs/guides/native-candidate.md b/docs/guides/native-candidate.md index eb3c3e896..f28d24825 100644 --- a/docs/guides/native-candidate.md +++ b/docs/guides/native-candidate.md @@ -93,6 +93,25 @@ still apply. A same-checkout branch namespace does not create an isolated source tree. These planning contracts do not qualify multiple independent source roots inside one VM pool. +An existing frontend branch mapping may retain a graph admitted before native +branch namespaces. If its current scoped review differs, the frontend first +checks native selection against that exact saved run, owner, namespace and plan: +restore selection for a stopped graph, or authenticated service selection for an +active graph. A native unbranched review of the same canonical project and unchanged +original Compose must then return the saved namespace. Only that retained run +continues with unbranched review and execution arguments and keeps its original +adapted route labels, before branch hostname rewriting. This preserves legacy +source contracts that did not enroll hostname changes. The selected generation, +restore selection and shared-source semantic compatibility are checked again +before admission; mappings, graph receipts and ownership namespaces are not +rewritten. Active legacy restart also rechecks its service, container, boot and +receipt generation after compatibility review, before cleanup eligibility. A +changed selection refuses before stopping the graph. This preflight check does +not make cleanup atomic with that generation; existing ownership and frontend +finalization checks still apply. Fresh and already branch-scoped graphs keep +their normal branch namespace. This compatibility path does not authorize +adopting an unrelated branch or project. + ## Manual candidate upgrade and rollback Candidate bundles are selected by their full path. There is no automatic candidate @@ -118,6 +137,141 @@ complete bundle to change the candidate. binary refusing newer state is an unsupported downgrade, not a successful rollback. Preserve that state; do not rewrite receipts or adopt another home's data. +Host cleanup or a reboot can remove the temporary `/private/tmp/hkl-` HOME +alias while the private candidate home and disks remain intact. Normal commands +report `provider_home_missing`; they do not initialize another pool or recreate the +alias during observation. Explicit `runtime recover --json` can restore only that +absent exact alias, after verifying the receipt-bound dead provider, both recorded +disks, free VM lock, closed disk handles, and no active provider command. Existing +files, directories, foreign links, a live or reused PID, and uncertain ownership +refuse without replacement. Recovery rechecks the receipt before exclusive creation, +then validates socket absence and flushes the disks before recording +`recovered-unclean`. If those final checks fail, it reports +`provider_home_restored_recovery_incomplete`, keeps the exact owned alias and retained +data, and requires inspection before retry. No guest is started by alias recovery. +Interrupted starts without a recorded process or both identified disks remain +outside this repair path. + +A physical macOS reboot may also renumber the mounted filesystem device. Strict +disk and source checks still refuse a changed device number; missing-HOME recovery +does not waive them. For an offline **stock pool**, inspect the separate migration: + +```sh +./hack-native --candidate-root /absolute/private/candidate-home runtime host-filesystem-recovery --json +./hack-native --candidate-root /absolute/private/candidate-home runtime recover-host-filesystem --expect-sha256 --accept-legacy-device-rebind --json +./hack-native --candidate-root /absolute/private/candidate-home runtime recover --json +``` + +Review the inspection before selecting its hash. This explicit legacy migration +requires an absent recorded provider whose start predates the current host boot, +no active provider commands, exclusive existing operation/VM locks and closed disk +handles. Both disk inodes, sizes and ext4 UUIDs, the exact source path/inode and its +ownership must be unchanged; only one common old-to-new device-number change is +allowed. The inspection is read-only. Publication atomically changes only the +owner's disk and source device numbers, retaining its phase and process record. +Normal commands keep their strict identity checks. + +Legacy receipts have no original host boot UUID or filesystem volume UUID. Calendar +timestamps corroborate a reboot, and matching retained file identities constrain +the migration, but neither proves original volume continuity. The opt-in explicitly +accepts that limitation; copied or relocated pools are outside this procedure. +Prepared-base pools and pending owner/network/activation updates require separate +recovery and are refused. A torn owner publication preserves `owner.pending` and +blocks another migration; do not delete or adopt that file manually. + +This does not start the VM, restore HOME, retire stale sockets or rewrite historical +graph receipts. Use ordinary recovery afterward. Historical shared-source graphs +retain the prior identity and may require verified cleanup plus a new generation; +successful metadata migration alone does not establish application recovery. + +After that explicit provider recovery and one audited VM boot, a retained graph +from the immediately preceding guest boot can still carry old host device numbers +in its publisher, relay-control, and dependency-socket receipts. Inspect and select +one run's host-pin recovery separately: + +```sh +./hack-native --candidate-root /absolute/private/candidate-home graph inspect-host-pin-recovery --run-id --json +./hack-native --candidate-root /absolute/private/candidate-home graph recover-host-pins --run-id --expect-selection --accept-legacy-device-rebind --json +./hack-native --candidate-root /absolute/private/candidate-home graph recover-cleanup --run-id --expect-receipt --json +./hack-native --candidate-root /absolute/private/candidate-home graph retire-recovered-publisher --run-id --expect-owner --json +``` + +The first command only inspects. The second publishes a private, exact-run +witness before cleanup; it changes no graph, publisher, control, or dependency +receipt. It requires the provider's repaired current disks/source, unchanged +recorded inodes and raw receipts, dead owners whose recorded starts precede the +current physical host boot, refused socket listeners, and the immediate guest +boot transition. It also binds retained volume names and actual labels. Cleanup +and retirement can use the witness only for those selected old pins; ordinary +reads and later publisher generations remain strict. Missing or foreign pins, +active listeners, changed resources, pending journals, or further guest boots +refuse without adopting another run. A completed older cleanup journal is +retained as history and does not itself block a later selected generation. + +Legacy receipts do not identify the original APFS volume. This explicit +device-number rebind cannot prove pre-reboot volume continuity. Completing these +commands retains the old run's data and proves cleanup of its dead generation; +it does not migrate `graph.source.shared`, restore the application, establish +route readiness, or claim overall v5 acceptance. Same-run source continuity +requires a separate explicit witness-bound transition after cleanup and +publisher retirement. If macOS removed the foreground or relay-control +directory itself during reboot, this command refuses: the old pin receipts no +longer exist, and absent pathnames cannot stand in for their recorded owner +identities. That case requires a separate selected absence-recovery procedure. + +When a **physical host reboot** removed both the deterministic foreground +publication root and this run's relay-control root, inspect the distinct +absence path with the private original provider Owner and its exact +pre-migration host-filesystem inspection: + +```sh +./hack-native --candidate-root /absolute/private/candidate-home graph inspect-absent-publication-cleanup --run-id --original-owner-file --host-inspection-file --json +./hack-native --candidate-root /absolute/private/candidate-home graph recover-absent-publication-cleanup --run-id --original-owner-file --host-inspection-file --expect-selection --retain-data --accept-unpinned-post-reboot --json +``` + +Inspection does not create a publication root. A failed selected action can +leave only its private lock reservation; inspection recognizes that exact +lock-only state. Recovery locks the foreground publication before acquiring +the VM lease, then durably records the selected absence before any cleanup. +It retains volumes, verifies the stopped receipt, and records a separate +absent-publisher retirement. Ordinary publication and cleanup cannot infer +ownership from missing paths. The old foreground PID and original physical +volume are not proved by legacy graph receipts; this path requires explicit +acceptance of that post-reboot limitation and refuses a changed boot, present +or foreign publication, stale selection, pending state, or changed resources. +An interrupted recovery can resume only its exact selected intent on the same +host and immediate guest boot. A later successful restore treats this witness +as history; it never grants cleanup of the new generation. + +This operation does not change historical shared-source device identity. +Source-mounted projects require the separate selected source-continuity step +before normal same-run restore. The command's stopped/data-retained result is +not proof of application startup or routing. + +The frontend project-run mapping also records directory device numbers. If native +inspection succeeds after provider recovery but ordinary project commands refuse +the mapping, inspect the exact instance with the current bundle: + +```sh +./hack-v5 doctor --path /absolute/project --branch my-instance --native-run-mapping inspect --json +./hack-v5 doctor --path /absolute/project --branch my-instance --native-run-mapping repair --expect-selection --accept-legacy-device-rebind --json +``` + +Omitting `--branch` uses the same linked-worktree default as project commands; +detached linked worktrees require an explicit instance. Inspection writes nothing. +Repair requires the exact selection and explicit acceptance of unproven original +filesystem volume continuity. It changes only the three scope directory device +numbers, preserving canonical paths, inodes, branch, run, owner, plan, environment +selection, profiles and AWS selector. The selected graph and current project share +must still match native authority. An audit copy retains the original private +mapping. Held locks, pending restart state, substituted directories, stale hashes +and changed native identities refuse publication. + +This is a metadata repair, not an app restart: it does not retire sockets, modify +graph history, remove data or establish browser readiness. Doctor's ordinary +`--fix` never applies it. Run the normal project command separately after reviewing +the result; its existing native graph and ownership checks still apply. + Compatibility is qualified for specific bundle hashes and state formats. The displayed version alone does not establish frontend/executor or downgrade compatibility. V4 and candidate homes remain separate; this procedure does not @@ -230,6 +384,12 @@ and route selections. Normal foreground `hack restart` performs this selection and retained-data restore. It verifies the old containers and networks are absent and the retained volumes still have their recorded identities; it does not silently create replacement data. +When the original Compose file is unchanged, retained startup also selects the +authenticated container image IDs from that stopped graph. Mutable tags are not +resolved again for those services. Missing images still refuse native admission; +this selection never replaces generation, source or resource checks. Changing the +original Compose file uses normal image resolution and the existing compatibility +rules, rather than silently adopting an old image for a new declaration. For a shared-source graph admitted with a retained compatibility contract, ordinary source-content edits may change the reviewed plan ID: restart checks the stable execution, exclusion, mount and dependency-cache inputs before cleanup and again @@ -387,12 +547,26 @@ absence while keeping the exact run mapping. It then retires that branch's owned lifecycle processes before running after hooks. A previously recovered stopped graph also retires remaining owned host processes without replaying hooks or guest cleanup; uncertain ownership refuses retirement and leaves the retained mapping intact. +Before capturing live bridge helpers for cleanup, the owner retires an exited +helper only after confirming its current guest boot, run, container, network and +reservation. An interrupted retirement requires a fresh exit observation or an +exact stopped slot fence with the allocation and socket absent. A live, replaced +or unconfirmed helper remains a refusal; named volumes and sibling runs retain +their existing ownership protections. +Successful `down` also requires acknowledgement of the exact frontend attempt +captured before cleanup. Missing, changed or unacknowledged finalization refuses +success even when compute is already stopped; the run mapping and volumes remain +available for explicit recovery. Restart retains its own captured finalization +barrier so its selected dead-frontend recovery can run before replacement startup. If the foreground owner concurrently retires its tmux session, cleanup accepts a fresh explicit absence proof. Missing ownership metadata or an unsuccessful query alone does not prove absence, and retirement preserves any replacement owner's state. The next ordinary `up` verifies that stopped receipt and restores the same run and volumes, including when its clean owned -VM has been stopped. Offline `graph retained-preflight` returns durable eligibility +VM has been stopped. It requires the retained frontend's exact finalization to be +acknowledged before hooks, storage preparation, VM startup or shared HTTPS admission. +This observation never recovers an interrupted frontend automatically, and the later +startup ownership check still refuses a concurrent replacement. Offline `graph retained-preflight` returns durable eligibility only; it never fabricates volume observations or a restore generation. Startup checks the saved environment, profiles and AWS selector before hooks and binds the exact owner and receipt selection to `runtime up` with paired `--expect-retained-run` and @@ -436,7 +610,16 @@ requires ownership inspection rather than automatic removal. Restart does not implicitly migrate a shared pool's network policy or interrupt other projects. When an interrupted frontend cannot write that final acknowledgement, an explicit `restart --recover-frontend --expect-finalization-attempt <32-hex>` can resume -the saved, already-cleaned restart intent. New attempts retain their frontend +the saved, already-cleaned restart intent. It can also preflight an exact graph +that was stopped before the frontend saved an intent: fresh native observations +must prove absent containers/networks and present retained volumes, and repeat +that proof after review before cleanup. The normal intent, retaining down, +frontend recovery and startup sequence still runs. An active graph keeps its +authenticated service selection and live-listener checks; uncertain observations +cannot select the stopped path. Stopped preflight retains the admitted +image IDs when the original Compose input is unchanged, matching retained startup +instead of resolving mutable tags again. Changed original input keeps normal +image resolution and the existing compatibility checks. New attempts retain their frontend PID and HTTPS port in the private token. Older v1 attempts additionally require `--expect-frontend-pid `; a guessed PID is not recovery evidence. Recovery requires the exact stopped graph receipt with its volumes @@ -576,6 +759,56 @@ Retained starts resolve automatic selections again against current ownership. This allocation does not resize a pool or enable multi-worktree source sharing. Shared HTTPS lifetime uses separate graph leases. +An empty shared HTTPS configuration left by failed startup is not a released +lease. The experimental source API `archiveEmptyNativeHttpsOwner` can archive +only the explicitly selected generation and configuration digest after proving +the original observed spawner exited, the selected endpoint/port/authority and +publications are absent, and the same pool incarnation remains running. It +preserves configuration bytes/inodes and Caddy data. Current startup and archival +share a short admission lock; uncertain intent publication retains its barriers +for inspection. Older binaries do not honor that lock and must remain quiescent +during this explicit recovery. Changed, active or ambiguous state refuses; this +is neither automatic recovery nor a public CLI command. + +A nonempty shared HTTPS owner from the immediately preceding guest boot requires +a separate retaining recovery. The selected graph must already have an exact +current completed dead-owner cleanup proof, a retired publisher, retained volumes +and no compute, bridge, publication or hostname authority. Stop other uses of the +recorded frontend, runtime and Caddy executables first; legacy binaries do not +honor the new admission lock. Recovery refuses a live or uncertain process using +any of these executable paths, including Caddy serving another candidate home. + +Explicit `restart --recover-frontend --expect-finalization-attempt ATTEMPT` can +select this transition for its exact v3 finalization lease. The internal native +operation binds every lease field: + +```sh +./hack-native --candidate-root /absolute/private/candidate-home runtime archive-previous-boot-shared-https \ + --run-id RUN --expect-owner-generation GENERATION --expect-lease-id LEASE \ + --expect-attempt ATTEMPT --expect-owner OWNER --expect-namespace NAMESPACE \ + --expect-plan PLAN --json +``` + +The operation holds shared HTTPS admission, retired-publisher protection and the +provider cleanup lease while checking the immediate boot transition, original +executable hashes, owner/lease files, CA and exclusive port availability. It +archives original owner and lease files by same-filesystem rename and signals no +process. A present control socket and parent must match their selected identities +and be inactive; the socket is renamed within its unchanged parent. If both +socket and parent are absent, +that absence must remain true throughout archival, alongside the independent +process, port and graph proofs. A partial or replaced socket path refuses. + +An interrupted archive retains its journal and admission barrier. Only the exact +selection with a verified dead recovery process can resume; do not remove lock +files or synthesize missing ownership records. Completion preserves Caddy data, +volumes and original lease evidence. Explicit frontend recovery consumes that +completion as a distinct recovery proof, never as an ordinary live-owner lease +release. Keep other shared HTTPS startup quiescent until frontend finalization +completes. A newer active owner blocks completed-archive replay; this recovery +does not adopt or stop it. Archival alone does not start the application or +establish browser access. + Normal foreground startup also selects `--auto-dependency-slots`. Each graph's logical dependency groups receive distinct physical pool slots under the provider lock. The complete assignment is recorded before listeners bind, so concurrent @@ -689,6 +922,58 @@ understand the new completed startup phase. Use the current owning bundle for retaining cleanup and journal archival before selecting an older bundle; do not substitute an older binary against active state. +If a required startup listener exits with a service declared `started` or +`healthy`, the stopped receipt preserves the owned service name, container ID and +exit observation before cleanup, including exit code zero. A successful +`completed` service is not recorded as a startup failure. This evidence contains +no command, environment or application log values. Older bundles may refuse the +new zero-exit evidence; use the owning bundle for diagnosis and retaining cleanup. + +A Docker container can remain `created` with a nonzero error when its first start +fails before a process runs. Retaining shutdown accepts this state only with +verified ownership, no running PID, no pause/restart/dead state and no OOM flag. +The original exit code remains failure evidence; it does not become a successful +application exit or a claim that Hack sent a stop request. + +If that failure interrupted enrolled cleanup, the candidate frontend can inspect +and select a separate retaining recovery after the exact native foreground has +exited. The private controls are `graph inspect-interrupted-start-cleanup +--run-id RUN --json` and `graph recover-interrupted-start-cleanup --run-id RUN +--expect-selection SHA256 --retain-data --json`. They require the same guest +boot, exact dead foreground, matching pending coordinator operation, and unchanged +selected resources and retained volumes. The relay publication must either match +its original pinned owner or have both owner and socket absent after foreground +exit. The absent pair is bound to the existing operation lock, unchanged parent +identity and exact selected coordinator owner, process, publication and attempt; +a mixed pair, live process or replaced lock refuses recovery. This is not a +general retry of a coordinator effect. Its own journal records each cleanup step. +Shared dependency caches stay retained, with their original cache bindings and +volume provenance rechecked at effect and completion boundaries. +An uncertain stop is never sent again merely because the container still appears +running. A crash between stop intent and request can therefore require later +terminal evidence before recovery can advance. +Likewise, a pending guest helper, probe or environment deletion advances only +after complete absence is independently verified. An interruption partway through +a guest cleanup script can remain fenced; this control does not replay that +script or guarantee automatic recovery from every internal deletion boundary. + +Successful recovery requires independent compute, helper, probe and environment +absence, then current cleanup acknowledgement, exact publisher retirement and +dependency reservation release. The frontend rechecks retained data and completes +its existing finalization contract. It still reports the original startup failure. +If retirement is interrupted after acknowledgement, a bounded native journal hint +prompts fresh inspection of the exact original selection before finishing only +the remaining publisher and reservation transitions. The hint grants no effect +authority, and a stopped graph alone does not establish complete retirement. +Completed recovery journals remain as evidence. A later interrupted startup on +the same retained run cannot reuse that earlier generation's selection: native +recovery reports `graph_interrupted_start_cleanup_stale`. Archiving a completed +journal at the verified restore boundary is a separate follow-up; never remove +the journal manually to enable another recovery. +A component acknowledgement is not application readiness, and this recovery does +not diagnose or fix an upstream engine panic. Changed or incomplete evidence +remains available for inspection; do not remove journals or ownership pins to retry. + Native `hack exec` and `hack run` make this request once with a 180-second frontend budget before selecting their command or reading managed environment values. `ps` and `logs` remain observations and do not request refresh. Authenticated relay @@ -915,7 +1200,10 @@ checks succeed. Confirmed graph cleanup preserves the original startup error. If inspection or cleanup cannot be confirmed, the error instead reports retained-state uncertainty alongside the sanitized startup diagnostic. Any published mapping remains intact; inspect owned runtime and bridge state before retrying. A native cleanup -error code is included when available, without its raw output or message. +error code is included when available, without its raw output or message. When a +foreground owner refuses cleanup, the CLI also preserves its bounded cause code +alongside `graph_owner_recovery`; raw owner diagnostics remain omitted. Failed +cleanup is never automatically replayed. ### Explicit cleanup after a dead foreground owner @@ -927,8 +1215,15 @@ operation stops verified guest dependency listeners, removes owned containers an network resources, retains persistent data, retires stale foreground and relay publications, and releases the graph's dependency reservation before unlocking the pool. It does not restart the pool. Pending startup, one-off or dependency -rebind state is refused. Completion records owner-death evidence independently of +rebind state is refused. A completed dependency refresh is accepted only for the +current boot and exact ready receipt generation, with matching terminal services +and helper markers. After owned absence is independently verified, its journal is +archived before stale publications are retired. An interrupted archive resumes +from the committed cleanup proof and exact journal bytes; it never replays refresh. +Completion records owner-death evidence independently of the live relay acknowledgement protocol. +When a restored publisher is admitted, older absence-retirement records remain +validated history; the exact current cleanup proof supplies retention authority. After completion, retained restore can create a fresh owner; `graph retire-recovered-publisher --run-id RUN --expect-owner OWNER` remains an @@ -940,12 +1235,33 @@ requirements. A graph with completed host-dependency startup and no inbound routes can also use its exact dead relay publication as the previous-boot proof; an empty bridge registry alone never grants cleanup authority. -A completed previous-boot recovery may remain as historical evidence after restore. -Its exact stopped receipt must still be in the bounded restore history, and the -current graph must have a different recorded container generation. Incomplete, -foreign, one-off, altered or no-longer-verifiable records refuse without modifying -them. The historical proof never authorizes cleanup of current resources; current -owner, boot, process and resource checks remain required, including on retries. +A completed previous-boot or same-boot recovery may remain as historical evidence +after restore. +When its exact stopped receipt remains in bounded restore history, that receipt +must match the completion proof. After history eviction, only a self-consistent +proof for the same graph and a different container generation can be treated as +inert history; the verified truncated history must independently contain a later +stopped generation. This does not establish the evicted completion hash as current +authority. Same-boot history also requires validated listener retirement. +Incomplete, foreign, one-off, conflicting or missing-history records +refuse without modification. Current owner, boot, process, receipt and resource +checks still authorize cleanup independently, including on retries. +A later previous-boot recovery may archive that validated older completion even +after its exact stopped receipt has been evicted. It preserves the original proof +bytes and inode by renaming it; it does not reconstruct an evicted receipt or use +historical evidence as current cleanup authority. Pending or incomplete proofs, +changed graph identities and an unchanged container generation still refuse. +Retention and recovered-publisher retirement prefer a completed previous-boot +proof that matches the exact current stopped receipt. An older same-boot proof +must still pass historical validation; its presence cannot redirect these +operations away from the current proof. Pending, incomplete, malformed or +conflicting same-boot records continue to refuse both operations. +An already completed prior-boot dependency archive is likewise historical +metadata. Restore may leave it untouched after validating its original generation, +boot, journal and exact artifact hashes, with no active or pending journal and an +independently confirmed current retention proof. Eviction of its old stopped +receipt does not require repeating completed archival. First-time or interrupted +archival still requires the exact selected cleanup proof and ownership checks. The same-boot completion proof does not authorize direct data removal: restore the graph and use ordinary cleanup for that operation. Native routed recovery has been qualified with repeated owner crashes, retained data and an unaffected sibling's diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 164130b76..ccf439c3f 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -1263,6 +1263,10 @@ hack doctor [options] | `--json` | Output JSON (machine-readable) | | `--browser-url https://app.hack` | HTTPS origin manually tested in the browser (no path or credentials) | | `--browser-result unknown|works|fails|permission-denied` | Your manual browser observation for --browser-url (default: unknown) | +| `--branch ` | Run against a branch-specific instance (compose name + hostnames) | +| `--native-run-mapping inspect|repair` | Inspect or explicitly repair a native run mapping after filesystem device renumbering | +| `--expect-selection <64-hex>` | Require the exact run-mapping recovery inspection selection | +| `--accept-legacy-device-rebind` | Explicitly accept legacy migration without proof of original filesystem volume continuity | | `--no-interactive` | Never prompt: apply documented defaults or fail with E_INTERACTIVE_REQUIRED (also via HACK_NO_INTERACTIVE=1) | | `--help, -h` | Show help | | `--version, -v` | Show version | diff --git a/packages/runtime-core/README.md b/packages/runtime-core/README.md index dd830284f..dfb45d5ce 100644 --- a/packages/runtime-core/README.md +++ b/packages/runtime-core/README.md @@ -259,6 +259,21 @@ state or changed provider identity refuses recovery. VM disks and graph data rem This is cooperative same-user recovery for a stopped pool, not adoption of an unreceipted active relay or proof of application readiness. +For an already-running VM whose retained graphs are all stopped, the separate +`hack-local runtime quiescent-dependency-socket-recovery --json` command selects +legacy unreceipted socket inodes without rebooting. Explicitly apply its selection +with `hack-local runtime recover-quiescent-dependency-sockets --expect-sha256 --json`. +This mode holds a pool publication gate before foreground retirement locks and the +provider lease. It pins the exact provider receipt, host/guest boot, every retained +graph receipt and volume identity, and requires absent compute, no foreground +publisher, no dependency assignment and no untracked guest containers. It rechecks +that proof before each inode-selected unlink and journals the selection separately +from stopped-pool recovery so partial retries retain their original scope. +An incomplete journal, live listener, replacement path or changed proof refuses +recovery. Owner bytes, source-rebind witnesses, retained data and the running VM +remain unchanged. This is explicit cooperative legacy cleanup, not proof of the +original socket creator or a security boundary against another same-user process. + TERM and INT are checked during initial graph startup, including readiness waits and before new service effects. Cancellation enters owned cleanup while preserving persistent data. Checks occur between bounded operations; an in-flight operation @@ -463,6 +478,10 @@ combined filesystem capacity, excluding kernel overhead). Partial paths also con capacity. Graph admission checks the complete requested batch under its runtime lease before graph effects; staging checks again. Verified retirement frees guest capacity, while immutable intent history retains its separate 4096-entry bound. +Graph cleanup uses that same history capacity, counting active and archived +records together, so repeated restores cannot exceed a smaller cleanup-only +ceiling after allocation succeeds. Matching ownership, container bindings and +unambiguous records are still required. No retained evidence is automatically deleted to admit new work. This environment budget does not increase the runtime's separate container, CPU or memory limits and does not establish 32-branch qualification. @@ -803,3 +822,189 @@ before fresh dependency ownership is admitted. This does not enable file-based `graph restart`/`restore` or normalized live-owner redelivery, and does not replay initializer cache-release effects. Native same-run restoration and cross-boot volume continuity require separate runtime qualification. + +An explicitly recovered post-reboot graph may retain a stopped shared-source +receipt whose `source.shared.device` names the prior host filesystem device. +After **completed** selected absent-publication cleanup and separate publisher +retirement, inspect that same run with `graph inspect-source-device-rebind +--run-id RUN --json`. Review the returned original stopped-receipt hash and +qualification, then commit only that selection with `graph +recover-source-device-rebind --run-id RUN --expect-selection SHA +--accept-legacy-device-rebind --json`. This private candidate command changes no +stopped receipt, cleanup proof, retirement proof or named volume. It writes one +immutable witness for the current Owner share, requiring the same canonical +project path, inode, guest path and unfiltered access, with only the host device +number changed. The current guest's writable virtiofs mount and selected volume +identities are checked independently. The explicit acceptance records that +legacy receipts cannot establish original physical volume continuity across the +host reboot. + +An exact completed cleanup and retirement for the current stopped receipt takes +precedence over older cleanup records. Pending operations, unresolved initializer +effects and mismatched receipt or retirement identities still refuse recovery; +historical records cannot authorize a later generation. + +When a normal foreground shutdown has already removed its publisher endpoints, +`graph retire-recovered-publisher` confirms the current boot's acknowledged +cleanup under the publisher lock. It checks the exact stopped receipt, cleanup +effect, guest inventories, retained volumes and absent endpoints. It changes no +publisher files or graph data. A confirmed but invalid acknowledgement refuses; +historical crash-recovery receipts cannot override it. + +Ordinary cleanup can be acknowledged before a later dependency-journal archival +step fails. If this leaves a dead foreground publication for the current stopped +graph, use the explicit private `graph retire-acknowledged-publisher --run-id RUN +--expect-owner OWNER --expect-receipt RECEIPT_SHA --expect-publisher PUBLISHER_SHA +--json` recovery. The SHA selectors identify the exact stopped receipt and +foreground owner-record bytes. Admission requires the current guest boot's +confirmed cleanup acknowledgement, no pending operation, independent absence of +compute by both immutable ID and reserved name, and all named volumes present. +Environment, startup, probe and bridge cleanup inventories are rechecked while +holding the publisher and provider locks, including at each publication effect. +The existing retirement journal preserves both original publication files and +resumes an interrupted socket-first rename. Active owners, replacement listeners, +stale selectors and incomplete acknowledgements refuse without falling back to +historical recovery. This operation changes no graph receipt, VM, dependency +journal or retained data; it does not authorize a restore by itself. +An old completed or cleanup-completed dependency journal is accepted as inert +evidence only; malformed, incomplete or pending rebind state refuses at each +effect boundary. Safe archival and the next restore remain separate checks. + +If ordinary cleanup closed all dependency sockets but a later archival failure +left its reservation, first complete that exact publisher retirement. Inspect +`graph dependency-reservations --json`, then use `graph +release-acknowledged-dependencies --run-id RUN --expect-owner OWNER +--expect-receipt RECEIPT_SHA --expect-publisher PUBLISHER_SHA +--expect-reservation RESERVATION_SHA --json`. The final selector is the inspected +reservation fingerprint. This requires the current boot's acknowledged cleanup, +the exact completed publisher retirement, a dead reservation process matching +that publisher, the same run/Owner/boot and precisely the receipt's dependency +slots. Every selected socket must already be absent; even a matching stale socket +refuses. The selected record moves atomically without overwriting into the graph's +`dependency-reservation-retired-SHA.json` history, preserving bytes and inode. +Retry verifies that same record and proof. Changed selectors, another publisher, +pending/foreign records, occupied history and any reappearing socket refuse. +No socket, VM, graph receipt, dependency journal, data or sibling is removed. +This explicit recovery frees the claim; normal `up` still performs its own checks. + +`restore-selection` then binds the witness's raw hash to the selected generation. +The source checks use only an in-memory device projection; original retention and +restore history consume the unchanged stopped receipt. The new receipt receives +the projected source after history retention, before the first new-attempt write. +The witness file remains pinned and is rechecked before that write. Once the +first new generation replaces the stopped receipt, the witness is audit history +and grants no authority to later ordinary down/up cycles. An incomplete +`source-device-rebind.pending` file, including an empty or partial write, is +preserved and refuses both explicit recovery and a new retired-publisher bind; +there is no automatic prefix-based adoption or deletion. The ignored native +synthetic-device fixture checks this same-run transition and later retained +marker reads, but only an actual host-reboot application run can qualify the +physical continuity claim. + +A selected source-device witness also preserves an existing dependency cache's +namespace when the Git common directory is on that same filesystem. The runtime +reconstructs the old scope hash from the unchanged common directory path/inode +and selected prior device number, and requires every retained cache scope to +match. Only the new attempt receives optional cache-scope continuity metadata; +the stopped receipt, original cache names/fingerprints, volumes and witness bytes +stay unchanged. All package inputs, image, command, environment, host mappings +and volume layouts are still freshly fingerprinted, with strict resource equality. +This is not a fallback for changed lockfiles or an unrelated repository. + +Later ordinary replays verify the immutable witness's raw hash and run/owner/share +linkage, the current common directory path/device/inode and retained scopes. +Missing, changed or pending provenance and replaced Git metadata refuse. A second +device transition requires a separately supported explicit recovery; the runtime +does not silently extend this projection. Older candidate binaries may refuse the +new optional source metadata, so a downgrade must be qualified separately. + +Completed post-reboot absence recovery also retires the prior boot's completed +dependency-rebind journal. A retained foreground restore can finish this archival +for an earlier candidate: it selects the exact original ready receipt and exact +completed stopped receipt from bounded history, verifies durable retirement, +unchanged provider/boot and retained volumes, and proves old and current compute +absent under the held foreground lock and provider lease. The journal's completed +generation must match the original receipt. Original bytes remain in a digest-bound, +resumable archive; a later current-boot journal is preserved for its own cleanup. +Changed evidence, pending state or live compute refuses archival. This does not +replay old dependencies or relax the fresh graph's endpoint/readiness checks. + +The ignored native regression +`foreground::native_test::retired_rebind_history::completed_prior_boot_rebind_archives_after_newer_stopped_generation` +uses an isolated capacity-two candidate, a pinned pre-archive native executable +(`7dc0811cd0179f40340e48c540699933dc099cf7`) and a current executable with +`retire-acknowledged-publisher` and `release-acknowledged-dependencies`. +The candidate must already be prepared and running with the pinned image loaded. +Supply `HACK_LOCAL_TEST_ROOT`, +`HACK_LOCAL_TEST_BINARY`, `HACK_LOCAL_TEST_LEGACY_BINARY`, +`HACK_LOCAL_TEST_LEGACY_SHA256`, `HACK_LOCAL_TEST_IMAGE`, +`HACK_GRAPH_RELAY_ARTIFACT` and `HACK_GRAPH_RELAY_SHA256`. The pinned image must +contain `/usr/local/bin/bun` for the HTTP server and retained marker. Precompile the +test, then run it under an external 300-second watchdog with one test thread. +Both executables must accept that exact candidate root; checkout-bound source +binaries must be built for it, or use relocatable `hack-native` executables with +a dedicated private root. The test stops and restarts the entire supplied pool, +so the root must contain only this fixture, with no application or sibling-owner +resources. Its host listener uses an ephemeral loopback port. Project and private +inputs live beneath that declared root and remain on failure, along with readiness +stage labels and bounded foreground failure output. The test has no fallback +`--remove-data` cleanup on unwind. Only its successful explicit graph cleanup +removes data and disposes its own input directories; the external harness must +stop the exact owned pool and preserve evidence after failure. +The prior-host-boot owner and selected dependency reservation are synthetic: +after proving the exact reservation owner dead and every recorded socket inactive +and unchanged, test-only setup projects its process timestamp and socket device +numbers. It preserves the real completed journal, socket inodes, bindings, graph +receipt and sibling reservations. Production inspection still rejects current-boot +timestamps, mismatched device/inode evidence and live owners or listeners. +It uses normal refresh, recovery, cleanup and restore paths, plus an injected +archive interruption at the owned verifier boundary. The original owner crash and +legacy archival refusal remain controls. Only the later legacy foreground owner +receives graceful SIGTERM: the test requires normal handler exit, process death +and exact dependency socket absence before independently retiring its publisher +and reservation through acknowledged cleanup. Reservation history must preserve +the original bytes and inode, with the selected active claim absent and sibling +state unchanged. A passing test proves +same-run retained marker and sibling isolation in that fixture; it does not +establish application or physical host-reboot acceptance. + +An interrupted legacy HTTPS frontend can leave its exact socket, Caddy receipt +and lock alongside an unpublished shared-owner configuration. Explicit private +`runtime recover-quiescent-https --expect-owner RECEIPT_SHA +--expect-configuration CONFIG_SHA --expect-frontend-pid OBSERVED_DEAD_PID --json` +archives this combined incident without signaling processes or changing CA data. +It requires the current pool/boot, strict configuration, exact authority path, +dead recorded processes and configured executables, no shared leases or release +history for that generation, an inactive Unix socket, and exclusive IPv4/IPv6 +wildcard and loopback listeners held across each move. The legacy receipt and +unpublished configuration must select the same Caddy path, binary hash and port. +Receipt/config selectors identify raw file bytes; the CA must retain its original DER fingerprint. A durable journal +and exclusive renames preserve original inodes and permit exact partial retries; +foreign targets, changed evidence and live or uncertain effects refuse. A crash +during initial journal publication may leave an incomplete journal: it refuses +retry before archival rather than guessing or discarding that evidence. Global +HTTPS evidence archival does not itself recover a project finalization token or +prove application readiness: run the explicit frontend recovery and normal +startup checks afterward. + +Legacy receipts may predate a host filesystem device-number change. The strict +command refuses those records. An explicit additional +`--accept-legacy-device-rebind RUN:WITNESS_SHA:SOCKET_DEV:SOCKET_INO:LOCK_DEV:LOCK_INO` +selects the existing graph source-device witness and both current HTTPS identities. +Both old receipt devices must match its old device; both current devices must match +its current device, with each respective inode unchanged. Current owner bytes, +host/guest boot, project share, stopped graph scope and witness bytes are rechecked +before each effect. Mixed changes refuse. The journal pins the current identities; +the legacy receipt stays byte-identical. This is explicit legacy migration with +**original host-volume continuity unproven**, not an automatic identity relaxation. + +Development admission distinguishes fresh capacity from a verified running pool. +A fresh VM retains the 58 GiB disk floor (32 GiB storage, 10 GiB overlay and +16 GiB host reserve). Reusing the exact owned running pool requires the 16 GiB +host reserve, with memory-pressure, memory-headroom and thermal checks unchanged. +`runtime probe --profile development --json` reports `disk_budget_basis`. Reuse +requires matching profile, creation receipt, native process/PID-file identity, +machine name and both disks' current identities and declared sizes. Every sample +and the acquired startup lease recheck the selected owner; changed ownership +refuses, and a reserve-qualified request cannot enter VM create or boot. Stopped, +missing or unproved capacity retains fresh-allocation requirements or refuses. diff --git a/packages/runtime-core/src/graph_cli.rs b/packages/runtime-core/src/graph_cli.rs index 9b9e3b52c..16c450cb2 100644 --- a/packages/runtime-core/src/graph_cli.rs +++ b/packages/runtime-core/src/graph_cli.rs @@ -25,6 +25,310 @@ pub fn command(candidate: &Candidate, args: &[&str]) -> Result run, + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::inspect_interrupted_start_cleanup(candidate, run); + } + #[cfg(not(target_os = "macos"))] + { + let _ = run; + return Err(invalid()); + } + } + if *action == "recover-interrupted-start-cleanup" { + let (run, expected) = match *args { + [ + "--run-id", + run, + "--expect-selection", + expected, + "--retain-data", + ] + | [ + "--run-id", + run, + "--expect-selection", + expected, + "--retain-data", + "--json", + ] => (run, expected), + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::recover_interrupted_start_cleanup(candidate, run, expected); + } + #[cfg(not(target_os = "macos"))] + { + let _ = (run, expected); + return Err(invalid()); + } + } + if *action == "release-acknowledged-dependencies" { + let (run, owner, receipt, publisher, reservation) = match *args { + [ + "--run-id", + run, + "--expect-owner", + owner, + "--expect-receipt", + receipt, + "--expect-publisher", + publisher, + "--expect-reservation", + reservation, + ] + | [ + "--run-id", + run, + "--expect-owner", + owner, + "--expect-receipt", + receipt, + "--expect-publisher", + publisher, + "--expect-reservation", + reservation, + "--json", + ] => (run, owner, receipt, publisher, reservation), + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::release_acknowledged_dependencies( + candidate, + graph::AcknowledgedPublisherSelection { + run, + owner, + receipt_sha256: receipt, + publisher_sha256: publisher, + }, + reservation, + ); + } + #[cfg(not(target_os = "macos"))] + { + let _ = (run, owner, receipt, publisher, reservation); + return Err(invalid()); + } + } + if *action == "retire-acknowledged-publisher" { + let (run, owner, receipt, publisher) = match *args { + [ + "--run-id", + run, + "--expect-owner", + owner, + "--expect-receipt", + receipt, + "--expect-publisher", + publisher, + ] + | [ + "--run-id", + run, + "--expect-owner", + owner, + "--expect-receipt", + receipt, + "--expect-publisher", + publisher, + "--json", + ] => (run, owner, receipt, publisher), + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::retire_acknowledged_publisher( + candidate, + graph::AcknowledgedPublisherSelection { + run, + owner, + receipt_sha256: receipt, + publisher_sha256: publisher, + }, + ); + } + #[cfg(not(target_os = "macos"))] + { + let _ = (run, owner, receipt, publisher); + return Err(invalid()); + } + } + if *action == "inspect-host-pin-recovery" { + let run = match *args { + ["--run-id", run] | ["--run-id", run, "--json"] => run, + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::inspect_host_pin_recovery(candidate, run); + } + #[cfg(not(target_os = "macos"))] + { + let _ = run; + return Err(invalid()); + } + } + if *action == "recover-host-pins" { + let (run, expected) = match *args { + [ + "--run-id", + run, + "--expect-selection", + expected, + "--accept-legacy-device-rebind", + ] + | [ + "--run-id", + run, + "--expect-selection", + expected, + "--accept-legacy-device-rebind", + "--json", + ] => (run, expected), + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::recover_host_pins(candidate, run, expected); + } + #[cfg(not(target_os = "macos"))] + { + let _ = (run, expected); + return Err(invalid()); + } + } + if *action == "inspect-absent-publication-cleanup" { + let (run, original, inspection) = match *args { + [ + "--run-id", + run, + "--original-owner-file", + original, + "--host-inspection-file", + inspection, + ] + | [ + "--run-id", + run, + "--original-owner-file", + original, + "--host-inspection-file", + inspection, + "--json", + ] => (run, original, inspection), + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::inspect_absent_publication_cleanup( + candidate, + run, + Path::new(original), + Path::new(inspection), + ); + } + #[cfg(not(target_os = "macos"))] + { + let _ = (run, original, inspection); + return Err(invalid()); + } + } + if *action == "recover-absent-publication-cleanup" { + let (run, original, inspection, expected) = match *args { + [ + "--run-id", + run, + "--original-owner-file", + original, + "--host-inspection-file", + inspection, + "--expect-selection", + expected, + "--retain-data", + "--accept-unpinned-post-reboot", + ] + | [ + "--run-id", + run, + "--original-owner-file", + original, + "--host-inspection-file", + inspection, + "--expect-selection", + expected, + "--retain-data", + "--accept-unpinned-post-reboot", + "--json", + ] => (run, original, inspection, expected), + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::recover_absent_publication_cleanup( + candidate, + run, + expected, + Path::new(original), + Path::new(inspection), + ); + } + #[cfg(not(target_os = "macos"))] + { + let _ = (run, original, inspection, expected); + return Err(invalid()); + } + } + if *action == "inspect-source-device-rebind" { + let run = match *args { + ["--run-id", run] | ["--run-id", run, "--json"] => run, + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::inspect_source_device_rebind(candidate, run); + } + #[cfg(not(target_os = "macos"))] + { + let _ = run; + return Err(invalid()); + } + } + if *action == "recover-source-device-rebind" { + let (run, expected) = match *args { + [ + "--run-id", + run, + "--expect-selection", + expected, + "--accept-legacy-device-rebind", + ] + | [ + "--run-id", + run, + "--expect-selection", + expected, + "--accept-legacy-device-rebind", + "--json", + ] => (run, expected), + _ => return Err(invalid()), + }; + #[cfg(target_os = "macos")] + { + return graph::recover_source_device_rebind(candidate, run, expected); + } + #[cfg(not(target_os = "macos"))] + { + let _ = (run, expected); + return Err(invalid()); + } + } if ["run-selection", "run-service"].contains(action) { return one_off::command(candidate, action, args); } diff --git a/packages/runtime-core/src/graph_cli/restore_tests.rs b/packages/runtime-core/src/graph_cli/restore_tests.rs index fc407a4bb..d4f1d8df3 100644 --- a/packages/runtime-core/src/graph_cli/restore_tests.rs +++ b/packages/runtime-core/src/graph_cli/restore_tests.rs @@ -5,6 +5,84 @@ const RUN: &str = "11111111111111111111111111111111"; const PLAN: &str = "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"; const GENERATION: &str = "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"; +#[test] +fn acknowledged_dependency_release_requires_all_selectors_and_forbids_other_effects() { + let fixture = Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let args = vec![ + "release-acknowledged-dependencies", + "--run-id", + RUN, + "--expect-owner", + RUN, + "--expect-receipt", + PLAN, + "--expect-publisher", + GENERATION, + "--expect-reservation", + PLAN, + ]; + for offset in [1, 3, 5, 7, 9] { + let mut missing = args.clone(); + missing.drain(offset..offset + 2); + assert_eq!( + command(&candidate, &missing).unwrap_err().code, + "graph_arguments" + ); + } + for effect in [ + "--remove-data", + "--environment-stdin", + "--accept-legacy-device-rebind", + ] { + let mut expanded = args.clone(); + expanded.push(effect); + assert_eq!( + command(&candidate, &expanded).unwrap_err().code, + "graph_arguments" + ); + } + assert!(!candidate.state_root.exists()); +} + +#[test] +fn acknowledged_retirement_requires_every_selector_and_forbids_other_effects() { + let fixture = Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let args = vec![ + "retire-acknowledged-publisher", + "--run-id", + RUN, + "--expect-owner", + RUN, + "--expect-receipt", + PLAN, + "--expect-publisher", + GENERATION, + ]; + for offset in [1, 3, 5, 7] { + let mut missing = args.clone(); + missing.drain(offset..offset + 2); + assert_eq!( + command(&candidate, &missing).unwrap_err().code, + "graph_arguments" + ); + } + for effect in [ + "--remove-data", + "--accept-legacy-device-rebind", + "--environment-stdin", + ] { + let mut expanded = args.clone(); + expanded.push(effect); + assert_eq!( + command(&candidate, &expanded).unwrap_err().code, + "graph_arguments" + ); + } + assert!(!candidate.state_root.exists()); +} + #[test] fn same_boot_recovery_requires_explicit_receipt_and_forbids_data_removal() { let fixture = Fixture::new(); diff --git a/packages/runtime-core/src/main.rs b/packages/runtime-core/src/main.rs index 68bdc05ed..1bf27c77f 100644 --- a/packages/runtime-core/src/main.rs +++ b/packages/runtime-core/src/main.rs @@ -27,6 +27,14 @@ Usage: hack-local graph recover-cleanup --run-id <32-hex> --expect-receipt [--json] hack-local graph recover-live-owner --run-id <32-hex> --expect-receipt [--json] hack-local graph retire-recovered-publisher --run-id <32-hex> --expect-owner <32-hex> [--json] + hack-local graph retire-acknowledged-publisher --run-id <32-hex> --expect-owner <32-hex> --expect-receipt <64-hex> --expect-publisher <64-hex> [--json] + hack-local graph release-acknowledged-dependencies --run-id <32-hex> --expect-owner <32-hex> --expect-receipt <64-hex> --expect-publisher <64-hex> --expect-reservation <64-hex> [--json] + hack-local graph inspect-host-pin-recovery --run-id <32-hex> [--json] + hack-local graph recover-host-pins --run-id <32-hex> --expect-selection <64-hex> --accept-legacy-device-rebind [--json] + hack-local graph inspect-absent-publication-cleanup --run-id <32-hex> --original-owner-file --host-inspection-file [--json] + hack-local graph recover-absent-publication-cleanup --run-id <32-hex> --original-owner-file --host-inspection-file --expect-selection <64-hex> --retain-data --accept-unpinned-post-reboot [--json] + hack-local graph inspect-source-device-rebind --run-id <32-hex> [--json] + hack-local graph recover-source-device-rebind --run-id <32-hex> --expect-selection <64-hex> --accept-legacy-device-rebind [--json] hack-local graph logs --run-id <32-hex> --service [--tail <1..1000>] [--json] hack-local graph exec --run-id <32-hex> --service [--workdir /path] [--timeout-seconds <1..120>] [--json] -- [args...] hack-local graph dependency-plan --dependencies [--json] @@ -42,6 +50,9 @@ Usage: hack-local runtime hostname-authority --socket [--json] hack-local runtime recover-hostname-authority --socket --expect-sha256 [--json] hack-local runtime managed-hostname-authority [--json] + hack-local runtime recover-quiescent-https --expect-owner <64-hex> --expect-configuration <64-hex> --expect-frontend-pid --json + hack-local runtime archive-previous-boot-shared-https --run-id <32-hex> --expect-owner-generation <32-hex> --expect-lease-id <32-hex> --expect-attempt <32-hex> --expect-owner <32-hex> --expect-namespace <64-hex> --expect-plan <64-hex> --json + hack-local runtime recover-quiescent-https --expect-owner <64-hex> --expect-configuration <64-hex> --expect-frontend-pid --accept-legacy-device-rebind --json hack-local runtime inspect-host-listener --pid --port --executable [--peer-port ] [--json] hack-local runtime serve-managed-hostnames [--certificate-name-limit <1..4096>] (owner pipe on stdin) hack-local runtime certificate-admission [--json] @@ -53,6 +64,8 @@ Usage: hack-local runtime publication-recovery [--json] hack-local runtime recover-publications --expect-sha256 [--json] hack-local runtime dependency-socket-recovery [--json] + hack-local runtime quiescent-dependency-socket-recovery [--json] + hack-local runtime recover-quiescent-dependency-sockets --expect-sha256 [--json] hack-local runtime recover-dependency-sockets --expect-sha256 [--json] hack-local runtime bridge-recovery [--json] hack-local runtime export-bridge-recovery --slot <1..8> --expect-sha256 [--json] @@ -86,6 +99,8 @@ Usage: hack-local runtime up --profile development --internet --json hack-local runtime network extend --allow-host [--allow-host ] --json hack-local runtime up|status|down|recover [--json] + hack-local runtime host-filesystem-recovery [--json] + hack-local runtime recover-host-filesystem --expect-sha256 --accept-legacy-device-rebind [--json] hack-local node serve|status|inspect hack-local node request hack-local --version @@ -454,10 +469,7 @@ fn run() -> Result<(), CandidateError> { }; let candidate = discover_candidate(&requested)?; if *action == "probe" { - print_json(&provider::admission::probe_for( - &candidate.checkout, - profile, - )?)?; + print_json(&provider::probe_with_profile(&candidate, profile)?)?; } else { print_json(&provider::up_with_profile(&candidate, profile)?)?; } @@ -502,6 +514,89 @@ fn run() -> Result<(), CandidateError> { )?, )?; } + #[cfg(target_os = "macos")] + [ + "runtime", + "archive-previous-boot-shared-https", + "--run-id", + run, + "--expect-owner-generation", + generation, + "--expect-lease-id", + lease, + "--expect-attempt", + attempt, + "--expect-owner", + owner, + "--expect-namespace", + namespace, + "--expect-plan", + plan, + "--json", + ] => { + print_json( + &hack_runtime_core::provider::shared_https_recovery::archive( + &discover_candidate(&requested)?, + hack_runtime_core::provider::shared_https_recovery::ArchiveSelection { + run, + owner_generation: generation, + lease_id: lease, + attempt, + owner, + namespace, + plan, + }, + )?, + )?; + } + [ + "runtime", + "recover-quiescent-https", + "--expect-owner", + owner, + "--expect-configuration", + configuration, + "--expect-frontend-pid", + pid, + "--json", + ] => { + print_json(&hack_runtime_core::provider::https_recovery::recover( + &discover_candidate(&requested)?, + owner, + configuration, + pid.parse().map_err(|_| { + CandidateError::new("invalid_arguments", "Expected a positive frontend PID.") + })?, + )?)?; + } + [ + "runtime", + "recover-quiescent-https", + "--expect-owner", + owner, + "--expect-configuration", + configuration, + "--expect-frontend-pid", + pid, + "--accept-legacy-device-rebind", + selection, + "--json", + ] => { + print_json( + &hack_runtime_core::provider::https_recovery::recover_legacy_device_rebind( + &discover_candidate(&requested)?, + owner, + configuration, + pid.parse().map_err(|_| { + CandidateError::new( + "invalid_arguments", + "Expected a positive frontend PID.", + ) + })?, + selection, + )?, + )?; + } ["runtime", "managed-hostname-authority"] | ["runtime", "managed-hostname-authority", "--json"] => { print_json( @@ -672,6 +767,34 @@ fn run() -> Result<(), CandidateError> { )?, )?; } + ["runtime", "quiescent-dependency-socket-recovery"] + | ["runtime", "quiescent-dependency-socket-recovery", "--json"] => { + print_json( + &hack_runtime_core::provider::quiescent_dependency_socket_recovery::inspect( + &discover_candidate(&requested)?, + )?, + )?; + } + [ + "runtime", + "recover-quiescent-dependency-sockets", + "--expect-sha256", + hash, + ] + | [ + "runtime", + "recover-quiescent-dependency-sockets", + "--expect-sha256", + hash, + "--json", + ] => { + print_json( + &hack_runtime_core::provider::quiescent_dependency_socket_recovery::recover( + &discover_candidate(&requested)?, + hash, + )?, + )?; + } [ "runtime", "recover-dependency-sockets", @@ -857,6 +980,32 @@ fn run() -> Result<(), CandidateError> { image, )?)?; } + ["runtime", "host-filesystem-recovery"] + | ["runtime", "host-filesystem-recovery", "--json"] => { + print_json(&hack_runtime_core::provider::host_filesystem::inspect( + &discover_candidate(&requested)?, + )?)?; + } + [ + "runtime", + "recover-host-filesystem", + "--expect-sha256", + hash, + "--accept-legacy-device-rebind", + ] + | [ + "runtime", + "recover-host-filesystem", + "--expect-sha256", + hash, + "--accept-legacy-device-rebind", + "--json", + ] => { + print_json(&hack_runtime_core::provider::host_filesystem::recover( + &discover_candidate(&requested)?, + hash, + )?)?; + } ["runtime", action] | ["runtime", action, "--json"] => { let candidate = discover_candidate(&requested)?; let result = match *action { diff --git a/packages/runtime-core/src/provider/admission.rs b/packages/runtime-core/src/provider/admission.rs index 6b89befda..be451deb6 100644 --- a/packages/runtime-core/src/provider/admission.rs +++ b/packages/runtime-core/src/provider/admission.rs @@ -42,6 +42,7 @@ pub struct Admission { pub memory_pressure_normal: bool, pub disk_free_bytes: Option, pub minimum_disk_free_bytes: u64, + pub disk_budget_basis: &'static str, pub one_minute_load: Option, pub load_ceiling: Option, pub thermal_normal: bool, @@ -167,6 +168,37 @@ pub fn probe(path: &Path) -> Result { } pub fn probe_for(path: &Path, profile: Profile) -> Result { + probe_with_disk_budget(path, profile, false) +} + +/// Only the lifecycle's independently verified live development pool may use this budget. +pub(super) fn probe_owned_development(path: &Path) -> Result { + probe_with_disk_budget(path, Profile::Development, true) +} + +fn disk_budget(profile: Profile, owned_live: bool) -> (u64, &'static str) { + if profile == Profile::Development && owned_live { + ( + DEVELOPMENT_HOST_DISK_RESERVE_GIB * 1024 * 1024 * 1024, + "verified-live-owned-pool-host-reserve", + ) + } else { + ( + disk_floor_bytes(profile), + if profile == Profile::Research { + "research-floor" + } else { + "new-pool-storage-overlay-host-reserve" + }, + ) + } +} + +fn probe_with_disk_budget( + path: &Path, + profile: Profile, + owned_live: bool, +) -> Result { let supported = cfg!(all(target_os = "macos", target_arch = "aarch64")); let mut result = Admission { profile, @@ -181,7 +213,8 @@ pub fn probe_for(path: &Path, profile: Profile) -> Result Result "Free disk is below the 100 GiB research floor.", + Profile::Development if owned_live => { + "Free disk is below the host reserve for the verified running development pool." + } Profile::Development => { "Free disk is below the development VM storage and host reserve budget." } @@ -301,10 +337,17 @@ pub fn qualification(path: &Path) -> Result, CandidateError> { } pub fn sample_for(path: &Path, profile: Profile) -> Result, CandidateError> { + sample_with(profile, || probe_for(path, profile)) +} + +pub(super) fn sample_with( + profile: Profile, + mut observe: impl FnMut() -> Result, +) -> Result, CandidateError> { let mut samples: Vec = Vec::new(); for index in 0..3 { - let sample = probe_for(path, profile)?; - if !sample.admitted { + let sample = observe()?; + if !sample.admitted || sample.profile != profile { return Err(CandidateError::new( "admission_rejected", sample.reasons.join(" "), @@ -340,6 +383,9 @@ mod tests { const GIB: u64 = 1024 * 1024 * 1024; assert_eq!(disk_floor_bytes(Profile::Research), 100 * GIB); assert_eq!(disk_floor_bytes(Profile::Development), 58 * GIB); + assert_eq!(disk_budget(Profile::Development, true).0, 16 * GIB); + assert_eq!(disk_budget(Profile::Research, true).0, 100 * GIB); + assert_eq!(disk_budget(Profile::Development, false).0, 58 * GIB); assert_eq!(load_ceiling(Profile::Research), Some(8.0)); assert_eq!(load_ceiling(Profile::Development), None); } diff --git a/packages/runtime-core/src/provider/dependency_socket.rs b/packages/runtime-core/src/provider/dependency_socket.rs index 0b2bfd157..071cac20f 100644 --- a/packages/runtime-core/src/provider/dependency_socket.rs +++ b/packages/runtime-core/src/provider/dependency_socket.rs @@ -9,6 +9,7 @@ use std::{ path::{Path, PathBuf}, }; +pub mod quiescent_recovery; pub mod recovery; #[derive(Clone, Copy, Debug, PartialEq, Eq, Serialize, Deserialize)] diff --git a/packages/runtime-core/src/provider/dependency_socket/quiescent_recovery.rs b/packages/runtime-core/src/provider/dependency_socket/quiescent_recovery.rs new file mode 100644 index 000000000..f2b154a72 --- /dev/null +++ b/packages/runtime-core/src/provider/dependency_socket/quiescent_recovery.rs @@ -0,0 +1,267 @@ +//! Explicit legacy socket retirement while an independently verified VM stays live. +//! This uses a separate receipt from stopped-pool recovery. It never proves who +//! created an unreceipted socket; the reviewed selection authorizes only those inodes. +#[cfg(any(test, target_os = "macos"))] +use super::{ + DependencySocketIntent, + recovery::{Socket, digest, hex, observed as observe_socket, path}, +}; +#[cfg(any(test, target_os = "macos"))] +use crate::provider::state; +use crate::{Candidate, CandidateError}; +#[cfg(any(test, target_os = "macos"))] +use serde::{Deserialize, Serialize}; +use serde_json::Value; +#[cfg(any(test, target_os = "macos"))] +use serde_json::json; +#[cfg(any(test, target_os = "macos"))] +use std::{ + fs::{self, File}, + path::{Path, PathBuf}, +}; + +#[cfg(any(test, target_os = "macos"))] +const MAX_RECEIPTS: usize = 16; +#[cfg(any(test, target_os = "macos"))] +const RECEIPT_LIMIT: u64 = 16 * 1024; + +fn refused() -> CandidateError { + CandidateError::new( + "quiescent_dependency_socket_recovery", + "Live pool quiescence or selected socket identity is uncertain; paths and data were retained.", + ) +} + +#[cfg(any(test, target_os = "macos"))] +#[derive(Clone, Debug, Deserialize, Serialize, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +struct Scope { + version: u8, + quiescence_sha256: String, + slots: u8, +} + +#[cfg(any(test, target_os = "macos"))] +#[derive(Clone, Debug, Deserialize, Serialize, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +struct Selection { + scope: Scope, + sockets: Vec, +} + +#[cfg(target_os = "macos")] +fn journal(candidate: &Candidate) -> PathBuf { + candidate + .state_root + .join("run/smolvm/quiescent-dependency-socket-recovery") +} + +#[cfg(any(test, target_os = "macos"))] +fn observed(directory: &Path, slot: u8) -> Result, CandidateError> { + observe_socket(directory, slot).map_err(|_| refused()) +} + +#[cfg(any(test, target_os = "macos"))] +fn current(scope: Scope, directory: &Path) -> Result { + DependencySocketIntent::new(scope.slots)?; + if scope.version != 1 || !hex(&scope.quiescence_sha256) { + return Err(refused()); + } + let mut sockets = Vec::new(); + for slot in 0..scope.slots { + if let Some(socket) = observed(directory, slot)? { + sockets.push(socket); + } + } + Ok(Selection { scope, sockets }) +} + +#[cfg(any(test, target_os = "macos"))] +fn matching( + selection: &Selection, + scope: &Scope, + directory: &Path, +) -> Result { + if selection.scope != *scope + || selection.sockets.is_empty() + || selection.sockets.len() > usize::from(scope.slots) + || selection.sockets.iter().any(|s| s.slot >= scope.slots) + || selection + .sockets + .windows(2) + .any(|pair| pair[0].slot >= pair[1].slot) + { + return Err(refused()); + } + let mut remaining = 0; + for slot in 0..scope.slots { + let saved = selection.sockets.iter().find(|item| item.slot == slot); + match (saved, observed(directory, slot)?) { + (Some(saved), Some(actual)) if *saved == actual => remaining += 1, + (Some(_), None) | (None, None) => {} + _ => return Err(refused()), + } + } + Ok(remaining) +} + +#[cfg(any(test, target_os = "macos"))] +fn receipts(root: &Path) -> Result, CandidateError> { + let entries = match fs::read_dir(root) { + Ok(entries) => entries, + Err(error) if error.kind() == std::io::ErrorKind::NotFound => return Ok(Vec::new()), + Err(_) => return Err(refused()), + }; + state::check_private_directory(root)?; + let mut paths = Vec::new(); + for entry in entries { + let entry = entry.map_err(|_| refused())?; + let name = entry.file_name().into_string().map_err(|_| refused())?; + // An incomplete journal or unknown file is retained and never interpreted + // as permission to delete. Only committed, bounded selections can resume. + if !name.strip_suffix(".json").is_some_and(hex) || paths.len() >= MAX_RECEIPTS { + return Err(refused()); + } + paths.push(entry.path()); + } + paths.sort(); + Ok(paths) +} + +#[cfg(any(test, target_os = "macos"))] +fn read(root: &Path, hash: &str) -> Result { + let selected: Selection = + state::read_bounded(&root.join(format!("{hash}.json")), RECEIPT_LIMIT)?; + if digest(&selected)? != hash { + return Err(refused()); + } + Ok(selected) +} + +#[cfg(any(test, target_os = "macos"))] +fn inspect_scope(root: &Path, scope: &Scope, directory: &Path) -> Result { + for receipt in receipts(root)? { + let hash = receipt + .file_stem() + .and_then(|s| s.to_str()) + .ok_or_else(refused)?; + let selected = read(root, hash)?; + if selected.scope == *scope { + let remaining = matching(&selected, scope, directory)?; + if remaining > 0 { + return Ok( + json!({"recoverable":true,"sha256":hash,"selected":selected.sockets.len(),"remaining":remaining,"resumable":true,"mode":"live-quiescent-legacy"}), + ); + } + } + } + let selected = current(scope.clone(), directory)?; + if selected.sockets.is_empty() { + return Ok(json!({"recoverable":false,"selected":0,"mode":"live-quiescent-legacy"})); + } + Ok( + json!({"recoverable":true,"sha256":digest(&selected)?,"selected":selected.sockets.len(),"remaining":selected.sockets.len(),"resumable":false,"mode":"live-quiescent-legacy"}), + ) +} + +#[cfg(any(test, target_os = "macos"))] +fn recover_scope( + root: &Path, + scope: &Scope, + directory: &Path, + expected: &str, + verify: impl Fn() -> Result<(), CandidateError>, +) -> Result { + if !hex(expected) { + return Err(refused()); + } + verify()?; + let paths = receipts(root)?; + let filename = root.join(format!("{expected}.json")); + let saved = if paths.contains(&filename) { + read(root, expected)? + } else { + let selected = current(scope.clone(), directory)?; + if selected.sockets.is_empty() + || digest(&selected)? != expected + || paths.len() >= MAX_RECEIPTS + { + return Err(refused()); + } + verify()?; + matching(&selected, scope, directory)?; + state::private_directory(root)?; + state::write(&filename, &selected)?; + selected + }; + let remaining = matching(&saved, scope, directory)?; + for socket in &saved.sockets { + verify()?; + matching(&saved, scope, directory)?; + if observed(directory, socket.slot)?.as_ref() == Some(socket) { + fs::remove_file(path(directory, socket.slot)).map_err(state::io)?; + File::open(directory) + .and_then(|f| f.sync_all()) + .map_err(state::io)?; + } + } + verify()?; + if matching(&saved, scope, directory)? != 0 { + return Err(refused()); + } + Ok( + json!({"recovered":true,"sha256":expected,"selected":saved.sockets.len(),"removed":remaining,"already_absent":saved.sockets.len()-remaining,"data_retained":true,"mode":"live-quiescent-legacy"}), + ) +} + +/// Select unlistened declared socket inodes only after holding the pool publication +/// gate, all retained graph retirement locks and the provider lease. Owner bytes, +/// guest boot, data volumes, empty reservations and absent compute are pinned. +#[cfg(target_os = "macos")] +pub fn inspect(candidate: &Candidate) -> Result { + let guard = crate::provider::graph::quiescent_dependency_recovery::Guard::acquire(candidate)?; + let scope = Scope { + version: 1, + quiescence_sha256: guard.sha256().into(), + slots: guard.owner().dependency_sockets.ok_or_else(refused)?.slots, + }; + let result = inspect_scope(&journal(candidate), &scope, &guard.owner().short_home)?; + guard.verify(candidate)?; + Ok(result) +} + +/// Explicit inode-selected recovery. A committed, separate journal precedes every +/// unlink; each effect rechecks the held quiescence proof. A partial retry accepts +/// exact remaining inodes or absence, and never restarts the VM or edits its Owner. +#[cfg(target_os = "macos")] +pub fn recover(candidate: &Candidate, expected: &str) -> Result { + if !hex(expected) { + return Err(refused()); + } + let guard = crate::provider::graph::quiescent_dependency_recovery::Guard::acquire(candidate)?; + let scope = Scope { + version: 1, + quiescence_sha256: guard.sha256().into(), + slots: guard.owner().dependency_sockets.ok_or_else(refused)?.slots, + }; + recover_scope( + &journal(candidate), + &scope, + &guard.owner().short_home, + expected, + || guard.verify(candidate), + ) +} + +#[cfg(not(target_os = "macos"))] +pub fn inspect(_: &Candidate) -> Result { + Err(refused()) +} +#[cfg(not(target_os = "macos"))] +pub fn recover(_: &Candidate, _: &str) -> Result { + Err(refused()) +} + +#[cfg(test)] +#[path = "quiescent_recovery_tests.rs"] +mod tests; diff --git a/packages/runtime-core/src/provider/dependency_socket/quiescent_recovery_tests.rs b/packages/runtime-core/src/provider/dependency_socket/quiescent_recovery_tests.rs new file mode 100644 index 000000000..8edfbe567 --- /dev/null +++ b/packages/runtime-core/src/provider/dependency_socket/quiescent_recovery_tests.rs @@ -0,0 +1,213 @@ +use super::*; +use std::{ + cell::Cell, + os::unix::{fs::PermissionsExt, net::UnixListener}, + sync::atomic::{AtomicU64, Ordering}, +}; + +static NEXT: AtomicU64 = AtomicU64::new(0); +struct Fixture(PathBuf); +impl Fixture { + fn new() -> Self { + let base = if cfg!(target_os = "macos") { + PathBuf::from("/private/tmp") + } else { + std::env::temp_dir() + }; + let root = base.join(format!( + "hk-qd-{}-{}", + std::process::id(), + NEXT.fetch_add(1, Ordering::Relaxed) + )); + fs::create_dir(&root).unwrap(); + fs::set_permissions(&root, fs::Permissions::from_mode(0o700)).unwrap(); + Self(root) + } + fn journal(&self) -> PathBuf { + self.0.join("journal") + } + fn scope(&self) -> Scope { + Scope { + version: 1, + quiescence_sha256: "a".repeat(64), + slots: 2, + } + } + fn bind(&self, slot: u8) -> UnixListener { + let target = path(&self.0, slot); + let listener = UnixListener::bind(&target).unwrap(); + fs::set_permissions(target, fs::Permissions::from_mode(0o600)).unwrap(); + listener + } + fn stale(&self, slot: u8) { + drop(self.bind(slot)); + } + fn selection(&self) -> Selection { + current(self.scope(), &self.0).unwrap() + } + fn hash(&self) -> String { + digest(&self.selection()).unwrap() + } + fn recover(&self, hash: &str) -> Result { + recover_scope(&self.journal(), &self.scope(), &self.0, hash, || Ok(())) + } +} +impl Drop for Fixture { + fn drop(&mut self) { + fs::remove_dir_all(&self.0).unwrap(); + } +} + +#[test] +fn selected_recovery_journals_before_unlink_and_is_idempotent() { + let fixture = Fixture::new(); + fixture.stale(0); + fixture.stale(1); + fs::write(fixture.0.join("retained-data"), b"unchanged").unwrap(); + let hash = fixture.hash(); + assert_eq!(fixture.recover(&hash).unwrap()["removed"], 2); + let receipt = fixture.journal().join(format!("{hash}.json")); + let bytes = fs::read(&receipt).unwrap(); + let retry = fixture.recover(&hash).unwrap(); + assert_eq!(retry["removed"], 0); + assert_eq!(retry["already_absent"], 2); + assert_eq!(fs::read(receipt).unwrap(), bytes); + assert_eq!( + fs::read(fixture.0.join("retained-data")).unwrap(), + b"unchanged" + ); + assert_eq!( + inspect_scope(&fixture.journal(), &fixture.scope(), &fixture.0).unwrap()["recoverable"], + false + ); +} + +#[test] +fn proof_drift_between_unlinks_stops_and_resumes_only_exact_remaining_inodes() { + let fixture = Fixture::new(); + fixture.stale(0); + fixture.stale(1); + let hash = fixture.hash(); + let calls = Cell::new(0); + let result = recover_scope( + &fixture.journal(), + &fixture.scope(), + &fixture.0, + &hash, + || { + calls.set(calls.get() + 1); + if calls.get() == 4 { + Err(refused()) + } else { + Ok(()) + } + }, + ); + assert!(result.is_err()); + assert!(!path(&fixture.0, 0).exists()); + assert!(path(&fixture.0, 1).exists()); + let selected = inspect_scope(&fixture.journal(), &fixture.scope(), &fixture.0).unwrap(); + assert_eq!(selected["sha256"], hash); + assert_eq!(selected["remaining"], 1); + assert_eq!(selected["resumable"], true); + assert_eq!(fixture.recover(&hash).unwrap()["removed"], 1); +} + +#[test] +fn owner_graph_boot_or_volume_proof_drift_refuses_before_any_effect() { + let fixture = Fixture::new(); + fixture.stale(0); + let hash = fixture.hash(); + let selected = fixture.selection(); + let changed = Scope { + quiescence_sha256: "b".repeat(64), + ..fixture.scope() + }; + assert!(recover_scope(&fixture.journal(), &changed, &fixture.0, &hash, || Ok(())).is_err()); + assert!( + recover_scope( + &fixture.journal(), + &fixture.scope(), + &fixture.0, + &hash, + || Err(refused()) + ) + .is_err() + ); + assert_eq!(fixture.selection(), selected); + assert!(!fixture.journal().exists()); +} + +#[test] +fn live_listener_and_foreign_replacement_are_never_removed() { + let fixture = Fixture::new(); + let listener = fixture.bind(0); + assert!(current(fixture.scope(), &fixture.0).is_err()); + drop(listener); + let hash = fixture.hash(); + fs::rename(path(&fixture.0, 0), fixture.0.join("original.sock")).unwrap(); + fixture.stale(0); + assert!(fixture.recover(&hash).is_err()); + assert!(path(&fixture.0, 0).exists()); + fs::remove_file(path(&fixture.0, 0)).unwrap(); + fs::write(path(&fixture.0, 0), b"foreign").unwrap(); + assert!(fixture.recover(&hash).is_err()); + assert_eq!(fs::read(path(&fixture.0, 0)).unwrap(), b"foreign"); +} + +#[test] +fn an_unselected_socket_or_pending_journal_blocks_retry() { + let fixture = Fixture::new(); + fixture.stale(0); + let hash = fixture.hash(); + state::private_directory(&fixture.journal()).unwrap(); + state::write( + &fixture.journal().join(format!("{hash}.json")), + &fixture.selection(), + ) + .unwrap(); + fixture.stale(1); + assert!(fixture.recover(&hash).is_err()); + assert!(path(&fixture.0, 0).exists()); + assert!(path(&fixture.0, 1).exists()); + fs::remove_file(path(&fixture.0, 1)).unwrap(); + let pending = fixture + .journal() + .join(format!("{}.pending", "c".repeat(64))); + fs::write(&pending, b"uncommitted").unwrap(); + assert!(fixture.recover(&hash).is_err()); + assert_eq!(fs::read(pending).unwrap(), b"uncommitted"); + assert!(path(&fixture.0, 0).exists()); +} + +#[test] +fn malformed_selection_and_stopped_pool_schema_cannot_authorize_live_recovery() { + let fixture = Fixture::new(); + fixture.stale(0); + let selected = fixture.selection(); + for sockets in [ + vec![selected.sockets[0].clone(), selected.sockets[0].clone()], + vec![], + ] { + assert!( + matching( + &Selection { + sockets, + ..selected.clone() + }, + &fixture.scope(), + &fixture.0 + ) + .is_err() + ); + } + assert!( + serde_json::from_value::( + json!({"version":1,"owner":"a","boot":"b","sockets":[]}) + ) + .is_err() + ); + assert!(fixture.recover("not-a-sha256").is_err()); + assert!(fixture.recover(&"0".repeat(64)).is_err()); + assert!(path(&fixture.0, 0).exists()); +} diff --git a/packages/runtime-core/src/provider/dependency_socket/recovery.rs b/packages/runtime-core/src/provider/dependency_socket/recovery.rs index 9aa655be4..3b0ff3df8 100644 --- a/packages/runtime-core/src/provider/dependency_socket/recovery.rs +++ b/packages/runtime-core/src/provider/dependency_socket/recovery.rs @@ -26,12 +26,12 @@ fn refused() -> CandidateError { ) } -fn digest(value: &T) -> Result { +pub(super) fn digest(value: &T) -> Result { let bytes = serde_json::to_vec(value).map_err(|_| refused())?; Ok(format!("{:x}", Sha256::digest(bytes))) } -fn hex(value: &str) -> bool { +pub(super) fn hex(value: &str) -> bool { value.len() == 64 && value .bytes() @@ -40,8 +40,8 @@ fn hex(value: &str) -> bool { #[derive(Clone, Debug, Deserialize, Serialize, PartialEq, Eq)] #[serde(deny_unknown_fields)] -struct Socket { - slot: u8, +pub(super) struct Socket { + pub(super) slot: u8, device: u64, inode: u64, } @@ -112,11 +112,11 @@ fn stopped(candidate: &Candidate) -> Result<(Selection, PathBuf), CandidateError )) } -fn path(directory: &Path, slot: u8) -> PathBuf { +pub(super) fn path(directory: &Path, slot: u8) -> PathBuf { directory.join(format!("dependency-{slot:02}.sock")) } -fn observed(directory: &Path, slot: u8) -> Result, CandidateError> { +pub(super) fn observed(directory: &Path, slot: u8) -> Result, CandidateError> { let target = path(directory, slot); let metadata = match fs::symlink_metadata(&target) { Ok(metadata) => metadata, diff --git a/packages/runtime-core/src/provider/environment_recovery.rs b/packages/runtime-core/src/provider/environment_recovery.rs index a2b9b8fa8..bac614abd 100644 --- a/packages/runtime-core/src/provider/environment_recovery.rs +++ b/packages/runtime-core/src/provider/environment_recovery.rs @@ -88,9 +88,10 @@ fn validate(intent: &Intent, slot: &str) -> Result<(), CandidateError> { Ok(()) } const MAX_INTENT_ENTRIES: usize = 4096; -// A graph can retain several generations of immutable service lease evidence. -// Keep this below the global bound while permitting repeated normal restarts. -const MAX_GRAPH_INTENT_ENTRIES: usize = 256; +// Cleanup must cover every allocation admitted by the global history budget. +// A smaller per-graph bound can strand a valid graph after repeated restores. +// Active and archived entries share this bound; none are discarded to fit it. +const MAX_GRAPH_INTENT_ENTRIES: usize = MAX_INTENT_ENTRIES; /// Retired and uncertain entries still reserve history capacity. This never removes evidence. pub(super) fn preflight_records( @@ -317,6 +318,12 @@ impl GraphInventory { pub(super) fn is_empty(&self) -> bool { self.intents.is_empty() } + pub(super) fn slots(&self) -> Vec { + self.intents + .iter() + .map(|intent| intent.slot.clone()) + .collect() + } } #[cfg(target_os = "macos")] @@ -399,6 +406,28 @@ pub(super) fn verify_graph_retired( Ok(()) } +/// Verify one selected allocation after an interrupted cleanup effect. Other +/// allocations in the same graph may still be present at this journal step. +#[cfg(target_os = "macos")] +pub(super) fn verify_graph_slot_retired( + guest: &OwnedGuest<'_>, + inventory: &GraphInventory, + slot: &str, +) -> Result<(), CandidateError> { + if inventory.incarnation != guest.incarnation() { + return Err(error()); + } + let intent = inventory + .intents + .iter() + .find(|intent| intent.slot == slot) + .ok_or_else(error)?; + let binding = intent.graph.as_ref().ok_or_else(error)?; + super::engine::require_container_absent(guest, &binding.container)?; + retention::absent(guest, slot)?; + guest.verify() +} + /// Move retired, value-free intents into a removed graph's retained evidence. The graph ID /// stays reserved by its active/archive directory, then by the verified-prune consumed record. /// A rename interrupted before final graph archival is recovered by validating both inventories. @@ -867,7 +896,10 @@ mod tests { run: run.clone(), container: container.clone(), }); - for index in 0..84 { + // Nineteen normal starts of a fourteen-service graph exceeded the old + // cleanup-only ceiling, despite remaining within allocation admission. + let retained_count = 19 * 14; + for index in 0..retained_count { intent.slot = format!("hack-env-lease-{}-{index:032x}", intent.boot); state::write( &root(&fixture.0).join(format!("{}.json", intent.slot)), @@ -875,6 +907,7 @@ mod tests { ) .unwrap(); } + assert!(preflight_records(&fixture.0, 14).is_ok()); let inventory = graph_inventory( &fixture.0, &intent.incarnation, @@ -883,7 +916,8 @@ mod tests { &fixture.0.state_root.join("graph-environment-archive"), ) .unwrap(); - assert_eq!(inventory.intents.len(), 84); + assert_eq!(inventory.intents.len(), retained_count); + assert_eq!(recorded_slots(&fixture.0).unwrap().len(), retained_count); } #[test] diff --git a/packages/runtime-core/src/provider/graph/absent_publication_cleanup.rs b/packages/runtime-core/src/provider/graph/absent_publication_cleanup.rs new file mode 100644 index 000000000..1a6676ef9 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/absent_publication_cleanup.rs @@ -0,0 +1,1521 @@ +//! Explicit data-retaining recovery when both ephemeral host publications were +//! lost across a physical host reboot. Missing paths never create ordinary +//! ownership authority: the selected, durable witness is the only admission. +use super::{ + Candidate, CandidateError, Engine, Kind, Receipt, dead_owner_cleanup, dependency_slots, + directory, foreground, host_pin_recovery, host_relay, initializer_cache, inspect_resource, + load, startup, state, +}; +use crate::provider::{ + ProjectShareIntent, artifact, host_pin::DeviceRebind, identity, lifecycle, state::Owner, +}; +use serde::{Deserialize, Serialize}; +use serde_json::{Value, json}; +use sha2::{Digest, Sha256}; +use std::{ + collections::BTreeMap, + fs, + os::unix::fs::MetadataExt, + path::{Path, PathBuf}, +}; + +const INTENT: &str = "absent-publication-cleanup.json"; +const RETIREMENT: &str = "absent-publication-retirement.json"; +const LIMIT: u64 = 2 * 1024 * 1024; +mod rebind_history; +pub(super) use rebind_history::archive_retired_rebind_under; + +fn refused() -> CandidateError { + CandidateError::new( + "graph_absent_publication_recovery", + "Selected post-reboot publication absence or retained graph identity changed; evidence and data were preserved.", + ) +} +fn digest(bytes: &[u8]) -> String { + format!("{:x}", Sha256::digest(bytes)) +} +fn absent(path: &Path) -> Result { + match fs::symlink_metadata(path) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => Ok(true), + Ok(_) => Ok(false), + Err(_) => Err(refused()), + } +} +fn require_absent(path: &Path) -> Result<(), CandidateError> { + if !absent(path)? { + return Err(refused()); + } + Ok(()) +} +fn no_pending(root: &Path) -> Result<(), CandidateError> { + for name in [ + "state.pending", + "dead-owner-cleanup.pending", + "relay-cleanup-bridges.pending", + "one-off.pending", + "one-off-normalization.pending", + "absent-publication-cleanup.pending", + "absent-publication-retirement.pending", + ] { + require_absent(&root.join(name))?; + } + Ok(()) +} +fn retain_interrupted_publications(root: &Path) -> Result<(), CandidateError> { + super::journal::retain_file( + root, + "absent-publication-cleanup.pending", + "absent-publication-cleanup-recovery", + 4 * 1024 * 1024, + )?; + super::journal::retain_file( + root, + "absent-publication-retirement.pending", + "absent-publication-retirement-recovery", + 65536, + )?; + Ok(()) +} +fn private_input(path: &Path, limit: u64) -> Result, CandidateError> { + if !path.is_absolute() || fs::canonicalize(path).map_err(|_| refused())? != path { + return Err(refused()); + } + host_pin_recovery::read_raw(path, limit).map_err(|_| refused()) +} +fn device_only_share( + old: Option<&ProjectShareIntent>, + current: Option<&ProjectShareIntent>, + old_device: u64, + new_device: u64, +) -> bool { + match (old, current) { + (None, None) => true, + (Some(previous), Some(current)) => { + let mut expected = previous.clone(); + expected.device = new_device; + previous.device == old_device && *current == expected + } + _ => false, + } +} + +/// This is the exact serialization shape of the pre-migration inspection. It +/// binds the raw original Owner, one common host-device change and the boot +/// used to approve it; it is not a physical-volume UUID proof. +#[derive(Clone, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct FilesystemInspection { + schema: String, + selection_sha256: String, + owner_sha256: String, + machine: String, + host_boot_micros: u64, + provider_start_micros: u64, + old_device: u64, + new_device: u64, + pool_inode: u64, + storage: identity::DiskIdentity, + overlay: identity::DiskIdentity, + project_share: Option, + qualification: String, +} +impl FilesystemInspection { + fn verify(&self, owner_bytes: &[u8], old: &Owner) -> Result<(), CandidateError> { + let mut unsigned = self.clone(); + unsigned.selection_sha256.clear(); + if self.schema != "hack.host-filesystem-recovery/v1" + || self.qualification != "explicit-legacy-migration-original-volume-continuity-unproven" + || self.owner_sha256 != digest(owner_bytes) + || self.selection_sha256 + != digest(&serde_json::to_vec(&unsigned).map_err(|_| refused())?) + || self.machine != old.machine + || self.provider_start_micros != old.process.as_ref().ok_or_else(refused)?.start_micros + || self.old_device == self.new_device + || self.storage.device != self.new_device + || self.overlay.device != self.new_device + || old.storage.as_ref().is_none_or(|disk| { + disk.device != self.old_device + || disk.inode != self.storage.inode + || disk.bytes != self.storage.bytes + || disk.uuid != self.storage.uuid + }) + || old.overlay.as_ref().is_none_or(|disk| { + disk.device != self.old_device + || disk.inode != self.overlay.inode + || disk.bytes != self.overlay.bytes + || disk.uuid != self.overlay.uuid + }) + || !device_only_share( + old.project_share.as_ref(), + self.project_share.as_ref(), + self.old_device, + self.new_device, + ) + { + return Err(refused()); + } + Ok(()) + } +} + +#[cfg(all(test, feature = "environment-launcher"))] +pub(in crate::provider::graph) fn fixture_inspection( + owner_bytes: &[u8], + current: &Owner, + old_device: u64, + host_boot_micros: u64, + pool_inode: u64, +) -> Vec { + let old: Owner = serde_json::from_slice(owner_bytes).unwrap(); + let mut selection = FilesystemInspection { + schema: "hack.host-filesystem-recovery/v1".into(), + selection_sha256: String::new(), + owner_sha256: digest(owner_bytes), + machine: old.machine.clone(), + host_boot_micros, + provider_start_micros: old.process.as_ref().unwrap().start_micros, + old_device, + new_device: current.storage.as_ref().unwrap().device, + pool_inode, + storage: current.storage.clone().unwrap(), + overlay: current.overlay.clone().unwrap(), + project_share: current.project_share.clone(), + qualification: "explicit-legacy-migration-original-volume-continuity-unproven".into(), + }; + selection.selection_sha256 = digest(&serde_json::to_vec(&selection).unwrap()); + serde_json::to_vec_pretty(&selection).unwrap() +} + +#[derive(Clone, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +pub(super) struct Selection { + version: u8, + candidate: PathBuf, + run: String, + owner: String, + namespace: String, + plan: String, + original_owner_path: PathBuf, + original_owner_sha256: String, + filesystem_inspection_path: PathBuf, + filesystem_inspection_sha256: String, + current_owner_sha256: String, + graph_sha256: String, + host_boot_micros: u64, + old_device: u64, + new_device: u64, + previous_guest_boot: String, + current_guest_boot: String, + foreground_root: PathBuf, + control_root: PathBuf, + source_shared: Option, + environment_inventory: Value, + scoped_bridge_projection: Value, + retained_volumes: BTreeMap, + dependency_reservation: Option, + qualification: String, +} +impl Selection { + pub(super) fn digest(&self) -> Result { + Ok(digest(&serde_json::to_vec(self).map_err(|_| refused())?)) + } + pub(super) fn matches_graph(&self, receipt: &Receipt) -> bool { + self.run == receipt.run + && self.owner == receipt.owner + && self.namespace == receipt.namespace + && self.plan == receipt.plan_id + && self.source_shared == receipt.source.as_ref().and_then(|s| s.shared.clone()) + } + pub(super) fn previous_boot(&self) -> &str { + &self.previous_guest_boot + } + pub(super) fn control_root(&self) -> &Path { + &self.control_root + } +} + +#[derive(Clone, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct Intent { + version: u8, + selection_sha256: String, + selection: Selection, + original: Receipt, + environment: Value, + bridges: super::bridges::cleanup::Selection, + #[serde(default, skip_serializing_if = "Option::is_none")] + prior_bridges: Option, + #[serde(default, skip_serializing_if = "Option::is_none")] + complete_sha256: Option, +} +#[derive(Clone, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct Retirement { + version: u8, + selection_sha256: String, + complete_sha256: String, + owner: String, + run: String, +} + +fn selected_volumes( + engine: &Engine<'_>, + receipt: &Receipt, +) -> Result, CandidateError> { + let mut volumes = BTreeMap::new(); + for (key, resource) in receipt + .resources + .iter() + .filter(|(_, resource)| resource.kind == Kind::Volume) + { + let actual = inspect_resource(engine, receipt, resource)?.ok_or_else(refused)?; + let created = actual["CreatedAt"] + .as_str() + .filter(|value| !value.is_empty()); + if created.is_none() { + return Err(refused()); + } + volumes.insert( + key.clone(), + json!({ + "expected_name": resource.name, + "observed_name": actual["Name"], + "observed_created_at": actual["CreatedAt"], + "observed_mountpoint": actual["Mountpoint"], + "observed_driver": actual["Driver"], + "observed_labels_sha256": digest( + &serde_json::to_vec(&actual["Labels"]).map_err(|_| refused())? + ), + "labels": super::expected_labels(receipt, resource), + }), + ); + } + Ok(volumes) +} + +fn current_owner( + candidate: &Candidate, + old: &Owner, + inspection: &FilesystemInspection, + engine: &Engine<'_>, +) -> Result<(Owner, String), CandidateError> { + lifecycle::host_filesystem::no_auxiliary_update(candidate)?; + if lifecycle::host_filesystem::host_boot_micros()? != inspection.host_boot_micros { + return Err(refused()); + } + let old_process = old.process.as_ref().ok_or_else(refused)?; + // SAFETY: geteuid has no arguments or side effects. + identity::verify( + old_process, + old_process, + &artifact::root(candidate).join("smolvm-bin"), + unsafe { libc::geteuid() }, + ) + .map_err(|_| refused())?; + DeviceRebind { + old: inspection.old_device, + current: inspection.new_device, + } + .definitely_dead_before_boot(old_process, inspection.host_boot_micros) + .map_err(|_| refused())?; + let owner = Owner::load(candidate)?; + lifecycle::verify_disks(candidate, &owner)?; + host_pin_recovery::verify_guest_identity(engine)?; + let pool = + fs::symlink_metadata(candidate.state_root.join("run/smolvm")).map_err(|_| refused())?; + if old.phase != "running" + || owner.phase != "running" + || old.token != owner.token + || owner.token != engine.guest().incarnation() + || old.guest_boot_id.as_deref() != owner.previous_guest_boot_id.as_deref() + || owner.guest_boot_id.as_deref() != Some(engine.guest().boot_id()) + || old.guest_boot_id.as_deref() == owner.guest_boot_id.as_deref() + || owner.storage.as_ref() != Some(&inspection.storage) + || owner.overlay.as_ref() != Some(&inspection.overlay) + || owner.project_share != inspection.project_share + || pool.ino() != inspection.pool_inode + || pool.dev() != inspection.new_device + { + return Err(refused()); + } + let mut expected = old.clone(); + expected.storage = owner.storage.clone(); + expected.overlay = owner.overlay.clone(); + expected.project_share = owner.project_share.clone(); + expected.process = owner.process.clone(); + expected.guest_boot_id = owner.guest_boot_id.clone(); + expected.previous_guest_boot_id = owner.previous_guest_boot_id.clone(); + expected.daemon_pid = owner.daemon_pid; + expected.daemon_start = owner.daemon_start; + if expected != owner { + return Err(refused()); + } + let owner_bytes = host_pin_recovery::read_raw( + &candidate.state_root.join("run/smolvm/owner.json"), + 1024 * 1024, + )?; + let raw_owner: Owner = serde_json::from_slice(&owner_bytes).map_err(|_| refused())?; + if raw_owner != owner { + return Err(refused()); + } + let owner_sha256 = digest(&owner_bytes); + Ok((owner, owner_sha256)) +} + +fn publication_roots(selection: &Selection, allow_lock_only: bool) -> Result<(), CandidateError> { + require_absent(&selection.control_root)?; + match fs::symlink_metadata(&selection.foreground_root) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound && !allow_lock_only => Ok(()), + Ok(_) if allow_lock_only => { + state::check_private_directory(&selection.foreground_root)?; + let mut entries = fs::read_dir(&selection.foreground_root).map_err(|_| refused())?; + let entry = entries.next().ok_or_else(refused)?.map_err(|_| refused())?; + if entry.file_name() != "operation.lock" || entries.next().is_some() { + return Err(refused()); + } + let lock = fs::symlink_metadata(entry.path()).map_err(|_| refused())?; + // SAFETY: geteuid has no arguments or side effects. + if !lock.is_file() + || lock.nlink() != 1 + || lock.mode() & 0o7777 != 0o600 + || lock.uid() != unsafe { libc::geteuid() } + { + return Err(refused()); + } + Ok(()) + } + _ => Err(refused()), + } +} + +fn select_content( + candidate: &Candidate, + run: &str, + original_owner_path: &Path, + filesystem_inspection_path: &Path, + engine: &Engine<'_>, + allow_lock_only: bool, +) -> Result<(Selection, Receipt), CandidateError> { + let (receipt, root) = load(candidate, engine, run)?; + no_pending(&root)?; + if receipt.phase != "ready-observed" + || receipt.relay_cleanup.is_some() + || receipt.relay_startup.is_none() + || !super::hex(&receipt.owner, 32) + { + return Err(refused()); + } + initializer_cache::require_resolved(&receipt)?; + startup::require_dependency_rebind_complete(&root, &receipt)?; + dead_owner_cleanup::require_historical_recovery(&root, &receipt)?; + let old_bytes = private_input(original_owner_path, 1024 * 1024)?; + let old: Owner = serde_json::from_slice(&old_bytes).map_err(|_| refused())?; + let inspection_bytes = private_input(filesystem_inspection_path, 64 * 1024)?; + let inspection: FilesystemInspection = + serde_json::from_slice(&inspection_bytes).map_err(|_| refused())?; + inspection.verify(&old_bytes, &old)?; + let (owner, current_owner_sha256) = current_owner(candidate, &old, &inspection, engine)?; + if old.checkout != candidate.checkout + || old.machine != inspection.machine + || owner.machine != inspection.machine + || old + .project_share + .as_ref() + .is_some_and(|share| share.device != inspection.old_device) + || receipt + .source + .as_ref() + .and_then(|source| source.shared.as_ref()) + != old.project_share.as_ref() + { + return Err(refused()); + } + let startup = receipt.relay_startup.as_ref().ok_or_else(refused)?; + let foreground_root = foreground::transport::root(candidate, run)?; + let graph_bytes = host_pin_recovery::read_raw(&root.join("state.json"), LIMIT)?; + if graph_bytes != serde_json::to_vec_pretty(&receipt).map_err(|_| refused())? { + return Err(refused()); + } + let graph_sha256 = digest(&graph_bytes); + let rebind = DeviceRebind { + old: inspection.old_device, + current: inspection.new_device, + }; + let dependency_reservation = + dependency_slots::inspect_legacy(candidate, run, rebind, inspection.host_boot_micros)?; + let previous_guest_boot = old.guest_boot_id.clone().ok_or_else(refused)?; + let selection = Selection { + version: 1, + candidate: candidate.checkout.clone(), + run: run.into(), + owner: receipt.owner.clone(), + namespace: receipt.namespace.clone(), + plan: receipt.plan_id.clone(), + original_owner_path: original_owner_path.to_path_buf(), + original_owner_sha256: digest(&old_bytes), + filesystem_inspection_path: filesystem_inspection_path.to_path_buf(), + filesystem_inspection_sha256: digest(&inspection_bytes), + current_owner_sha256, + graph_sha256, + host_boot_micros: inspection.host_boot_micros, + old_device: inspection.old_device, + new_device: inspection.new_device, + previous_guest_boot: previous_guest_boot.clone(), + current_guest_boot: owner.guest_boot_id.ok_or_else(refused)?, + foreground_root, + control_root: startup.control_root.clone(), + source_shared: receipt + .source + .as_ref() + .and_then(|source| source.shared.clone()), + environment_inventory: serde_json::to_value(super::environment::cleanup_inventory( + candidate, engine, &receipt, &root, + )?) + .map_err(|_| refused())?, + scoped_bridge_projection: super::bridges::cleanup::scoped_absence_projection( + candidate, + engine, + &receipt, + &previous_guest_boot, + )?, + retained_volumes: selected_volumes(engine, &receipt)?, + dependency_reservation, + qualification: "explicit-unpinned-post-host-reboot-original-volume-continuity-unproven" + .into(), + }; + publication_roots(&selection, allow_lock_only)?; + host_pin_recovery::verify_volume_projections(engine, &receipt, &selection.retained_volumes)?; + for resource in receipt.resources.values() { + if resource.kind != Kind::Volume && inspect_resource(engine, &receipt, resource)?.is_none() + { + return Err(refused()); + } + } + Ok((selection, receipt)) +} + +pub fn inspect( + candidate: &Candidate, + run: &str, + original_owner_path: &Path, + filesystem_inspection_path: &Path, +) -> Result { + let existing = Reservation::inspect_existing(candidate, run)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let (selected, _) = select_content( + candidate, + run, + original_owner_path, + filesystem_inspection_path, + &engine, + existing.is_some(), + )?; + if let Some(reservation) = &existing { + reservation.verify()?; + } + Ok(json!({ + "run": run, + "selection_sha256": selected.digest()?, + "host_boot_micros": selected.host_boot_micros, + "qualification": selected.qualification, + "data_retained": true, + })) +} + +struct Reservation { + root: PathBuf, + identity: (u64, u64), + lock: state::Lock, +} +impl Reservation { + fn inspect_existing(candidate: &Candidate, run: &str) -> Result, CandidateError> { + let root = foreground::transport::root(candidate, run)?; + match fs::symlink_metadata(&root) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => Ok(None), + Ok(_) => { + state::check_private_directory(&root)?; + let lock = state::Lock::acquire_existing(&root)?; + let metadata = fs::symlink_metadata(&root).map_err(|_| refused())?; + let value = Self { + root, + identity: (metadata.dev(), metadata.ino()), + lock, + }; + value.verify()?; + Ok(Some(value)) + } + Err(_) => Err(refused()), + } + } + fn acquire(candidate: &Candidate, run: &str) -> Result { + let root = foreground::transport::root(candidate, run)?; + let lock = match fs::symlink_metadata(&root) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => { + state::Lock::acquire(&root)? + } + Ok(_) => { + state::check_private_directory(&root)?; + state::Lock::acquire_existing(&root)? + } + Err(_) => return Err(refused()), + }; + let metadata = fs::symlink_metadata(&root).map_err(|_| refused())?; + let value = Self { + root, + identity: (metadata.dev(), metadata.ino()), + lock, + }; + value.verify()?; + Ok(value) + } + fn verify(&self) -> Result<(), CandidateError> { + let metadata = fs::symlink_metadata(&self.root).map_err(|_| refused())?; + if !metadata.is_dir() + || (metadata.dev(), metadata.ino()) != self.identity + || state::check_private_directory(&self.root).is_err() + { + return Err(refused()); + } + host_pin_recovery::exact_lock_path(&self.root, &self.lock).map_err(|_| refused())?; + let lock = fs::symlink_metadata(self.root.join("operation.lock")).map_err(|_| refused())?; + // SAFETY: geteuid has no arguments or side effects. + if lock.mode() & 0o7777 != 0o600 + || lock.uid() != unsafe { libc::geteuid() } + || lock.nlink() != 1 + { + return Err(refused()); + } + // This verifier is intentionally limited to the reserved root; it + // does not infer control-root or graph ownership from the lock. + let mut entries = fs::read_dir(&self.root).map_err(|_| refused())?; + let entry = entries.next().ok_or_else(refused)?.map_err(|_| refused())?; + if entry.file_name() != "operation.lock" || entries.next().is_some() { + return Err(refused()); + } + Ok(()) + } +} + +fn verify_static_with( + candidate: &Candidate, + run: &str, + engine: &Engine<'_>, + selected: &Selection, + original: &Receipt, + foreground_root: &Path, + verify_lock: impl Fn() -> Result<(), CandidateError>, +) -> Result { + if selected.version != 1 + || selected.candidate != candidate.checkout + || selected.run != run + || selected.foreground_root != foreground_root + || selected.qualification + != "explicit-unpinned-post-host-reboot-original-volume-continuity-unproven" + || !selected.matches_graph(original) + || selected.digest()?.len() != 64 + { + return Err(refused()); + } + verify_lock()?; + publication_roots(selected, true)?; + lifecycle::host_filesystem::no_auxiliary_update(candidate)?; + if lifecycle::host_filesystem::host_boot_micros()? != selected.host_boot_micros { + return Err(refused()); + } + let old_bytes = private_input(&selected.original_owner_path, 1024 * 1024)?; + let inspection_bytes = private_input(&selected.filesystem_inspection_path, 64 * 1024)?; + if digest(&old_bytes) != selected.original_owner_sha256 + || digest(&inspection_bytes) != selected.filesystem_inspection_sha256 + { + return Err(refused()); + } + let old: Owner = serde_json::from_slice(&old_bytes).map_err(|_| refused())?; + let inspection: FilesystemInspection = + serde_json::from_slice(&inspection_bytes).map_err(|_| refused())?; + inspection.verify(&old_bytes, &old)?; + let (owner, owner_sha256) = current_owner(candidate, &old, &inspection, engine)?; + if owner_sha256 != selected.current_owner_sha256 + || owner.token != selected.owner + || owner.guest_boot_id.as_deref() != Some(&selected.current_guest_boot) + || owner.previous_guest_boot_id.as_deref() != Some(&selected.previous_guest_boot) + || inspection.old_device != selected.old_device + || inspection.new_device != selected.new_device + { + return Err(refused()); + } + let (receipt, root) = load(candidate, engine, run)?; + no_pending(&root)?; + let receipt_bytes = host_pin_recovery::read_raw(&root.join("state.json"), LIMIT)?; + if receipt_bytes != serde_json::to_vec_pretty(&receipt).map_err(|_| refused())? { + return Err(refused()); + } + if dead_owner_cleanup::immutable(&receipt)? != dead_owner_cleanup::immutable(original)? + || !selected.matches_graph(&receipt) + || !["ready-observed", "cleanup-intent", "stopped-data-retained"] + .contains(&receipt.phase.as_str()) + { + return Err(refused()); + } + if receipt.phase == "ready-observed" && digest(&receipt_bytes) != selected.graph_sha256 { + return Err(refused()); + } + if receipt.phase == "ready-observed" + && (serde_json::to_value(super::environment::cleanup_inventory( + candidate, engine, &receipt, &root, + )?) + .map_err(|_| refused())? + != selected.environment_inventory + || super::bridges::cleanup::scoped_absence_projection( + candidate, + engine, + &receipt, + &selected.previous_guest_boot, + )? != selected.scoped_bridge_projection) + { + return Err(refused()); + } + host_pin_recovery::verify_volume_projections(engine, &receipt, &selected.retained_volumes)?; + let rebind = DeviceRebind { + old: selected.old_device, + current: selected.new_device, + }; + if let Some(legacy) = &selected.dependency_reservation { + dependency_slots::verify_legacy_remaining( + candidate, + run, + rebind, + legacy, + receipt.phase != "ready-observed", + )?; + } else if dependency_slots::inspect_legacy(candidate, run, rebind, selected.host_boot_micros)? + .is_some() + { + return Err(refused()); + } + verify_lock()?; + Ok(receipt) +} + +fn verify_static( + candidate: &Candidate, + run: &str, + engine: &Engine<'_>, + selected: &Selection, + original: &Receipt, + reservation: &Reservation, +) -> Result { + verify_static_with( + candidate, + run, + engine, + selected, + original, + &reservation.root, + || reservation.verify(), + ) +} + +fn read_intent(root: &Path) -> Result, CandidateError> { + let path = root.join(INTENT); + if absent(&path)? { + return Ok(None); + } + let intent: Intent = state::read_bounded(&path, 4 * 1024 * 1024).map_err(|_| refused())?; + if intent.version != 1 + || intent.selection_sha256 != intent.selection.digest()? + || intent.selection.graph_sha256 + != digest(&serde_json::to_vec_pretty(&intent.original).map_err(|_| refused())?) + || !intent.selection.matches_graph(&intent.original) + { + return Err(refused()); + } + Ok(Some(intent)) +} +fn completed_receipt( + candidate: &Candidate, + engine: &Engine<'_>, + intent: &Intent, + current: &Receipt, + root: &Path, +) -> Result { + if current.phase != "stopped-data-retained" { + return Err(refused()); + } + let environment = super::environment::cleanup_inventory(candidate, engine, current, root)?; + if serde_json::to_value(&environment).map_err(|_| refused())? != intent.environment { + return Err(refused()); + } + super::bridges::cleanup::verify_recovery_file( + root, + &intent.bridges, + intent.prior_bridges.as_ref(), + )?; + host_relay::inspect_cleanup_absence( + candidate, + engine, + current, + &environment, + &intent.bridges, + &intent.selection, + )?; + Ok(digest(&host_pin_recovery::read_raw( + &root.join("state.json"), + LIMIT, + )?)) +} + +fn retire( + candidate: &Candidate, + engine: &Engine<'_>, + intent: &Intent, + current: &Receipt, + root: &Path, + reservation: &Reservation, +) -> Result<(), CandidateError> { + reservation.verify()?; + let complete = intent.complete_sha256.as_deref().ok_or_else(refused)?; + if completed_receipt(candidate, engine, intent, current, root)? != complete { + return Err(refused()); + } + let retirement = Retirement { + version: 1, + selection_sha256: intent.selection_sha256.clone(), + complete_sha256: complete.into(), + owner: intent.selection.owner.clone(), + run: intent.selection.run.clone(), + }; + let path = root.join(RETIREMENT); + if absent(&path)? { + verify_static( + candidate, + &intent.selection.run, + engine, + &intent.selection, + &intent.original, + reservation, + )?; + state::write(&path, &retirement)?; + } else { + let recorded: Retirement = state::read_bounded(&path, 65536).map_err(|_| refused())?; + if recorded.version != retirement.version + || recorded.selection_sha256 != retirement.selection_sha256 + || recorded.complete_sha256 != retirement.complete_sha256 + || recorded.owner != retirement.owner + || recorded.run != retirement.run + { + return Err(refused()); + } + } + reservation.verify() +} + +/// A pending absence intent blocks ordinary publication. Completed retirement +/// permits only the explicit restored-generation bind; it is not a Pin. +pub(super) fn publication_allowed( + candidate: &Candidate, + run: &str, + retired: bool, +) -> Result<(), CandidateError> { + let root = directory(candidate, run)?; + require_absent(&root.join("absent-publication-cleanup.pending"))?; + require_absent(&root.join("absent-publication-retirement.pending"))?; + if absent(&root.join(INTENT))? { + return Ok(()); + } + if !retired { + return Err(refused()); + } + let intent = read_intent(&root)?.ok_or_else(refused)?; + let complete = intent.complete_sha256.as_deref().ok_or_else(refused)?; + let retirement: Retirement = + state::read_bounded(&root.join(RETIREMENT), 65536).map_err(|_| refused())?; + let current: Receipt = state::read_bounded(&root.join("state.json"), LIMIT)?; + if retirement.version != 1 + || retirement.selection_sha256 != intent.selection_sha256 + || retirement.complete_sha256 != complete + || retirement.owner != current.owner + || retirement.run != run + || current.phase != "stopped-data-retained" + { + return Err(refused()); + } + let current_sha256 = digest(&host_pin_recovery::read_raw( + &root.join("state.json"), + LIMIT, + )?); + if current_sha256 != complete { + if current_sha256 != digest(&serde_json::to_vec_pretty(¤t).map_err(|_| refused())?) { + return Err(refused()); + } + require_historical_retention(&root, ¤t, &intent)?; + // Historical absence retirement is not the current cleanup authority. + // Dispatch current retention outside `retained` to avoid recursion and + // admit only an independently confirmed current recovery or relay ACK. + super::cleanup_enrollment::retention(&root, ¤t)?; + } + Ok(()) +} + +/// Prefer only an exact current completion over historical cleanup sidecars. +/// A later generation must still pass the ordinary enrollment checks. +pub(super) fn retained_current(root: &Path, receipt: &Receipt) -> Result { + let Some(intent) = read_intent(root)? else { + return Ok(false); + }; + no_pending(root)?; + require_absent(&root.join("live-owner-cleanup.pending"))?; + super::initializer_cache::require_resolved(receipt)?; + let requested = digest(&serde_json::to_vec_pretty(receipt).map_err(|_| refused())?); + if intent.complete_sha256.as_deref() != Some(requested.as_str()) { + return Ok(false); + } + if !intent.selection.matches_graph(receipt) || receipt.relay_cleanup.is_some() { + return Err(refused()); + } + let bridges: super::bridges::cleanup::Selection = + state::read_bounded(&root.join("relay-cleanup-bridges.json"), 65536)?; + if bridges != intent.bridges { + return Err(refused()); + } + dead_owner_cleanup::require_historical_recovery(root, receipt)?; + super::live_owner_cleanup::require_historical_recovery(root, receipt)?; + retained(root, receipt) +} + +/// Retention is bound to the completed stopped receipt, not to missing paths. +pub(super) fn retained(root: &Path, receipt: &Receipt) -> Result { + let Some(intent) = read_intent(root)? else { + return Ok(false); + }; + let Some(complete) = intent.complete_sha256.as_deref() else { + return Err(refused()); + }; + let retirement: Retirement = + state::read_bounded(&root.join(RETIREMENT), 65536).map_err(|_| refused())?; + let current: Receipt = state::read_bounded(&root.join("state.json"), LIMIT)?; + let current_sha256 = digest(&host_pin_recovery::read_raw( + &root.join("state.json"), + LIMIT, + )?); + let requested_sha256 = digest(&serde_json::to_vec_pretty(receipt).map_err(|_| refused())?); + if current_sha256 != requested_sha256 + || receipt.phase != "stopped-data-retained" + || retirement.version != 1 + || retirement.selection_sha256 != intent.selection_sha256 + || retirement.complete_sha256 != complete + || retirement.run != receipt.run + || retirement.owner != receipt.owner + || serde_json::to_vec_pretty(¤t).map_err(|_| refused())? + != serde_json::to_vec_pretty(receipt).map_err(|_| refused())? + { + return Err(refused()); + } + if requested_sha256 != complete { + require_historical_retention(root, receipt, &intent)?; + super::cleanup_enrollment::retention_receipt(receipt, false)?; + return Ok(false); + } + Ok(true) +} + +fn require_historical_retention( + root: &Path, + receipt: &Receipt, + intent: &Intent, +) -> Result<(), CandidateError> { + if intent.original.run != receipt.run + || intent.original.owner != receipt.owner + || intent.original.namespace != receipt.namespace + || !super::restore_history::confirms_prior_generation(root, receipt)? + { + return Err(refused()); + } + Ok(()) +} + +/// Exact completed, retired first-generation proof for a later independently +/// selected source-continuity transition. The proof grants no authority for a +/// subsequent graph generation and does not change the original graph source. +#[derive(Clone, Serialize)] +#[allow(dead_code)] // Consumed by the independently selected source-continuity unit. +pub(super) struct CompletedSourceProof { + pub original_ready_sha256: String, + pub completed_stopped_sha256: String, + pub intent_raw_sha256: String, + pub retirement_raw_sha256: String, + pub original_owner_sha256: String, + pub current_owner_sha256: String, + pub host_boot_micros: u64, + pub previous_guest_boot: String, + pub current_guest_boot: String, + pub old_share: Option, + pub current_share: Option, + pub retained_volumes: BTreeMap, +} + +/// The caller holds the foreground retirement lock before its Engine lease and +/// must recheck this proof at the final publication boundary. A historical +/// sidecar with a different stopped receipt cannot lend authority. +#[allow(dead_code)] // Exposed for the subsequent source-continuity unit. +pub(super) fn verify_completed_under( + candidate: &Candidate, + engine: &Engine<'_>, + run: &str, + receipt: &Receipt, + retired: &foreground::transport::Retired, +) -> Result, CandidateError> { + let root = directory(candidate, run)?; + if absent(&root.join(INTENT))? { + return Ok(None); + } + retired.verify()?; + let intent = read_intent(&root)?.ok_or_else(refused)?; + let complete = intent.complete_sha256.as_deref().ok_or_else(refused)?; + let (current, _) = load(candidate, engine, run)?; + if current.phase != "stopped-data-retained" + || digest(&host_pin_recovery::read_raw( + &root.join("state.json"), + LIMIT, + )?) != complete + || serde_json::to_vec_pretty(¤t).map_err(|_| refused())? + != serde_json::to_vec_pretty(receipt).map_err(|_| refused())? + { + return Err(refused()); + } + let foreground_root = foreground::transport::root(candidate, run)?; + verify_static_with( + candidate, + run, + engine, + &intent.selection, + &intent.original, + &foreground_root, + || retired.verify(), + )?; + if completed_receipt(candidate, engine, &intent, ¤t, &root)? != complete { + return Err(refused()); + } + let retirement_bytes = host_pin_recovery::read_raw(&root.join(RETIREMENT), 65536)?; + let retirement: Retirement = + serde_json::from_slice(&retirement_bytes).map_err(|_| refused())?; + if retirement.version != 1 + || retirement.selection_sha256 != intent.selection_sha256 + || retirement.complete_sha256 != complete + || retirement.run != run + || retirement.owner != current.owner + { + return Err(refused()); + } + let old_bytes = private_input(&intent.selection.original_owner_path, 1024 * 1024)?; + let old: Owner = serde_json::from_slice(&old_bytes).map_err(|_| refused())?; + let inspection_bytes = private_input(&intent.selection.filesystem_inspection_path, 64 * 1024)?; + let inspection: FilesystemInspection = + serde_json::from_slice(&inspection_bytes).map_err(|_| refused())?; + inspection.verify(&old_bytes, &old)?; + retired.verify()?; + Ok(Some(CompletedSourceProof { + original_ready_sha256: intent.selection.graph_sha256.clone(), + completed_stopped_sha256: complete.into(), + intent_raw_sha256: digest(&host_pin_recovery::read_raw( + &root.join(INTENT), + 4 * 1024 * 1024, + )?), + retirement_raw_sha256: digest(&retirement_bytes), + original_owner_sha256: intent.selection.original_owner_sha256.clone(), + current_owner_sha256: intent.selection.current_owner_sha256.clone(), + host_boot_micros: intent.selection.host_boot_micros, + previous_guest_boot: intent.selection.previous_guest_boot.clone(), + current_guest_boot: intent.selection.current_guest_boot.clone(), + old_share: old.project_share, + current_share: inspection.project_share, + retained_volumes: intent.selection.retained_volumes.clone(), + })) +} + +/// Explicit, hash-selected cleanup and separate durable absent-publisher +/// retirement. This command never removes a selected graph volume. +pub fn recover( + candidate: &Candidate, + run: &str, + expected: &str, + original_owner_path: &Path, + filesystem_inspection_path: &Path, +) -> Result { + if !super::hex(expected, 64) { + return Err(refused()); + } + // Foreground publication takes this lock before the VM operation lease. + // Following the same order prevents a new publisher from racing admission. + let reservation = Reservation::acquire(candidate, run)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let root = directory(candidate, run)?; + retain_interrupted_publications(&root)?; + let mut intent = if let Some(existing) = read_intent(&root)? { + if existing.selection_sha256 != expected + || existing.selection.original_owner_path != original_owner_path + || existing.selection.filesystem_inspection_path != filesystem_inspection_path + { + return Err(refused()); + } + existing + } else { + let (selection, original) = select_content( + candidate, + run, + original_owner_path, + filesystem_inspection_path, + &engine, + true, + )?; + if selection.digest()? != expected { + return Err(refused()); + } + let environment = + super::environment::cleanup_inventory(candidate, &engine, &original, &root)?; + let environment_value = serde_json::to_value(&environment).map_err(|_| refused())?; + if environment_value != selection.environment_inventory { + return Err(refused()); + } + host_relay::cleanup_preflight(&engine, &original, &root, false)?; + let bridges = super::bridges::cleanup::capture_previous_boot_absence( + candidate, &engine, &original, &selection, + )?; + if super::bridges::cleanup::selection_absence_projection(&bridges)? + != selection.scoped_bridge_projection + { + return Err(refused()); + } + let prior_bridges = + super::bridges::cleanup::capture_prior_generation(&root, &bridges, &original)?; + if prior_bridges.is_some() + && !super::restore_history::confirms_prior_generation(&root, &original)? + { + return Err(refused()); + } + let selected = Intent { + version: 1, + selection_sha256: expected.into(), + selection, + original, + environment: environment_value, + bridges, + prior_bridges, + complete_sha256: None, + }; + reservation.verify()?; + state::write(&root.join(INTENT), &selected)?; + #[cfg(test)] + super::fault_pause(&root, run, "absent-after-intent")?; + selected + }; + let current = verify_static( + candidate, + run, + &engine, + &intent.selection, + &intent.original, + &reservation, + )?; + if intent.complete_sha256.is_none() { + if current.phase == "stopped-data-retained" { + // Cleanup may have completed immediately before the intent's final + // commit. Exact confirmation permits a deterministic retry. + intent.complete_sha256 = Some(completed_receipt( + candidate, &engine, &intent, ¤t, &root, + )?); + state::write(&root.join(INTENT), &intent)?; + } else { + super::bridges::cleanup::verify_remaining_absence( + candidate, + &engine, + ¤t, + &intent.bridges, + &intent.selection, + )?; + super::bridges::cleanup::verify_recovery_file( + &root, + &intent.bridges, + intent.prior_bridges.as_ref(), + )?; + super::bridges::cleanup::recover_persist(&root, &intent.bridges)?; + reservation.verify()?; + let cleaned = super::cleanup_owned_fenced( + candidate, + &engine, + current, + &root, + false, + true, + || { + verify_static( + candidate, + run, + &engine, + &intent.selection, + &intent.original, + &reservation, + ) + .map(|_| ()) + }, + )?; + reservation.verify()?; + intent.complete_sha256 = Some(completed_receipt( + candidate, &engine, &intent, &cleaned, &root, + )?); + state::write(&root.join(INTENT), &intent)?; + } + } + let (current, _) = load(candidate, &engine, run)?; + let complete = completed_receipt(candidate, &engine, &intent, ¤t, &root)?; + if intent.complete_sha256.as_deref() != Some(complete.as_str()) { + return Err(refused()); + } + if let Some(reservation_record) = &intent.selection.dependency_reservation { + dependency_slots::recover_cleaned( + candidate, + ¤t, + Some(( + DeviceRebind { + old: intent.selection.old_device, + current: intent.selection.new_device, + }, + reservation_record, + )), + )?; + } else { + dependency_slots::recover_cleaned(candidate, ¤t, None)?; + } + retire(candidate, &engine, &intent, ¤t, &root, &reservation)?; + archive_retired_rebind_under(candidate, &engine, ¤t, &|| reservation.verify())?; + Ok(json!({ + "run": run, + "phase": "stopped-data-retained", + "publisher_retired": true, + "data_retained": true, + "selection_sha256": expected, + "qualification": intent.selection.qualification, + })) +} + +#[cfg(test)] +mod tests { + use super::*; + use std::os::unix::fs::PermissionsExt; + + fn candidate() -> (super::super::tests::Fixture, Candidate) { + let fixture = super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + (fixture, candidate) + } + + #[test] + fn current_absent_completion_precedes_historical_enrollment_but_never_partial_proofs() { + let (_fixture, candidate) = candidate(); + let root = directory(&candidate, &"a".repeat(32)).unwrap(); + state::private_directory(&root).unwrap(); + let root = &root; + let original: Receipt = serde_json::from_value(json!({ + "version":1,"run":"a".repeat(32),"owner":"b".repeat(32), + "namespace":"c".repeat(64),"plan_id":"d".repeat(64), + "phase":"ready-observed","readiness":{},"resources":{"container:web":{ + "kind":"container","key":"web","name":"owned-web","id":"b".repeat(64),"image":null,"phase":"absent"}}, + "relay_startup":{"control_only":true,"guest_root":null, + "control_root":"/private/control","artifact":"e".repeat(64),"services":{}} + })) + .unwrap(); + let mut current = original.clone(); + current.phase = "stopped-data-retained".into(); + let original_hash = digest(&serde_json::to_vec_pretty(&original).unwrap()); + let complete = digest(&serde_json::to_vec_pretty(¤t).unwrap()); + let selection: Selection = serde_json::from_value(json!({ + "version":1,"candidate":root,"run":current.run,"owner":current.owner, + "namespace":current.namespace,"plan":current.plan_id, + "original_owner_path":"/private/owner","original_owner_sha256":"1".repeat(64), + "filesystem_inspection_path":"/private/inspection","filesystem_inspection_sha256":"2".repeat(64), + "current_owner_sha256":"3".repeat(64),"graph_sha256":original_hash, + "host_boot_micros":1,"old_device":2,"new_device":3, + "previous_guest_boot":"previous","current_guest_boot":"current", + "foreground_root":"/private/foreground","control_root":"/private/control", + "source_shared":null,"environment_inventory":{},"scoped_bridge_projection":{}, + "retained_volumes":{},"dependency_reservation":null,"qualification":"test" + })).unwrap(); + let intent = Intent { + version: 1, + selection_sha256: selection.digest().unwrap(), + selection, + original: original.clone(), + environment: json!({}), + bridges: serde_json::from_value(json!({"version":1,"owner":current.owner, + "boot":"current","run":current.run,"plan":current.plan_id, + "capacity":1,"serial":1,"selected":{}})) + .unwrap(), + prior_bridges: None, + complete_sha256: Some(complete.clone()), + }; + let retirement = Retirement { + version: 1, + selection_sha256: intent.selection_sha256.clone(), + complete_sha256: complete, + owner: current.owner.clone(), + run: current.run.clone(), + }; + state::write(&root.join("state.json"), ¤t).unwrap(); + state::write(&root.join(INTENT), &intent).unwrap(); + state::write(&root.join(RETIREMENT), &retirement).unwrap(); + state::write(&root.join("relay-cleanup-bridges.json"), &intent.bridges).unwrap(); + // A completed older recovery is retained as exact history, not current authority. + let mut prior = current.clone(); + prior.resources.get_mut("container:web").unwrap().id = Some("a".repeat(64)); + let mut prior_original = prior.clone(); + prior_original.phase = "ready-observed".into(); + let dead = json!({ + "version":1,"original_sha256":digest(&serde_json::to_vec_pretty(&prior_original).unwrap()), + "owner_sha256":"4".repeat(64),"old_boot":"older","new_boot":"previous", + "original":prior_original,"environment":null,"bridges":null, + "complete_sha256":digest(&serde_json::to_vec_pretty(&prior).unwrap()) + }); + state::write(&root.join("dead-owner-cleanup.json"), &dead).unwrap(); + state::write( + &root.join("restore-history.json"), + &json!({ + "version":1,"run":current.run,"owner":current.owner,"namespace":current.namespace, + "truncated":false,"entries":[prior],"legacy":{} + }), + ) + .unwrap(); + assert_eq!( + super::super::dead_owner_cleanup::retained(root, ¤t) + .unwrap_err() + .code, + "graph_relay_enrollment" + ); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_ok()); + let live = json!({ + "version":1,"boot":"previous","original":prior_original, + "original_sha256":dead["original_sha256"],"foreground_sha256":"f".repeat(64), + "relay":{"bytes":[],"record_id":[1,1]},"environment":{},"bridges":intent.bridges, + "prior_bridges":null,"listeners_retired":true,"complete_sha256":dead["complete_sha256"] + }); + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_ok()); + for (name, valid) in [ + ("dead-owner-cleanup.json", &dead), + ("live-owner-cleanup.json", &live), + ] { + for field in ["complete_sha256", "original_sha256"] { + let mut invalid = valid.clone(); + invalid[field] = if field == "complete_sha256" { + Value::Null + } else { + json!("0".repeat(64)) + }; + state::write(&root.join(name), &invalid).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + } + let mut invalid = valid.clone(); + invalid["original"]["owner"] = json!("0".repeat(32)); + state::write(&root.join(name), &invalid).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + state::write(&root.join(name), valid).unwrap(); + } + let mut unretired = live.clone(); + unretired["listeners_retired"] = json!(false); + state::write(&root.join("live-owner-cleanup.json"), &unretired).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + let history = fs::read(root.join("restore-history.json")).unwrap(); + fs::remove_file(root.join("restore-history.json")).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + state::write( + &root.join("restore-history.json"), + &serde_json::from_slice::(&history).unwrap(), + ) + .unwrap(); + let mut changed_bridges = serde_json::to_value(&intent.bridges).unwrap(); + changed_bridges["serial"] = json!(99); + state::write(&root.join("relay-cleanup-bridges.json"), &changed_bridges).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + state::write(&root.join("relay-cleanup-bridges.json"), &intent.bridges).unwrap(); + let before = fs::read(root.join("state.json")).unwrap(); + for pending in [ + "state.pending", + "dead-owner-cleanup.pending", + "live-owner-cleanup.pending", + "absent-publication-cleanup.pending", + "absent-publication-retirement.pending", + ] { + fs::write(root.join(pending), b"interrupted").unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + fs::remove_file(root.join(pending)).unwrap(); + } + for change in 0..4 { + let mut invalid = intent.clone(); + match change { + 0 => invalid.complete_sha256 = None, + 1 => invalid.complete_sha256 = Some("6".repeat(64)), + 2 => invalid.selection.owner = "7".repeat(32), + _ => invalid.selection.plan = "8".repeat(64), + } + state::write(&root.join(INTENT), &invalid).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + } + state::write(&root.join(INTENT), &intent).unwrap(); + let mut invalid = retirement.clone(); + invalid.complete_sha256 = "9".repeat(64); + state::write(&root.join(RETIREMENT), &invalid).unwrap(); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + assert_eq!(fs::read(root.join("state.json")).unwrap(), before); + + // An older absence retirement must not shadow a newer exact same-boot + // completion when the restored foreground publication is selected. + state::write(&root.join(RETIREMENT), &retirement).unwrap(); + super::super::restore_history::retain(root, ¤t).unwrap(); + let mut newer = current.clone(); + newer.resources.get_mut("container:web").unwrap().id = Some("c".repeat(64)); + let mut ready = newer.clone(); + ready.phase = "ready-observed".into(); + ready.resources.get_mut("container:web").unwrap().phase = "started".into(); + let live = json!({"version":1,"boot":"current","original":ready, + "original_sha256":digest(&serde_json::to_vec_pretty(&ready).unwrap()), + "foreground_sha256":"f".repeat(64),"relay":{"bytes":[],"record_id":[1,1]}, + "environment":{},"bridges":intent.bridges,"prior_bridges":null, + "listeners_retired":true, + "complete_sha256":digest(&serde_json::to_vec_pretty(&newer).unwrap())}); + state::write(&root.join("state.json"), &newer).unwrap(); + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + let before = fs::read(root.join("state.json")).unwrap(); + publication_allowed(&candidate, &newer.run, true).unwrap(); + assert!(publication_allowed(&candidate, &newer.run, false).is_err()); + // Equivalent JSON does not preserve the exact selected receipt bytes. + fs::write(root.join("state.json"), serde_json::to_vec(&newer).unwrap()).unwrap(); + assert!(publication_allowed(&candidate, &newer.run, true).is_err()); + state::write(&root.join("state.json"), &newer).unwrap(); + publication_allowed(&candidate, &newer.run, true).unwrap(); + let mut altered = live.clone(); + altered["complete_sha256"] = json!("9".repeat(64)); + state::write(&root.join("live-owner-cleanup.json"), &altered).unwrap(); + assert!(publication_allowed(&candidate, &newer.run, true).is_err()); + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + let history = root.join("restore-history.json"); + fs::rename(&history, root.join("selected-history.json")).unwrap(); + assert!(publication_allowed(&candidate, &newer.run, true).is_err()); + fs::rename(root.join("selected-history.json"), history).unwrap(); + for pending in [ + "absent-publication-cleanup.pending", + "absent-publication-retirement.pending", + ] { + fs::write(root.join(pending), b"interrupted").unwrap(); + assert!(publication_allowed(&candidate, &newer.run, true).is_err()); + fs::remove_file(root.join(pending)).unwrap(); + } + publication_allowed(&candidate, &newer.run, true).unwrap(); + assert_eq!(fs::read(root.join("state.json")).unwrap(), before); + } + + #[test] + fn read_only_inspection_of_existing_reservation_never_creates_a_root() { + let (_fixture, candidate) = candidate(); + let run = "a".repeat(32); + let root = foreground::transport::root(&candidate, &run).unwrap(); + assert!( + Reservation::inspect_existing(&candidate, &run) + .unwrap() + .is_none() + ); + assert!(absent(&root).unwrap()); + let acquired = Reservation::acquire(&candidate, &run).unwrap(); + acquired.verify().unwrap(); + drop(acquired); + let inspected = Reservation::inspect_existing(&candidate, &run) + .unwrap() + .unwrap(); + inspected.verify().unwrap(); + assert_eq!( + fs::read_dir(&root).unwrap().count(), + 1, + "only the lock-only reservation is present" + ); + drop(inspected); + fs::remove_file(root.join("operation.lock")).unwrap(); + fs::remove_dir(root).unwrap(); + } + + #[test] + fn substituted_foreground_lock_path_refuses_without_deleting_replacement() { + let (_fixture, candidate) = candidate(); + let run = "b".repeat(32); + let acquired = Reservation::acquire(&candidate, &run).unwrap(); + let root = acquired.root.clone(); + let path = acquired.root.join("operation.lock"); + let saved = acquired.root.join("operation.held"); + fs::rename(&path, &saved).unwrap(); + fs::write(&path, b"foreign replacement").unwrap(); + fs::set_permissions(&path, fs::Permissions::from_mode(0o600)).unwrap(); + let replacement = fs::symlink_metadata(&path).unwrap(); + assert!(acquired.verify().is_err()); + let after = fs::symlink_metadata(&path).unwrap(); + assert_eq!( + (after.dev(), after.ino()), + (replacement.dev(), replacement.ino()) + ); + assert_eq!(fs::read(&path).unwrap(), b"foreign replacement"); + drop(acquired); + fs::remove_file(&path).unwrap(); + fs::rename(saved, &path).unwrap(); + fs::remove_file(&path).unwrap(); + fs::remove_dir(&root).unwrap(); + } + + #[test] + fn interrupted_intent_blocks_every_ordinary_publisher_mode() { + let (_fixture, candidate) = candidate(); + let run = "c".repeat(32); + let root = directory(&candidate, &run).unwrap(); + state::private_directory(&root).unwrap(); + fs::write(root.join("absent-publication-cleanup.pending"), b"partial").unwrap(); + assert!(publication_allowed(&candidate, &run, false).is_err()); + assert!(publication_allowed(&candidate, &run, true).is_err()); + assert!(absent(&root.join(INTENT)).unwrap()); + } + + #[test] + fn legacy_share_rebind_refuses_every_non_device_change() { + let old = ProjectShareIntent { + project: PathBuf::from("/private/tmp/selected"), + guest_path: "/workspace".into(), + device: 10, + inode: 92, + unfiltered_source: true, + }; + let mut current = old.clone(); + current.device = 11; + assert!(device_only_share(Some(&old), Some(¤t), 10, 11)); + assert!(!device_only_share(None, Some(¤t), 10, 11)); + assert!(!device_only_share(Some(&old), None, 10, 11)); + assert!(!device_only_share(Some(&old), Some(¤t), 9, 11)); + for changed in [ + { + let mut value = current.clone(); + value.project = PathBuf::from("/private/tmp/other"); + value + }, + { + let mut value = current.clone(); + value.guest_path = "/different".into(); + value + }, + { + let mut value = current.clone(); + value.inode += 1; + value + }, + { + let mut value = current.clone(); + value.unfiltered_source = false; + value + }, + ] { + assert!(!device_only_share(Some(&old), Some(&changed), 10, 11)); + } + } +} diff --git a/packages/runtime-core/src/provider/graph/absent_publication_cleanup/rebind_history.rs b/packages/runtime-core/src/provider/graph/absent_publication_cleanup/rebind_history.rs new file mode 100644 index 000000000..8900fdd9e --- /dev/null +++ b/packages/runtime-core/src/provider/graph/absent_publication_cleanup/rebind_history.rs @@ -0,0 +1,646 @@ +//! Metadata retirement for a completed prior-boot rebind. Historical receipts +//! select preserved evidence only; current ownership and live absence still +//! fence every archive write/rename under the caller's foreground lock. +use super::*; + +fn require_retirement( + root: &Path, + current: &Receipt, + intent: &Intent, +) -> Result<(), CandidateError> { + let complete = intent.complete_sha256.as_deref().ok_or_else(refused)?; + let retirement: Retirement = state::read_bounded(&root.join(RETIREMENT), 65536)?; + if intent.original.phase != "ready-observed" + || !super::super::hex(complete, 64) + || retirement.version != 1 + || retirement.selection_sha256 != intent.selection_sha256 + || retirement.complete_sha256 != complete + || retirement.run != current.run + || retirement.owner != current.owner + || current.run != intent.original.run + || current.owner != intent.original.owner + || current.namespace != intent.original.namespace + || current.plan_id != intent.original.plan_id + || current.phase != "stopped-data-retained" + { + return Err(refused()); + } + Ok(()) +} + +fn completion(root: &Path, current: &Receipt, intent: &Intent) -> Result { + require_retirement(root, current, intent)?; + let complete = intent.complete_sha256.as_deref().ok_or_else(refused)?; + let completed = + if digest(&serde_json::to_vec_pretty(current).map_err(|_| refused())?) == complete { + current.clone() + } else { + super::super::restore_history::completed_for_recovery(root, current, complete)? + .ok_or_else(refused)? + }; + if intent.original.phase != "ready-observed" + || completed.phase != "stopped-data-retained" + || !intent.selection.matches_graph(&completed) + || dead_owner_cleanup::immutable(&completed)? + != dead_owner_cleanup::immutable(&intent.original)? + || intent.original.resources.len() != completed.resources.len() + || intent.original.resources.iter().any(|(key, before)| { + completed.resources.get(key).is_none_or(|after| { + before.kind != after.kind + || before.id != after.id + || before.name != after.name + || (after.kind != Kind::Volume && after.phase != "absent") + }) + }) + || completed + .resources + .iter() + .filter(|(_, r)| r.kind == Kind::Volume) + .any(|(key, old)| { + current.resources.get(key).is_none_or(|now| { + now.kind != Kind::Volume || old.name != now.name || old.cache != now.cache + }) + }) + || completed + .resources + .values() + .filter(|r| r.kind == Kind::Volume) + .count() + != current + .resources + .values() + .filter(|r| r.kind == Kind::Volume) + .count() + { + return Err(refused()); + } + Ok(completed) +} + +/// Completed archival grants only a metadata no-op. Current retained compute +/// must still have its own exact recovery or acknowledged ordinary cleanup. +fn already_archived( + root: &Path, + current: &Receipt, + intent: &Intent, + verify_lock: &dyn Fn() -> Result<(), CandidateError>, +) -> Result { + verify_lock()?; + if !startup::retired_dependency_rebind_archive_complete( + root, + &intent.original, + &intent.selection.previous_guest_boot, + )? { + return Ok(false); + } + let verify = || { + require_retirement(root, current, intent)?; + no_pending(root)?; + for name in [ + "live-owner-cleanup.pending", + "source-device-rebind.pending", + "restore-history.pending", + ] { + require_absent(&root.join(name))?; + } + let recorded = read_intent(root)?.ok_or_else(refused)?; + if serde_json::to_vec_pretty(&recorded).map_err(|_| refused())? + != serde_json::to_vec_pretty(intent).map_err(|_| refused())? + || host_pin_recovery::read_raw(&root.join(INTENT), 4 * 1024 * 1024)? + != serde_json::to_vec_pretty(intent).map_err(|_| refused())? + || host_pin_recovery::read_raw(&root.join("state.json"), LIMIT)? + != serde_json::to_vec_pretty(current).map_err(|_| refused())? + { + return Err(refused()); + } + super::super::cleanup_enrollment::retention(root, current) + }; + verify()?; + verify_lock()?; + if !startup::retired_dependency_rebind_archive_complete( + root, + &intent.original, + &intent.selection.previous_guest_boot, + )? { + return Err(refused()); + } + verify()?; + Ok(true) +} + +fn compute_absent(engine: &Engine<'_>, receipt: &Receipt) -> Result<(), CandidateError> { + for resource in receipt + .resources + .values() + .filter(|r| r.kind != Kind::Volume) + { + let mut by_name = resource.clone(); + by_name.id = None; + if resource.phase != "absent" + || inspect_resource(engine, receipt, resource)?.is_some() + || inspect_resource(engine, receipt, &by_name)?.is_some() + { + return Err(refused()); + } + } + Ok(()) +} + +/// Called after completed absence retirement, or before the first effect of a +/// newly selected retained restore. The verifier must bind the held foreground +/// lock descriptor to its pathname, including after interrupted archival. +pub(in crate::provider::graph) fn archive_retired_rebind_under( + candidate: &Candidate, + engine: &Engine<'_>, + current: &Receipt, + verify_lock: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + let root = directory(candidate, ¤t.run)?; + let Some(intent) = read_intent(&root)? else { + return Ok(()); + }; + let boot = startup::dependency_rebind_boot(&root, current)?; + // A later current-boot journal belongs to its own normal cleanup. The old + // absence intent cannot authorize retiring that generation. + if boot.as_deref() == Some(engine.guest().boot_id()) { + return Ok(()); + } + let generation = super::super::service_exec_generation(&intent.original)?; + let history = root.join(format!("dependency-rebind-history-{generation}")); + if boot.is_none() && absent(&root.join("dependency-rebind.pending"))? && absent(&history)? { + return Ok(()); + } + if boot + .as_deref() + .is_some_and(|boot| boot != intent.selection.previous_guest_boot) + { + return Err(refused()); + } + if already_archived(&root, current, &intent, verify_lock)? { + return Ok(()); + } + let completed = completion(&root, current, &intent)?; + let current_bytes = serde_json::to_vec_pretty(current).map_err(|_| refused())?; + let intent_bytes = host_pin_recovery::read_raw(&root.join(INTENT), 4 * 1024 * 1024)?; + let retirement_bytes = host_pin_recovery::read_raw(&root.join(RETIREMENT), 65536)?; + if serde_json::to_vec_pretty(&intent).map_err(|_| refused())? != intent_bytes { + return Err(refused()); + } + let verify = || { + verify_lock()?; + no_pending(&root)?; + for name in [ + "live-owner-cleanup.pending", + "source-device-rebind.pending", + "restore-history.pending", + ] { + require_absent(&root.join(name))?; + } + if host_pin_recovery::read_raw(&root.join("state.json"), LIMIT)? != current_bytes + || host_pin_recovery::read_raw(&root.join(INTENT), 4 * 1024 * 1024)? != intent_bytes + || host_pin_recovery::read_raw(&root.join(RETIREMENT), 65536)? != retirement_bytes + || serde_json::to_vec_pretty(&completion(&root, current, &intent)?) + .map_err(|_| refused())? + != serde_json::to_vec_pretty(&completed).map_err(|_| refused())? + { + return Err(refused()); + } + let selection = &intent.selection; + if selection.candidate != candidate.checkout || selection.version != 1 { + return Err(refused()); + } + let old_bytes = private_input(&selection.original_owner_path, 1024 * 1024)?; + let inspection_bytes = private_input(&selection.filesystem_inspection_path, 65536)?; + if digest(&old_bytes) != selection.original_owner_sha256 + || digest(&inspection_bytes) != selection.filesystem_inspection_sha256 + { + return Err(refused()); + } + let old: Owner = serde_json::from_slice(&old_bytes).map_err(|_| refused())?; + let inspection: FilesystemInspection = + serde_json::from_slice(&inspection_bytes).map_err(|_| refused())?; + inspection.verify(&old_bytes, &old)?; + let (owner, owner_sha) = current_owner(candidate, &old, &inspection, engine)?; + if owner_sha != selection.current_owner_sha256 + || owner.token != current.owner + || owner.guest_boot_id.as_deref() != Some(&selection.current_guest_boot) + || owner.previous_guest_boot_id.as_deref() != Some(&selection.previous_guest_boot) + || inspection.host_boot_micros != selection.host_boot_micros + || inspection.old_device != selection.old_device + || inspection.new_device != selection.new_device + { + return Err(refused()); + } + super::super::source_device_rebind::verify_cache_scope_cleanup_origin( + candidate, + current, + &digest(&intent_bytes), + &digest(&retirement_bytes), + )?; + host_pin_recovery::verify_volume_projections( + engine, + &completed, + &selection.retained_volumes, + )?; + host_pin_recovery::verify_volume_projections(engine, current, &selection.retained_volumes)?; + compute_absent(engine, &completed)?; + compute_absent(engine, current)?; + verify_lock() + }; + startup::archive_retired_dependency_rebind( + &root, + &intent.original, + &completed, + &intent.selection.previous_guest_boot, + &verify, + ) +} + +#[cfg(test)] +mod tests { + use super::*; + + fn archived_fixture() -> ( + crate::provider::graph::tests::Fixture, + Receipt, + Intent, + PathBuf, + ) { + let fixture = crate::provider::graph::tests::Fixture::new(); + let root = &fixture.0; + let original: Receipt = serde_json::from_value(json!({ + "version":1,"run":"a".repeat(32),"owner":"b".repeat(32),"namespace":"c".repeat(64), + "plan_id":"d".repeat(64),"phase":"ready-observed","readiness":{}, + "relay_startup":{"control_only":true,"guest_root":null,"control_root":"/private/control","artifact":"e".repeat(64),"services":{}}, + "resources":{"container:web":{"kind":"container","key":"web","name":"owned-web","id":"e".repeat(64),"image":null,"phase":"started"}, + "volume:data":{"kind":"volume","key":"data","name":"owned-data","id":null,"image":null,"phase":"created"}} + })).unwrap(); + let mut stopped = original.clone(); + stopped.phase = "stopped-data-retained".into(); + stopped.resources.get_mut("container:web").unwrap().phase = "absent".into(); + let selection: Selection = serde_json::from_value(json!({ + "version":1,"candidate":root,"run":original.run,"owner":original.owner, + "namespace":original.namespace,"plan":original.plan_id, + "original_owner_path":"/private/owner","original_owner_sha256":"1".repeat(64), + "filesystem_inspection_path":"/private/inspection","filesystem_inspection_sha256":"2".repeat(64), + "current_owner_sha256":"3".repeat(64),"graph_sha256":digest(&serde_json::to_vec_pretty(&original).unwrap()), + "host_boot_micros":1,"old_device":2,"new_device":3,"previous_guest_boot":"old","current_guest_boot":"current", + "foreground_root":"/private/foreground","control_root":"/private/control","source_shared":null, + "environment_inventory":{},"scoped_bridge_projection":{},"retained_volumes":{},"dependency_reservation":null,"qualification":"test" + })).unwrap(); + let intent = Intent { version:1,selection_sha256:selection.digest().unwrap(),selection, + original:original.clone(),environment:json!({}),bridges:serde_json::from_value(json!({ + "version":1,"owner":original.owner,"boot":"current","run":original.run,"plan":original.plan_id,"capacity":1,"serial":1,"selected":{} + })).unwrap(),prior_bridges:None,complete_sha256:Some(digest(&serde_json::to_vec_pretty(&stopped).unwrap())) }; + state::write(&root.join(INTENT), &intent).unwrap(); + state::write( + &root.join(RETIREMENT), + &Retirement { + version: 1, + selection_sha256: intent.selection_sha256.clone(), + complete_sha256: intent.complete_sha256.clone().unwrap(), + owner: original.owner.clone(), + run: original.run.clone(), + }, + ) + .unwrap(); + for generation in 1..=12 { + let mut newer = stopped.clone(); + newer.resources.get_mut("container:web").unwrap().id = + Some(format!("{generation:064x}")); + crate::provider::graph::restore_history::retain(root, &newer).unwrap(); + } + let mut current = stopped.clone(); + current.resources.get_mut("container:web").unwrap().id = Some("f".repeat(64)); + let mut ready = current.clone(); + ready.phase = "ready-observed".into(); + ready.resources.get_mut("container:web").unwrap().phase = "started".into(); + state::write(&root.join("state.json"), ¤t).unwrap(); + state::write(&root.join("live-owner-cleanup.json"), &json!({ + "version":1,"boot":"current","original":ready, + "original_sha256":digest(&serde_json::to_vec_pretty(&ready).unwrap()), + "foreground_sha256":"f".repeat(64),"relay":{"bytes":[],"record_id":[1,1]}, + "environment":{},"bridges":intent.bridges,"prior_bridges":null, + "listeners_retired":true,"complete_sha256":digest(&serde_json::to_vec_pretty(¤t).unwrap()) + })).unwrap(); + let generation = crate::provider::graph::service_exec_generation(&original).unwrap(); + let history = root.join(format!("dependency-rebind-history-{generation}")); + state::private_directory(&history).unwrap(); + state::write(&history.join("dependency-rebind.json"), &json!({ + "version":1,"operation":"1".repeat(32),"run":original.run,"owner":original.owner, + "boot":"old","expected_generation":"2".repeat(64),"phase":"completed", + "slots":{"0":{"before":"3".repeat(64),"after":"4".repeat(64),"bindings":[["web","content"]]}}, + "processes":{},"completed_generation":generation + })).unwrap(); + state::write(&history.join("proof.json"), &json!({ + "version":1,"run":original.run,"owner":original.owner,"boot":"old", + "original_generation":generation, + "cleaned_generation":crate::provider::graph::service_exec_generation(&stopped).unwrap(), + "artifacts":{"dependency-rebind.json":digest(&fs::read(history.join("dependency-rebind.json")).unwrap())}, + "complete":true + })).unwrap(); + (fixture, current, intent, history) + } + + fn retained_bytes(root: &Path, history: &Path) -> Vec>> { + [ + root.join(INTENT), + root.join(RETIREMENT), + root.join("state.json"), + root.join("live-owner-cleanup.json"), + root.join("restore-history.json"), + root.join("dependency-rebind.json"), + root.join("dependency-rebind.pending"), + history.join("dependency-rebind.json"), + history.join("proof.json"), + history.join("proof.pending"), + ] + .iter() + .map(|path| fs::read(path).ok()) + .collect() + } + + #[test] + fn completed_archive_is_read_only_after_stopped_receipt_eviction() { + let (fixture, current, intent, history) = archived_fixture(); + let root = &fixture.0; + assert!(completion(root, ¤t, &intent).is_err()); + let before = retained_bytes(root, &history); + let checks = std::cell::Cell::new(0); + assert!( + already_archived(root, ¤t, &intent, &|| { + checks.set(checks.get() + 1); + Ok(()) + }) + .unwrap() + ); + assert_eq!(checks.get(), 2); + assert_eq!(retained_bytes(root, &history), before); + } + + #[test] + fn retirement_requires_the_original_ready_generation() { + let (fixture, current, mut intent, history) = archived_fixture(); + let before = retained_bytes(&fixture.0, &history); + require_retirement(&fixture.0, ¤t, &intent).unwrap(); + for phase in ["stopped-data-retained", "cleanup-intent", "failed"] { + intent.original.phase = phase.into(); + assert!(require_retirement(&fixture.0, ¤t, &intent).is_err()); + } + assert_eq!(retained_bytes(&fixture.0, &history), before); + } + + #[test] + fn first_or_interrupted_archive_keeps_exact_stopped_receipt_requirement() { + for fault in [ + "missing", + "incomplete", + "active", + "pending", + "proof-pending", + ] { + let (fixture, current, intent, history) = archived_fixture(); + let root = &fixture.0; + match fault { + "missing" => fs::remove_file(history.join("proof.json")).unwrap(), + "incomplete" => { + let mut proof: Value = state::read(&history.join("proof.json")).unwrap(); + proof["complete"] = json!(false); + state::write(&history.join("proof.json"), &proof).unwrap(); + } + "active" => fs::copy( + history.join("dependency-rebind.json"), + root.join("dependency-rebind.json"), + ) + .map(|_| ()) + .unwrap(), + "pending" => { + state::write(&root.join("dependency-rebind.pending"), &json!({})).unwrap() + } + "proof-pending" => { + state::write(&history.join("proof.pending"), &json!({})).unwrap() + } + _ => unreachable!(), + } + let before = retained_bytes(root, &history); + assert!( + !already_archived(root, ¤t, &intent, &|| Ok(())).unwrap(), + "{fault}" + ); + assert!(completion(root, ¤t, &intent).is_err(), "{fault}"); + assert_eq!(retained_bytes(root, &history), before); + } + } + + #[test] + fn archive_noop_refuses_changed_proofs_or_unconfirmed_current_cleanup() { + for fault in [ + "proof-owner", + "proof-boot", + "proof-generation", + "proof-hash", + "journal-phase", + "journal-generation", + "current-owner", + "current-state", + "retirement", + "live-incomplete", + "live-digest", + "pending", + "lock", + ] { + let (fixture, mut current, intent, history) = archived_fixture(); + let root = &fixture.0; + match fault { + "proof-owner" | "proof-boot" | "proof-generation" | "proof-hash" => { + let mut proof: Value = state::read(&history.join("proof.json")).unwrap(); + match fault { + "proof-owner" => proof["owner"] = json!("9".repeat(32)), + "proof-boot" => proof["boot"] = json!("foreign"), + "proof-generation" => proof["original_generation"] = json!("9".repeat(64)), + _ => proof["artifacts"]["dependency-rebind.json"] = json!("9".repeat(64)), + } + state::write(&history.join("proof.json"), &proof).unwrap(); + } + "journal-phase" | "journal-generation" => { + let mut journal: Value = + state::read(&history.join("dependency-rebind.json")).unwrap(); + if fault == "journal-phase" { + journal["phase"] = json!("intent"); + } else { + journal["completed_generation"] = json!("9".repeat(64)); + } + state::write(&history.join("dependency-rebind.json"), &journal).unwrap(); + let mut proof: Value = state::read(&history.join("proof.json")).unwrap(); + proof["artifacts"]["dependency-rebind.json"] = json!(digest( + &fs::read(history.join("dependency-rebind.json")).unwrap() + )); + state::write(&history.join("proof.json"), &proof).unwrap(); + } + "current-owner" => { + current.owner = "9".repeat(32); + state::write(&root.join("state.json"), ¤t).unwrap(); + } + "current-state" => { + let mut changed = current.clone(); + changed.resources.get_mut("container:web").unwrap().id = Some("9".repeat(64)); + state::write(&root.join("state.json"), &changed).unwrap(); + } + "retirement" => { + let mut retirement: Value = state::read(&root.join(RETIREMENT)).unwrap(); + retirement["complete_sha256"] = json!("9".repeat(64)); + state::write(&root.join(RETIREMENT), &retirement).unwrap(); + } + "live-incomplete" | "live-digest" => { + let mut live: Value = + state::read(&root.join("live-owner-cleanup.json")).unwrap(); + if fault == "live-incomplete" { + live["complete_sha256"] = Value::Null; + } else { + live["original_sha256"] = json!("9".repeat(64)); + } + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + } + "pending" => { + state::write(&root.join("live-owner-cleanup.pending"), &json!({})).unwrap() + } + "lock" => {} + _ => unreachable!(), + } + let before = retained_bytes(root, &history); + let checks = std::cell::Cell::new(0); + assert!( + already_archived(root, ¤t, &intent, &|| { + checks.set(checks.get() + 1); + if fault == "lock" && checks.get() == 2 { + Err(refused()) + } else { + Ok(()) + } + }) + .is_err(), + "{fault}" + ); + assert_eq!(retained_bytes(root, &history), before); + } + } + + #[test] + fn archive_noop_rechecks_evidence_after_final_lock_verification() { + for fault in ["current-state", "current-cleanup", "archive", "pending"] { + let (fixture, current, intent, history) = archived_fixture(); + let root = &fixture.0; + let checks = std::cell::Cell::new(0); + let expected = std::cell::RefCell::new(retained_bytes(root, &history)); + assert!( + already_archived(root, ¤t, &intent, &|| { + checks.set(checks.get() + 1); + if checks.get() == 2 { + match fault { + "current-state" => { + let mut changed = current.clone(); + changed.resources.get_mut("container:web").unwrap().id = + Some("9".repeat(64)); + state::write(&root.join("state.json"), &changed).unwrap(); + } + "current-cleanup" => { + let mut live: Value = + state::read(&root.join("live-owner-cleanup.json")).unwrap(); + live["complete_sha256"] = Value::Null; + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + } + "archive" => { + let mut proof: Value = + state::read(&history.join("proof.json")).unwrap(); + proof["complete"] = json!(false); + state::write(&history.join("proof.json"), &proof).unwrap(); + } + "pending" => { + state::write(&root.join("dependency-rebind.pending"), &json!({})) + .unwrap() + } + _ => unreachable!(), + } + *expected.borrow_mut() = retained_bytes(root, &history); + } + Ok(()) + }) + .is_err(), + "{fault}" + ); + assert_eq!(checks.get(), 2); + assert_eq!(retained_bytes(root, &history), *expected.borrow()); + } + } + + #[test] + fn historical_completion_requires_exact_stopped_digest_and_retirement() { + let fixture = crate::provider::graph::tests::Fixture::new(); + let root = &fixture.0; + let original: Receipt = serde_json::from_value(json!({ + "version":1,"run":"a".repeat(32),"owner":"b".repeat(32),"namespace":"c".repeat(64), + "plan_id":"d".repeat(64),"phase":"ready-observed","readiness":{}, + "resources":{"container:web":{"kind":"container","key":"web","name":"owned-web","id":"e".repeat(64),"image":null,"phase":"present"}, + "volume:data":{"kind":"volume","key":"data","name":"owned-data","id":null,"image":null,"phase":"present"}} + })).unwrap(); + let mut stopped = original.clone(); + stopped.phase = "stopped-data-retained".into(); + stopped.resources.get_mut("container:web").unwrap().phase = "absent".into(); + let mut current = stopped.clone(); + current.resources.get_mut("container:web").unwrap().id = Some("f".repeat(64)); + let selection: Selection = serde_json::from_value(json!({ + "version":1,"candidate":root,"run":original.run,"owner":original.owner, + "namespace":original.namespace,"plan":original.plan_id, + "original_owner_path":"/private/owner","original_owner_sha256":"1".repeat(64), + "filesystem_inspection_path":"/private/inspection","filesystem_inspection_sha256":"2".repeat(64), + "current_owner_sha256":"3".repeat(64),"graph_sha256":digest(&serde_json::to_vec_pretty(&original).unwrap()), + "host_boot_micros":1,"old_device":2,"new_device":3,"previous_guest_boot":"old","current_guest_boot":"current", + "foreground_root":"/private/foreground","control_root":"/private/control","source_shared":null, + "environment_inventory":{},"scoped_bridge_projection":{},"retained_volumes":{},"dependency_reservation":null,"qualification":"test" + })).unwrap(); + let intent = Intent { version:1,selection_sha256:selection.digest().unwrap(),selection, + original:original.clone(),environment:json!({}),bridges:serde_json::from_value(json!({ + "version":1,"owner":original.owner,"boot":"current","run":original.run,"plan":original.plan_id,"capacity":1,"serial":1,"selected":{} + })).unwrap(),prior_bridges:None,complete_sha256:Some(digest(&serde_json::to_vec_pretty(&stopped).unwrap())) }; + let retirement = Retirement { + version: 1, + selection_sha256: intent.selection_sha256.clone(), + complete_sha256: intent.complete_sha256.clone().unwrap(), + owner: current.owner.clone(), + run: current.run.clone(), + }; + state::write(&root.join(RETIREMENT), &retirement).unwrap(); + assert!(completion(root, ¤t, &intent).is_err()); + crate::provider::graph::restore_history::retain(root, &stopped).unwrap(); + assert_eq!( + serde_json::to_vec(&completion(root, ¤t, &intent).unwrap()).unwrap(), + serde_json::to_vec(&stopped).unwrap() + ); + for change in 0..4 { + let mut changed = current.clone(); + match change { + 0 => changed.owner = "0".repeat(32), + 1 => changed.plan_id = "0".repeat(64), + 2 => changed.phase = "ready-observed".into(), + _ => changed.resources.get_mut("volume:data").unwrap().name = "foreign-data".into(), + } + assert!(completion(root, &changed, &intent).is_err()); + } + let mut changed = retirement.clone(); + changed.selection_sha256 = "0".repeat(64); + state::write(&root.join(RETIREMENT), &changed).unwrap(); + assert!(completion(root, ¤t, &intent).is_err()); + state::write(&root.join(RETIREMENT), &retirement).unwrap(); + let mut changed = intent.clone(); + changed.complete_sha256 = Some("0".repeat(64)); + assert!(completion(root, ¤t, &changed).is_err()); + let mut changed = intent.clone(); + changed + .original + .resources + .get_mut("container:web") + .unwrap() + .id = Some("0".repeat(64)); + assert!(completion(root, ¤t, &changed).is_err()); + } +} diff --git a/packages/runtime-core/src/provider/graph/acknowledged_publisher.rs b/packages/runtime-core/src/provider/graph/acknowledged_publisher.rs new file mode 100644 index 000000000..8f204354d --- /dev/null +++ b/packages/runtime-core/src/provider/graph/acknowledged_publisher.rs @@ -0,0 +1,476 @@ +//! Publication-only recovery after ordinary cleanup was acknowledged, but a +//! later archival step failed. Historical recovery receipts grant no authority. +use super::*; +use sha2::{Digest, Sha256}; +use std::path::Path; + +const LIMIT: u64 = 2 * 1024 * 1024; + +pub struct AcknowledgedPublisherSelection<'a> { + pub run: &'a str, + pub owner: &'a str, + pub receipt_sha256: &'a str, + pub publisher_sha256: &'a str, +} + +fn refused() -> CandidateError { + error( + "graph_acknowledged_publisher", + "Publisher retirement requires the exact stopped receipt, current-boot cleanup acknowledgement, absent compute and unchanged retained data; evidence preserved.", + ) +} + +fn no_pending(root: &Path) -> Result<(), CandidateError> { + for name in [ + "state.pending", + "restore-history.pending", + "one-off.json", + "one-off.pending", + "one-off-normalization.json", + "dependency-rebind.pending", + "source-device-rebind.pending", + "dead-owner-cleanup.pending", + "live-owner-cleanup.pending", + "absent-publication-cleanup.pending", + "absent-publication-retirement.pending", + "relay-cleanup-bridges.pending", + "retired-data-removal.json", + "retired-data-removal.pending", + ] { + match fs::symlink_metadata(root.join(name)) { + Err(e) if e.kind() == std::io::ErrorKind::NotFound => {} + _ => return Err(refused()), + } + } + Ok(()) +} + +/// Cleanup proof is independent of the publisher selected for archival. +struct CleanupSelection<'a> { + run: &'a str, + owner: &'a str, + receipt_sha256: &'a str, +} +impl<'a> From<&AcknowledgedPublisherSelection<'a>> for CleanupSelection<'a> { + fn from(selected: &AcknowledgedPublisherSelection<'a>) -> Self { + Self { + run: selected.run, + owner: selected.owner, + receipt_sha256: selected.receipt_sha256, + } + } +} + +fn eligible( + receipt: &Receipt, + selected: &CleanupSelection<'_>, + boot: &str, +) -> Result<(), CandidateError> { + let marker = receipt.relay_cleanup.as_ref().ok_or_else(refused)?; + let startup = receipt.relay_startup.as_ref().ok_or_else(refused)?; + let context = host_relay::context(selected.owner, boot)?; + if receipt.run != selected.run + || receipt.owner != selected.owner + || receipt.phase != "stopped-data-retained" + || receipt.normalized_input.is_none() + || !marker.valid() + || marker.phase != cleanup_enrollment::Phase::Confirmed + || marker.runtime != context.runtime + || marker.boot != context.boot + || marker.control_root != startup.control_root + || receipt + .resources + .values() + .any(|resource| match resource.kind { + Kind::Volume => resource.phase != "created", + _ => resource.phase != "absent", + }) + { + return Err(refused()); + } + initializer_cache::require_resolved(receipt) +} + +fn exact_receipt(root: &Path, expected: &str) -> Result<(), CandidateError> { + no_pending(root)?; + let bytes = host_pin_recovery::read_raw(&root.join("state.json"), LIMIT)?; + if format!("{:x}", Sha256::digest(bytes)) != expected { + return Err(refused()); + } + Ok(()) +} + +fn acknowledged_effect(receipt: &Receipt, effect: [u8; 32]) -> Result<(), CandidateError> { + if receipt + .relay_cleanup + .as_ref() + .is_none_or(|marker| marker.effect != effect) + { + return Err(refused()); + } + Ok(()) +} + +fn verify_cleanup( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: &Receipt, + root: &Path, + selected: &CleanupSelection<'_>, +) -> Result<(), CandidateError> { + eligible(receipt, selected, engine.guest().boot_id())?; + exact_receipt(root, selected.receipt_sha256)?; + startup::require_dependency_rebind_complete(root, receipt)?; + let environment = environment::cleanup_inventory(candidate, engine, receipt, root)?; + let bridges = bridges::cleanup::read(engine, receipt, root)?; + acknowledged_effect( + receipt, + host_relay::cleanup_effect( + receipt, + engine.guest().boot_id(), + false, + &(&environment, &bridges), + )?, + )?; + host_relay::require_acknowledged_enrollment(receipt)?; + host_relay::inspect_cleanup(candidate, engine, receipt, false, &environment, &bridges)?; + engine.guest().verify() +} + +/// Normal foreground shutdown removes its publication after acknowledged +/// cleanup. Confirm that current proof under the publisher lock; historical +/// recovery sidecars and archived publisher journals cannot select this path. +/// A selected acknowledgement never falls back to historical recovery on error. +pub(super) fn confirm_retired( + candidate: &Candidate, + run: &str, + owner: &str, +) -> Result, CandidateError> { + let engine = Engine::connect_cleanup_wait(candidate)?; + let (receipt, root) = load(candidate, &engine, run)?; + if !uses_acknowledgement(&receipt, engine.guest().boot_id())? + || super::live_owner_cleanup::current_completion(&root, &receipt)? + || super::dead_owner_cleanup::current_completion(&root, &receipt)? + { + return Ok(None); + } + // Selection grants no effects. Release the Engine lease before acquiring + // the publication lock, matching serve-restore's lock order. The exact + // selected receipt and current boot are validated again under both locks. + drop(engine); + let retired = foreground::transport::Retired::acquire(candidate, run)?.ok_or_else(refused)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let bytes = serde_json::to_vec_pretty(&receipt).map_err(|_| refused())?; + let receipt_sha256 = format!("{:x}", Sha256::digest(&bytes)); + let selected = CleanupSelection { + run, + owner, + receipt_sha256: &receipt_sha256, + }; + verify_cleanup(candidate, &engine, &receipt, &root, &selected)?; + retired.verify()?; + exact_receipt(&root, &receipt_sha256)?; + engine.guest().verify()?; + Ok(Some( + json!({"run":run,"publisher_retired":true,"data_retained":true, + "same_boot":true,"acknowledged_cleanup":true}), + )) +} + +fn uses_acknowledgement(receipt: &Receipt, boot: &str) -> Result { + let context = host_relay::context(&receipt.owner, boot)?; + // Current-boot enrollment cannot borrow historical recovery on failure. + // A prior-boot marker may use only the existing exact completed recovery + // proof, including its already-retired publisher, on a later boot. + Ok(receipt + .relay_cleanup + .as_ref() + .is_some_and(|marker| marker.boot == context.boot)) +} + +/// Retire only a selected dead publication after independent current-boot ACK +/// and guest absence checks. The existing immutable publisher journal preserves +/// both original files and resumes either rename interruption. No graph receipt, +/// VM resource, dependency journal or persistent volume is changed here. +pub fn retire( + candidate: &Candidate, + selected: AcknowledgedPublisherSelection<'_>, +) -> Result { + if !hex(selected.run, 32) + || !hex(selected.owner, 32) + || !hex(selected.receipt_sha256, 64) + || !hex(selected.publisher_sha256, 64) + { + return Err(refused()); + } + // Match foreground restore's publication-before-Engine lock order; both + // locks stay held through observation and each retirement rename. + let publication = foreground::transport::root(candidate, selected.run)?; + let lock = state::Lock::acquire_existing(&publication)?; + host_pin_recovery::exact_lock_path(&publication, &lock)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let (receipt, root) = load(candidate, &engine, selected.run)?; + let verify = || { + host_pin_recovery::exact_lock_path(&publication, &lock)?; + verify_cleanup( + candidate, + &engine, + &receipt, + &root, + &CleanupSelection::from(&selected), + ) + }; + verify()?; + // This helper checks the selected publisher bytes, dead process, refused + // listener and exact lock/path identities before writing or moving files. + // Its durable journal permits retry after either original name has moved. + foreground::transport::retire_recovered_publisher_locked_fenced( + candidate, + selected.run, + selected.publisher_sha256, + selected.receipt_sha256, + None, + &lock, + &verify, + )?; + verify()?; + Ok( + json!({"run":selected.run,"publisher_retired":true,"data_retained":true,"same_boot":true,"acknowledged_cleanup":true}), + ) +} + +/// Release a selected same-boot dependency claim only after exact publisher +/// retirement and current acknowledged cleanup. Every socket must be absent; +/// the record is exclusively moved to history, never deleted or overwritten. +pub fn release_dependencies( + candidate: &Candidate, + selected: AcknowledgedPublisherSelection<'_>, + expected_reservation: &str, +) -> Result { + if !hex(selected.run, 32) + || !hex(selected.owner, 32) + || !hex(selected.receipt_sha256, 64) + || !hex(selected.publisher_sha256, 64) + || !hex(expected_reservation, 64) + { + return Err(refused()); + } + let retired = + foreground::transport::Retired::acquire(candidate, selected.run)?.ok_or_else(refused)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let (receipt, root) = load(candidate, &engine, selected.run)?; + let verify = || { + retired.verify_recovery( + candidate, + selected.run, + selected.publisher_sha256, + selected.receipt_sha256, + )?; + verify_cleanup( + candidate, + &engine, + &receipt, + &root, + &CleanupSelection::from(&selected), + ) + }; + verify()?; + let process = retired.publisher_process( + candidate, + selected.run, + selected.publisher_sha256, + selected.receipt_sha256, + )?; + dependency_slots::archive_acknowledged( + candidate, + &receipt, + engine.guest().boot_id(), + &process, + expected_reservation, + &verify, + )?; + verify()?; + Ok( + json!({"run":selected.run,"reservation_released":true,"record_retained":true,"data_retained":true,"same_boot":true}), + ) +} + +#[cfg(test)] +mod tests { + use super::*; + + fn receipt() -> Receipt { + let mut value: Receipt = serde_json::from_value(json!({"version":1,"run":"a".repeat(32),"owner":"b".repeat(32),"namespace":"c".repeat(64),"plan_id":"d".repeat(64),"phase":"stopped-data-retained","normalized_input":{"namespace":"c".repeat(64),"original_compose_sha256":"d".repeat(64),"normalized_compose_sha256":"e".repeat(64)},"readiness":{},"resources":{"container:web":{"kind":"container","key":"web","name":"owned-web","id":"f".repeat(64),"image":"sha256:".to_owned()+&"f".repeat(64),"phase":"absent"},"volume:data":{"kind":"volume","key":"data","name":"owned-data","id":null,"image":null,"phase":"created"}},"relay_startup":{"control_only":true,"guest_root":null,"control_root":"/private/owned","artifact":"e".repeat(64),"services":{}},"relay_cleanup":{"version":1,"runtime":vec![1;16],"boot":vec![2;16],"operation":vec![3;16],"effect":vec![4;32],"control_root":"/private/owned","phase":"confirmed"}})).unwrap(); + let context = host_relay::context(&value.owner, "fixture-boot").unwrap(); + let marker = value.relay_cleanup.as_mut().unwrap(); + marker.runtime = context.runtime; + marker.boot = context.boot; + value + } + + #[test] + fn current_acknowledgement_is_selected_even_when_invalid_and_never_uses_old_recovery() { + let mut value = receipt(); + assert!(uses_acknowledgement(&value, "fixture-boot").unwrap()); + assert!(!uses_acknowledgement(&value, "later-boot").unwrap()); + // Selection is deliberately separate from validation: a stale boot, + // incomplete graph or malformed confirmed marker must refuse in this + // path, not borrow authority from historical cleanup sidecars. + value.phase = "cleanup-intent".into(); + value.relay_cleanup.as_mut().unwrap().version = 0; + assert!(uses_acknowledgement(&value, "fixture-boot").unwrap()); + let selected = CleanupSelection { + run: &value.run, + owner: &value.owner, + receipt_sha256: &"1".repeat(64), + }; + assert!(eligible(&value, &selected, "fixture-boot").is_err()); + for phase in [ + cleanup_enrollment::Phase::Pending, + cleanup_enrollment::Phase::Dormant, + ] { + value.relay_cleanup.as_mut().unwrap().phase = phase; + assert!(uses_acknowledgement(&value, "fixture-boot").unwrap()); + assert!( + eligible( + &value, + &CleanupSelection { + run: &value.run, + owner: &value.owner, + receipt_sha256: &"1".repeat(64) + }, + "fixture-boot" + ) + .is_err() + ); + } + value.relay_cleanup = None; + assert!(!uses_acknowledgement(&value, "fixture-boot").unwrap()); + } + + #[test] + fn acknowledgement_eligibility_refuses_other_generations_and_partial_cleanup() { + let value = receipt(); + let selected = CleanupSelection { + run: &value.run, + owner: &value.owner, + receipt_sha256: &"1".repeat(64), + }; + eligible(&value, &selected, "fixture-boot").unwrap(); + assert!(eligible(&value, &selected, "other-boot").is_err()); + for field in [ + "run", + "owner", + "phase", + "normalized", + "marker", + "ack", + "root", + "compute", + "volume", + ] { + let mut changed = value.clone(); + match field { + "run" => changed.run = "9".repeat(32), + "owner" => changed.owner = "9".repeat(32), + "phase" => changed.phase = "failed-retained".into(), + "normalized" => changed.normalized_input = None, + "marker" => changed.relay_cleanup = None, + "ack" => { + changed.relay_cleanup.as_mut().unwrap().phase = + cleanup_enrollment::Phase::Pending + } + "root" => { + changed.relay_startup.as_mut().unwrap().control_root = "/private/other".into() + } + "compute" => { + changed.resources.get_mut("container:web").unwrap().phase = "started".into() + } + "volume" => { + changed.resources.get_mut("volume:data").unwrap().phase = "absent".into() + } + _ => unreachable!(), + } + assert!( + eligible(&changed, &selected, "fixture-boot").is_err(), + "{field}" + ); + } + } + + #[test] + fn pending_operations_and_changed_receipt_bytes_are_preserved() { + let fixture = super::super::tests::Fixture::new(); + let value = receipt(); + state::write(&fixture.0.join("state.json"), &value).unwrap(); + let before = fs::read(fixture.0.join("state.json")).unwrap(); + let selected = format!("{:x}", Sha256::digest(&before)); + exact_receipt(&fixture.0, &selected).unwrap(); + assert!(exact_receipt(&fixture.0, &"0".repeat(64)).is_err()); + for name in [ + "state.pending", + "restore-history.pending", + "dependency-rebind.pending", + "source-device-rebind.pending", + "one-off.json", + "retired-data-removal.json", + "relay-cleanup-bridges.pending", + ] { + fs::write(fixture.0.join(name), b"preserve").unwrap(); + assert!(exact_receipt(&fixture.0, &selected).is_err(), "{name}"); + assert_eq!(fs::read(fixture.0.join(name)).unwrap(), b"preserve"); + assert_eq!(fs::read(fixture.0.join("state.json")).unwrap(), before); + fs::remove_file(fixture.0.join(name)).unwrap(); + } + fs::write(fixture.0.join("state.json"), b"changed").unwrap(); + assert!(exact_receipt(&fixture.0, &selected).is_err()); + assert_eq!(fs::read(fixture.0.join("state.json")).unwrap(), b"changed"); + } + + #[test] + fn completed_old_boot_journal_is_inert_but_uncertain_rebind_refuses() { + let fixture = super::super::tests::Fixture::new(); + let value = receipt(); + let mut journal = json!({"version":1,"operation":"1".repeat(32),"run":value.run,"owner":value.owner,"boot":"prior-boot","expected_generation":"2".repeat(64),"phase":"completed","slots":{},"processes":{},"completed_services":["deps"],"completed_generation":"3".repeat(64)}); + let path = fixture.0.join("dependency-rebind.json"); + for phase in ["completed", "cleaned"] { + journal["phase"] = json!(phase); + state::write(&path, &journal).unwrap(); + let before = fs::read(&path).unwrap(); + startup::require_dependency_rebind_complete(&fixture.0, &value).unwrap(); + assert_eq!(fs::read(&path).unwrap(), before); + fs::write(fixture.0.join("dependency-rebind.pending"), b"partial").unwrap(); + assert!(startup::require_dependency_rebind_complete(&fixture.0, &value).is_err()); + assert_eq!(fs::read(&path).unwrap(), before); + fs::remove_file(fixture.0.join("dependency-rebind.pending")).unwrap(); + } + journal["phase"] = json!("provisioning"); + state::write(&path, &journal).unwrap(); + assert!(startup::require_dependency_rebind_complete(&fixture.0, &value).is_err()); + fs::write(&path, b"malformed").unwrap(); + assert!(startup::require_dependency_rebind_complete(&fixture.0, &value).is_err()); + assert_eq!(fs::read(&path).unwrap(), b"malformed"); + } + + #[test] + fn acknowledgement_is_bound_to_the_selected_cleanup_effect() { + let mut value = receipt(); + let calculate = |receipt: &Receipt| { + host_relay::cleanup_effect(receipt, "fixture-boot", false, &json!(["inventory"])) + .unwrap() + }; + let effect = calculate(&value); + value.relay_cleanup.as_mut().unwrap().effect = effect; + acknowledged_effect(&value, calculate(&value)).unwrap(); + let mut changed = value.clone(); + changed.relay_cleanup.as_mut().unwrap().effect[0] ^= 1; + assert!(acknowledged_effect(&changed, calculate(&changed)).is_err()); + changed = value.clone(); + changed.resources.get_mut("container:web").unwrap().id = Some("9".repeat(64)); + assert!(acknowledged_effect(&changed, calculate(&changed)).is_err()); + let wrong_inventory = + host_relay::cleanup_effect(&value, "fixture-boot", false, &json!(["other-inventory"])) + .unwrap(); + assert!(acknowledged_effect(&value, wrong_inventory).is_err()); + } +} diff --git a/packages/runtime-core/src/provider/graph/bridges/cleanup.rs b/packages/runtime-core/src/provider/graph/bridges/cleanup.rs index 8d918fe5f..f362922d1 100644 --- a/packages/runtime-core/src/provider/graph/bridges/cleanup.rs +++ b/packages/runtime-core/src/provider/graph/bridges/cleanup.rs @@ -11,12 +11,21 @@ pub(crate) struct Selection { previous_boot: Option, #[serde(default, skip_serializing_if = "Option::is_none")] predecessor_owner: Option, + #[serde(default, skip_serializing_if = "Option::is_none")] + predecessor_absence: Option, run: String, plan: String, capacity: u8, serial: u64, selected: BTreeMap, } +#[derive(Clone, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct AbsentPredecessor { + selection_sha256: String, + previous_boot: String, + control_root: PathBuf, +} #[derive(Clone, PartialEq, Eq, Serialize, Deserialize)] #[serde(deny_unknown_fields)] @@ -42,10 +51,39 @@ fn strict_store(candidate: &Candidate, engine: &Engine<'_>) -> Result, +) -> Result<(), CandidateError> { + if !strict_store(candidate, engine)?.slots.is_empty() { + return Err(invalid()); + } + Ok(()) +} + fn validate_selection( engine: &Engine<'_>, receipt: &Receipt, selection: &Selection, +) -> Result<(), CandidateError> { + validate_selection_recovery(engine, receipt, selection, None) +} +fn validate_selection_recovery( + engine: &Engine<'_>, + receipt: &Receipt, + selection: &Selection, + host_pin: Option<&super::super::host_pin_recovery::Witness>, +) -> Result<(), CandidateError> { + validate_selection_with_absence(engine, receipt, selection, host_pin, None) +} +fn validate_selection_with_absence( + engine: &Engine<'_>, + receipt: &Receipt, + selection: &Selection, + host_pin: Option<&super::super::host_pin_recovery::Witness>, + absence: Option<&super::super::absent_publication_cleanup::Selection>, ) -> Result<(), CandidateError> { let capacity = engine.guest().bridge_intent().map_or(0, |v| v.slots); if selection.version != 1 @@ -67,11 +105,30 @@ fn validate_selection( if predecessor_owner( receipt, selection.previous_boot.as_deref().ok_or_else(invalid)?, + host_pin, )? != *expected { return Err(invalid()); } } + match (&selection.predecessor_absence, absence) { + (None, None) => {} + (None, Some(witness)) + if !selection.selected.is_empty() + && selection.predecessor_owner.is_none() + && selection.previous_boot.as_deref() == Some(witness.previous_boot()) + && witness.matches_graph(receipt) => {} + (Some(recorded), Some(witness)) + if recorded.selection_sha256 == witness.digest()? + && recorded.previous_boot == witness.previous_boot() + && selection.previous_boot.as_deref() == Some(witness.previous_boot()) + && recorded.control_root == witness.control_root() + && witness.matches_graph(receipt) + && selection.selected.is_empty() + && selection.predecessor_owner.is_none() + && pinned_predecessor_eligible(receipt) => {} + _ => return Err(invalid()), + } Ok(()) } @@ -83,7 +140,9 @@ fn validate_bindings( if let Some(previous) = &selection.previous_boot { if previous == &selection.boot || previous_boot != Some(previous.as_str()) - || (selection.selected.is_empty() && selection.predecessor_owner.is_none()) + || (selection.selected.is_empty() + && selection.predecessor_owner.is_none() + && selection.predecessor_absence.is_none()) { return Err(invalid()); } @@ -96,6 +155,23 @@ fn validate_bindings( }) { return Err(invalid()); } + if selection + .predecessor_absence + .as_ref() + .is_some_and(|absence| { + !hex(&absence.selection_sha256, 64) + || absence.previous_boot != selection.previous_boot.as_deref().unwrap_or_default() + || receipt + .relay_startup + .as_ref() + .is_none_or(|startup| startup.control_root != absence.control_root) + || selection.predecessor_owner.is_some() + || !selection.selected.is_empty() + || !pinned_predecessor_eligible(receipt) + }) + { + return Err(invalid()); + } let store = Store { version: 1, owner: selection.owner.clone(), @@ -144,6 +220,7 @@ pub(crate) fn capture( boot: engine.guest().boot_id().into(), previous_boot: None, predecessor_owner: None, + predecessor_absence: None, run: receipt.run.clone(), plan: receipt.plan_id.clone(), capacity: engine.guest().bridge_intent().map_or(0, |v| v.slots), @@ -167,6 +244,97 @@ pub(crate) fn capture( Ok(selection) } +// A completed guest relay has no live process/socket identity to capture. Retire +// only that exact current-boot reservation through the ordinary per-slot stop +// path before enrolling graph cleanup; surviving relays still require capture. +// A partial stop remains in the registry and can be resumed on the next call, +// but a missing allocation is sufficient only after the host committed stopped. +pub(crate) fn retire_exited_for_cleanup( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: &Receipt, +) -> Result<(), CandidateError> { + let mut store = strict_store(candidate, engine)?; + let selected = retirement_targets(&store, receipt, engine.guest().boot_id(), |slot, a| { + if a.phase != "running" { + match relay::retirement_absence(engine, slot, a)? { + "absent" => return Ok("absent".into()), + "present" => {} + _ => return Err(invalid()), + } + } + relay::operate(engine, slot, a, "inspect-retirement", None) + })?; + for (slot, expected) in selected { + require_exact_retirement_target(&store, slot, &expected)?; + stop_slot(candidate, engine, &mut store, slot)?; + store.slots.remove(&slot); + save(candidate, &store)?; + } + Ok(()) +} + +fn require_exact_retirement_target( + store: &Store, + slot: u8, + expected: &Assignment, +) -> Result<(), CandidateError> { + if store.slots.get(&slot) == Some(expected) { + Ok(()) + } else { + Err(error( + "bridge_reservation_changed", + "The selected reservation changed before cleanup; nothing was released.", + )) + } +} + +fn retirement_targets( + store: &Store, + receipt: &Receipt, + boot: &str, + mut observe: impl FnMut(u8, &Assignment) -> Result, +) -> Result, CandidateError> { + let mut selected = Vec::new(); + for (slot, a) in store.slots.iter().filter(|(_, a)| a.run == receipt.run) { + if !matches!(a.phase.as_str(), "running" | "stopping" | "stopped") { + continue; + } + // Legacy raw relays retain their existing live-capture path. This + // recovery applies only to launch-fenced reservations. + if a.relay + .as_ref() + .is_none_or(|relay| relay.transport != relay::Transport::ReservationV1) + { + continue; + } + if a.boot_id != boot + || !receipt + .resources + .get(&format!("container:{}", a.service)) + .is_some_and(|r| { + r.kind == Kind::Container + && r.key == a.service + && r.id.as_deref() == Some(a.container_id.as_str()) + }) + || !receipt + .resources + .values() + .any(|r| r.kind == Kind::Network && r.id.as_deref() == Some(a.network_id.as_str())) + { + return Err(invalid()); + } + match (a.phase.as_str(), observe(*slot, a)?.as_str()) { + (_, "exited") | ("stopped", "absent") => { + selected.push((*slot, a.clone())); + } + ("running", "running") => {} + _ => return Err(invalid()), + } + } + Ok(selected) +} + // An empty ingress registry does not imply an empty host-dependency graph. // Completed dependency startup may use the same exact dead-publication proof; // pending startup never obtains cleanup authority from an empty registry. @@ -183,17 +351,39 @@ fn pinned_predecessor_eligible(receipt: &Receipt) -> bool { }) } -fn predecessor_owner(receipt: &Receipt, previous: &str) -> Result { +fn predecessor_owner( + receipt: &Receipt, + previous: &str, + host_pin: Option<&super::super::host_pin_recovery::Witness>, +) -> Result { let startup = receipt .relay_startup .as_ref() .filter(|_| pinned_predecessor_eligible(receipt)) .ok_or_else(invalid)?; - let pin = crate::provider::relay_owner::publication::PinnedEndpoint::load( - &startup.control_root, - super::super::host_relay::context(&receipt.owner, previous)?, - )?; - pin.verify_dead()?; + let context = super::super::host_relay::context(&receipt.owner, previous)?; + let pin = if let Some(selected) = host_pin { + if !selected.matches_graph(receipt) || selected.control_root() != startup.control_root { + return Err(invalid()); + } + let pin = crate::provider::relay_owner::publication::PinnedEndpoint::load_legacy_recovery( + &startup.control_root, + context, + selected.rebind(), + selected.host_boot_micros(), + )?; + if &pin.legacy_summary() != selected.control() { + return Err(invalid()); + } + pin + } else { + let pin = crate::provider::relay_owner::publication::PinnedEndpoint::load( + &startup.control_root, + context, + )?; + pin.verify_dead()?; + pin + }; Ok(pin .fingerprint() .iter() @@ -208,6 +398,7 @@ pub(crate) fn capture_previous_boot( engine: &Engine<'_>, receipt: &Receipt, previous: &str, + host_pin: Option<&super::super::host_pin_recovery::Witness>, ) -> Result { let store = strict_store(candidate, engine)?; let mut selection = Selection { @@ -216,6 +407,7 @@ pub(crate) fn capture_previous_boot( boot: engine.guest().boot_id().into(), previous_boot: Some(previous.into()), predecessor_owner: None, + predecessor_absence: None, run: receipt.run.clone(), plan: receipt.plan_id.clone(), capacity: engine.guest().bridge_intent().map_or(0, |v| v.slots), @@ -248,6 +440,7 @@ pub(crate) fn capture_previous_boot( || prior.boot != previous || prior.previous_boot.is_some() || prior.predecessor_owner.is_some() + || prior.predecessor_absence.is_some() || prior.run != selection.run || prior.plan != selection.plan || prior.capacity != selection.capacity @@ -271,22 +464,123 @@ pub(crate) fn capture_previous_boot( }) .collect(); } else { - selection.predecessor_owner = Some(predecessor_owner(receipt, previous)?); + selection.predecessor_owner = Some(predecessor_owner(receipt, previous, host_pin)?); } } - validate_selection(engine, receipt, &selection)?; + validate_selection_recovery(engine, receipt, &selection, host_pin)?; + Ok(selection) +} + +/// A missing relay-control root has no Pin fingerprint. Select a distinct, +/// witness-bound predecessor only with the exact current bridge registry. +pub(crate) fn capture_previous_boot_absence( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: &Receipt, + witness: &super::super::absent_publication_cleanup::Selection, +) -> Result { + let store = strict_store(candidate, engine)?; + let mut selection = Selection { + version: 1, + owner: engine.guest().incarnation().into(), + boot: engine.guest().boot_id().into(), + previous_boot: Some(witness.previous_boot().into()), + predecessor_owner: None, + predecessor_absence: None, + run: receipt.run.clone(), + plan: receipt.plan_id.clone(), + capacity: engine + .guest() + .bridge_intent() + .map_or(0, |value| value.slots), + serial: store.next_launch_serial, + selected: store + .slots + .into_iter() + .filter(|(_, assignment)| assignment.run == receipt.run) + .map(|(slot, assignment)| { + ( + slot, + Selected { + assignment, + helper: None, + }, + ) + }) + .collect(), + }; + if selection.selected.is_empty() { + selection.predecessor_absence = Some(AbsentPredecessor { + selection_sha256: witness.digest()?, + previous_boot: witness.previous_boot().into(), + control_root: witness.control_root().to_path_buf(), + }); + } + validate_selection_with_absence(engine, receipt, &selection, None, Some(witness))?; Ok(selection) } +/// Selection binds only this run's assignments. Independent sibling +/// allocations may advance the global serial without changing its authority. +pub(crate) fn scoped_absence_projection( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: &Receipt, + previous_boot: &str, +) -> Result { + let store = strict_store(candidate, engine)?; + let selected = store + .slots + .iter() + .filter(|(_, assignment)| assignment.run == receipt.run) + .collect::>(); + Ok(json!({ + "owner": engine.guest().incarnation(), + "current_boot": engine.guest().boot_id(), + "previous_boot": previous_boot, + "run": receipt.run, + "plan": receipt.plan_id, + "capacity": engine.guest().bridge_intent().map_or(0, |value| value.slots), + "selected": selected, + })) +} + +pub(crate) fn selection_absence_projection(selection: &Selection) -> Result { + Ok(json!({ + "owner": selection.owner, + "current_boot": selection.boot, + "previous_boot": selection.previous_boot, + "run": selection.run, + "plan": selection.plan, + "capacity": selection.capacity, + "selected": selection.selected.iter() + .map(|(slot, selected)| (slot, &selected.assignment)) + .collect::>(), + })) +} + +pub(crate) fn verify_remaining_absence( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: &Receipt, + selection: &Selection, + witness: &super::super::absent_publication_cleanup::Selection, +) -> Result<(), CandidateError> { + validate_selection_with_absence(engine, receipt, selection, None, Some(witness))?; + let store = strict_store(candidate, engine)?; + remaining_matches(&store, receipt, selection) +} + /// Pending recovery may resume only the exact remaining reservations. Released /// slots may be absent; a replacement or new reservation is never adopted. -pub(crate) fn verify_remaining( +pub(crate) fn verify_remaining_recovery( candidate: &Candidate, engine: &Engine<'_>, receipt: &Receipt, selection: &Selection, + host_pin: Option<&super::super::host_pin_recovery::Witness>, ) -> Result<(), CandidateError> { - validate_selection(engine, receipt, selection)?; + validate_selection_recovery(engine, receipt, selection, host_pin)?; if selection.previous_boot.is_none() { return Ok(()); } @@ -311,6 +605,7 @@ pub(crate) fn verify_live_remaining( fn live_remaining_matches(store: &Store, selection: &Selection) -> Result<(), CandidateError> { if selection.previous_boot.is_some() || selection.predecessor_owner.is_some() + || selection.predecessor_absence.is_some() || store.owner != selection.owner || store.next_launch_serial < selection.serial { @@ -381,6 +676,7 @@ fn prior_generation( && prior.boot == current.previous_boot.as_deref().unwrap_or_default() && prior.previous_boot.is_none() && prior.predecessor_owner.is_none() + && prior.predecessor_absence.is_none() && prior.run == current.run && prior.plan == current.plan && prior.capacity == current.capacity @@ -403,6 +699,7 @@ fn prior_generation( && prior.boot == current.previous_boot.as_deref().unwrap_or_default() && prior.previous_boot.is_none() && prior.predecessor_owner.is_none() + && prior.predecessor_absence.is_none() && prior.run == current.run && prior.plan == current.plan && prior.capacity == current.capacity @@ -529,19 +826,38 @@ pub(crate) fn read( engine: &Engine<'_>, receipt: &Receipt, root: &std::path::Path, +) -> Result { + read_recovery(engine, receipt, root, None) +} +pub(crate) fn read_recovery( + engine: &Engine<'_>, + receipt: &Receipt, + root: &std::path::Path, + host_pin: Option<&super::super::host_pin_recovery::Witness>, ) -> Result { let selection = state::read_bounded(&selection_path(root)?, 65536)?; - validate_selection(engine, receipt, &selection)?; + validate_selection_recovery(engine, receipt, &selection, host_pin)?; + Ok(selection) +} +pub(crate) fn read_recovery_absence( + engine: &Engine<'_>, + receipt: &Receipt, + root: &std::path::Path, + witness: &super::super::absent_publication_cleanup::Selection, +) -> Result { + let selection = state::read_bounded(&selection_path(root)?, 65536)?; + validate_selection_with_absence(engine, receipt, &selection, None, Some(witness))?; Ok(selection) } -pub(crate) fn verify( +pub(crate) fn verify_recovery( candidate: &Candidate, engine: &Engine<'_>, receipt: &Receipt, selection: &Selection, + host_pin: Option<&super::super::host_pin_recovery::Witness>, ) -> Result<(), CandidateError> { - validate_selection(engine, receipt, selection)?; + validate_selection_recovery(engine, receipt, selection, host_pin)?; let store = strict_store(candidate, engine)?; if store.next_launch_serial < selection.serial || store.slots.values().any(|a| a.run == receipt.run) @@ -556,6 +872,40 @@ pub(crate) fn verify( engine.guest().verify()?; Ok(()) } +pub(crate) fn verify_recovery_absence( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: &Receipt, + selection: &Selection, + witness: &super::super::absent_publication_cleanup::Selection, +) -> Result<(), CandidateError> { + validate_selection_with_absence(engine, receipt, selection, None, Some(witness))?; + let store = strict_store(candidate, engine)?; + if store.next_launch_serial < selection.serial + || store + .slots + .values() + .any(|assignment| assignment.run == receipt.run) + { + return Err(invalid()); + } + for (slot, selected) in &selection.selected { + if let Some(helper) = &selected.helper { + relay::verify_cleanup(engine, *slot, &selected.assignment, helper)?; + } + } + engine.guest().verify() +} + +#[cfg(test)] +pub(crate) fn verify( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: &Receipt, + selection: &Selection, +) -> Result<(), CandidateError> { + verify_recovery(candidate, engine, receipt, selection, None) +} #[cfg(test)] mod tests { @@ -570,6 +920,7 @@ mod tests { boot: "boot".into(), previous_boot: None, predecessor_owner: None, + predecessor_absence: None, run: "b".repeat(32), plan: "c".repeat(64), capacity: 0, @@ -622,6 +973,169 @@ mod tests { (selected, store) } + fn receipt_for_assignment(selected: &Selection, assignment: &Assignment) -> Receipt { + serde_json::from_value(json!({ + "version":1, + "run":selected.run, + "owner":selected.owner, + "namespace":"3".repeat(64), + "plan_id":selected.plan, + "phase":"ready-observed", + "readiness":{"web":"healthy"}, + "resources":{ + "container:web":{ + "kind":"container","key":"web","name":"owned", + "id":assignment.container_id,"image":null,"phase":"started" + }, + "network:default":{ + "kind":"network","key":"default","name":"owned", + "id":assignment.network_id,"image":null,"phase":"created" + } + } + })) + .unwrap() + } + + #[test] + fn exited_current_run_relay_is_selected_before_live_bridge_capture() { + let (selected, mut store) = live_selection(); + let assignment = store.slots[&0].clone(); + let receipt = receipt_for_assignment(&selected, &assignment); + let mut foreign = assignment.clone(); + foreign.run = "9".repeat(32); + foreign.service = "foreign".into(); + foreign.reservation = "8".repeat(32); + foreign.relay.as_mut().unwrap().launch_serial = 2; + store.next_launch_serial = 2; + store.slots.insert(1, foreign.clone()); + let mut seen = Vec::new(); + let targets = retirement_targets(&store, &receipt, &selected.boot, |slot, _| { + seen.push(slot); + Ok("exited".into()) + }) + .unwrap(); + assert_eq!(seen, vec![0]); + assert_eq!(targets, vec![(0, assignment.clone())]); + assert_eq!(store.slots[&1], foreign); + assert!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { + Ok("running".into()) + }) + .unwrap() + .is_empty() + ); + // An interrupted exact stop still needs fresh exit or fenced absence. + store.slots.get_mut(&0).unwrap().phase = "stopping".into(); + assert_eq!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { + Ok("exited".into()) + }) + .unwrap()[0] + .0, + 0 + ); + store.slots.get_mut(&0).unwrap().phase = "stopped".into(); + assert_eq!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { + Ok("absent".into()) + }) + .unwrap()[0] + .0, + 0 + ); + store.slots.get_mut(&0).unwrap().phase = "stopping".into(); + assert!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { + Ok("absent".into()) + }) + .is_err() + ); + for phase in ["stopping", "stopped"] { + store.slots.get_mut(&0).unwrap().phase = phase.into(); + assert!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { + Ok("running".into()) + }) + .is_err() + ); + assert!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { Err(invalid()) }) + .is_err() + ); + } + } + + #[test] + fn exited_relay_retirement_refuses_changed_identity_and_uncertain_observation() { + for mutation in 0..4 { + let (selected, mut store) = live_selection(); + let receipt = receipt_for_assignment(&selected, &store.slots[&0]); + let assignment = store.slots.get_mut(&0).unwrap(); + match mutation { + 0 => assignment.boot_id = "22222222-2222-2222-2222-222222222222".into(), + 1 => assignment.container_id = "8".repeat(64), + 2 => assignment.network_id = "8".repeat(64), + 3 => assignment.service = "replaced".into(), + _ => unreachable!(), + } + assert!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { + Ok("exited".into()) + }) + .is_err(), + "mutation {mutation}" + ); + } + let (selected, store) = live_selection(); + let receipt = receipt_for_assignment(&selected, &store.slots[&0]); + let (_, mut legacy) = live_selection(); + legacy + .slots + .get_mut(&0) + .unwrap() + .relay + .as_mut() + .unwrap() + .transport = relay::Transport::Raw; + assert!( + retirement_targets(&legacy, &receipt, &selected.boot, |_, _| { + panic!("raw relay remains on its existing cleanup path") + }) + .unwrap() + .is_empty() + ); + assert!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { Err(invalid()) }) + .is_err() + ); + assert!( + retirement_targets(&store, &receipt, &selected.boot, |_, _| { + Ok("unknown".into()) + }) + .is_err() + ); + let expected = &store.slots[&0]; + assert!(require_exact_retirement_target(&store, 0, expected).is_ok()); + for mutation in 0..4 { + let (_, mut replaced) = live_selection(); + let current = replaced.slots.get_mut(&0).unwrap(); + match mutation { + 0 => current.reservation = "8".repeat(32), + 1 => current.run = "8".repeat(32), + 2 => current.generation = "8".repeat(64), + 3 => current.relay.as_mut().unwrap().launch_serial += 1, + _ => unreachable!(), + } + assert_eq!( + require_exact_retirement_target(&replaced, 0, expected) + .unwrap_err() + .code, + "bridge_reservation_changed", + "mutation {mutation}" + ); + } + } + #[test] fn same_boot_remaining_allows_cleanup_progress_without_selecting_siblings() { let (selected, mut store) = live_selection(); diff --git a/packages/runtime-core/src/provider/graph/cleanup_enrollment.rs b/packages/runtime-core/src/provider/graph/cleanup_enrollment.rs index 6c187f2e0..4bc12667c 100644 --- a/packages/runtime-core/src/provider/graph/cleanup_enrollment.rs +++ b/packages/runtime-core/src/provider/graph/cleanup_enrollment.rs @@ -116,9 +116,21 @@ pub(super) fn ordinary_mutation(root: &Path, receipt: &Receipt) -> Result<(), Ca Ok(()) } pub(super) fn retention(root: &Path, receipt: &Receipt) -> Result<(), CandidateError> { + #[cfg(target_os = "macos")] + if super::absent_publication_cleanup::retained_current(root, receipt)? { + return Ok(()); + } + #[cfg(target_os = "macos")] + if super::dead_owner_cleanup::current_completed_precedence(root, receipt)? { + if !super::dead_owner_cleanup::retained(root, receipt)? { + return Err(refused()); + } + return Ok(()); + } #[cfg(target_os = "macos")] if super::live_owner_cleanup::retained(root, receipt)? || super::dead_owner_cleanup::retained(root, receipt)? + || super::absent_publication_cleanup::retained(root, receipt)? { return Ok(()); } diff --git a/packages/runtime-core/src/provider/graph/config.rs b/packages/runtime-core/src/provider/graph/config.rs index df8b3747d..c84f57fcb 100644 --- a/packages/runtime-core/src/provider/graph/config.rs +++ b/packages/runtime-core/src/provider/graph/config.rs @@ -155,7 +155,10 @@ pub(super) fn prepare_delivery( "Dependency caches require an explicitly verified source publication.", ) })?; - let scope = dependency_cache::scope(&plan.source)?; + let scope = source.binding.cache_scope.as_ref().map_or_else( + || dependency_cache::scope(&plan.source), + |scope| scope.scope(&plan.source), + )?; dependency_cache::resolve( plan, source.current_manifest.as_ref().unwrap_or(&source.manifest), diff --git a/packages/runtime-core/src/provider/graph/dead_owner_cleanup.rs b/packages/runtime-core/src/provider/graph/dead_owner_cleanup.rs index f2d1a8135..c5bd84e25 100644 --- a/packages/runtime-core/src/provider/graph/dead_owner_cleanup.rs +++ b/packages/runtime-core/src/provider/graph/dead_owner_cleanup.rs @@ -4,6 +4,7 @@ use super::*; use crate::provider::{identity, lifecycle, state::Owner}; use sha2::{Digest, Sha256}; +use std::os::unix::fs::MetadataExt; const FILE: &str = "dead-owner-cleanup.json"; #[derive(Serialize, Deserialize)] #[serde(deny_unknown_fields)] @@ -22,6 +23,19 @@ struct Intent { #[serde(default, skip_serializing_if = "Option::is_none")] one_off_sha256: Option, } +/// Dispatch only an exact completed recovery generation; history is inert. +pub(super) fn current_completion( + root: &std::path::Path, + receipt: &Receipt, +) -> Result { + if !exists(&root.join(FILE))? { + return Ok(false); + } + let intent: Intent = state::read(&root.join(FILE))?; + let complete = selected(receipt)?; + Ok(intent.complete_sha256.as_deref() == Some(complete.as_str())) +} + fn refused() -> CandidateError { error( "graph_dead_owner_recovery", @@ -37,7 +51,7 @@ fn selected(receipt: &Receipt) -> Result { &serde_json::to_vec_pretty(receipt).map_err(|_| refused())?, )) } -fn immutable(receipt: &Receipt) -> Result { +pub(super) fn immutable(receipt: &Receipt) -> Result { let mut value = serde_json::to_value(receipt).map_err(|_| refused())?; value.as_object_mut().ok_or_else(refused)?.remove("phase"); for group in ["resources", "probes"] { @@ -90,8 +104,8 @@ fn retain_interrupted_write(root: &std::path::Path) -> Result<(), CandidateError Ok(()) } -/// Supersede only a completed prior recovery whose exact stopped receipt is in -/// bounded history and whose containers were replaced. An atomic rename keeps +/// Supersede only a completed prior recovery whose historical generation is +/// independently validated against bounded history. An atomic rename keeps /// the old value-free proof if selection or publication is interrupted. fn archive_completed_prior( root: &std::path::Path, @@ -132,21 +146,25 @@ fn archive_completed_prior( return Err(refused()); } let complete = prior.complete_sha256.as_deref().ok_or_else(refused)?; - let stopped = - restore_history::completed_for_recovery(root, current, complete)?.ok_or_else(refused)?; - validate( - &prior, - &stopped, - &prior.original_sha256, - &prior.owner_sha256, - )?; - if !stopped.resources.iter().any(|(key, resource)| { - resource.kind == Kind::Container - && resource.id.as_deref().is_some_and(|old| { - current.resources.get(key).and_then(|now| now.id.as_deref()) != Some(old) - }) - }) { - return Err(refused()); + if let Some(stopped) = restore_history::completed_for_recovery(root, current, complete)? { + validate( + &prior, + &stopped, + &prior.original_sha256, + &prior.owner_sha256, + )?; + if !stopped.resources.iter().any(|(key, resource)| { + resource.kind == Kind::Container + && resource.id.as_deref().is_some_and(|old| { + current.resources.get(key).and_then(|now| now.id.as_deref()) != Some(old) + }) + }) { + return Err(refused()); + } + } else { + // Bounded history may evict the exact completion. Its existing strict + // truncated-history proof must still establish a superseded generation. + require_historical_recovery(root, current)?; } let mut archived = 0; for entry in fs::read_dir(root).map_err(state::io)? { @@ -172,8 +190,9 @@ fn archive_completed_prior( } /// A completed recovery from an older container generation is diagnostic history, -/// not a pending operation. Require its exact stopped receipt in restore history; -/// this never supplies authority to retire current resources or mutates old proof. +/// not a pending operation. Prefer its exact stopped receipt; after history +/// truncation, validate the superseded original and a newer stopped generation. +/// Neither path supplies current cleanup authority or mutates the old proof. pub(super) fn require_historical_recovery( root: &std::path::Path, current: &Receipt, @@ -195,15 +214,29 @@ pub(super) fn require_historical_recovery( { return Err(refused()); } - let stopped = - restore_history::completed_for_recovery(root, current, complete)?.ok_or_else(refused)?; + let stopped = restore_history::completed_for_recovery(root, current, complete)?; + let historical = stopped.as_ref().unwrap_or(&prior.original); validate( &prior, - &stopped, + historical, &prior.original_sha256, &prior.owner_sha256, )?; - if !stopped.resources.iter().any(|(key, old)| { + if stopped.is_none() + && (prior.original.version != current.version + || prior.original.run != current.run + || prior.original.owner != current.owner + || prior.original.namespace != current.namespace + || prior.original.plan_id != current.plan_id + || !restore_history::confirms_truncated_newer_generation( + root, + current, + &prior.original, + )?) + { + return Err(refused()); + } + if !historical.resources.iter().any(|(key, old)| { old.kind == Kind::Container && old.id.as_ref().is_some_and(|id| { current @@ -267,6 +300,152 @@ pub fn recover_cleanup( execute(candidate, run, expected) } +/// A current completed dead-owner cleanup outranks only a *validated historical* +/// live-owner sidecar for retention or publisher retirement. This selection +/// does not replace either operation's exact receipt and inventory checks. +pub(super) fn current_completed_precedence( + root: &std::path::Path, + receipt: &Receipt, +) -> Result { + if !current_completion(root, receipt)? { + return Ok(false); + } + super::live_owner_cleanup::require_historical_recovery(root, receipt)?; + Ok(true) +} + +/// Holds the retired publisher and provider cleanup leases while an exact +/// previous-boot shared HTTPS owner is archived. Neither pathname absence nor +/// a merely stopped graph grants this authority. +pub(in crate::provider) struct HttpsArchiveGuard<'a> { + retired: foreground::transport::Retired, + engine: Engine<'a>, + run: String, + owner: String, + namespace: String, + plan: String, + old_boot: String, +} + +impl<'a> HttpsArchiveGuard<'a> { + pub(in crate::provider) fn acquire( + candidate: &'a Candidate, + run: &str, + owner: &str, + namespace: &str, + plan: &str, + old_boot: &str, + ) -> Result { + // Admission is held by the caller. The publisher lock precedes the + // provider lease, matching the existing retained-data recovery order. + let retired = + foreground::transport::Retired::acquire(candidate, run)?.ok_or_else(refused)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let guard = Self { + retired, + engine, + run: run.into(), + owner: owner.into(), + namespace: namespace.into(), + plan: plan.into(), + old_boot: old_boot.into(), + }; + guard.verify(candidate)?; + Ok(guard) + } + + pub(in crate::provider) fn current_boot(&self) -> &str { + self.engine.guest().boot_id() + } + + /// Recheck under both held leases immediately before each archive effect. + pub(in crate::provider) fn verify(&self, candidate: &Candidate) -> Result<(), CandidateError> { + self.retired.verify()?; + let (receipt, root) = load(candidate, &self.engine, &self.run)?; + no_pending(&root)?; + if receipt.phase != "stopped-data-retained" + || receipt.owner != self.owner + || receipt.namespace != self.namespace + || receipt.plan_id != self.plan + || !current_completed_precedence(&root, &receipt)? + || !retained(&root, &receipt)? + { + return Err(refused()); + } + initializer_cache::require_resolved(&receipt)?; + let intent: Intent = state::read(&root.join(FILE))?; + let complete = selected(&receipt)?; + validate( + &intent, + &receipt, + &intent.original_sha256, + &intent.owner_sha256, + )?; + let pool = Owner::load(candidate)?; + if intent.complete_sha256.as_deref() != Some(complete.as_str()) + || intent.old_boot != self.old_boot + || intent.new_boot.as_deref() != Some(self.engine.guest().boot_id()) + || pool.previous_guest_boot_id.as_deref() != Some(self.old_boot.as_str()) + || pool.guest_boot_id.as_deref() != Some(self.engine.guest().boot_id()) + || pool.token != receipt.owner + || intent.one_off_sha256.is_some() + { + return Err(refused()); + } + let witness = super::host_pin_recovery::load_witness(candidate, &self.run)?.filter(|w| { + w.graph_sha256() == intent.original_sha256 + && w.publisher_sha256() == intent.owner_sha256 + && w.matches_graph(&intent.original) + }); + self.retired.verify_recovery_with_rebind( + candidate, + &self.run, + &intent.owner_sha256, + &complete, + witness.as_ref().map(|w| w.rebind()), + )?; + for resource in receipt.resources.values() { + let observed = inspect_resource(&self.engine, &receipt, resource)?; + if (resource.kind == Kind::Volume && observed.is_none()) + || (resource.kind != Kind::Volume + && (resource.phase != "absent" || observed.is_some())) + { + return Err(refused()); + } + } + let environment = environment::cleanup_inventory(candidate, &self.engine, &receipt, &root)?; + if intent.environment.as_ref() + != Some(&serde_json::to_value(&environment).map_err(|_| refused())?) + { + return Err(refused()); + } + let bridges = intent.bridges.as_ref().ok_or_else(refused)?; + bridges::cleanup::verify_recovery_file(&root, bridges, intent.prior_bridges.as_ref())?; + let active = bridges::inspect_bridges_using(candidate, &self.engine, &self.run)?; + if active["slots"] + .as_object() + .is_none_or(|slots| !slots.is_empty()) + { + return Err(refused()); + } + host_relay::inspect_cleanup_recovery( + candidate, + &self.engine, + &receipt, + false, + &environment, + bridges, + witness.as_ref(), + )?; + super::super::publication::require_no_claims_locked(candidate, &receipt.owner)?; + let authority = super::super::hostname_authority::managed::inspect(candidate)?; + if authority["authority"]["present"] != false { + return Err(refused()); + } + self.engine.guest().verify() + } +} + /// The completed dead-owner cleanup receipt, not missing endpoint names, grants /// a separate explicit publisher retirement. This never removes graph volumes. pub fn retire_recovered_publisher( @@ -277,11 +456,28 @@ pub fn retire_recovered_publisher( if !hex(expected_owner, 32) { return Err(refused()); } - if let Some(result) = super::live_owner_cleanup::retire(candidate, run, expected_owner)? { + if let Some(result) = + super::acknowledged_publisher::confirm_retired(candidate, run, expected_owner)? + { return Ok(result); } + let dead_owner_current = { + let engine = Engine::connect_cleanup_wait(candidate)?; + let (receipt, root) = load(candidate, &engine, run)?; + current_completed_precedence(&root, &receipt)? + }; + if !dead_owner_current { + if let Some(result) = super::live_owner_cleanup::retire(candidate, run, expected_owner)? { + return Ok(result); + } + } let engine = Engine::connect_cleanup_wait(candidate)?; let (receipt, root) = load(candidate, &engine, run)?; + // Dispatch crossed a cleanup lease boundary. Revalidate the current proof + // and historical sidecar under the lease that protects retirement effects. + if dead_owner_current && !current_completed_precedence(&root, &receipt)? { + return Err(refused()); + } no_pending(&root)?; if receipt.phase != "stopped-data-retained" || receipt.owner != expected_owner @@ -322,18 +518,98 @@ pub fn retire_recovered_publisher( } let bridges = intent.bridges.as_ref().ok_or_else(refused)?; bridges::cleanup::verify_recovery_file(&root, bridges, intent.prior_bridges.as_ref())?; + let legacy = super::host_pin_recovery::load_witness(candidate, run)?.filter(|witness| { + witness.graph_sha256() == intent.original_sha256 + && witness.publisher_sha256() == intent.owner_sha256 + && witness.matches_graph(&intent.original) + }); if intent.new_boot.as_deref() == Some(engine.guest().boot_id()) { - host_relay::inspect_cleanup(candidate, &engine, &receipt, false, &environment, bridges)?; - foreground::retire_publisher_path(candidate, run, &intent.owner_sha256, &complete)?; + if let Some(witness) = legacy { + if let Some(retired) = foreground::transport::Retired::acquire(candidate, run)? { + let control_root = witness.control_root().join("relay-control"); + let control_lock = state::Lock::acquire_existing(&control_root)?; + host_relay::inspect_cleanup_recovery( + candidate, + &engine, + &receipt, + false, + &environment, + bridges, + Some(&witness), + )?; + let lock_path = fs::symlink_metadata(control_root.join("operation.lock")) + .map_err(|_| refused())?; + if (lock_path.dev(), lock_path.ino()) != control_lock.identity()? { + return Err(refused()); + } + retired.verify_recovery_with_rebind( + candidate, + run, + &intent.owner_sha256, + &complete, + Some(witness.rebind()), + )?; + } else { + let foreground_root = foreground::transport::root(candidate, run)?; + let foreground_lock = state::Lock::acquire_existing(&foreground_root)?; + let guard = super::host_pin_recovery::acquire_for_cleanup( + candidate, + run, + &engine, + super::host_pin_recovery::CleanupProof { + original: &intent.original, + sha256: &intent.original_sha256, + current_is_original: false, + allow_absent_reservation: true, + publisher_may_be_partial: true, + }, + witness, + )?; + host_relay::inspect_cleanup_recovery( + candidate, + &engine, + &receipt, + false, + &environment, + bridges, + Some(guard.witness()), + )?; + guard.verify_lock()?; + let lock_path = fs::symlink_metadata(foreground_root.join("operation.lock")) + .map_err(|_| refused())?; + if (lock_path.dev(), lock_path.ino()) != foreground_lock.identity()? { + return Err(refused()); + } + foreground::transport::retire_recovered_publisher_locked( + candidate, + run, + &intent.owner_sha256, + &complete, + Some(guard.witness().rebind()), + &foreground_lock, + )?; + } + } else { + host_relay::inspect_cleanup( + candidate, + &engine, + &receipt, + false, + &environment, + bridges, + )?; + foreground::retire_publisher_path(candidate, run, &intent.owner_sha256, &complete)?; + } } else { // A later VM boot has a new bridge registry generation. Recheck the // retained graph and immutable prior proof, then require the publisher // to have been fully retired under the original recovery boot. - foreground::verify_recovered_publisher_retired( + foreground::transport::verify_recovered_publisher_retired_recovery( candidate, run, &intent.owner_sha256, &complete, + legacy.as_ref().map(|witness| witness.rebind()), )?; } Ok(json!({"run":run,"publisher_retired":true,"data_retained":true})) @@ -427,7 +703,13 @@ fn fresh_boot( } fn execute(candidate: &Candidate, run: &str, expected: &str) -> Result { let engine = Engine::connect_cleanup_wait(candidate)?; - let dead = foreground::DeadOwner::acquire(candidate, run)?; + let selected_pin = super::host_pin_recovery::selected_for_old_publisher(candidate, run)?; + let dead = foreground::DeadOwner::acquire_recovery( + candidate, + run, + selected_pin.as_ref().map(|pin| pin.rebind()), + selected_pin.as_ref().map(|pin| pin.host_boot_micros()), + )?; let (receipt, root) = load(candidate, &engine, run)?; if exists(&root.join("one-off-normalization.json"))? { let intent: Intent = state::read(&root.join(FILE))?; @@ -448,6 +730,36 @@ fn execute(candidate: &Candidate, run: &str, expected: &str) -> Result = if exists(&root.join(FILE))? { + Some(state::read(&root.join(FILE))?) + } else { + None + }; + let pin_guard = if let Some(witness) = selected_pin { + let original = existing_intent + .as_ref() + .map_or(&receipt, |intent| &intent.original); + let original_sha = existing_intent + .as_ref() + .map_or(expected, |intent| intent.original_sha256.as_str()); + Some(super::host_pin_recovery::acquire_for_cleanup( + candidate, + run, + &engine, + super::host_pin_recovery::CleanupProof { + original, + sha256: original_sha, + current_is_original: existing_intent.is_none(), + allow_absent_reservation: existing_intent + .as_ref() + .is_some_and(|intent| intent.complete_sha256.is_some()), + publisher_may_be_partial: false, + }, + witness, + )?) + } else { + None + }; let mut intent: Intent = if exists(&root.join(FILE))? { state::read(&root.join(FILE))? } else { @@ -479,8 +791,13 @@ fn execute(candidate: &Candidate, run: &str, expected: &str) -> Result Result Result Result Result Result Result Result (super::super::tests::Fixture, Receipt, Receipt, Intent) { + let (fixture, mut current, stopped, proof) = historical_fixture(); + for generation in 1..=12 { + let mut newer = stopped.clone(); + newer.resources.get_mut("container:init").unwrap().id = + Some(format!("{generation:064x}")); + restore_history::retain(&fixture.0, &newer).unwrap(); + } + current.resources.get_mut("container:init").unwrap().id = Some(format!("{:064x}", 13)); + (fixture, current, stopped, proof) + } + + fn retirement_precedence_fixture() -> (super::super::tests::Fixture, Receipt, Value) { + let fixture = super::super::tests::Fixture::new(); + let root = &fixture.0; + let mut original = partial(); + original.phase = "ready-observed".into(); + original.resources.get_mut("container:init").unwrap().id = Some(format!("{:064x}", 1)); + let mut prior_stopped = original.clone(); + prior_stopped.phase = "stopped-data-retained".into(); + for resource in prior_stopped.resources.values_mut() { + if resource.kind != Kind::Volume { + resource.phase = "absent".into(); + } + } + let live = json!({ + "version":1,"boot":"historical-boot","original":original, + "original_sha256":selected(&original).unwrap(), + "foreground_sha256":"f".repeat(64), + "relay":{"bytes":[],"record_id":[1,1]}, + "environment":[], + "bridges":{"version":1,"owner":original.owner,"boot":"historical-boot", + "run":original.run,"plan":original.plan_id,"capacity":0,"serial":0, + "selected":{}}, + "prior_bridges":null,"listeners_retired":true, + "complete_sha256":selected(&prior_stopped).unwrap() + }); + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + for generation in 1..=12 { + let mut stopped = prior_stopped.clone(); + stopped.resources.get_mut("container:init").unwrap().id = + Some(format!("{generation:064x}")); + restore_history::retain(root, &stopped).unwrap(); + } + let mut current = prior_stopped; + current.resources.get_mut("container:init").unwrap().id = Some(format!("{:064x}", 13)); + let mut dead = intent(¤t); + dead.complete_sha256 = Some(selected(¤t).unwrap()); + state::write(&root.join(FILE), &dead).unwrap(); + state::write(&root.join("state.json"), ¤t).unwrap(); + (fixture, current, live) + } + + #[test] + fn current_dead_owner_retirement_precedes_only_valid_historical_live_proof() { + let (fixture, current, _live) = retirement_precedence_fixture(); + let root = &fixture.0; + let live_before = fs::read(root.join("live-owner-cleanup.json")).unwrap(); + let dead_before = fs::read(root.join(FILE)).unwrap(); + assert!(current_completed_precedence(root, ¤t).unwrap()); + super::super::cleanup_enrollment::retention(root, ¤t).unwrap(); + assert_eq!( + fs::read(root.join("live-owner-cleanup.json")).unwrap(), + live_before + ); + assert_eq!(fs::read(root.join(FILE)).unwrap(), dead_before); + + let mut stale_dead: Intent = state::read(&root.join(FILE)).unwrap(); + stale_dead.complete_sha256 = Some("8".repeat(64)); + state::write(&root.join(FILE), &stale_dead).unwrap(); + assert!(!current_completed_precedence(root, ¤t).unwrap()); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + state::write(&root.join(FILE), &intent(¤t)).unwrap(); + assert!(!current_completed_precedence(root, ¤t).unwrap()); + assert!(super::super::cleanup_enrollment::retention(root, ¤t).is_err()); + } + + #[test] + fn current_dead_owner_retirement_refuses_pending_incomplete_or_foreign_live_proof() { + for fault in [ + "pending", + "incomplete", + "foreign-owner", + "current-live", + "malformed", + ] { + let (fixture, current, mut live) = retirement_precedence_fixture(); + let root = &fixture.0; + match fault { + "pending" => { + fs::write(root.join("live-owner-cleanup.pending"), b"partial proof").unwrap(); + } + "incomplete" => live["complete_sha256"] = Value::Null, + "foreign-owner" => { + live["original"]["owner"] = json!("e".repeat(32)); + let changed: Receipt = + serde_json::from_value(live["original"].clone()).unwrap(); + live["original_sha256"] = json!(selected(&changed).unwrap()); + } + "current-live" => { + let mut same = current.clone(); + same.phase = "ready-observed".into(); + live["original"] = serde_json::to_value(&same).unwrap(); + live["original_sha256"] = json!(selected(&same).unwrap()); + live["complete_sha256"] = json!(selected(¤t).unwrap()); + } + "malformed" => { + fs::write(root.join("live-owner-cleanup.json"), b"invalid proof").unwrap(); + } + _ => unreachable!(), + } + if !matches!(fault, "pending" | "malformed") { + state::write(&root.join("live-owner-cleanup.json"), &live).unwrap(); + } + let live_before = fs::read(root.join("live-owner-cleanup.json")).unwrap(); + let dead_before = fs::read(root.join(FILE)).unwrap(); + assert!( + current_completed_precedence(root, ¤t).is_err(), + "{fault}" + ); + assert!( + super::super::cleanup_enrollment::retention(root, ¤t).is_err(), + "{fault}" + ); + assert_eq!( + fs::read(root.join("live-owner-cleanup.json")).unwrap(), + live_before + ); + assert_eq!(fs::read(root.join(FILE)).unwrap(), dead_before); + } + } + + #[test] + fn repeated_recovery_archives_evicted_completion_without_rewriting_prior_proof() { + let (fixture, current, _, proof) = evicted_historical_fixture(); + let root = &fixture.0; + let source = root.join(FILE); + let before = fs::read(&source).unwrap(); + let metadata = fs::metadata(&source).unwrap(); + let history = fs::read(root.join("restore-history.json")).unwrap(); + let complete = proof.complete_sha256.as_deref().unwrap(); + assert!( + restore_history::completed_for_recovery(root, ¤t, complete) + .unwrap() + .is_none() + ); + + archive_completed_prior( + root, + ¤t, + &selected(¤t).unwrap(), + &"2".repeat(64), + "current-boot", + ) + .unwrap(); + + let archived = root.join(format!("dead-owner-cleanup-retired-{complete}.json")); + let archived_metadata = fs::metadata(&archived).unwrap(); + assert!(!source.exists()); + assert_eq!(fs::read(archived).unwrap(), before); + assert_eq!(archived_metadata.dev(), metadata.dev()); + assert_eq!(archived_metadata.ino(), metadata.ino()); + assert_eq!( + fs::read(root.join("restore-history.json")).unwrap(), + history + ); + } + + #[test] + fn repeated_recovery_refuses_evicted_proof_with_changed_identity_or_generation() { + for fault in [ + "owner", + "run", + "plan_id", + "unchanged-generation", + "pending", + "incomplete", + "mutated-resource-id", + ] { + let (fixture, mut current, _, mut proof) = evicted_historical_fixture(); + let root = &fixture.0; + match fault { + "owner" | "run" | "plan_id" => { + let mut original = serde_json::to_value(&proof.original).unwrap(); + original[fault] = json!("e".repeat(if fault == "plan_id" { 64 } else { 32 })); + proof.original = serde_json::from_value(original).unwrap(); + proof.original_sha256 = selected(&proof.original).unwrap(); + state::write(&root.join(FILE), &proof).unwrap(); + } + "unchanged-generation" => { + current.resources.get_mut("container:init").unwrap().id = + proof.original.resources["container:init"].id.clone(); + } + "pending" => { + fs::write(root.join("dead-owner-cleanup.pending"), b"partial proof").unwrap(); + } + "incomplete" => { + proof.complete_sha256 = None; + state::write(&root.join(FILE), &proof).unwrap(); + } + "mutated-resource-id" => { + proof + .original + .resources + .get_mut("container:init") + .unwrap() + .id = Some("7".repeat(64)); + state::write(&root.join(FILE), &proof).unwrap(); + } + _ => unreachable!(), + } + let before = fs::read(root.join(FILE)).unwrap(); + let history = fs::read(root.join("restore-history.json")).unwrap(); + assert!( + archive_completed_prior( + root, + ¤t, + &selected(¤t).unwrap(), + &"2".repeat(64), + "current-boot", + ) + .is_err(), + "{fault}" + ); + assert_eq!(fs::read(root.join(FILE)).unwrap(), before, "{fault}"); + assert_eq!( + fs::read(root.join("restore-history.json")).unwrap(), + history + ); + assert!( + !root + .join(format!( + "dead-owner-cleanup-retired-{}.json", + proof.complete_sha256.as_deref().unwrap_or("missing") + )) + .exists(), + "{fault}" + ); + } + } + + #[test] + fn evicted_previous_boot_completion_is_read_only_superseded_history() { + let (fixture, current, _, proof) = evicted_historical_fixture(); + let before = fs::read(fixture.0.join(FILE)).unwrap(); + let history_before = fs::read(fixture.0.join("restore-history.json")).unwrap(); + assert!( + restore_history::completed_for_recovery( + &fixture.0, + ¤t, + proof.complete_sha256.as_deref().unwrap() + ) + .unwrap() + .is_none() + ); + require_historical_recovery(&fixture.0, ¤t).unwrap(); + assert!(!current_completion(&fixture.0, ¤t).unwrap()); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), before); + assert_eq!( + fs::read(fixture.0.join("restore-history.json")).unwrap(), + history_before + ); + } + + #[test] + fn truncated_history_does_not_admit_invalid_or_current_recovery_proof() { + for (field, value) in [ + ("version", json!(9)), + ("complete_sha256", Value::Null), + ("complete_sha256", json!("invalid")), + ("original_sha256", json!("8".repeat(64))), + ("owner_sha256", json!("invalid")), + ("new_boot", Value::Null), + ("new_boot", json!("")), + ("old_boot", json!("")), + ("old_boot", json!("successor-boot")), + ("one_off_sha256", json!("8".repeat(64))), + ] { + let (fixture, current, _, proof) = evicted_historical_fixture(); + let mut changed = serde_json::to_value(proof).unwrap(); + changed[field] = value; + state::write(&fixture.0.join(FILE), &changed).unwrap(); + let bytes = fs::read(fixture.0.join(FILE)).unwrap(); + let history = fs::read(fixture.0.join("restore-history.json")).unwrap(); + assert!( + require_historical_recovery(&fixture.0, ¤t).is_err(), + "{field}" + ); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), bytes); + assert_eq!( + fs::read(fixture.0.join("restore-history.json")).unwrap(), + history + ); + } + for field in ["version", "owner", "run", "namespace", "plan_id"] { + let (fixture, current, _, mut proof) = evicted_historical_fixture(); + let mut changed = serde_json::to_value(&proof.original).unwrap(); + changed[field] = if field == "version" { + json!(9) + } else { + json!("e".repeat(64)) + }; + proof.original = serde_json::from_value(changed).unwrap(); + proof.original_sha256 = selected(&proof.original).unwrap(); + state::write(&fixture.0.join(FILE), &proof).unwrap(); + let bytes = fs::read(fixture.0.join(FILE)).unwrap(); + assert!( + require_historical_recovery(&fixture.0, ¤t).is_err(), + "{field}" + ); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), bytes); + } + let (fixture, mut current, _, proof) = evicted_historical_fixture(); + current.resources = proof.original.resources; + let bytes = fs::read(fixture.0.join(FILE)).unwrap(); + assert!(require_historical_recovery(&fixture.0, ¤t).is_err()); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), bytes); + } + + #[test] + fn evicted_completion_requires_verified_truncation_and_newer_stopped_evidence() { + for fault in [ + "missing", + "untruncated", + "same-generation", + "foreign", + "malformed", + "pending", + ] { + let (fixture, current, stopped, _) = evicted_historical_fixture(); + let path = fixture.0.join("restore-history.json"); + let mut history: Value = state::read(&path).unwrap(); + match fault { + "missing" => fs::remove_file(&path).unwrap(), + "pending" => fs::write( + fixture.0.join("dead-owner-cleanup.pending"), + b"partial evidence", + ) + .unwrap(), + _ => { + match fault { + "untruncated" => history["truncated"] = json!(false), + "same-generation" => { + let mut same_generation = stopped; + same_generation + .readiness + .insert("init".into(), project::execution::Condition::Started); + *history["entries"] + .as_array_mut() + .unwrap() + .last_mut() + .unwrap() = serde_json::to_value(&same_generation).unwrap() + } + "foreign" => history["entries"][0]["owner"] = json!("e".repeat(32)), + "malformed" => history["version"] = json!(9), + _ => unreachable!(), + } + state::write(&path, &history).unwrap(); + } + } + let bytes = fs::read(fixture.0.join(FILE)).unwrap(); + let history = fs::read(&path).ok(); + assert!( + require_historical_recovery(&fixture.0, ¤t).is_err(), + "{fault}" + ); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), bytes); + assert_eq!(fs::read(&path).ok(), history); + if fault == "pending" { + assert_eq!( + fs::read(fixture.0.join("dead-owner-cleanup.pending")).unwrap(), + b"partial evidence" + ); + } + } + } + #[test] fn completed_previous_boot_proof_is_read_only_history_for_later_live_recovery() { let (fixture, mut current, _, _) = historical_fixture(); diff --git a/packages/runtime-core/src/provider/graph/dependency_cache.rs b/packages/runtime-core/src/provider/graph/dependency_cache.rs index a4af6ad93..ca6a44b63 100644 --- a/packages/runtime-core/src/provider/graph/dependency_cache.rs +++ b/packages/runtime-core/src/provider/graph/dependency_cache.rs @@ -257,7 +257,7 @@ fn metadata_text(path: &Path) -> Result { Ok(value.to_owned()) } /// Uses bounded Git metadata only; never invokes Git or reads configuration/credentials. -pub fn scope(project: &Path) -> Result { +fn scope_root(project: &Path) -> Result { let project = owned_directory(project)?; let dotgit = project.join(".git"); let metadata = match fs::symlink_metadata(&dotgit) { @@ -285,13 +285,148 @@ pub fn scope(project: &Path) -> Result { } _ => return Err(refused()), }; - let metadata = fs::metadata(&common).map_err(|_| refused())?; - hash(&( - "hack-dependency-cache-scope-v1", + let directory = fs::OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_DIRECTORY) + .open(&common) + .map_err(|_| refused())?; + let metadata = directory.metadata().map_err(|_| refused())?; + let observed = fs::symlink_metadata(&common).map_err(|_| refused())?; + if !metadata.is_dir() + || !observed.is_dir() + || metadata.dev() != observed.dev() + || metadata.ino() != observed.ino() + { + return Err(refused()); + } + Ok(ScopeRoot { common, - metadata.dev(), - metadata.ino(), - )) + device: metadata.dev(), + inode: metadata.ino(), + }) +} + +struct ScopeRoot { + common: PathBuf, + device: u64, + inode: u64, +} +impl ScopeRoot { + fn hash(&self, device: u64) -> Result { + hash(&( + "hack-dependency-cache-scope-v1", + &self.common, + device, + self.inode, + )) + } +} + +pub fn scope(project: &Path) -> Result { + let root = scope_root(project)?; + root.hash(root.device) +} + +/// Continuity of an existing cache namespace after a selected device rebind. +/// This records no package inputs and grants no volume creation or adoption. +/// Every replay still recomputes the full fingerprint and compares exact resources. +#[derive(Clone, Debug, Serialize, Deserialize, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +pub struct ReplayScope { + common: PathBuf, + device: u64, + inode: u64, + original_device: u64, + origin_sha256: String, +} +impl ReplayScope { + pub(super) fn valid(&self) -> bool { + self.common.is_absolute() + && self.common.as_os_str().len() <= 4096 + && self.common.to_str().is_some_and(|path| { + !path.chars().any(char::is_control) + && self + .common + .components() + .all(|c| matches!(c, Component::RootDir | Component::Normal(_))) + }) + && self.inode != 0 + && hex(&self.origin_sha256) + } + + pub(super) fn scope(&self, project: &Path) -> Result { + let root = scope_root(project)?; + if !self.valid() + || root.common != self.common + || root.device != self.device + || root.inode != self.inode + { + return Err(refused()); + } + root.hash(self.original_device) + } + + /// Called only by an already selected source-device witness. Reconstruct the + /// old hash using the same Git common path/inode and require every retained + /// cache to match it. Git metadata on a different filesystem is not projected. + #[cfg(any(target_os = "macos", test))] + pub(super) fn project( + old: &crate::provider::ProjectShareIntent, + current: &crate::provider::ProjectShareIntent, + previous: Option<&Self>, + retained: &[&str], + origin_sha256: &str, + ) -> Result, CandidateError> { + if retained.is_empty() { + return if previous.is_none() { + Ok(None) + } else { + Err(refused()) + }; + } + if !hex(origin_sha256) + || previous.is_some() + || old.device == current.device + || old.project != current.project + || old.inode != current.inode + { + return Err(refused()); + } + let root = scope_root(¤t.project)?; + if root.device != current.device { + return Err(refused()); + } + let original_device = old.device; + let expected = root.hash(original_device)?; + if retained.iter().any(|scope| *scope != expected) { + return Err(refused()); + } + Ok(Some(Self { + common: root.common, + device: root.device, + inode: root.inode, + original_device, + origin_sha256: origin_sha256.into(), + })) + } + + #[cfg(target_os = "macos")] + pub(super) fn matches_origin( + &self, + sha256: &str, + old: &crate::provider::ProjectShareIntent, + current: &crate::provider::ProjectShareIntent, + ) -> bool { + self.valid() + && self.origin_sha256 == sha256 + && self.original_device == old.device + && self.device == current.device + && old.device != current.device + && old.project == current.project + && old.inode == current.inode + && old.guest_path == current.guest_path + && old.unfiltered_source == current.unfiltered_source + } } #[cfg(test)] @@ -364,6 +499,130 @@ mod tests { } .unwrap(); } + + #[test] + fn selected_device_replay_keeps_exact_scope_and_fingerprint_inputs() { + let (source, _home, plan, manifest) = prepared(); + let root = scope_root(&source.0).unwrap(); + let current = crate::provider::ProjectShareIntent::approve(&source.0, true).unwrap(); + let mut old = current.clone(); + old.device += 7; + let retained = root.hash(old.device).unwrap(); + let projected = ReplayScope::project(&old, ¤t, None, &[&retained], &"e".repeat(64)) + .unwrap() + .unwrap(); + assert_eq!(projected.scope(&source.0).unwrap(), retained); + assert_ne!(scope(&source.0).unwrap(), retained); + let persisted: ReplayScope = + serde_json::from_slice(&serde_json::to_vec(&projected).unwrap()).unwrap(); + assert_eq!(persisted.scope(&source.0).unwrap(), retained); + let expected = resolve(&plan, &manifest, &retained).unwrap(); + assert_eq!( + resolve(&plan, &manifest, &persisted.scope(&source.0).unwrap()).unwrap(), + expected + ); + + let mut changed = manifest.clone(); + changed + .entries + .iter_mut() + .find(|entry| entry.path == "bun.lock") + .unwrap() + .sha256 = Some("f".repeat(64)); + resign(&mut changed); + assert_ne!(resolve(&plan, &changed, &retained).unwrap(), expected); + for field in ["environment", "extra_hosts", "entrypoint"] { + let mut executable = executable(&plan); + let service = executable.get_mut("deps").unwrap(); + match field { + "environment" => service.environment.push("NEW=public".into()), + "extra_hosts" => { + service + .extra_hosts + .insert("changed.example".into(), "host-gateway".into()); + } + "entrypoint" => service.entrypoint = Some(vec!["/changed".into()]), + _ => unreachable!(), + } + assert_ne!( + super::resolve(&plan, &manifest, &retained, &executable).unwrap(), + expected + ); + } + assert!( + ReplayScope::project(&old, ¤t, None, &[&"a".repeat(64)], &"e".repeat(64)) + .is_err() + ); + let mut other_filesystem = current.clone(); + other_filesystem.device += 1; + assert!( + ReplayScope::project(&old, &other_filesystem, None, &[&retained], &"e".repeat(64)) + .is_err() + ); + } + + #[test] + fn replay_scope_refuses_replaced_git_root_and_a_second_device_change() { + let fixture = Fixture::new(); + fs::create_dir(fixture.0.join(".git")).unwrap(); + let root = scope_root(&fixture.0).unwrap(); + let current = crate::provider::ProjectShareIntent { + project: fixture.0.clone(), + guest_path: "/mnt/hack-projects/exact".into(), + device: root.device, + inode: fs::metadata(&fixture.0).unwrap().ino(), + unfiltered_source: true, + }; + let mut old = current.clone(); + old.device += 2; + let retained = root.hash(old.device).unwrap(); + let projected = ReplayScope::project(&old, ¤t, None, &[&retained], &"e".repeat(64)) + .unwrap() + .unwrap(); + let mut previous = projected.clone(); + previous.device = old.device + 1; + let mut intermediate = old.clone(); + intermediate.device = previous.device; + assert!( + ReplayScope::project( + &intermediate, + ¤t, + Some(&previous), + &[&retained], + &"e".repeat(64) + ) + .is_err() + ); + previous.inode += 1; + assert!( + ReplayScope::project( + &intermediate, + ¤t, + Some(&previous), + &[&retained], + &"e".repeat(64) + ) + .is_err() + ); + + fs::rename(fixture.0.join(".git"), fixture.0.join("original-git")).unwrap(); + fs::create_dir(fixture.0.join(".git")).unwrap(); + assert!(projected.scope(&fixture.0).is_err()); + assert!(ReplayScope::project(&old, ¤t, None, &[&retained], &"e".repeat(64)).is_err()); + assert!( + ReplayScope::project( + &old, + ¤t, + Some(&projected), + &[&retained], + &"e".repeat(64) + ) + .is_err() + ); + fs::remove_dir(fixture.0.join(".git")).unwrap(); + symlink(fixture.0.join("original-git"), fixture.0.join(".git")).unwrap(); + assert!(projected.scope(&fixture.0).is_err()); + } #[test] fn subpath_layout_changes_cache_identity_without_compose_digest_shortcut() { let (_source, _home, mut plan, manifest) = prepared(); diff --git a/packages/runtime-core/src/provider/graph/dependency_slots.rs b/packages/runtime-core/src/provider/graph/dependency_slots.rs index 922cf1b74..95b684afa 100644 --- a/packages/runtime-core/src/provider/graph/dependency_slots.rs +++ b/packages/runtime-core/src/provider/graph/dependency_slots.rs @@ -2,6 +2,7 @@ //! mapped only under the provider lease. Receipts survive owner death; absence of //! a listener is never permission to steal a recorded or foreign transport. use super::{Candidate, CandidateError, Engine, Receipt, hex, state}; +use crate::provider::host_pin::DeviceRebind; use crate::provider::identity::{self, ProcessIdentity}; use serde::{Deserialize, Serialize}; use sha2::{Digest, Sha256}; @@ -13,6 +14,10 @@ use std::{ path::{Path, PathBuf}, time::{Duration, Instant}, }; +mod acknowledged; +pub(super) use acknowledged::archive_acknowledged; +#[cfg(test)] +pub(super) mod fixture_prior_boot; #[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] #[serde(deny_unknown_fields)] @@ -47,6 +52,14 @@ fn absent(path: &Path) -> Result { } } fn read(path: &Path) -> Result { + read_with_bytes(path).map(|(record, _)| record) +} +fn read_with_bytes(path: &Path) -> Result<(Record, Vec), CandidateError> { + read_with_identity(path).map(|(record, bytes, _)| (record, bytes)) +} +/// Parsed record, original bytes and the no-follow descriptor's file identity. +type ObservedRecord = (Record, Vec, (u64, u64)); +fn read_with_identity(path: &Path) -> Result { let mut file = OpenOptions::new() .read(true) .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK) @@ -70,7 +83,125 @@ fn read(path: &Path) -> Result { if bytes.len() as u64 != m.len() { return Err(refused()); } - serde_json::from_slice(&bytes).map_err(|_| refused()) + let current = fs::symlink_metadata(path).map_err(|_| refused())?; + if (current.dev(), current.ino()) != (m.dev(), m.ino()) || current.nlink() != 1 { + return Err(refused()); + } + let record = serde_json::from_slice(&bytes).map_err(|_| refused())?; + Ok((record, bytes, (m.dev(), m.ino()))) +} +#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +pub(super) struct LegacyReservation { + pub record_sha256: String, + pub sockets: BTreeMap, + pub process: ProcessIdentity, +} +pub(super) fn inspect_legacy( + candidate: &Candidate, + run: &str, + rebind: DeviceRebind, + host_boot_micros: u64, +) -> Result, CandidateError> { + let owner = state::Owner::load(candidate)?; + let directory = root(candidate); + let values = records( + &directory, + &owner.token, + owner.dependency_sockets.map_or(0, |v| v.slots), + )?; + let Some(record) = values.iter().find(|r| r.run == run) else { + return Ok(None); + }; + let (again, bytes) = read_with_bytes(&directory.join(format!("{run}.json")))?; + if &again != record || record.sockets.len() != record.slots.len() { + return Err(refused()); + } + // SAFETY: geteuid has no arguments or side effects. + identity::verify( + &record.process, + &record.process, + &record.process.executable, + unsafe { libc::geteuid() }, + )?; + rebind.definitely_dead_before_boot(&record.process, host_boot_micros)?; + for slot in record.slots.values() { + let expected = record.sockets.get(slot).ok_or_else(refused)?; + let path = socket(&owner.short_home, *slot); + let metadata = fs::symlink_metadata(&path).map_err(|_| refused())?; + if !metadata.file_type().is_socket() + || metadata.uid() != record.process.uid + || metadata.mode() & 0o7777 != 0o600 + || metadata.nlink() != 1 + || !rebind.matches(*expected, (metadata.dev(), metadata.ino())) + { + return Err(refused()); + } + crate::provider::relay_owner::publication::dead::no_listener(&path) + .map_err(|_| refused())?; + } + Ok(Some(LegacyReservation { + record_sha256: format!("{:x}", Sha256::digest(bytes)), + sockets: record.sockets.clone(), + process: record.process.clone(), + })) +} +pub(super) fn verify_legacy_remaining( + candidate: &Candidate, + run: &str, + rebind: DeviceRebind, + selected: &LegacyReservation, + allow_absent_record: bool, +) -> Result<(), CandidateError> { + let owner = state::Owner::load(candidate)?; + let path = root(candidate).join(format!("{run}.json")); + if absent(&path)? { + if !allow_absent_record { + return Err(refused()); + } + for slot in selected.sockets.keys() { + if !absent(&socket(&owner.short_home, *slot))? { + return Err(refused()); + } + } + return Ok(()); + } + let values = records( + &root(candidate), + &owner.token, + owner.dependency_sockets.map_or(0, |v| v.slots), + )?; + let record = values + .iter() + .find(|record| record.run == run) + .ok_or_else(refused)?; + let (same, bytes) = read_with_bytes(&path)?; + if &same != record + || selected.record_sha256 != format!("{:x}", Sha256::digest(bytes)) + || selected.sockets != record.sockets + || selected.process != record.process + { + return Err(refused()); + } + for slot in record.slots.values() { + let path = socket(&owner.short_home, *slot); + if absent(&path)? { + continue; + } + let expected = record.sockets.get(slot).ok_or_else(refused)?; + let observed = fs::symlink_metadata(&path).map_err(|_| refused())?; + if !observed.file_type().is_socket() + || observed.uid() != record.process.uid + || observed.mode() & 0o7777 != 0o600 + || observed.nlink() != 1 + || !rebind.matches(*expected, (observed.dev(), observed.ino())) + { + return Err(refused()); + } + crate::provider::relay_owner::publication::dead::no_listener(&path) + .map_err(|_| refused())?; + } + Ok(()) } fn valid(record: &Record, owner: &str, capacity: u8) -> bool { record.version == 1 @@ -339,7 +470,11 @@ pub fn inspect(candidate: &Candidate) -> Result Result<(), CandidateError> { +fn recover_record( + candidate: &Candidate, + record: &Record, + rebind: Option, +) -> Result<(), CandidateError> { if identity::alive(record.process.pid)? { return Err(refused()); } @@ -347,10 +482,14 @@ fn recover_record(candidate: &Candidate, record: &Record) -> Result<(), Candidat if record.owner != owner.token { return Err(refused()); } - recover_paths(&owner.short_home, record)?; + recover_paths(&owner.short_home, record, rebind)?; remove_record(&root(candidate), record) } -fn recover_paths(home: &Path, record: &Record) -> Result<(), CandidateError> { +fn recover_paths( + home: &Path, + record: &Record, + rebind: Option, +) -> Result<(), CandidateError> { // Validate every selected path before removing any. Unexpected or unrecorded // paths are never adopted, even when nobody currently listens on them. let mut selected = Vec::new(); @@ -364,7 +503,11 @@ fn recover_paths(home: &Path, record: &Record) -> Result<(), CandidateError> { || m.mode() & 0o7777 != 0o600 || m.uid() != record.process.uid || m.nlink() != 1 - || record.sockets.get(slot) != Some(&(m.dev(), m.ino())) + || !record.sockets.get(slot).is_some_and(|expected| { + rebind.map_or(*expected == (m.dev(), m.ino()), |v| { + v.matches(*expected, (m.dev(), m.ino())) + }) + }) { return Err(refused()); } @@ -402,13 +545,14 @@ pub fn recover_orphan( if fingerprint(record)? != expected || !absent(&super::directory(candidate, run)?)? { return Err(refused()); } - recover_record(candidate, record)?; + recover_record(candidate, record, None)?; Ok(serde_json::json!({"run":run,"released":true})) } /// Caller holds Engine lease and has completed dead-owner cleanup/verification. pub(super) fn recover_cleaned( candidate: &Candidate, receipt: &Receipt, + rebind: Option<(DeviceRebind, &LegacyReservation)>, ) -> Result<(), CandidateError> { let directory = root(candidate); if absent(&directory.join(format!("{}.json", receipt.run)))? { @@ -438,7 +582,17 @@ pub(super) fn recover_cleaned( { return Err(refused()); } - recover_record(candidate, record) + if let Some((_, selected)) = rebind { + let (again, bytes) = read_with_bytes(&directory.join(format!("{}.json", receipt.run)))?; + if &again != record + || selected.record_sha256 != format!("{:x}", Sha256::digest(bytes)) + || selected.sockets != record.sockets + || selected.process != record.process + { + return Err(refused()); + } + } + recover_record(candidate, record, rebind.map(|(value, _)| value)) } #[cfg(test)] @@ -555,14 +709,14 @@ mod tests { let mut b = fixture_record('2', 1); let a_listener = listener(&fixture.0, &mut a, 0); let b_listener = listener(&fixture.0, &mut b, 1); - assert!(recover_paths(&fixture.0, &a).is_err()); + assert!(recover_paths(&fixture.0, &a, None).is_err()); drop(a_listener); let old = a.sockets[&0]; a.sockets.insert(0, (old.0, old.1 + 1)); - assert!(recover_paths(&fixture.0, &a).is_err()); + assert!(recover_paths(&fixture.0, &a, None).is_err()); assert!(socket(&fixture.0, 0).exists()); a.sockets.insert(0, old); - recover_paths(&fixture.0, &a).unwrap(); + recover_paths(&fixture.0, &a, None).unwrap(); assert!(!socket(&fixture.0, 0).exists()); assert!(UnixStream::connect(socket(&fixture.0, 1)).is_ok()); drop(b_listener); @@ -575,10 +729,10 @@ mod tests { let first = listener(&fixture.0, &mut a, 0); let second = listener(&fixture.0, &mut a, 1); drop(first); - assert!(recover_paths(&fixture.0, &a).is_err()); + assert!(recover_paths(&fixture.0, &a, None).is_err()); assert!(socket(&fixture.0, 0).exists()); drop(second); - recover_paths(&fixture.0, &a).unwrap(); + recover_paths(&fixture.0, &a, None).unwrap(); assert!(!socket(&fixture.0, 0).exists()); assert!(!socket(&fixture.0, 1).exists()); } diff --git a/packages/runtime-core/src/provider/graph/dependency_slots/acknowledged.rs b/packages/runtime-core/src/provider/graph/dependency_slots/acknowledged.rs new file mode 100644 index 000000000..d62afffc1 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/dependency_slots/acknowledged.rs @@ -0,0 +1,606 @@ +//! A completed current-boot cleanup may strand a reservation if later archival +//! fails. Exact retirement and ACK authority come from the caller, not this file. +use super::*; +use std::ffi::CString; + +struct Paths<'a> { + directory: &'a Path, + home: &'a Path, + canonical_home: &'a Path, + history: &'a Path, +} + +struct HomePin<'a> { + alias: &'a Path, + target: &'a Path, + alias_id: (u64, u64), + target_id: (u64, u64), +} +impl<'a> HomePin<'a> { + fn capture(alias: &'a Path, target: &'a Path) -> Result { + let a = fs::symlink_metadata(alias).map_err(|_| refused())?; + let t = fs::symlink_metadata(target).map_err(|_| refused())?; + let pin = Self { + alias, + target, + alias_id: (a.dev(), a.ino()), + target_id: (t.dev(), t.ino()), + }; + pin.verify()?; + Ok(pin) + } + fn verify(&self) -> Result<(), CandidateError> { + let a = fs::symlink_metadata(self.alias).map_err(|_| refused())?; + let t = fs::symlink_metadata(self.target).map_err(|_| refused())?; + state::check_private_directory(self.target)?; + // SAFETY: geteuid has no parameters or side effects. + if !a.file_type().is_symlink() + || a.uid() != unsafe { libc::geteuid() } + || a.nlink() != 1 + || (a.dev(), a.ino()) != self.alias_id + || !t.is_dir() + || (t.dev(), t.ino()) != self.target_id + || fs::read_link(self.alias).map_err(|_| refused())? != self.target + { + return Err(refused()); + } + Ok(()) + } +} +struct Selection<'a> { + owner: &'a str, + run: &'a str, + boot: &'a str, + process: &'a ProcessIdentity, + slots: BTreeSet, + expected: &'a str, + capacity: u8, +} + +fn validate(record: &Record, selected: &Selection<'_>, home: &Path) -> Result<(), CandidateError> { + if !valid(record, selected.owner, selected.capacity) + || record.run != selected.run + || record.boot != selected.boot + || record.process != *selected.process + || record.slots.values().copied().collect::>() != selected.slots + || fingerprint(record)? != selected.expected + || identity::alive(record.process.pid)? + { + return Err(refused()); + } + // Even a matching stale socket is retained: this operation moves only the + // record after ordinary cleanup already closed every recorded listener. + for slot in record.slots.values() { + if !absent(&socket(home, *slot))? { + return Err(refused()); + } + } + Ok(()) +} + +fn rename_exclusive(source: &Path, destination: &Path) -> Result<(), CandidateError> { + let source = CString::new(source.as_os_str().as_encoded_bytes()).map_err(|_| refused())?; + let destination = + CString::new(destination.as_os_str().as_encoded_bytes()).map_err(|_| refused())?; + // SAFETY: both C strings remain valid during this macOS call. RENAME_EXCL + // atomically refuses an occupied target; no existing evidence is replaced. + if unsafe { libc::renamex_np(source.as_ptr(), destination.as_ptr(), libc::RENAME_EXCL) } != 0 { + return Err(refused()); + } + Ok(()) +} + +fn sync_directory(path: &Path, expected: (u64, u64)) -> Result<(), CandidateError> { + let file = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK | libc::O_DIRECTORY) + .open(path) + .map_err(|_| refused())?; + let m = file.metadata().map_err(|_| refused())?; + if (m.dev(), m.ino()) != expected || !m.is_dir() { + return Err(refused()); + } + file.sync_all().map_err(state::io) +} + +fn archive( + paths: Paths<'_>, + selected: Selection<'_>, + verify: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + verify()?; + state::check_private_directory(paths.directory)?; + state::check_private_directory(paths.history)?; + let home = HomePin::capture(paths.home, paths.canonical_home)?; + let directory_metadata = fs::symlink_metadata(paths.directory).map_err(|_| refused())?; + let history_metadata = fs::symlink_metadata(paths.history).map_err(|_| refused())?; + let parents = [ + (directory_metadata.dev(), directory_metadata.ino()), + (history_metadata.dev(), history_metadata.ino()), + ]; + let verify_paths = || { + home.verify()?; + for (path, expected) in [paths.directory, paths.history].into_iter().zip(parents) { + state::check_private_directory(path)?; + let m = fs::symlink_metadata(path).map_err(|_| refused())?; + if (m.dev(), m.ino()) != expected { + return Err(refused()); + } + } + Ok(()) + }; + records(paths.directory, selected.owner, selected.capacity)?; + let source = paths.directory.join(format!("{}.json", selected.run)); + let target = paths.history.join(format!( + "dependency-reservation-retired-{}.json", + selected.expected + )); + if !absent(&source.with_extension("pending"))? || !absent(&target.with_extension("pending"))? { + return Err(refused()); + } + if absent(&source)? { + // A retry must find the exact retained record; ordinary absence is not + // authority and a new same-run reservation is never adopted. + let (record, bytes, id) = read_with_identity(&target)?; + verify_paths()?; + validate(&record, &selected, paths.canonical_home)?; + verify()?; + verify_paths()?; + for (path, expected) in [paths.directory, paths.history].into_iter().zip(parents) { + sync_directory(path, expected)?; + } + let (again, current, current_id) = read_with_identity(&target)?; + if again != record || current != bytes || current_id != id || !absent(&source)? { + return Err(refused()); + } + validate(&again, &selected, paths.canonical_home)?; + return Ok(()); + } + if !absent(&target)? { + return Err(refused()); + } + let (record, bytes, id) = read_with_identity(&source)?; + verify_paths()?; + validate(&record, &selected, paths.canonical_home)?; + verify()?; + let (again, current, current_id) = read_with_identity(&source)?; + verify_paths()?; + if again != record || current != bytes || current_id != id { + return Err(refused()); + } + validate(&again, &selected, paths.canonical_home)?; + let file = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK) + .open(&source) + .map_err(|_| refused())?; + let m = file.metadata().map_err(|_| refused())?; + if !m.is_file() || (m.dev(), m.ino()) != id || m.nlink() != 1 { + return Err(refused()); + } + file.sync_all().map_err(state::io)?; + verify()?; + // The provider/publication locks stay held. Recheck the selected inode after + // the final external fence, including tests that replace it at that window. + let (again, current, current_id) = read_with_identity(&source)?; + verify_paths()?; + if again != record || current != bytes || current_id != id { + return Err(refused()); + } + validate(&again, &selected, paths.canonical_home)?; + rename_exclusive(&source, &target)?; + sync_directory(paths.history, parents[1])?; + sync_directory(paths.directory, parents[0])?; + let (retained, retained_bytes, retained_id) = read_with_identity(&target)?; + verify_paths()?; + if retained != record || retained_bytes != bytes || retained_id != id || !absent(&source)? { + return Err(refused()); + } + validate(&retained, &selected, paths.canonical_home)?; + verify()?; + verify_paths()?; + let (again, current, current_id) = read_with_identity(&target)?; + if again != record || current != bytes || current_id != id || !absent(&source)? { + return Err(refused()); + } + validate(&again, &selected, paths.canonical_home) +} + +/// Caller holds the provider and retired-publisher locks and supplies an exact +/// current ACK fence. No socket, graph receipt, named data or sibling is changed. +pub(in crate::provider::graph) fn archive_acknowledged( + candidate: &Candidate, + receipt: &Receipt, + boot: &str, + process: &ProcessIdentity, + expected: &str, + verify: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + let owner = state::Owner::load(candidate)?; + if receipt.phase != "stopped-data-retained" || receipt.owner != owner.token { + return Err(refused()); + } + let directory = root(candidate); + let history = super::super::directory(candidate, &receipt.run)?; + let canonical_home = candidate.state_root.join("run/smolvm/home"); + let slots = receipt + .relay_startup + .as_ref() + .ok_or_else(refused)? + .services + .values() + .flat_map(|service| service.bindings.values().map(|binding| binding.slot)) + .collect(); + archive( + Paths { + directory: &directory, + home: &owner.short_home, + canonical_home: &canonical_home, + history: &history, + }, + Selection { + owner: &owner.token, + run: &receipt.run, + boot, + process, + slots, + expected, + capacity: owner.dependency_sockets.ok_or_else(refused)?.slots, + }, + verify, + ) +} + +#[cfg(test)] +mod tests { + use super::*; + use std::{cell::Cell, os::unix::fs::PermissionsExt, process::Command}; + + struct Fixture { + _root: super::super::super::tests::Fixture, + directory: PathBuf, + home: PathBuf, + canonical_home: PathBuf, + history: PathBuf, + record: Record, + expected: String, + } + impl Fixture { + fn new() -> Self { + let root = super::super::super::tests::Fixture::new(); + let directory = root.0.join("assignments"); + let canonical_home = root.0.join("home"); + let home = PathBuf::from(format!( + "/private/tmp/hkdr-{}", + &format!( + "{:x}", + Sha256::digest(root.0.as_os_str().as_encoded_bytes()) + )[..16] + )); + let history = root.0.join("graph"); + for path in [&directory, &canonical_home, &history] { + state::private_directory(path).unwrap(); + } + std::os::unix::fs::symlink(&canonical_home, &home).unwrap(); + let mut child = Command::new("/bin/sleep").arg("30").spawn().unwrap(); + let process = identity::observe(child.id() as i32).unwrap(); + child.kill().unwrap(); + child.wait().unwrap(); + let record = Record { + version: 1, + owner: "a".repeat(32), + boot: "b".repeat(36), + run: "1".repeat(32), + token: "c".repeat(32), + process, + slots: BTreeMap::from([(0, 0)]), + sockets: BTreeMap::from([(0, (1, 2))]), + }; + state::write(&directory.join(format!("{}.json", record.run)), &record).unwrap(); + let expected = fingerprint(&record).unwrap(); + Self { + _root: root, + directory, + home, + canonical_home, + history, + record, + expected, + } + } + fn source(&self) -> PathBuf { + self.directory.join(format!("{}.json", self.record.run)) + } + fn target(&self) -> PathBuf { + self.history.join(format!( + "dependency-reservation-retired-{}.json", + self.expected + )) + } + fn archive( + &self, + verify: &dyn Fn() -> Result<(), CandidateError>, + ) -> Result<(), CandidateError> { + archive( + Paths { + directory: &self.directory, + home: &self.home, + canonical_home: &self.canonical_home, + history: &self.history, + }, + Selection { + owner: &self.record.owner, + run: &self.record.run, + boot: &self.record.boot, + process: &self.record.process, + slots: BTreeSet::from([0]), + expected: &self.expected, + capacity: 2, + }, + verify, + ) + } + } + impl Drop for Fixture { + fn drop(&mut self) { + fs::remove_file(&self.home).unwrap(); + } + } + + #[test] + fn exact_dead_claim_is_retained_with_same_bytes_inode_and_sibling() { + let fixture = Fixture::new(); + let before = fs::read(fixture.source()).unwrap(); + let m = fs::symlink_metadata(fixture.source()).unwrap(); + let mut sibling = fixture.record.clone(); + sibling.run = "2".repeat(32); + sibling.slots = BTreeMap::from([(0, 1)]); + sibling.sockets = BTreeMap::from([(1, (1, 3))]); + let path = fixture.directory.join(format!("{}.json", sibling.run)); + state::write(&path, &sibling).unwrap(); + let sibling_before = fs::read(&path).unwrap(); + fixture.archive(&|| Ok(())).unwrap(); + assert!(!fixture.source().exists()); + assert_eq!(fs::read(fixture.target()).unwrap(), before); + let retained = fs::symlink_metadata(fixture.target()).unwrap(); + assert_eq!((retained.dev(), retained.ino()), (m.dev(), m.ino())); + assert_eq!(fs::read(path).unwrap(), sibling_before); + fixture.archive(&|| Ok(())).unwrap(); + } + + #[test] + fn mismatched_process_boot_slots_and_live_owner_refuse() { + let fixture = Fixture::new(); + let base = &fixture.record; + let selected = Selection { + owner: &base.owner, + run: &base.run, + boot: &base.boot, + process: &base.process, + slots: BTreeSet::from([0]), + expected: &fixture.expected, + capacity: 2, + }; + for change in ["process", "boot", "slots", "run", "token", "live"] { + let mut changed = base.clone(); + match change { + "process" => changed.process.start_micros += 1, + "boot" => changed.boot = "d".repeat(36), + "slots" => changed.slots.insert(0, 1).map(|_| ()).unwrap(), + "run" => changed.run = "9".repeat(32), + "token" => changed.token = "9".repeat(32), + "live" => changed.process = identity::observe(std::process::id() as i32).unwrap(), + _ => unreachable!(), + } + assert!( + validate(&changed, &selected, &fixture.home).is_err(), + "{change}" + ); + } + let live = identity::observe(std::process::id() as i32).unwrap(); + let mut changed = base.clone(); + changed.process = live.clone(); + let expected = fingerprint(&changed).unwrap(); + let live_selection = Selection { + process: &live, + expected: &expected, + ..selected + }; + assert!(validate(&changed, &live_selection, &fixture.home).is_err()); + assert!(fixture.source().exists()); + } + + #[test] + fn every_socket_pending_record_and_occupied_history_is_preserved() { + use std::os::unix::net::UnixListener; + for reason in ["socket", "pending", "history"] { + let fixture = Fixture::new(); + let before = fs::read(fixture.source()).unwrap(); + let path = match reason { + "socket" => socket(&fixture.home, 0), + "pending" => fixture.source().with_extension("pending"), + "history" => fixture.target(), + _ => unreachable!(), + }; + let listener = if reason == "socket" { + Some(UnixListener::bind(&path).unwrap()) + } else { + fs::write(&path, b"foreign or partial").unwrap(); + fs::set_permissions(&path, fs::Permissions::from_mode(0o600)).unwrap(); + None + }; + let id = fs::symlink_metadata(&path).unwrap(); + assert!(fixture.archive(&|| Ok(())).is_err(), "{reason}"); + assert_eq!(fs::read(fixture.source()).unwrap(), before); + let after = fs::symlink_metadata(path).unwrap(); + assert_eq!((id.dev(), id.ino()), (after.dev(), after.ino())); + drop(listener); + } + } + + #[test] + fn failed_fences_preserve_claim_or_allow_exact_post_move_retry() { + for fault in 1..=4 { + let fixture = Fixture::new(); + let before = fs::read(fixture.source()).unwrap(); + let calls = Cell::new(0); + assert!( + fixture + .archive(&|| { + calls.set(calls.get() + 1); + if calls.get() == fault { + Err(refused()) + } else { + Ok(()) + } + }) + .is_err(), + "fence {fault}" + ); + if fault < 4 { + assert_eq!(fs::read(fixture.source()).unwrap(), before); + assert!(!fixture.target().exists()); + } else { + assert!(!fixture.source().exists()); + assert_eq!(fs::read(fixture.target()).unwrap(), before); + } + fixture.archive(&|| Ok(())).unwrap(); + assert_eq!(fs::read(fixture.target()).unwrap(), before); + } + } + + #[test] + fn same_bytes_replacement_and_late_history_collision_refuse() { + for replacement in ["source", "history", "socket"] { + let fixture = Fixture::new(); + let before = fs::read(fixture.source()).unwrap(); + let calls = Cell::new(0); + let displaced = fixture.history.join("displaced.json"); + assert!( + fixture + .archive(&|| { + calls.set(calls.get() + 1); + if calls.get() == 3 { + match replacement { + "source" => { + fs::rename(fixture.source(), &displaced).unwrap(); + fs::write(fixture.source(), &before).unwrap(); + fs::set_permissions( + fixture.source(), + fs::Permissions::from_mode(0o600), + ) + .unwrap(); + } + "history" => { + fs::write(fixture.target(), b"foreign").unwrap(); + } + "socket" => { + fs::write(socket(&fixture.home, 0), b"foreign").unwrap(); + } + _ => unreachable!(), + } + } + Ok(()) + }) + .is_err(), + "{replacement}" + ); + assert_eq!(fs::read(fixture.source()).unwrap(), before); + if replacement == "history" { + assert_eq!(fs::read(fixture.target()).unwrap(), b"foreign"); + } + if replacement == "source" { + assert_eq!(fs::read(displaced).unwrap(), before); + } + if replacement == "socket" { + assert_eq!(fs::read(socket(&fixture.home, 0)).unwrap(), b"foreign"); + } + } + } + + #[test] + fn alias_home_and_reservation_parent_replacement_refuse_before_release() { + for replacement in ["alias", "home", "assignments", "history"] { + let fixture = Fixture::new(); + let before = fs::read(fixture.source()).unwrap(); + let calls = Cell::new(0); + let displaced = fixture._root.0.join("displaced"); + assert!( + fixture + .archive(&|| { + calls.set(calls.get() + 1); + if calls.get() == 3 { + match replacement { + "alias" => { + fs::rename(&fixture.home, &displaced).unwrap(); + std::os::unix::fs::symlink( + &fixture.canonical_home, + &fixture.home, + ) + .unwrap(); + } + "home" => { + fs::rename(&fixture.canonical_home, &displaced).unwrap(); + state::private_directory(&fixture.canonical_home).unwrap(); + } + "assignments" => { + fs::rename(&fixture.directory, &displaced).unwrap(); + state::private_directory(&fixture.directory).unwrap(); + fs::write(fixture.source(), &before).unwrap(); + fs::set_permissions( + fixture.source(), + fs::Permissions::from_mode(0o600), + ) + .unwrap(); + } + "history" => { + fs::rename(&fixture.history, &displaced).unwrap(); + state::private_directory(&fixture.history).unwrap(); + } + _ => unreachable!(), + } + } + Ok(()) + }) + .is_err(), + "{replacement}" + ); + assert_eq!(fs::read(fixture.source()).unwrap(), before); + assert!(!fixture.target().exists()); + } + } + + #[test] + fn retry_and_final_success_recheck_retained_record_after_external_fence() { + for retry in [false, true] { + let fixture = Fixture::new(); + let before = fs::read(fixture.source()).unwrap(); + if retry { + fixture.archive(&|| Ok(())).unwrap(); + } + let calls = Cell::new(0); + let displaced = fixture.history.join("original.json"); + assert!( + fixture + .archive(&|| { + calls.set(calls.get() + 1); + if calls.get() == if retry { 2 } else { 4 } { + fs::rename(fixture.target(), &displaced).unwrap(); + fs::write(fixture.target(), &before).unwrap(); + fs::set_permissions( + fixture.target(), + fs::Permissions::from_mode(0o600), + ) + .unwrap(); + } + Ok(()) + }) + .is_err(), + "retry={retry}" + ); + assert_eq!(fs::read(fixture.target()).unwrap(), before); + assert_eq!(fs::read(displaced).unwrap(), before); + assert!(!fixture.source().exists()); + } + } +} diff --git a/packages/runtime-core/src/provider/graph/dependency_slots/fixture_prior_boot.rs b/packages/runtime-core/src/provider/graph/dependency_slots/fixture_prior_boot.rs new file mode 100644 index 000000000..0216db7d6 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/dependency_slots/fixture_prior_boot.rs @@ -0,0 +1,325 @@ +//! Test-only pre-host-boot projection for an exact dead, previously bound claim. +//! The journal and guest observations remain real; this is not a physical reboot. +use super::*; + +pub(in crate::provider::graph) fn synthesize( + candidate: &Candidate, + receipt: &Receipt, + expected_process: &ProcessIdentity, + rebind: DeviceRebind, + host_boot_micros: u64, +) -> Result<(), CandidateError> { + let _lease = state::Lock::acquire_existing(&candidate.state_root.join("run/smolvm"))?; + let owner = state::Owner::load(candidate)?; + let graph = super::super::directory(candidate, &receipt.run)?; + let graph_bytes = + super::super::host_pin_recovery::read_raw(&graph.join("state.json"), 4 * 1024 * 1024)?; + if receipt.phase != "ready-observed" + || receipt.owner != owner.token + || owner.phase != "running" + || owner.storage.as_ref().map(|disk| disk.device) != Some(rebind.current) + || rebind.old == rebind.current + || host_boot_micros <= 1 + || graph_bytes != serde_json::to_vec_pretty(receipt).map_err(|_| refused())? + { + return Err(refused()); + } + let values = records( + &root(candidate), + &owner.token, + owner.dependency_sockets.map_or(0, |intent| intent.slots), + )?; + let record = values + .iter() + .find(|r| r.run == receipt.run) + .ok_or_else(refused)?; + // SAFETY: geteuid has no arguments or side effects. + identity::verify( + &record.process, + expected_process, + &expected_process.executable, + unsafe { libc::geteuid() }, + )?; + let slots = receipt + .relay_startup + .as_ref() + .ok_or_else(refused)? + .services + .values() + .flat_map(|service| service.bindings.values().map(|binding| binding.slot)) + .collect::>(); + if owner.previous_guest_boot_id.as_deref() != Some(record.boot.as_str()) + || owner.guest_boot_id.as_deref() == Some(record.boot.as_str()) + || record.slots.values().copied().collect::>() != slots + || record.sockets.len() != record.slots.len() + || record.process.start_micros < host_boot_micros + || identity::alive(record.process.pid)? + { + return Err(refused()); + } + let path = root(candidate).join(format!("{}.json", receipt.run)); + let (again, bytes, id) = read_with_identity(&path)?; + if &again != record || !absent(&path.with_extension("pending"))? { + return Err(refused()); + } + // Preflight every actual socket before translating only its recorded device. + // A replaced path, live listener or incomplete binding is never fabricated. + for slot in record.slots.values() { + let path = socket(&owner.short_home, *slot); + let metadata = fs::symlink_metadata(&path).map_err(|_| refused())?; + if !metadata.file_type().is_socket() + || metadata.uid() != record.process.uid + || metadata.mode() & 0o7777 != 0o600 + || metadata.nlink() != 1 + || record.sockets.get(slot) != Some(&(metadata.dev(), metadata.ino())) + || metadata.dev() != rebind.current + { + return Err(refused()); + } + crate::provider::relay_owner::publication::dead::no_listener(&path)?; + } + let mut projected = record.clone(); + projected.process.start_micros = host_boot_micros - 1; + for identity in projected.sockets.values_mut() { + identity.0 = rebind.old; + } + // Preserve PID/UID/executable, owner/run/token/boot, slots and socket inodes. + let (same, same_bytes, same_id) = read_with_identity(&path)?; + if same != *record || same_bytes != bytes || same_id != id { + return Err(refused()); + } + state::write(&path, &projected)?; + // Exercise the unchanged production inspector before historical retirement. + inspect_legacy(candidate, &receipt.run, rebind, host_boot_micros)?.ok_or_else(refused)?; + Ok(()) +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::provider::state::Owner; + use serde_json::json; + use std::os::unix::{fs::PermissionsExt, net::UnixListener}; + + struct Fixture { + root: super::super::super::tests::Fixture, + candidate: Candidate, + owner: Owner, + receipt: Receipt, + record: Record, + rebind: DeviceRebind, + host_boot: u64, + } + impl Fixture { + fn new() -> Self { + let root = super::super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&root.0).unwrap(); + let provider = candidate.state_root.join("run/smolvm"); + state::private_directory(&provider.join("home")).unwrap(); + drop(state::Lock::acquire(&provider).unwrap()); + let token = super::super::super::probes::token().unwrap(); + let short_home = Path::new("/private/tmp").join(format!("hkl-{}", &token[..12])); + std::os::unix::fs::symlink(provider.join("home"), &short_home).unwrap(); + let current = fs::metadata(provider.join("home")).unwrap().dev(); + let owner: Owner = serde_json::from_value(json!({ + "version":1,"checkout":candidate.checkout,"token":token, + "machine":format!("hack-{}",&token[..12]),"short_home":short_home, + "created":true,"phase":"running","process":null, + "dependency_sockets":{"slots":3}, + "storage":{"device":current,"inode":1,"bytes":1,"uuid":"synthetic"}, + "overlay":null,"guest_boot_id":"d".repeat(36), + "previous_guest_boot_id":"b".repeat(36),"daemon_pid":null, + "daemon_start":null,"rootfs_digest":null + })) + .unwrap(); + state::write(&provider.join("owner.json"), &owner).unwrap(); + let host_boot = + crate::provider::lifecycle::host_filesystem::host_boot_micros().unwrap(); + let mut process = identity::observe(std::process::id() as i32).unwrap(); + process.pid = i32::MAX; + assert!(!identity::alive(process.pid).unwrap()); + let mut record = Record { + version: 1, + owner: token, + boot: "b".repeat(36), + run: "a".repeat(32), + token: "c".repeat(32), + process, + slots: BTreeMap::from([(0, 0), (1, 1)]), + sockets: BTreeMap::new(), + }; + for slot in 0..3 { + let path = socket(&short_home, slot); + let listener = UnixListener::bind(&path).unwrap(); + fs::set_permissions(&path, fs::Permissions::from_mode(0o600)).unwrap(); + let m = fs::symlink_metadata(&path).unwrap(); + if slot < 2 { + record.sockets.insert(slot, (m.dev(), m.ino())); + } + drop(listener); + } + state::private_directory(&super::super::root(&candidate)).unwrap(); + state::write( + &super::super::root(&candidate).join(format!("{}.json", record.run)), + &record, + ) + .unwrap(); + let mut sibling = record.clone(); + sibling.run = "e".repeat(32); + sibling.slots = BTreeMap::from([(0, 2)]); + let m = fs::symlink_metadata(socket(&short_home, 2)).unwrap(); + sibling.sockets = BTreeMap::from([(2, (m.dev(), m.ino()))]); + state::write( + &super::super::root(&candidate).join(format!("{}.json", sibling.run)), + &sibling, + ) + .unwrap(); + let receipt: Receipt = serde_json::from_value(json!({ + "version":1,"run":record.run,"owner":owner.token,"namespace":"c".repeat(64), + "plan_id":"d".repeat(64),"phase":"ready-observed","readiness":{},"resources":{}, + "relay_startup":{"guest_root":null,"control_root":"/private/fixture-control", + "artifact":"f".repeat(64),"services":{"web":{"generation":"a".repeat(32), + "phase":"released","started_at":null,"bindings":{ + "default":{"slot":0,"port":25252,"process":null}, + "second":{"slot":1,"port":25253,"process":null}}}}} + })) + .unwrap(); + let graph = super::super::super::directory(&candidate, &record.run).unwrap(); + state::private_directory(&graph).unwrap(); + state::write(&graph.join("state.json"), &receipt).unwrap(); + fs::write(graph.join("dependency-rebind.json"), b"unchanged journal").unwrap(); + Self { + root, + candidate, + owner, + receipt, + record, + rebind: DeviceRebind { + old: current + 1, + current, + }, + host_boot, + } + } + fn path(&self) -> PathBuf { + super::super::root(&self.candidate).join(format!("{}.json", self.record.run)) + } + fn project(&self) -> Result<(), CandidateError> { + synthesize( + &self.candidate, + &self.receipt, + &self.record.process, + self.rebind, + self.host_boot, + ) + } + } + impl Drop for Fixture { + fn drop(&mut self) { + let expected = self.candidate.state_root.join("run/smolvm/home"); + assert_eq!(fs::read_link(&self.owner.short_home).unwrap(), expected); + fs::remove_file(&self.owner.short_home).unwrap(); + assert!(self.root.0.exists()); + } + } + + #[test] + fn projects_every_selected_socket_and_only_prior_boot_fields() { + let fixture = Fixture::new(); + let sibling = + super::super::root(&fixture.candidate).join(format!("{}.json", "e".repeat(32))); + let sibling_before = fs::read(&sibling).unwrap(); + let graph = + super::super::super::directory(&fixture.candidate, &fixture.record.run).unwrap(); + let graph_before = fs::read(graph.join("state.json")).unwrap(); + fixture.project().unwrap(); + let mut expected = fixture.record.clone(); + expected.process.start_micros = fixture.host_boot - 1; + for id in expected.sockets.values_mut() { + id.0 = fixture.rebind.old; + } + assert_eq!(read(&fixture.path()).unwrap(), expected); + assert_eq!(fs::read(sibling).unwrap(), sibling_before); + assert_eq!(fs::read(graph.join("state.json")).unwrap(), graph_before); + assert_eq!( + fs::read(graph.join("dependency-rebind.json")).unwrap(), + b"unchanged journal" + ); + } + + #[test] + fn production_inspector_still_refuses_current_boot_and_device_mismatches() { + let fixture = Fixture::new(); + let inspect = || { + inspect_legacy( + &fixture.candidate, + &fixture.record.run, + fixture.rebind, + fixture.host_boot, + ) + }; + assert_eq!(inspect().unwrap_err().code, "host_pin_recovery"); + fixture.project().unwrap(); + let projected: Record = read(&fixture.path()).unwrap(); + for case in 0..3 { + let mut changed: Record = projected.clone(); + let code = if case == 0 { + changed.process.start_micros = fixture.host_boot; + "host_pin_recovery" + } else { + let id = changed.sockets.get_mut(&1).unwrap(); + if case == 1 { + id.0 = fixture.rebind.current; + } else { + id.1 += 1; + } + "dependency_reservation" + }; + state::write(&fixture.path(), &changed).unwrap(); + let before = fs::read(fixture.path()).unwrap(); + assert_eq!(inspect().unwrap_err().code, code); + assert_eq!(fs::read(fixture.path()).unwrap(), before); + assert!(socket(&fixture.owner.short_home, 0).exists()); + assert!(socket(&fixture.owner.short_home, 1).exists()); + } + } + + #[test] + fn setup_refuses_changed_schema_identity_and_live_listener_before_writing() { + for case in 0..7 { + let fixture = Fixture::new(); + let path = fixture.path(); + let mut record = fixture.record.clone(); + let mut live = None; + match case { + 0 => record.owner = "0".repeat(32), + 1 => record.boot = "0".repeat(36), + 2 => record.sockets.get_mut(&1).unwrap().1 += 1, + 3 => record.process = identity::observe(std::process::id() as i32).unwrap(), + 4 => { + fs::remove_file(socket(&fixture.owner.short_home, 1)).unwrap(); + live = Some(UnixListener::bind(socket(&fixture.owner.short_home, 1)).unwrap()); + fs::set_permissions( + socket(&fixture.owner.short_home, 1), + fs::Permissions::from_mode(0o600), + ) + .unwrap(); + let m = fs::symlink_metadata(socket(&fixture.owner.short_home, 1)).unwrap(); + record.sockets.insert(1, (m.dev(), m.ino())); + } + 6 => record.process.executable = PathBuf::from("/tmp/unrelated-executable"), + _ => {} + } + state::write(&path, &record).unwrap(); + if case == 5 { + let mut invalid = serde_json::to_value(&record).unwrap(); + invalid["unknown"] = json!(true); + state::write(&path, &invalid).unwrap(); + } + let before = fs::read(&path).unwrap(); + assert!(fixture.project().is_err()); + assert_eq!(fs::read(path).unwrap(), before); + drop(live); + } + } +} diff --git a/packages/runtime-core/src/provider/graph/foreground.rs b/packages/runtime-core/src/provider/graph/foreground.rs index 7f46e84be..e602f8906 100644 --- a/packages/runtime-core/src/provider/graph/foreground.rs +++ b/packages/runtime-core/src/provider/graph/foreground.rs @@ -20,7 +20,6 @@ mod tests; pub(in crate::provider::graph) mod transport; pub(in crate::provider::graph) use transport::DeadOwner; pub(in crate::provider::graph) use transport::retire_recovered_publisher as retire_publisher_path; -pub(in crate::provider::graph) use transport::verify_recovered_publisher_retired; use transport::{Publication, WireRequest}; fn refused() -> CandidateError { CandidateError::new( @@ -164,10 +163,20 @@ pub fn restore_selection(candidate: &Candidate, run: &str) -> Result( deadline, &mut runtime, generation, + &|| publication.verify(), ) } else if let Some((compose, _)) = normalized { super::run_normalized_with_host_dependencies_until( diff --git a/packages/runtime-core/src/provider/graph/foreground/native_test.rs b/packages/runtime-core/src/provider/graph/foreground/native_test.rs index 016db1b0f..041c5eceb 100644 --- a/packages/runtime-core/src/provider/graph/foreground/native_test.rs +++ b/packages/runtime-core/src/provider/graph/foreground/native_test.rs @@ -542,5 +542,10 @@ mod dependency_rebind; mod dependency_slots; mod startup_cancellation; +mod absent_publication_recovery; +mod host_pin_recovery; +#[cfg(feature = "native-http-probe")] +mod retired_rebind_history; mod retired_recovery_cleanup; mod same_boot_recovery; +mod source_device_rebind; diff --git a/packages/runtime-core/src/provider/graph/foreground/native_test/absent_publication_recovery.rs b/packages/runtime-core/src/provider/graph/foreground/native_test/absent_publication_recovery.rs new file mode 100644 index 000000000..9c4061fea --- /dev/null +++ b/packages/runtime-core/src/provider/graph/foreground/native_test/absent_publication_recovery.rs @@ -0,0 +1,587 @@ +//! Owned capacity-two VM qualification for absent post-host-reboot publication. +//! The host reboot evidence is explicitly simulated in this isolated fixture; +//! this does not prove physical-volume continuity on a real reboot. +use super::dependency_rebind::{checked_cli, exec, snapshot}; +use super::*; +use crate::provider::{lifecycle, state::Owner}; +use std::{ + io::Write, + os::unix::fs::{MetadataExt, OpenOptionsExt, PermissionsExt}, +}; + +pub(super) fn ready(owner: &mut Process, run: &str, deadline: Instant) { + loop { + if let Some(status) = owner.poll() { + let error: Option = serde_json::from_slice(&owner.err).ok(); + let code = error + .as_ref() + .and_then(|value| value["code"].as_str()) + .unwrap_or("unstructured"); + let cause_code = error + .as_ref() + .and_then(|value| value["cause_code"].as_str()) + .unwrap_or("none"); + panic!( + "foreground owner exited before ready: status={:?}, code={code}, cause_code={cause_code}", + status.code() + ); + } + if let Some(end) = owner.out.iter().position(|byte| *byte == b'\n') { + let value: Value = serde_json::from_slice(&owner.out[..end]).unwrap(); + assert_eq!(value["kind"], "graph_foreground_ready"); + assert_eq!(value["run"], run); + return; + } + assert!(Instant::now() < deadline, "foreground readiness deadline"); + std::thread::sleep(Duration::from_millis(10)); + } +} + +pub(super) fn refused_cli( + binary: &Path, + candidate: &Candidate, + args: &[&str], + expected_code: &str, + deadline: Instant, +) { + let mut process = Process::start(binary, candidate, args, None); + let status = process.wait(deadline); + assert_eq!( + status.code(), + Some(2), + "expected a CLI refusal, not a crash" + ); + assert!(process.out.is_empty(), "refusal emitted a result"); + let error: Value = serde_json::from_slice(&process.err).expect("structured CLI error"); + assert_eq!(error["code"], expected_code); +} + +pub(super) fn inspect_args<'a>(run: &'a str, old: &'a str, prior: &'a str) -> [&'a str; 9] { + [ + "graph", + "inspect-absent-publication-cleanup", + "--run-id", + run, + "--original-owner-file", + old, + "--host-inspection-file", + prior, + "--json", + ] +} +pub(super) fn recover_args<'a>( + run: &'a str, + old: &'a str, + prior: &'a str, + hash: &'a str, +) -> [&'a str; 13] { + [ + "graph", + "recover-absent-publication-cleanup", + "--run-id", + run, + "--original-owner-file", + old, + "--host-inspection-file", + prior, + "--expect-selection", + hash, + "--retain-data", + "--accept-unpinned-post-reboot", + "--json", + ] +} + +/// Scoped fixture file mutation restores only bytes it wrote, including on a +/// panic. It never replaces managed provider or graph state. +struct PrivateBytes<'a> { + path: &'a Path, + original: Vec, + substituted: Vec, +} +impl<'a> PrivateBytes<'a> { + fn replace(path: &'a Path, substituted: Vec) -> Self { + let original = fs::read(path).unwrap(); + fs::write(path, &substituted).unwrap(); + Self { + path, + original, + substituted, + } + } +} +impl Drop for PrivateBytes<'_> { + fn drop(&mut self) { + if fs::read(self.path).ok().as_deref() == Some(self.substituted.as_slice()) { + fs::write(self.path, &self.original).unwrap(); + } + } +} + +struct Moved { + original: PathBuf, + moved: PathBuf, +} +impl Moved { + fn new(original: PathBuf) -> Self { + let moved = original.with_extension("absent-fixture-held"); + assert!(!moved.exists()); + fs::rename(&original, &moved).unwrap(); + Self { original, moved } + } +} +impl Drop for Moved { + fn drop(&mut self) { + if self.moved.exists() { + assert!( + !self.original.exists(), + "foreign path replaced selected fixture" + ); + fs::rename(&self.moved, &self.original).unwrap(); + } + } +} + +struct FaultChild(Option); +impl Drop for FaultChild { + fn drop(&mut self) { + if let Some(mut child) = self.0.take() { + let _ = child.kill(); + let _ = child.wait(); + } + } +} + +#[test] +#[ignore = "Only the parent owned VM fixture starts this exact fault helper"] +fn absent_publication_after_bridge_fault_child() { + let candidate = + Candidate::discover(Path::new(&std::env::var("HACK_LOCAL_TEST_ROOT").unwrap())).unwrap(); + graph::absent_publication_cleanup::recover( + &candidate, + &std::env::var("HACK_LOCAL_GRAPH_RUN").unwrap(), + &std::env::var("HACK_LOCAL_GRAPH_SELECTION").unwrap(), + Path::new(&std::env::var("HACK_LOCAL_GRAPH_OLD_OWNER").unwrap()), + Path::new(&std::env::var("HACK_LOCAL_GRAPH_INSPECTION").unwrap()), + ) + .unwrap(); +} + +#[test] +#[ignore = "Owned capacity-two native VM, pinned image and external 300s watchdog required"] +fn explicit_absent_publication_cleanup_retains_selected_volume_and_sibling() { + let deadline = Instant::now() + Duration::from_secs(270); + let candidate = + Candidate::discover(Path::new(&std::env::var("HACK_LOCAL_TEST_ROOT").unwrap())).unwrap(); + let binary = PathBuf::from(std::env::var("HACK_LOCAL_TEST_BINARY").unwrap()); + let image = std::env::var("HACK_LOCAL_TEST_IMAGE").unwrap(); + let fixtures = [graph::tests::Fixture::new(), graph::tests::Fixture::new()]; + let runs = [ + graph::probes::token().unwrap(), + graph::probes::token().unwrap(), + ]; + let private = graph::tests::Fixture::new(); + let mut plans = Vec::new(); + let mut selections = Vec::new(); + let mut dependencies = Vec::new(); + for (index, fixture) in fixtures.iter().enumerate() { + state::write(&fixture.0.join("compose.yaml"), &json!({ + "services":{"web":{"image":image,"read_only":true,"network_mode":"none", + "init":true,"user":"0:0","entrypoint":["/bin/sleep","300"],"command":[], + "volumes":["data:/data"],"healthcheck":{"test":["CMD","/bin/hack-graph-startup-app","complete"], + "interval":"200ms","timeout":"2s","retries":10,"start_period":"500ms"}}}, + "volumes":{"data":{}} + })).unwrap(); + let review = project::plan( + &candidate, + project::PlanOptions { + branch: None, + project: &fixture.0, + compose_file: Path::new("compose.yaml"), + profiles: &[], + }, + ) + .unwrap(); + let selection = private.0.join(format!("dependencies-{index}.json")); + state::write(&selection, &json!({"version":1,"plan":review.plan_id, + "artifact":"/tmp/unused-control-only-artifact","artifact_sha256":"a".repeat(64),"dependencies":[]})).unwrap(); + let dependency = checked_cli( + "dependency-plan", + &binary, + &candidate, + &[ + "graph", + "dependency-plan", + "--dependencies", + selection.to_str().unwrap(), + "--json", + ], + deadline, + ); + plans.push(review); + selections.push(selection); + dependencies.push(dependency); + } + let start = |index: usize| { + let normalized = fixtures[index].0.join("compose.yaml"); + Process::start( + &binary, + &candidate, + &[ + "graph", + "serve", + "--project", + fixtures[index].0.to_str().unwrap(), + "--file", + "compose.yaml", + "--expect-plan", + &plans[index].plan_id, + "--run-id", + &runs[index], + "--ready", + "web=healthy", + "--timeout-seconds", + "90", + "--dependencies", + selections[index].to_str().unwrap(), + "--expect-dependencies", + dependencies[index]["dependency_plan_id"].as_str().unwrap(), + "--normalized-file", + normalized.to_str().unwrap(), + "--expect-original", + &plans[index].plan.compose_sha256, + "--expect-namespace", + &plans[index].plan.namespace, + "--json", + ], + None, + ) + }; + let mut selected_owner = start(0); + let mut selected_cleanup = Cleanup { + binary: &binary, + candidate: &candidate, + run: &runs[0], + done: false, + }; + ready(&mut selected_owner, &runs[0], deadline); + assert_eq!( + exec(&binary, &candidate, &runs[0], "web", "write-data", deadline)["exit_code"], + 0 + ); + let original = snapshot(&candidate, &runs[0], deadline).receipt; + assert_eq!(original.phase, "ready-observed"); + let old_owner = Owner::load(&candidate).unwrap(); + let old_boot = old_owner.guest_boot_id.clone().unwrap(); + selected_owner.child.kill().unwrap(); + selected_owner.wait(deadline); + assert_eq!(lifecycle::down(&candidate).unwrap().phase, "stopped"); + let new_boot = lifecycle::up(&candidate).unwrap().guest_boot_id.unwrap(); + assert_ne!(old_boot, new_boot); + + let mut sibling_owner = start(1); + let mut sibling_cleanup = Cleanup { + binary: &binary, + candidate: &candidate, + run: &runs[1], + done: false, + }; + ready(&mut sibling_owner, &runs[1], deadline); + assert_eq!( + exec(&binary, &candidate, &runs[1], "web", "write-data", deadline)["exit_code"], + 0 + ); + let sibling = serde_json::to_vec(&snapshot(&candidate, &runs[1], deadline).receipt).unwrap(); + + let current = Owner::load(&candidate).unwrap(); + let boot = lifecycle::host_filesystem::host_boot_micros().unwrap(); + let new_device = current.storage.as_ref().unwrap().device; + let old_device = new_device.checked_add(1).unwrap(); + let mut synthetic = old_owner.clone(); + synthetic.storage.as_mut().unwrap().device = old_device; + synthetic.overlay.as_mut().unwrap().device = old_device; + if let Some(share) = synthetic.project_share.as_mut() { + share.device = old_device; + } + let prior_process = synthetic.process.as_mut().unwrap(); + prior_process.pid = i32::MAX; + prior_process.start_micros = boot - 1; + let old_path = private.0.join("pre-host-boot-owner.json"); + state::write(&old_path, &synthetic).unwrap(); + let old_bytes = fs::read(&old_path).unwrap(); + let pool = fs::symlink_metadata(candidate.state_root.join("run/smolvm")).unwrap(); + let inspection_path = private.0.join("pre-migration-inspection.json"); + state::write( + &inspection_path, + &serde_json::from_slice::(&graph::absent_publication_cleanup::fixture_inspection( + &old_bytes, + ¤t, + old_device, + boot, + pool.ino(), + )) + .unwrap(), + ) + .unwrap(); + let graph_root = graph::directory(&candidate, &runs[0]).unwrap(); + let raw_graph = fs::read(graph_root.join("state.json")).unwrap(); + let provider_owner_path = candidate.state_root.join("run/smolvm/owner.json"); + let raw_provider_owner = fs::read(&provider_owner_path).unwrap(); + let publisher_root = transport::root(&candidate, &runs[0]).unwrap(); + let control_root = original + .relay_startup + .as_ref() + .unwrap() + .control_root + .clone(); + assert!(publisher_root.exists() && control_root.exists()); + fs::remove_dir_all(&publisher_root).unwrap(); + fs::remove_dir_all(&control_root).unwrap(); + + let old = old_path.to_str().unwrap(); + let original_inspection = inspection_path.to_str().unwrap(); + let original_selection = checked_cli( + "inspect-private-original-inspection", + &binary, + &candidate, + &inspect_args(&runs[0], old, original_inspection), + deadline, + ); + let original_selection_hash = original_selection["selection_sha256"].as_str().unwrap(); + assert_eq!(original_selection_hash.len(), 64); + let inspection_bytes = fs::read(&inspection_path).unwrap(); + fs::set_permissions(&inspection_path, fs::Permissions::from_mode(0o644)).unwrap(); + assert_eq!( + fs::symlink_metadata(&inspection_path).unwrap().mode() & 0o7777, + 0o644 + ); + refused_cli( + &binary, + &candidate, + &inspect_args(&runs[0], old, original_inspection), + "graph_absent_publication_recovery", + deadline, + ); + assert!( + !publisher_root.exists(), + "read-only refusal did not reserve a publisher" + ); + assert_eq!(fs::read(&provider_owner_path).unwrap(), raw_provider_owner); + assert_eq!(fs::read(graph_root.join("state.json")).unwrap(), raw_graph); + + // A new private path retains the exact inspected bytes, while its path is + // deliberately part of the fresh selection hash. The historical 0644 + // receipt remains untouched after the copy. + let private_inspection = private.0.join("private-inspection-copy.json"); + let mut copy = fs::OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .open(&private_inspection) + .unwrap(); + copy.write_all(&inspection_bytes).unwrap(); + copy.sync_all().unwrap(); + fs::File::open(&private.0).unwrap().sync_all().unwrap(); + assert_eq!(fs::read(&private_inspection).unwrap(), inspection_bytes); + assert_eq!( + fs::symlink_metadata(&private_inspection).unwrap().mode() & 0o7777, + 0o600 + ); + assert_eq!( + fs::symlink_metadata(&inspection_path).unwrap().mode() & 0o7777, + 0o644 + ); + let prior = private_inspection.to_str().unwrap(); + let initial = checked_cli( + "inspect-exact-private-inspection-copy", + &binary, + &candidate, + &inspect_args(&runs[0], old, prior), + deadline, + ); + let selection = initial["selection_sha256"].as_str().unwrap(); + assert_eq!(selection.len(), 64); + assert_ne!( + selection, original_selection_hash, + "selection binds the new canonical path" + ); + assert_eq!(fs::read(&provider_owner_path).unwrap(), raw_provider_owner); + assert_eq!(fs::read(graph_root.join("state.json")).unwrap(), raw_graph); + let intent_path = graph_root.join("absent-publication-cleanup.json"); + assert!(!intent_path.exists()); + refused_cli( + &binary, + &candidate, + &recover_args(&runs[0], old, prior, original_selection_hash), + "graph_absent_publication_recovery", + deadline, + ); + refused_cli( + &binary, + &candidate, + &recover_args(&runs[0], old, prior, &"0".repeat(64)), + "graph_absent_publication_recovery", + deadline, + ); + assert!(!intent_path.exists()); + assert_eq!(fs::read(graph_root.join("state.json")).unwrap(), raw_graph); + let refreshed = checked_cli( + "inspect-after-stale-selection", + &binary, + &candidate, + &inspect_args(&runs[0], old, prior), + deadline, + ); + assert_eq!(refreshed["selection_sha256"], selection); + { + let mut whitespace = old_bytes.clone(); + whitespace.push(b' '); + let _changed = PrivateBytes::replace(&old_path, whitespace); + refused_cli( + &binary, + &candidate, + &recover_args(&runs[0], old, prior, selection), + "graph_absent_publication_recovery", + deadline, + ); + assert!(!intent_path.exists()); + } + { + let _missing = Moved::new(private_inspection.clone()); + refused_cli( + &binary, + &candidate, + &recover_args(&runs[0], old, prior, selection), + "graph_absent_publication_recovery", + deadline, + ); + assert!(!intent_path.exists()); + } + { + let state_path = graph_root.join("state.json"); + let mut changed = raw_graph.clone(); + changed.push(b' '); + let _changed = PrivateBytes::replace(&state_path, changed); + refused_cli( + &binary, + &candidate, + &recover_args(&runs[0], old, prior, selection), + "graph_absent_publication_recovery", + deadline, + ); + assert!(!intent_path.exists()); + } + { + let lock = publisher_root.join("operation.lock"); + let _original = Moved::new(lock.clone()); + fs::write(&lock, b"substituted lock path").unwrap(); + refused_cli( + &binary, + &candidate, + &recover_args(&runs[0], old, prior, selection), + "graph_absent_publication_recovery", + deadline, + ); + assert!(!intent_path.exists()); + fs::remove_file(&lock).unwrap(); + } + let child = Command::new(std::env::current_exe().unwrap()) + .args(["--ignored", "--exact", + "provider::graph::foreground::native_test::absent_publication_recovery::absent_publication_after_bridge_fault_child"]) + .env("HACK_LOCAL_TEST_ROOT", std::env::var("HACK_LOCAL_TEST_ROOT").unwrap()) + .env("HACK_LOCAL_GRAPH_RUN", &runs[0]) + .env("HACK_LOCAL_GRAPH_SELECTION", selection) + .env("HACK_LOCAL_GRAPH_OLD_OWNER", old) + .env("HACK_LOCAL_GRAPH_INSPECTION", prior) + .env("HACK_LOCAL_GRAPH_FAULT", "absent-after-bridge-release") + .stdin(Stdio::null()).stdout(Stdio::null()).stderr(Stdio::null()) + .spawn().unwrap(); + let mut fault = FaultChild(Some(child)); + let marker = graph_root.join("fault-absent-after-bridge-release.json"); + let fault_deadline = Instant::now() + Duration::from_secs(25); + while !marker.exists() { + assert!( + fault.0.as_mut().unwrap().try_wait().unwrap().is_none(), + "fault child exited before selected boundary" + ); + assert!( + Instant::now() < fault_deadline, + "selected fault boundary deadline" + ); + std::thread::sleep(Duration::from_millis(20)); + } + assert_eq!(state::read::(&marker).unwrap()["run"], runs[0]); + assert!( + intent_path.exists(), + "absence intent was durable before bridge release" + ); + assert_eq!( + state::read::(&graph_root.join("state.json")) + .unwrap() + .phase, + "cleanup-intent", + "graph intent was durable before bridge release" + ); + assert!( + graph::absent_publication_cleanup::publication_allowed(&candidate, &runs[0], false) + .is_err() + ); + assert!( + graph::absent_publication_cleanup::publication_allowed(&candidate, &runs[0], true).is_err() + ); + fault.0.as_mut().unwrap().kill().unwrap(); + fault.0.as_mut().unwrap().wait().unwrap(); + fault.0 = None; + let result = checked_cli( + "recover-absent-publications", + &binary, + &candidate, + &recover_args(&runs[0], old, prior, selection), + deadline, + ); + assert_eq!(result["phase"], "stopped-data-retained"); + assert_eq!(result["publisher_retired"], true); + let stopped = snapshot(&candidate, &runs[0], deadline).receipt; + assert_eq!( + stopped.resources["volume:data"].name, + original.resources["volume:data"].name + ); + let repeated = checked_cli( + "idempotent-absent-retirement", + &binary, + &candidate, + &recover_args(&runs[0], old, prior, selection), + deadline, + ); + assert_eq!(repeated["publisher_retired"], true); + assert!(sibling_owner.poll().is_none()); + assert_eq!( + serde_json::to_vec(&snapshot(&candidate, &runs[1], deadline).receipt).unwrap(), + sibling + ); + assert_eq!( + exec(&binary, &candidate, &runs[1], "web", "read-data", deadline)["exit_code"], + 0 + ); + + // Keep the selected retained volume for the later source-continuity + // fixture. The external harness tears down this isolated test pool. + selected_cleanup.done = true; + let removed = checked_cli( + "remove-sibling-data", + &binary, + &candidate, + &[ + "graph", + "cleanup", + "--run-id", + &runs[1], + "--remove-data", + "--json", + ], + deadline, + ); + assert_eq!(removed["phase"], "removed"); + assert!(sibling_owner.wait(deadline).success()); + sibling_cleanup.done = true; +} diff --git a/packages/runtime-core/src/provider/graph/foreground/native_test/host_pin_recovery.rs b/packages/runtime-core/src/provider/graph/foreground/native_test/host_pin_recovery.rs new file mode 100644 index 000000000..1ce87e513 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/foreground/native_test/host_pin_recovery.rs @@ -0,0 +1,537 @@ +//! Real CLI path for explicitly witnessed legacy host-pin cleanup. An isolated +//! test VM supplies the guest; only fixture receipt device numbers are changed. +use super::dependency_rebind::{checked_cli, exec, snapshot}; +use super::*; +use std::{ + io::Write, + os::unix::fs::{MetadataExt, OpenOptionsExt}, +}; + +fn ready(owner: &mut Process, run: &str, deadline: Instant) { + loop { + assert!( + owner.poll().is_none(), + "foreground owner exited before ready" + ); + if let Some(end) = owner.out.iter().position(|byte| *byte == b'\n') { + let value: Value = serde_json::from_slice(&owner.out[..end]).unwrap(); + assert_eq!(value["kind"], "graph_foreground_ready"); + assert_eq!(value["run"], run); + return; + } + assert!(Instant::now() < deadline, "foreground readiness deadline"); + std::thread::sleep(Duration::from_millis(10)); + } +} + +fn legacy_pin(path: &Path, old: u64, before_boot: u64, control: bool) -> Vec { + let mut value: Value = state::read(path).unwrap(); + for key in if control { + ["parent", "endpoint"] + } else { + ["parent", "socket"] + } { + let pair = value[key].as_array_mut().unwrap(); + assert_ne!(pair[0].as_u64().unwrap(), old); + pair[0] = json!(old); + } + // This is a synthetic previous-physical-boot receipt. The actual owner + // exited before the test changes its private fixture bytes. + value["process"]["pid"] = json!(i32::MAX); + value["process"]["start_micros"] = json!(before_boot - 1); + state::write(path, &value).unwrap(); + fs::read(path).unwrap() +} + +fn refused_cli( + binary: &Path, + candidate: &Candidate, + args: &[&str], + expected_code: &str, + deadline: Instant, +) { + let mut process = Process::start(binary, candidate, args, None); + let status = process.wait(deadline); + assert_eq!( + status.code(), + Some(2), + "expected a CLI refusal, not a crash" + ); + assert!(process.out.is_empty(), "refused action emitted a result"); + let error: Value = serde_json::from_slice(&process.err).expect("structured CLI error"); + assert_eq!(error["code"], expected_code); + assert!(error["message"].as_str().is_some_and(|v| !v.is_empty())); +} + +fn publish_args<'a>(run: &'a str, expected: &'a str) -> [&'a str; 8] { + [ + "graph", + "recover-host-pins", + "--run-id", + run, + "--expect-selection", + expected, + "--accept-legacy-device-rebind", + "--json", + ] +} + +struct TemporarilyMoved { + original: PathBuf, + moved: PathBuf, +} +impl TemporarilyMoved { + fn new(original: PathBuf) -> Self { + let moved = original.with_extension("host-pin-test-held"); + assert!(!moved.exists()); + fs::rename(&original, &moved).unwrap(); + Self { original, moved } + } +} +impl Drop for TemporarilyMoved { + fn drop(&mut self) { + if self.moved.exists() { + assert!(!self.original.exists(), "foreign replacement at held path"); + fs::rename(&self.moved, &self.original).unwrap(); + } + } +} + +/// Preserve exact pre-mutation bytes outside managed state. A killed test can +/// be recovered from this private fixture artifact; normal unwinding restores +/// only the bytes this test itself wrote, never a foreign replacement. +struct ProviderReceiptRestore { + path: PathBuf, + original: Vec, + current: Vec, +} +impl ProviderReceiptRestore { + fn new(candidate: &Candidate, path: PathBuf) -> Self { + let original = fs::read(&path).unwrap(); + let backup = candidate + .checkout + .join("host-pin-test-provider-owner-original.json"); + let mut artifact = fs::OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .open(&backup) + .unwrap(); + artifact.write_all(&original).unwrap(); + artifact.sync_all().unwrap(); + fs::File::open(&candidate.checkout) + .unwrap() + .sync_all() + .unwrap(); + Self { + path, + current: original.clone(), + original, + } + } + fn replace(&mut self, bytes: Vec) { + assert_eq!(fs::read(&self.path).unwrap(), self.current); + fs::write(&self.path, &bytes).unwrap(); + self.current = bytes; + } + fn restore(&mut self) { + assert_eq!(fs::read(&self.path).unwrap(), self.current); + fs::write(&self.path, &self.original).unwrap(); + assert_eq!(fs::read(&self.path).unwrap(), self.original); + self.current = self.original.clone(); + } +} +impl Drop for ProviderReceiptRestore { + fn drop(&mut self) { + if self.current != self.original + && fs::read(&self.path).ok().as_deref() == Some(self.current.as_slice()) + { + let _ = fs::write(&self.path, &self.original); + } + } +} + +#[test] +#[ignore = "Owned capacity-two native VM, pinned static image and external 300s watchdog required"] +fn explicit_host_pin_witness_recovers_only_selected_previous_boot_run() { + let deadline = Instant::now() + Duration::from_secs(270); + let candidate = + Candidate::discover(Path::new(&std::env::var("HACK_LOCAL_TEST_ROOT").unwrap())).unwrap(); + let binary = PathBuf::from(std::env::var("HACK_LOCAL_TEST_BINARY").unwrap()); + let image = std::env::var("HACK_LOCAL_TEST_IMAGE").unwrap(); + let fixtures = [graph::tests::Fixture::new(), graph::tests::Fixture::new()]; + let runs = [ + graph::probes::token().unwrap(), + graph::probes::token().unwrap(), + ]; + let selection_root = graph::tests::Fixture::new(); + let mut plans = Vec::new(); + let mut selections = Vec::new(); + let mut dependencies = Vec::new(); + for (index, fixture) in fixtures.iter().enumerate() { + state::write(&fixture.0.join("compose.yaml"), &json!({ + "services":{"web":{"image":image,"read_only":true,"network_mode":"none", + "init":true,"user":"0:0","entrypoint":["/bin/sleep","300"],"command":[], + "volumes":["data:/data"],"healthcheck":{"test":["CMD","/bin/hack-graph-startup-app","complete"], + "interval":"200ms","timeout":"2s","retries":10,"start_period":"500ms"}}}, + "volumes":{"data":{}} + })).unwrap(); + let review = project::plan( + &candidate, + project::PlanOptions { + branch: None, + project: &fixture.0, + compose_file: Path::new("compose.yaml"), + profiles: &[], + }, + ) + .unwrap(); + let selection = selection_root.0.join(format!("dependencies-{index}.json")); + state::write(&selection, &json!({"version":1,"plan":review.plan_id, + "artifact":"/tmp/unused-control-only-artifact","artifact_sha256":"a".repeat(64),"dependencies":[]})).unwrap(); + let dependency = checked_cli( + "dependency-plan", + &binary, + &candidate, + &[ + "graph", + "dependency-plan", + "--dependencies", + selection.to_str().unwrap(), + "--json", + ], + deadline, + ); + plans.push(review); + selections.push(selection); + dependencies.push(dependency); + } + let start = |index: usize, generation: Option<&str>| { + let normalized = fixtures[index].0.join("compose.yaml"); + let mut args = vec![ + "graph", + if generation.is_some() { + "serve-restore" + } else { + "serve" + }, + "--project", + fixtures[index].0.to_str().unwrap(), + "--file", + "compose.yaml", + "--expect-plan", + &plans[index].plan_id, + "--run-id", + &runs[index], + "--ready", + "web=healthy", + "--timeout-seconds", + "90", + "--dependencies", + selections[index].to_str().unwrap(), + "--expect-dependencies", + dependencies[index]["dependency_plan_id"].as_str().unwrap(), + "--normalized-file", + normalized.to_str().unwrap(), + "--expect-original", + &plans[index].plan.compose_sha256, + "--expect-namespace", + &plans[index].plan.namespace, + "--json", + ]; + if let Some(generation) = generation { + args.extend(["--expect-generation", generation]); + } + Process::start(&binary, &candidate, &args, None) + }; + let mut selected_owner = start(0, None); + let mut selected_cleanup = Cleanup { + binary: &binary, + candidate: &candidate, + run: &runs[0], + done: false, + }; + ready(&mut selected_owner, &runs[0], deadline); + assert_eq!( + exec(&binary, &candidate, &runs[0], "web", "write-data", deadline)["exit_code"], + 0 + ); + let original = snapshot(&candidate, &runs[0], deadline).receipt; + assert_eq!(original.phase, "ready-observed"); + let expected_receipt = format!( + "{:x}", + Sha256::digest(serde_json::to_vec_pretty(&original).unwrap()) + ); + let old_boot = crate::provider::lifecycle::status(&candidate) + .unwrap() + .guest_boot_id + .unwrap(); + selected_owner.child.kill().unwrap(); + selected_owner.wait(deadline); + assert_eq!( + crate::provider::lifecycle::down(&candidate).unwrap().phase, + "stopped" + ); + let new_boot = crate::provider::lifecycle::up(&candidate) + .unwrap() + .guest_boot_id + .unwrap(); + assert_ne!(old_boot, new_boot); + + let mut sibling_owner = start(1, None); + let mut sibling_cleanup = Cleanup { + binary: &binary, + candidate: &candidate, + run: &runs[1], + done: false, + }; + ready(&mut sibling_owner, &runs[1], deadline); + assert_eq!( + exec(&binary, &candidate, &runs[1], "web", "write-data", deadline)["exit_code"], + 0 + ); + let sibling = serde_json::to_vec(&snapshot(&candidate, &runs[1], deadline).receipt).unwrap(); + + let root = graph::directory(&candidate, &runs[0]).unwrap(); + let graph_bytes = fs::read(root.join("state.json")).unwrap(); + let startup = original.relay_startup.as_ref().unwrap(); + let publisher_root = transport::root(&candidate, &runs[0]).unwrap(); + let publisher = publisher_root.join("owner.json"); + let control = startup.control_root.join("relay-control/owner.json"); + let current = fs::symlink_metadata(&publisher_root).unwrap().dev(); + let old = current.checked_add(1).unwrap(); + let before_boot = crate::provider::lifecycle::host_filesystem::host_boot_micros().unwrap(); + let publisher_bytes = legacy_pin(&publisher, old, before_boot, false); + let control_bytes = legacy_pin(&control, old, before_boot, true); + + // Ordinary recovery remains strict; the new witness is required to use + // the legacy device-only pins and is scoped to the selected run. + assert!(graph::foreground::transport::Pin::load(&candidate, &runs[0]).is_err()); + let inspected = checked_cli( + "inspect-host-pins", + &binary, + &candidate, + &[ + "graph", + "inspect-host-pin-recovery", + "--run-id", + &runs[0], + "--json", + ], + deadline, + ); + let selection = inspected["selection_sha256"].as_str().unwrap(); + assert_eq!(selection.len(), 64); + assert_eq!(inspected["old_device"], old); + assert_eq!(inspected["new_device"], current); + assert_eq!(fs::read(root.join("state.json")).unwrap(), graph_bytes); + assert_eq!(fs::read(&publisher).unwrap(), publisher_bytes); + assert_eq!(fs::read(&control).unwrap(), control_bytes); + + let witness_path = root.join("host-pin-recovery.json"); + let inspect_args = [ + "graph", + "inspect-host-pin-recovery", + "--run-id", + &runs[0], + "--json", + ]; + // A stale selection must not publish a witness or rewrite any source + // receipt. The three temporary path removals also preserve exact bytes. + refused_cli( + &binary, + &candidate, + &publish_args(&runs[0], &"0".repeat(64)), + "graph_host_pin_recovery", + deadline, + ); + assert!(!witness_path.exists()); + assert_eq!(fs::read(root.join("state.json")).unwrap(), graph_bytes); + assert_eq!(fs::read(&publisher).unwrap(), publisher_bytes); + assert_eq!(fs::read(&control).unwrap(), control_bytes); + { + let _missing = TemporarilyMoved::new(publisher_root.clone()); + refused_cli( + &binary, + &candidate, + &inspect_args, + "provider_state", + deadline, + ); + assert!(!witness_path.exists()); + } + assert_eq!(fs::read(&publisher).unwrap(), publisher_bytes); + { + let _missing = TemporarilyMoved::new(startup.control_root.join("relay-control")); + refused_cli( + &binary, + &candidate, + &inspect_args, + "provider_state", + deadline, + ); + assert!(!witness_path.exists()); + } + assert_eq!(fs::read(&control).unwrap(), control_bytes); + + // The provider receipt is valid JSON throughout. A changed raw byte + // selection and a changed guest boot must both refuse without writes. + let provider_path = candidate.state_root.join("run/smolvm/owner.json"); + let mut owner_restore = ProviderReceiptRestore::new(&candidate, provider_path.clone()); + let provider_bytes = owner_restore.original.clone(); + let mut whitespace = provider_bytes.clone(); + whitespace.push(b' '); + owner_restore.replace(whitespace.clone()); + refused_cli( + &binary, + &candidate, + &publish_args(&runs[0], selection), + "graph_host_pin_recovery", + deadline, + ); + assert_eq!(fs::read(&provider_path).unwrap(), whitespace); + assert!(!witness_path.exists()); + owner_restore.restore(); + let mut changed_boot: Value = serde_json::from_slice(&provider_bytes).unwrap(); + changed_boot["guest_boot_id"] = json!("f".repeat(32)); + assert_ne!(changed_boot["guest_boot_id"], new_boot); + owner_restore.replace(serde_json::to_vec_pretty(&changed_boot).unwrap()); + let boot_bytes = fs::read(&provider_path).unwrap(); + refused_cli( + &binary, + &candidate, + &inspect_args, + "graph_host_pin_recovery", + deadline, + ); + assert_eq!(fs::read(&provider_path).unwrap(), boot_bytes); + assert!(!witness_path.exists()); + owner_restore.restore(); + assert_eq!(fs::read(root.join("state.json")).unwrap(), graph_bytes); + assert_eq!(fs::read(&publisher).unwrap(), publisher_bytes); + assert_eq!(fs::read(&control).unwrap(), control_bytes); + + let refreshed = checked_cli( + "inspect-after-negative-controls", + &binary, + &candidate, + &inspect_args, + deadline, + ); + let selection = refreshed["selection_sha256"].as_str().unwrap(); + let recovered = checked_cli( + "publish-host-pin-witness", + &binary, + &candidate, + &publish_args(&runs[0], selection), + deadline, + ); + assert_eq!(recovered["witness_published"], true); + assert_eq!(fs::read(root.join("state.json")).unwrap(), graph_bytes); + assert_eq!(fs::read(&publisher).unwrap(), publisher_bytes); + assert_eq!(fs::read(&control).unwrap(), control_bytes); + assert_eq!( + serde_json::to_vec(&snapshot(&candidate, &runs[1], deadline).receipt).unwrap(), + sibling + ); + + let cleanup = checked_cli( + "legacy-dead-owner-cleanup", + &binary, + &candidate, + &[ + "graph", + "recover-cleanup", + "--run-id", + &runs[0], + "--expect-receipt", + &expected_receipt, + "--json", + ], + deadline, + ); + assert_eq!(cleanup["phase"], "stopped-data-retained"); + let retained = snapshot(&candidate, &runs[0], deadline).receipt; + assert_eq!( + retained.resources["volume:data"].name, + original.resources["volume:data"].name + ); + for _ in 0..2 { + let retired = checked_cli( + "legacy-publisher-retirement", + &binary, + &candidate, + &[ + "graph", + "retire-recovered-publisher", + "--run-id", + &runs[0], + "--expect-owner", + &retained.owner, + "--json", + ], + deadline, + ); + assert_eq!(retired["publisher_retired"], true); + } + let restoration = graph::foreground::restore_selection(&candidate, &runs[0]).unwrap(); + selected_owner = start(0, Some(restoration["generation"].as_str().unwrap())); + ready(&mut selected_owner, &runs[0], deadline); + assert_eq!( + exec(&binary, &candidate, &runs[0], "web", "read-data", deadline)["exit_code"], + 0, + "selected marker must survive cleanup, retirement and same-run restore" + ); + let restored = snapshot(&candidate, &runs[0], deadline).receipt; + assert_eq!( + restored.resources["volume:data"].name, + original.resources["volume:data"].name + ); + assert_ne!( + restored.resources["container:web"].id, + original.resources["container:web"].id + ); + assert!(sibling_owner.poll().is_none()); + assert_eq!( + serde_json::to_vec(&snapshot(&candidate, &runs[1], deadline).receipt).unwrap(), + sibling + ); + assert_eq!( + exec(&binary, &candidate, &runs[1], "web", "read-data", deadline)["exit_code"], + 0 + ); + let selected_removed = checked_cli( + "remove-selected-data", + &binary, + &candidate, + &[ + "graph", + "cleanup", + "--run-id", + &runs[0], + "--remove-data", + "--json", + ], + deadline, + ); + assert_eq!(selected_removed["phase"], "removed"); + assert!(selected_owner.wait(deadline).success()); + selected_cleanup.done = true; + let removed = checked_cli( + "remove-sibling-data", + &binary, + &candidate, + &[ + "graph", + "cleanup", + "--run-id", + &runs[1], + "--remove-data", + "--json", + ], + deadline, + ); + assert_eq!(removed["phase"], "removed"); + assert!(sibling_owner.wait(deadline).success()); + sibling_cleanup.done = true; +} diff --git a/packages/runtime-core/src/provider/graph/foreground/native_test/retired_rebind_history.rs b/packages/runtime-core/src/provider/graph/foreground/native_test/retired_rebind_history.rs new file mode 100644 index 000000000..790996676 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/foreground/native_test/retired_rebind_history.rs @@ -0,0 +1,834 @@ +//! Cross-version owned regression for the historical dependency-rebind journal. +//! A pinned pre-archive binary creates the previously admitted state; this +//! source resumes its exact completed proof without fabricating a journal. +use super::absent_publication_recovery::{inspect_args, ready, recover_args}; +use super::dependency_rebind::{checked_cli, snapshot}; +use super::*; +use crate::provider::{graph::startup::native_test::RestartableBackend, lifecycle, state::Owner}; +use sha2::{Digest, Sha256}; +use std::os::unix::fs::MetadataExt; + +mod fixture; +use fixture::{Files, graceful_stop, ready_at}; + +fn file_hash(path: &Path) -> String { + let mut file = fs::File::open(path).unwrap(); + let mut digest = Sha256::new(); + let mut bytes = [0u8; 65536]; + loop { + let count = file.read(&mut bytes).unwrap(); + if count == 0 { + break; + } + digest.update(&bytes[..count]); + } + format!("{:x}", digest.finalize()) +} + +fn selection( + path: &Path, + plan: &str, + artifact: &Path, + artifact_hash: &str, + endpoint: &crate::provider::host_endpoint::HostEndpoint, + port: u16, +) { + let bindings = [ + json!({"service":"web","binding":"default","slot":0,"guest_port":25252, + "host_pid":endpoint.process_identity().pid,"host_port":port, + "host_executable":std::env::current_exe().unwrap()}), + ]; + state::write( + path, + &json!({"version":1,"plan":plan,"artifact":artifact, + "artifact_sha256":artifact_hash,"dependencies":bindings}), + ) + .unwrap(); +} + +fn marker( + binary: &Path, + candidate: &Candidate, + run: &str, + value: &str, + write: bool, + deadline: Instant, +) { + let value = serde_json::to_string(value).unwrap(); + let script = if write { + format!("await Bun.write('/data/retired-rebind-marker',{value});") + } else { + format!( + "if((await Bun.file('/data/retired-rebind-marker').text())!=={value})process.exit(91);" + ) + }; + let mut process = Process::start( + binary, + candidate, + &[ + "graph", + "exec", + "--run-id", + run, + "--service", + "web", + "--timeout-seconds", + "12", + "--json", + "--", + "/usr/local/bin/bun", + "-e", + &script, + ], + None, + ); + assert!( + process.wait(deadline).success(), + "owned marker exec refused" + ); + let result: Value = serde_json::from_slice(&process.out).unwrap(); + assert_eq!(result["exit_code"], 0, "owned marker check failed"); + assert_eq!(result["truncated"], false); +} + +#[test] +#[ignore = "Owned capacity-two VM, pinned prior/current binaries, image and relay artifact, external 300s watchdog required"] +fn completed_prior_boot_rebind_archives_after_newer_stopped_generation() { + let deadline = Instant::now() + Duration::from_secs(270); + let candidate = + Candidate::discover(Path::new(&std::env::var("HACK_LOCAL_TEST_ROOT").unwrap())).unwrap(); + let binary = PathBuf::from(std::env::var("HACK_LOCAL_TEST_BINARY").unwrap()); + let legacy = PathBuf::from(std::env::var("HACK_LOCAL_TEST_LEGACY_BINARY").unwrap()); + let legacy_hash = std::env::var("HACK_LOCAL_TEST_LEGACY_SHA256").unwrap(); + assert_eq!(legacy_hash.len(), 64); + assert_eq!( + file_hash(&legacy), + legacy_hash, + "prior binary bytes changed" + ); + assert_ne!( + file_hash(&binary), + legacy_hash, + "historical setup used current binary" + ); + let image = std::env::var("HACK_LOCAL_TEST_IMAGE").unwrap(); + let artifact = PathBuf::from(std::env::var("HACK_GRAPH_RELAY_ARTIFACT").unwrap()); + let artifact_hash = std::env::var("HACK_GRAPH_RELAY_SHA256").unwrap(); + assert_eq!(file_hash(&artifact), artifact_hash); + // Inputs stay under this declared private candidate on every failed path. + // There is deliberately no panic-time graph cleanup or directory Drop. + let files = Files::new(&candidate.checkout); + let selected = &files.selected; + let sibling = &files.sibling; + let private = &files.private; + state::write( + &selected.join("compose.yaml"), + &json!({ + "services":{ + "web":{"image":image,"read_only":true,"init":true,"user":"0:0", + "entrypoint":["/usr/local/bin/bun","-e", "Bun.serve({hostname:'0.0.0.0',port:3000,fetch(){return new Response('ready')}})"], + "command":[],"volumes":["data:/data"],"networks":["private"], + "healthcheck":{"x-hack-http":{"port":3000,"path":"/","interval_ms":100, + "timeout_ms":500,"retries":20,"start_period_ms":0}}} + },"networks":{"private":{"internal":true}},"volumes":{"data":{}} + }), + ) + .unwrap(); + state::write( + &sibling.join("compose.yaml"), + &json!({ + "services":{"web":{"image":image,"read_only":true,"network_mode":"none", + "init":true,"user":"0:0","entrypoint":["/usr/local/bin/bun","-e","setInterval(()=>{},1000)"],"command":[], + "volumes":["data:/data"],"healthcheck":{"test":["CMD", + "/usr/local/bin/bun","-e","process.exit(0)"],"interval":"200ms", + "timeout":"2s","retries":10,"start_period":"500ms"}}}, + "volumes":{"data":{}} + }), + ) + .unwrap(); + let selected_plan = project::plan( + &candidate, + project::PlanOptions { + branch: None, + project: selected, + compose_file: Path::new("compose.yaml"), + profiles: &[], + }, + ) + .unwrap(); + let sibling_plan = project::plan( + &candidate, + project::PlanOptions { + branch: None, + project: sibling, + compose_file: Path::new("compose.yaml"), + profiles: &[], + }, + ) + .unwrap(); + let run = graph::probes::token().unwrap(); + let sibling_run = graph::probes::token().unwrap(); + let selected_dependencies = private.join("selected-dependencies.json"); + let sibling_dependencies = private.join("sibling-dependencies.json"); + let selected_compose = selected.join("compose.yaml"); + state::write( + &sibling_dependencies, + &json!({"version":1,"plan":sibling_plan.plan_id, + "artifact":"/tmp/unused-control-only-artifact","artifact_sha256":"a".repeat(64), + "dependencies":[]}), + ) + .unwrap(); + let sibling_plan_id = checked_cli( + "sibling-dependency-plan", + &binary, + &candidate, + &[ + "graph", + "dependency-plan", + "--dependencies", + sibling_dependencies.to_str().unwrap(), + "--json", + ], + deadline, + )["dependency_plan_id"] + .as_str() + .unwrap() + .to_owned(); + let mut backend = RestartableBackend::start(0); + selection( + &selected_dependencies, + &selected_plan.plan_id, + &artifact, + &artifact_hash, + &backend.endpoint(), + backend.port, + ); + let mut selected_plan_id = checked_cli( + "selected-dependency-plan", + &legacy, + &candidate, + &[ + "graph", + "dependency-plan", + "--dependencies", + selected_dependencies.to_str().unwrap(), + "--json", + ], + deadline, + )["dependency_plan_id"] + .as_str() + .unwrap() + .to_owned(); + let start_selected = |binary: &Path, generation: Option<&str>, plan_id: &str| { + let mut args = vec![ + "graph", + if generation.is_some() { + "serve-restore" + } else { + "serve" + }, + "--project", + selected.to_str().unwrap(), + "--file", + "compose.yaml", + "--expect-plan", + selected_plan.plan_id.as_str(), + "--run-id", + run.as_str(), + "--ready", + "web=healthy", + "--timeout-seconds", + "90", + "--dependencies", + selected_dependencies.to_str().unwrap(), + "--expect-dependencies", + plan_id, + "--normalized-file", + selected_compose.to_str().unwrap(), + "--expect-original", + selected_plan.plan.compose_sha256.as_str(), + "--expect-namespace", + selected_plan.plan.namespace.as_str(), + "--json", + ]; + if let Some(generation) = generation { + args.extend(["--expect-generation", generation]); + } + Process::start(binary, &candidate, &args, None) + }; + let mut owner = start_selected(&legacy, None, &selected_plan_id); + ready_at("initial-legacy", &files, &mut owner, &run, deadline); + marker( + &legacy, + &candidate, + &run, + "selected-marker-v1", + true, + deadline, + ); + let before = snapshot(&candidate, &run, deadline).receipt; + assert!( + snapshot(&candidate, &run, deadline) + .guest_endpoints + .contains_key("web") + ); + let old_owner = Owner::load(&candidate).unwrap(); + let old_boot = old_owner.guest_boot_id.clone().unwrap(); + let port = backend.port; + backend.stop(); + backend = RestartableBackend::start(port); + checked_cli( + "real-listener-refresh", + &legacy, + &candidate, + &["graph", "refresh-dependencies", "--run-id", &run, "--json"], + deadline, + ); + let refreshed = snapshot(&candidate, &run, deadline).receipt; + let graph_root = graph::directory(&candidate, &run).unwrap(); + let journal_path = graph_root.join("dependency-rebind.json"); + let journal: Value = state::read(&journal_path).unwrap(); + assert_eq!(journal["phase"], "completed"); + assert_eq!(journal["boot"], old_boot); + assert_eq!( + journal["completed_generation"], + graph::service_exec_generation(&refreshed).unwrap() + ); + assert_eq!( + before.resources["volume:data"].name, + refreshed.resources["volume:data"].name + ); + let prior_process = crate::provider::identity::observe(owner.child.id() as i32).unwrap(); + owner.child.kill().unwrap(); + owner.wait(deadline); + assert_eq!(lifecycle::down(&candidate).unwrap().phase, "stopped"); + let new_boot = lifecycle::up(&candidate).unwrap().guest_boot_id.unwrap(); + assert_ne!(new_boot, old_boot); + let mut sibling_owner = Process::start( + &binary, + &candidate, + &[ + "graph", + "serve", + "--project", + sibling.to_str().unwrap(), + "--file", + "compose.yaml", + "--expect-plan", + &sibling_plan.plan_id, + "--run-id", + &sibling_run, + "--ready", + "web=healthy", + "--timeout-seconds", + "90", + "--dependencies", + sibling_dependencies.to_str().unwrap(), + "--expect-dependencies", + &sibling_plan_id, + "--normalized-file", + sibling.join("compose.yaml").to_str().unwrap(), + "--expect-original", + &sibling_plan.plan.compose_sha256, + "--expect-namespace", + &sibling_plan.plan.namespace, + "--json", + ], + None, + ); + ready_at( + "sibling", + &files, + &mut sibling_owner, + &sibling_run, + deadline, + ); + marker( + &binary, + &candidate, + &sibling_run, + "sibling-marker-v1", + true, + deadline, + ); + let sibling_before = snapshot(&candidate, &sibling_run, deadline).receipt; + let current_owner = Owner::load(&candidate).unwrap(); + let host_boot = lifecycle::host_filesystem::host_boot_micros().unwrap(); + let old_device = current_owner + .storage + .as_ref() + .unwrap() + .device + .checked_add(1) + .unwrap(); + // The real refresh leaves a current-host-boot reservation. Project its exact + // dead process timestamp and socket devices alongside the synthetic VM owner; + // retain the real journal, socket inodes, receipt, bindings and sibling claim. + graph::dependency_slots::fixture_prior_boot::synthesize( + &candidate, + &refreshed, + &prior_process, + crate::provider::host_pin::DeviceRebind { + old: old_device, + current: current_owner.storage.as_ref().unwrap().device, + }, + host_boot, + ) + .unwrap(); + let mut synthetic = old_owner; + synthetic.storage.as_mut().unwrap().device = old_device; + synthetic.overlay.as_mut().unwrap().device = old_device; + if let Some(share) = synthetic.project_share.as_mut() { + share.device = old_device; + } + synthetic.process.as_mut().unwrap().pid = i32::MAX; + synthetic.process.as_mut().unwrap().start_micros = host_boot - 1; + let old_path = private.join("prior-owner.json"); + state::write(&old_path, &synthetic).unwrap(); + let pool = fs::symlink_metadata(candidate.state_root.join("run/smolvm")).unwrap(); + let inspection_path = private.join("prior-inspection.json"); + state::write( + &inspection_path, + &serde_json::from_slice::(&graph::absent_publication_cleanup::fixture_inspection( + &fs::read(&old_path).unwrap(), + ¤t_owner, + old_device, + host_boot, + pool.ino(), + )) + .unwrap(), + ) + .unwrap(); + let publisher = transport::root(&candidate, &run).unwrap(); + let control = refreshed + .relay_startup + .as_ref() + .unwrap() + .control_root + .clone(); + fs::remove_dir_all(&publisher).unwrap(); + fs::remove_dir_all(&control).unwrap(); + let old = old_path.to_str().unwrap(); + let prior = inspection_path.to_str().unwrap(); + let inspected = checked_cli( + "inspect-absence-with-real-journal", + &legacy, + &candidate, + &inspect_args(&run, old, prior), + deadline, + ); + let selection_hash = inspected["selection_sha256"].as_str().unwrap(); + let stopped = checked_cli( + "legacy-absence-retirement", + &legacy, + &candidate, + &recover_args(&run, old, prior, selection_hash), + deadline, + ); + assert_eq!(stopped["phase"], "stopped-data-retained"); + let first_stopped = snapshot(&candidate, &run, deadline).receipt; + assert!( + journal_path.exists(), + "legacy binary must retain the old journal" + ); + assert!( + graph_root + .join("absent-publication-retirement.json") + .exists() + ); + assert_eq!( + first_stopped.resources["volume:data"].name, + before.resources["volume:data"].name + ); + // Rotate the external listener so the old journal cannot describe the next + // generation even if its completed services happen to match. + backend.stop(); + backend = RestartableBackend::start(port); + selection( + &selected_dependencies, + &selected_plan.plan_id, + &artifact, + &artifact_hash, + &backend.endpoint(), + backend.port, + ); + selected_plan_id = checked_cli( + "new-dependency-plan", + &legacy, + &candidate, + &[ + "graph", + "dependency-plan", + "--dependencies", + selected_dependencies.to_str().unwrap(), + "--json", + ], + deadline, + )["dependency_plan_id"] + .as_str() + .unwrap() + .to_owned(); + let first_generation = checked_cli( + "legacy-restore-selection", + &legacy, + &candidate, + &["graph", "restore-selection", "--run-id", &run, "--json"], + deadline, + )["generation"] + .as_str() + .unwrap() + .to_owned(); + let mut first_owner = start_selected(&legacy, Some(&first_generation), &selected_plan_id); + ready_at( + "first-legacy-restore", + &files, + &mut first_owner, + &run, + deadline, + ); + let first_ready = snapshot(&candidate, &run, deadline); + assert_eq!(first_ready.receipt.phase, "ready-observed"); + assert_ne!( + first_ready.receipt.resources["container:web"].id, + refreshed.resources["container:web"].id + ); + assert!(first_ready.journal_incomplete); + assert!(first_ready.guest_endpoints.is_empty()); + let foreground_root = transport::root(&candidate, &run).unwrap(); + let publisher_sha = file_hash(&foreground_root.join("owner.json")); + let first_cleanup = Process::start( + &legacy, + &candidate, + &["graph", "cleanup", "--run-id", &run, "--json"], + None, + ); + let mut first_cleanup = first_cleanup; + let cleanup_status = first_cleanup.wait(deadline); + assert_eq!(cleanup_status.code(), Some(2)); + let cleanup_error: Value = serde_json::from_slice(&first_cleanup.err).unwrap(); + assert_eq!(cleanup_error["code"], "graph_owner_recovery"); + assert!(first_owner.poll().is_none(), "failed cleanup retired owner"); + let reservation_path = candidate + .state_root + .join(format!("run/dependency-assignments/{run}.json")); + let reservation_bytes = fs::read(&reservation_path).unwrap(); + let reservation_metadata = fs::symlink_metadata(&reservation_path).unwrap(); + let reservation: Value = serde_json::from_slice(&reservation_bytes).unwrap(); + let second_process = crate::provider::identity::observe(first_owner.child.id() as i32).unwrap(); + assert_eq!(reservation["process"], json!(second_process)); + assert_eq!(reservation["run"], run); + let pool_owner = Owner::load(&candidate).unwrap(); + assert_eq!(reservation["owner"], pool_owner.token); + assert_eq!( + reservation["boot"], + pool_owner.guest_boot_id.clone().unwrap() + ); + let slots = first_ready + .receipt + .relay_startup + .as_ref() + .unwrap() + .services + .values() + .flat_map(|service| service.bindings.values().map(|binding| binding.slot)) + .collect::>(); + let recorded_slots = reservation["slots"] + .as_object() + .unwrap() + .values() + .map(|slot| u8::try_from(slot.as_u64().unwrap()).unwrap()) + .collect::>(); + assert_eq!(recorded_slots, slots); + for slot in &slots { + let socket = pool_owner + .short_home + .join(format!("dependency-{slot:02}.sock")); + let metadata = fs::symlink_metadata(socket).unwrap(); + assert_eq!( + reservation["sockets"][slot.to_string()], + json!([metadata.dev(), metadata.ino()]) + ); + } + // The original crash and legacy archival-error controls stay intact. Only + // this later owner retires through its installed signal handler so its + // ManagedOwner closes exact sockets before supported claim retirement. + assert_eq!( + graceful_stop(&files, &mut first_owner, &second_process, deadline).code(), + Some(2) + ); + let stop_error: Value = serde_json::from_slice(&first_owner.err).unwrap(); + assert_eq!(stop_error["code"], "graph_dependency_refresh_refused"); + for slot in &slots { + let socket = pool_owner + .short_home + .join(format!("dependency-{slot:02}.sock")); + assert_eq!( + fs::symlink_metadata(socket).unwrap_err().kind(), + std::io::ErrorKind::NotFound, + "graceful legacy teardown retained a selected socket" + ); + } + assert_eq!(fs::read(&reservation_path).unwrap(), reservation_bytes); + let retained_metadata = fs::symlink_metadata(&reservation_path).unwrap(); + assert_eq!( + (retained_metadata.dev(), retained_metadata.ino()), + (reservation_metadata.dev(), reservation_metadata.ino()) + ); + let newer_stopped = snapshot(&candidate, &run, deadline).receipt; + assert_eq!(newer_stopped.phase, "stopped-data-retained"); + let stopped_sha = file_hash(&graph_root.join("state.json")); + assert!(journal_path.exists()); + assert_eq!( + newer_stopped.resources["volume:data"].name, + first_stopped.resources["volume:data"].name + ); + assert_eq!( + serde_json::to_value( + snapshot(&candidate, &sibling_run, deadline) + .receipt + .resources + ) + .unwrap(), + serde_json::to_value(&sibling_before.resources).unwrap(), + ); + assert!(sibling_owner.poll().is_none()); + let retired_publisher = checked_cli( + "retire-acknowledged-publisher", + &binary, + &candidate, + &[ + "graph", + "retire-acknowledged-publisher", + "--run-id", + &run, + "--expect-owner", + &newer_stopped.owner, + "--expect-receipt", + &stopped_sha, + "--expect-publisher", + &publisher_sha, + "--json", + ], + deadline, + ); + assert_eq!(retired_publisher["publisher_retired"], true); + assert_eq!(retired_publisher["data_retained"], true); + assert_eq!(retired_publisher["acknowledged_cleanup"], true); + assert!(!foreground_root.join("owner.json").exists()); + assert!(!foreground_root.join("control.sock").exists()); + let inspected_reservations = checked_cli( + "select-acknowledged-dependency-reservation", + &binary, + &candidate, + &["graph", "dependency-reservations", "--json"], + deadline, + ); + let selected_claims = inspected_reservations["reservations"] + .as_array() + .unwrap() + .iter() + .filter(|claim| claim["run"] == run) + .collect::>(); + assert_eq!(selected_claims.len(), 1); + assert_eq!(selected_claims[0]["owner_alive"], false); + assert_eq!(selected_claims[0]["slots"], reservation["slots"]); + let reservation_sha = selected_claims[0]["reservation"].as_str().unwrap(); + let released = checked_cli( + "release-acknowledged-dependencies", + &binary, + &candidate, + &[ + "graph", + "release-acknowledged-dependencies", + "--run-id", + &run, + "--expect-owner", + &newer_stopped.owner, + "--expect-receipt", + &stopped_sha, + "--expect-publisher", + &publisher_sha, + "--expect-reservation", + reservation_sha, + "--json", + ], + deadline, + ); + assert_eq!(released["reservation_released"], true); + assert_eq!(released["record_retained"], true); + assert_eq!(released["data_retained"], true); + assert_eq!(released["same_boot"], true); + assert_eq!( + fs::symlink_metadata(&reservation_path).unwrap_err().kind(), + std::io::ErrorKind::NotFound + ); + let archived_reservation = graph_root.join(format!( + "dependency-reservation-retired-{reservation_sha}.json" + )); + assert_eq!(fs::read(&archived_reservation).unwrap(), reservation_bytes); + let archived_metadata = fs::symlink_metadata(archived_reservation).unwrap(); + assert_eq!( + (archived_metadata.dev(), archived_metadata.ino()), + (reservation_metadata.dev(), reservation_metadata.ino()) + ); + assert_eq!(file_hash(&graph_root.join("state.json")), stopped_sha); + let remaining_claims = checked_cli( + "verify-acknowledged-dependency-release", + &binary, + &candidate, + &["graph", "dependency-reservations", "--json"], + deadline, + ); + assert_eq!( + remaining_claims["reservations"], + json!( + inspected_reservations["reservations"] + .as_array() + .unwrap() + .iter() + .filter(|claim| claim["run"] != run) + .collect::>() + ) + ); + assert_eq!( + serde_json::to_value(snapshot(&candidate, &sibling_run, deadline).receipt).unwrap(), + serde_json::to_value(&sibling_before).unwrap() + ); + assert!(sibling_owner.poll().is_none()); + let history_path = graph_root.join(format!( + "dependency-rebind-history-{}", + graph::service_exec_generation(&refreshed).unwrap() + )); + { + let retired = transport::Retired::acquire(&candidate, &run) + .unwrap() + .unwrap(); + let engine = graph::Engine::connect_cleanup_wait(&candidate).unwrap(); + let interrupted = graph::absent_publication_cleanup::archive_retired_rebind_under( + &candidate, + &engine, + &newer_stopped, + &|| { + retired.verify()?; + if history_path.join("dependency-rebind.json").exists() { + return Err(CandidateError::new( + "test_archive_fault", + "owned archival pause", + )); + } + Ok(()) + }, + ); + assert!(interrupted.is_err()); + retired.verify().unwrap(); + } + assert!(!journal_path.exists()); + assert!(history_path.join("dependency-rebind.json").exists()); + assert_eq!( + state::read::(&history_path.join("proof.json")).unwrap()["complete"], + false + ); + let current_generation = checked_cli( + "current-restore-selection", + &binary, + &candidate, + &["graph", "restore-selection", "--run-id", &run, "--json"], + deadline, + )["generation"] + .as_str() + .unwrap() + .to_owned(); + selected_plan_id = checked_cli( + "current-dependency-plan", + &binary, + &candidate, + &[ + "graph", + "dependency-plan", + "--dependencies", + selected_dependencies.to_str().unwrap(), + "--json", + ], + deadline, + )["dependency_plan_id"] + .as_str() + .unwrap() + .to_owned(); + let mut recovered_owner = start_selected(&binary, Some(¤t_generation), &selected_plan_id); + ready_at( + "current-restore", + &files, + &mut recovered_owner, + &run, + deadline, + ); + let final_snapshot = snapshot(&candidate, &run, deadline); + assert_eq!(final_snapshot.receipt.phase, "ready-observed"); + assert!(!final_snapshot.journal_incomplete); + assert!(final_snapshot.guest_endpoints.contains_key("web")); + assert_eq!( + state::read::(&history_path.join("proof.json")).unwrap()["complete"], + true + ); + assert!(!journal_path.exists()); + marker( + &binary, + &candidate, + &run, + "selected-marker-v1", + false, + deadline, + ); + marker( + &binary, + &candidate, + &sibling_run, + "sibling-marker-v1", + false, + deadline, + ); + assert_eq!( + serde_json::to_value( + snapshot(&candidate, &sibling_run, deadline) + .receipt + .resources + ) + .unwrap(), + serde_json::to_value(&sibling_before.resources).unwrap(), + ); + let removed = checked_cli( + "selected-cleanup", + &binary, + &candidate, + &[ + "graph", + "cleanup", + "--run-id", + &run, + "--remove-data", + "--json", + ], + deadline, + ); + assert_eq!(removed["phase"], "removed"); + assert!(recovered_owner.wait(deadline).success()); + let sibling_removed = checked_cli( + "sibling-cleanup", + &binary, + &candidate, + &[ + "graph", + "cleanup", + "--run-id", + &sibling_run, + "--remove-data", + "--json", + ], + deadline, + ); + assert_eq!(sibling_removed["phase"], "removed"); + assert!(sibling_owner.wait(deadline).success()); + backend.stop(); + // Delete inputs only after both explicit managed data removals succeeded. + files.dispose().unwrap(); +} diff --git a/packages/runtime-core/src/provider/graph/foreground/native_test/retired_rebind_history/fixture.rs b/packages/runtime-core/src/provider/graph/foreground/native_test/retired_rebind_history/fixture.rs new file mode 100644 index 000000000..8d64a5909 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/foreground/native_test/retired_rebind_history/fixture.rs @@ -0,0 +1,362 @@ +//! This regression keeps its inputs on unwind. Only its verified success path +//! disposes them; graph/VM cleanup remains explicit in the caller. +use super::*; +use std::{ + io, + os::unix::fs::{DirBuilderExt, OpenOptionsExt}, + panic::{AssertUnwindSafe, catch_unwind, resume_unwind}, +}; + +pub(super) struct Files { + root: PathBuf, + identities: Vec<(PathBuf, u64, u64)>, + pub(super) selected: PathBuf, + pub(super) sibling: PathBuf, + pub(super) private: PathBuf, +} + +impl Files { + pub(super) fn new(parent: &Path) -> Self { + let root = parent.join(format!( + "retired-rebind-fixture-{}", + graph::probes::token().unwrap() + )); + let selected = root.join("selected"); + let sibling = root.join("sibling"); + let private = root.join("private"); + let mut identities = Vec::new(); + for path in [&root, &selected, &sibling, &private] { + fs::DirBuilder::new().mode(0o700).create(path).unwrap(); + let metadata = fs::symlink_metadata(path).unwrap(); + identities.push((path.clone(), metadata.dev(), metadata.ino())); + } + Self { + root, + identities, + selected, + sibling, + private, + } + } + + pub(super) fn dispose(self) -> io::Result<()> { + for (path, device, inode) in &self.identities { + let metadata = fs::symlink_metadata(path)?; + if !metadata.is_dir() || (metadata.dev(), metadata.ino()) != (*device, *inode) { + return Err(io::Error::other("owned fixture directory changed")); + } + } + fs::remove_dir_all(&self.root) + } +} + +fn write_new(path: &Path, bytes: &[u8]) -> io::Result<()> { + let mut file = fs::OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .open(path)?; + file.write_all(bytes) +} + +pub(super) fn ready_at( + stage: &str, + files: &Files, + owner: &mut Process, + run: &str, + deadline: Instant, +) { + // The four callers supply fixed stage names. These private files retain the + // bounded foreground pipes before unwinding can retire the owned child. + eprintln!("retired-rebind stage={stage} waiting"); + write_new(&files.private.join(format!("{stage}.waiting")), b"").unwrap(); + let result = catch_unwind(AssertUnwindSafe(|| ready(owner, run, deadline))); + if let Err(panic) = result { + eprintln!("retired-rebind stage={stage} failed"); + for (suffix, bytes) in [("stdout", &owner.out), ("stderr", &owner.err)] { + if let Err(error) = write_new(&files.private.join(format!("{stage}.{suffix}")), bytes) { + eprintln!("retired-rebind stage={stage} diagnostic write failed: {error}"); + } + } + resume_unwind(panic); + } + write_new(&files.private.join(format!("{stage}.ready")), b"").unwrap(); + eprintln!("retired-rebind stage={stage} ready"); +} + +/// Only the second legacy owner uses graceful retirement. Its registered +/// foreground signal handler must unwind the real managed transport owner even +/// when historical journal archival still refuses; PID death alone is not proof +/// that a reservation or socket may be released. +pub(super) fn graceful_stop( + files: &Files, + owner: &mut Process, + process: &crate::provider::identity::ProcessIdentity, + deadline: Instant, +) -> ExitStatus { + let stage = "second-legacy-stop"; + write_new(&files.private.join(format!("{stage}.waiting")), b"").unwrap(); + let result = catch_unwind(AssertUnwindSafe(|| { + assert!( + owner.poll().is_none(), + "selected legacy owner already exited" + ); + assert_eq!( + crate::provider::identity::observe(owner.child.id() as i32).unwrap(), + *process, + "selected legacy process changed" + ); + // SAFETY: the retained, unreaped Child owns this exact PID; it cannot be + // recycled. SIGTERM is handled by foreground::signals::Events. + assert_eq!(unsafe { libc::kill(process.pid, libc::SIGTERM) }, 0); + let status = owner.wait(deadline.min(Instant::now() + Duration::from_secs(20))); + assert!( + status.code().is_some(), + "legacy signal handler did not exit normally" + ); + assert!(!crate::provider::identity::alive(process.pid).unwrap()); + status + })); + for (suffix, bytes) in [("stdout", &owner.out), ("stderr", &owner.err)] { + write_new(&files.private.join(format!("{stage}.{suffix}")), bytes).unwrap(); + } + let status = match result { + Ok(status) => status, + Err(panic) => resume_unwind(panic), + }; + write_new(&files.private.join(format!("{stage}.stopped")), b"").unwrap(); + status +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::provider::relay_owner::{ + Context, + managed::{ManagedOwner, ManagedSlot}, + }; + use std::os::unix::{ + fs::PermissionsExt, + net::{UnixListener, UnixStream}, + }; + + #[test] + fn graceful_stop_uses_foreground_signals_and_real_owner_without_stealing_replacements() { + const CHILD_ROOT: &str = "HACK_RETIRED_REBIND_STOP_TEST_ROOT"; + const CHILD_PRIVATE: &str = "HACK_RETIRED_REBIND_STOP_TEST_PRIVATE"; + const CHILD_SOCKET_HOME: &str = "HACK_RETIRED_REBIND_STOP_TEST_SOCKET_HOME"; + const CHILD_RUN: &str = "HACK_RETIRED_REBIND_STOP_TEST_RUN"; + if let Some(root) = std::env::var_os(CHILD_ROOT) { + let candidate = Candidate::discover(Path::new(&root)).unwrap(); + let private = PathBuf::from(std::env::var_os(CHILD_PRIVATE).unwrap()); + let socket_home = PathBuf::from(std::env::var_os(CHILD_SOCKET_HOME).unwrap()); + let run = std::env::var(CHILD_RUN).unwrap(); + let mut publication = transport::Publication::bind(&candidate, &run).unwrap(); + let events = super::super::super::super::signals::Events::new(&publication).unwrap(); + let control = socket_home.join("owner-control"); + state::private_directory(&control).unwrap(); + let managed = ManagedOwner::start( + Context { + runtime: [1; 16], + boot: [2; 16], + }, + &control, + vec![ManagedSlot { + slot: 0, + path: socket_home.join("dependency-00.sock"), + canonical_parent: socket_home.clone(), + }], + ) + .unwrap(); + managed.verify_alive().unwrap(); + write_new(&private.join("owner-ready"), b"").unwrap(); + // wait reports EVFILT_SIGNAL directly; the process-wide pending + // flag is an independent fast path, not a required postcondition. + assert!(events.wait().unwrap()); + // Exercise the same production Drop that a legacy archival refusal + // reaches after its foreground SIGTERM handler returns. + drop(managed); + publication.finish().unwrap(); + return; + } + for replacement in [false, true] { + let parent = graph::tests::Fixture::new(); + let candidate = Candidate::discover(&parent.0).unwrap(); + let files = Files::new(&candidate.checkout); + let run = graph::probes::token().unwrap(); + // Use an owned short control directory, as the real managed owner + // does; macOS Unix addresses cannot hold the host's full TMPDIR. + let socket_home = PathBuf::from(format!("/private/tmp/hkrr-{}", &run[..16])); + fs::DirBuilder::new() + .mode(0o700) + .create(&socket_home) + .unwrap(); + let sibling = socket_home.join("sibling.sock"); + let sibling_listener = UnixListener::bind(&sibling).unwrap(); + let sibling_inode = fs::symlink_metadata(&sibling).unwrap().ino(); + let child = Command::new(std::env::current_exe().unwrap()) + .args([ + "--exact", + "provider::graph::foreground::native_test::retired_rebind_history::fixture::tests::graceful_stop_uses_foreground_signals_and_real_owner_without_stealing_replacements", + "--nocapture", + "--test-threads=1", + ]) + .env(CHILD_ROOT, &candidate.checkout) + .env(CHILD_PRIVATE, &files.private) + .env(CHILD_SOCKET_HOME, &socket_home) + .env(CHILD_RUN, &run) + .stdin(Stdio::null()) + .stdout(Stdio::piped()) + .stderr(Stdio::piped()) + .spawn().unwrap(); + let mut owner = Process { + child, + out: Vec::new(), + err: Vec::new(), + }; + nonblocking(owner.child.stdout.as_ref().unwrap().as_raw_fd()); + nonblocking(owner.child.stderr.as_ref().unwrap().as_raw_fd()); + let deadline = Instant::now() + Duration::from_secs(10); + while !files.private.join("owner-ready").exists() { + assert!( + owner.poll().is_none(), + "owned signal test exited: {}", + String::from_utf8_lossy(&owner.err) + ); + assert!( + Instant::now() < deadline, + "owned signal test readiness deadline" + ); + std::thread::sleep(Duration::from_millis(10)); + } + let process = crate::provider::identity::observe(owner.child.id() as i32).unwrap(); + let socket = socket_home.join("dependency-00.sock"); + let replacement_listener = replacement.then(|| { + fs::rename(&socket, socket_home.join("displaced-owned.sock")).unwrap(); + UnixListener::bind(&socket).unwrap() + }); + let inode = fs::symlink_metadata(&socket).unwrap().ino(); + assert!( + graceful_stop(&files, &mut owner, &process, deadline).success(), + "{}", + String::from_utf8_lossy(&owner.err) + ); + if replacement { + assert_eq!(fs::symlink_metadata(&socket).unwrap().ino(), inode); + assert!(UnixStream::connect(&socket).is_ok()); + assert!(socket_home.join("displaced-owned.sock").exists()); + } else { + assert_eq!( + fs::symlink_metadata(&socket).unwrap_err().kind(), + io::ErrorKind::NotFound + ); + } + assert_eq!(fs::symlink_metadata(&sibling).unwrap().ino(), sibling_inode); + assert!(UnixStream::connect(&sibling).is_ok()); + assert!(files.private.join("second-legacy-stop.stopped").exists()); + drop(replacement_listener); + drop(sibling_listener); + // The child's Publication::finish already removed its endpoint names. + // Dispose only this test's now-empty publication directory/lock. + fs::remove_dir_all(transport::root(&candidate, &run).unwrap()).unwrap(); + fs::remove_dir_all(socket_home).unwrap(); + files.dispose().unwrap(); + } + } + + fn inputs(files: &Files) -> Vec<(PathBuf, Vec, u64)> { + [ + files.selected.join("compose.yaml"), + files.sibling.join("compose.yaml"), + files.private.join("selected-dependencies.json"), + files.private.join("prior-owner.json"), + ] + .into_iter() + .map(|path| { + let bytes = path.file_name().unwrap().as_encoded_bytes().to_vec(); + write_new(&path, &bytes).unwrap(); + let inode = fs::symlink_metadata(&path).unwrap().ino(); + (path, bytes, inode) + }) + .collect() + } + + #[test] + fn unwind_preserves_inputs_and_foreground_failure_diagnostics() { + let parent = graph::tests::Fixture::new(); + let candidate = Candidate::discover(&parent.0).unwrap(); + let files = Files::new(&candidate.checkout); + let preserved = inputs(&files); + let private = files.private.clone(); + let binary = private.join("refused-cli"); + write_new( + &binary, + b"#!/bin/sh\nprintf '{\"code\":\"dependency_reservation\"}\\n' >&2\nexit 2\n", + ) + .unwrap(); + fs::set_permissions(&binary, fs::Permissions::from_mode(0o700)).unwrap(); + let failed = catch_unwind(AssertUnwindSafe(move || { + let mut owner = Process::start(&binary, &candidate, &[], None); + ready_at( + "first-legacy-restore", + &files, + &mut owner, + "fixture-run", + Instant::now() + Duration::from_secs(5), + ); + })); + assert!(failed.is_err()); + for (path, bytes, inode) in preserved { + assert_eq!(fs::read(&path).unwrap(), bytes); + assert_eq!(fs::symlink_metadata(path).unwrap().ino(), inode); + } + assert!(private.join("first-legacy-restore.waiting").exists()); + assert!(!private.join("first-legacy-restore.ready").exists()); + assert_eq!( + fs::read(private.join("first-legacy-restore.stderr")).unwrap(), + b"{\"code\":\"dependency_reservation\"}\n" + ); + } + + #[test] + fn explicit_success_disposal_removes_only_owned_inputs() { + let parent = graph::tests::Fixture::new(); + let sentinel = parent.0.join("other-owner-data"); + write_new(&sentinel, b"preserved").unwrap(); + let files = Files::new(&parent.0); + let paths = inputs(&files); + let root = files.root.clone(); + files.dispose().unwrap(); + assert!(!root.exists()); + for (path, _, _) in paths { + assert!(!path.exists()); + } + assert_eq!(fs::read(sentinel).unwrap(), b"preserved"); + } + + #[test] + fn disposal_refuses_replaced_directory_and_preserves_originals() { + let parent = graph::tests::Fixture::new(); + let files = Files::new(&parent.0); + let preserved = inputs(&files); + let original = files.root.join("original-private"); + fs::rename(&files.private, &original).unwrap(); + fs::create_dir(&files.private).unwrap(); + write_new(&files.private.join("other-owner-data"), b"preserved").unwrap(); + let replacement = files.private.clone(); + assert!(files.dispose().is_err()); + assert_eq!( + fs::read(replacement.join("other-owner-data")).unwrap(), + b"preserved" + ); + for (path, bytes, inode) in preserved { + let path = if path.parent() == Some(replacement.as_path()) { + original.join(path.file_name().unwrap()) + } else { + path + }; + assert_eq!(fs::read(&path).unwrap(), bytes); + assert_eq!(fs::symlink_metadata(path).unwrap().ino(), inode); + } + } +} diff --git a/packages/runtime-core/src/provider/graph/foreground/native_test/same_boot_recovery.rs b/packages/runtime-core/src/provider/graph/foreground/native_test/same_boot_recovery.rs index ce1810cf8..269e9b6a6 100644 --- a/packages/runtime-core/src/provider/graph/foreground/native_test/same_boot_recovery.rs +++ b/packages/runtime-core/src/provider/graph/foreground/native_test/same_boot_recovery.rs @@ -419,6 +419,77 @@ fn exercise(previous_boot: bool) { deadline, ); assert!(owners[0].wait(deadline).success()); + let graph_root = graph::directory(&candidate, &runs[0]).unwrap(); + let stopped_path = graph_root.join("state.json"); + let stopped_bytes = fs::read(&stopped_path).unwrap(); + let stopped = snapshot(&candidate, &runs[0], deadline); + assert_eq!(stopped.receipt.phase, "stopped-data-retained"); + let volume = stopped.observations["volume:data"].clone(); + let mount = volume["Mountpoint"].as_str().unwrap(); + let sibling_path = graph::directory(&candidate, &runs[1]) + .unwrap() + .join("state.json"); + let sibling_bytes = fs::read(&sibling_path).unwrap(); + for _ in 0..2 { + let result = checked_cli( + "confirm-ordinary-retirement", + &binary, + &candidate, + &[ + "graph", + "retire-recovered-publisher", + "--run-id", + &runs[0], + "--expect-owner", + &stopped.receipt.owner, + "--json", + ], + deadline, + ); + assert_eq!(result["publisher_retired"], true); + assert_eq!(result["data_retained"], true); + assert_eq!(fs::read(&stopped_path).unwrap(), stopped_bytes); + let observed = snapshot(&candidate, &runs[0], deadline); + assert_eq!(observed.observations["volume:data"], volume); + let engine = graph::Engine::connect_cleanup(&candidate).unwrap(); + assert_eq!( + engine + .guest() + .execute_cleanup( + "set -eu; test ! -L \"$1\"; test ! -L \"$1/hack-rebind-marker\"; test \"$(cat \"$1/hack-rebind-marker\")\" = retained-dependency-rebind-v1; printf preserved", + &[mount], + ) + .unwrap(), + "preserved" + ); + assert!(owners[1].poll().is_none()); + assert_eq!(fs::read(&sibling_path).unwrap(), sibling_bytes); + } + let other_owner = if stopped.receipt.owner == "f".repeat(32) { + "e".repeat(32) + } else { + "f".repeat(32) + }; + let mut wrong_owner = Process::start( + &binary, + &candidate, + &[ + "graph", + "retire-recovered-publisher", + "--run-id", + &runs[0], + "--expect-owner", + &other_owner, + "--json", + ], + None, + ); + assert_eq!(wrong_owner.wait(deadline).code(), Some(2)); + assert!(wrong_owner.out.is_empty()); + let error: Value = serde_json::from_slice(&wrong_owner.err).unwrap(); + assert_eq!(error["code"], "graph_acknowledged_publisher"); + assert_eq!(fs::read(&stopped_path).unwrap(), stopped_bytes); + assert_eq!(fs::read(&sibling_path).unwrap(), sibling_bytes); let selection = graph::foreground::restore_selection(&candidate, &runs[0]).unwrap(); owners[0] = start(0, Some(selection["generation"].as_str().unwrap())); ready(&mut owners[0], &runs[0], deadline); diff --git a/packages/runtime-core/src/provider/graph/foreground/native_test/source_device_rebind.rs b/packages/runtime-core/src/provider/graph/foreground/native_test/source_device_rebind.rs new file mode 100644 index 000000000..e7d6660ee --- /dev/null +++ b/packages/runtime-core/src/provider/graph/foreground/native_test/source_device_rebind.rs @@ -0,0 +1,535 @@ +//! Owned synthetic-device fixture. Only a physical host reboot can qualify +//! original APFS volume continuity for an application; this fixture checks the +//! selected source transition, original receipt/history and retained marker. +use super::absent_publication_recovery::{inspect_args, ready, recover_args, refused_cli}; +use super::dependency_rebind::{checked_cli, exec, snapshot}; +use super::*; +use crate::provider::{ProjectShareIntent, lifecycle, state::Owner}; +use base64::Engine as _; +use std::os::unix::fs::MetadataExt; + +fn shared_source_marker( + binary: &Path, + candidate: &Candidate, + run: &str, + expected: &[u8], + deadline: Instant, +) { + let mut process = Process::start( + binary, + candidate, + &[ + "graph", + "exec", + "--run-id", + run, + "--service", + "web", + "--timeout-seconds", + "12", + "--json", + "--", + "/bin/sh", + "-ec", + "cat /app/source-marker.txt", + ], + None, + ); + let status = process.wait(deadline); + let result: Value = serde_json::from_slice(&process.out).unwrap_or(Value::Null); + assert!(status.success(), "shared-source exec CLI refused"); + assert_eq!(result["exit_code"], 0, "shared-source process failed"); + assert_eq!(result["truncated"], false); + let stdout = base64::engine::general_purpose::STANDARD + .decode(result["stdout_base64"].as_str().unwrap()) + .unwrap(); + let stderr = base64::engine::general_purpose::STANDARD + .decode(result["stderr_base64"].as_str().unwrap()) + .unwrap(); + assert!(stderr.is_empty()); + assert_eq!(stdout, expected); +} + +struct ReplacedBytes { + path: PathBuf, + original: Vec, + substituted: Vec, +} +impl ReplacedBytes { + fn new(path: &Path, bytes: Vec) -> Self { + let original = fs::read(path).unwrap(); + fs::write(path, &bytes).unwrap(); + Self { + path: path.to_owned(), + original, + substituted: bytes, + } + } +} +impl Drop for ReplacedBytes { + fn drop(&mut self) { + if fs::read(&self.path).ok().as_deref() == Some(self.substituted.as_slice()) { + fs::write(&self.path, &self.original).unwrap(); + } + } +} + +#[test] +#[ignore = "Owned development VM, private synthetic prior-device fault, pinned image and 300s watchdog required"] +fn selected_source_device_rebind_restores_same_run_and_two_ordinary_generations() { + let deadline = Instant::now() + Duration::from_secs(270); + let candidate = + Candidate::discover(Path::new(&std::env::var("HACK_LOCAL_TEST_ROOT").unwrap())).unwrap(); + let binary = PathBuf::from(std::env::var("HACK_LOCAL_TEST_BINARY").unwrap()); + let image = std::env::var("HACK_LOCAL_TEST_IMAGE").unwrap(); + let image_archive = std::env::var("HACK_LOCAL_TEST_IMAGE_ARCHIVE").unwrap(); + let image_sha256 = std::env::var("HACK_LOCAL_TEST_IMAGE_SHA256").unwrap(); + let fixture = graph::tests::Fixture::new(); + let project = fixture.0.join("project"); + fs::create_dir(&project).unwrap(); + state::write( + &project.join("compose.yaml"), + &json!({ + "services":{"web":{"image":image,"read_only":true,"network_mode":"none", + "init":true,"user":"0:0","entrypoint":["/bin/sleep","300"],"command":[], + "volumes":["data:/data",".:/app:ro"], + "healthcheck":{"test":["CMD","/bin/hack-graph-startup-app","complete"], + "interval":"200ms","timeout":"2s","retries":10,"start_period":"500ms"}}}, + "volumes":{"data":{}} + }), + ) + .unwrap(); + let marker_path = project.join("source-marker.txt"); + let initial_marker = b"source-initial\n"; + fs::write(&marker_path, initial_marker).unwrap(); + // This fixture starts an otherwise uninitialized owned pool with one exact + // approved source root; the external harness supplies the candidate binary. + assert_eq!( + lifecycle::status(&candidate).unwrap().phase, + "uninitialized" + ); + let share = ProjectShareIntent::approve(&project, true).unwrap(); + lifecycle::up_with_project_share( + &candidate, + crate::provider::Profile::Development, + None, + None, + None, + Some(share.clone()), + ) + .unwrap(); + checked_cli( + "load-pinned-image", + &binary, + &candidate, + &[ + "runtime", + "load-image", + "--archive", + &image_archive, + "--sha256", + &image_sha256, + "--image-id", + &image, + ], + deadline, + ); + let run = graph::probes::token().unwrap(); + let review = project::plan( + &candidate, + project::PlanOptions { + branch: None, + project: &project, + compose_file: Path::new("compose.yaml"), + profiles: &[], + }, + ) + .unwrap(); + let normalized = project.join("compose.yaml"); + let dependencies_path = fixture.0.join("dependencies.json"); + state::write( + &dependencies_path, + &json!({"version":1,"plan":review.plan_id, + "artifact":"/tmp/unused-control-only-artifact","artifact_sha256":"a".repeat(64), + "dependencies":[]}), + ) + .unwrap(); + let dependency = checked_cli( + "dependency-plan", + &binary, + &candidate, + &[ + "graph", + "dependency-plan", + "--dependencies", + dependencies_path.to_str().unwrap(), + "--json", + ], + deadline, + ); + let start = |generation: Option<&str>| { + let mut args = vec![ + "graph", + if generation.is_some() { + "serve-restore" + } else { + "serve" + }, + "--project", + project.to_str().unwrap(), + "--file", + "compose.yaml", + "--expect-plan", + &review.plan_id, + "--run-id", + &run, + "--ready", + "web=healthy", + "--timeout-seconds", + "90", + "--dependencies", + dependencies_path.to_str().unwrap(), + "--expect-dependencies", + dependency["dependency_plan_id"].as_str().unwrap(), + "--shared-source", + "--normalized-file", + normalized.to_str().unwrap(), + "--expect-original", + &review.plan.compose_sha256, + "--expect-namespace", + &review.plan.namespace, + "--json", + ]; + if let Some(generation) = generation { + args.extend(["--expect-generation", generation]); + } + Process::start(&binary, &candidate, &args, None) + }; + let mut owner = start(None); + let mut cleanup = Cleanup { + binary: &binary, + candidate: &candidate, + run: &run, + done: false, + }; + ready(&mut owner, &run, deadline); + shared_source_marker(&binary, &candidate, &run, initial_marker, deadline); + assert_eq!( + exec(&binary, &candidate, &run, "web", "write-data", deadline)["exit_code"], + 0 + ); + let ready_receipt = snapshot(&candidate, &run, deadline).receipt; + let original_volume = ready_receipt.resources["volume:data"].name.clone(); + let original_container = ready_receipt.resources["container:web"].id.clone(); + let old_owner = Owner::load(&candidate).unwrap(); + owner.child.kill().unwrap(); + owner.wait(deadline); + assert_eq!(lifecycle::down(&candidate).unwrap().phase, "stopped"); + lifecycle::up(&candidate).unwrap(); + let current = Owner::load(&candidate).unwrap(); + assert_eq!(current.project_share.as_ref(), Some(&share)); + let old_device = share.device.checked_add(1).unwrap(); + let boot = lifecycle::host_filesystem::host_boot_micros().unwrap(); + let mut synthetic_owner = old_owner.clone(); + synthetic_owner.storage.as_mut().unwrap().device = old_device; + synthetic_owner.overlay.as_mut().unwrap().device = old_device; + synthetic_owner.project_share.as_mut().unwrap().device = old_device; + synthetic_owner.process.as_mut().unwrap().pid = i32::MAX; + synthetic_owner.process.as_mut().unwrap().start_micros = boot - 1; + let old_owner_path = fixture.0.join("prior-owner.json"); + state::write(&old_owner_path, &synthetic_owner).unwrap(); + let old_bytes = fs::read(&old_owner_path).unwrap(); + let pool = fs::symlink_metadata(candidate.state_root.join("run/smolvm")).unwrap(); + let inspection_path = fixture.0.join("prior-inspection.json"); + state::write( + &inspection_path, + &serde_json::from_slice::(&graph::absent_publication_cleanup::fixture_inspection( + &old_bytes, + ¤t, + old_device, + boot, + pool.ino(), + )) + .unwrap(), + ) + .unwrap(); + let graph_root = graph::directory(&candidate, &run).unwrap(); + // Test-only host-device fault: a real reboot changes st_dev without + // rewriting the original graph. This isolated synthetic receipt is fixed + // before selection and remains byte-for-byte immutable thereafter. + let mut synthetic_ready = ready_receipt.clone(); + synthetic_ready + .source + .as_mut() + .unwrap() + .shared + .as_mut() + .unwrap() + .device = old_device; + state::write(&graph_root.join("state.json"), &synthetic_ready).unwrap(); + let original_ready_bytes = fs::read(graph_root.join("state.json")).unwrap(); + let publisher_root = transport::root(&candidate, &run).unwrap(); + fs::remove_dir_all(&publisher_root).unwrap(); + fs::remove_dir_all( + synthetic_ready + .relay_startup + .as_ref() + .unwrap() + .control_root + .clone(), + ) + .unwrap(); + let prior = inspection_path.to_str().unwrap(); + let old = old_owner_path.to_str().unwrap(); + let absent_selection = checked_cli( + "select-absence", + &binary, + &candidate, + &inspect_args(&run, old, prior), + deadline, + ); + let absent_hash = absent_selection["selection_sha256"].as_str().unwrap(); + let stopped = checked_cli( + "recover-absence", + &binary, + &candidate, + &recover_args(&run, old, prior, absent_hash), + deadline, + ); + assert_eq!(stopped["phase"], "stopped-data-retained"); + let original_stopped_bytes = fs::read(graph_root.join("state.json")).unwrap(); + let original_stopped: graph::Receipt = serde_json::from_slice(&original_stopped_bytes).unwrap(); + assert_eq!( + original_stopped + .source + .as_ref() + .unwrap() + .shared + .as_ref() + .unwrap() + .device, + old_device + ); + let original_history_hash = format!("{:x}", Sha256::digest(&original_stopped_bytes)); + let witness_path = graph_root.join("source-device-rebind.json"); + let inspect = checked_cli( + "select-source", + &binary, + &candidate, + &[ + "graph", + "inspect-source-device-rebind", + "--run-id", + &run, + "--json", + ], + deadline, + ); + let witness_hash = inspect["selection_sha256"].as_str().unwrap(); + refused_cli( + &binary, + &candidate, + &[ + "graph", + "recover-source-device-rebind", + "--run-id", + &run, + "--expect-selection", + &"0".repeat(64), + "--accept-legacy-device-rebind", + "--json", + ], + "graph_source_device_rebind", + deadline, + ); + assert!(!witness_path.exists()); + let pending = witness_path.with_extension("pending"); + fs::write(&pending, b"foreign incomplete witness").unwrap(); + refused_cli( + &binary, + &candidate, + &[ + "graph", + "recover-source-device-rebind", + "--run-id", + &run, + "--expect-selection", + witness_hash, + "--accept-legacy-device-rebind", + "--json", + ], + "graph_source_device_rebind", + deadline, + ); + assert_eq!(fs::read(&pending).unwrap(), b"foreign incomplete witness"); + fs::remove_file(&pending).unwrap(); // Only this fixture's known pending file. + checked_cli( + "commit-source", + &binary, + &candidate, + &[ + "graph", + "recover-source-device-rebind", + "--run-id", + &run, + "--expect-selection", + witness_hash, + "--accept-legacy-device-rebind", + "--json", + ], + deadline, + ); + assert_eq!( + fs::read(graph_root.join("state.json")).unwrap(), + original_stopped_bytes + ); + let committed_bytes = fs::read(&witness_path).unwrap(); + { + let mut changed = committed_bytes.clone(); + changed.push(b' '); + let _changed = ReplacedBytes::new(&witness_path, changed); + refused_cli( + &binary, + &candidate, + &["graph", "restore-selection", "--run-id", &run, "--json"], + "graph_source_device_rebind", + deadline, + ); + } + { + let _changed = ReplacedBytes::new( + &witness_path, + graph::source_device_rebind::fixture_wrong_volume_projection(&committed_bytes), + ); + refused_cli( + &binary, + &candidate, + &["graph", "restore-selection", "--run-id", &run, "--json"], + "graph_source_device_rebind", + deadline, + ); + } + assert_eq!(fs::read(&witness_path).unwrap(), committed_bytes); + let selection = checked_cli( + "restore-selection-source", + &binary, + &candidate, + &["graph", "restore-selection", "--run-id", &run, "--json"], + deadline, + ); + let selected_generation = selection["generation"].as_str().unwrap(); + fs::write(&pending, b"appeared after selection").unwrap(); + let mut blocked = start(Some(selected_generation)); + assert_eq!(blocked.wait(deadline).code(), Some(2)); + assert!(blocked.out.is_empty()); + let refusal: Value = serde_json::from_slice(&blocked.err).unwrap(); + assert_eq!(refusal["code"], "graph_source_device_rebind"); + assert_eq!(fs::read(&pending).unwrap(), b"appeared after selection"); + assert!(!publisher_root.join("control.sock").exists()); + assert!(!publisher_root.join("owner.json").exists()); + fs::remove_file(&pending).unwrap(); // The fixture owns this exact pending file. + owner = start(Some(selected_generation)); + ready(&mut owner, &run, deadline); + shared_source_marker(&binary, &candidate, &run, initial_marker, deadline); + assert_eq!( + exec(&binary, &candidate, &run, "web", "read-data", deadline)["exit_code"], + 0 + ); + let restored = snapshot(&candidate, &run, deadline).receipt; + assert_eq!(restored.resources["volume:data"].name, original_volume); + assert_ne!(restored.resources["container:web"].id, original_container); + assert_eq!( + restored + .source + .as_ref() + .unwrap() + .shared + .as_ref() + .unwrap() + .device, + share.device + ); + assert!(fs::read(graph_root.join("restore-history.json")).is_ok()); + assert_eq!(fs::read(&witness_path).unwrap(), committed_bytes); + assert_eq!( + fs::read(graph_root.join("state.json")).unwrap(), + serde_json::to_vec_pretty(&restored).unwrap() + ); + assert!( + graph::restore_history::completed_for_recovery( + &graph_root, + &restored, + &original_history_hash + ) + .unwrap() + .is_some() + ); + assert_ne!( + fs::read(graph_root.join("state.json")).unwrap(), + original_ready_bytes + ); + // The selected witness is now historical. Each later cleanup/restore takes + // an ordinary generation and reads the same named retained data. + for generation in 0..2 { + let cleaned = checked_cli( + "ordinary-cleanup", + &binary, + &candidate, + &["graph", "cleanup", "--run-id", &run, "--json"], + deadline, + ); + assert_eq!(cleaned["phase"], "stopped-data-retained"); + assert!(owner.wait(deadline).success()); + let next = checked_cli( + "ordinary-selection", + &binary, + &candidate, + &["graph", "restore-selection", "--run-id", &run, "--json"], + deadline, + ); + owner = start(Some(next["generation"].as_str().unwrap())); + ready(&mut owner, &run, deadline); + shared_source_marker(&binary, &candidate, &run, initial_marker, deadline); + assert_eq!( + exec(&binary, &candidate, &run, "web", "read-data", deadline)["exit_code"], + 0 + ); + let observed = snapshot(&candidate, &run, deadline).receipt; + assert_eq!( + observed.resources["volume:data"].name, original_volume, + "ordinary generation {generation}" + ); + assert_eq!( + observed + .source + .as_ref() + .unwrap() + .shared + .as_ref() + .unwrap() + .device, + share.device + ); + } + // A source edit changes the reviewed plan's metadata identity even if its + // bytes are changed back. Prove the live share after all pinned-plan restores. + fs::write(&marker_path, b"source-host-edit\n").unwrap(); + shared_source_marker(&binary, &candidate, &run, b"source-host-edit\n", deadline); + let removed = checked_cli( + "remove-selected-data", + &binary, + &candidate, + &[ + "graph", + "cleanup", + "--run-id", + &run, + "--remove-data", + "--json", + ], + deadline, + ); + assert_eq!(removed["phase"], "removed"); + assert!(owner.wait(deadline).success()); + cleanup.done = true; +} diff --git a/packages/runtime-core/src/provider/graph/foreground/tests.rs b/packages/runtime-core/src/provider/graph/foreground/tests.rs index 3ba36a097..2c9e033b3 100644 --- a/packages/runtime-core/src/provider/graph/foreground/tests.rs +++ b/packages/runtime-core/src/provider/graph/foreground/tests.rs @@ -47,6 +47,25 @@ fn fresh_restore_publication_requires_prior_retirement_and_exclusive_lock() { drop(second); fs::remove_dir_all(path).unwrap(); } +#[cfg(target_os = "macos")] +#[test] +fn retired_publication_refuses_incomplete_source_device_witness_before_binding() { + let (_fixture, candidate, run) = fixture(); + let publisher = root(&candidate, &run); + let mut first = Publication::bind(&candidate, &run).unwrap(); + first.finish().unwrap(); + drop(first); + let graph_root = super::super::directory(&candidate, &run).unwrap(); + crate::provider::state::private_directory(&graph_root).unwrap(); + let pending = graph_root.join("source-device-rebind.pending"); + fs::write(&pending, b"foreign incomplete witness").unwrap(); + + assert!(Publication::bind_retired(&candidate, &run).is_err()); + assert!(!publisher.join("control.sock").exists()); + assert!(!publisher.join("owner.json").exists()); + assert_eq!(fs::read(&pending).unwrap(), b"foreign incomplete witness"); + fs::remove_dir_all(publisher).unwrap(); +} #[test] fn publication_serializes_and_pins_exact_server_socket() { let (_fixture, candidate, run) = fixture(); diff --git a/packages/runtime-core/src/provider/graph/foreground/transport.rs b/packages/runtime-core/src/provider/graph/foreground/transport.rs index 8ab4a43f3..5e43651d3 100644 --- a/packages/runtime-core/src/provider/graph/foreground/transport.rs +++ b/packages/runtime-core/src/provider/graph/foreground/transport.rs @@ -1,5 +1,6 @@ use super::{Candidate, CandidateError, refused}; use crate::provider::{ + host_pin::DeviceRebind, identity::{self, ProcessIdentity}, state, }; @@ -239,6 +240,7 @@ fn selected_retirement_path( owner: &str, socket: bool, expected: (u64, u64), + device_rebind: Option, ) -> Result<(PathBuf, bool), CandidateError> { let original = root.join(if socket { "control.sock" } else { "owner.json" }); let retired = retired_path(root, owner, socket); @@ -247,8 +249,9 @@ fn selected_retirement_path( (None, Some(found)) => (retired, found, false), _ => return Err(retirement_refused()), }; - if id(¤t) != expected - || !private(¤t) + if !device_rebind.map_or(id(¤t) == expected, |v| { + v.matches(expected, id(¤t)) + }) || !private(¤t) || (socket && !current.file_type().is_socket()) || (!socket && (!current.is_file() || current.nlink() != 1 || current.len() > 8192)) { @@ -260,19 +263,26 @@ fn selected_retirement_path( fn verify_retirement( candidate: &Candidate, run: &str, - expected_owner: &str, - expected_receipt: &str, + expected: (&str, &str), root: &Path, lock: &state::Lock, intent: &Retirement, + device_rebind: Option, ) -> Result<(bool, bool), CandidateError> { + let root_id = id(&fs::symlink_metadata(root).map_err(|_| retirement_refused())?); + let lock_path_id = + id(&fs::symlink_metadata(root.join("operation.lock")).map_err(|_| retirement_refused())?); + let parent_matches = device_rebind.map_or(intent.parent == root_id, |v| { + v.matches(intent.parent, root_id) + }); if intent.version != 1 || intent.candidate != candidate.checkout || intent.run != run - || intent.owner_sha256 != expected_owner - || intent.receipt_sha256 != expected_receipt - || intent.parent != id(&fs::symlink_metadata(root).map_err(|_| retirement_refused())?) + || intent.owner_sha256 != expected.0 + || intent.receipt_sha256 != expected.1 + || !parent_matches || intent.lock != lock.identity().map_err(|_| retirement_refused())? + || lock_path_id != intent.lock || intent.owner.parent != intent.parent || intent.owner.socket != intent.socket || intent.owner.candidate != candidate.checkout @@ -282,9 +292,9 @@ fn verify_retirement( return Err(retirement_refused()); } let (socket_path, socket_original) = - selected_retirement_path(root, expected_owner, true, intent.socket)?; + selected_retirement_path(root, expected.0, true, intent.socket, device_rebind)?; let (record_path, record_original) = - selected_retirement_path(root, expected_owner, false, intent.record)?; + selected_retirement_path(root, expected.0, false, intent.record, None)?; // Socket retirement precedes record retirement. The inverse is foreign. if socket_original && !record_original { return Err(retirement_refused()); @@ -304,7 +314,7 @@ fn verify_retirement( .read_to_end(&mut bytes) .map_err(|_| retirement_refused())?; let record: Record = serde_json::from_slice(&bytes).map_err(|_| retirement_refused())?; - if record != intent.owner || format!("{:x}", Sha256::digest(&bytes)) != expected_owner { + if record != intent.owner || format!("{:x}", Sha256::digest(&bytes)) != expected.0 { return Err(retirement_refused()); } Ok((socket_original, record_original)) @@ -318,6 +328,15 @@ pub(in crate::provider::graph) fn retire_recovered_publisher( run: &str, expected_owner: &str, expected_receipt: &str, +) -> Result<(), CandidateError> { + retire_recovered_publisher_recovery(candidate, run, expected_owner, expected_receipt, None) +} +pub(in crate::provider::graph) fn retire_recovered_publisher_recovery( + candidate: &Candidate, + run: &str, + expected_owner: &str, + expected_receipt: &str, + device_rebind: Option, ) -> Result<(), CandidateError> { if !super::super::hex(expected_owner, 64) || !super::super::hex(expected_receipt, 64) { return Err(retirement_refused()); @@ -325,17 +344,58 @@ pub(in crate::provider::graph) fn retire_recovered_publisher( let root = root(candidate, run)?; state::check_private_directory(&root).map_err(|_| retirement_refused())?; let lock = state::Lock::acquire_existing(&root).map_err(|_| retirement_refused())?; + retire_recovered_publisher_locked( + candidate, + run, + expected_owner, + expected_receipt, + device_rebind, + &lock, + ) +} +pub(in crate::provider::graph) fn retire_recovered_publisher_locked( + candidate: &Candidate, + run: &str, + expected_owner: &str, + expected_receipt: &str, + device_rebind: Option, + lock: &state::Lock, +) -> Result<(), CandidateError> { + retire_recovered_publisher_locked_fenced( + candidate, + run, + expected_owner, + expected_receipt, + device_rebind, + lock, + &|| Ok(()), + ) +} + +/// The caller's current cleanup proof remains valid at each publication effect. +pub(in crate::provider::graph) fn retire_recovered_publisher_locked_fenced( + candidate: &Candidate, + run: &str, + expected_owner: &str, + expected_receipt: &str, + device_rebind: Option, + lock: &state::Lock, + verify_cleanup: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + verify_cleanup()?; + let root = root(candidate, run)?; let path = retirement_path(&root, expected_owner); let pending = path.with_extension("pending"); let name = pending .file_name() .and_then(|value| value.to_str()) .ok_or_else(retirement_refused)?; + verify_cleanup()?; super::super::journal::retain_file(&root, name, "publisher-retirement-interrupted", 8192)?; let intent: Retirement = if metadata(&path)?.is_some() { state::read(&path).map_err(|_| retirement_refused())? } else { - let pin = Pin::read(candidate, run)?; + let pin = Pin::read_with_rebind(candidate, run, device_rebind)?; if format!("{:x}", Sha256::digest(&pin.bytes)) != expected_owner || identity::alive(pin.record.process.pid).unwrap_or(true) { @@ -354,19 +414,21 @@ pub(in crate::provider::graph) fn retire_recovered_publisher( record: pin.record_id, owner: pin.record, }; + verify_cleanup()?; state::write(&path, &selected).map_err(|_| retirement_refused())?; selected }; let (socket_original, record_original) = verify_retirement( candidate, run, - expected_owner, - expected_receipt, + (expected_owner, expected_receipt), &root, - &lock, + lock, &intent, + device_rebind, )?; if socket_original { + verify_cleanup()?; fs::rename( root.join("control.sock"), retired_path(&root, expected_owner, true), @@ -377,14 +439,15 @@ pub(in crate::provider::graph) fn retire_recovered_publisher( .map_err(|_| retirement_refused())?; } if record_original { + verify_cleanup()?; let _ = verify_retirement( candidate, run, - expected_owner, - expected_receipt, + (expected_owner, expected_receipt), &root, - &lock, + lock, &intent, + device_rebind, )?; fs::rename( root.join("owner.json"), @@ -398,26 +461,43 @@ pub(in crate::provider::graph) fn retire_recovered_publisher( let (socket_original, record_original) = verify_retirement( candidate, run, - expected_owner, - expected_receipt, + (expected_owner, expected_receipt), &root, - &lock, + lock, &intent, + device_rebind, )?; if socket_original || record_original { return Err(retirement_refused()); } + verify_cleanup()?; Ok(()) } /** A later boot may resume frontend recovery only after retirement was fully * recorded. This read-only proof never begins or completes a partial move. */ +#[cfg(test)] pub(in crate::provider::graph) fn verify_recovered_publisher_retired( candidate: &Candidate, run: &str, expected_owner: &str, expected_receipt: &str, +) -> Result<(), CandidateError> { + verify_recovered_publisher_retired_recovery( + candidate, + run, + expected_owner, + expected_receipt, + None, + ) +} +pub(in crate::provider::graph) fn verify_recovered_publisher_retired_recovery( + candidate: &Candidate, + run: &str, + expected_owner: &str, + expected_receipt: &str, + device_rebind: Option, ) -> Result<(), CandidateError> { if !super::super::hex(expected_owner, 64) || !super::super::hex(expected_receipt, 64) { return Err(retirement_refused()); @@ -430,11 +510,11 @@ pub(in crate::provider::graph) fn verify_recovered_publisher_retired( let (socket_original, record_original) = verify_retirement( candidate, run, - expected_owner, - expected_receipt, + (expected_owner, expected_receipt), &root, &lock, &intent, + device_rebind, )?; if socket_original || record_original { return Err(retirement_refused()); @@ -485,11 +565,12 @@ fn peer(stream: &UnixStream) -> Result { identity::verify(&observed, &observed, &observed.executable, uid).map_err(|_| refused())?; Ok(observed) } -pub(super) struct Pin { +pub(in crate::provider::graph) struct Pin { root: PathBuf, record: Record, bytes: Vec, record_id: (u64, u64), + device_rebind: Option, } impl Pin { pub fn load(candidate: &Candidate, run: &str) -> Result { @@ -498,6 +579,13 @@ impl Pin { Ok(pin) } fn read(candidate: &Candidate, run: &str) -> Result { + Self::read_with_rebind(candidate, run, None) + } + fn read_with_rebind( + candidate: &Candidate, + run: &str, + device_rebind: Option, + ) -> Result { let root = root(candidate, run)?; state::check_private_directory(&root).map_err(|_| refused())?; let mut file = OpenOptions::new() @@ -530,6 +618,7 @@ impl Pin { record, bytes, record_id: id(&metadata), + device_rebind, }; pin.verify_files()?; Ok(pin) @@ -540,10 +629,10 @@ impl Pin { let socket = fs::symlink_metadata(self.root.join("control.sock")).map_err(|_| refused())?; let file = fs::symlink_metadata(self.root.join("owner.json")).map_err(|_| refused())?; if !parent.is_dir() - || id(&parent) != self.record.parent + || !self.matches_pin(self.record.parent, id(&parent)) || !socket.file_type().is_socket() || !private(&socket) - || id(&socket) != self.record.socket + || !self.matches_pin(self.record.socket, id(&socket)) || !file.is_file() || !private(&file) || file.nlink() != 1 @@ -570,6 +659,46 @@ impl Pin { } Ok(()) } + fn matches_pin(&self, recorded: (u64, u64), observed: (u64, u64)) -> bool { + self.device_rebind + .map_or(recorded == observed, |v| v.matches(recorded, observed)) + } + pub(in crate::provider::graph) fn legacy_summary( + candidate: &Candidate, + run: &str, + device_rebind: DeviceRebind, + host_boot_micros: u64, + ) -> Result { + let pin = Self::read_with_rebind(candidate, run, Some(device_rebind))?; + // SAFETY: geteuid has no arguments or side effects. + identity::verify( + &pin.record.process, + &pin.record.process, + &pin.record.process.executable, + unsafe { libc::geteuid() }, + )?; + device_rebind.definitely_dead_before_boot(&pin.record.process, host_boot_micros)?; + no_listener(&pin.root.join("control.sock"))?; + Ok(LegacyPublisher { + owner_sha256: format!("{:x}", Sha256::digest(&pin.bytes)), + parent: pin.record.parent, + socket: pin.record.socket, + record: pin.record_id, + process: pin.record.process, + }) + } + pub(in crate::provider::graph) fn legacy_recorded_device( + candidate: &Candidate, + run: &str, + ) -> Result { + let root = root(candidate, run)?; + state::check_private_directory(&root)?; + let record: Record = state::read_bounded(&root.join("owner.json"), 8192)?; + if record.version != 1 || record.candidate != candidate.checkout || record.run != run { + return Err(refused()); + } + Ok(record.parent.0) + } pub fn verify(&self) -> Result<(), CandidateError> { self.verify_files()?; let expected = &self.record.process; @@ -647,6 +776,16 @@ impl Pin { Ok(stream) } } + +#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +pub(in crate::provider::graph) struct LegacyPublisher { + pub owner_sha256: String, + pub parent: (u64, u64), + pub socket: (u64, u64), + pub record: (u64, u64), + pub process: ProcessIdentity, +} pub(super) struct Publication { listener: UnixListener, pin: Pin, @@ -660,6 +799,10 @@ impl Publication { Self::bind_mode(candidate, run, true) } fn bind_mode(candidate: &Candidate, run: &str, retired: bool) -> Result { + // Publication precedes the Engine lease. Recovery must exclude new runs + // at this boundary, rather than relying on a later provider lock alone. + let gate = super::super::publication_gate::Guard::acquire(candidate)?; + super::super::absent_publication_cleanup::publication_allowed(candidate, run, retired)?; let root = root(candidate, run)?; let lock = if retired { state::check_private_directory(&root).map_err(|_| refused())?; @@ -668,12 +811,19 @@ impl Publication { state::private_directory(&root).map_err(|_| refused())?; state::Lock::acquire(&root).map_err(|_| refused())? }; + // The foreground lock serializes a concurrent absence-intent writer. + super::super::absent_publication_cleanup::publication_allowed(candidate, run, retired)?; + #[cfg(target_os = "macos")] + if retired { + super::super::source_device_rebind::require_no_pending(candidate, run)?; + } for name in ["control.sock", "owner.json"] { match fs::symlink_metadata(root.join(name)) { Err(e) if e.kind() == std::io::ErrorKind::NotFound => {} _ => return Err(refused()), } } + gate.verify(candidate)?; let listener = UnixListener::bind(root.join("control.sock")).map_err(|_| refused())?; fs::set_permissions(root.join("control.sock"), fs::Permissions::from_mode(0o600)) .map_err(|_| refused())?; @@ -706,8 +856,10 @@ impl Publication { record, bytes, record_id, + device_rebind: None, }; pin.verify()?; + gate.verify(candidate)?; Ok(Self { listener, pin, @@ -718,6 +870,7 @@ impl Publication { self.listener.as_raw_fd() } pub fn verify(&self) -> Result<(), CandidateError> { + super::super::host_pin_recovery::exact_lock_path(&self.pin.root, &self._lock)?; self.pin.verify() } pub fn accept(&self) -> Result, CandidateError> { @@ -929,6 +1082,34 @@ pub(in crate::provider::graph) struct Retired { _lock: state::Lock, } impl Retired { + /// Obtain the process identity only from the exact completed retirement. + /// This does not infer ownership from pathname absence or PID death. + pub fn publisher_process( + &self, + candidate: &Candidate, + run: &str, + owner_sha256: &str, + complete_sha256: &str, + ) -> Result { + self.verify_recovery(candidate, run, owner_sha256, complete_sha256)?; + let intent: Retirement = state::read(&retirement_path(&self.root, owner_sha256)) + .map_err(|_| retirement_refused())?; + let (socket_original, record_original) = verify_retirement( + candidate, + run, + (owner_sha256, complete_sha256), + &self.root, + &self._lock, + &intent, + None, + )?; + if socket_original || record_original { + return Err(retirement_refused()); + } + self.verify()?; + Ok(intent.owner.process) + } + pub fn acquire(candidate: &Candidate, run: &str) -> Result, CandidateError> { let root = root(candidate, run)?; state::check_private_directory(&root)?; @@ -980,6 +1161,16 @@ impl Retired { run: &str, owner_sha256: &str, complete_sha256: &str, + ) -> Result<(), CandidateError> { + self.verify_recovery_with_rebind(candidate, run, owner_sha256, complete_sha256, None) + } + pub fn verify_recovery_with_rebind( + &self, + candidate: &Candidate, + run: &str, + owner_sha256: &str, + complete_sha256: &str, + device_rebind: Option, ) -> Result<(), CandidateError> { self.verify()?; if !super::super::hex(owner_sha256, 64) @@ -996,11 +1187,11 @@ impl Retired { let (socket_original, record_original) = verify_retirement( candidate, run, - owner_sha256, - complete_sha256, + (owner_sha256, complete_sha256), &self.root, &self._lock, &intent, + device_rebind, )?; if socket_original || record_original { return Err(retirement_refused()); @@ -1018,16 +1209,39 @@ impl DeadOwner { pub(in crate::provider::graph) fn acquire( candidate: &Candidate, run: &str, + ) -> Result { + Self::acquire_recovery(candidate, run, None, None) + } + pub(in crate::provider::graph) fn acquire_recovery( + candidate: &Candidate, + run: &str, + device_rebind: Option, + host_boot_micros: Option, ) -> Result { let lock = state::Lock::acquire_existing(&root(candidate, run)?)?; let value = Self { - pin: Pin::read(candidate, run)?, + pin: Pin::read_with_rebind(candidate, run, device_rebind)?, _lock: lock, }; + if let Some(rebind) = device_rebind { + rebind.definitely_dead_before_boot( + &value.pin.record.process, + host_boot_micros.ok_or_else(refused)?, + )?; + } value.verify()?; Ok(value) } pub(in crate::provider::graph) fn verify(&self) -> Result<(), CandidateError> { + let pathname = + fs::symlink_metadata(self.pin.root.join("operation.lock")).map_err(|_| refused())?; + if !pathname.is_file() + || pathname.nlink() != 1 + || !private(&pathname) + || id(&pathname) != self._lock.identity()? + { + return Err(refused()); + } self.pin.verify_files()?; if identity::alive(self.pin.record.process.pid)? { return Err(refused()); diff --git a/packages/runtime-core/src/provider/graph/foreground/transport_tests.rs b/packages/runtime-core/src/provider/graph/foreground/transport_tests.rs index 5b21b6a10..8b5dac3bc 100644 --- a/packages/runtime-core/src/provider/graph/foreground/transport_tests.rs +++ b/packages/runtime-core/src/provider/graph/foreground/transport_tests.rs @@ -1,6 +1,67 @@ use super::*; use serde_json::json; +#[test] +fn publication_refuses_a_replaced_operation_lock_and_preserves_it() { + let fixture = super::super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let run = "a".repeat(32); + let publication = Publication::bind(&candidate, &run).unwrap(); + publication.verify().unwrap(); + let directory = root(&candidate, &run).unwrap(); + fs::rename( + directory.join("operation.lock"), + directory.join("original.lock"), + ) + .unwrap(); + let replacement = state::Lock::acquire(&directory).unwrap(); + let replaced = fs::read(directory.join("operation.lock")).unwrap(); + assert!(publication.verify().is_err()); + assert_eq!( + fs::read(directory.join("operation.lock")).unwrap(), + replaced + ); + assert!(directory.join("original.lock").exists()); + drop(replacement); +} + +#[test] +fn pool_gate_excludes_publication_before_engine_for_every_run() { + let fixture = super::super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let gate = super::super::super::publication_gate::Guard::acquire(&candidate).unwrap(); + for run in ["a".repeat(32), "b".repeat(32)] { + assert!(Publication::bind(&candidate, &run).is_err()); + assert!(!root(&candidate, &run).unwrap().exists()); + } + gate.verify(&candidate).unwrap(); + drop(gate); + let mut first = Publication::bind(&candidate, &"a".repeat(32)).unwrap(); + // A publisher keeps its own run lock, not the pool gate for its lifetime. + let mut second = Publication::bind(&candidate, &"b".repeat(32)).unwrap(); + first.finish().unwrap(); + second.finish().unwrap(); + for run in ["a".repeat(32), "b".repeat(32)] { + fs::remove_dir_all(root(&candidate, &run).unwrap()).unwrap(); + } +} + +#[test] +fn pool_gate_refuses_a_replaced_lock_path() { + let fixture = super::super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let gate = super::super::super::publication_gate::Guard::acquire(&candidate).unwrap(); + let directory = candidate.state_root.join("run/graph-publication-gate"); + fs::rename( + directory.join("operation.lock"), + directory.join("original.lock"), + ) + .unwrap(); + let _replacement = state::Lock::acquire(&directory).unwrap(); + assert!(gate.verify(&candidate).is_err()); + assert!(directory.join("original.lock").exists()); +} + #[cfg(target_os = "macos")] fn abandoned_publisher() -> ( super::super::super::tests::Fixture, @@ -70,6 +131,114 @@ fn recovered_publisher_retirement_is_exact_and_idempotent() { fs::remove_dir_all(root).unwrap(); } +#[cfg(target_os = "macos")] +#[test] +fn cleanup_fence_refuses_effects_and_resumes_each_retirement_rename() { + for fail_at in [1, 3, 4, 5, 6] { + let (_fixture, candidate, run, owner, directory) = abandoned_publisher(); + let receipt = "f".repeat(64); + let before = fs::read(directory.join("owner.json")).unwrap(); + let sibling_run = "b".repeat(32); + let mut sibling = Publication::bind(&candidate, &sibling_run).unwrap(); + let sibling_root = root(&candidate, &sibling_run).unwrap(); + let sibling_before = fs::read(sibling_root.join("owner.json")).unwrap(); + let lock = state::Lock::acquire_existing(&directory).unwrap(); + let calls = std::cell::Cell::new(0); + let fence = || { + calls.set(calls.get() + 1); + if calls.get() == fail_at { + Err(retirement_refused()) + } else { + Ok(()) + } + }; + assert!( + retire_recovered_publisher_locked_fenced( + &candidate, &run, &owner, &receipt, None, &lock, &fence, + ) + .is_err() + ); + assert_eq!(calls.get(), fail_at); + let owner_path = if fail_at == 6 { + retired_path(&directory, &owner, false) + } else { + directory.join("owner.json") + }; + assert_eq!(fs::read(&owner_path).unwrap(), before); + let journal = retirement_path(&directory, &owner); + assert_eq!(journal.exists(), fail_at >= 4); + assert_eq!(directory.join("control.sock").exists(), fail_at < 5); + assert_eq!( + retired_path(&directory, &owner, true).exists(), + fail_at >= 5 + ); + if journal.exists() { + let retained = fs::read(&journal).unwrap(); + assert!( + retire_recovered_publisher_locked_fenced( + &candidate, + &run, + &owner, + &"e".repeat(64), + None, + &lock, + &|| Ok(()), + ) + .is_err() + ); + assert_eq!(fs::read(&journal).unwrap(), retained); + assert_eq!(fs::read(&owner_path).unwrap(), before); + } + retire_recovered_publisher_locked_fenced( + &candidate, + &run, + &owner, + &receipt, + None, + &lock, + &|| Ok(()), + ) + .unwrap(); + assert!(!directory.join("control.sock").exists()); + assert!(!directory.join("owner.json").exists()); + assert_eq!( + fs::read(retired_path(&directory, &owner, false)).unwrap(), + before + ); + let archived_owner = fs::symlink_metadata(retired_path(&directory, &owner, false)).unwrap(); + let archived_socket = fs::symlink_metadata(retired_path(&directory, &owner, true)).unwrap(); + let journal_bytes = fs::read(&journal).unwrap(); + retire_recovered_publisher_locked_fenced( + &candidate, + &run, + &owner, + &receipt, + None, + &lock, + &|| Ok(()), + ) + .unwrap(); + assert_eq!(fs::read(&journal).unwrap(), journal_bytes); + assert_eq!( + id(&fs::symlink_metadata(retired_path(&directory, &owner, false)).unwrap()), + id(&archived_owner) + ); + assert_eq!( + id(&fs::symlink_metadata(retired_path(&directory, &owner, true)).unwrap()), + id(&archived_socket) + ); + assert_eq!( + fs::read(sibling_root.join("owner.json")).unwrap(), + sibling_before + ); + sibling.verify().unwrap(); + sibling.finish().unwrap(); + fs::remove_dir_all(sibling_root).unwrap(); + drop(lock); + fs::remove_dir_all(directory).unwrap(); + } +} + #[cfg(target_os = "macos")] #[test] fn retired_recovery_refuses_pending_missing_and_partial_proof_without_mutation() { @@ -201,6 +370,78 @@ fn interrupted_socket_move_resumes_but_replaced_owner_refuses() { fs::remove_dir_all(root).unwrap(); } +#[cfg(target_os = "macos")] +#[test] +fn selected_legacy_device_socket_move_resumes_only_exact_retirement() { + let (_fixture, candidate, run, _original_owner, root) = abandoned_publisher(); + let receipt = "f".repeat(64); + let current = id(&fs::symlink_metadata(&root).unwrap()).0; + let rebind = DeviceRebind { + old: current.checked_add(1).unwrap(), + current, + }; + let mut record: Record = state::read(&root.join("owner.json")).unwrap(); + assert_eq!(record.parent.0, current); + assert_eq!(record.socket.0, current); + record.parent.0 = rebind.old; + record.socket.0 = rebind.old; + let bytes = serde_json::to_vec(&record).unwrap(); + fs::write(root.join("owner.json"), &bytes).unwrap(); + let owner = format!("{:x}", Sha256::digest(&bytes)); + assert!(Pin::read(&candidate, &run).is_err()); + let lock = state::Lock::acquire_existing(&root).unwrap(); + let pin = Pin::read_with_rebind(&candidate, &run, Some(rebind)).unwrap(); + let intent = Retirement { + version: 1, + candidate: candidate.checkout.clone(), + run: run.clone(), + receipt_sha256: receipt.clone(), + owner_sha256: owner.clone(), + parent: pin.record.parent, + lock: lock.identity().unwrap(), + socket: pin.record.socket, + record: pin.record_id, + owner: pin.record, + }; + state::write(&retirement_path(&root, &owner), &intent).unwrap(); + fs::rename(root.join("control.sock"), retired_path(&root, &owner, true)).unwrap(); + drop(lock); + assert!(retire_recovered_publisher(&candidate, &run, &owner, &receipt).is_err()); + retire_recovered_publisher_recovery(&candidate, &run, &owner, &receipt, Some(rebind)).unwrap(); + let retired = Retired::acquire(&candidate, &run).unwrap().unwrap(); + retired + .verify_recovery_with_rebind(&candidate, &run, &owner, &receipt, Some(rebind)) + .unwrap(); + assert!( + retired + .verify_recovery(&candidate, &run, &owner, &receipt) + .is_err() + ); + drop(retired); + fs::remove_dir_all(root).unwrap(); +} + +#[cfg(target_os = "macos")] +#[test] +fn dead_owner_guard_refuses_replaced_foreground_lock_path() { + let (_fixture, candidate, run, _owner, root) = abandoned_publisher(); + let dead = DeadOwner::acquire(&candidate, &run).unwrap(); + dead.verify().unwrap(); + let pathname = root.join("operation.lock"); + let saved = root.join("held-operation.lock"); + fs::rename(&pathname, &saved).unwrap(); + fs::write(&pathname, b"replacement").unwrap(); + fs::set_permissions(&pathname, fs::Permissions::from_mode(0o600)).unwrap(); + let replacement = id(&fs::symlink_metadata(&pathname).unwrap()); + assert!(dead.verify().is_err()); + assert_eq!(id(&fs::symlink_metadata(&pathname).unwrap()), replacement); + assert_eq!(fs::read(&pathname).unwrap(), b"replacement"); + drop(dead); + fs::remove_file(&pathname).unwrap(); + fs::rename(&saved, &pathname).unwrap(); + fs::remove_dir_all(root).unwrap(); +} + #[cfg(target_os = "macos")] #[test] fn replaced_socket_or_inherited_listener_refuses_retirement() { diff --git a/packages/runtime-core/src/provider/graph/host_pin_recovery.rs b/packages/runtime-core/src/provider/graph/host_pin_recovery.rs new file mode 100644 index 000000000..1846c7018 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/host_pin_recovery.rs @@ -0,0 +1,704 @@ +//! Explicit, selected legacy host-pin migration for one retained graph run. +//! +//! The private witness precedes dead-owner cleanup. It never edits historical graph, +//! publisher, control, or dependency receipts and grants no ordinary replay authority. +use super::{Candidate, CandidateError, Engine, Kind, Receipt, directory, foreground, load, state}; +use crate::provider::{ + ProjectShareIntent, + host_pin::DeviceRebind, + lifecycle, + relay_owner::publication::{LegacyControl, PinnedEndpoint}, +}; +use serde::{Deserialize, Serialize}; +use serde_json::{Value, json}; +use sha2::{Digest, Sha256}; +use std::{ + collections::BTreeMap, + fs::{self, OpenOptions}, + io::{Read, Seek, SeekFrom}, + os::unix::fs::{MetadataExt, OpenOptionsExt}, + path::{Path, PathBuf}, +}; + +use super::{dependency_slots::LegacyReservation, foreground::transport::LegacyPublisher}; + +const FILE: &str = "host-pin-recovery.json"; +const LIMIT: u64 = 64 * 1024; +const GUEST_IDENTITY_ACK: &str = "host-pin-guest-identity-v1\n"; + +fn refused() -> CandidateError { + CandidateError::new( + "graph_host_pin_recovery", + "Legacy host pin selection is incomplete or changed; receipts and data were preserved.", + ) +} +fn digest(bytes: &[u8]) -> String { + format!("{:x}", Sha256::digest(bytes)) +} + +/// Guest::boot_id is sourced from the host Owner receipt. Execute a read-only +/// guest check so a changed receipt cannot impersonate the running kernel boot. +/// execute_cleanup itself compares /proc boot ID and /storage owner under the +/// held Engine/Guest lease before this fixed acknowledgement is emitted. +pub(super) fn verify_guest_identity(engine: &Engine<'_>) -> Result<(), CandidateError> { + let acknowledged = engine + .guest() + .execute_cleanup("printf 'host-pin-guest-identity-v1\\n'", &[]) + .map_err(|_| refused())?; + if acknowledged != GUEST_IDENTITY_ACK { + return Err(refused()); + } + Ok(()) +} +pub(super) fn read_raw(path: &Path, limit: u64) -> Result, CandidateError> { + read_raw_with(path, limit, || {}) +} +fn read_raw_with( + path: &Path, + limit: u64, + after_read: impl FnOnce(), +) -> Result, CandidateError> { + let parent = path.parent().ok_or_else(refused)?; + state::check_private_directory(parent)?; + let mut file = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK) + .open(path) + .map_err(|_| refused())?; + let metadata = file.metadata().map_err(|_| refused())?; + // SAFETY: geteuid has no arguments or side effects. + if !metadata.is_file() + || metadata.nlink() != 1 + || metadata.mode() & 0o7777 != 0o600 + || metadata.uid() != unsafe { libc::geteuid() } + || metadata.len() == 0 + || metadata.len() > limit + { + return Err(refused()); + } + let mut bytes = Vec::new(); + file.by_ref() + .take(limit + 1) + .read_to_end(&mut bytes) + .map_err(|_| refused())?; + after_read(); + file.seek(SeekFrom::Start(0)).map_err(|_| refused())?; + let mut confirmation = Vec::new(); + file.by_ref() + .take(limit + 1) + .read_to_end(&mut confirmation) + .map_err(|_| refused())?; + let retained = file.metadata().map_err(|_| refused())?; + let pathname = fs::symlink_metadata(path).map_err(|_| refused())?; + if bytes != confirmation + || bytes.len() as u64 != metadata.len() + || [retained.dev(), pathname.dev()] != [metadata.dev(); 2] + || [retained.ino(), pathname.ino()] != [metadata.ino(); 2] + || [retained.len(), pathname.len()] != [metadata.len(); 2] + || [retained.uid(), pathname.uid()] != [metadata.uid(); 2] + || [retained.mode(), pathname.mode()] != [metadata.mode(); 2] + || [retained.nlink(), pathname.nlink()] != [metadata.nlink(); 2] + || [retained.mtime(), pathname.mtime()] != [metadata.mtime(); 2] + || [retained.mtime_nsec(), pathname.mtime_nsec()] != [metadata.mtime_nsec(); 2] + || [retained.ctime(), pathname.ctime()] != [metadata.ctime(); 2] + || [retained.ctime_nsec(), pathname.ctime_nsec()] != [metadata.ctime_nsec(); 2] + { + return Err(refused()); + } + Ok(bytes) +} +fn absent(path: &Path) -> Result { + match fs::symlink_metadata(path) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => Ok(true), + _ => Err(refused()), + } +} + +pub(super) fn exact_lock_path(root: &Path, held: &state::Lock) -> Result<(), CandidateError> { + let pathname = fs::symlink_metadata(root.join("operation.lock")).map_err(|_| refused())?; + if !pathname.is_file() + || pathname.nlink() != 1 + || (pathname.dev(), pathname.ino()) != held.identity()? + { + return Err(refused()); + } + Ok(()) +} + +#[derive(Clone, Debug, PartialEq, Eq, Deserialize, Serialize)] +#[serde(deny_unknown_fields)] +pub(super) struct Witness { + version: u8, + candidate: PathBuf, + run: String, + owner: String, + namespace: String, + plan: String, + graph_sha256: String, + provider_sha256: String, + host_boot_micros: u64, + previous_guest_boot: String, + current_guest_boot: String, + rebind: DeviceRebind, + publisher: LegacyPublisher, + control_root: PathBuf, + control: LegacyControl, + reservation: Option, + source_shared: Option, + retained_volumes: BTreeMap, + qualification: String, +} + +impl Witness { + pub(super) fn rebind(&self) -> DeviceRebind { + self.rebind + } + pub(super) fn host_boot_micros(&self) -> u64 { + self.host_boot_micros + } + pub(super) fn publisher_sha256(&self) -> &str { + &self.publisher.owner_sha256 + } + pub(super) fn reservation(&self) -> Option<&LegacyReservation> { + self.reservation.as_ref() + } + pub(super) fn control(&self) -> &LegacyControl { + &self.control + } + pub(super) fn control_root(&self) -> &Path { + &self.control_root + } + pub(super) fn graph_sha256(&self) -> &str { + &self.graph_sha256 + } + pub(super) fn matches_graph(&self, receipt: &Receipt) -> bool { + self.run == receipt.run + && self.owner == receipt.owner + && self.namespace == receipt.namespace + && self.plan == receipt.plan_id + && self.source_shared == receipt.source.as_ref().and_then(|s| s.shared.clone()) + } +} + +pub(super) struct Guard { + witness: Witness, + witness_path: PathBuf, + control_root: PathBuf, + control_lock: state::Lock, +} + +pub(super) struct CleanupProof<'a> { + pub original: &'a Receipt, + pub sha256: &'a str, + pub current_is_original: bool, + pub allow_absent_reservation: bool, + pub publisher_may_be_partial: bool, +} +impl Guard { + pub(super) fn witness(&self) -> &Witness { + &self.witness + } + pub(super) fn verify_lock(&self) -> Result<(), CandidateError> { + let retained: Witness = + serde_json::from_slice(&read_raw(&self.witness_path, LIMIT)?).map_err(|_| refused())?; + if retained != self.witness { + return Err(refused()); + } + let context = + super::host_relay::context(&self.witness.owner, &self.witness.previous_guest_boot)?; + let control = PinnedEndpoint::load_legacy_recovery( + &self.witness.control_root, + context, + self.witness.rebind, + self.witness.host_boot_micros, + )?; + if control.legacy_summary() != self.witness.control { + return Err(refused()); + } + exact_lock_path(&self.control_root, &self.control_lock) + } +} + +fn path(candidate: &Candidate, run: &str) -> Result { + Ok(directory(candidate, run)?.join(FILE)) +} +pub(super) fn load_witness( + candidate: &Candidate, + run: &str, +) -> Result, CandidateError> { + let file = path(candidate, run)?; + absent(&file.with_extension("pending"))?; + if absent(&file).is_ok() { + return Ok(None); + } + let bytes = read_raw(&file, LIMIT)?; + let witness: Witness = serde_json::from_slice(&bytes).map_err(|_| refused())?; + if witness.version != 1 || witness.candidate != candidate.checkout || witness.run != run { + return Err(refused()); + } + Ok(Some(witness)) +} + +/// Only the exact old publisher may select the historical overlay. A successor +/// with current-device pins uses ordinary strict recovery despite retained history. +pub(super) fn selected_for_old_publisher( + candidate: &Candidate, + run: &str, +) -> Result, CandidateError> { + let Some(witness) = load_witness(candidate, run)? else { + return Ok(None); + }; + let recorded = foreground::transport::Pin::legacy_recorded_device(candidate, run)?; + if recorded == witness.rebind.current { + return Ok(None); + } + if recorded != witness.rebind.old { + return Err(refused()); + } + Ok(Some(witness)) +} + +pub(super) fn verify_volume_projections( + engine: &Engine<'_>, + receipt: &Receipt, + selected: &BTreeMap, +) -> Result<(), CandidateError> { + let keys = receipt + .resources + .iter() + .filter(|(_, value)| value.kind == Kind::Volume) + .map(|(key, _)| key.clone()) + .collect::>(); + if keys.len() != selected.len() || keys.iter().any(|key| !selected.contains_key(key)) { + return Err(refused()); + } + for key in keys { + let resource = &receipt.resources[&key]; + let actual = super::inspect_resource(engine, receipt, resource)?.ok_or_else(refused)?; + let expected = selected.get(&key).ok_or_else(refused)?; + if expected["expected_name"] != resource.name + || expected["observed_name"] != actual["Name"] + || expected["observed_created_at"] != actual["CreatedAt"] + || expected["observed_mountpoint"] != actual["Mountpoint"] + || expected["observed_driver"] != actual["Driver"] + || expected["observed_labels_sha256"] + != digest(&serde_json::to_vec(&actual["Labels"]).map_err(|_| refused())?) + { + return Err(refused()); + } + let labels = expected["labels"].as_object().ok_or_else(refused)?; + if labels + .iter() + .any(|(name, value)| actual["Labels"].get(name) != Some(value)) + { + return Err(refused()); + } + } + Ok(()) +} + +/// Caller already holds the Engine lease and exact foreground owner lock. +/// The returned control lock remains held across graph cleanup effects. +pub(super) fn acquire_for_cleanup( + candidate: &Candidate, + run: &str, + engine: &Engine<'_>, + proof: CleanupProof<'_>, + witness: Witness, +) -> Result { + lifecycle::host_filesystem::no_auxiliary_update(candidate)?; + if witness.graph_sha256 != proof.sha256 + || !witness.matches_graph(proof.original) + || witness.candidate != candidate.checkout + || witness.run != run + || witness.qualification + != "explicit-legacy-device-rebind-original-volume-continuity-unproven" + || lifecycle::host_filesystem::host_boot_micros()? != witness.host_boot_micros + { + return Err(refused()); + } + witness + .rebind + .definitely_dead_before_boot(&witness.publisher.process, witness.host_boot_micros)?; + if proof.current_is_original + && digest(&read_raw( + &directory(candidate, run)?.join("state.json"), + 2 * 1024 * 1024, + )?) != witness.graph_sha256 + { + return Err(refused()); + } + let owner = state::Owner::load(candidate)?; + lifecycle::verify_disks(candidate, &owner)?; + verify_guest_identity(engine)?; + let bytes = read_raw( + &candidate.state_root.join("run/smolvm/owner.json"), + 1024 * 1024, + )?; + if digest(&bytes) != witness.provider_sha256 + || owner.token != witness.owner + || owner.guest_boot_id.as_deref() != Some(&witness.current_guest_boot) + || owner.previous_guest_boot_id.as_deref() != Some(&witness.previous_guest_boot) + || engine.guest().incarnation() != witness.owner + || engine.guest().boot_id() != witness.current_guest_boot + || owner + .storage + .as_ref() + .is_none_or(|disk| disk.device != witness.rebind.current) + || owner + .overlay + .as_ref() + .is_none_or(|disk| disk.device != witness.rebind.current) + { + return Err(refused()); + } + if !proof.publisher_may_be_partial { + let publisher = foreground::transport::Pin::legacy_summary( + candidate, + run, + witness.rebind, + witness.host_boot_micros, + )?; + if publisher != witness.publisher { + return Err(refused()); + } + } + let control_root = witness.control_root.join("relay-control"); + let control_lock = state::Lock::acquire_existing(&control_root)?; + absent(&foreground::transport::root(candidate, run)?.join("owner.pending"))?; + absent(&control_root.join("owner.pending"))?; + let context = super::host_relay::context(&witness.owner, &witness.previous_guest_boot)?; + let control = PinnedEndpoint::load_legacy_recovery( + &witness.control_root, + context, + witness.rebind, + witness.host_boot_micros, + )? + .legacy_summary(); + if control != witness.control { + return Err(refused()); + } + if let Some(selected) = &witness.reservation { + super::dependency_slots::verify_legacy_remaining( + candidate, + run, + witness.rebind, + selected, + proof.allow_absent_reservation, + )?; + } else if super::dependency_slots::inspect_legacy( + candidate, + run, + witness.rebind, + witness.host_boot_micros, + )? + .is_some() + { + return Err(refused()); + } + verify_volume_projections(engine, proof.original, &witness.retained_volumes)?; + let guard = Guard { + witness_path: path(candidate, run)?, + witness, + control_root, + control_lock, + }; + guard.verify_lock()?; + Ok(guard) +} + +struct Selected { + witness: Witness, + foreground_root: PathBuf, + foreground_lock: state::Lock, + control_lock_root: PathBuf, + control_lock: state::Lock, +} + +impl Selected { + fn verify( + &self, + candidate: &Candidate, + run: &str, + engine: &Engine<'_>, + ) -> Result<(), CandidateError> { + for (root, lock) in [ + (&self.foreground_root, &self.foreground_lock), + (&self.control_lock_root, &self.control_lock), + ] { + exact_lock_path(root, lock)?; + } + if self.foreground_root != foreground::transport::root(candidate, run)? + || self.control_lock_root != self.witness.control_root.join("relay-control") + || self.witness != select_content(candidate, run, engine)? + { + return Err(refused()); + } + Ok(()) + } +} + +fn select( + candidate: &Candidate, + run: &str, + engine: &Engine<'_>, +) -> Result { + let (receipt, _) = load(candidate, engine, run)?; + let startup = receipt.relay_startup.as_ref().ok_or_else(refused)?; + let foreground_root = foreground::transport::root(candidate, run)?; + let foreground_lock = state::Lock::acquire_existing(&foreground_root)?; + let control_lock_root = startup.control_root.join("relay-control"); + let control_lock = state::Lock::acquire_existing(&control_lock_root)?; + let selected = Selected { + witness: select_content(candidate, run, engine)?, + foreground_root, + foreground_lock, + control_lock_root, + control_lock, + }; + selected.verify(candidate, run, engine)?; + Ok(selected) +} + +fn select_content( + candidate: &Candidate, + run: &str, + engine: &Engine<'_>, +) -> Result { + lifecycle::host_filesystem::no_auxiliary_update(candidate)?; + let owner = state::Owner::load(candidate)?; + lifecycle::verify_disks(candidate, &owner)?; + verify_guest_identity(engine)?; + let host_boot_micros = lifecycle::host_filesystem::host_boot_micros()?; + let current_guest_boot = owner.guest_boot_id.clone().ok_or_else(refused)?; + let previous_guest_boot = owner.previous_guest_boot_id.clone().ok_or_else(refused)?; + if current_guest_boot == previous_guest_boot + || current_guest_boot != engine.guest().boot_id() + || owner.token != engine.guest().incarnation() + { + return Err(refused()); + } + let (receipt, root) = load(candidate, engine, run)?; + for name in [ + "state.pending", + "dead-owner-cleanup.pending", + "relay-cleanup-bridges.pending", + ] { + absent(&root.join(name))?; + } + if receipt.phase != "ready-observed" || receipt.relay_cleanup.is_some() { + return Err(refused()); + } + let startup = receipt.relay_startup.as_ref().ok_or_else(refused)?; + let graph_sha256 = digest(&read_raw(&root.join("state.json"), 2 * 1024 * 1024)?); + let provider_sha256 = digest(&read_raw( + &candidate.state_root.join("run/smolvm/owner.json"), + 1024 * 1024, + )?); + let publisher_root = foreground::transport::root(candidate, run)?; + absent(&publisher_root.join("owner.pending"))?; + absent(&startup.control_root.join("relay-control/owner.pending"))?; + let current_device = fs::symlink_metadata(&publisher_root) + .map_err(|_| refused())? + .dev(); + let rebind = DeviceRebind { + old: foreground::transport::Pin::legacy_recorded_device(candidate, run)?, + current: current_device, + }; + if rebind.old == rebind.current { + return Err(refused()); + } + if owner + .storage + .as_ref() + .is_none_or(|disk| disk.device != rebind.current) + || owner + .overlay + .as_ref() + .is_none_or(|disk| disk.device != rebind.current) + { + return Err(refused()); + } + let publisher = + foreground::transport::Pin::legacy_summary(candidate, run, rebind, host_boot_micros)?; + let context = super::host_relay::context(&receipt.owner, &previous_guest_boot)?; + let control = PinnedEndpoint::load_legacy_recovery( + &startup.control_root, + context, + rebind, + host_boot_micros, + )? + .legacy_summary(); + let reservation = + super::dependency_slots::inspect_legacy(candidate, run, rebind, host_boot_micros)?; + if receipt + .source + .as_ref() + .and_then(|source| source.shared.as_ref()) + .is_some_and(|shared| shared.device != rebind.old) + || owner.project_share.as_ref() + != receipt + .source + .as_ref() + .and_then(|s| s.shared.as_ref()) + .map(|shared| { + // Historical shared source is pinned to the old host device. Only the + // selected device may differ from the currently approved project share. + let mut expected = shared.clone(); + expected.device = rebind.current; + expected + }) + .as_ref() + { + return Err(refused()); + } + if let Some(share) = &owner.project_share { + share.validate()?; + } + let mut retained_volumes = BTreeMap::new(); + for (key, resource) in receipt + .resources + .iter() + .filter(|(_, value)| value.kind == Kind::Volume) + { + let actual = super::inspect_resource(engine, &receipt, resource)?.ok_or_else(refused)?; + let labels = super::expected_labels(&receipt, resource) + .as_object() + .ok_or_else(refused)? + .keys() + .map(|label| (label.clone(), actual["Labels"][label].clone())) + .collect::>(); + retained_volumes.insert( + key.clone(), + json!({"expected_name":resource.name,"observed_name":actual["Name"], + "observed_created_at":actual["CreatedAt"], + "observed_mountpoint":actual["Mountpoint"], + "observed_driver":actual["Driver"], + "observed_labels_sha256":digest(&serde_json::to_vec(&actual["Labels"]).map_err(|_| refused())?), + "labels":labels}), + ); + } + Ok(Witness { + version: 1, + candidate: candidate.checkout.clone(), + run: run.into(), + owner: receipt.owner, + namespace: receipt.namespace, + plan: receipt.plan_id, + graph_sha256, + provider_sha256, + host_boot_micros, + previous_guest_boot, + current_guest_boot, + rebind, + publisher, + control_root: startup.control_root.clone(), + control, + reservation, + source_shared: receipt.source.and_then(|s| s.shared), + retained_volumes, + qualification: "explicit-legacy-device-rebind-original-volume-continuity-unproven".into(), + }) +} + +fn selected_sha256(witness: &Witness) -> Result { + Ok(digest(&serde_json::to_vec(witness).map_err(|_| refused())?)) +} + +/// Read-only exact selection. Existing completed recovery history is retained. +pub fn inspect(candidate: &Candidate, run: &str) -> Result { + let engine = Engine::connect_cleanup_wait(candidate)?; + let selected = select(candidate, run, &engine)?; + Ok( + json!({"run":run,"selection_sha256":selected_sha256(&selected.witness)?,"old_device":selected.witness.rebind.old,"new_device":selected.witness.rebind.current,"qualification":selected.witness.qualification}), + ) +} + +/// Explicitly publish one private immutable witness before cleanup can consume an overlay. +pub fn recover(candidate: &Candidate, run: &str, expected: &str) -> Result { + if expected.len() != 64 || !expected.bytes().all(|b| b.is_ascii_hexdigit()) { + return Err(refused()); + } + let engine = Engine::connect_cleanup_wait(candidate)?; + let selected = select(candidate, run, &engine)?; + if selected_sha256(&selected.witness)? != expected { + return Err(refused()); + } + if let Some(existing) = load_witness(candidate, run)? { + if existing == selected.witness { + return Ok( + json!({"run":run,"selection_sha256":expected,"witness_published":true,"already_published":true,"qualification":selected.witness.qualification}), + ); + } + return Err(refused()); + } + selected.verify(candidate, run, &engine)?; + state::write(&path(candidate, run)?, &selected.witness)?; + Ok( + json!({"run":run,"selection_sha256":expected,"witness_published":true,"qualification":selected.witness.qualification}), + ) +} + +#[cfg(test)] +mod tests { + use super::*; + use std::{io::Write, os::unix::fs::PermissionsExt}; + + #[test] + fn raw_selection_refuses_same_inode_in_place_substitution() { + let fixture = super::super::tests::Fixture::new(); + let path = fixture.0.join("selected.json"); + fs::write(&path, b"first").unwrap(); + fs::set_permissions(&path, fs::Permissions::from_mode(0o600)).unwrap(); + let before = fs::symlink_metadata(&path).unwrap(); + let result = read_raw_with(&path, 64, || { + let mut file = OpenOptions::new().write(true).open(&path).unwrap(); + file.write_all(b"other").unwrap(); + file.sync_all().unwrap(); + }); + assert!(result.is_err()); + assert_eq!(fs::symlink_metadata(&path).unwrap().ino(), before.ino()); + } + + #[test] + fn device_overlay_requires_only_one_number_change_and_exact_inode() { + let translated = DeviceRebind { + old: 10, + current: 20, + }; + assert!(translated.matches((10, 77), (20, 77))); + for (old, current) in [ + ((20, 77), (20, 77)), + ((10, 77), (10, 77)), + ((10, 77), (20, 78)), + ((10, 0), (20, 0)), + ] { + assert!(!translated.matches(old, current)); + } + } + + #[test] + fn held_lock_path_replacement_refuses_and_preserves_foreign_replacement() { + let fixture = super::super::tests::Fixture::new(); + let lock = state::Lock::acquire(&fixture.0).unwrap(); + exact_lock_path(&fixture.0, &lock).unwrap(); + let pathname = fixture.0.join("operation.lock"); + let saved = fixture.0.join("held-operation.lock"); + fs::rename(&pathname, &saved).unwrap(); + fs::write(&pathname, b"foreign replacement").unwrap(); + fs::set_permissions(&pathname, fs::Permissions::from_mode(0o600)).unwrap(); + let replacement = fs::symlink_metadata(&pathname).unwrap(); + assert!(exact_lock_path(&fixture.0, &lock).is_err()); + let after = fs::symlink_metadata(&pathname).unwrap(); + assert_eq!( + (after.dev(), after.ino()), + (replacement.dev(), replacement.ino()) + ); + assert_eq!(fs::read(&pathname).unwrap(), b"foreign replacement"); + drop(lock); + // The fixture is private and isolated; restore the original pathname + // before its directory is torn down. + fs::remove_file(&pathname).unwrap(); + fs::rename(&saved, &pathname).unwrap(); + } +} diff --git a/packages/runtime-core/src/provider/graph/host_relay.rs b/packages/runtime-core/src/provider/graph/host_relay.rs index 94a5145d8..9c2af9652 100644 --- a/packages/runtime-core/src/provider/graph/host_relay.rs +++ b/packages/runtime-core/src/provider/graph/host_relay.rs @@ -2,7 +2,7 @@ use super::*; use crate::provider::{ host_endpoint::HostEndpoint, - relay_owner::{Context, Grant, GraphScope, RelayOwner}, + relay_owner::{Context, Grant, GraphScope, RelayOwner, publication::dead}, }; use sha2::{Digest, Sha256}; @@ -89,6 +89,7 @@ pub(super) fn cleanup_with_relay_expected( } let scope = graph_scope(context(&receipt.owner, engine.guest().boot_id())?, run)?; let environment = environment::cleanup_inventory(candidate, &engine, &receipt, &root)?; + bridges::cleanup::retire_exited_for_cleanup(candidate, &engine, &receipt)?; let bridge_selection = bridges::cleanup::capture(candidate, &engine, &receipt)?; let effect = cleanup_effect( &receipt, @@ -132,7 +133,7 @@ pub(super) fn cleanup_with_relay_expected( &bridge_selection, ) })?; - finish_confirmation(&mut coordinator, &mut cleaned, &root)?; + finish_confirmation(&mut coordinator, &mut cleaned, &root, None, None)?; Ok(cleaned) } @@ -205,7 +206,7 @@ pub fn resume_relay_cleanup( &bridge_selection, ) })?; - finish_confirmation(&mut coordinator, &mut cleaned, &root)?; + finish_confirmation(&mut coordinator, &mut cleaned, &root, None, None)?; Ok(cleaned) } @@ -218,6 +219,46 @@ pub fn confirm_relay_cleanup( remove_data: bool, control_root: &std::path::Path, selection: crate::provider::relay_owner::lifecycle_intent::Selection, +) -> Result { + confirm_relay_cleanup_fenced( + candidate, + run, + remove_data, + control_root, + selection, + None, + None, + ) +} + +pub(super) fn confirm_relay_cleanup_selected( + candidate: &Candidate, + run: &str, + remove_data: bool, + control_root: &std::path::Path, + selection: crate::provider::relay_owner::lifecycle_intent::Selection, + expected_identity: &str, + relay_selection_sha256: &str, +) -> Result { + confirm_relay_cleanup_fenced( + candidate, + run, + remove_data, + control_root, + selection, + Some(expected_identity), + Some(relay_selection_sha256), + ) +} + +fn confirm_relay_cleanup_fenced( + candidate: &Candidate, + run: &str, + remove_data: bool, + control_root: &std::path::Path, + selection: crate::provider::relay_owner::lifecycle_intent::Selection, + expected_identity: Option<&str>, + relay_selection_sha256: Option<&str>, ) -> Result { use crate::provider::relay_owner::lifecycle_intent::{Coordinator, Phase}; let engine = Engine::connect_cleanup(candidate)?; @@ -243,17 +284,53 @@ pub fn confirm_relay_cleanup( { return Err(refused()); } + let relay = match (expected_identity, relay_selection_sha256) { + (Some(identity), Some(sha256)) => { + use crate::provider::relay_owner::{lifecycle_intent::Inspection, publication::dead}; + let inspected = Inspection::load(control_root, expected)?; + let witness = + dead::CleanupWitness::acquire(control_root, expected, &inspected.process)?; + if inspected.recovery_fingerprint != identity + || inspected.owner == [0; 16] + || inspected.publication == [0; 32] + || witness.selection_sha256()? != sha256 + || witness + .present_identity() + .is_some_and(|(owner, publication)| { + inspected.owner != owner || inspected.publication != publication + }) + { + return Err(refused()); + } + Some(witness) + } + (None, None) => None, + _ => return Err(refused()), + }; let mut coordinator = Coordinator::resume(control_root, selection)?; + if let Some(identity) = expected_identity { + coordinator.verify_recovery_identity( + &marker.selection(), + graph_scope(expected, run)?, + identity, + )?; + } let inspect = || { - inspect_cleanup( + let observed = inspect_cleanup( candidate, &engine, &receipt, remove_data, &environment, &bridge_selection, - ) + )?; + if expected_identity.is_some() { + interrupted_start_cleanup::verify_retained_volumes(&engine, &receipt)?; + } + verify_selected_relay(relay.as_ref())?; + Ok(observed) }; + verify_selected_relay(relay.as_ref())?; match coordinator.phase() { Phase::EffectStarted => coordinator.confirm(effect, inspect)?, Phase::Confirmed if coordinator.acknowledgement_pending() => { @@ -261,14 +338,40 @@ pub fn confirm_relay_cleanup( } _ => return Err(refused()), } - finish_confirmation(&mut coordinator, &mut receipt, &root)?; + if let Some(identity) = expected_identity { + coordinator.verify_recovery_identity( + &marker.selection(), + graph_scope(expected, run)?, + identity, + )?; + } + verify_selected_relay(relay.as_ref())?; + finish_confirmation( + &mut coordinator, + &mut receipt, + &root, + relay.as_ref(), + expected_identity.map(|_| &engine), + )?; + verify_selected_relay(relay.as_ref())?; Ok(receipt) } +pub(super) fn verify_selected_relay( + relay: Option<&dead::CleanupWitness>, +) -> Result<(), CandidateError> { + if let Some(witness) = relay { + witness.verify()?; + } + Ok(()) +} + fn finish_confirmation( coordinator: &mut crate::provider::relay_owner::lifecycle_intent::Coordinator, receipt: &mut Receipt, root: &std::path::Path, + relay: Option<&dead::CleanupWitness>, + retained_engine: Option<&Engine<'_>>, ) -> Result<(), CandidateError> { use crate::provider::relay_owner::lifecycle_intent::Phase; if coordinator.phase() != Phase::Confirmed || !coordinator.acknowledgement_pending() { @@ -276,14 +379,28 @@ fn finish_confirmation( } #[cfg(test)] fault_pause(root, &receipt.run, "relay-before-confirmed-receipt")?; - let marker = receipt.relay_cleanup.as_mut().ok_or_else(refused)?; - if marker.operation != coordinator.operation() { + if receipt + .relay_cleanup + .as_ref() + .ok_or_else(refused)? + .operation + != coordinator.operation() + { return Err(refused()); } - marker.phase = cleanup_enrollment::Phase::Confirmed; + if let Some(engine) = retained_engine { + interrupted_start_cleanup::verify_retained_volumes(engine, receipt)?; + } + verify_selected_relay(relay)?; + receipt.relay_cleanup.as_mut().ok_or_else(refused)?.phase = + cleanup_enrollment::Phase::Confirmed; state::write(&root.join("state.json"), receipt)?; #[cfg(test)] fault_pause(root, &receipt.run, "relay-before-ack")?; + if let Some(engine) = retained_engine { + interrupted_start_cleanup::verify_retained_volumes(engine, receipt)?; + } + verify_selected_relay(relay)?; coordinator.acknowledge() } @@ -489,6 +606,66 @@ pub(super) fn inspect_cleanup( remove_data: bool, environment: &super::super::environment_recovery::GraphInventory, bridge_selection: &bridges::cleanup::Selection, +) -> Result<[u8; 32], CandidateError> { + inspect_cleanup_recovery( + candidate, + engine, + expected, + remove_data, + environment, + bridge_selection, + None, + ) +} +pub(super) fn inspect_cleanup_recovery( + candidate: &Candidate, + engine: &Engine<'_>, + expected: &Receipt, + remove_data: bool, + environment: &super::super::environment_recovery::GraphInventory, + bridge_selection: &bridges::cleanup::Selection, + host_pin: Option<&super::host_pin_recovery::Witness>, +) -> Result<[u8; 32], CandidateError> { + inspect_cleanup_selected( + candidate, + engine, + expected, + remove_data, + environment, + bridge_selection, + SelectedCleanupProof::Existing(host_pin), + ) +} +pub(super) fn inspect_cleanup_absence( + candidate: &Candidate, + engine: &Engine<'_>, + expected: &Receipt, + environment: &super::super::environment_recovery::GraphInventory, + bridge_selection: &bridges::cleanup::Selection, + witness: &super::absent_publication_cleanup::Selection, +) -> Result<[u8; 32], CandidateError> { + inspect_cleanup_selected( + candidate, + engine, + expected, + false, + environment, + bridge_selection, + SelectedCleanupProof::Absence(witness), + ) +} +enum SelectedCleanupProof<'a> { + Existing(Option<&'a super::host_pin_recovery::Witness>), + Absence(&'a super::absent_publication_cleanup::Selection), +} +fn inspect_cleanup_selected( + candidate: &Candidate, + engine: &Engine<'_>, + expected: &Receipt, + remove_data: bool, + environment: &super::super::environment_recovery::GraphInventory, + bridge_selection: &bridges::cleanup::Selection, + proof: SelectedCleanupProof<'_>, ) -> Result<[u8; 32], CandidateError> { let (receipt, root) = archive::load_confirmation(candidate, engine, &expected.run)?; match fs::symlink_metadata(root.join("state.pending")) { @@ -533,10 +710,33 @@ pub(super) fn inspect_cleanup( require_retained_volume(resource, value.is_some(), remove_data)?; observations.insert(key, value); } - if &bridges::cleanup::read(engine, &receipt, &root)? != bridge_selection { + let recorded = match proof { + SelectedCleanupProof::Absence(witness) => { + bridges::cleanup::read_recovery_absence(engine, &receipt, &root, witness)? + } + SelectedCleanupProof::Existing(host_pin) => { + bridges::cleanup::read_recovery(engine, &receipt, &root, host_pin)? + } + }; + if &recorded != bridge_selection { return Err(refused()); } - bridges::cleanup::verify(candidate, engine, &receipt, bridge_selection)?; + match proof { + SelectedCleanupProof::Absence(witness) => bridges::cleanup::verify_recovery_absence( + candidate, + engine, + &receipt, + bridge_selection, + witness, + )?, + SelectedCleanupProof::Existing(host_pin) => bridges::cleanup::verify_recovery( + candidate, + engine, + &receipt, + bridge_selection, + host_pin, + )?, + } environment::verify_cleanup(candidate, engine, &receipt, &root, environment)?; probes::verify_cleanup(engine, &receipt)?; startup::verify_cleanup(engine, &receipt)?; diff --git a/packages/runtime-core/src/provider/graph/hostname_change.rs b/packages/runtime-core/src/provider/graph/hostname_change.rs index 7f8242089..45a96fcef 100644 --- a/packages/runtime-core/src/provider/graph/hostname_change.rs +++ b/packages/runtime-core/src/provider/graph/hostname_change.rs @@ -213,6 +213,7 @@ pub(super) mod tests { }; let binding = source::SourceBinding { shared: None, + cache_scope: None, shared_contract: Some( project::live_source::Contract::from_plan(&plan.plan, snapshot.receipt()).unwrap(), ), diff --git a/packages/runtime-core/src/provider/graph/interrupted_start_cleanup.rs b/packages/runtime-core/src/provider/graph/interrupted_start_cleanup.rs new file mode 100644 index 000000000..e4ebcbd85 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/interrupted_start_cleanup.rs @@ -0,0 +1,521 @@ +//! Exact same-boot cleanup after a failed container start and an interrupted +//! enrolled cleanup. The old coordinator effect is observed, never executed again. +use super::*; +use crate::provider::{ + environment_recovery::GraphInventory, + relay_owner::{ + lifecycle_intent::{Inspection, Phase}, + publication::dead, + }, +}; +use sha2::{Digest, Sha256}; +use std::path::Path; + +mod recovery; +#[cfg(test)] +mod tests; + +const JOURNAL: &str = "interrupted-start-cleanup.json"; +const LIMIT: u64 = 2 * 1024 * 1024; + +fn refused() -> CandidateError { + error( + "graph_interrupted_start_cleanup", + "Interrupted start cleanup selection, owner or resource state changed; no effect was replayed.", + ) +} +fn digest(bytes: &[u8]) -> String { + format!("{:x}", Sha256::digest(bytes)) +} +fn no_pending(root: &Path) -> Result<(), CandidateError> { + state::check_private_directory(root)?; + for entry in fs::read_dir(root).map_err(state::io)? { + let name = entry.map_err(state::io)?.file_name(); + if name + .to_str() + .is_none_or(|value| value.ends_with(".pending")) + { + return Err(refused()); + } + } + for name in [ + "one-off.json", + "one-off-normalization.json", + "retired-data-removal.json", + ] { + match fs::symlink_metadata(root.join(name)) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => {} + _ => return Err(refused()), + } + } + Ok(()) +} +fn journal(root: &Path) -> Result, CandidateError> { + let path = root.join(JOURNAL); + match fs::symlink_metadata(&path) { + Err(e) if e.kind() == std::io::ErrorKind::NotFound => Ok(None), + Err(e) => Err(state::io(e)), + Ok(_) => state::read_bounded(&path, LIMIT).map(Some), + } +} + +#[derive(Clone, Debug, Serialize, Deserialize, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +struct Selection { + version: u8, + run: String, + owner: String, + boot: String, + receipt_sha256: String, + coordinator_sha256: String, + coordinator_identity_sha256: String, + foreground_sha256: String, + relay_sha256: String, + bridge_sha256: String, + environment_sha256: String, + reservation_sha256: String, + failed_key: String, + failed_id: String, +} +impl Selection { + fn digest(&self) -> Result { + Ok(digest(&serde_json::to_vec(self).map_err(|_| refused())?)) + } +} + +struct Proof<'a> { + candidate: &'a Candidate, + owner: foreground::transport::DeadOwner, + engine: Engine<'a>, + relay: dead::CleanupWitness, + relay_owner: [u8; 16], + relay_publication: [u8; 32], + root: PathBuf, + current: Receipt, + original: Receipt, + bridges: bridges::cleanup::Selection, + environment: GraphInventory, + selected: Selection, +} +impl Proof<'_> { + fn verify(&self) -> Result<(), CandidateError> { + self.owner.verify_retirement_ready()?; + self.relay.verify()?; + if self.relay.selection_sha256()? != self.selected.relay_sha256 { + return Err(refused()); + } + self.engine.guest().verify()?; + no_pending(&self.root)?; + let (current, root) = load(self.candidate, &self.engine, &self.selected.run)?; + let context = host_relay::context(&self.selected.owner, &self.selected.boot)?; + let marker = current.relay_cleanup.as_ref().ok_or_else(refused)?; + if root != self.root + || !["cleanup-intent", "stopped-data-retained"].contains(¤t.phase.as_str()) + || dead_owner_cleanup::immutable(¤t)? + != dead_owner_cleanup::immutable(&self.original)? + || current.owner != self.selected.owner + || self.engine.guest().boot_id() != self.selected.boot + || marker.phase != cleanup_enrollment::Phase::Pending + || !marker.valid() + || marker.runtime != context.runtime + || marker.boot != context.boot + || marker.control_root + != self + .original + .relay_cleanup + .as_ref() + .ok_or_else(refused)? + .control_root + || marker.operation + != self + .original + .relay_cleanup + .as_ref() + .ok_or_else(refused)? + .operation + || marker.effect + != self + .original + .relay_cleanup + .as_ref() + .ok_or_else(refused)? + .effect + { + return Err(refused()); + } + startup::require_dependency_rebind_complete(&root, ¤t)?; + initializer_cache::require_resolved(¤t)?; + if serde_json::to_value(environment::cleanup_inventory( + self.candidate, + &self.engine, + ¤t, + &root, + )?) + .map_err(|_| refused())? + != serde_json::to_value(&self.environment).map_err(|_| refused())? + { + return Err(refused()); + } + if digest(&host_pin_recovery::read_raw( + &root.join("relay-cleanup-bridges.json"), + 65536, + )?) != self.selected.bridge_sha256 + { + return Err(refused()); + } + bridges::cleanup::verify_recovery( + self.candidate, + &self.engine, + ¤t, + &self.bridges, + None, + )?; + verify_retained_volumes(&self.engine, ¤t)?; + let failed = current + .resources + .get(&self.selected.failed_key) + .ok_or_else(refused)?; + if let Some(value) = inspect_resource(&self.engine, ¤t, failed)? { + failed_created(failed, &value)?; + } + Ok(()) + } +} + +fn retained_volume_observation( + receipt: &Receipt, + resource: &Resource, + observed: Option<&Value>, +) -> Result<(), CandidateError> { + let value = observed.ok_or_else(refused)?; + if resource.kind != Kind::Volume + || resource.phase != "created" + || value["CreatedAt"].as_str().is_none_or(str::is_empty) + { + return Err(refused()); + } + match (&resource.cache, &resource.cache_provenance) { + (None, None) => Ok(()), + (Some(cache), Some(provenance)) + if cache.valid() + && resource.name == cache.name() + && provenance.valid(receipt, resource) => + { + Ok(()) + } + _ => Err(refused()), + } +} + +fn verify_retained_volume( + engine: &Engine<'_>, + receipt: &Receipt, + resource: &Resource, +) -> Result<(), CandidateError> { + let observed = inspect_resource(engine, receipt, resource)?; + retained_volume_observation(receipt, resource, observed.as_ref())?; + cache_provenance::verify(engine, receipt, resource) +} + +pub(super) fn verify_retained_volumes( + engine: &Engine<'_>, + receipt: &Receipt, +) -> Result<(), CandidateError> { + for resource in receipt + .resources + .values() + .filter(|resource| resource.kind == Kind::Volume) + { + verify_retained_volume(engine, receipt, resource)?; + } + Ok(()) +} + +fn failed_created(resource: &Resource, value: &Value) -> Result<(), CandidateError> { + let state = &value["State"]; + let prepared = shutdown::prepare(resource, value)?; + if resource.kind != Kind::Container + || resource.id.as_deref() != Some(prepared.id.as_str()) + || prepared.running + || state["Status"] != "created" + || !state["ExitCode"] + .as_u64() + .is_some_and(|n| (1..=255).contains(&n)) + || state["Pid"] != 0 + || state["OOMKilled"] != false + || value["RestartCount"] != 0 + || !state["StartedAt"] + .as_str() + .is_some_and(|s| s.starts_with("0001-")) + { + return Err(refused()); + } + Ok(()) +} + +fn failed_start( + receipt: &Receipt, + engine: &Engine<'_>, +) -> Result<(String, String), CandidateError> { + let mut failed = None; + for (key, resource) in &receipt.resources { + if resource.kind != Kind::Container { + continue; + } + let observed = inspect_resource(engine, receipt, resource)?; + let Some(value) = observed else { + if resource.id.is_some() || resource.phase != "reserved" { + return Err(refused()); + } + continue; + }; + if resource.id.as_deref() != value["Id"].as_str() { + return Err(refused()); + } + let state = &value["State"]; + let prepared = shutdown::prepare(resource, &value)?; + let restart_count = value["RestartCount"].as_u64().ok_or_else(refused)?; + if restart_count != 0 { + return Err(refused()); + } + if state["Status"] == "created" + && state["ExitCode"] + .as_u64() + .is_some_and(|n| (1..=255).contains(&n)) + { + if failed.is_some() + || resource.phase != "uncertain" + || failed_created(resource, &value).is_err() + { + return Err(refused()); + } + failed = Some((key.clone(), resource.id.clone().ok_or_else(refused)?)); + } else if prepared.running { + if resource.phase != "started" + || !state["FinishedAt"] + .as_str() + .is_some_and(|s| s.starts_with("0001-")) + { + return Err(refused()); + } + } else if state["Status"] != "exited" + || state["ExitCode"] != 0 + || receipt.readiness.get(&resource.key) != Some(&Condition::Completed) + { + return Err(refused()); + } + } + failed.ok_or_else(refused) +} + +fn no_current_shutdown(root: &Path, receipt: &Receipt) -> Result<(), CandidateError> { + let path = root.join("shutdown.json"); + if !path.exists() && !path.is_symlink() { + return Ok(()); + } + let value: Value = state::read_bounded(&path, 64 * 1024)?; + if value["version"] != 1 + || value["run"] != receipt.run + || value["owner"] != receipt.owner + || value["plan"] != receipt.plan_id + { + return Err(refused()); + } + let records = value["containers"].as_object().ok_or_else(refused)?; + if records.len() > MAX_SERVICES + || records.iter().any(|(key, terminal)| { + !receipt + .resources + .get(key) + .is_some_and(|r| r.kind == Kind::Container) + || !terminal["id"].as_str().is_some_and(|id| hex(id, 64)) + || receipt.resources[key].id.as_deref() == terminal["id"].as_str() + }) + { + return Err(refused()); + } + Ok(()) +} + +fn select<'a>(candidate: &'a Candidate, run: &str) -> Result, CandidateError> { + let prior_root = directory(candidate, run)?; + let existing = journal(&prior_root)?; + let reservation_sha256 = if let Some(prior) = &existing { + prior.selection.reservation_sha256.clone() + } else { + let inventory = dependency_slots::inspect(candidate)?; + let records = inventory["reservations"].as_array().ok_or_else(refused)?; + let matches = records + .iter() + .filter(|value| value["run"] == run) + .collect::>(); + if matches.len() != 1 || matches[0]["owner_alive"] != false { + return Err(refused()); + } + matches[0]["reservation"] + .as_str() + .filter(|value| hex(value, 64)) + .ok_or_else(refused)? + .to_owned() + }; + let owner = foreground::transport::DeadOwner::acquire(candidate, run)?; + owner.verify_retirement_ready()?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let (current, root) = load(candidate, &engine, run)?; + no_pending(&root)?; + let prior = journal(&root)?; + if serde_json::to_value(&prior).map_err(|_| refused())? + != serde_json::to_value(&existing).map_err(|_| refused())? + { + return Err(refused()); + } + let original = prior + .as_ref() + .map_or_else(|| current.clone(), |j| j.original.clone()); + if let Some(record) = &prior { + record.validate_phase(¤t)?; + } + if current.relay_startup.is_none() + || current.normalized_input.is_none() + || current + .relay_cleanup + .as_ref() + .is_none_or(|m| m.phase != cleanup_enrollment::Phase::Pending) + || !["cleanup-intent", "stopped-data-retained"].contains(¤t.phase.as_str()) + || (prior.is_none() && current.phase != "cleanup-intent") + || dead_owner_cleanup::immutable(¤t)? != dead_owner_cleanup::immutable(&original)? + { + return Err(refused()); + } + let receipt_bytes = host_pin_recovery::read_raw(&root.join("state.json"), LIMIT)?; + let original_sha256 = digest(&serde_json::to_vec_pretty(&original).map_err(|_| refused())?); + if prior.is_none() && digest(&receipt_bytes) != original_sha256 { + return Err(refused()); + } + let marker = current.relay_cleanup.as_ref().ok_or_else(refused)?; + let context = host_relay::context(¤t.owner, engine.guest().boot_id())?; + let relay = dead::CleanupWitness::acquire(&marker.control_root, context, owner.process())?; + let relay_sha256 = relay.selection_sha256()?; + let inspected = Inspection::load(&marker.control_root, context)?; + let present_identity = relay.present_identity(); + if inspected.phase != Phase::EffectStarted + || !inspected.acknowledgement_pending + || !inspected.selection_observed + || inspected.graph != Some(host_relay::graph_scope(context, run)?) + || inspected.selection.operation != marker.operation + || inspected.selection.effect != marker.effect + || inspected.selection.context != context + || inspected.owner == [0; 16] + || &inspected.process != owner.process() + || inspected.publication == [0; 32] + || present_identity.is_some_and(|(relay_owner, relay_publication)| { + inspected.owner != relay_owner || inspected.publication != relay_publication + }) + { + return Err(refused()); + } + let coordinator_sha256 = digest(&host_pin_recovery::read_raw( + &marker.control_root.join("relay-lifecycle/state.json"), + 128 * 1024, + )?); + let bridges = bridges::cleanup::read(&engine, &original, &root)?; + bridges::cleanup::verify_recovery(candidate, &engine, ¤t, &bridges, None)?; + let bridge_sha256 = digest(&host_pin_recovery::read_raw( + &root.join("relay-cleanup-bridges.json"), + 65536, + )?); + let environment = environment::cleanup_inventory(candidate, &engine, &original, &root)?; + let environment_sha256 = digest(&serde_json::to_vec(&environment).map_err(|_| refused())?); + if host_relay::cleanup_effect( + &original, + engine.guest().boot_id(), + false, + &(&environment, &bridges), + )? != marker.effect + { + return Err(refused()); + } + let (failed_key, failed_id) = if let Some(j) = &prior { + ( + j.selection.failed_key.clone(), + j.selection.failed_id.clone(), + ) + } else { + no_current_shutdown(&root, &original)?; + failed_start(&original, &engine)? + }; + let selected = Selection { + version: 1, + run: run.into(), + owner: current.owner.clone(), + boot: engine.guest().boot_id().into(), + receipt_sha256: original_sha256, + coordinator_sha256, + coordinator_identity_sha256: inspected.recovery_fingerprint, + foreground_sha256: owner.fingerprint(), + relay_sha256, + bridge_sha256, + environment_sha256, + reservation_sha256, + failed_key, + failed_id, + }; + if prior.as_ref().is_some_and(|j| { + j.selection != selected || j.selection.digest().ok() != Some(j.selection_sha256.clone()) + }) { + return Err(refused()); + } + let proof = Proof { + candidate, + owner, + engine, + relay, + relay_owner: inspected.owner, + relay_publication: inspected.publication, + root, + current, + original, + bridges, + environment, + selected, + }; + proof.verify()?; + Ok(proof) +} + +/// Read-only, value-free selection for exactly one interrupted failed start. +pub fn inspect(candidate: &Candidate, run: &str) -> Result { + let root = directory(candidate, run)?; + if let Some(record) = journal(&root)? { + record.validate_basic(run, &record.selection_sha256)?; + if record.graph_complete { + return recovery::inspect_completed(candidate, run, &record); + } + } + let proof = select(candidate, run)?; + if let Some(journal) = journal(&proof.root)? { + recovery::validate_retry(&proof, &journal)?; + } + Ok( + json!({"run":run,"phase":proof.current.phase,"eligible":true, + "selection_sha256":proof.selected.digest()?,"data_retained":true,"same_boot":true}), + ) +} + +/// Presentation hint only. The exact selected recovery still performs every +/// native proof before any effect; a malformed journal blocks inspection. +pub(super) fn incomplete(root: &Path, run: &str) -> Result { + let Some(record) = journal(root)? else { + return Ok(false); + }; + record.validate_basic(run, &record.selection_sha256)?; + Ok(!record.reservation_released) +} + +/// Recover only a selected same-boot failed start, retaining every named volume. +pub fn recover(candidate: &Candidate, run: &str, expected: &str) -> Result { + if !hex(expected, 64) { + return Err(refused()); + } + recovery::recover(candidate, run, expected) +} diff --git a/packages/runtime-core/src/provider/graph/interrupted_start_cleanup/recovery.rs b/packages/runtime-core/src/provider/graph/interrupted_start_cleanup/recovery.rs new file mode 100644 index 000000000..f39495924 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/interrupted_start_cleanup/recovery.rs @@ -0,0 +1,676 @@ +//! A fresh journal for one interrupted cleanup. A pending step is never sent +//! again: only independently observed terminal or absent state advances it. +use super::*; +use crate::provider::relay_owner::lifecycle_intent::Coordinator; + +fn stale_generation() -> CandidateError { + error( + "graph_interrupted_start_cleanup_stale", + "Completed interrupted cleanup evidence belongs to an earlier resource generation; no old cleanup effect was replayed.", + ) +} + +#[derive(Clone, Debug, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +pub(super) struct Journal { + pub(super) version: u8, + pub(super) selection: Selection, + pub(super) selection_sha256: String, + pub(super) original: Receipt, + pub(super) steps: Vec, + pub(super) cursor: usize, + pub(super) pending: bool, + pub(super) graph_complete: bool, + pub(super) publisher_retired: bool, + pub(super) reservation_released: bool, +} + +#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] +#[serde(rename_all = "snake_case", deny_unknown_fields)] +pub(super) enum Step { + Stop(String), + Delete(String), + Startup, + Probe(String), + Environment(String), +} + +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub(super) enum StopDecision { + AlreadyTerminal, + IssueOneStop, +} + +pub(super) fn stop_decision( + resource: &Resource, + observed: &Value, + pending: bool, +) -> Result { + let prepared = shutdown::prepare(resource, observed)?; + if prepared.id != resource.id.as_deref().ok_or_else(refused)? { + return Err(refused()); + } + if pending { + shutdown::terminal(resource, observed, true)?; + Ok(StopDecision::AlreadyTerminal) + } else if prepared.running { + Ok(StopDecision::IssueOneStop) + } else { + shutdown::terminal(resource, observed, false)?; + Ok(StopDecision::AlreadyTerminal) + } +} + +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub(super) enum Acknowledgement { + Confirm, + Complete, +} + +pub(super) fn acknowledgement( + marker: cleanup_enrollment::Phase, + coordinator: Phase, + pending: bool, + selected: bool, +) -> Result { + if !selected { + return Err(refused()); + } + match (marker, coordinator, pending) { + (cleanup_enrollment::Phase::Pending, Phase::EffectStarted, true) + | (cleanup_enrollment::Phase::Pending, Phase::Confirmed, true) + | (cleanup_enrollment::Phase::Confirmed, Phase::Confirmed, true) => { + Ok(Acknowledgement::Confirm) + } + (cleanup_enrollment::Phase::Confirmed, Phase::Confirmed, false) => { + Ok(Acknowledgement::Complete) + } + _ => Err(refused()), + } +} + +pub(super) fn steps(receipt: &Receipt, slots: &[String]) -> Result, CandidateError> { + if receipt.resources.len() > MAX_SERVICES + MAX_NETWORKS + MAX_SERVICES * 8 + || receipt.probes.len() > MAX_SERVICES + { + return Err(refused()); + } + let mut result = Vec::new(); + for (key, resource) in &receipt.resources { + if resource.kind == Kind::Container && resource.id.is_some() { + result.push(Step::Stop(key.clone())); + } + } + for kind in [Kind::Container, Kind::Network] { + for (key, resource) in &receipt.resources { + if resource.kind == kind { + result.push(Step::Delete(key.clone())); + } + } + } + result.push(Step::Startup); + for name in receipt.probes.keys() { + result.push(Step::Probe(name.clone())); + } + for slot in slots { + result.push(Step::Environment(slot.clone())); + } + Ok(result) +} + +impl Journal { + fn new(proof: &Proof<'_>) -> Result { + Ok(Self { + version: 1, + selection: proof.selected.clone(), + selection_sha256: proof.selected.digest()?, + original: proof.original.clone(), + steps: steps(&proof.original, &proof.environment.slots())?, + cursor: 0, + pending: false, + graph_complete: false, + publisher_retired: false, + reservation_released: false, + }) + } + pub(super) fn validate_basic(&self, run: &str, expected: &str) -> Result<(), CandidateError> { + if self.version != 1 + || self.selection.version != 1 + || self.selection.run != run + || self.original.run != run + || self.original.owner != self.selection.owner + || self.selection.digest()? != self.selection_sha256 + || self.selection_sha256 != expected + || self.cursor > self.steps.len() + || (self.graph_complete && (self.cursor != self.steps.len() || self.pending)) + || (self.publisher_retired && !self.graph_complete) + || (self.reservation_released && !self.publisher_retired) + || !self + .original + .resources + .get(&self.selection.failed_key) + .is_some_and(|resource| { + resource.kind == Kind::Container + && resource.phase == "uncertain" + && resource.id.as_deref() == Some(self.selection.failed_id.as_str()) + }) + || digest(&serde_json::to_vec_pretty(&self.original).map_err(|_| refused())?) + != self.selection.receipt_sha256 + { + return Err(refused()); + } + Ok(()) + } + pub(super) fn validate_phase(&self, current: &Receipt) -> Result<(), CandidateError> { + match current.phase.as_str() { + "cleanup-intent" if !self.graph_complete => Ok(()), + "stopped-data-retained" if self.cursor == self.steps.len() && !self.pending => Ok(()), + _ => Err(refused()), + } + } + fn write(&self, root: &Path) -> Result<(), CandidateError> { + state::write(&root.join(JOURNAL), self) + } +} + +pub(super) fn validate_retry(proof: &Proof<'_>, journal: &Journal) -> Result<(), CandidateError> { + journal.validate_basic(&proof.selected.run, &proof.selected.digest()?)?; + journal.validate_phase(&proof.current)?; + if journal.selection != proof.selected + || journal.steps != steps(&proof.original, &proof.environment.slots())? + || journal.original.run != proof.original.run + || dead_owner_cleanup::immutable(&journal.original)? + != dead_owner_cleanup::immutable(&proof.original)? + { + return Err(refused()); + } + Ok(()) +} + +fn fence( + proof: &Proof<'_>, + coordinator: &Coordinator, + journal: &Journal, +) -> Result<(), CandidateError> { + proof.verify()?; + journal.validate_basic(&proof.selected.run, &proof.selected.digest()?)?; + let current = super::journal(&proof.root)?.ok_or_else(refused)?; + if serde_json::to_value(¤t).map_err(|_| refused())? + != serde_json::to_value(journal).map_err(|_| refused())? + { + return Err(refused()); + } + let marker = proof.current.relay_cleanup.as_ref().ok_or_else(refused)?; + let graph = host_relay::graph_scope(marker.selection().context, &proof.selected.run)?; + coordinator.verify_selected_effect( + &marker.selection(), + graph, + proof.relay_owner, + proof.owner.process(), + proof.relay_publication, + &proof.selected.coordinator_identity_sha256, + )?; + if digest(&host_pin_recovery::read_raw( + &marker.control_root.join("relay-lifecycle/state.json"), + 128 * 1024, + )?) != proof.selected.coordinator_sha256 + { + return Err(refused()); + } + Ok(()) +} + +fn absent(proof: &Proof<'_>, resource: &Resource) -> Result { + if inspect_resource(&proof.engine, &proof.current, resource)?.is_some() { + return Ok(false); + } + let mut by_name = resource.clone(); + by_name.id = None; + Ok(inspect_resource(&proof.engine, &proof.current, &by_name)?.is_none()) +} + +fn advance_stop( + proof: &Proof<'_>, + coordinator: &Coordinator, + journal: &mut Journal, + key: &str, +) -> Result<(), CandidateError> { + let resource = proof.original.resources.get(key).ok_or_else(refused)?; + let value = inspect_resource(&proof.engine, &proof.current, resource)?.ok_or_else(refused)?; + if key == proof.selected.failed_key { + failed_created(resource, &value)?; + } + if stop_decision(resource, &value, journal.pending)? == StopDecision::IssueOneStop { + let prepared = shutdown::prepare(resource, &value)?; + journal.pending = true; + journal.write(&proof.root)?; + fence(proof, coordinator, journal)?; + proof + .engine + .stop_containers(&[(prepared.id.clone(), u64::from(prepared.grace_seconds))])?; + let latest = + inspect_resource(&proof.engine, &proof.current, resource)?.ok_or_else(refused)?; + shutdown::terminal(resource, &latest, true)?; + } + journal.cursor += 1; + journal.pending = false; + journal.write(&proof.root) +} + +fn advance_delete( + proof: &Proof<'_>, + coordinator: &Coordinator, + journal: &mut Journal, + key: &str, +) -> Result<(), CandidateError> { + let resource = proof.original.resources.get(key).ok_or_else(refused)?; + if resource.kind != Kind::Container && resource.kind != Kind::Network { + return Err(refused()); + } + if journal.pending || resource.id.is_none() { + if !absent(proof, resource)? { + return Err(refused()); + } + } else { + let value = + inspect_resource(&proof.engine, &proof.current, resource)?.ok_or_else(refused)?; + let id = value["Id"] + .as_str() + .filter(|id| Some(*id) == resource.id.as_deref()) + .ok_or_else(refused)?; + if resource.kind == Kind::Container { + shutdown::terminal(resource, &value, false)?; + } + journal.pending = true; + journal.write(&proof.root)?; + fence(proof, coordinator, journal)?; + proof.engine.request( + Method::DELETE, + &format!( + "/v1.53/{}/{id}{}", + resource.kind.collection(), + if resource.kind == Kind::Container { + "?v=true" + } else { + "" + } + ), + None, + )?; + if !absent(proof, resource)? { + return Err(refused()); + } + } + journal.cursor += 1; + journal.pending = false; + journal.write(&proof.root) +} + +fn advance_other( + proof: &Proof<'_>, + coordinator: &Coordinator, + journal: &mut Journal, + step: &Step, +) -> Result<(), CandidateError> { + let verify = || -> Result<(), CandidateError> { + match step { + Step::Startup => startup::verify_cleanup(&proof.engine, &proof.current), + Step::Probe(name) => probes::verify_one_absent(&proof.engine, &proof.current, name), + Step::Environment(slot) => { + super::super::super::environment_recovery::verify_graph_slot_retired( + proof.engine.guest(), + &proof.environment, + slot, + ) + } + _ => Err(refused()), + } + }; + if journal.pending { + verify()?; + } else if verify().is_err() { + journal.pending = true; + journal.write(&proof.root)?; + fence(proof, coordinator, journal)?; + match step { + Step::Startup => startup::cleanup_guest(&proof.engine, &proof.current, true)?, + Step::Probe(name) => probes::retire_one(&proof.engine, &proof.current, name)?, + Step::Environment(slot) => super::super::super::environment_recovery::retire( + proof.candidate, + proof.engine.guest(), + slot, + None, + )?, + _ => return Err(refused()), + } + verify()?; + } + journal.cursor += 1; + journal.pending = false; + journal.write(&proof.root) +} + +fn complete_graph(proof: &Proof<'_>, journal: &mut Journal) -> Result<(), CandidateError> { + if journal.cursor != journal.steps.len() || journal.pending { + return Err(refused()); + } + let mut cleaned = proof.current.clone(); + for resource in cleaned.resources.values_mut() { + match resource.kind { + Kind::Volume => { + if resource.phase != "created" { + return Err(refused()); + } + } + _ => resource.phase = "absent".into(), + } + } + for probe in cleaned.probes.values_mut() { + probe.phase = "retired".into(); + } + cleaned.phase = "stopped-data-retained".into(); + for resource in proof + .original + .resources + .values() + .filter(|r| r.kind != Kind::Volume) + { + if !absent(proof, resource)? { + return Err(refused()); + } + } + startup::verify_cleanup(&proof.engine, &cleaned)?; + probes::verify_cleanup(&proof.engine, &cleaned)?; + environment::verify_cleanup( + proof.candidate, + &proof.engine, + &cleaned, + &proof.root, + &proof.environment, + )?; + verify_retained_volumes(&proof.engine, &proof.current)?; + state::write(&proof.root.join("state.json"), &cleaned)?; + host_relay::inspect_cleanup( + proof.candidate, + &proof.engine, + &cleaned, + false, + &proof.environment, + &proof.bridges, + )?; + journal.graph_complete = true; + journal.write(&proof.root) +} + +pub(super) fn recover( + candidate: &Candidate, + run: &str, + expected: &str, +) -> Result { + let root = directory(candidate, run)?; + if let Some(existing) = journal(&root)? { + existing.validate_basic(run, expected)?; + if existing.graph_complete { + return finish(candidate, run, &existing); + } + } + let proof = select(candidate, run)?; + if proof.selected.digest()? != expected { + return Err(refused()); + } + let mut journal = match journal(&proof.root)? { + Some(existing) => { + validate_retry(&proof, &existing)?; + existing + } + None => { + let created = Journal::new(&proof)?; + created.write(&proof.root)?; + created + } + }; + let marker = proof.current.relay_cleanup.as_ref().ok_or_else(refused)?; + let coordinator = Coordinator::resume(&marker.control_root, marker.selection())?; + fence(&proof, &coordinator, &journal)?; + while journal.cursor < journal.steps.len() { + fence(&proof, &coordinator, &journal)?; + let step = journal.steps[journal.cursor].clone(); + match &step { + Step::Stop(key) => advance_stop(&proof, &coordinator, &mut journal, key)?, + Step::Delete(key) => advance_delete(&proof, &coordinator, &mut journal, key)?, + _ => advance_other(&proof, &coordinator, &mut journal, &step)?, + } + } + if !journal.graph_complete { + fence(&proof, &coordinator, &journal)?; + complete_graph(&proof, &mut journal)?; + } + drop(coordinator); + drop(proof); + finish(candidate, run, &journal) +} + +fn validate_finished<'a>( + candidate: &'a Candidate, + run: &str, + record: &Journal, +) -> Result<(Engine<'a>, Receipt, PathBuf, bool, dead::CleanupWitness), CandidateError> { + record.validate_basic(run, &record.selection_sha256)?; + if !record.graph_complete { + return Err(refused()); + } + let selection = &record.selection; + let engine = Engine::connect_cleanup_wait(candidate)?; + let (receipt, root) = load(candidate, &engine, run)?; + if receipt.phase != "stopped-data-retained" { + return Err(stale_generation()); + } + no_pending(&root)?; + let persisted = journal(&root)?.ok_or_else(refused)?; + if serde_json::to_value(&persisted).map_err(|_| refused())? + != serde_json::to_value(record).map_err(|_| refused())? + || receipt.phase != "stopped-data-retained" + || receipt.owner != selection.owner + || engine.guest().boot_id() != selection.boot + { + return Err(refused()); + } + let mut expected_receipt = record.original.clone(); + expected_receipt.phase = "stopped-data-retained".into(); + for resource in expected_receipt.resources.values_mut() { + if resource.kind != Kind::Volume { + resource.phase = "absent".into(); + } + } + for probe in expected_receipt.probes.values_mut() { + probe.phase = "retired".into(); + } + expected_receipt + .relay_cleanup + .as_mut() + .ok_or_else(refused)? + .phase = receipt.relay_cleanup.as_ref().ok_or_else(refused)?.phase; + if serde_json::to_value(&receipt).map_err(|_| refused())? + != serde_json::to_value(&expected_receipt).map_err(|_| refused())? + { + return Err(refused()); + } + verify_retained_volumes(&engine, &receipt)?; + let inventory = environment::cleanup_inventory(candidate, &engine, &receipt, &root)?; + if record.steps != steps(&record.original, &inventory.slots())? + || digest(&serde_json::to_vec(&inventory).map_err(|_| refused())?) + != selection.environment_sha256 + { + return Err(refused()); + } + let marker = receipt.relay_cleanup.as_ref().ok_or_else(refused)?; + let original_marker = record.original.relay_cleanup.as_ref().ok_or_else(refused)?; + if !marker.valid() + || marker.runtime != original_marker.runtime + || marker.boot != original_marker.boot + || marker.operation != original_marker.operation + || marker.effect != original_marker.effect + || marker.control_root != original_marker.control_root + { + return Err(refused()); + } + let context = host_relay::context(&selection.owner, &selection.boot)?; + let inspected = Inspection::load(&marker.control_root, context)?; + let relay = dead::CleanupWitness::acquire(&marker.control_root, context, &inspected.process)?; + if inspected.graph != Some(host_relay::graph_scope(context, run)?) + || inspected.selection.context != marker.selection().context + || inspected.selection.operation != marker.operation + || inspected.selection.effect != marker.effect + || inspected.recovery_fingerprint != selection.coordinator_identity_sha256 + || inspected.owner == [0; 16] + || inspected.publication == [0; 32] + || relay.selection_sha256()? != selection.relay_sha256 + || relay + .present_identity() + .is_some_and(|(relay_owner, relay_publication)| { + inspected.owner != relay_owner || inspected.publication != relay_publication + }) + || digest(&host_pin_recovery::read_raw( + &root.join("relay-cleanup-bridges.json"), + 65536, + )?) != selection.bridge_sha256 + { + return Err(refused()); + } + if marker.phase == cleanup_enrollment::Phase::Pending + && inspected.phase == Phase::EffectStarted + && digest(&host_pin_recovery::read_raw( + &marker.control_root.join("relay-lifecycle/state.json"), + 128 * 1024, + )?) != selection.coordinator_sha256 + { + return Err(refused()); + } + let status = acknowledgement( + marker.phase, + inspected.phase, + inspected.acknowledgement_pending, + inspected.selection_observed, + )?; + if record.publisher_retired && status != Acknowledgement::Complete { + return Err(refused()); + } + let bridges = bridges::cleanup::read(&engine, &receipt, &root)?; + host_relay::inspect_cleanup(candidate, &engine, &receipt, false, &inventory, &bridges)?; + Ok(( + engine, + receipt, + root, + status == Acknowledgement::Confirm, + relay, + )) +} + +pub(super) fn inspect_completed( + candidate: &Candidate, + run: &str, + record: &Journal, +) -> Result { + let (_engine, _receipt, _root, _needs_confirmation, _relay) = + validate_finished(candidate, run, record)?; + Ok( + json!({"run":run,"phase":"stopped-data-retained","eligible":true, + "selection_sha256":record.selection_sha256,"data_retained":true,"same_boot":true}), + ) +} + +fn finish(candidate: &Candidate, run: &str, record: &Journal) -> Result { + let (engine, mut receipt, root, needs_confirmation, relay) = + validate_finished(candidate, run, record)?; + let selection = &record.selection; + if needs_confirmation { + let marker = receipt.relay_cleanup.as_ref().ok_or_else(refused)?; + let control_root = marker.control_root.clone(); + let selected = marker.selection(); + relay.verify()?; + drop(relay); + drop(engine); + receipt = host_relay::confirm_relay_cleanup_selected( + candidate, + run, + false, + &control_root, + selected, + &selection.coordinator_identity_sha256, + &selection.relay_sha256, + )?; + } else { + if receipt.relay_cleanup.as_ref().ok_or_else(refused)?.phase + != cleanup_enrollment::Phase::Confirmed + { + return Err(refused()); + } + drop(relay); + drop(engine); + } + let engine = Engine::connect_cleanup_wait(candidate)?; + let (current, current_root) = load(candidate, &engine, run)?; + if current_root != root + || serde_json::to_value(¤t).map_err(|_| refused())? + != serde_json::to_value(&receipt).map_err(|_| refused())? + { + return Err(refused()); + } + let environment = environment::cleanup_inventory(candidate, &engine, ¤t, &root)?; + let bridges = bridges::cleanup::read(&engine, ¤t, &root)?; + host_relay::inspect_cleanup(candidate, &engine, ¤t, false, &environment, &bridges)?; + startup::archive_dependency_rebind_after_cleanup( + &root, + &record.original, + ¤t, + &selection.boot, + )?; + let receipt_sha256 = digest(&host_pin_recovery::read_raw( + &root.join("state.json"), + LIMIT, + )?); + drop(engine); + let publisher = acknowledged_publisher::AcknowledgedPublisherSelection { + run, + owner: &selection.owner, + receipt_sha256: &receipt_sha256, + publisher_sha256: &selection.foreground_sha256, + }; + acknowledged_publisher::retire(candidate, publisher)?; + let mut progress = record.clone(); + if !progress.publisher_retired { + verify_journal_exact(&root, &progress)?; + progress.publisher_retired = true; + progress.write(&root)?; + } + let publisher = acknowledged_publisher::AcknowledgedPublisherSelection { + run, + owner: &selection.owner, + receipt_sha256: &receipt_sha256, + publisher_sha256: &selection.foreground_sha256, + }; + acknowledged_publisher::release_dependencies( + candidate, + publisher, + &selection.reservation_sha256, + )?; + if !progress.reservation_released { + verify_journal_exact(&root, &progress)?; + progress.reservation_released = true; + progress.write(&root)?; + } + Ok( + json!({"run":run,"phase":"stopped-data-retained","recovered":true, + "data_retained":true,"same_boot":true,"publisher_retired":true,"reservation_released":true}), + ) +} + +fn verify_journal_exact(root: &Path, expected: &Journal) -> Result<(), CandidateError> { + let observed = super::journal(root)?.ok_or_else(refused)?; + if serde_json::to_value(observed).map_err(|_| refused())? + != serde_json::to_value(expected).map_err(|_| refused())? + { + return Err(refused()); + } + Ok(()) +} diff --git a/packages/runtime-core/src/provider/graph/interrupted_start_cleanup/tests.rs b/packages/runtime-core/src/provider/graph/interrupted_start_cleanup/tests.rs new file mode 100644 index 000000000..b1b136b1d --- /dev/null +++ b/packages/runtime-core/src/provider/graph/interrupted_start_cleanup/tests.rs @@ -0,0 +1,322 @@ +use super::recovery::{ + Acknowledgement, Journal, Step, StopDecision, acknowledgement, steps, stop_decision, +}; +use super::*; +use crate::provider::{identity, relay_owner::Context}; +use std::process::{Command, Stdio}; +use std::sync::atomic::{AtomicU64, Ordering}; + +struct ShortRoot(PathBuf); +impl ShortRoot { + fn new() -> Self { + static NEXT: AtomicU64 = AtomicU64::new(0); + let root = fs::canonicalize("/tmp").unwrap().join(format!( + "hgack-{}-{}", + std::process::id(), + NEXT.fetch_add(1, Ordering::Relaxed) + )); + state::private_directory(&root).unwrap(); + Self(root) + } +} +impl Drop for ShortRoot { + fn drop(&mut self) { + let _ = fs::remove_dir_all(&self.0); + } +} + +fn original() -> Receipt { + serde_json::from_value(json!({ + "version": 1, + "run": "a".repeat(32), + "owner": "b".repeat(32), + "namespace": "private", + "plan_id": "c".repeat(64), + "phase": "cleanup-intent", + "readiness": {}, + "resources": { + "container:failed": {"kind":"container","key":"failed","name":"failed","id":"d".repeat(64),"image":"sha256:image","phase":"uncertain"}, + "container:reserved": {"kind":"container","key":"reserved","name":"reserved","id":null,"image":"sha256:image","phase":"reserved"}, + "network:private": {"kind":"network","key":"private","name":"private","id":"e".repeat(64),"image":null,"phase":"created"}, + "volume:data": {"kind":"volume","key":"data","name":"data","id":null,"image":null,"phase":"created"} + } + })).unwrap() +} + +fn selected(original: &Receipt) -> Selection { + Selection { + version: 1, + run: original.run.clone(), + owner: original.owner.clone(), + boot: "f".repeat(32), + receipt_sha256: digest(&serde_json::to_vec_pretty(original).unwrap()), + coordinator_sha256: "1".repeat(64), + coordinator_identity_sha256: "7".repeat(64), + foreground_sha256: "2".repeat(64), + relay_sha256: "3".repeat(64), + bridge_sha256: "4".repeat(64), + environment_sha256: "5".repeat(64), + reservation_sha256: "6".repeat(64), + failed_key: "container:failed".into(), + failed_id: "d".repeat(64), + } +} + +#[test] +fn retained_plain_volume_requires_created_receipt_and_present_inspection() { + let receipt = original(); + let volume = &receipt.resources["volume:data"]; + let observed = json!({"CreatedAt":"2026-09-18T00:00:00Z"}); + assert!(retained_volume_observation(&receipt, volume, Some(&observed)).is_ok()); + assert!(retained_volume_observation(&receipt, volume, None).is_err()); + assert!(retained_volume_observation(&receipt, volume, Some(&json!({}))).is_err()); + let mut changed = volume.clone(); + changed.phase = "absent".into(); + assert!(retained_volume_observation(&receipt, &changed, Some(&observed)).is_err()); +} + +#[test] +fn retained_cache_requires_exact_binding_and_provenance() { + let mut receipt = original(); + let mut volume = receipt.resources["volume:data"].clone(); + let cache = dependency_cache::CacheBinding { + scope: "a".repeat(64), + fingerprint: "b".repeat(64), + image: format!("sha256:{}", "c".repeat(64)), + }; + volume.name = cache.name(); + volume.cache = Some(cache); + volume.cache_provenance = Some( + serde_json::from_value(json!({ + "version":1, + "origin":"adopted", + "boot":"prior-boot", + "identity":{"created_at":"2026-09-18T00:00:00Z","directory":"1:2"}, + "initializers":[], + "completed":{} + })) + .unwrap(), + ); + receipt + .resources + .insert("volume:data".into(), volume.clone()); + let observed = json!({"CreatedAt":"2026-09-18T00:00:00Z"}); + assert!(retained_volume_observation(&receipt, &volume, Some(&observed)).is_ok()); + assert!(retained_volume_observation(&receipt, &volume, None).is_err()); + + let mut missing_provenance = volume.clone(); + missing_provenance.cache_provenance = None; + assert!(retained_volume_observation(&receipt, &missing_provenance, Some(&observed)).is_err()); + let mut stale_name = volume.clone(); + stale_name.name = "replacement".into(); + assert!(retained_volume_observation(&receipt, &stale_name, Some(&observed)).is_err()); + let mut stale_binding = volume.clone(); + stale_binding.cache.as_mut().unwrap().fingerprint = "d".repeat(64); + assert!(retained_volume_observation(&receipt, &stale_binding, Some(&observed)).is_err()); + let mut malformed_provenance = volume; + malformed_provenance.cache_provenance = Some( + serde_json::from_value(json!({ + "version":1, + "origin":"adopted", + "boot":"prior-boot", + "identity":{"created_at":"2026-09-18T00:00:00Z","directory":"missing"}, + "initializers":[], + "completed":{} + })) + .unwrap(), + ); + assert!(retained_volume_observation(&receipt, &malformed_provenance, Some(&observed)).is_err()); +} + +#[test] +fn step_plan_never_deletes_retained_volume_and_keeps_reserved_name_check() { + let planned = steps(&original(), &["slot-one".into()]).unwrap(); + assert_eq!( + planned, + vec![ + Step::Stop("container:failed".into()), + Step::Delete("container:failed".into()), + Step::Delete("container:reserved".into()), + Step::Delete("network:private".into()), + Step::Startup, + Step::Environment("slot-one".into()), + ] + ); +} + +#[test] +fn journal_binds_original_and_selection_and_refuses_invalid_completion() { + let original = original(); + let selection = selected(&original); + let expected = selection.digest().unwrap(); + let steps = steps(&original, &[]).unwrap(); + let mut journal = Journal { + version: 1, + selection, + selection_sha256: expected.clone(), + original, + steps, + cursor: 0, + pending: false, + graph_complete: false, + publisher_retired: false, + reservation_released: false, + }; + assert!(journal.validate_basic(&"a".repeat(32), &expected).is_ok()); + assert!(journal.validate_phase(&journal.original).is_ok()); + let mut stopped = journal.original.clone(); + stopped.phase = "stopped-data-retained".into(); + assert!(journal.validate_phase(&stopped).is_err()); + journal.cursor = journal.steps.len(); + assert!(journal.validate_phase(&stopped).is_ok()); + journal.pending = true; + assert!(journal.validate_phase(&stopped).is_err()); + journal.pending = false; + assert!(journal.validate_basic(&"b".repeat(32), &expected).is_err()); + journal.cursor = 0; + journal.graph_complete = true; + assert!(journal.validate_basic(&"a".repeat(32), &expected).is_err()); + journal.graph_complete = false; + journal + .original + .resources + .get_mut("volume:data") + .unwrap() + .name = "replacement".into(); + assert!(journal.validate_basic(&"a".repeat(32), &expected).is_err()); +} + +#[test] +fn uncertain_stop_never_reissues_to_still_running_or_failed_created_target() { + let receipt = original(); + let resource = &receipt.resources["container:failed"]; + let created = json!({ + "Id":resource.id,"RestartCount":0, + "Config":{"StopTimeout":10,"StopSignal":null}, + "State":{"Paused":false,"Restarting":false,"Dead":false,"OOMKilled":false, + "ExitCode":128,"Running":false,"Pid":0,"Status":"created", + "StartedAt":"0001-01-01T00:00:00Z"} + }); + assert!(failed_created(resource, &created).is_ok()); + assert_eq!( + stop_decision(resource, &created, false).unwrap(), + StopDecision::AlreadyTerminal + ); + assert!(stop_decision(resource, &created, true).is_err()); + let mut now_running = created.clone(); + now_running["State"]["Status"] = json!("running"); + now_running["State"]["Running"] = json!(true); + now_running["State"]["Pid"] = json!(42); + now_running["State"]["ExitCode"] = json!(0); + now_running["State"]["StartedAt"] = json!("2026-10-02T12:00:00Z"); + assert!(failed_created(resource, &now_running).is_err()); + assert_eq!( + stop_decision(resource, &now_running, false).unwrap(), + StopDecision::IssueOneStop + ); + assert!(stop_decision(resource, &now_running, true).is_err()); + let mut exited = now_running; + exited["State"]["Status"] = json!("exited"); + exited["State"]["Running"] = json!(false); + exited["State"]["Pid"] = json!(0); + assert_eq!( + stop_decision(resource, &exited, true).unwrap(), + StopDecision::AlreadyTerminal + ); +} + +#[test] +fn all_ack_crash_windows_require_the_exact_marker_pairing() { + use cleanup_enrollment::Phase as Marker; + assert_eq!( + acknowledgement(Marker::Pending, Phase::EffectStarted, true, true).unwrap(), + Acknowledgement::Confirm + ); + // Coordinator confirmation can persist before the graph marker is written. + assert_eq!( + acknowledgement(Marker::Pending, Phase::Confirmed, true, true).unwrap(), + Acknowledgement::Confirm + ); + assert_eq!( + acknowledgement(Marker::Confirmed, Phase::Confirmed, true, true).unwrap(), + Acknowledgement::Confirm + ); + assert_eq!( + acknowledgement(Marker::Confirmed, Phase::Confirmed, false, true).unwrap(), + Acknowledgement::Complete + ); + for (marker, coordinator, pending, selected) in [ + (Marker::Pending, Phase::Confirmed, false, true), + (Marker::Pending, Phase::Intent, true, true), + (Marker::Confirmed, Phase::EffectStarted, true, true), + (Marker::Confirmed, Phase::Confirmed, false, false), + ] { + assert!(acknowledgement(marker, coordinator, pending, selected).is_err()); + } +} + +#[test] +fn incomplete_hint_tracks_publication_and_reservation_without_overwriting_history() { + let fixture = super::super::tests::Fixture::new(); + let original = original(); + let selection = selected(&original); + let sha = selection.digest().unwrap(); + let planned = steps(&original, &[]).unwrap(); + let mut record = Journal { + version: 1, + selection, + selection_sha256: sha, + original, + cursor: planned.len(), + steps: planned, + pending: false, + graph_complete: true, + publisher_retired: false, + reservation_released: false, + }; + state::write(&fixture.0.join(JOURNAL), &record).unwrap(); + assert!(incomplete(&fixture.0, &"a".repeat(32)).unwrap()); + record.publisher_retired = true; + state::write(&fixture.0.join(JOURNAL), &record).unwrap(); + assert!(incomplete(&fixture.0, &"a".repeat(32)).unwrap()); + record.reservation_released = true; + state::write(&fixture.0.join(JOURNAL), &record).unwrap(); + assert!(!incomplete(&fixture.0, &"a".repeat(32)).unwrap()); + record.publisher_retired = false; + state::write(&fixture.0.join(JOURNAL), &record).unwrap(); + assert!(incomplete(&fixture.0, &"a".repeat(32)).is_err()); +} + +#[test] +fn ack_boundary_rechecks_absent_publication_after_confirmed_receipt_write() { + let fixture = ShortRoot::new(); + let control = fixture.0.join("relay-control"); + state::private_directory(&control).unwrap(); + drop(state::Lock::acquire(&control).unwrap()); + let mut child = Command::new("/bin/cat") + .stdin(Stdio::piped()) + .stdout(Stdio::null()) + .stderr(Stdio::null()) + .spawn() + .unwrap(); + let process = identity::observe(child.id() as i32).unwrap(); + drop(child.stdin.take()); + assert!(child.wait().unwrap().success()); + assert!(state::check_private_directory(&fixture.0).is_ok()); + assert!(state::Lock::acquire_existing(&control).is_ok()); + assert!(identity::verify(&process, &process, &process.executable, process.uid).is_ok()); + assert!(!identity::alive(process.pid).unwrap()); + let witness = dead::CleanupWitness::acquire( + &fixture.0, + Context { + runtime: [1; 16], + boot: [2; 16], + }, + &process, + ) + .unwrap(); + state::write(&fixture.0.join("state.json"), &json!({"phase":"confirmed"})).unwrap(); + assert!(host_relay::verify_selected_relay(Some(&witness)).is_ok()); + fs::write(control.join("control.sock"), b"replacement").unwrap(); + assert!(host_relay::verify_selected_relay(Some(&witness)).is_err()); +} diff --git a/packages/runtime-core/src/provider/graph/live_owner_cleanup.rs b/packages/runtime-core/src/provider/graph/live_owner_cleanup.rs index 33f1aa01e..84fed4e83 100644 --- a/packages/runtime-core/src/provider/graph/live_owner_cleanup.rs +++ b/packages/runtime-core/src/provider/graph/live_owner_cleanup.rs @@ -20,6 +20,19 @@ struct Intent { listeners_retired: bool, complete_sha256: Option, } +/// Dispatch only an exact completed recovery generation; history is inert. +pub(super) fn current_completion( + root: &std::path::Path, + receipt: &Receipt, +) -> Result { + if !exists(&root.join(FILE))? { + return Ok(false); + } + let intent: Intent = state::read(&root.join(FILE))?; + let complete = digest(receipt)?; + Ok(intent.complete_sha256.as_deref() == Some(complete.as_str())) +} + fn refused() -> CandidateError { error( "graph_live_owner_recovery", @@ -45,7 +58,6 @@ fn no_pending(root: &Path, receipt: &Receipt) -> Result<(), CandidateError> { "one-off.json", "one-off.pending", "one-off-normalization.json", - "dependency-rebind.json", "dependency-rebind.pending", "dead-owner-cleanup.pending", "relay-cleanup-bridges.pending", @@ -54,6 +66,7 @@ fn no_pending(root: &Path, receipt: &Receipt) -> Result<(), CandidateError> { return Err(refused()); } } + startup::require_dependency_rebind_complete(root, receipt)?; dead_owner_cleanup::require_historical_recovery(root, receipt) } fn ready(receipt: &Receipt) -> Result<(), CandidateError> { @@ -294,6 +307,7 @@ pub fn recover_live_owner( if intent.foreground_sha256 != foreground.fingerprint() || intent.relay != relay.selection() { return Err(refused()); } + startup::require_dependency_rebind_recovery_complete(&root, &intent.original, &intent.boot)?; foreground.verify_retirement_ready()?; relay.verify()?; host_relay::cleanup_preflight(&engine, &receipt, &root, false)?; @@ -325,6 +339,8 @@ pub fn recover_live_owner( relay.verify()?; intent.complete_sha256 = Some(digest(&cleaned)?); save(&root, &intent)?; + // Commit independently verified owned absence before archival. A crash during + // a rename resumes through the completed intent, never through host adoption. drop(relay); drop(foreground); finish_retirement(candidate, &engine, &cleaned, &root, &intent)?; @@ -368,6 +384,9 @@ fn finish_retirement( } startup::verify_cleanup(engine, receipt)?; probes::verify_cleanup(engine, receipt)?; + archive_completed_refresh(root, &intent.original, receipt, &intent.boot, &|| { + engine.guest().verify() + })?; let startup = receipt.relay_startup.as_ref().ok_or_else(refused)?; dead::retire( &startup.control_root, @@ -380,10 +399,79 @@ fn finish_retirement( &intent.foreground_sha256, &digest(receipt)?, )?; - dependency_slots::recover_cleaned(candidate, receipt)?; + dependency_slots::recover_cleaned(candidate, receipt, None)?; engine.guest().verify() } +/// Completed cleanup grants archival, not refresh replay. The same exact ready +/// generation is checked again on retry, including a journal already moved by +/// an interrupted owned rename. +fn archive_completed_refresh( + root: &Path, + original: &Receipt, + cleaned: &Receipt, + boot: &str, + verify: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + let generation = service_exec_generation(original)?; + if exists(&root.join("dependency-rebind.json"))? + || exists(&root.join(format!("dependency-rebind-history-{generation}")))? + { + startup::archive_retired_dependency_rebind(root, original, cleaned, boot, verify) + } else { + startup::require_dependency_rebind_recovery_complete(root, original, boot) + } +} + +/// A superseded cleanup is inert history. Prefer its exact stopped generation; +/// verified truncation may establish replacement of a validated original after +/// that completion is evicted. Neither path supplies current cleanup authority. +pub(super) fn require_historical_recovery( + root: &Path, + current: &Receipt, +) -> Result<(), CandidateError> { + if exists(&root.join("live-owner-cleanup.pending"))? { + return Err(refused()); + } + if !exists(&root.join(FILE))? { + return Ok(()); + } + let prior: Intent = state::read(&root.join(FILE))?; + let complete = prior.complete_sha256.as_deref().ok_or_else(refused)?; + let stopped = restore_history::completed_for_recovery(root, current, complete)?; + let historical = stopped.as_ref().unwrap_or(&prior.original); + validate(&prior, historical, &prior.original_sha256, &prior.boot)?; + if stopped.is_none() + && (prior.original.version != current.version + || prior.original.run != current.run + || prior.original.owner != current.owner + || prior.original.namespace != current.namespace + || prior.original.plan_id != current.plan_id + || !restore_history::confirms_truncated_newer_generation( + root, + current, + &prior.original, + )?) + { + return Err(refused()); + } + if !prior.listeners_retired + || !historical.resources.iter().any(|(key, old)| { + old.kind == Kind::Container + && old.id.as_ref().is_some_and(|id| { + current + .resources + .get(key) + .and_then(|now| now.id.as_ref()) + .is_some_and(|current_id| current_id != id) + }) + }) + { + return Err(refused()); + } + Ok(()) +} + pub(super) fn retained(root: &Path, receipt: &Receipt) -> Result { if !exists(&root.join(FILE))? { return Ok(false); @@ -460,6 +548,26 @@ mod tests { fn receipt() -> Receipt { serde_json::from_value(json!({"version":1,"run":"a".repeat(32),"owner":"b".repeat(32),"namespace":"c".repeat(64),"plan_id":"d".repeat(64),"phase":"ready-observed","readiness":{},"resources":{},"relay_startup":{"control_only":true,"guest_root":null,"control_root":"/private/owned","artifact":"e".repeat(64),"services":{}}})).unwrap() } + #[test] + fn current_recovery_dispatch_does_not_select_historical_completion() { + let fixture = super::super::tests::Fixture::new(); + let mut receipt = receipt(); + receipt.phase = "stopped-data-retained".into(); + assert!(!current_completion(&fixture.0, &receipt).unwrap()); + let mut proof = completed(receipt.clone()); + proof.complete_sha256 = Some(digest(&receipt).unwrap()); + let path = fixture.0.join(FILE); + state::write(&path, &proof).unwrap(); + let before = fs::read(&path).unwrap(); + assert!(current_completion(&fixture.0, &receipt).unwrap()); + receipt.owner = "9".repeat(32); + assert!(!current_completion(&fixture.0, &receipt).unwrap()); + assert_eq!(fs::read(&path).unwrap(), before); + fs::write(&path, b"unconfirmed").unwrap(); + assert!(current_completion(&fixture.0, &receipt).is_err()); + assert_eq!(fs::read(&path).unwrap(), b"unconfirmed"); + } + #[test] fn only_fully_ready_receipts_are_admitted() { let mut value = receipt(); @@ -497,6 +605,80 @@ mod tests { } } #[test] + fn completed_terminal_refresh_is_not_a_pending_mutation() { + let fixture = super::super::tests::Fixture::new(); + let mut receipt = receipt(); + receipt + .readiness + .insert("init".into(), Condition::Completed); + receipt.relay_startup.as_mut().unwrap().services.insert( + "init".into(), + serde_json::from_value(json!({"generation":"f".repeat(32), + "phase":"completed","started_at":null,"bindings":{}})) + .unwrap(), + ); + let journal = json!({"version":1,"operation":"e".repeat(32), + "run":receipt.run,"owner":receipt.owner,"boot":"fixture-boot", + "expected_generation":"a".repeat(64),"phase":"completed", + "slots":{},"processes":{},"completed_services":["init"], + "completed_generation":service_exec_generation(&receipt).unwrap()}); + let path = fixture.0.join("dependency-rebind.json"); + state::write(&path, &journal).unwrap(); + let before = fs::read(&path).unwrap(); + no_pending(&fixture.0, &receipt).unwrap(); + startup::require_dependency_rebind_recovery_complete(&fixture.0, &receipt, "fixture-boot") + .unwrap(); + assert_eq!(fs::read(&path).unwrap(), before); + let mut partial = journal.clone(); + partial["phase"] = json!("provisioning"); + state::write(&path, &partial).unwrap(); + let before = fs::read(&path).unwrap(); + assert!(no_pending(&fixture.0, &receipt).is_err()); + assert_eq!(fs::read(&path).unwrap(), before); + let mut cleaned = receipt.clone(); + cleaned.phase = "stopped-data-retained".into(); + let generation = service_exec_generation(&receipt).unwrap(); + let history = fixture + .0 + .join(format!("dependency-rebind-history-{generation}")); + let mut replaced = journal.clone(); + replaced["completed_generation"] = json!("9".repeat(64)); + state::write(&path, &replaced).unwrap(); + let before = fs::read(&path).unwrap(); + assert!( + archive_completed_refresh(&fixture.0, &receipt, &cleaned, "fixture-boot", &|| Ok(()),) + .is_err() + ); + assert_eq!(fs::read(&path).unwrap(), before); + assert!(!history.exists()); + state::write(&path, &journal).unwrap(); + let before = fs::read(&path).unwrap(); + let interrupted = || { + if history.join("proof.json").exists() && path.exists() { + Err(refused()) + } else { + Ok(()) + } + }; + assert!(archive_completed_refresh( + &fixture.0, &receipt, &cleaned, "fixture-boot", &interrupted, + ).is_err()); + assert_eq!(fs::read(&path).unwrap(), before); + fs::rename(&path, history.join("dependency-rebind.json")).unwrap(); + for _ in 0..2 { + archive_completed_refresh(&fixture.0, &receipt, &cleaned, "fixture-boot", &|| Ok(())) + .unwrap(); + } + assert_eq!( + fs::read(history.join("dependency-rebind.json")).unwrap(), + before + ); + assert_eq!( + state::read::(&history.join("proof.json")).unwrap()["complete"], + true + ); + } + #[test] fn immutable_selection_allows_cleanup_progress_but_not_resource_replacement() { let value = receipt(); let mut changed = value.clone(); @@ -538,6 +720,71 @@ mod tests { } } #[test] + fn historical_recovery_survives_eviction_without_lending_current_authority() { + let fixture = super::super::tests::Fixture::new(); + let prior = completed(generation(1, "ready-observed")); + state::write(&fixture.0.join(FILE), &prior).unwrap(); + for number in 1..=12 { + restore_history::retain(&fixture.0, &generation(number, "stopped-data-retained")) + .unwrap(); + } + let current = generation(13, "ready-observed"); + assert!( + restore_history::completed_for_recovery( + &fixture.0, + ¤t, + prior.complete_sha256.as_deref().unwrap(), + ) + .unwrap() + .is_none() + ); + let before = fs::read(fixture.0.join(FILE)).unwrap(); + require_historical_recovery(&fixture.0, ¤t).unwrap(); + assert!(!current_completion(&fixture.0, ¤t).unwrap()); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), before); + let stopped = generation(13, "stopped-data-retained"); + state::write(&fixture.0.join("state.json"), &stopped).unwrap(); + assert!(retained(&fixture.0, &stopped).is_err()); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), before); + } + #[test] + fn evicted_historical_recovery_requires_valid_retired_original_and_newer_history() { + let fixture = super::super::tests::Fixture::new(); + let prior = completed(generation(1, "ready-observed")); + let current = generation(13, "ready-observed"); + state::write(&fixture.0.join(FILE), &prior).unwrap(); + assert!(require_historical_recovery(&fixture.0, ¤t).is_err()); + restore_history::retain(&fixture.0, &generation(2, "stopped-data-retained")).unwrap(); + assert!(require_historical_recovery(&fixture.0, ¤t).is_err()); + for number in 3..=12 { + restore_history::retain(&fixture.0, &generation(number, "stopped-data-retained")) + .unwrap(); + } + for case in 0..7 { + let mut bad = completed(generation(1, "ready-observed")); + match case { + 0 => bad.complete_sha256 = None, + 1 => bad.complete_sha256 = Some("invalid".into()), + 2 => bad.listeners_retired = false, + 3 => bad.original_sha256 = "0".repeat(64), + 4 => { + bad.original.owner = "0".repeat(32); + bad.original_sha256 = digest(&bad.original).unwrap(); + } + 5 => bad.boot.clear(), + 6 => bad = completed(current.clone()), + _ => unreachable!(), + } + state::write(&fixture.0.join(FILE), &bad).unwrap(); + let before = fs::read(fixture.0.join(FILE)).unwrap(); + assert!(require_historical_recovery(&fixture.0, ¤t).is_err()); + assert_eq!(fs::read(fixture.0.join(FILE)).unwrap(), before); + } + state::write(&fixture.0.join(FILE), &prior).unwrap(); + state::write(&fixture.0.join("live-owner-cleanup.pending"), &json!({})).unwrap(); + assert!(require_historical_recovery(&fixture.0, ¤t).is_err()); + } + #[test] fn repeated_recovery_does_not_depend_on_evicted_restore_history_or_grow_archives() { let fixture = super::super::tests::Fixture::new(); for cycle in 0..20u64 { diff --git a/packages/runtime-core/src/provider/graph/mod.rs b/packages/runtime-core/src/provider/graph/mod.rs index 0236bce3c..2cb499f47 100644 --- a/packages/runtime-core/src/provider/graph/mod.rs +++ b/packages/runtime-core/src/provider/graph/mod.rs @@ -12,15 +12,45 @@ pub use normalized::run_normalized_with_host_dependencies_until; pub use normalized::{ NormalizedInputIdentity, NormalizedRunOptions, compile_normalized_inputs, run_normalized, }; +#[cfg(target_os = "macos")] +mod absent_publication_cleanup; mod admission; mod cache_provenance; mod dependency_hosts; #[cfg(target_os = "macos")] mod dependency_slots; #[cfg(target_os = "macos")] +mod host_pin_recovery; +#[cfg(target_os = "macos")] +mod interrupted_start_cleanup; +#[cfg(target_os = "macos")] +pub(crate) mod publication_gate; +#[cfg(target_os = "macos")] +pub(crate) mod quiescent_dependency_recovery; +#[cfg(target_os = "macos")] +pub use absent_publication_cleanup::{ + inspect as inspect_absent_publication_cleanup, recover as recover_absent_publication_cleanup, +}; +#[cfg(target_os = "macos")] +mod source_device_rebind; +#[cfg(target_os = "macos")] pub use dependency_slots::{ inspect as dependency_reservations, recover_orphan as recover_dependency_reservation, }; +#[cfg(target_os = "macos")] +pub use host_pin_recovery::{inspect as inspect_host_pin_recovery, recover as recover_host_pins}; +#[cfg(target_os = "macos")] +pub use interrupted_start_cleanup::{ + inspect as inspect_interrupted_start_cleanup, recover as recover_interrupted_start_cleanup, +}; +#[cfg(all(test, target_os = "macos"))] +pub(in crate::provider) use source_device_rebind::fixture_https_witness; +#[cfg(target_os = "macos")] +pub(in crate::provider) use source_device_rebind::https_devices; +#[cfg(target_os = "macos")] +pub use source_device_rebind::{ + inspect as inspect_source_device_rebind, recover as recover_source_device_rebind, +}; mod initializer_cache; mod volume_subpaths; pub use dependency_hosts::dependency_address; @@ -36,12 +66,21 @@ mod bridge_recovery; mod bridges; pub use bridge_recovery::{export_bridge_recovery, inspect_bridge_recovery}; pub(in crate::provider) use bridges::{initialize_owner_registry, verify_owner_registry}; +#[cfg(target_os = "macos")] +mod acknowledged_publisher; mod cleanup_enrollment; #[cfg(target_os = "macos")] +pub use acknowledged_publisher::{ + AcknowledgedPublisherSelection, release_dependencies as release_acknowledged_dependencies, + retire as retire_acknowledged_publisher, +}; +#[cfg(target_os = "macos")] mod dead_owner_cleanup; #[cfg(target_os = "macos")] mod live_owner_cleanup; #[cfg(target_os = "macos")] +pub(in crate::provider) use dead_owner_cleanup::HttpsArchiveGuard; +#[cfg(target_os = "macos")] pub use dead_owner_cleanup::recover_cleanup; #[cfg(target_os = "macos")] pub use dead_owner_cleanup::retire_recovered_publisher; @@ -217,6 +256,8 @@ pub struct Receipt { pub struct Snapshot { pub receipt: Receipt, pub journal_incomplete: bool, + #[cfg(target_os = "macos")] + pub interrupted_start_cleanup_incomplete: bool, pub observations: BTreeMap, pub guest_endpoints: BTreeMap, pub storage_references: Value, @@ -1436,6 +1477,14 @@ pub fn source_compatibility( } let engine = Engine::connect_cleanup(candidate)?; let (receipt, root) = load(candidate, &engine, run)?; + #[cfg(target_os = "macos")] + let source_rebind = source_device_rebind::select(&engine, &receipt, &root)?; + #[cfg(target_os = "macos")] + let source_receipt = source_rebind + .as_ref() + .map_or(&receipt, |selected| selected.source_receipt()); + #[cfg(not(target_os = "macos"))] + let source_receipt = &receipt; if !hex(plan_id, 64) || project::identity(plan)? != plan_id || !normalized.matches_plan(plan) @@ -1453,7 +1502,7 @@ pub fn source_compatibility( .services .values() .any(|service| service.active && service.dependency_cache.is_some()); - let revision = receipt + let revision = source_receipt .source .as_ref() .map(|binding| binding.revision.as_str()); @@ -1466,8 +1515,13 @@ pub fn source_compatibility( } if !unchanged_normalized_review(&receipt, plan_id, normalized) { let revision = revision.ok_or_else(refused)?; - if hostname_change::prepare(&engine, plan, &receipt, normalized)?.is_none() { - source::prepare_shared_changed(&engine, plan, &receipt, cached.then_some(revision))?; + if hostname_change::prepare(&engine, plan, source_receipt, normalized)?.is_none() { + source::prepare_shared_changed( + &engine, + plan, + source_receipt, + cached.then_some(revision), + )?; } } Ok(json!({ @@ -1489,6 +1543,8 @@ fn inspect_using( let journal_incomplete = root.join("state.pending").exists() || root.join("state.pending").is_symlink() || startup::require_dependency_rebind_complete(&root, &receipt).is_err(); + #[cfg(target_os = "macos")] + let interrupted_start_cleanup_incomplete = interrupted_start_cleanup::incomplete(&root, run)?; let mut networks = BTreeMap::new(); for (key, resource) in receipt .resources @@ -1543,6 +1599,8 @@ fn inspect_using( Ok(Snapshot { receipt, journal_incomplete, + #[cfg(target_os = "macos")] + interrupted_start_cleanup_incomplete, observations, guest_endpoints, storage_references, @@ -1563,11 +1621,28 @@ pub fn cleanup( // The caller retains the VM mutation lease through this entire effect and any // additional relay confirmation. Never reconnect from inside this function. fn cleanup_owned( + candidate: &Candidate, + engine: &Engine<'_>, + receipt: Receipt, + root: &std::path::Path, + remove_data: bool, +) -> Result { + cleanup_owned_fenced(candidate, engine, receipt, root, remove_data, false, || { + Ok(()) + }) +} + +/// An explicit recovery can recheck its independent host proof immediately +/// before each destructive graph effect. Ordinary callers use the same engine +/// lease and state machine without an additional recovery witness. +fn cleanup_owned_fenced( candidate: &Candidate, engine: &Engine<'_>, mut receipt: Receipt, root: &std::path::Path, remove_data: bool, + early_intent: bool, + fence: impl Fn() -> Result<(), CandidateError>, ) -> Result { if root.join("state.pending").exists() || root.join("state.pending").is_symlink() { return Err(error( @@ -1575,11 +1650,35 @@ fn cleanup_owned( "Pending graph journal is retained; cleanup is blocked until journal reconciliation.", )); } + fence()?; + // Absence recovery must commit the graph cleanup intent before changing a + // selected relay/bridge inventory. A crash can then resume from an exact + // durable intent even when an early external effect already completed. + let environment_slots = if early_intent { + let slots = environment::cleanup_slots(candidate, engine, &receipt)?; + receipt.phase = "cleanup-intent".into(); + state::write(&root.join("state.json"), &receipt)?; + fence()?; + slots + } else { + Vec::new() + }; startup::cleanup_guest(engine, &receipt, false)?; + fence()?; bridges::release_run(candidate, engine, &receipt)?; - let environment_slots = environment::cleanup_slots(candidate, engine, &receipt)?; - receipt.phase = "cleanup-intent".into(); - state::write(&root.join("state.json"), &receipt)?; + #[cfg(test)] + if early_intent { + fault_pause(root, &receipt.run, "absent-after-bridge-release")?; + } + let environment_slots = if early_intent { + environment_slots + } else { + let slots = environment::cleanup_slots(candidate, engine, &receipt)?; + receipt.phase = "cleanup-intent".into(); + state::write(&root.join("state.json"), &receipt)?; + slots + }; + fence()?; shutdown::stop_owned(engine, &receipt, root)?; for kind in [Kind::Container, Kind::Network, Kind::Volume] { if kind == Kind::Volume && !remove_data { @@ -1601,12 +1700,14 @@ fn cleanup_owned( state::write(&root.join("state.json"), &receipt)?; continue; } + fence()?; if let Some(value) = inspect_resource(engine, &receipt, &resource)? { let target = if kind == Kind::Volume { resource.name.as_str() } else { value["Id"].as_str().expect("verified id") }; + fence()?; engine.request( Method::DELETE, &format!( @@ -1635,12 +1736,15 @@ fn cleanup_owned( fault_pause(root, &receipt.run, "cleanup-after-remove")?; } } + fence()?; startup::cleanup_guest(engine, &receipt, true)?; startup::verify_cleanup(engine, &receipt)?; probes::cleanup(engine, &mut receipt)?; for slot in environment_slots { + fence()?; super::environment_recovery::retire(candidate, engine.guest(), &slot, None)?; } + fence()?; receipt.phase = if remove_data { "removed" } else { @@ -1648,6 +1752,7 @@ fn cleanup_owned( } .into(); state::write(&root.join("state.json"), &receipt)?; + fence()?; Ok(receipt) } @@ -1685,6 +1790,7 @@ pub fn restart(candidate: &Candidate, options: RunOptions<'_>) -> Result Duration::from_secs(600) { return Err(error("graph_budget", "Invalid graph timeout.")); } + let project_path = options.project.project; let inputs = project::inputs::compile( candidate, options.project, @@ -1773,6 +1879,9 @@ pub fn restart(candidate: &Candidate, options: RunOptions<'_>) -> Result, receipt: &mut Receipt) -> Result<(), Ok(()) } +/// One selected probe effect for interrupted-start recovery. The caller journals +/// the step before calling and must verify absence before advancing that journal. +#[cfg(target_os = "macos")] +fn inactive_exec(value: Option<&Value>) -> bool { + value.is_none_or(|observed| observed["Running"].as_bool() == Some(false)) +} + +#[cfg(target_os = "macos")] +pub(super) fn retire_one( + engine: &Engine<'_>, + receipt: &Receipt, + name: &str, +) -> Result<(), CandidateError> { + validate(receipt)?; + if !receipt.probes.contains_key(name) + || inspect_resource( + engine, + receipt, + &receipt.resources[&format!("container:{name}")], + )? + .is_some() + || !inactive_exec(exec(engine, receipt, name)?.as_ref()) + { + return Err(error( + "graph_probe_running", + "Selected probe still has an owned consumer.", + )); + } + execute(engine, receipt, name, "remove") +} + +/// An interrupted probe removal advances only from fresh guest absence. +#[cfg(target_os = "macos")] +pub(super) fn verify_one_absent( + engine: &Engine<'_>, + receipt: &Receipt, + name: &str, +) -> Result<(), CandidateError> { + validate(receipt)?; + let probe = receipt.probes.get(name).ok_or_else(|| { + error( + "graph_probe_receipt", + "Selected probe is missing from the graph receipt.", + ) + })?; + if inspect_resource( + engine, + receipt, + &receipt.resources[&format!("container:{name}")], + )? + .is_some() + || !inactive_exec(exec(engine, receipt, name)?.as_ref()) + { + return Err(error( + "graph_probe_cleanup_uncertain", + "Selected probe remains active.", + )); + } + let output = engine + .guest() + .execute_cleanup(PROBE_ABSENT, &[&path(probe)])?; + if output != "probe-storage-absent\n" { + return Err(error( + "graph_probe_cleanup_uncertain", + "Selected probe storage remains.", + )); + } + Ok(()) +} + /// Immutable probe selection for a cleanup effect. Retirement changes only phase; /// every other field remains bound so lost/replaced probe records cannot look empty. #[cfg(target_os = "macos")] @@ -617,6 +687,22 @@ mod storage_tests { use super::STORAGE; use std::{fs, process::Command}; + #[cfg(target_os = "macos")] + #[test] + fn recovery_probe_exec_requires_explicit_nonrunning_observation() { + use serde_json::json; + assert!(super::inactive_exec(None)); + assert!(super::inactive_exec(Some(&json!({"Running":false})))); + for observed in [ + json!({}), + json!({"Running":null}), + json!({"Running":"false"}), + json!({"Running":true}), + ] { + assert!(!super::inactive_exec(Some(&observed))); + } + } + #[cfg(target_os = "macos")] #[test] fn absence_observation_refuses_live_and_dangling_allocations_without_writes() { diff --git a/packages/runtime-core/src/provider/graph/publication_gate.rs b/packages/runtime-core/src/provider/graph/publication_gate.rs new file mode 100644 index 000000000..a43db0dcc --- /dev/null +++ b/packages/runtime-core/src/provider/graph/publication_gate.rs @@ -0,0 +1,50 @@ +//! Serialize publication effects with explicit pool-wide quiescence recovery. +//! Publishers release this gate after binding; their ordinary run lock remains held. +//! Recovery retains it before taking run locks and the provider engine lease. +use crate::{Candidate, CandidateError, provider::state}; +use std::{fs, os::unix::fs::MetadataExt, path::PathBuf}; + +pub(crate) struct Guard { + root: PathBuf, + directory: (u64, u64), + lock: state::Lock, +} + +impl Guard { + pub(crate) fn acquire(candidate: &Candidate) -> Result { + let root = candidate.state_root.join("run/graph-publication-gate"); + let lock = state::Lock::acquire(&root)?; + let metadata = fs::symlink_metadata(&root).map_err(state::io)?; + let guard = Self { + root, + directory: (metadata.dev(), metadata.ino()), + lock, + }; + guard.verify(candidate)?; + Ok(guard) + } + + pub(crate) fn verify(&self, candidate: &Candidate) -> Result<(), CandidateError> { + let root = candidate.state_root.join("run/graph-publication-gate"); + state::check_private_directory(&root)?; + let directory = fs::symlink_metadata(&root).map_err(state::io)?; + let lock = fs::symlink_metadata(root.join("operation.lock")).map_err(state::io)?; + if self.root != root + || !directory.is_dir() + || directory.file_type().is_symlink() + || (directory.dev(), directory.ino()) != self.directory + || !lock.is_file() + || lock.file_type().is_symlink() + || lock.nlink() != 1 + || lock.uid() != unsafe { libc::geteuid() } + || lock.mode() & 0o077 != 0 + || (lock.dev(), lock.ino()) != self.lock.identity()? + { + return Err(CandidateError::new( + "graph_publication_gate", + "Graph publication gate changed; recovery and publication were refused.", + )); + } + Ok(()) + } +} diff --git a/packages/runtime-core/src/provider/graph/quiescent_dependency_recovery.rs b/packages/runtime-core/src/provider/graph/quiescent_dependency_recovery.rs new file mode 100644 index 000000000..75f3cb627 --- /dev/null +++ b/packages/runtime-core/src/provider/graph/quiescent_dependency_recovery.rs @@ -0,0 +1,453 @@ +//! Exact pool quiescence proof for explicit legacy dependency-socket recovery. +//! This grants no graph cleanup or data deletion authority. In particular, a +//! private unrecorded socket inode does not prove who originally created it. +use super::{ + Candidate, CandidateError, Engine, Kind, Method, Receipt, directory, host_pin_recovery, + inspect_resource, load_at, publication_gate, restore, state, +}; +use crate::provider::{lifecycle, state::Owner}; +use serde::Serialize; +use sha2::{Digest, Sha256}; +use std::{ + collections::BTreeMap, + fs, + os::unix::fs::MetadataExt, + path::{Path, PathBuf}, +}; + +const MAX_GRAPHS: usize = 64; +const MAX_GRAPH_FILES: usize = 256; +const MAX_STATE_BYTES: u64 = 2 * 1024 * 1024; + +fn refused() -> CandidateError { + CandidateError::new( + "dependency_socket_recovery", + "Pool quiescence or selected runtime identity changed; dependency sockets and data were retained.", + ) +} + +fn id(path: &Path) -> Result<(u64, u64), CandidateError> { + let metadata = fs::symlink_metadata(path).map_err(|_| refused())?; + Ok((metadata.dev(), metadata.ino())) +} + +fn absent(path: &Path) -> Result { + match fs::symlink_metadata(path) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => Ok(true), + Ok(_) => Ok(false), + Err(_) => Err(refused()), + } +} + +type GraphInventory = (Option<(u64, u64)>, Vec); + +fn graph_runs(candidate: &Candidate) -> Result { + let root = candidate.state_root.join("run/graphs"); + if absent(&root)? { + return Ok((None, Vec::new())); + } + state::check_private_directory(&root).map_err(|_| refused())?; + let identity = id(&root)?; + let entries = fs::read_dir(&root) + .map_err(|_| refused())? + .take(MAX_GRAPHS + 1) + .collect::, _>>() + .map_err(|_| refused())?; + if entries.len() > MAX_GRAPHS { + return Err(refused()); + } + let mut runs = Vec::with_capacity(entries.len()); + for entry in entries { + let run = entry.file_name().into_string().map_err(|_| refused())?; + if !super::hex(&run, 32) || directory(candidate, &run)? != entry.path() { + return Err(refused()); + } + state::check_private_directory(&entry.path()).map_err(|_| refused())?; + runs.push(run); + } + runs.sort(); + if runs.windows(2).any(|pair| pair[0] == pair[1]) { + return Err(refused()); + } + Ok((Some(identity), runs)) +} + +struct Foreground { + root: PathBuf, + directory: (u64, u64), + lock: state::Lock, +} + +impl Foreground { + fn acquire(candidate: &Candidate, run: &str) -> Result { + let root = super::foreground::transport::root(candidate, run)?; + let lock = if absent(&root)? { + state::Lock::acquire(&root)? + } else { + state::check_private_directory(&root).map_err(|_| refused())?; + state::Lock::acquire_existing(&root)? + }; + let result = Self { + directory: id(&root)?, + root, + lock, + }; + result.verify()?; + Ok(result) + } + + fn verify(&self) -> Result<(), CandidateError> { + state::check_private_directory(&self.root).map_err(|_| refused())?; + if id(&self.root)? != self.directory { + return Err(refused()); + } + host_pin_recovery::exact_lock_path(&self.root, &self.lock).map_err(|_| refused())?; + let entries = fs::read_dir(&self.root) + .map_err(|_| refused())? + .take(2) + .collect::, _>>() + .map_err(|_| refused())?; + if entries.len() != 1 || entries[0].file_name().to_str() != Some("operation.lock") { + return Err(refused()); + } + Ok(()) + } +} + +#[derive(Clone, PartialEq, Eq, Serialize)] +struct GraphProof { + run: String, + graph_root: (u64, u64), + state_file: (u64, u64), + state_sha256: String, + foreground_root: (u64, u64), + retained_volumes: BTreeMap, +} + +#[derive(Clone, PartialEq, Eq, Serialize)] +struct Selection { + version: u8, + checkout: PathBuf, + owner_sha256: String, + owner_file: (u64, u64), + owner_root: (u64, u64), + owner_token: String, + guest_boot: String, + host_boot_micros: u64, + home: (u64, u64), + graph_directory: Option<(u64, u64)>, + dependency_assignments: Option<(u64, u64)>, + graphs: Vec, +} + +fn no_pending(root: &Path) -> Result<(), CandidateError> { + let entries = fs::read_dir(root) + .map_err(|_| refused())? + .take(MAX_GRAPH_FILES + 1) + .collect::, _>>() + .map_err(|_| refused())?; + if entries.len() > MAX_GRAPH_FILES + || entries.iter().any(|entry| { + entry + .file_name() + .to_str() + .is_none_or(|name| name.ends_with(".pending")) + }) + { + return Err(refused()); + } + Ok(()) +} + +fn no_dependency_assignments(candidate: &Candidate) -> Result, CandidateError> { + let root = candidate.state_root.join("run/dependency-assignments"); + if absent(&root)? { + return Ok(None); + } + state::check_private_directory(&root).map_err(|_| refused())?; + if fs::read_dir(&root).map_err(|_| refused())?.next().is_some() { + return Err(refused()); + } + Ok(Some(id(&root)?)) +} + +fn no_guest_containers(engine: &Engine<'_>) -> Result<(), CandidateError> { + let inventory = engine.request( + Method::GET, + "/v1.53/containers/json?all=true&limit=129", + None, + )?; + let entries = inventory.as_array().ok_or_else(refused)?; + // A nonempty list includes either a still-live consumer or an unrecorded + // stopped object whose mounts cannot be attributed to this selection. + if !entries.is_empty() { + return Err(refused()); + } + Ok(()) +} + +fn no_compute(engine: &Engine<'_>, receipt: &Receipt) -> Result<(), CandidateError> { + for resource in receipt.resources.values() { + if resource.kind == Kind::Volume { + continue; + } + if resource.phase != "absent" || inspect_resource(engine, receipt, resource)?.is_some() { + return Err(refused()); + } + let suffix = if resource.kind == Kind::Container { + "/json" + } else { + "" + }; + let path = format!( + "/v1.53/{}/{}{}", + resource.kind.collection(), + resource.name, + suffix + ); + match engine.request(Method::GET, &path, None) { + Err(error) if error.code == "engine_not_found" => {} + _ => return Err(refused()), + } + } + Ok(()) +} + +fn stopped_startup(receipt: &Receipt) -> Result<&super::startup::Startup, CandidateError> { + if receipt.phase != "stopped-data-retained" || receipt.relay_cleanup.is_some() { + return Err(refused()); + } + receipt.relay_startup.as_ref().ok_or_else(refused) +} + +fn graph_proof( + candidate: &Candidate, + engine: &Engine<'_>, + run: &str, + foreground: &Foreground, +) -> Result { + foreground.verify()?; + let root = directory(candidate, run)?; + no_pending(&root)?; + let (receipt, _) = load_at(root.clone(), run, engine.guest().incarnation())?; + let startup = stopped_startup(&receipt)?; + if !absent(&startup.control_root)? { + return Err(refused()); + } + let state_path = root.join("state.json"); + let state_bytes = host_pin_recovery::read_raw(&state_path, MAX_STATE_BYTES)?; + let raw_receipt: Receipt = serde_json::from_slice(&state_bytes).map_err(|_| refused())?; + if serde_json::to_value(&raw_receipt).map_err(|_| refused())? + != serde_json::to_value(&receipt).map_err(|_| refused())? + { + return Err(refused()); + } + no_compute(engine, &receipt)?; + let retained_volumes = restore::observed_volumes(engine, &receipt)?; + Ok(GraphProof { + run: run.into(), + graph_root: id(&root)?, + state_file: id(&state_path)?, + state_sha256: format!("{:x}", Sha256::digest(state_bytes)), + foreground_root: foreground.directory, + retained_volumes, + }) +} + +fn observe( + candidate: &Candidate, + engine: &Engine<'_>, + foreground: &[Foreground], +) -> Result<(Owner, Selection), CandidateError> { + lifecycle::host_filesystem::no_auxiliary_update(candidate)?; + let owner = Owner::load(candidate)?; + if owner.phase != "running" + || owner.token != engine.guest().incarnation() + || owner.guest_boot_id.as_deref() != Some(engine.guest().boot_id()) + || owner.dependency_sockets.is_none() + { + return Err(refused()); + } + host_pin_recovery::verify_guest_identity(engine)?; + let owner_root = candidate.state_root.join("run/smolvm"); + let owner_path = owner_root.join("owner.json"); + let owner_bytes = host_pin_recovery::read_raw(&owner_path, 1024 * 1024)?; + if serde_json::from_slice::(&owner_bytes).map_err(|_| refused())? != owner { + return Err(refused()); + } + let home = owner_root.join("home"); + state::check_private_directory(&home).map_err(|_| refused())?; + let home_id = id(&home)?; + let alias = fs::metadata(&owner.short_home).map_err(|_| refused())?; + if (alias.dev(), alias.ino()) != home_id { + return Err(refused()); + } + let (graph_directory, runs) = graph_runs(candidate)?; + if runs.len() != foreground.len() { + return Err(refused()); + } + no_guest_containers(engine)?; + super::bridges::cleanup::require_quiescent(candidate, engine)?; + let mut graphs = Vec::with_capacity(runs.len()); + for (run, guard) in runs.iter().zip(foreground) { + graphs.push(graph_proof(candidate, engine, run, guard)?); + } + let selected = Selection { + version: 1, + checkout: candidate.checkout.clone(), + owner_sha256: format!("{:x}", Sha256::digest(owner_bytes)), + owner_file: id(&owner_path)?, + owner_root: id(&owner_root)?, + owner_token: owner.token.clone(), + guest_boot: engine.guest().boot_id().into(), + host_boot_micros: lifecycle::host_filesystem::host_boot_micros()?, + home: home_id, + graph_directory, + dependency_assignments: no_dependency_assignments(candidate)?, + graphs, + }; + Ok((owner, selected)) +} + +/// Holds the publication gate, every selected foreground lock, and then the +/// provider Engine lease. Rechecking this exact observation is required before +/// each socket unlink and after the last unlink. +pub(crate) struct Guard<'a> { + candidate: &'a Candidate, + gate: publication_gate::Guard, + foreground: Vec, + engine: Engine<'a>, + owner: Owner, + selected: Selection, + sha256: String, +} + +impl<'a> Guard<'a> { + pub(crate) fn acquire(candidate: &'a Candidate) -> Result { + let gate = publication_gate::Guard::acquire(candidate)?; + let (_, runs) = graph_runs(candidate)?; + let mut foreground = Vec::with_capacity(runs.len()); + for run in &runs { + foreground.push(Foreground::acquire(candidate, run)?); + } + gate.verify(candidate)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let (owner, selected) = observe(candidate, &engine, &foreground)?; + let bytes = serde_json::to_vec(&selected).map_err(|_| refused())?; + let sha256 = format!("{:x}", Sha256::digest(bytes)); + let guard = Self { + candidate, + gate, + foreground, + engine, + owner, + selected, + sha256, + }; + guard.verify(candidate)?; + Ok(guard) + } + + pub(crate) fn verify(&self, candidate: &Candidate) -> Result<(), CandidateError> { + if self.candidate.state_root != candidate.state_root + || self.candidate.checkout != candidate.checkout + { + return Err(refused()); + } + self.gate.verify(candidate)?; + for guard in &self.foreground { + guard.verify()?; + } + self.engine.guest().verify()?; + let (owner, selected) = observe(candidate, &self.engine, &self.foreground)?; + if owner != self.owner || selected != self.selected { + return Err(refused()); + } + self.gate.verify(candidate)?; + Ok(()) + } + + pub(crate) fn sha256(&self) -> &str { + &self.sha256 + } + + pub(crate) fn owner(&self) -> &Owner { + &self.owner + } +} + +#[cfg(test)] +mod tests { + use super::*; + use serde_json::json; + + #[test] + fn graph_inventory_accepts_only_absent_or_bounded_private_run_directories() { + let fixture = super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + assert_eq!(graph_runs(&candidate).unwrap(), (None, Vec::new())); + let graphs = candidate.state_root.join("run/graphs"); + state::private_directory(&graphs).unwrap(); + assert_eq!(graph_runs(&candidate).unwrap().1, Vec::::new()); + fs::write(graphs.join("unknown"), b"not a graph").unwrap(); + assert!(graph_runs(&candidate).is_err()); + fs::remove_file(graphs.join("unknown")).unwrap(); + state::private_directory(&graphs.join("a".repeat(32))).unwrap(); + assert_eq!(graph_runs(&candidate).unwrap().1, vec!["a".repeat(32)]); + } + + #[test] + fn pending_and_active_graphs_never_satisfy_stopped_policy() { + let fixture = super::super::tests::Fixture::new(); + let root = fixture.0.join("graph"); + state::private_directory(&root).unwrap(); + assert!(no_pending(&root).is_ok()); + fs::write(root.join("state.pending"), b"interrupted").unwrap(); + assert!(no_pending(&root).is_err()); + fs::remove_file(root.join("state.pending")).unwrap(); + fs::write(root.join("unknown.pending"), b"interrupted").unwrap(); + assert!(no_pending(&root).is_err()); + + let mut receipt: Receipt = serde_json::from_value(json!({ + "version":1,"run":"a".repeat(32),"owner":"b".repeat(32), + "namespace":"c".repeat(64),"plan_id":"d".repeat(64), + "phase":"ready-observed","readiness":{},"resources":{}, + "relay_startup":{"control_only":true,"guest_root":null, + "control_root":"/private/missing","artifact":"e".repeat(64),"services":{}} + })) + .unwrap(); + assert!(stopped_startup(&receipt).is_err()); + receipt.phase = "stopped-data-retained".into(); + assert!(stopped_startup(&receipt).is_ok()); + receipt.relay_startup = None; + assert!(stopped_startup(&receipt).is_err()); + } + + #[test] + fn dependency_assignment_directory_must_be_empty_and_private() { + let fixture = super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + assert_eq!(no_dependency_assignments(&candidate).unwrap(), None); + let assignments = candidate.state_root.join("run/dependency-assignments"); + state::private_directory(&assignments).unwrap(); + assert!(no_dependency_assignments(&candidate).unwrap().is_some()); + fs::write(assignments.join("stale.json"), b"selected").unwrap(); + assert!(no_dependency_assignments(&candidate).is_err()); + } + + #[test] + fn foreground_reservation_refuses_replaced_lock_path() { + let fixture = super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let run = "a".repeat(32); + let guard = Foreground::acquire(&candidate, &run).unwrap(); + guard.verify().unwrap(); + let root = guard.root.clone(); + fs::rename(root.join("operation.lock"), root.join("old.lock")).unwrap(); + let replacement = state::Lock::acquire(&root).unwrap(); + assert!(guard.verify().is_err()); + drop(replacement); + drop(guard); + fs::remove_dir_all(root).unwrap(); + } +} diff --git a/packages/runtime-core/src/provider/graph/relay-fence.sh b/packages/runtime-core/src/provider/graph/relay-fence.sh index 8950f14cb..74d812784 100644 --- a/packages/runtime-core/src/provider/graph/relay-fence.sh +++ b/packages/runtime-core/src/provider/graph/relay-fence.sh @@ -93,7 +93,7 @@ if test "$serial" -ne 0; then test "$serial" -eq "$seen" case "$phase" in cancelled|stopped|discarded) :;; *) exit 1;; esac ;; - inspect) + inspect|inspect-retirement) test "$serial" -eq "$seen" case "$phase" in preparing|discarding|discarded) check_staging; printf 'exited\n'; exit;; diff --git a/packages/runtime-core/src/provider/graph/relay-retired-absence.sh b/packages/runtime-core/src/provider/graph/relay-retired-absence.sh new file mode 100644 index 000000000..01dc5f21d --- /dev/null +++ b/packages/runtime-core/src/provider/graph/relay-retired-absence.sh @@ -0,0 +1,50 @@ +set -efu +allocation=$1; slot=$2; serial=$3 +case "$allocation" in *[!0-9a-f]*|'') exit 1;; esac +test "${#allocation}" = 32 +case "$slot" in *[!0-9]*|'') exit 1;; esac +test "$slot" -ge 0 +test "$slot" -le 31 +case "$serial" in *[!0-9]*|'') exit 1;; esac +test "$serial" -gt 0 +test "$(findmnt -n -o FSTYPE --target /run/hack-local)" = tmpfs +base=/run/hack-local +private_dir() { + test ! -L "$1" + test -d "$1" + test "$(stat -c %u:%g:%a "$1")" = 0:0:700 +} +private_file() { + test ! -L "$1" + test -f "$1" + test "$(stat -c %u:%g:%a:%h "$1")" = 0:0:600:1 + test "$(stat -c %s "$1")" -le 256 +} +private_dir "$base" +private_dir "$base/graph-relays" +root="$base/graph-relays/$allocation" +socket=$(printf '%s/bridge-%02d.sock' "$base" "$slot") +test ! -L "$root" +if test -e "$root"; then + test -d "$root" + printf 'present\n' + exit +fi +test ! -e "$socket" +test ! -L "$socket" +private_dir "$base/relay-slots" +control="$base/relay-slots/slot-$slot" +private_dir "$control" +private_file "$control/lock" +exec 9<> "$control/lock" +flock -w 7 9 +test ! -e "$control/pending" +test ! -L "$control/pending" +private_file "$control/state" +set -- $(cat "$control/state") +test "$#" = 3 +test "$1" = "$serial" +test "$2" = "$allocation" +test "$3" = stopped +printf '%s %s %s\n' "$1" "$2" "$3" | cmp -s - "$control/state" +printf 'absent\n' diff --git a/packages/runtime-core/src/provider/graph/relay.rs b/packages/runtime-core/src/provider/graph/relay.rs index e330bc6f6..d9a3fb69e 100644 --- a/packages/runtime-core/src/provider/graph/relay.rs +++ b/packages/runtime-core/src/provider/graph/relay.rs @@ -117,7 +117,7 @@ pub(super) fn operate( }; let confirmed = match action { "start" => result == "running\n", - "inspect" => ["running\n", "exited\n"].contains(&result.as_str()), + "inspect" | "inspect-retirement" => ["running\n", "exited\n"].contains(&result.as_str()), "stop" => result == "stopped\n", "remove" => result == "removed\n", _ => false, @@ -131,6 +131,40 @@ pub(super) fn operate( Ok(result.trim().to_owned()) } +/// Read-only proof for a partially removed reservation. An absent allocation +/// is accepted only with the exact slot's stopped fence in this guest boot. +#[cfg(target_os = "macos")] +pub(super) fn retirement_absence( + engine: &Engine<'_>, + slot: u8, + assignment: &bridges::Assignment, +) -> Result<&'static str, CandidateError> { + let relay = assignment.relay.as_ref().ok_or_else(|| { + error( + "graph_bridge_observation", + "Missing relay intent during cleanup.", + ) + })?; + let slot_arg = slot.to_string(); + let serial_arg = relay.launch_serial.to_string(); + let args = [ + assignment.reservation.as_str(), + slot_arg.as_str(), + serial_arg.as_str(), + ]; + let result = engine + .guest() + .execute_cleanup(include_str!("relay-retired-absence.sh"), &args)?; + match result.as_str() { + "absent\n" => Ok("absent"), + "present\n" => Ok("present"), + _ => Err(error( + "graph_bridge_observation", + "Unconfirmed retired relay allocation; cleanup made no change.", + )), + } +} + /// Actual relay-child and listener identity captured before their guest receipts /// are removed. The assignment separately binds boot, reservation and launch serial. #[cfg(target_os = "macos")] @@ -322,6 +356,123 @@ pub(super) fn verify_cleanup( #[cfg(test)] mod tests { use super::*; + + #[cfg(target_os = "macos")] + #[test] + fn retirement_inspection_refuses_reused_pid_before_live_classification() { + use std::process::Command; + let fixture = super::super::tests::Fixture::new(); + let process = fixture.0.join("recorded-process-stat"); + std::fs::write(&process, "present").unwrap(); + let source = include_str!("relay.sh"); + let begin = source + .find("if test \"$action\" = inspect-retirement && test -e") + .unwrap(); + let end = begin + source[begin..].find("if alive; then").unwrap(); + let guard = source[begin..end].replace("/proc/$pid/stat", "$1/recorded-process-stat"); + let program = format!( + "set -efu\naction=$2; current=$3; pid=42; born=100\nstart_ticks() {{ printf '%s\\n' \"$current\"; }}\n{guard}printf 'accepted\\n'\n" + ); + let run = |action: &str, current: &str| { + Command::new("/bin/sh") + .args(["-c", &program, "retirement-pid-test"]) + .arg(&fixture.0) + .args([action, current]) + .output() + .unwrap() + }; + assert!(!run("inspect-retirement", "101").status.success()); + assert_eq!(run("inspect-retirement", "100").stdout, b"accepted\n"); + assert_eq!(run("inspect", "101").stdout, b"accepted\n"); + } + + #[cfg(target_os = "macos")] + #[test] + fn retired_absence_requires_exact_stopped_fence_and_no_replaced_paths() { + use std::os::unix::fs::{MetadataExt, PermissionsExt, symlink}; + use std::process::Command; + + let fixture = super::super::tests::Fixture::new(); + let base = fixture.0.join("guest-relay-absence"); + let controls = base.join("relay-slots/slot-0"); + std::fs::create_dir_all(base.join("graph-relays")).unwrap(); + std::fs::create_dir_all(&controls).unwrap(); + for dir in [ + &base, + &base.join("graph-relays"), + &base.join("relay-slots"), + &controls, + ] { + std::fs::set_permissions(dir, std::fs::Permissions::from_mode(0o700)).unwrap(); + } + let state = controls.join("state"); + let pending = controls.join("pending"); + let lock = controls.join("lock"); + let allocation = "a".repeat(32); + std::fs::write(&lock, "").unwrap(); + std::fs::set_permissions(&lock, std::fs::Permissions::from_mode(0o600)).unwrap(); + let owner = base.metadata().unwrap(); + let script = include_str!("relay-retired-absence.sh") + .replace( + "test \"$(findmnt -n -o FSTYPE --target /run/hack-local)\" = tmpfs", + ":", + ) + .replace("/run/hack-local", base.to_str().unwrap()) + // Translate GNU guest stat formats to macOS stat while preserving + // the real permission, link-count and size checks in this fixture. + .replace("stat -c %u:%g:%a:%h", "/usr/bin/stat -f %u:%g:%Lp:%l") + .replace("stat -c %u:%g:%a", "/usr/bin/stat -f %u:%g:%Lp") + .replace("stat -c %s", "/usr/bin/stat -f %z") + .replace("0:0", &format!("{}:{}", owner.uid(), owner.gid())) + .replace("flock -w 7 9", ":"); + let run = |allocation: &str, serial: &str| { + Command::new("/bin/sh") + .args([ + "-c", + &script, + "retired-absence-test", + allocation, + "0", + serial, + ]) + .output() + .unwrap() + }; + let write_state = |value: String| { + std::fs::write(&state, value).unwrap(); + std::fs::set_permissions(&state, std::fs::Permissions::from_mode(0o600)).unwrap(); + }; + write_state(format!("7 {allocation} stopped\n")); + let valid = run(&allocation, "7"); + assert!(valid.status.success(), "{:?}", valid.stderr); + assert_eq!(valid.stdout, b"absent\n"); + assert!(!run(&"b".repeat(32), "7").status.success()); + assert!(!run(&allocation, "8").status.success()); + for phase in [ + "launching", + "closing", + "preparing", + "discarded", + "cancelled", + ] { + write_state(format!("7 {allocation} {phase}\n")); + assert!(!run(&allocation, "7").status.success(), "{phase}"); + } + write_state(format!("7 {allocation} stopped\n")); + std::fs::write(&pending, format!("7 {allocation} stopped\n")).unwrap(); + assert!(!run(&allocation, "7").status.success()); + std::fs::remove_file(&pending).unwrap(); + let socket = base.join("bridge-00.sock"); + symlink("replaced", &socket).unwrap(); + assert!(!run(&allocation, "7").status.success()); + std::fs::remove_file(&socket).unwrap(); + let allocation_root = base.join("graph-relays").join(&allocation); + symlink("replaced", &allocation_root).unwrap(); + assert!(!run(&allocation, "7").status.success()); + std::fs::remove_file(&allocation_root).unwrap(); + std::fs::create_dir(&allocation_root).unwrap(); + assert_eq!(run(&allocation, "7").stdout, b"present\n"); + } #[test] fn observer_descriptor_loop_expands_only_its_controlled_glob() { use std::os::unix::fs::symlink; diff --git a/packages/runtime-core/src/provider/graph/relay.sh b/packages/runtime-core/src/provider/graph/relay.sh index 070cb66e3..68e3db593 100644 --- a/packages/runtime-core/src/provider/graph/relay.sh +++ b/packages/runtime-core/src/provider/graph/relay.sh @@ -115,13 +115,17 @@ if test "$action" = remove; then fi check_binary read_process +if test "$action" = inspect-retirement && test -e "/proc/$pid/stat"; then + # A reused PID is not proof that the recorded helper exited safely. + test "$(start_ticks "$pid")" = "$born" +fi if alive; then test "$(stat -Lc %d:%i /proc/$pid/exe)" = "$(stat -c %d:%i "$root/relay")" - if test "$action" = inspect; then check_socket; printf 'running\n'; exit; fi + if test "$action" = inspect || test "$action" = inspect-retirement; then check_socket; printf 'running\n'; exit; fi test "$action" = stop "$root/relay" --stop "$pid" "$born" ! alive -elif test "$action" = inspect; then +elif test "$action" = inspect || test "$action" = inspect-retirement; then printf 'exited\n'; exit fi test "$action" = stop diff --git a/packages/runtime-core/src/provider/graph/restore.rs b/packages/runtime-core/src/provider/graph/restore.rs index 4b8c7532a..f60b9dfab 100644 --- a/packages/runtime-core/src/provider/graph/restore.rs +++ b/packages/runtime-core/src/provider/graph/restore.rs @@ -71,6 +71,8 @@ pub(super) struct FreshOwnerRestore<'a> { pub identity: NormalizedInputIdentity, pub generation: &'a str, pub deadline: std::time::Instant, + #[cfg(target_os = "macos")] + pub verify_publication: &'a dyn Fn() -> Result<(), CandidateError>, } #[cfg(target_os = "macos")] pub(super) fn restore_normalized_foreground( @@ -80,6 +82,7 @@ pub(super) fn restore_normalized_foreground( deadline: std::time::Instant, runtime: &mut HostRelayRuntime, generation: &str, + verify_publication: &dyn Fn() -> Result<(), CandidateError>, ) -> Result { startup::Driver::check_cancelled(runtime)?; check_environment_deadline(deadline)?; @@ -109,6 +112,7 @@ pub(super) fn restore_normalized_foreground( identity, generation, deadline, + verify_publication, }), ) } @@ -149,11 +153,10 @@ fn verify_fresh_change( /// Binds the stopped selection to current boot and exact observed retained data. /// Existing receipts identify managed volumes by name/labels; this additionally /// fences replacement between explicit selection and restoration. -pub(super) fn restore_generation( +pub(super) fn observed_volumes( engine: &Engine<'_>, receipt: &Receipt, -) -> Result { - use sha2::{Digest, Sha256}; +) -> Result, CandidateError> { let mut volumes = BTreeMap::new(); for (key, resource) in &receipt.resources { if resource.kind != Kind::Volume { @@ -172,12 +175,44 @@ pub(super) fn restore_generation( "Retained volume changed during selection.", )); } - volumes.insert(key, (resource.name.as_str(), created.to_owned(), directory)); + volumes.insert( + key.clone(), + (resource.name.clone(), created.to_owned(), directory), + ); } - let bytes = serde_json::to_vec(&(receipt, engine.guest().boot_id(), volumes)) + Ok(volumes) +} + +pub(super) fn restore_generation( + engine: &Engine<'_>, + receipt: &Receipt, +) -> Result { + use sha2::{Digest, Sha256}; + let volumes = observed_volumes(engine, receipt)?; + let bytes = serde_json::to_vec(&(receipt, engine.guest().boot_id(), &volumes)) .map_err(|_| error("graph_receipt", "Restore selection unavailable."))?; Ok(format!("{:x}", Sha256::digest(bytes))) } + +#[cfg(target_os = "macos")] +pub(super) fn restore_generation_with_source_rebind( + engine: &Engine<'_>, + receipt: &Receipt, + selected: Option<&super::source_device_rebind::Selected>, +) -> Result { + use sha2::{Digest, Sha256}; + let original = restore_generation(engine, receipt)?; + let Some(selected) = selected else { + return Ok(original); + }; + let bytes = serde_json::to_vec(&( + "hack-graph-restore-source-device-rebind-v1", + original, + selected.raw_sha256(), + )) + .map_err(|_| error("graph_receipt", "Restore selection unavailable."))?; + Ok(format!("{:x}", Sha256::digest(bytes))) +} fn restore_inputs( candidate: &Candidate, options: RunOptions<'_>, @@ -198,9 +233,28 @@ fn restore_inputs( )); } let (mut receipt, root) = load(candidate, &engine, options.run_id)?; + #[cfg(target_os = "macos")] + let source_rebind = super::source_device_rebind::select(&engine, &receipt, &root)?; + #[cfg(target_os = "macos")] + if source_rebind.is_some() && fresh.is_none() { + return Err(error( + "graph_source_device_rebind", + "Source-device continuity requires a new foreground restore owner.", + )); + } let hostname_change = if let Some(fresh) = &fresh { - let change = - hostname_change::prepare(&engine, &inputs.review.plan, &receipt, &fresh.identity)?; + #[cfg(target_os = "macos")] + let source_receipt = source_rebind + .as_ref() + .map_or(&receipt, |selected| selected.source_receipt()); + #[cfg(not(target_os = "macos"))] + let source_receipt = &receipt; + let change = hostname_change::prepare( + &engine, + &inputs.review.plan, + source_receipt, + &fresh.identity, + )?; if change.is_some() && (!options.shared_source || options.live_source @@ -216,6 +270,10 @@ fn restore_inputs( None }; if let Some(fresh) = &fresh { + #[cfg(target_os = "macos")] + let generation = + restore_generation_with_source_rebind(&engine, &receipt, source_rebind.as_ref())?; + #[cfg(not(target_os = "macos"))] let generation = restore_generation(&engine, &receipt)?; if let Some(change) = &hostname_change { verify_fresh_change( @@ -256,11 +314,17 @@ fn restore_inputs( )); } super::super::source_job::check_reservations(&engine)?; + #[cfg(target_os = "macos")] + let source_receipt = source_rebind + .as_ref() + .map_or(&receipt, |selected| selected.source_receipt()); + #[cfg(not(target_os = "macos"))] + let source_receipt = &receipt; let ordinary_source = if hostname_change.is_none() { source::prepare_replay( &engine, &inputs, - &receipt, + source_receipt, options.source_revision, options.live_source, options.shared_source, @@ -361,7 +425,12 @@ fn restore_inputs( } if let Some(fresh) = &fresh { check_environment_deadline(fresh.deadline)?; - if restore_generation(&engine, &receipt)? != fresh.generation { + #[cfg(target_os = "macos")] + let generation = + restore_generation_with_source_rebind(&engine, &receipt, source_rebind.as_ref())?; + #[cfg(not(target_os = "macos"))] + let generation = restore_generation(&engine, &receipt)?; + if generation != fresh.generation { return Err(error( "graph_restore_refused", "Retained restore selection changed before effects.", @@ -371,10 +440,30 @@ fn restore_inputs( if let Some(driver) = startup.as_ref() { driver.check_cancelled()?; } + source::verify_cache_scope(source, options.project.project)?; + #[cfg(target_os = "macos")] + super::source_device_rebind::verify_cache_scope_origin(candidate, &receipt)?; + #[cfg(target_os = "macos")] + if let Some(fresh) = &fresh { + super::absent_publication_cleanup::archive_retired_rebind_under( + candidate, + &engine, + &receipt, + fresh.verify_publication, + )?; + } // Retain the complete acknowledged old generation before replacing its // boot-bound dependency owner. Historical cleanup is verified in its own context. if fresh.is_some() { + #[cfg(target_os = "macos")] + if let Some(selected) = &source_rebind { + selected.reverify(&engine)?; + } super::restore_history::retain(&root, &receipt)?; + #[cfg(target_os = "macos")] + if let Some(selected) = &source_rebind { + selected.apply_to_new_attempt(&mut receipt)?; + } receipt.relay_startup = None; receipt.relay_cleanup = None; } else { @@ -435,6 +524,10 @@ fn restore_inputs( session.fault_pause("restore-intent")?; } let result = (|| { + #[cfg(target_os = "macos")] + if let Some(selected) = &source_rebind { + selected.reverify(&session.engine)?; + } if fresh.is_some() && let Some(driver) = session.startup.as_mut() { @@ -492,6 +585,8 @@ mod tests { identity: next.normalized_input.unwrap(), generation: &generation, deadline: Instant::now() + Duration::from_secs(30), + #[cfg(target_os = "macos")] + verify_publication: &|| Ok(()), }; let reviewed = "8".repeat(64); assert!(verify_fresh(&receipt, &reviewed, &fresh, &generation).is_err()); @@ -534,6 +629,8 @@ mod tests { identity: identity.clone(), generation: &generation, deadline: Instant::now() + Duration::from_secs(30), + #[cfg(target_os = "macos")] + verify_publication: &|| Ok(()), }; verify_fresh(&receipt, &receipt.plan_id, &fresh, &generation).unwrap(); assert!(verify_fresh(&receipt, &"2".repeat(64), &fresh, &generation).is_err()); @@ -568,6 +665,8 @@ mod tests { identity: identity.clone(), generation: &generation, deadline: Instant::now() + Duration::from_secs(30), + #[cfg(target_os = "macos")] + verify_publication: &|| Ok(()), }; let changed = "2".repeat(64); assert!(verify_fresh(&receipt, &changed, &fresh, &generation).is_err()); diff --git a/packages/runtime-core/src/provider/graph/restore_history.rs b/packages/runtime-core/src/provider/graph/restore_history.rs index c87530454..8826cc754 100644 --- a/packages/runtime-core/src/provider/graph/restore_history.rs +++ b/packages/runtime-core/src/provider/graph/restore_history.rs @@ -197,6 +197,34 @@ pub(super) fn completed_for_recovery( .transpose() } +/// Truncated diagnostic history may outlive a superseded recovery's exact +/// completion. Its latest stopped generation must still prove replacement of +/// that original container; this never supplies current cleanup authority. +#[cfg(target_os = "macos")] +pub(super) fn confirms_truncated_newer_generation( + root: &Path, + current: &Receipt, + original: &Receipt, +) -> Result { + let Some(history) = verified_for_recovery(root, current)? else { + return Ok(false); + }; + Ok(history.truncated + && history.entries.last().is_some_and(|stopped| { + original.resources.iter().any(|(key, resource)| { + resource.kind == Kind::Container + && resource.id.as_deref().is_some_and(|old| { + stopped + .resources + .get(key) + .filter(|new| new.kind == Kind::Container) + .and_then(|new| new.id.as_deref()) + .is_some_and(|new| new != old) + }) + }) + })) +} + /// A superseded bridge sidecar must bind to the most recent fully stopped /// generation, not merely to some older container in the bounded history. #[cfg(target_os = "macos")] diff --git a/packages/runtime-core/src/provider/graph/shutdown.rs b/packages/runtime-core/src/provider/graph/shutdown.rs index c385db8c8..0d9cff30e 100644 --- a/packages/runtime-core/src/provider/graph/shutdown.rs +++ b/packages/runtime-core/src/provider/graph/shutdown.rs @@ -222,7 +222,9 @@ pub(super) fn prepare(resource: &Resource, inspected: &Value) -> Result {} (false, Some("exited"), 0) => {} - (false, Some("created"), 0) if state["ExitCode"] == 0 && state["OOMKilled"] == false => {} + // Docker may assign a nonzero exit code when startup fails before the + // container ever runs. Preserve that failure without requesting a stop. + (false, Some("created"), 0) if state["OOMKilled"] == false => {} _ => return Err(refused()), } Ok(Prepared { @@ -399,13 +401,83 @@ mod tests { value["State"]["Running"] = json!(false); value["State"]["Status"] = json!("created"); value["State"]["Pid"] = json!(0); - assert!(!prepare(&resource, &value).unwrap().running); - assert_eq!(terminal(&resource, &value, false).unwrap().exit_code, 0); - assert!(terminal(&resource, &value, true).is_err()); + for code in [0, 126, 127, 128, 255] { + value["State"]["ExitCode"] = json!(code); + assert!(!prepare(&resource, &value).unwrap().running); + let evidence = terminal(&resource, &value, false).unwrap(); + assert_eq!(evidence.exit_code, code); + assert!(!evidence.stop_requested && !evidence.oom_killed); + assert!(terminal(&resource, &value, true).is_err()); + } value["State"]["Pid"] = json!(42); assert!(terminal(&resource, &value, false).is_err()); } + #[test] + fn created_start_failure_does_not_strand_running_siblings_or_lose_failure_evidence() { + let (resource, mut value) = fixture(); + value["State"]["Running"] = json!(false); + value["State"]["Status"] = json!("created"); + value["State"]["Pid"] = json!(0); + value["State"]["ExitCode"] = json!(128); + value["State"]["Error"] = json!("synthetic private startup details"); + let failed = prepare(&resource, &value).unwrap(); + let mut selected = vec![failed.clone()]; + for index in 1..=9 { + let (mut sibling, mut observed) = fixture(); + sibling.id = Some(format!("{index:064x}")); + observed["Id"] = json!(sibling.id); + selected.push(prepare(&sibling, &observed).unwrap()); + } + let stop_ids = selected + .iter() + .filter(|prepared| prepared.running) + .map(|prepared| &prepared.id) + .collect::>(); + assert_eq!(stop_ids.len(), 9); + assert!(!stop_ids.contains(&&failed.id)); + let evidence = terminal(&resource, &value, failed.running).unwrap(); + assert_eq!(evidence.exit_code, 128); + assert!(!evidence.stop_requested && !evidence.oom_killed); + let encoded = serde_json::to_string(&evidence).unwrap(); + assert!(!encoded.contains("synthetic private startup details")); + assert_eq!( + serde_json::from_str::(&encoded).unwrap(), + evidence + ); + } + + #[test] + fn created_failure_still_refuses_uncertain_or_contradictory_state() { + let (resource, mut value) = fixture(); + value["State"]["Running"] = json!(false); + value["State"]["Status"] = json!("created"); + value["State"]["Pid"] = json!(0); + value["State"]["ExitCode"] = json!(128); + for (field, invalid) in [ + ("Running", json!(true)), + ("Running", json!(null)), + ("Pid", json!(1)), + ("Pid", json!(-1)), + ("Paused", json!(true)), + ("Restarting", json!(true)), + ("Dead", json!(true)), + ("OOMKilled", json!(true)), + ("OOMKilled", json!(null)), + ("ExitCode", json!(-1)), + ("ExitCode", json!(256)), + ("ExitCode", json!(1.5)), + ("ExitCode", json!("128")), + ("ExitCode", json!(null)), + ] { + let mut changed = value.clone(); + changed["State"][field] = invalid; + assert!(prepare(&resource, &changed).is_err(), "{field}"); + assert!(terminal(&resource, &changed, false).is_err(), "{field}"); + } + assert!(terminal(&resource, &value, true).is_err()); + } + #[test] fn accepts_docker_defaults_and_known_signals_rejects_malformed_policy() { for signal in [ diff --git a/packages/runtime-core/src/provider/graph/source.rs b/packages/runtime-core/src/provider/graph/source.rs index cd8260efb..8c09c49b5 100644 --- a/packages/runtime-core/src/provider/graph/source.rs +++ b/packages/runtime-core/src/provider/graph/source.rs @@ -3,10 +3,13 @@ use super::*; use crate::project::{PlanData, snapshot::ContentRevision}; use sha2::{Digest, Sha256}; +use std::path::Path; #[derive(Clone, Debug, Serialize, Deserialize, PartialEq, Eq)] #[serde(deny_unknown_fields)] pub struct SourceBinding { + #[serde(default, skip_serializing_if = "Option::is_none")] + pub cache_scope: Option, #[serde(default, skip_serializing_if = "Option::is_none")] pub shared: Option, #[serde(default, skip_serializing_if = "Option::is_none")] @@ -30,6 +33,10 @@ impl SourceBinding { .all(|v| hex(v, 64)) && !(self.live.is_some() && self.shared.is_some()) && (self.shared_contract.is_none() || self.shared.is_some()) + && self + .cache_scope + .as_ref() + .is_none_or(|scope| self.shared.is_some() && scope.valid()) && self.live.as_ref().is_none_or(|live| live.workspace.valid()) && self .shared @@ -75,10 +82,18 @@ fn shared_paths( )?; selected.insert(mount.source.clone(), path); } - engine.guest().execute("set -eu; test \"$(findmnt -n -o FSTYPE --mountpoint \"$1\")\" = virtiofs; case \",$(findmnt -n -o OPTIONS --mountpoint \"$1\"),\" in *,rw,*) ;; *) exit 1;; esac", &[&share.guest_path], None)?; + verify_shared_mount(engine, share)?; Ok(selected) } +pub(super) fn verify_shared_mount( + engine: &Engine<'_>, + share: &super::super::ProjectShareIntent, +) -> Result<(), CandidateError> { + engine.guest().execute("set -eu; test \"$(findmnt -n -o FSTYPE --mountpoint \"$1\")\" = virtiofs; case \",$(findmnt -n -o OPTIONS --mountpoint \"$1\"),\" in *,rw,*) ;; *) exit 1;; esac", &[&share.guest_path], None)?; + Ok(()) +} + pub(super) fn requested(plan: &PlanData, revision: Option<&str>) -> Result<(), CandidateError> { if (mounts(plan).next().is_some() || plan @@ -220,6 +235,7 @@ pub(super) fn prepare_mode( current_manifest: None, binding: SourceBinding { shared: Some(share.clone()), + cache_scope: None, shared_contract: Some(project::live_source::Contract::from_plan(plan, &manifest)?), live: None, revision: manifest.revision.clone(), @@ -300,6 +316,19 @@ pub(super) fn prepare_replay( shared: bool, non_secret_values: &BTreeMap, ) -> Result, CandidateError> { + #[cfg(target_os = "macos")] + super::source_device_rebind::verify_cache_scope_origin(engine.guest().candidate(), receipt)?; + #[cfg(not(target_os = "macos"))] + if receipt + .source + .as_ref() + .is_some_and(|binding| binding.cache_scope.is_some()) + { + return Err(error( + "graph_source_device_rebind", + "Cache device replay requires its native witness.", + )); + } if shared || receipt.source.as_ref().is_some_and(|s| s.shared.is_some()) { if !shared || live { return Err(error( @@ -315,7 +344,7 @@ pub(super) fn prepare_replay( revision, )?)); } - let source = prepare_mode( + let mut source = prepare_mode( engine.guest().candidate(), engine, &inputs.review.plan, @@ -323,6 +352,7 @@ pub(super) fn prepare_replay( false, true, )?; + inherit_cache_scope(&mut source, receipt, &inputs.review.plan.source)?; unchanged(&source, receipt)?; return Ok(source); } @@ -577,6 +607,7 @@ fn published_inputs( manifest: publication.manifest.clone(), binding: SourceBinding { shared: None, + cache_scope: None, shared_contract: None, live: None, revision: revision.into(), @@ -598,6 +629,36 @@ pub(super) fn unchanged(source: &Option, receipt: &Receipt) -> Result<() Ok(()) } +fn inherit_cache_scope( + source: &mut Option, + receipt: &Receipt, + project: &Path, +) -> Result<(), CandidateError> { + if let Some(scope) = receipt + .source + .as_ref() + .and_then(|binding| binding.cache_scope.as_ref()) + { + scope.scope(project)?; + source + .as_mut() + .ok_or_else(|| error("graph_source_changed", "Missing replay source."))? + .binding + .cache_scope = Some(scope.clone()); + } + Ok(()) +} + +pub(super) fn verify_cache_scope( + source: Option<&Inputs>, + project: &Path, +) -> Result<(), CandidateError> { + if let Some(scope) = source.and_then(|source| source.binding.cache_scope.as_ref()) { + scope.scope(project)?; + } + Ok(()) +} + #[cfg(test)] mod tests { use super::*; @@ -631,6 +692,7 @@ mod tests { }; let binding = SourceBinding { shared: None, + cache_scope: None, shared_contract: None, revision: baseline.receipt().revision.clone(), archive_sha256: publication.archive_sha256.clone(), @@ -770,6 +832,7 @@ mod tests { probes: BTreeMap::new(), source: Some(SourceBinding { shared: None, + cache_scope: None, shared_contract: None, revision: manifest.revision, archive_sha256: "d".repeat(64), @@ -888,6 +951,7 @@ mod tests { manifest: manifest.clone(), binding: SourceBinding { shared: None, + cache_scope: None, shared_contract: None, live: None, revision: manifest.revision.clone(), @@ -965,6 +1029,7 @@ mod tests { assert_eq!(serde_json::to_value(&receipt).unwrap(), value); let binding = SourceBinding { shared: None, + cache_scope: None, shared_contract: None, live: None, revision: "a".repeat(64), @@ -994,6 +1059,59 @@ mod tests { assert_eq!(decoded.source, receipt.source); } + #[test] + fn ordinary_replay_propagates_verified_scope_without_adopting_changed_source() { + use std::os::unix::fs::MetadataExt; + let fixture = super::super::tests::Fixture::new(); + fs::write(fixture.0.join("compose.yaml"), "services: {}\n").unwrap(); + let share = super::super::super::ProjectShareIntent::approve(&fixture.0, true).unwrap(); + let metadata = fs::metadata(&fixture.0).unwrap(); + let scope: dependency_cache::ReplayScope = serde_json::from_value(json!({ + "common":fixture.0,"device":metadata.dev(),"inode":metadata.ino(),"original_device":metadata.dev()+7,"origin_sha256":"e".repeat(64), + })).unwrap(); + let mut binding = SourceBinding { + cache_scope: Some(scope.clone()), + shared: Some(share), + shared_contract: None, + live: None, + revision: "a".repeat(64), + archive_sha256: "b".repeat(64), + selection_sha256: "c".repeat(64), + }; + assert!(binding.valid()); + let receipt:Receipt = serde_json::from_value(json!({"version":1,"run":"a".repeat(32),"owner":"b".repeat(32),"namespace":"c".repeat(64),"plan_id":"d".repeat(64),"phase":"stopped-data-retained","readiness":{},"resources":{},"source":binding})).unwrap(); + binding.cache_scope = None; + let mut source = Some(Inputs { + binding, + current_manifest: None, + manifest: ContentRevision { + schema_version: 1, + revision: "a".repeat(64), + selection_sha256: "c".repeat(64), + total_bytes: 0, + entries: vec![], + }, + paths: BTreeMap::new(), + }); + assert!(unchanged(&source, &receipt).is_err()); + inherit_cache_scope(&mut source, &receipt, &fixture.0).unwrap(); + unchanged(&source, &receipt).unwrap(); + verify_cache_scope(source.as_ref(), &fixture.0).unwrap(); + source + .as_mut() + .unwrap() + .binding + .shared + .as_mut() + .unwrap() + .inode += 1; + assert!(unchanged(&source, &receipt).is_err()); + assert!(inherit_cache_scope(&mut None, &receipt, &fixture.0).is_err()); + let foreign = super::super::tests::Fixture::new(); + assert!(verify_cache_scope(source.as_ref(), &foreign.0).is_err()); + assert!(inherit_cache_scope(&mut source, &receipt, &foreign.0).is_err()); + } + #[test] fn shared_source_contract_round_trips_without_rewriting_legacy_bindings() { let fixture = super::super::tests::Fixture::new(); @@ -1025,6 +1143,7 @@ mod tests { shared: Some( super::super::super::ProjectShareIntent::approve(&fixture.0, true).unwrap(), ), + cache_scope: None, shared_contract: Some(contract.clone()), live: None, revision: snapshot.receipt().revision.clone(), diff --git a/packages/runtime-core/src/provider/graph/source_device_rebind.rs b/packages/runtime-core/src/provider/graph/source_device_rebind.rs new file mode 100644 index 000000000..d59ded9af --- /dev/null +++ b/packages/runtime-core/src/provider/graph/source_device_rebind.rs @@ -0,0 +1,1060 @@ +//! One explicit source-device translation for an acknowledged stopped graph. +//! +//! The stopped receipt and its cleanup proofs remain immutable. This witness may +//! project the share device and a proven retained cache scope into the next generation. +use super::{ + Candidate, CandidateError, Engine, Kind, Receipt, cleanup_enrollment, directory, foreground, + host_pin_recovery, inspect_resource, load, restore, source, state, +}; +use crate::provider::{ProjectShareIntent, lifecycle}; +use serde::{Deserialize, Serialize}; +use serde_json::{Value, json}; +use sha2::{Digest, Sha256}; +use std::{ + collections::BTreeMap, + fs::{self, File, OpenOptions}, + io::{Read, Seek, SeekFrom}, + os::unix::fs::{MetadataExt, OpenOptionsExt}, + path::{Path, PathBuf}, +}; + +const FILE: &str = "source-device-rebind.json"; +const LIMIT: u64 = 256 * 1024; +const STATE_LIMIT: u64 = 2 * 1024 * 1024; + +fn refused() -> CandidateError { + CandidateError::new( + "graph_source_device_rebind", + "Selected source-device continuity or retained graph identity changed; original receipts and data were preserved.", + ) +} + +fn sha256(bytes: &[u8]) -> String { + format!("{:x}", Sha256::digest(bytes)) +} + +fn absent(path: &Path) -> Result { + match fs::symlink_metadata(path) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => Ok(true), + Err(_) => Err(refused()), + Ok(_) => Ok(false), + } +} + +pub(super) fn require_no_pending(candidate: &Candidate, run: &str) -> Result<(), CandidateError> { + if absent( + &directory(candidate, run)? + .join(FILE) + .with_extension("pending"), + )? { + Ok(()) + } else { + Err(refused()) + } +} + +fn no_pending(root: &Path) -> Result<(), CandidateError> { + for name in [ + "state.pending", + "absent-publication-cleanup.pending", + "absent-publication-retirement.pending", + "source-device-rebind.pending", + ] { + if !absent(&root.join(name))? { + return Err(refused()); + } + } + Ok(()) +} + +fn encoded(value: &impl Serialize) -> Result, CandidateError> { + serde_json::to_vec_pretty(value).map_err(|_| refused()) +} + +struct PinnedRaw { + file: File, + bytes: Vec, + file_id: (u64, u64), + parent_id: (u64, u64), + path: PathBuf, +} + +impl PinnedRaw { + fn reverify(&self) -> Result<(), CandidateError> { + let parent = self.path.parent().ok_or_else(refused)?; + state::check_private_directory(parent).map_err(|_| refused())?; + let parent_meta = fs::symlink_metadata(parent).map_err(|_| refused())?; + let meta = self.file.metadata().map_err(|_| refused())?; + let path_meta = fs::symlink_metadata(&self.path).map_err(|_| refused())?; + // SAFETY: geteuid has no arguments or side effects. + let uid = unsafe { libc::geteuid() }; + if (parent_meta.dev(), parent_meta.ino()) != self.parent_id + || (meta.dev(), meta.ino()) != self.file_id + || (path_meta.dev(), path_meta.ino()) != self.file_id + || !meta.is_file() + || meta.nlink() != 1 + || meta.uid() != uid + || path_meta.uid() != uid + || meta.mode() & 0o7777 != 0o600 + || path_meta.mode() & 0o7777 != 0o600 + || path_meta.nlink() != 1 + || meta.len() != self.bytes.len() as u64 + { + return Err(refused()); + } + let mut file = &self.file; + file.seek(SeekFrom::Start(0)).map_err(|_| refused())?; + let mut bytes = Vec::new(); + file.take(self.bytes.len() as u64 + 1) + .read_to_end(&mut bytes) + .map_err(|_| refused())?; + if bytes != self.bytes { + return Err(refused()); + } + Ok(()) + } +} + +fn pin_raw(path: &Path, limit: u64, allow_empty: bool) -> Result { + let parent = path.parent().ok_or_else(refused)?; + state::check_private_directory(parent).map_err(|_| refused())?; + let parent_meta = fs::symlink_metadata(parent).map_err(|_| refused())?; + let mut file = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK) + .open(path) + .map_err(|_| refused())?; + let before = file.metadata().map_err(|_| refused())?; + // SAFETY: geteuid has no arguments or side effects. + if !before.is_file() + || before.nlink() != 1 + || before.uid() != unsafe { libc::geteuid() } + || before.mode() & 0o7777 != 0o600 + || before.len() > limit + || (!allow_empty && before.len() == 0) + { + return Err(refused()); + } + let mut bytes = Vec::new(); + file.by_ref() + .take(limit + 1) + .read_to_end(&mut bytes) + .map_err(|_| refused())?; + file.seek(SeekFrom::Start(0)).map_err(|_| refused())?; + let mut confirmation = Vec::new(); + file.by_ref() + .take(limit + 1) + .read_to_end(&mut confirmation) + .map_err(|_| refused())?; + let after = file.metadata().map_err(|_| refused())?; + let path_after = fs::symlink_metadata(path).map_err(|_| refused())?; + let parent_after = fs::symlink_metadata(parent).map_err(|_| refused())?; + if bytes != confirmation + || bytes.len() as u64 != before.len() + || [after.dev(), path_after.dev()] != [before.dev(); 2] + || [after.ino(), path_after.ino()] != [before.ino(); 2] + || [after.len(), path_after.len()] != [before.len(); 2] + || [after.mtime(), path_after.mtime()] != [before.mtime(); 2] + || [after.mtime_nsec(), path_after.mtime_nsec()] != [before.mtime_nsec(); 2] + || [after.ctime(), path_after.ctime()] != [before.ctime(); 2] + || [after.ctime_nsec(), path_after.ctime_nsec()] != [before.ctime_nsec(); 2] + || path_after.nlink() != 1 + || (parent_after.dev(), parent_after.ino()) != (parent_meta.dev(), parent_meta.ino()) + { + return Err(refused()); + } + Ok(PinnedRaw { + file, + bytes, + file_id: (before.dev(), before.ino()), + parent_id: (parent_meta.dev(), parent_meta.ino()), + path: path.to_owned(), + }) +} + +fn raw( + path: &Path, + limit: u64, + allow_empty: bool, +) -> Result<(Vec, (u64, u64)), CandidateError> { + let pin = pin_raw(path, limit, allow_empty)?; + Ok((pin.bytes, pin.file_id)) +} + +#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct Witness { + version: u8, + run: String, + owner: String, + namespace: String, + plan: String, + original_ready_sha256: String, + stopped_raw_sha256: String, + absent_intent_raw_sha256: String, + retirement_raw_sha256: String, + original_owner_raw_sha256: String, + current_owner_raw_sha256: String, + host_boot_micros: u64, + previous_guest_boot: String, + current_guest_boot: String, + old_share: ProjectShareIntent, + current_share: ProjectShareIntent, + retained_volume_projections: BTreeMap, + retained_volumes: BTreeMap, +} + +impl Witness { + fn project(&self, receipt: &Receipt) -> Result { + let mut projected = receipt.clone(); + let shared = projected + .source + .as_mut() + .and_then(|source| source.shared.as_mut()) + .ok_or_else(refused)?; + if *shared != self.old_share { + return Err(refused()); + } + shared.device = self.current_share.device; + if *shared != self.current_share { + return Err(refused()); + } + let scopes = receipt + .resources + .values() + .filter_map(|resource| resource.cache.as_ref().map(|cache| cache.scope.as_str())) + .collect::>(); + let binding = projected.source.as_mut().ok_or_else(refused)?; + binding.cache_scope = super::dependency_cache::ReplayScope::project( + &self.old_share, + &self.current_share, + binding.cache_scope.as_ref(), + &scopes, + &sha256(&encoded(self)?), + )?; + Ok(projected) + } + + fn unchanged_identity(&self, receipt: &Receipt) -> bool { + self.version == 1 + && self.run == receipt.run + && self.owner == receipt.owner + && self.namespace == receipt.namespace + && self.plan == receipt.plan_id + && self.old_share.device != self.current_share.device + && self.old_share.project == self.current_share.project + && self.old_share.guest_path == self.current_share.guest_path + && self.old_share.inode == self.current_share.inode + && self.old_share.unfiltered_source == self.current_share.unfiltered_source + && receipt.source.as_ref().and_then(|s| s.shared.as_ref()) == Some(&self.old_share) + } +} + +/// Corroborate an explicitly selected legacy HTTPS device translation. This +/// witness does not establish original host-volume continuity or authorize any +/// HTTPS effect; the caller must independently select and pin both current inodes. +pub(in crate::provider) fn https_devices( + candidate: &Candidate, + run: &str, + expected: &str, +) -> Result<(u64, u64), CandidateError> { + if !super::hex(run, 32) || !super::hex(expected, 64) { + return Err(refused()); + } + let root = directory(candidate, run)?; + no_pending(&root)?; + for entry in fs::read_dir(&root).map_err(|_| refused())? { + let name = entry.map_err(|_| refused())?.file_name(); + if name.to_str().is_none_or(|s| s.ends_with(".pending")) { + return Err(refused()); + } + } + let pin = pin_raw(&root.join(FILE), LIMIT, false)?; + let witness: Witness = serde_json::from_slice(&pin.bytes).map_err(|_| refused())?; + let owner_root = candidate.state_root.join("run/smolvm"); + let owner_pin = pin_raw(&owner_root.join("owner.json"), STATE_LIMIT, false)?; + let owner: state::Owner = serde_json::from_slice(&owner_pin.bytes).map_err(|_| refused())?; + let receipt_pin = pin_raw(&root.join("state.json"), STATE_LIMIT, false)?; + let receipt: Receipt = serde_json::from_slice(&receipt_pin.bytes).map_err(|_| refused())?; + if sha256(&pin.bytes) != expected + || pin.bytes != encoded(&witness)? + || witness.version != 1 + || witness.run != run + || witness.owner != owner.token + || sha256(&owner_pin.bytes) != witness.current_owner_raw_sha256 + || owner.project_share.as_ref() != Some(&witness.current_share) + || owner.guest_boot_id.as_deref() != Some(witness.current_guest_boot.as_str()) + || lifecycle::host_filesystem::host_boot_micros()? != witness.host_boot_micros + || witness.old_share.device == witness.current_share.device + || witness.old_share.project != witness.current_share.project + || witness.old_share.guest_path != witness.current_share.guest_path + || witness.old_share.inode != witness.current_share.inode + || witness.old_share.unfiltered_source != witness.current_share.unfiltered_source + || receipt.run != witness.run + || receipt.owner != witness.owner + || receipt.namespace != witness.namespace + || receipt.plan_id != witness.plan + || receipt.phase != "stopped-data-retained" + || receipt.source.as_ref().and_then(|s| s.shared.as_ref()) != Some(&witness.current_share) + || receipt + .resources + .values() + .any(|r| r.kind != Kind::Volume && r.phase != "absent") + { + return Err(refused()); + } + witness.current_share.validate().map_err(|_| refused())?; + pin.reverify()?; + owner_pin.reverify()?; + receipt_pin.reverify()?; + Ok((witness.old_share.device, witness.current_share.device)) +} + +#[cfg(test)] +pub(in crate::provider) fn fixture_https_witness( + candidate: &Candidate, + receipt: &Receipt, + old_device: u64, +) -> Vec { + let owner = state::Owner::load(candidate).unwrap(); + let current_share = owner.project_share.clone().unwrap(); + let mut old_share = current_share.clone(); + old_share.device = old_device; + encoded(&Witness { + version: 1, + run: receipt.run.clone(), + owner: owner.token, + namespace: receipt.namespace.clone(), + plan: receipt.plan_id.clone(), + original_ready_sha256: "1".repeat(64), + stopped_raw_sha256: sha256(&encoded(receipt).unwrap()), + absent_intent_raw_sha256: "3".repeat(64), + retirement_raw_sha256: "4".repeat(64), + original_owner_raw_sha256: "5".repeat(64), + current_owner_raw_sha256: sha256( + &fs::read(candidate.state_root.join("run/smolvm/owner.json")).unwrap(), + ), + host_boot_micros: lifecycle::host_filesystem::host_boot_micros().unwrap(), + previous_guest_boot: "old".into(), + current_guest_boot: owner.guest_boot_id.unwrap(), + old_share, + current_share, + retained_volume_projections: BTreeMap::new(), + retained_volumes: BTreeMap::new(), + }) + .unwrap() +} + +/// A later ordinary replay uses the original witness as immutable provenance, +/// never as authority to change the current Owner, guest boot or source device. +pub(super) fn verify_cache_scope_origin( + candidate: &Candidate, + receipt: &Receipt, +) -> Result<(), CandidateError> { + let Some(binding) = &receipt.source else { + return Ok(()); + }; + let Some(scope) = &binding.cache_scope else { + return Ok(()); + }; + let root = directory(candidate, &receipt.run)?; + no_pending(&root)?; + let pin = pin_raw(&root.join(FILE), LIMIT, false)?; + let witness: Witness = serde_json::from_slice(&pin.bytes).map_err(|_| refused())?; + if encoded(&witness)? != pin.bytes + || witness.version != 1 + || witness.run != receipt.run + || witness.owner != receipt.owner + || witness.namespace != receipt.namespace + || witness.plan != receipt.plan_id + || binding.shared.as_ref() != Some(&witness.current_share) + || !scope.matches_origin( + &sha256(&pin.bytes), + &witness.old_share, + &witness.current_share, + ) + { + return Err(refused()); + } + let expected = scope.scope(&witness.current_share.project)?; + if receipt + .resources + .values() + .filter_map(|resource| resource.cache.as_ref()) + .any(|cache| cache.scope != expected) + { + return Err(refused()); + } + pin.reverify() +} + +/// Historical cache provenance also pins the exact cleanup/retirement bytes +/// used to archive its old journal; another same-schema proof is insufficient. +pub(super) fn verify_cache_scope_cleanup_origin( + candidate: &Candidate, + receipt: &Receipt, + intent_sha256: &str, + retirement_sha256: &str, +) -> Result<(), CandidateError> { + verify_cache_scope_origin(candidate, receipt)?; + let Some(scope) = receipt + .source + .as_ref() + .and_then(|source| source.cache_scope.as_ref()) + else { + return Ok(()); + }; + let root = directory(candidate, &receipt.run)?; + let pin = pin_raw(&root.join(FILE), LIMIT, false)?; + let witness: Witness = serde_json::from_slice(&pin.bytes).map_err(|_| refused())?; + if encoded(&witness)? != pin.bytes + || !scope.matches_origin( + &sha256(&pin.bytes), + &witness.old_share, + &witness.current_share, + ) + || witness.absent_intent_raw_sha256 != intent_sha256 + || witness.retirement_raw_sha256 != retirement_sha256 + { + return Err(refused()); + } + pin.reverify() +} + +/// Pinned complete witness for the original stopped receipt only. A subsequent +/// graph generation cannot use this object as cleanup or source authority. +pub(super) struct Selected { + witness: Witness, + pin: PinnedRaw, + original: Receipt, + projected: Receipt, + root: PathBuf, +} + +impl Selected { + pub(super) fn raw_sha256(&self) -> String { + sha256(&self.pin.bytes) + } + + pub(super) fn source_receipt(&self) -> &Receipt { + &self.projected + } + + pub(super) fn apply_to_new_attempt(&self, receipt: &mut Receipt) -> Result<(), CandidateError> { + if encoded(receipt)? != encoded(&self.original)? { + return Err(refused()); + } + receipt.source = self.projected.source.clone(); + Ok(()) + } + + /// Recheck raw bytes, pathname identity, Owner, guest and volumes before the + /// first new-attempt effect/write. The caller still holds the Engine lease. + pub(super) fn reverify(&self, engine: &Engine<'_>) -> Result<(), CandidateError> { + self.pin.reverify()?; + validate_current(engine, &self.root, &self.original, &self.witness)?; + let current: Receipt = state::read(&self.root.join("state.json"))?; + if encoded(¤t)? != encoded(&self.original)? + || encoded(&self.witness.project(¤t)?)? != encoded(&self.projected)? + { + return Err(refused()); + } + Ok(()) + } +} + +fn validate_current( + engine: &Engine<'_>, + root: &Path, + receipt: &Receipt, + witness: &Witness, +) -> Result<(), CandidateError> { + no_pending(root)?; + if receipt.phase != "stopped-data-retained" || !witness.unchanged_identity(receipt) { + return Err(refused()); + } + let state_bytes = raw(&root.join("state.json"), STATE_LIMIT, false)?.0; + if state_bytes != encoded(receipt)? || sha256(&state_bytes) != witness.stopped_raw_sha256 { + return Err(refused()); + } + for (name, expected) in [ + ( + "absent-publication-cleanup.json", + &witness.absent_intent_raw_sha256, + ), + ( + "absent-publication-retirement.json", + &witness.retirement_raw_sha256, + ), + ] { + if sha256(&raw(&root.join(name), STATE_LIMIT, false)?.0) != *expected { + return Err(refused()); + } + } + let candidate = engine.guest().candidate(); + let owner_root = candidate.state_root.join("run/smolvm"); + if sha256(&raw(&owner_root.join("owner.json"), STATE_LIMIT, false)?.0) + != witness.current_owner_raw_sha256 + || lifecycle::host_filesystem::host_boot_micros()? != witness.host_boot_micros + || engine.guest().boot_id() != witness.current_guest_boot + || engine.guest().project_share() != Some(&witness.current_share) + { + return Err(refused()); + } + witness.current_share.validate().map_err(|_| refused())?; + source::verify_shared_mount(engine, &witness.current_share).map_err(|_| refused())?; + if restore::observed_volumes(engine, receipt)? != witness.retained_volumes { + return Err(refused()); + } + host_pin_recovery::verify_volume_projections( + engine, + receipt, + &witness.retained_volume_projections, + ) + .map_err(|_| refused())?; + for resource in receipt.resources.values() { + if resource.kind != Kind::Volume { + let observed = inspect_resource(engine, receipt, resource)?; + let mut by_name = resource.clone(); + by_name.id = None; + if resource.phase != "absent" + || observed.is_some() + || inspect_resource(engine, receipt, &by_name)?.is_some() + { + return Err(refused()); + } + } + } + cleanup_enrollment::retention(root, receipt)?; + Ok(()) +} + +/// A committed witness applies only to its original stopped state. Once the +/// next generation has a current share it remains audit data, never authority. +pub(super) fn select( + engine: &Engine<'_>, + receipt: &Receipt, + root: &Path, +) -> Result, CandidateError> { + if !absent(&root.join(FILE).with_extension("pending"))? { + return Err(refused()); + } + let path = root.join(FILE); + if absent(&path)? { + return Ok(None); + } + let pin = pin_raw(&path, LIMIT, false)?; + let witness: Witness = serde_json::from_slice(&pin.bytes).map_err(|_| refused())?; + if pin.bytes != encoded(&witness)? { + return Err(refused()); + } + let state_sha = sha256(&raw(&root.join("state.json"), STATE_LIMIT, false)?.0); + if state_sha != witness.stopped_raw_sha256 { + // The witness is historical only when normal source ownership has moved on. + if receipt.owner == witness.owner + && receipt.run == witness.run + && receipt.namespace == witness.namespace + && receipt.source.as_ref().and_then(|s| s.shared.as_ref()) + == engine.guest().project_share() + { + return Ok(None); + } + return Err(refused()); + } + validate_current(engine, root, receipt, &witness)?; + let projected = witness.project(receipt)?; + Ok(Some(Selected { + witness, + pin, + original: receipt.clone(), + projected, + root: root.to_owned(), + })) +} + +fn selected_witness( + candidate: &Candidate, + run: &str, + retired: &foreground::transport::Retired, + engine: &Engine<'_>, +) -> Result<(Witness, Receipt, PathBuf), CandidateError> { + let (receipt, root) = load(candidate, engine, run)?; + let proof = super::absent_publication_cleanup::verify_completed_under( + candidate, engine, run, &receipt, retired, + )? + .ok_or_else(refused)?; + let old_share = proof.old_share.ok_or_else(refused)?; + let current_share = proof.current_share.ok_or_else(refused)?; + let state_bytes = raw(&root.join("state.json"), STATE_LIMIT, false)?.0; + if state_bytes != encoded(&receipt)? + || receipt.phase != "stopped-data-retained" + || sha256(&state_bytes) != proof.completed_stopped_sha256 + || engine.guest().project_share() != Some(¤t_share) + || lifecycle::host_filesystem::host_boot_micros()? != proof.host_boot_micros + || engine.guest().boot_id() != proof.current_guest_boot + { + return Err(refused()); + } + let witness = Witness { + version: 1, + run: receipt.run.clone(), + owner: receipt.owner.clone(), + namespace: receipt.namespace.clone(), + plan: receipt.plan_id.clone(), + original_ready_sha256: proof.original_ready_sha256, + stopped_raw_sha256: proof.completed_stopped_sha256, + absent_intent_raw_sha256: proof.intent_raw_sha256, + retirement_raw_sha256: proof.retirement_raw_sha256, + original_owner_raw_sha256: proof.original_owner_sha256, + current_owner_raw_sha256: proof.current_owner_sha256, + host_boot_micros: proof.host_boot_micros, + previous_guest_boot: proof.previous_guest_boot, + current_guest_boot: proof.current_guest_boot, + old_share, + current_share, + retained_volume_projections: proof.retained_volumes, + retained_volumes: restore::observed_volumes(engine, &receipt)?, + }; + validate_current(engine, &root, &receipt, &witness)?; + retired.verify()?; + Ok((witness, receipt, root)) +} + +fn selected_output(witness: &Witness, committed: bool) -> Result { + Ok(json!({ + "run": witness.run, + "owner": witness.owner, + "namespace": witness.namespace, + "plan": witness.plan, + "stopped_receipt_sha256": witness.stopped_raw_sha256, + "selection_sha256": sha256(&encoded(witness)?), + "committed": committed, + "qualification": "explicit-legacy-migration-original-volume-continuity-unproven", + })) +} + +/// Read-only selection holds the retired publisher lock before the Engine lease. +pub fn inspect(candidate: &Candidate, run: &str) -> Result { + let retired = foreground::transport::Retired::acquire(candidate, run)?.ok_or_else(refused)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + let (witness, receipt, root) = selected_witness(candidate, run, &retired, &engine)?; + let committed = if absent(&root.join(FILE))? { + false + } else { + let selected = select(&engine, &receipt, &root)?.ok_or_else(refused)?; + selected.witness == witness + }; + if !absent(&root.join(FILE).with_extension("pending"))? + || !committed && !absent(&root.join(FILE))? + { + return Err(refused()); + } + retired.verify()?; + selected_output(&witness, committed) +} + +/// Publish one immutable selected device-only source witness; no graph receipt +/// or cleanup/retirement digest is rewritten. +pub fn recover(candidate: &Candidate, run: &str, expected: &str) -> Result { + if !super::hex(expected, 64) { + return Err(refused()); + } + let retired = foreground::transport::Retired::acquire(candidate, run)?.ok_or_else(refused)?; + let engine = Engine::connect_cleanup_wait(candidate)?; + require_no_pending(candidate, run)?; + let (witness, receipt, root) = selected_witness(candidate, run, &retired, &engine)?; + let bytes = encoded(&witness)?; + if bytes.len() as u64 > LIMIT || sha256(&bytes) != expected { + return Err(refused()); + } + let path = root.join(FILE); + if !absent(&path)? { + let selected = select(&engine, &receipt, &root)?.ok_or_else(refused)?; + if selected.pin.bytes != bytes { + return Err(refused()); + } + return selected_output(&witness, true); + } + retired.verify()?; + validate_current(&engine, &root, &receipt, &witness)?; + state::write(&path, &witness)?; + let selected = select(&engine, &receipt, &root)?.ok_or_else(refused)?; + if selected.pin.bytes != bytes { + return Err(refused()); + } + retired.verify()?; + selected_output(&witness, true) +} + +#[cfg(all(test, feature = "environment-launcher"))] +pub(in crate::provider::graph) fn fixture_wrong_volume_projection(bytes: &[u8]) -> Vec { + let mut witness: Witness = serde_json::from_slice(bytes).unwrap(); + witness + .retained_volume_projections + .get_mut("volume:data") + .unwrap()["observed_created_at"] = json!("foreign"); + encoded(&witness).unwrap() +} + +#[cfg(test)] +mod tests { + use super::*; + use std::os::unix::fs::PermissionsExt; + + #[test] + fn https_translation_requires_selected_witness_current_owner_boot_share_and_scope() { + let fixture = super::super::tests::Fixture::new(); + fs::write(fixture.0.join("compose.yaml"), "services: {}\n").unwrap(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let share = ProjectShareIntent::approve(&fixture.0, true).unwrap(); + let mut owner = state::Owner::create( + &candidate, + crate::provider::Profile::Research, + None, + crate::provider::NetworkIntent::Isolated, + ) + .unwrap(); + owner.project_share = Some(share.clone()); + owner.guest_boot_id = Some("new".into()); + owner.save(&candidate).unwrap(); + struct Alias(PathBuf, PathBuf); + impl Drop for Alias { + fn drop(&mut self) { + if fs::read_link(&self.0).ok().as_ref() == Some(&self.1) { + let _ = fs::remove_file(&self.0); + } + } + } + let _alias = Alias( + owner.short_home.clone(), + candidate.state_root.join("run/smolvm/home"), + ); + let mut old_share = share.clone(); + old_share.device += 7; + let run = "a".repeat(32); + let parent = candidate.state_root.join("run/graphs"); + fs::create_dir(&parent).unwrap(); + fs::set_permissions(&parent, fs::Permissions::from_mode(0o700)).unwrap(); + let root = parent.join(&run); + fs::create_dir(&root).unwrap(); + fs::set_permissions(&root, fs::Permissions::from_mode(0o700)).unwrap(); + let receipt: Receipt = serde_json::from_value(json!({ + "version":1,"run":run,"owner":owner.token,"namespace":"c".repeat(64), + "plan_id":"d".repeat(64),"phase":"stopped-data-retained","readiness":{}, + "source":{"shared":share,"revision":"e".repeat(64),"archive_sha256":"f".repeat(64),"selection_sha256":"0".repeat(64)}, + "resources":{} + })).unwrap(); + private_file(&root.join("state.json"), &encoded(&receipt).unwrap()); + let witness = Witness { + version: 1, + run: run.clone(), + owner: owner.token.clone(), + namespace: receipt.namespace.clone(), + plan: receipt.plan_id.clone(), + original_ready_sha256: "1".repeat(64), + stopped_raw_sha256: "2".repeat(64), + absent_intent_raw_sha256: "3".repeat(64), + retirement_raw_sha256: "4".repeat(64), + original_owner_raw_sha256: "5".repeat(64), + current_owner_raw_sha256: sha256( + &fs::read(candidate.state_root.join("run/smolvm/owner.json")).unwrap(), + ), + host_boot_micros: lifecycle::host_filesystem::host_boot_micros().unwrap(), + previous_guest_boot: "old".into(), + current_guest_boot: "new".into(), + old_share, + current_share: share.clone(), + retained_volume_projections: BTreeMap::new(), + retained_volumes: BTreeMap::new(), + }; + let path = root.join(FILE); + let bytes = encoded(&witness).unwrap(); + private_file(&path, &bytes); + let expected = sha256(&bytes); + assert_eq!( + https_devices(&candidate, &run, &expected).unwrap(), + (share.device + 7, share.device) + ); + assert!(https_devices(&candidate, &run, &"f".repeat(64)).is_err()); + private_file(&root.join("unknown.pending"), b"retained pending intent"); + assert!(https_devices(&candidate, &run, &expected).is_err()); + fs::remove_file(root.join("unknown.pending")).unwrap(); + for control in ["namespace", "owner", "boot", "share", "one_device"] { + let mut changed = witness.clone(); + match control { + "namespace" => changed.namespace = "9".repeat(64), + "owner" => changed.current_owner_raw_sha256 = "9".repeat(64), + "boot" => changed.host_boot_micros += 1, + "share" => changed.current_share.inode += 1, + _ => changed.old_share.device = changed.current_share.device, + } + let changed = encoded(&changed).unwrap(); + fs::write(&path, &changed).unwrap(); + assert!( + https_devices(&candidate, &run, &sha256(&changed)).is_err(), + "{control}" + ); + } + } + + fn private_file(path: &Path, bytes: &[u8]) { + let mut file = OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .open(path) + .unwrap(); + use std::io::Write; + file.write_all(bytes).unwrap(); + } + + #[test] + fn cached_projection_adds_continuity_only_to_the_new_attempt() { + let fixture = super::super::tests::Fixture::new(); + fs::write(fixture.0.join("compose.yaml"), "services: {}\n").unwrap(); + let current_share = ProjectShareIntent::approve(&fixture.0, true).unwrap(); + let mut old_share = current_share.clone(); + old_share.device += 7; + let legacy_scope = sha256( + &serde_json::to_vec(&( + "hack-dependency-cache-scope-v1", + &fixture.0, + old_share.device, + old_share.inode, + )) + .unwrap(), + ); + let receipt: Receipt = serde_json::from_value(json!({ + "version":1,"run":"a".repeat(32),"owner":"b".repeat(32),"namespace":"c".repeat(64), + "plan_id":"d".repeat(64),"phase":"stopped-data-retained","readiness":{}, + "source":{"shared":old_share,"revision":"e".repeat(64),"archive_sha256":"f".repeat(64),"selection_sha256":"0".repeat(64)}, + "resources":{"volume:deps":{"kind":"volume","key":"deps","name":format!("hack-cache-v5-{}","9".repeat(64)),"phase":"present","id":null,"image":null,"cache":{"scope":legacy_scope,"fingerprint":"9".repeat(64),"image":format!("sha256:{}","8".repeat(64))}}} + })).unwrap(); + let original = encoded(&receipt).unwrap(); + let witness = Witness { + version: 1, + run: receipt.run.clone(), + owner: receipt.owner.clone(), + namespace: receipt.namespace.clone(), + plan: receipt.plan_id.clone(), + original_ready_sha256: "1".repeat(64), + stopped_raw_sha256: "2".repeat(64), + absent_intent_raw_sha256: "3".repeat(64), + retirement_raw_sha256: "4".repeat(64), + original_owner_raw_sha256: "5".repeat(64), + current_owner_raw_sha256: "6".repeat(64), + host_boot_micros: 7, + previous_guest_boot: "old".into(), + current_guest_boot: "new".into(), + old_share, + current_share, + retained_volume_projections: BTreeMap::new(), + retained_volumes: BTreeMap::new(), + }; + let projected = witness.project(&receipt).unwrap(); + let scope = projected + .source + .as_ref() + .unwrap() + .cache_scope + .as_ref() + .unwrap(); + assert_eq!(scope.scope(&fixture.0).unwrap(), legacy_scope); + assert_eq!( + projected.source.as_ref().unwrap().shared, + Some(witness.current_share.clone()) + ); + assert!(super::super::same_resource_bindings( + &projected.resources, + &receipt.resources + )); + assert_eq!(encoded(&receipt).unwrap(), original); + let decoded: Receipt = serde_json::from_slice(&encoded(&projected).unwrap()).unwrap(); + assert_eq!(decoded.source, projected.source); + + let home = super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&home.0).unwrap(); + let root = directory(&candidate, &receipt.run).unwrap(); + state::private_directory(&root).unwrap(); + assert!(verify_cache_scope_origin(&candidate, &projected).is_err()); + private_file(&root.join(FILE), &encoded(&witness).unwrap()); + verify_cache_scope_origin(&candidate, &projected).unwrap(); + verify_cache_scope_cleanup_origin( + &candidate, + &projected, + &witness.absent_intent_raw_sha256, + &witness.retirement_raw_sha256, + ) + .unwrap(); + assert!( + verify_cache_scope_cleanup_origin( + &candidate, + &projected, + &"7".repeat(64), + &witness.retirement_raw_sha256 + ) + .is_err() + ); + assert!( + verify_cache_scope_cleanup_origin( + &candidate, + &projected, + &witness.absent_intent_raw_sha256, + &"7".repeat(64) + ) + .is_err() + ); + let mut changed_plan = projected.clone(); + changed_plan.plan_id = "7".repeat(64); + assert!(verify_cache_scope_origin(&candidate, &changed_plan).is_err()); + let mut foreign_origin = witness.clone(); + foreign_origin.owner = "7".repeat(32); + fs::remove_file(root.join(FILE)).unwrap(); + private_file(&root.join(FILE), &encoded(&foreign_origin).unwrap()); + assert!(verify_cache_scope_origin(&candidate, &projected).is_err()); + fs::remove_file(root.join(FILE)).unwrap(); + private_file(&root.join(FILE), &encoded(&witness).unwrap()); + let mut foreign_cache = projected.clone(); + foreign_cache + .resources + .get_mut("volume:deps") + .unwrap() + .cache + .as_mut() + .unwrap() + .scope = "0".repeat(64); + assert!(verify_cache_scope_origin(&candidate, &foreign_cache).is_err()); + private_file(&root.join(FILE).with_extension("pending"), b"partial"); + assert!(verify_cache_scope_origin(&candidate, &projected).is_err()); + assert_eq!( + fs::read(root.join(FILE).with_extension("pending")).unwrap(), + b"partial" + ); + + let mut foreign = receipt.clone(); + foreign + .resources + .get_mut("volume:deps") + .unwrap() + .cache + .as_mut() + .unwrap() + .scope = "0".repeat(64); + assert!(witness.project(&foreign).is_err()); + assert_eq!(encoded(&receipt).unwrap(), original); + } + + #[test] + fn pinned_witness_refuses_equal_bytes_at_a_replacement_path_or_parent() { + let fixture = super::super::tests::Fixture::new(); + let parent = fixture.0.join("witness"); + fs::create_dir(&parent).unwrap(); + fs::set_permissions(&parent, fs::Permissions::from_mode(0o700)).unwrap(); + let path = parent.join(FILE); + private_file(&path, b"proof"); + let pin = pin_raw(&path, LIMIT, false).unwrap(); + pin.reverify().unwrap(); + + fs::set_permissions(&path, fs::Permissions::from_mode(0o644)).unwrap(); + assert!(pin.reverify().is_err()); + fs::set_permissions(&path, fs::Permissions::from_mode(0o600)).unwrap(); + pin.reverify().unwrap(); + + fs::rename(&path, parent.join("archived")).unwrap(); + private_file(&path, b"proof"); + assert!(pin.reverify().is_err()); + + fs::rename(&parent, fixture.0.join("old-witness")).unwrap(); + fs::create_dir(&parent).unwrap(); + fs::set_permissions(&parent, fs::Permissions::from_mode(0o700)).unwrap(); + private_file(&path, b"proof"); + assert!(pin.reverify().is_err()); + } + + #[test] + fn pending_witness_is_preserved_and_refuses_publication() { + let fixture = super::super::tests::Fixture::new(); + let candidate = Candidate::discover(&fixture.0).unwrap(); + let run = "a".repeat(32); + let root = directory(&candidate, &run).unwrap(); + state::private_directory(&root).unwrap(); + let pending = root.join(FILE).with_extension("pending"); + private_file(&pending, b""); + assert!(require_no_pending(&candidate, &run).is_err()); + assert!(no_pending(&root).is_err()); + assert_eq!(fs::read(&pending).unwrap(), b""); + fs::remove_file(&pending).unwrap(); + private_file(&pending, b"foreign"); + assert!(require_no_pending(&candidate, &run).is_err()); + assert!(no_pending(&root).is_err()); + assert_eq!(fs::read(&pending).unwrap(), b"foreign"); + } + + #[test] + fn source_projection_changes_only_the_selected_device() { + let old_share = ProjectShareIntent { + project: PathBuf::from("/private/tmp/example-project"), + guest_path: "/mnt/hack-projects/exact".into(), + device: 17, + inode: 23, + unfiltered_source: true, + }; + let mut current_share = old_share.clone(); + current_share.device = 19; + let receipt: Receipt = serde_json::from_value(json!({ + "version": 1, + "run": "a".repeat(32), + "owner": "b".repeat(32), + "namespace": "c".repeat(64), + "plan_id": "d".repeat(64), + "phase": "stopped-data-retained", + "readiness": {}, + "resources": {}, + "source": { + "shared": old_share, + "revision": "e".repeat(64), + "archive_sha256": "f".repeat(64), + "selection_sha256": "0".repeat(64), + }, + })) + .unwrap(); + let witness = Witness { + version: 1, + run: receipt.run.clone(), + owner: receipt.owner.clone(), + namespace: receipt.namespace.clone(), + plan: receipt.plan_id.clone(), + original_ready_sha256: "1".repeat(64), + stopped_raw_sha256: "2".repeat(64), + absent_intent_raw_sha256: "3".repeat(64), + retirement_raw_sha256: "4".repeat(64), + original_owner_raw_sha256: "5".repeat(64), + current_owner_raw_sha256: "6".repeat(64), + host_boot_micros: 7, + previous_guest_boot: "old".into(), + current_guest_boot: "new".into(), + old_share, + current_share, + retained_volume_projections: BTreeMap::new(), + retained_volumes: BTreeMap::new(), + }; + assert!(witness.unchanged_identity(&receipt)); + let projected = witness.project(&receipt).unwrap(); + let mut expected = receipt.clone(); + expected + .source + .as_mut() + .unwrap() + .shared + .as_mut() + .unwrap() + .device = 19; + assert_eq!(encoded(&projected).unwrap(), encoded(&expected).unwrap()); + assert_eq!(receipt.source.unwrap().shared.unwrap().device, 17); + + let mut foreign = witness.clone(); + foreign.current_share.inode += 1; + assert!(!foreign.unchanged_identity(&expected)); + assert!(foreign.project(&expected).is_err()); + } +} diff --git a/packages/runtime-core/src/provider/graph/startup.rs b/packages/runtime-core/src/provider/graph/startup.rs index fd046759f..6397a5185 100644 --- a/packages/runtime-core/src/provider/graph/startup.rs +++ b/packages/runtime-core/src/provider/graph/startup.rs @@ -389,9 +389,19 @@ pub(super) fn require_dependency_rebind_complete( } } +/// Same-boot recovery accepts only a completed refresh for its exact selection. +#[cfg(target_os = "macos")] +pub(super) fn require_dependency_rebind_recovery_complete( + root: &Path, + receipt: &Receipt, + boot: &str, +) -> Result<(), CandidateError> { + runtime::require_dependency_rebind_recovery_complete(root, receipt, boot) +} + +#[cfg(target_os = "macos")] /// Explicit owned cleanup preserves dependency refresh evidence before a later /// restore may create new helper generations. This never replays the refresh. -#[cfg(target_os = "macos")] pub(super) fn archive_dependency_rebind_after_cleanup( root: &Path, original: &Receipt, @@ -401,6 +411,34 @@ pub(super) fn archive_dependency_rebind_after_cleanup( runtime::archive_dependency_rebind_after_cleanup(root, original, cleaned, boot) } +#[cfg(target_os = "macos")] +pub(super) fn dependency_rebind_boot( + root: &Path, + receipt: &Receipt, +) -> Result, CandidateError> { + runtime::dependency_rebind_boot(root, receipt) +} + +#[cfg(target_os = "macos")] +pub(super) fn archive_retired_dependency_rebind( + root: &Path, + original: &Receipt, + cleaned: &Receipt, + boot: &str, + verify: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + runtime::archive_retired_dependency_rebind(root, original, cleaned, boot, verify) +} + +#[cfg(target_os = "macos")] +pub(super) fn retired_dependency_rebind_archive_complete( + root: &Path, + original: &Receipt, + boot: &str, +) -> Result { + runtime::retired_dependency_rebind_archive_complete(root, original, boot) +} + pub(super) fn guest_directory(run: &str, generation: &str) -> String { format!("/storage/hack-graph-startup/{run}/{generation}") } diff --git a/packages/runtime-core/src/provider/graph/startup/native_test.rs b/packages/runtime-core/src/provider/graph/startup/native_test.rs index 93e430abc..53b5782c1 100644 --- a/packages/runtime-core/src/provider/graph/startup/native_test.rs +++ b/packages/runtime-core/src/provider/graph/startup/native_test.rs @@ -795,6 +795,8 @@ fn qualify_active_rebind(refreshable: bool, cancel_after_fence: bool) { Instant::now() + Duration::from_secs(60), &mut fresh, &restored_generation, + // This owned driver fixture has no historical absence intent. + &|| Ok(()), ); let assertions = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { let restored = restored.as_ref().unwrap(); diff --git a/packages/runtime-core/src/provider/graph/startup/runtime.rs b/packages/runtime-core/src/provider/graph/startup/runtime.rs index 5025035d8..085c22c73 100644 --- a/packages/runtime-core/src/provider/graph/startup/runtime.rs +++ b/packages/runtime-core/src/provider/graph/startup/runtime.rs @@ -20,6 +20,14 @@ pub(super) fn require_dependency_rebind_complete( rebind::require_complete(root, receipt) } +pub(super) fn require_dependency_rebind_recovery_complete( + root: &Path, + receipt: &Receipt, + boot: &str, +) -> Result<(), CandidateError> { + rebind::require_recovery_complete(root, receipt, boot) +} + pub(super) fn archive_dependency_rebind_after_cleanup( root: &Path, original: &Receipt, @@ -28,6 +36,34 @@ pub(super) fn archive_dependency_rebind_after_cleanup( ) -> Result<(), CandidateError> { rebind::archive_after_cleanup(root, original, cleaned, boot) } + +#[cfg(target_os = "macos")] +pub(super) fn dependency_rebind_boot( + root: &Path, + receipt: &Receipt, +) -> Result, CandidateError> { + rebind::boot(root, receipt) +} + +#[cfg(target_os = "macos")] +pub(super) fn archive_retired_dependency_rebind( + root: &Path, + original: &Receipt, + cleaned: &Receipt, + boot: &str, + verify: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + rebind::archive::retired_completed(root, original, cleaned, boot, verify) +} + +#[cfg(target_os = "macos")] +pub(super) fn retired_dependency_rebind_archive_complete( + root: &Path, + original: &Receipt, + boot: &str, +) -> Result { + rebind::archive::retired_archive_complete(root, original, boot) +} use sha2::{Digest, Sha256}; use std::{ io::Read, @@ -659,7 +695,22 @@ impl Driver for HostRelayRuntime { state::write(&root.join("state.json"), receipt)?; return Err(stage_refused("graph_startup_application_failed")); } - verify_exited_listener(receipt.readiness.get(name), &observed)?; + if let Err(error) = verify_exited_listener(receipt.readiness.get(name), &observed) { + if error.code == "graph_startup_listener_unexpected_exit" + && observation == (Observation::Exited { code: 0 }) + { + // A successful process exit can still lose a required + // listener. Keep its identity before cleanup removes it. + return Err(startup_failure::preserve_listener_error( + root, + receipt, + name, + observation, + error, + )); + } + return Err(error); + } } } Ok(()) @@ -1013,6 +1064,12 @@ mod route_tests { .code, "graph_startup_listener_unexpected_exit" ); + assert_eq!( + verify_exited_listener(Some(&Condition::Healthy), &exited) + .unwrap_err() + .code, + "graph_startup_listener_unexpected_exit" + ); let running = json!({"State":{"Running":true,"ExitCode":0,"OOMKilled":false}}); assert_eq!( verify_exited_listener(Some(&Condition::Completed), &running) diff --git a/packages/runtime-core/src/provider/graph/startup/runtime/rebind.rs b/packages/runtime-core/src/provider/graph/startup/runtime/rebind.rs index fa0da765c..c50bf2960 100644 --- a/packages/runtime-core/src/provider/graph/startup/runtime/rebind.rs +++ b/packages/runtime-core/src/provider/graph/startup/runtime/rebind.rs @@ -6,7 +6,7 @@ use serde::{Deserialize, Serialize}; use std::{collections::BTreeSet, fs, path::PathBuf}; const JOURNAL: &str = "dependency-rebind.json"; -mod archive; +pub(super) mod archive; #[cfg(test)] mod nested_tests; pub(super) use archive::after_cleanup as archive_after_cleanup; @@ -331,6 +331,30 @@ fn existing(root: &Path, receipt: &Receipt) -> Result, Can } } } +#[cfg(target_os = "macos")] +pub(super) fn boot(root: &Path, receipt: &Receipt) -> Result, CandidateError> { + Ok(existing(root, receipt)?.map(|journal| journal.boot)) +} +/// A completed refresh is evidence for this exact ready generation only. It +/// grants no host endpoint adoption or cleanup authority by itself. +pub(super) fn require_recovery_complete( + root: &Path, + receipt: &Receipt, + boot: &str, +) -> Result<(), CandidateError> { + require_complete(root, receipt)?; + if let Some(journal) = existing(root, receipt)? { + if receipt.phase != "ready-observed" + || boot.is_empty() + || journal.boot != boot + || journal.completed_generation.as_deref() + != Some(service_exec_generation(receipt)?.as_str()) + { + return Err(stage_refused("graph_dependency_rebind_incomplete")); + } + } + Ok(()) +} pub(super) fn require_complete(root: &Path, receipt: &Receipt) -> Result<(), CandidateError> { if pending_journal(root)? { return Err(stage_refused("graph_dependency_rebind_incomplete")); diff --git a/packages/runtime-core/src/provider/graph/startup/runtime/rebind/archive.rs b/packages/runtime-core/src/provider/graph/startup/runtime/rebind/archive.rs index 374cc0aaf..d2b6ef879 100644 --- a/packages/runtime-core/src/provider/graph/startup/runtime/rebind/archive.rs +++ b/packages/runtime-core/src/provider/graph/startup/runtime/rebind/archive.rs @@ -68,6 +68,71 @@ fn sync(root: &Path) -> Result<(), CandidateError> { .map_err(state::io) } +/// A fully moved archive is inert metadata. Validate its completed journal and +/// every retained artifact without needing the evictable old stopped receipt. +/// The caller independently verifies current cleanup authority and holds its lease. +#[cfg(any(target_os = "macos", test))] +pub(in crate::provider::graph::startup::runtime) fn retired_archive_complete( + root: &Path, + original: &Receipt, + boot: &str, +) -> Result { + state::check_private_directory(root)?; + for name in [JOURNAL, "dependency-rebind.pending"] { + if bytes(&root.join(name))?.is_some() { + return Ok(false); + } + } + let generation = service_exec_generation(original)?; + let history = root.join(format!("dependency-rebind-history-{generation}")); + crate::reject_aliased_state(&history)?; + if !history.try_exists().map_err(state::io)? { + return Ok(false); + } + state::check_private_directory(&history)?; + if bytes(&history.join("proof.pending"))?.is_some() { + return Ok(false); + } + let Some(value) = bytes(&history.join("proof.json"))? else { + return Ok(false); + }; + let proof: Proof = serde_json::from_slice(&value).map_err(|_| rejected())?; + if !proof.complete { + return Ok(false); + } + let names = [JOURNAL, "dependency-rebind.pending"]; + if boot.is_empty() + || proof.version != 1 + || proof.run != original.run + || proof.owner != original.owner + || proof.boot != boot + || proof.original_generation != generation + || !hex(&proof.cleaned_generation, 64) + || !proof.artifacts.contains_key(JOURNAL) + || proof.artifacts.len() > names.len() + || proof + .artifacts + .iter() + .any(|(name, hash)| !names.contains(&name.as_str()) || !hex(hash, 64)) + { + return Err(rejected()); + } + let journal = existing(&history, original)?.ok_or_else(rejected)?; + if journal.phase != "completed" + || journal.boot != boot + || journal.completed_generation.as_deref() != Some(generation.as_str()) + { + return Err(rejected()); + } + for (name, expected) in &proof.artifacts { + let value = bytes(&history.join(name))?.ok_or_else(rejected)?; + if digest(&value) != *expected { + return Err(rejected()); + } + } + Ok(true) +} + /// Caller retains the provider/dead-owner lease and has verified exactly owned /// container/helper absence. Neither a phase label nor this proof removes a VM /// resource or authorizes future endpoint selection. @@ -77,6 +142,49 @@ pub(in crate::provider::graph::startup::runtime) fn after_cleanup( cleaned: &Receipt, boot: &str, ) -> Result<(), CandidateError> { + after_cleanup_fenced(root, original, cleaned, boot, &|| Ok(())) +} + +/// A completed prior-boot journal must match the selected original ready +/// generation before first archival admission. A durable proof alone permits +/// exact rename recovery after the active journal has already moved. +#[cfg(any(target_os = "macos", test))] +pub(in crate::provider::graph::startup::runtime) fn retired_completed( + root: &Path, + original: &Receipt, + cleaned: &Receipt, + boot: &str, + verify: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + verify()?; + let generation = service_exec_generation(original)?; + let proof = root.join(format!("dependency-rebind-history-{generation}/proof.json")); + let journal_root = if bytes(&root.join(JOURNAL))?.is_some() { + root.to_owned() + } else { + proof.parent().ok_or_else(rejected)?.to_owned() + }; + let journal = existing(&journal_root, original)?.ok_or_else(rejected)?; + if journal.phase != "completed" + || journal.boot != boot + || journal.completed_generation.as_deref() != Some(generation.as_str()) + { + return Err(rejected()); + } + if bytes(&proof)?.is_none() { + require_complete(root, original)?; + } + after_cleanup_fenced(root, original, cleaned, boot, verify) +} + +fn after_cleanup_fenced( + root: &Path, + original: &Receipt, + cleaned: &Receipt, + boot: &str, + verify: &dyn Fn() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + verify()?; state::check_private_directory(root)?; if boot.is_empty() || original.run != cleaned.run @@ -118,6 +226,7 @@ pub(in crate::provider::graph::startup::runtime) fn after_cleanup( state::private_directory(&history)?; // A partial proof write is preserved before the same exact owned cleanup // reconstructs it. It is never interpreted as a capability or state receipt. + verify()?; crate::provider::graph::journal::retain_file( &history, "proof.pending", @@ -160,6 +269,7 @@ pub(in crate::provider::graph::startup::runtime) fn after_cleanup( artifacts: active, complete: false, }; + verify()?; state::write(&proof_path, &proof)?; proof }; @@ -179,6 +289,10 @@ pub(in crate::provider::graph::startup::runtime) fn after_cleanup( if digest(&value) != proof.artifacts[name] { return Err(rejected()); } + verify()?; + if bytes(&root.join(name))?.as_deref() != Some(value.as_slice()) { + return Err(rejected()); + } fs::rename(root.join(name), history.join(name)).map_err(state::io)?; sync(&history)?; sync(root)?; @@ -186,7 +300,8 @@ pub(in crate::provider::graph::startup::runtime) fn after_cleanup( } if !proof.complete { proof.complete = true; + verify()?; state::write(&proof_path, &proof)?; } - Ok(()) + verify() } diff --git a/packages/runtime-core/src/provider/graph/startup/runtime/rebind/tests.rs b/packages/runtime-core/src/provider/graph/startup/runtime/rebind/tests.rs index e50efa991..96b6c2f1d 100644 --- a/packages/runtime-core/src/provider/graph/startup/runtime/rebind/tests.rs +++ b/packages/runtime-core/src/provider/graph/startup/runtime/rebind/tests.rs @@ -461,6 +461,24 @@ fn incomplete_journal_blocks_ready_operations_and_is_never_replayed() { journal.completed_generation = Some(service_exec_generation(&receipt).unwrap()); state::write(&root.join(JOURNAL), &journal).unwrap(); require_complete(&root, &receipt).unwrap(); + let evidence = fs::read(root.join(JOURNAL)).unwrap(); + require_recovery_complete(&root, &receipt, "fixture-boot").unwrap(); + for boot in ["", "other-boot"] { + assert!(require_recovery_complete(&root, &receipt, boot).is_err()); + assert_eq!(fs::read(root.join(JOURNAL)).unwrap(), evidence); + } + // The ordinary readiness validator can accept a newer receipt, but recovery + // must select exactly the generation that committed this refresh. + let mut newer = receipt.clone(); + newer.plan_id = "9".repeat(64); + require_complete(&root, &newer).unwrap(); + assert!(require_recovery_complete(&root, &newer, "fixture-boot").is_err()); + let mut incomplete = serde_json::to_value(&journal).unwrap(); + incomplete["phase"] = json!("provisioning"); + state::write(&root.join(JOURNAL), &incomplete).unwrap(); + assert!(require_recovery_complete(&root, &receipt, "fixture-boot").is_err()); + state::write(&root.join(JOURNAL), &journal).unwrap(); + assert_eq!(fs::read(root.join(JOURNAL)).unwrap(), evidence); let mut changed = receipt.clone(); changed .relay_startup @@ -672,6 +690,124 @@ fn terminal_refresh_proofs_do_not_pin_old_helpers_across_owned_restore() { fs::remove_dir_all(root).unwrap(); } +#[test] +fn retired_completed_rebind_requires_exact_generation_and_resumes_from_proof() { + let fixture = crate::provider::graph::tests::Fixture::new(); + let root = &fixture.0; + let mut original = receipt(); + original + .readiness + .insert("init".into(), Condition::Completed); + original.relay_startup = Some(startup(&"6".repeat(64), &["web", "init"])); + let init = original + .relay_startup + .as_mut() + .unwrap() + .services + .get_mut("init") + .unwrap(); + init.phase = Phase::Completed; + init.bindings.get_mut("content").unwrap().process = None; + let process = original.relay_startup.as_ref().unwrap().services["web"].bindings["content"] + .process + .unwrap(); + original.resources.insert( + "container:web".into(), + serde_json::from_value(json!({ + "kind":"container","key":"web","name":"owned-web","id":"7".repeat(64), + "image":"sha256:".to_string()+&"8".repeat(64),"phase":"present" + })) + .unwrap(), + ); + let mut cleaned = original.clone(); + cleaned.phase = "stopped-data-retained".into(); + cleaned.resources.get_mut("container:web").unwrap().phase = "absent".into(); + let generation = service_exec_generation(&original).unwrap(); + let mut journal = RebindJournal { + version: 1, + operation: "4".repeat(32), + run: original.run.clone(), + owner: original.owner.clone(), + boot: "old-boot".into(), + expected_generation: "5".repeat(64), + phase: "completed".into(), + slots: BTreeMap::from([( + 0, + JournalSlot { + before: "5".repeat(64), + after: "6".repeat(64), + bindings: vec![ + ("web".into(), "content".into()), + ("init".into(), "content".into()), + ], + terminal_only: false, + }, + )]), + processes: BTreeMap::from([("web".into(), BTreeMap::from([("content".into(), process)]))]), + completed_services: BTreeSet::from(["init".into()]), + completed_generation: Some(generation.clone()), + }; + state::write(&root.join(JOURNAL), &journal).unwrap(); + let original_bytes = fs::read(root.join(JOURNAL)).unwrap(); + for wrong_boot in ["current-boot", ""] { + assert!( + archive::retired_completed(root, &original, &cleaned, wrong_boot, &|| Ok(())).is_err() + ); + assert_eq!(fs::read(root.join(JOURNAL)).unwrap(), original_bytes); + } + journal.completed_generation = Some("9".repeat(64)); + state::write(&root.join(JOURNAL), &journal).unwrap(); + assert!(archive::retired_completed(root, &original, &cleaned, "old-boot", &|| Ok(())).is_err()); + journal.completed_generation = Some(generation.clone()); + state::write(&root.join(JOURNAL), &journal).unwrap(); + let history = root.join(format!("dependency-rebind-history-{generation}")); + // Lost foreground ownership before the first write leaves evidence intact. + assert!( + archive::retired_completed(root, &original, &cleaned, "old-boot", &|| Err(rejected())) + .is_err() + ); + assert!(!history.exists()); + let verify = || { + if history.join("proof.json").exists() && root.join(JOURNAL).exists() { + Err(rejected()) + } else { + Ok(()) + } + }; + assert!(archive::retired_completed(root, &original, &cleaned, "old-boot", &verify).is_err()); + assert_eq!(fs::read(root.join(JOURNAL)).unwrap(), original_bytes); + assert_eq!( + state::read::(&history.join("proof.json")).unwrap()["complete"], + false + ); + // Simulate the owned archive's first successful rename followed by a crash. + fs::rename(root.join(JOURNAL), history.join(JOURNAL)).unwrap(); + assert!(require_complete(root, &original).is_err()); + crate::provider::graph::startup::archive_retired_dependency_rebind( + root, + &original, + &cleaned, + "old-boot", + &|| Ok(()), + ) + .unwrap(); + archive::retired_completed(root, &original, &cleaned, "old-boot", &|| Ok(())).unwrap(); + assert_eq!(fs::read(history.join(JOURNAL)).unwrap(), original_bytes); + assert_eq!( + state::read::(&history.join("proof.json")).unwrap()["complete"], + true + ); + require_complete(root, &cleaned).unwrap(); + let mut changed = cleaned.clone(); + changed.resources.get_mut("container:web").unwrap().id = Some("a".repeat(64)); + assert!(archive::retired_completed(root, &original, &changed, "old-boot", &|| Ok(())).is_err()); + let mut proof: Value = state::read(&history.join("proof.json")).unwrap(); + proof["boot"] = json!("substituted-boot"); + state::write(&history.join("proof.json"), &proof).unwrap(); + assert!(archive::retired_completed(root, &original, &cleaned, "old-boot", &|| Ok(())).is_err()); + assert_eq!(fs::read(history.join(JOURNAL)).unwrap(), original_bytes); +} + #[test] fn completed_journal_requires_explicit_terminal_state_and_never_fabricates_helpers() { let root = std::env::temp_dir().canonicalize().unwrap().join(format!( diff --git a/packages/runtime-core/src/provider/graph/startup_failure.rs b/packages/runtime-core/src/provider/graph/startup_failure.rs index df30185e3..e8b853990 100644 --- a/packages/runtime-core/src/provider/graph/startup_failure.rs +++ b/packages/runtime-core/src/provider/graph/startup_failure.rs @@ -1,7 +1,12 @@ //! Value-free historical failure evidence, independent of container retention. use super::{Kind, Receipt, error, hex}; -use crate::{CandidateError, project::execution::Observation}; +use crate::{ + CandidateError, + project::execution::{Condition, Observation}, +}; use serde::{Deserialize, Serialize}; +#[cfg(any(target_os = "macos", test))] +use std::path::Path; #[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] #[serde(deny_unknown_fields)] @@ -19,7 +24,12 @@ impl Failure { .bytes() .all(|b| b.is_ascii_alphanumeric() || b"_-".contains(&b)) && hex(&self.container, 64) - && self.observation.failed() + && (self.observation.failed() + || (self.observation == Observation::Exited { code: 0 } + && matches!( + receipt.readiness.get(&self.service), + Some(Condition::Started | Condition::Healthy) + ))) && receipt .resources .get(&format!("container:{}", self.service)) @@ -46,6 +56,21 @@ pub(super) fn record( receipt.startup_failure = Some(failure); Ok(()) } +#[cfg(any(target_os = "macos", test))] +pub(super) fn preserve_listener_error( + root: &Path, + receipt: &mut Receipt, + service: &str, + observation: Observation, + error: CandidateError, +) -> CandidateError { + match record(receipt, service, observation) + .and_then(|()| super::state::write(&root.join("state.json"), receipt)) + { + Ok(()) => error, + Err(diagnostic) => error.with_cause_code(diagnostic.code.into()), + } +} fn refused() -> CandidateError { error( "graph_receipt", @@ -85,12 +110,66 @@ mod tests { #[test] fn only_failed_owned_service_observations_are_recorded() { let mut receipt = receipt(); - for state in [Observation::Created, Observation::Exited { code: 0 }] { - assert!(record(&mut receipt, "redis", state).is_err()); - } + assert!(record(&mut receipt, "redis", Observation::Created).is_err()); assert!(record(&mut receipt, "foreign", Observation::Dead).is_err()); receipt.resources.get_mut("container:redis").unwrap().id = Some("invalid".into()); assert!(record(&mut receipt, "redis", Observation::Dead).is_err()); assert!(receipt.startup_failure.is_none()); } + #[test] + fn required_listener_zero_exit_survives_cleanup() { + for condition in [Condition::Started, Condition::Healthy] { + let mut receipt = receipt(); + receipt.readiness.insert("redis".into(), condition); + record(&mut receipt, "redis", Observation::Exited { code: 0 }).unwrap(); + receipt.phase = "stopped-data-retained".into(); + receipt.resources.get_mut("container:redis").unwrap().id = None; + let root = super::super::tests::Fixture::new(); + crate::provider::state::write(&root.0.join("state.json"), &receipt).unwrap(); + let (retained, _) = + super::super::load_at(root.0.clone(), &receipt.run, &receipt.owner).unwrap(); + let failure = retained.startup_failure.as_ref().unwrap(); + assert_eq!(failure.service, "redis"); + assert_eq!(failure.container, "e".repeat(64)); + assert_eq!(failure.observation, Observation::Exited { code: 0 }); + } + } + #[test] + fn successful_completion_is_not_startup_failure_evidence() { + let mut receipt = receipt(); + receipt + .readiness + .insert("redis".into(), Condition::Completed); + assert!(record(&mut receipt, "redis", Observation::Exited { code: 0 }).is_err()); + receipt.readiness.clear(); + assert!(record(&mut receipt, "redis", Observation::Exited { code: 0 }).is_err()); + assert!(receipt.startup_failure.is_none()); + } + #[test] + fn failed_evidence_write_preserves_original_error_and_pending_bytes() { + let root = super::super::tests::Fixture::new(); + let pending = root.0.join("state.pending"); + std::fs::write(&pending, b"retained pending evidence").unwrap(); + let mut receipt = receipt(); + let error = preserve_listener_error( + &root.0, + &mut receipt, + "redis", + Observation::Exited { code: 0 }, + CandidateError::new("graph_startup_listener_unexpected_exit", "Startup failed."), + ); + assert_eq!(error.code, "graph_startup_listener_unexpected_exit"); + assert_eq!(error.cause_code.as_deref(), Some("provider_state")); + assert_eq!(error.message, "Startup failed."); + assert_eq!( + std::fs::read(&pending).unwrap(), + b"retained pending evidence" + ); + assert!(!root.0.join("state.json").exists()); + assert!( + !serde_json::to_string(&error) + .unwrap() + .contains("state.pending") + ); + } } diff --git a/packages/runtime-core/src/provider/graph/tests.rs b/packages/runtime-core/src/provider/graph/tests.rs index 3e2b4a6f1..aa61c788c 100644 --- a/packages/runtime-core/src/provider/graph/tests.rs +++ b/packages/runtime-core/src/provider/graph/tests.rs @@ -1812,6 +1812,7 @@ fn dependency_cache_graphs_share_only_verified_binding_and_pin_restore_targets() current_manifest: None, binding: SourceBinding { shared: None, + cache_scope: None, shared_contract: None, live: None, revision: manifest.revision.clone(), @@ -2181,6 +2182,7 @@ fn shared_source_preserves_app_writes_but_freezes_cache_installer_inputs() { current_manifest: None, binding: SourceBinding { shared: Some(share), + cache_scope: None, shared_contract: Some( project::live_source::Contract::from_plan(&review.plan, &manifest).unwrap(), ), diff --git a/packages/runtime-core/src/provider/host_pin.rs b/packages/runtime-core/src/provider/host_pin.rs new file mode 100644 index 000000000..f583198d4 --- /dev/null +++ b/packages/runtime-core/src/provider/host_pin.rs @@ -0,0 +1,40 @@ +//! Exact device-number translation for explicitly selected pre-reboot host pins. +//! This is never used by ordinary ownership, source replay, or guest observations. +use super::identity::{self, ProcessIdentity}; +use crate::CandidateError; +use serde::{Deserialize, Serialize}; + +#[derive(Clone, Copy, Debug, Deserialize, Eq, PartialEq, Serialize)] +#[serde(deny_unknown_fields)] +pub(crate) struct DeviceRebind { + pub old: u64, + pub current: u64, +} + +impl DeviceRebind { + pub fn matches(self, recorded: (u64, u64), observed: (u64, u64)) -> bool { + self.old != self.current + && recorded.0 == self.old + && observed.0 == self.current + && recorded.1 != 0 + && recorded.1 == observed.1 + } + + pub fn definitely_dead_before_boot( + self, + process: &ProcessIdentity, + host_boot_micros: u64, + ) -> Result<(), CandidateError> { + if host_boot_micros == 0 + || process.start_micros == 0 + || process.start_micros >= host_boot_micros + || identity::alive(process.pid)? + { + return Err(CandidateError::new( + "host_pin_recovery", + "Legacy host pin owner is not proved dead before this host boot.", + )); + } + Ok(()) + } +} diff --git a/packages/runtime-core/src/provider/https_recovery.rs b/packages/runtime-core/src/provider/https_recovery.rs new file mode 100644 index 000000000..370bfc2c4 --- /dev/null +++ b/packages/runtime-core/src/provider/https_recovery.rs @@ -0,0 +1,701 @@ +//! Explicit archival of a quiescent legacy HTTPS owner and an unpublished shared owner. +//! No process is signaled, CA data is untouched, and every moved inode remains in history. +use super::{hostname_authority, identity, state}; +use crate::{Candidate, CandidateError}; +use base64::Engine as _; +mod ports; +pub(in crate::provider) use ports::port_absent; +use serde::{Deserialize, Serialize}; +use sha2::{Digest, Sha256}; +#[cfg(target_os = "macos")] +use std::ffi::CString; +use std::{ + fs::{self, OpenOptions}, + io::{Read, Write}, + os::unix::{ + fs::{FileTypeExt, MetadataExt, OpenOptionsExt}, + net::UnixStream, + }, + path::{Path, PathBuf}, +}; + +fn refused() -> CandidateError { + CandidateError::new( + "https_recovery_refused", + "HTTPS recovery could not prove quiescence and exact evidence; retained state was preserved.", + ) +} +fn digest(bytes: &[u8]) -> String { + format!("{:x}", Sha256::digest(bytes)) +} +fn valid_hash(value: &str) -> bool { + value.len() == 64 + && value + .bytes() + .all(|b| b.is_ascii_digit() || (b'a'..=b'f').contains(&b)) +} +#[derive(Clone, Copy, Debug, Serialize, Deserialize, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +struct Inode { + dev: u64, + ino: u64, +} +fn inode(path: &Path) -> Result { + let m = fs::symlink_metadata(path).map_err(|_| refused())?; + if m.uid() != unsafe { libc::geteuid() } || m.file_type().is_symlink() { + return Err(refused()); + } + Ok(Inode { + dev: m.dev(), + ino: m.ino(), + }) +} +fn absent(path: &Path) -> Result { + match fs::symlink_metadata(path) { + Ok(_) => Ok(false), + Err(e) if e.kind() == std::io::ErrorKind::NotFound => Ok(true), + Err(_) => Err(refused()), + } +} +fn private_directory(path: &Path) -> Result { + let id = inode(path)?; + let m = fs::symlink_metadata(path).map_err(|_| refused())?; + if !m.is_dir() + || m.mode() & 0o777 != 0o700 + || fs::canonicalize(path).map_err(|_| refused())? != path + { + return Err(refused()); + } + Ok(id) +} +fn bounded(path: &Path, limit: u64) -> Result, CandidateError> { + read_file(path, limit, true) +} +fn certificate(path: &Path) -> Result, CandidateError> { + read_file(path, 65_536, false) +} +fn read_file(path: &Path, limit: u64, private: bool) -> Result, CandidateError> { + let mut f = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK) + .open(path) + .map_err(|_| refused())?; + let m = f.metadata().map_err(|_| refused())?; + if !m.is_file() + || m.nlink() != 1 + || (private && m.mode() & 0o777 != 0o600) + || (!private && m.mode() & 0o022 != 0) + || m.uid() != unsafe { libc::geteuid() } + || m.len() == 0 + || m.len() > limit + { + return Err(refused()); + } + let mut bytes = Vec::new(); + (&mut f) + .take(limit + 1) + .read_to_end(&mut bytes) + .map_err(|_| refused())?; + let after = f.metadata().map_err(|_| refused())?; + if bytes.len() as u64 != m.len() + || m.mtime() != after.mtime() + || m.ctime() != after.ctime() + || m.mtime_nsec() != after.mtime_nsec() + || m.ctime_nsec() != after.ctime_nsec() + || fs::canonicalize(path).map_err(|_| refused())? != path + || inode(path)? + != (Inode { + dev: m.dev(), + ino: m.ino(), + }) + { + return Err(refused()); + } + Ok(bytes) +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Receipt { + version: u8, + listener: Listener, + caddy_sha256: String, + ca_sha256: String, + authority: Authority, + lock: Inode, + owner: Challenge, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields)] +struct Listener { + pid: i32, + start_micros: u64, + uid: u32, + executable: PathBuf, + port: u16, + fingerprint: String, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields)] +struct Authority { + pid: i32, + socket: PathBuf, + sha256: String, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Challenge { + public_key: String, + dev: u64, + ino: u64, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Configuration { + version: u8, + owner_generation: String, + binding: Binding, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Binding { + runtime: Runtime, + frontend: Binary, + runtime_sha256: String, + pool: Pool, + caddy_binary: PathBuf, + caddy_sha256: String, + https_port: u16, + certificate_name_limit: u16, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields)] +struct Runtime { + binary: PathBuf, + home: PathBuf, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields)] +struct Binary { + binary: PathBuf, + sha256: String, +} +#[derive(Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Pool { + owner: String, + boot_id: String, +} +fn valid_hex(value: &str, limit: usize) -> bool { + value.len() == limit + && value + .bytes() + .all(|b| b.is_ascii_digit() || (b'a'..=b'f').contains(&b)) +} +#[derive(Serialize, Deserialize, Clone, Debug, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +struct Entry { + name: String, + id: Inode, +} +#[derive(Serialize, Deserialize, Clone, Debug, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +struct Journal { + version: u8, + home: PathBuf, + parent: Inode, + owner_sha256: String, + configuration_sha256: String, + frontend_pid: i32, + legacy_device_rebind: Option, + ca: Inode, + ca_file_sha256: String, + configuration: Inode, + leases: Inode, + entries: Vec, +} + +#[derive(Clone, Debug, Serialize, Deserialize, PartialEq, Eq)] +#[serde(deny_unknown_fields)] +struct LegacyDeviceRebind { + run: String, + witness_sha256: String, + socket: Inode, + lock: Inode, +} +impl LegacyDeviceRebind { + fn parse(value: &str) -> Result { + let parts: Vec<_> = value.split(':').collect(); + if parts.len() != 6 || !valid_hex(parts[0], 32) || !valid_hash(parts[1]) { + return Err(refused()); + } + let number = |part: &str| { + part.parse::() + .ok() + .filter(|n| *n > 0) + .ok_or_else(refused) + }; + Ok(Self { + run: parts[0].into(), + witness_sha256: parts[1].into(), + socket: Inode { + dev: number(parts[2])?, + ino: number(parts[3])?, + }, + lock: Inode { + dev: number(parts[4])?, + ino: number(parts[5])?, + }, + }) + } + fn matches(&self, receipt: &Receipt, old: u64, current: u64) -> bool { + old != current + && receipt.owner.dev == old + && receipt.lock.dev == old + && self.socket.dev == current + && self.lock.dev == current + && self.socket.ino == receipt.owner.ino + && self.lock.ino == receipt.lock.ino + } +} +fn selected(root: &Path, entry: &Entry, suffix: &str) -> Result { + let source = root.join(&entry.name); + let archived = destination(root, &entry.name, suffix); + match (absent(&source)?, absent(&archived)?) { + (false, true) if inode(&source)? == entry.id => Ok(source), + (true, false) if inode(&archived)? == entry.id => Ok(archived), + _ => Err(refused()), + } +} +fn destination(root: &Path, name: &str, suffix: &str) -> PathBuf { + // The archived socket must still be probeable within sockaddr_un's path bound. + // A short-name collision refuses; the full journal and exact inode remain authority. + if name == "owner.sock" { + root.join(format!("s{}", &suffix[..8])) + } else { + root.join(format!("{name}.retired-{suffix}")) + } +} +fn names(path: &Path) -> Result, CandidateError> { + let mut names = fs::read_dir(path) + .map_err(|_| refused())? + .map(|e| { + e.map_err(|_| refused()) + .and_then(|e| e.file_name().into_string().map_err(|_| refused())) + }) + .collect::, _>>()?; + names.sort(); + Ok(names) +} +fn observe( + root: &Path, + suffix: &str, + journal: &Journal, +) -> Result<(Receipt, Configuration, PathBuf), CandidateError> { + if private_directory(root)? != journal.parent + || journal + .entries + .iter() + .map(|e| e.name.as_str()) + .collect::>() + != [ + "owner.sock", + "active-owner.json", + "shared-owner", + "owner.lock", + ] + { + return Err(refused()); + } + let paths = journal + .entries + .iter() + .map(|e| selected(root, e, suffix)) + .collect::, _>>()?; + let bytes = bounded(&paths[1], 4096)?; + let receipt: Receipt = serde_json::from_slice(&bytes).map_err(|_| refused())?; + if digest(&bytes) != journal.owner_sha256 + || receipt.version != 1 + || receipt.listener.pid <= 1 + || receipt.authority.pid <= 1 + || receipt.listener.start_micros == 0 + || receipt.listener.port == 0 + || receipt.listener.uid != unsafe { libc::geteuid() } + || !receipt.listener.executable.is_absolute() + || !receipt.authority.socket.is_absolute() + || !valid_hash(&receipt.listener.fingerprint) + || !valid_hash(&receipt.caddy_sha256) + || !valid_hash(&receipt.ca_sha256) + || !valid_hash(&receipt.authority.sha256) + || receipt.owner.public_key.is_empty() + { + return Err(refused()); + } + let socket = fs::symlink_metadata(&paths[0]).map_err(|_| refused())?; + let socket_id = journal.entries[0].id; + let lock_id = private_directory(&paths[3])?; + let exact = socket_id + == (Inode { + dev: receipt.owner.dev, + ino: receipt.owner.ino, + }) + && lock_id == receipt.lock; + let explicitly_selected = journal.legacy_device_rebind.as_ref().is_some_and(|s| { + s.socket == socket_id + && s.lock == lock_id + && s.matches(&receipt, receipt.owner.dev, socket_id.dev) + }); + if !socket.file_type().is_socket() + || socket.mode() & 0o777 != 0o600 + || !(exact && journal.legacy_device_rebind.is_none() || explicitly_selected) + || !names(&paths[3])?.is_empty() + { + return Err(refused()); + } + private_directory(&paths[2])?; + if names(&paths[2])? != ["configuration.json", "leases"] { + return Err(refused()); + } + if private_directory(&paths[2].join("leases"))? != journal.leases + || inode(&paths[2].join("configuration.json"))? != journal.configuration + || inode(&root.join("data/caddy/pki/authorities/local/root.crt"))? != journal.ca + || digest(&certificate( + &root.join("data/caddy/pki/authorities/local/root.crt"), + )?) != journal.ca_file_sha256 + { + return Err(refused()); + } + if !names(&paths[2].join("leases"))?.is_empty() { + return Err(refused()); + } + let config_bytes = bounded(&paths[2].join("configuration.json"), 16384)?; + if digest(&config_bytes) != journal.configuration_sha256 { + return Err(refused()); + } + let config: Configuration = serde_json::from_slice(&config_bytes).map_err(|_| refused())?; + if config.version != 1 + || config.binding.runtime.home != journal.home + || config.binding.https_port != receipt.listener.port + || !valid_hex(&config.owner_generation, 32) + || !valid_hex(&config.binding.pool.owner, 32) + || !valid_hash(&config.binding.runtime_sha256) + || !valid_hash(&config.binding.frontend.sha256) + || !valid_hash(&config.binding.caddy_sha256) + || config.binding.caddy_sha256 != receipt.caddy_sha256 + || config.binding.caddy_binary != receipt.listener.executable + || !(1..=4096).contains(&config.binding.certificate_name_limit) + || [ + &config.binding.runtime.home, + &config.binding.runtime.binary, + &config.binding.frontend.binary, + &config.binding.caddy_binary, + ] + .iter() + .any(|p| !p.is_absolute() || p.as_os_str().len() >= 4096) + { + return Err(refused()); + } + let generation = &config.owner_generation; + if !absent(&root.join("released-leases"))? { + private_directory(&root.join("released-leases"))?; + } + if !absent(&root.join("released-leases").join(generation))? { + return Err(refused()); + } + Ok((receipt, config, paths[0].clone())) +} +#[cfg(target_os = "macos")] +fn rename_exclusive(source: &Path, destination: &Path) -> Result<(), CandidateError> { + let source = CString::new(source.as_os_str().as_encoded_bytes()).map_err(|_| refused())?; + let destination = + CString::new(destination.as_os_str().as_encoded_bytes()).map_err(|_| refused())?; + // SAFETY: both C strings remain valid for this call. RENAME_EXCL never overwrites a destination. + if unsafe { libc::renamex_np(source.as_ptr(), destination.as_ptr(), libc::RENAME_EXCL) } != 0 { + return Err(refused()); + } + Ok(()) +} +#[cfg(not(target_os = "macos"))] +fn rename_exclusive(_source: &Path, _destination: &Path) -> Result<(), CandidateError> { + Err(refused()) +} +fn sync_directory(path: &Path, expected: &Inode) -> Result<(), CandidateError> { + let file = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK | libc::O_DIRECTORY) + .open(path) + .map_err(|_| refused())?; + let m = file.metadata().map_err(|_| refused())?; + if !m.is_dir() + || (Inode { + dev: m.dev(), + ino: m.ino(), + }) != *expected + { + return Err(refused()); + } + file.sync_all().map_err(|_| refused()) +} +fn archive_under(root: &Path, journal: &Journal, verify: F) -> Result<(), CandidateError> +where + F: Fn(&Receipt, &Configuration, &Path) -> Result, +{ + let suffix = digest(&serde_json::to_vec(journal).map_err(|_| refused())?); + for entry in &journal.entries { + let (receipt, config, socket) = observe(root, &suffix, journal)?; + let _guard = verify(&receipt, &config, &socket)?; + // External observations can take time: recheck every selected entry and parent before effect. + observe(root, &suffix, journal)?; + let source = selected(root, entry, &suffix)?; + if source == root.join(&entry.name) { + rename_exclusive(&source, &destination(root, &entry.name, &suffix))?; + sync_directory(root, &journal.parent)?; + } + } + let (receipt, config, socket) = observe(root, &suffix, journal)?; + verify(&receipt, &config, &socket)?; + Ok(()) +} +/** Compare DER bytes of the previously validated retained PEM; this does not install trust. */ +fn ca_hash(path: &Path) -> Result { + let bytes = certificate(path)?; + let pem = std::str::from_utf8(&bytes).map_err(|_| refused())?.trim(); + let body = pem + .strip_prefix("-----BEGIN CERTIFICATE-----") + .and_then(|p| p.strip_suffix("-----END CERTIFICATE-----")) + .ok_or_else(refused)?; + let encoded: Vec = body + .bytes() + .filter(|b| *b != b'\r' && *b != b'\n') + .collect(); + let der = base64::engine::general_purpose::STANDARD + .decode(encoded) + .map_err(|_| refused())?; + if der.is_empty() { + return Err(refused()); + } + Ok(digest(&der)) +} + +pub(in crate::provider) fn executable_hash(binary: &Path) -> Result { + if fs::canonicalize(binary).map_err(|_| refused())? != binary { + return Err(refused()); + } + let mut file = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK) + .open(binary) + .map_err(|_| refused())?; + let before = file.metadata().map_err(|_| refused())?; + if !before.is_file() + || before.mode() & 0o022 != 0 + || before.mode() & 0o111 == 0 + || before.len() > 512 * 1024 * 1024 + { + return Err(refused()); + } + let mut hasher = Sha256::new(); + let mut buffer = [0; 65536]; + loop { + let count = file.read(&mut buffer).map_err(|_| refused())?; + if count == 0 { + break; + } + hasher.update(&buffer[..count]); + } + let after = file.metadata().map_err(|_| refused())?; + if before.len() != after.len() + || before.mtime() != after.mtime() + || before.mtime_nsec() != after.mtime_nsec() + || before.ctime() != after.ctime() + || before.ctime_nsec() != after.ctime_nsec() + || inode(binary)? + != (Inode { + dev: before.dev(), + ino: before.ino(), + }) + { + return Err(refused()); + } + Ok(format!("{:x}", hasher.finalize())) +} + +fn socket_absent(socket: &Path) -> Result<(), CandidateError> { + match UnixStream::connect(socket) { + Err(e) if e.kind() == std::io::ErrorKind::ConnectionRefused => Ok(()), + _ => Err(refused()), + } +} + +/// Archive only exact dead legacy HTTPS evidence and a selected unpublished shared-owner generation. +/// Repeated calls resume by inode; occupied targets and any live/uncertain effects refuse. +pub fn recover( + candidate: &Candidate, + owner_hash: &str, + config_hash: &str, + frontend_pid: i32, +) -> Result { + recover_selected(candidate, owner_hash, config_hash, frontend_pid, None) +} + +/// Explicit legacy migration only; original host-volume continuity remains unproven. +pub fn recover_legacy_device_rebind( + candidate: &Candidate, + owner_hash: &str, + config_hash: &str, + frontend_pid: i32, + selection: &str, +) -> Result { + recover_selected( + candidate, + owner_hash, + config_hash, + frontend_pid, + Some(LegacyDeviceRebind::parse(selection)?), + ) +} + +fn recover_selected( + candidate: &Candidate, + owner_hash: &str, + config_hash: &str, + frontend_pid: i32, + legacy_device_rebind: Option, +) -> Result { + if !cfg!(target_os = "macos") + || !valid_hash(owner_hash) + || !valid_hash(config_hash) + || frontend_pid <= 1 + { + return Err(refused()); + } + let _lock = state::Lock::acquire_existing(&candidate.state_root.join("run/smolvm"))?; + let home_pin = private_directory(&candidate.checkout)?; + let operation_pin = _lock.identity()?; + let root = candidate.checkout.join("native-https"); + let parent = private_directory(&root)?; + let journal_path = root.join(format!("recovery-{owner_hash}-{config_hash}.json")); + let journal = if absent(&journal_path)? { + Journal { + version: 1, + home: candidate.checkout.clone(), + parent, + owner_sha256: owner_hash.into(), + configuration_sha256: config_hash.into(), + frontend_pid, + legacy_device_rebind: legacy_device_rebind.clone(), + ca: inode(&root.join("data/caddy/pki/authorities/local/root.crt"))?, + ca_file_sha256: digest(&certificate( + &root.join("data/caddy/pki/authorities/local/root.crt"), + )?), + configuration: inode(&root.join("shared-owner/configuration.json"))?, + leases: private_directory(&root.join("shared-owner/leases"))?, + entries: [ + "owner.sock", + "active-owner.json", + "shared-owner", + "owner.lock", + ] + .iter() + .map(|name| { + Ok(Entry { + name: (*name).into(), + id: inode(&root.join(name))?, + }) + }) + .collect::>()?, + } + } else { + serde_json::from_slice::(&bounded(&journal_path, 16384)?).map_err(|_| refused())? + }; + if journal.version != 1 + || journal.home != candidate.checkout + || journal.parent != parent + || journal.owner_sha256 != owner_hash + || journal.configuration_sha256 != config_hash + || journal.frontend_pid != frontend_pid + || journal.legacy_device_rebind != legacy_device_rebind + { + return Err(refused()); + } + let verify = |receipt: &Receipt, config: &Configuration, socket: &Path| { + let operation = inode(&candidate.state_root.join("run/smolvm/operation.lock"))?; + if private_directory(&candidate.checkout)? != home_pin + || (operation.dev, operation.ino) != operation_pin + { + return Err(refused()); + } + if identity::alive(frontend_pid)? + || identity::alive(receipt.listener.pid)? + || identity::alive(receipt.authority.pid)? + { + return Err(refused()); + } + for (binary, hash) in [ + ( + &config.binding.frontend.binary, + &config.binding.frontend.sha256, + ), + ( + &config.binding.runtime.binary, + &config.binding.runtime_sha256, + ), + ] { + if identity::executable_running(binary)? || executable_hash(binary)? != *hash { + return Err(refused()); + } + } + let pool = state::Owner::load(candidate)?; + let authority = hostname_authority::managed::inspect(candidate)?; + if config.binding.pool.owner != pool.token + || pool.guest_boot_id.as_deref() != Some(config.binding.pool.boot_id.as_str()) + || authority["authority"]["present"] != false + || authority["socket"].as_str() != receipt.authority.socket.to_str() + { + return Err(refused()); + } + #[cfg(target_os = "macos")] + if let Some(selection) = &journal.legacy_device_rebind { + let (old, current) = + super::graph::https_devices(candidate, &selection.run, &selection.witness_sha256)?; + if !selection.matches(receipt, old, current) { + return Err(refused()); + } + } + if ca_hash(&root.join("data/caddy/pki/authorities/local/root.crt"))? != receipt.ca_sha256 { + return Err(refused()); + } + socket_absent(socket)?; + // Wildcard probes cover IPv4 and IPv6 listeners, unlike a loopback-only probe. + port_absent(receipt.listener.port) + }; + let suffix = digest(&serde_json::to_vec(&journal).map_err(|_| refused())?); + let (receipt, config, socket) = observe(&root, &suffix, &journal)?; + verify(&receipt, &config, &socket)?; + if absent(&journal_path)? { + let mut file = OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .custom_flags(libc::O_NOFOLLOW) + .open(&journal_path) + .map_err(|_| refused())?; + file.write_all(&serde_json::to_vec(&journal).map_err(|_| refused())?) + .map_err(|_| refused())?; + file.sync_all().map_err(|_| refused())?; + sync_directory(&root, &journal.parent)?; + } + let journal_pin = inode(&journal_path)?; + let journal_bytes = serde_json::to_vec(&journal).map_err(|_| refused())?; + archive_under(&root, &journal, |receipt, config, socket| { + if inode(&journal_path)? != journal_pin || bounded(&journal_path, 16384)? != journal_bytes { + return Err(refused()); + } + verify(receipt, config, socket) + })?; + Ok( + serde_json::json!({"https_evidence_archived":true,"owner_sha256":owner_hash,"configuration_sha256":config_hash,"entries":journal.entries.len(),"ca_preserved":true,"processes_signaled":0,"qualification":if journal.legacy_device_rebind.is_some(){"explicit-legacy-device-migration-original-volume-continuity-unproven"}else{"exact-recorded-device-and-inode"}}), + ) +} + +#[cfg(all(test, target_os = "macos"))] +mod tests; diff --git a/packages/runtime-core/src/provider/https_recovery/ports.rs b/packages/runtime-core/src/provider/https_recovery/ports.rs new file mode 100644 index 000000000..646e833f2 --- /dev/null +++ b/packages/runtime-core/src/provider/https_recovery/ports.rs @@ -0,0 +1,102 @@ +//! MacOS wildcard and loopback claims held across HTTPS recovery effects. +use super::{CandidateError, refused}; +pub(in crate::provider) fn port_absent( + port: u16, +) -> Result, CandidateError> { + let mut guards = Vec::from(bind_pair(port, true)?); + // macOS permits a later SO_REUSEADDR loopback bind beside a wildcard listener. + // Reserve the actual local HTTPS path too; neither guard enables SO_REUSEPORT. + guards.extend(bind_pair(port, false)?); + Ok(guards) +} +fn bind_pair(port: u16, wildcard: bool) -> Result<[std::os::fd::OwnedFd; 2], CandidateError> { + use std::os::fd::{AsRawFd, FromRawFd, OwnedFd}; + let make = |family| { + // SAFETY: socket takes scalar constants; a successful new descriptor is uniquely owned below. + let fd = unsafe { libc::socket(family, libc::SOCK_STREAM, 0) }; + if fd < 0 { + return Err(refused()); + } + let fd = unsafe { OwnedFd::from_raw_fd(fd) }; + // SAFETY: fd is owned; fcntl uses only scalar arguments and does not retain pointers. + if unsafe { libc::fcntl(fd.as_raw_fd(), libc::F_SETFD, libc::FD_CLOEXEC) } < 0 { + return Err(refused()); + } + Ok(fd) + }; + let v4 = make(libc::AF_INET)?; + let v6 = make(libc::AF_INET6)?; + let only: libc::c_int = 1; + // SAFETY: only points to one live integer of the declared size. No address reuse is enabled. + if unsafe { + libc::setsockopt( + v6.as_raw_fd(), + libc::IPPROTO_IPV6, + libc::IPV6_V6ONLY, + (&only as *const libc::c_int).cast(), + std::mem::size_of_val(&only) as libc::socklen_t, + ) + } != 0 + { + return Err(refused()); + } + if !wildcard { + for fd in [&v4, &v6] { + // SAFETY: only is a live integer; this permits our specific loopback guard beside the wildcard guard. + // SO_REUSEPORT is never enabled, so another listener cannot share this exact address. + if unsafe { + libc::setsockopt( + fd.as_raw_fd(), + libc::SOL_SOCKET, + libc::SO_REUSEADDR, + (&only as *const libc::c_int).cast(), + std::mem::size_of_val(&only) as libc::socklen_t, + ) + } != 0 + { + return Err(refused()); + } + } + } + let mut address4: libc::sockaddr_in = unsafe { std::mem::zeroed() }; + address4.sin_family = libc::AF_INET as _; + address4.sin_port = port.to_be(); + let mut address6: libc::sockaddr_in6 = unsafe { std::mem::zeroed() }; + address6.sin6_family = libc::AF_INET6 as _; + address6.sin6_port = port.to_be(); + if !wildcard { + address4.sin_addr.s_addr = u32::from_ne_bytes([127, 0, 0, 1]); + address6.sin6_addr.s6_addr[15] = 1; + } + #[cfg(target_os = "macos")] + { + address4.sin_len = std::mem::size_of_val(&address4) as u8; + address6.sin6_len = std::mem::size_of_val(&address6) as u8; + } + // SAFETY: both zero-initialized wildcard addresses have the matching family and live sized storage. + if unsafe { + libc::bind( + v4.as_raw_fd(), + (&address4 as *const libc::sockaddr_in).cast(), + std::mem::size_of_val(&address4) as libc::socklen_t, + ) + } != 0 + || unsafe { + libc::bind( + v6.as_raw_fd(), + (&address6 as *const libc::sockaddr_in6).cast(), + std::mem::size_of_val(&address6) as libc::socklen_t, + ) + } != 0 + { + return Err(refused()); + } + // SAFETY: both descriptors are owned bound TCP sockets; listen retains the exact address claim. + for fd in [&v4, &v6] { + if unsafe { libc::listen(fd.as_raw_fd(), 1) } != 0 { + return Err(refused()); + } + } + // Retain these listeners across archival; never accept a connection. + Ok([v4, v6]) +} diff --git a/packages/runtime-core/src/provider/https_recovery/tests.rs b/packages/runtime-core/src/provider/https_recovery/tests.rs new file mode 100644 index 000000000..aa58a82ce --- /dev/null +++ b/packages/runtime-core/src/provider/https_recovery/tests.rs @@ -0,0 +1,570 @@ +use super::*; +use std::{ + cell::Cell, + os::unix::{fs::PermissionsExt, net::UnixListener}, + time::{SystemTime, UNIX_EPOCH}, +}; + +struct Fixture { + home: PathBuf, + root: PathBuf, + journal: Journal, +} +impl Drop for Fixture { + fn drop(&mut self) { + let _ = fs::remove_dir_all(&self.home); + } +} +fn write(path: &Path, bytes: &[u8]) { + let mut file = OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .open(path) + .unwrap(); + file.write_all(bytes).unwrap(); +} +fn directory(path: &Path) { + fs::create_dir(path).unwrap(); + fs::set_permissions(path, fs::Permissions::from_mode(0o700)).unwrap(); +} +fn fixture() -> Fixture { + static NEXT: std::sync::atomic::AtomicU64 = std::sync::atomic::AtomicU64::new(0); + let home = PathBuf::from(format!( + "/private/tmp/hr-{}-{}-{}", + std::process::id(), + SystemTime::now() + .duration_since(UNIX_EPOCH) + .unwrap() + .as_nanos(), + NEXT.fetch_add(1, std::sync::atomic::Ordering::Relaxed) + )); + directory(&home); + let root = home.join("native-https"); + directory(&root); + directory(&root.join("owner.lock")); + directory(&root.join("shared-owner")); + directory(&root.join("shared-owner/leases")); + let listener = UnixListener::bind(root.join("owner.sock")).unwrap(); + fs::set_permissions(root.join("owner.sock"), fs::Permissions::from_mode(0o600)).unwrap(); + drop(listener); + let socket = inode(&root.join("owner.sock")).unwrap(); + let lock = inode(&root.join("owner.lock")).unwrap(); + let receipt = serde_json::json!({"version":1,"listener":{"pid":999999,"start_micros":1,"uid":unsafe{libc::geteuid()},"executable":"/unused/caddy","port":18443,"fingerprint":"1".repeat(64)},"caddySha256":"2".repeat(64),"caSha256":"3".repeat(64),"authority":{"pid":999998,"socket":"/private/tmp/unused/route.sock","sha256":"4".repeat(64)},"lock":lock,"owner":{"publicKey":"public-key","dev":socket.dev,"ino":socket.ino}}); + let config = serde_json::json!({"version":1,"ownerGeneration":"5".repeat(32),"binding":{"runtime":{"home":home,"binary":"/unused/native"},"frontend":{"binary":"/unused/frontend","sha256":"6".repeat(64)},"runtimeSha256":"7".repeat(64),"caddyBinary":"/unused/caddy","caddySha256":"2".repeat(64),"certificateNameLimit":256,"httpsPort":18443,"pool":{"owner":"8".repeat(32),"bootId":"12345678-1234-1234-1234-123456789012"}}}); + let receipt = serde_json::to_vec(&receipt).unwrap(); + let config = serde_json::to_vec(&config).unwrap(); + write(&root.join("active-owner.json"), &receipt); + write(&root.join("shared-owner/configuration.json"), &config); + for name in [ + "data", + "data/caddy", + "data/caddy/pki", + "data/caddy/pki/authorities", + "data/caddy/pki/authorities/local", + ] { + directory(&root.join(name)); + } + let ca = root.join("data/caddy/pki/authorities/local/root.crt"); + write(&ca, b"retained certificate bytes"); + let journal = Journal { + version: 1, + home: home.clone(), + parent: inode(&root).unwrap(), + owner_sha256: digest(&receipt), + configuration_sha256: digest(&config), + frontend_pid: 999997, + legacy_device_rebind: None, + ca: inode(&ca).unwrap(), + ca_file_sha256: digest(&fs::read(&ca).unwrap()), + configuration: inode(&root.join("shared-owner/configuration.json")).unwrap(), + leases: inode(&root.join("shared-owner/leases")).unwrap(), + entries: [ + "owner.sock", + "active-owner.json", + "shared-owner", + "owner.lock", + ] + .iter() + .map(|name| Entry { + name: (*name).into(), + id: inode(&root.join(name)).unwrap(), + }) + .collect(), + }; + Fixture { + home, + root, + journal, + } +} +fn suffix(f: &Fixture) -> String { + digest(&serde_json::to_vec(&f.journal).unwrap()) +} +fn verify_socket(_: &Receipt, _: &Configuration, socket: &Path) -> Result<(), CandidateError> { + socket_absent(socket) +} +fn all_original(f: &Fixture) { + for entry in &f.journal.entries { + assert_eq!(inode(&f.root.join(&entry.name)).unwrap(), entry.id); + } +} + +#[test] +fn archives_exact_inodes_and_bytes_without_changing_ca_and_retry_is_idempotent() { + let f = fixture(); + let before = fs::read(f.root.join("active-owner.json")).unwrap(); + archive_under(&f.root, &f.journal, verify_socket).unwrap(); + for entry in &f.journal.entries { + assert!(absent(&f.root.join(&entry.name)).unwrap()); + assert_eq!( + inode(&destination(&f.root, &entry.name, &suffix(&f))).unwrap(), + entry.id + ); + } + assert_eq!( + fs::read( + f.root + .join(format!("active-owner.json.retired-{}", suffix(&f))) + ) + .unwrap(), + before + ); + assert_eq!( + inode(&f.root.join("data/caddy/pki/authorities/local/root.crt")).unwrap(), + f.journal.ca + ); + archive_under(&f.root, &f.journal, verify_socket).unwrap(); +} +#[test] +fn interruption_at_each_move_resumes_without_losing_original_evidence() { + for cut in 0..=4 { + let f = fixture(); + let count = Cell::new(0); + assert!( + archive_under(&f.root, &f.journal, |_, _, socket| { + let n = count.get(); + count.set(n + 1); + if n == cut { + return Err(refused()); + } + socket_absent(socket) + }) + .is_err() + ); + archive_under(&f.root, &f.journal, verify_socket).unwrap(); + for entry in &f.journal.entries { + assert_eq!( + inode(&destination(&f.root, &entry.name, &suffix(&f))).unwrap(), + entry.id + ); + } + } +} +#[test] +fn occupied_archive_target_is_never_overwritten() { + for entry_name in [ + "owner.sock", + "active-owner.json", + "shared-owner", + "owner.lock", + ] { + let f = fixture(); + let target = destination(&f.root, entry_name, &suffix(&f)); + write(&target, b"foreign evidence"); + assert!(archive_under(&f.root, &f.journal, verify_socket).is_err()); + all_original(&f); + assert_eq!(fs::read(&target).unwrap(), b"foreign evidence"); + } +} +#[test] +fn refuses_unknown_shared_files_leases_and_released_generation() { + for path in [ + "shared-owner/endpoint.json", + "shared-owner/leases/unknown.json", + "released-leases/55555555555555555555555555555555", + ] { + let f = fixture(); + if path.starts_with("released-leases") { + directory(&f.root.join("released-leases")); + } + write(&f.root.join(path), b"keep"); + assert!(archive_under(&f.root, &f.journal, verify_socket).is_err()); + all_original(&f); + } +} +#[test] +fn rejects_replaced_config_socket_lock_parent_and_changed_ca() { + for path in [ + "shared-owner/configuration.json", + "owner.sock", + "owner.lock", + "data/caddy/pki/authorities/local/root.crt", + ] { + let f = fixture(); + let p = f.root.join(path); + fs::rename(&p, p.with_extension("preserved")).unwrap(); + write(&p, b"replacement"); + assert!(archive_under(&f.root, &f.journal, verify_socket).is_err()); + } + let f = fixture(); + let prior = f.home.join("preserved-parent"); + fs::rename(&f.root, &prior).unwrap(); + directory(&f.root); + assert!(archive_under(&f.root, &f.journal, verify_socket).is_err()); + assert!(prior.join("active-owner.json").exists()); +} +#[test] +fn external_observation_cannot_change_ca_before_first_effect() { + let f = fixture(); + let ca = f.root.join("data/caddy/pki/authorities/local/root.crt"); + assert!( + archive_under(&f.root, &f.journal, |_, _, _| { + fs::write(&ca, b"changed certificate").unwrap(); + Ok(()) + }) + .is_err() + ); + all_original(&f); +} +#[test] +fn late_foreign_archive_target_is_preserved_before_any_move() { + let f = fixture(); + let target = destination(&f.root, "owner.sock", &suffix(&f)); + assert!( + archive_under(&f.root, &f.journal, |_, _, _| { + write(&target, b"late evidence"); + Ok(()) + }) + .is_err() + ); + all_original(&f); + assert_eq!(fs::read(&target).unwrap(), b"late evidence"); +} +#[test] +fn rejects_a_live_inherited_unix_listener_at_the_pinned_inode() { + let f = fixture(); + let stale = f.root.join("owner.sock"); + fs::remove_file(&stale).unwrap(); + let live = UnixListener::bind(&stale).unwrap(); + fs::set_permissions(&stale, fs::Permissions::from_mode(0o600)).unwrap(); + assert!(socket_absent(&stale).is_err()); + drop(live); + socket_absent(&stale).unwrap(); +} +#[test] +fn wildcard_port_probe_refuses_ipv4_and_ipv6_listeners() { + let ipv4 = std::net::TcpListener::bind((std::net::Ipv4Addr::LOCALHOST, 0)).unwrap(); + assert!(port_absent(ipv4.local_addr().unwrap().port()).is_err()); + drop(ipv4); + let ipv6 = std::net::TcpListener::bind((std::net::Ipv6Addr::LOCALHOST, 0)).unwrap(); + let port = ipv6.local_addr().unwrap().port(); + assert!(port_absent(port).is_err()); + drop(ipv6); + port_absent(port).unwrap(); +} +#[test] +fn symlinked_configuration_and_wrong_selectors_refuse_before_mutation() { + let f = fixture(); + let config = f.root.join("shared-owner/configuration.json"); + let target = config.with_extension("original"); + fs::rename(&config, &target).unwrap(); + std::os::unix::fs::symlink(&target, &config).unwrap(); + assert!(archive_under(&f.root, &f.journal, verify_socket).is_err()); + let candidate = Candidate::discover(&f.home).unwrap(); + for (hash, pid) in [(String::from("wrong"), 999997), ("0".repeat(64), 1)] { + assert!(recover(&candidate, &hash, &f.journal.configuration_sha256, pid).is_err()); + } + assert!(f.root.join("owner.sock").exists()); +} + +struct Alias { + path: PathBuf, + target: PathBuf, +} +impl Drop for Alias { + fn drop(&mut self) { + if fs::read_link(&self.path).ok().as_ref() == Some(&self.target) { + let _ = fs::remove_file(&self.path); + } + } +} +fn complete_fixture() -> (Fixture, Candidate, Alias) { + let f = fixture(); + let candidate = Candidate::discover(&f.home).unwrap(); + let mut owner = state::Owner::create( + &candidate, + super::super::Profile::Research, + None, + super::super::NetworkIntent::Isolated, + ) + .unwrap(); + let alias = Alias { + path: owner.short_home.clone(), + target: candidate.state_root.join("run/smolvm/home"), + }; + owner.guest_boot_id = Some("12345678-1234-1234-1234-123456789012".into()); + owner.save(&candidate).unwrap(); + drop(state::Lock::acquire(&candidate.state_root.join("run/smolvm")).unwrap()); + let binary = f.home.join("tool"); + fs::copy("/bin/sleep", &binary).unwrap(); + fs::set_permissions(&binary, fs::Permissions::from_mode(0o755)).unwrap(); + let hash = executable_hash(&binary).unwrap(); + let path = f.root.join("shared-owner/configuration.json"); + let mut config: serde_json::Value = serde_json::from_slice(&fs::read(&path).unwrap()).unwrap(); + config["binding"]["runtime"]["binary"] = serde_json::json!(binary); + config["binding"]["frontend"]["binary"] = serde_json::json!(binary); + config["binding"]["runtimeSha256"] = serde_json::json!(hash); + config["binding"]["frontend"]["sha256"] = serde_json::json!(hash); + config["binding"]["pool"]["owner"] = serde_json::json!(owner.token); + fs::write(&path, serde_json::to_vec(&config).unwrap()).unwrap(); + let ca = f.root.join("data/caddy/pki/authorities/local/root.crt"); + // DER fingerprint control, not a certificate-chain validation fixture. + let der = vec![42; 30000]; + let encoded = base64::engine::general_purpose::STANDARD.encode(&der); + fs::write( + &ca, + format!("-----BEGIN CERTIFICATE-----\n{encoded}\n-----END CERTIFICATE-----\n"), + ) + .unwrap(); + fs::set_permissions(&ca, fs::Permissions::from_mode(0o644)).unwrap(); + let receipt_path = f.root.join("active-owner.json"); + let mut receipt: serde_json::Value = + serde_json::from_slice(&fs::read(&receipt_path).unwrap()).unwrap(); + receipt["caSha256"] = serde_json::json!(digest(&der)); + receipt["authority"]["socket"] = + hostname_authority::managed::inspect(&candidate).unwrap()["socket"].clone(); + let free = std::net::TcpListener::bind((std::net::Ipv4Addr::LOCALHOST, 0)).unwrap(); + receipt["listener"]["port"] = serde_json::json!(free.local_addr().unwrap().port()); + drop(free); + config["binding"]["httpsPort"] = receipt["listener"]["port"].clone(); + fs::write(&path, serde_json::to_vec(&config).unwrap()).unwrap(); + fs::write(&receipt_path, serde_json::to_vec(&receipt).unwrap()).unwrap(); + (f, candidate, alias) +} +fn hashes(f: &Fixture) -> (String, String) { + ( + digest(&fs::read(f.root.join("active-owner.json")).unwrap()), + digest(&fs::read(f.root.join("shared-owner/configuration.json")).unwrap()), + ) +} +#[test] +fn full_command_preserves_large_public_mode_ca_and_exact_retry() { + let (f, candidate, _alias) = complete_fixture(); + let (owner, config) = hashes(&f); + let ca = f.root.join("data/caddy/pki/authorities/local/root.crt"); + let before = fs::read(&ca).unwrap(); + let id = inode(&ca).unwrap(); + assert_eq!( + recover(&candidate, &owner, &config, 999997).unwrap()["https_evidence_archived"], + true + ); + recover(&candidate, &owner, &config, 999997).unwrap(); + assert_eq!(fs::read(&ca).unwrap(), before); + assert_eq!(inode(&ca).unwrap(), id); + assert_eq!(fs::metadata(&ca).unwrap().mode() & 0o777, 0o644); +} +#[test] +fn full_command_refuses_wrong_boot_foreign_authority_unknown_schema_and_live_frontend() { + for control in ["boot", "authority", "schema", "frontend"] { + let (f, candidate, _alias) = complete_fixture(); + let config_path = f.root.join("shared-owner/configuration.json"); + let receipt_path = f.root.join("active-owner.json"); + let mut config: serde_json::Value = + serde_json::from_slice(&fs::read(&config_path).unwrap()).unwrap(); + let mut receipt: serde_json::Value = + serde_json::from_slice(&fs::read(&receipt_path).unwrap()).unwrap(); + match control { + "boot" => { + config["binding"]["pool"]["bootId"] = + serde_json::json!("87654321-1234-1234-1234-123456789012") + } + "authority" => { + receipt["authority"]["socket"] = + serde_json::json!("/private/tmp/foreign/route.sock") + } + "schema" => config["unknown"] = serde_json::json!(true), + _ => {} + } + fs::write(&config_path, serde_json::to_vec(&config).unwrap()).unwrap(); + fs::write(&receipt_path, serde_json::to_vec(&receipt).unwrap()).unwrap(); + let mut child = if control == "frontend" { + Some( + std::process::Command::new(f.home.join("tool")) + .arg("30") + .spawn() + .unwrap(), + ) + } else { + None + }; + let (owner, hash) = hashes(&f); + assert!(recover(&candidate, &owner, &hash, 999997).is_err()); + all_original(&f); + assert!( + !names(&f.root) + .unwrap() + .iter() + .any(|p| p.starts_with("recovery-")) + ); + if let Some(child) = child.as_mut() { + child.kill().unwrap(); + child.wait().unwrap(); + } + } +} +#[test] +fn exclusive_port_guards_retain_both_families_across_effects() { + let free = std::net::TcpListener::bind((std::net::Ipv4Addr::LOCALHOST, 0)).unwrap(); + let port = free.local_addr().unwrap().port(); + drop(free); + let guards = port_absent(port).unwrap(); + assert!(std::net::TcpListener::bind((std::net::Ipv4Addr::LOCALHOST, port)).is_err()); + assert!(std::net::TcpListener::bind((std::net::Ipv6Addr::LOCALHOST, port)).is_err()); + drop(guards); + port_absent(port).unwrap(); +} + +#[test] +fn legacy_device_rebind_requires_both_selected_identities_and_preserves_receipt_bytes() { + let mut f = fixture(); + let p = f.root.join("active-owner.json"); + let mut value: serde_json::Value = serde_json::from_slice(&fs::read(&p).unwrap()).unwrap(); + let current = f.journal.entries[0].id.dev; + let old = current + 7; + value["owner"]["dev"] = serde_json::json!(old); + value["lock"]["dev"] = serde_json::json!(old); + let bytes = serde_json::to_vec(&value).unwrap(); + fs::write(&p, &bytes).unwrap(); + f.journal.owner_sha256 = digest(&bytes); + assert!(archive_under(&f.root, &f.journal, verify_socket).is_err()); + all_original(&f); + let selection = LegacyDeviceRebind::parse(&format!( + "{}:{}:{}:{}:{}:{}", + "a".repeat(32), + "b".repeat(64), + current, + f.journal.entries[0].id.ino, + current, + f.journal.entries[3].id.ino + )) + .unwrap(); + let receipt: Receipt = serde_json::from_slice(&bytes).unwrap(); + assert!(selection.matches(&receipt, old, current)); + assert!(!selection.matches(&receipt, old + 1, current)); + assert!(!selection.matches(&receipt, old, current + 1)); + let mut wrong = selection.clone(); + wrong.lock.ino += 1; + f.journal.legacy_device_rebind = Some(wrong); + assert!(archive_under(&f.root, &f.journal, verify_socket).is_err()); + all_original(&f); + f.journal.legacy_device_rebind = Some(selection.clone()); + archive_under(&f.root, &f.journal, |receipt, _, socket| { + if !selection.matches(receipt, old, current) { + return Err(refused()); + } + socket_absent(socket) + }) + .unwrap(); + assert_eq!( + fs::read(destination(&f.root, "active-owner.json", &suffix(&f))).unwrap(), + bytes + ); + assert_eq!( + inode(&destination(&f.root, "owner.sock", &suffix(&f))).unwrap(), + selection.socket + ); +} + +#[test] +fn full_command_refuses_unconfirmed_legacy_witness_before_publishing_journal() { + let (f, candidate, _alias) = complete_fixture(); + let (owner, config) = hashes(&f); + let selection = format!( + "{}:{}:{}:{}:{}:{}", + "a".repeat(32), + "b".repeat(64), + f.journal.entries[0].id.dev, + f.journal.entries[0].id.ino, + f.journal.entries[3].id.dev, + f.journal.entries[3].id.ino + ); + assert!(recover_legacy_device_rebind(&candidate, &owner, &config, 999997, &selection).is_err()); + all_original(&f); + assert!( + !names(&f.root) + .unwrap() + .iter() + .any(|p| p.starts_with("recovery-")) + ); +} + +#[test] +fn full_legacy_command_archives_with_real_current_witness_and_exact_retry() { + let (f, candidate, _alias) = complete_fixture(); + fs::write(f.home.join("compose.yaml"), "services: {}\n").unwrap(); + let share = crate::provider::ProjectShareIntent::approve(&f.home, true).unwrap(); + let mut pool = state::Owner::load(&candidate).unwrap(); + pool.project_share = Some(share.clone()); + pool.save(&candidate).unwrap(); + let run = "a".repeat(32); + let graphs = candidate.state_root.join("run/graphs"); + directory(&graphs); + let graph_root = graphs.join(&run); + directory(&graph_root); + let graph: crate::provider::graph::Receipt = serde_json::from_value(serde_json::json!({ + "version":1,"run":run,"owner":pool.token,"namespace":"c".repeat(64),"plan_id":"d".repeat(64), + "phase":"stopped-data-retained","readiness":{},"resources":{}, + "source":{"shared":share,"revision":"e".repeat(64),"archive_sha256":"f".repeat(64),"selection_sha256":"0".repeat(64)} + })).unwrap(); + write( + &graph_root.join("state.json"), + &serde_json::to_vec_pretty(&graph).unwrap(), + ); + let current = f.journal.entries[0].id.dev; + let witness = crate::provider::graph::fixture_https_witness(&candidate, &graph, current + 7); + write(&graph_root.join("source-device-rebind.json"), &witness); + let p = f.root.join("active-owner.json"); + let mut receipt: serde_json::Value = serde_json::from_slice(&fs::read(&p).unwrap()).unwrap(); + receipt["owner"]["dev"] = serde_json::json!(current + 7); + receipt["lock"]["dev"] = serde_json::json!(current + 7); + let before = serde_json::to_vec(&receipt).unwrap(); + fs::write(&p, &before).unwrap(); + let (owner, config) = hashes(&f); + assert!(recover(&candidate, &owner, &config, 999997).is_err()); + let selection = format!( + "{run}:{}:{current}:{}:{current}:{}", + digest(&witness), + f.journal.entries[0].id.ino, + f.journal.entries[3].id.ino + ); + let result = + recover_legacy_device_rebind(&candidate, &owner, &config, 999997, &selection).unwrap(); + assert_eq!( + result["qualification"], + "explicit-legacy-device-migration-original-volume-continuity-unproven" + ); + assert_eq!( + recover_legacy_device_rebind(&candidate, &owner, &config, 999997, &selection).unwrap(), + result + ); + let journal: Journal = serde_json::from_slice( + &fs::read(f.root.join(format!("recovery-{owner}-{config}.json"))).unwrap(), + ) + .unwrap(); + let suffix = digest(&serde_json::to_vec(&journal).unwrap()); + assert_eq!( + fs::read(destination(&f.root, "active-owner.json", &suffix)).unwrap(), + before + ); + for entry in &f.journal.entries { + assert_eq!( + inode(&destination(&f.root, &entry.name, &suffix)).unwrap(), + entry.id + ); + } + assert_eq!( + inode(&f.root.join("data/caddy/pki/authorities/local/root.crt")).unwrap(), + f.journal.ca + ); +} diff --git a/packages/runtime-core/src/provider/lifecycle.rs b/packages/runtime-core/src/provider/lifecycle.rs index 39b2fb4ec..79e93d142 100644 --- a/packages/runtime-core/src/provider/lifecycle.rs +++ b/packages/runtime-core/src/provider/lifecycle.rs @@ -1,8 +1,11 @@ +mod admission_pool; +pub mod host_filesystem; mod interrupted; mod prepared_boot; #[cfg(any(target_os = "macos", test))] mod private_child; mod relay_process; +mod short_home; use super::{ admission, agent, artifact, identity, process, state::{self, Owner, io}, @@ -757,6 +760,15 @@ pub fn up_with_profile( up_with_bridge(candidate, profile, None) } +/// Resource admission distinguishes a proven running pool from a new VM allocation. +pub fn probe_with_profile( + candidate: &Candidate, + profile: super::Profile, +) -> Result { + let selected = admission_pool::select(candidate, profile)?; + admission_pool::probe(candidate, profile, selected.as_ref()) +} + pub fn up_with_bridge( candidate: &Candidate, profile: super::Profile, @@ -1038,19 +1050,26 @@ fn start_pool( "Existing capacity belongs to another profile. No resize, replacement or adoption was attempted.", )); } - // Admission before locks, aliases, provider commands, disks or VM effects. - let admission = admission::probe_for(&candidate.checkout, profile)?; + // Admission before aliases, provider commands, disks or VM effects. Only + // independently verified live ownership avoids charging a second VM's disks. + let admission_owner = admission_pool::select(candidate, profile)?; + let admission = admission_pool::probe(candidate, profile, admission_owner.as_ref())?; if !admission.admitted { return Err(CandidateError::new( "admission_rejected", admission.reasons.join(" "), )); } - let samples = admission::sample_for(&candidate.checkout, profile)?; + let samples = admission::sample_with(profile, || { + admission_pool::probe(candidate, profile, admission_owner.as_ref()) + })?; artifact::verify(candidate)?; artifact::verify_engine(candidate)?; super::network_tools::verify(candidate)?; let lock = startup_lease(&root(candidate), STARTUP_LEASE_WAIT)?; + if let Some(selected) = &admission_owner { + selected.reverify(candidate)?; + } #[cfg(target_os = "macos")] if let Some(guard) = &retained_guard { guard.verify(candidate)?; @@ -1110,6 +1129,9 @@ fn start_pool( super::dependency_socket::verify(candidate, &owner)?; state::write(&root(candidate).join("admission.json"), &samples)?; if owner.phase == "running" { + if let Some(selected) = &admission_owner { + selected.reverify(candidate)?; + } #[cfg(target_os = "macos")] if let Some(guard) = &retained_guard { guard.verify(candidate)?; @@ -1118,6 +1140,12 @@ fn start_pool( audit_boot(candidate, &owner)?; return status(candidate); } + if admission_owner.is_some() { + return Err(CandidateError::new( + "admission_owner_changed", + "Reserve-qualified ownership no longer names a running pool; no create or boot was admitted.", + )); + } if ![ "initializing", "stopped", @@ -1823,6 +1851,29 @@ fn finish_absent( value: &str, record_disks: bool, ) -> Result<(), CandidateError> { + let vm_lock = lock_absent_disks(candidate, owner)?; + finish_absent_locked(candidate, owner, value, record_disks, &vm_lock) +} + +/// Closing alone can leave a flock held by a fork/dup copy until that copy closes. +/// Explicitly unlock this open-file description when the recovery scope ends. +struct VmLock(File); +impl std::ops::Deref for VmLock { + type Target = File; + fn deref(&self) -> &File { + &self.0 + } +} +impl Drop for VmLock { + fn drop(&mut self) { + use std::os::fd::AsRawFd; + // SAFETY: the owned descriptor remains live through this Drop call. + unsafe { libc::flock(self.0.as_raw_fd(), libc::LOCK_UN) }; + } +} + +/// Keep this guard through alias restoration and the stopped receipt. +fn lock_absent_disks(candidate: &Candidate, owner: &Owner) -> Result { let directory = owner.real_data_dir(candidate)?; let vm_lock = OpenOptions::new() .read(true) @@ -1836,13 +1887,14 @@ fn finish_absent( return Err(CandidateError::new("foreign_state", "Unsafe VM lock.")); } use std::os::fd::AsRawFd; - // Keep the exclusive VM lock through the stopped receipt; closing the FD releases it. + // Keep the exclusive VM lock through the stopped receipt. if unsafe { libc::flock(vm_lock.as_raw_fd(), libc::LOCK_EX | libc::LOCK_NB) } != 0 { return Err(CandidateError::new( "stop_uncertain", "Provider VM lock is still held.", )); } + let vm_lock = VmLock(vm_lock); let handles = process::capture( process::clean_command(Path::new("/usr/sbin/lsof")) .args(["-n", "-P", "-t", "--"]) @@ -1858,6 +1910,16 @@ fn finish_absent( )); } verify_disks(candidate, owner)?; + Ok(vm_lock) +} + +fn finish_absent_locked( + candidate: &Candidate, + owner: &mut Owner, + value: &str, + record_disks: bool, + _vm_lock: &File, +) -> Result<(), CandidateError> { if record_disks && owner.storage.is_none() { owner.storage = Some(identity::disk( &owner.real_data_dir(candidate)?.join("storage.raw"), @@ -1907,7 +1969,13 @@ fn finish_absent( } pub fn recover(candidate: &Candidate) -> Result { - let initial = status(candidate)?; + let initial = match status(candidate) { + Ok(initial) => initial, + Err(error) if error.code == "provider_home_missing" => { + return short_home::recover(candidate); + } + Err(error) => return Err(error), + }; if initial.phase == "uninitialized" { return Ok(initial); } diff --git a/packages/runtime-core/src/provider/lifecycle/admission_pool.rs b/packages/runtime-core/src/provider/lifecycle/admission_pool.rs new file mode 100644 index 000000000..e25dcfb5c --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/admission_pool.rs @@ -0,0 +1,75 @@ +//! A host-reserve budget is valid only for one independently verified live pool. +use super::{Owner, admission, io, root, verify_disk_allocation, verify_live}; +use crate::{Candidate, CandidateError}; +use std::fs; + +pub(super) struct Selection(Owner); + +pub(super) fn select( + candidate: &Candidate, + profile: crate::provider::Profile, +) -> Result, CandidateError> { + if profile != crate::provider::Profile::Development { + return Ok(None); + } + match fs::symlink_metadata(root(candidate).join("owner.json")) { + Err(e) if e.kind() == std::io::ErrorKind::NotFound => return Ok(None), + Err(e) => return Err(io(e)), + Ok(_) => {} + } + let owner = Owner::load(candidate)?; + if owner.profile != profile || owner.phase != "running" { + return Ok(None); + } + if !owner.created || owner.storage.is_none() || owner.overlay.is_none() { + return Err(changed()); + } + verify_live(candidate, &owner)?; + verify_disk_allocation( + owner.storage.as_ref().unwrap(), + owner.overlay.as_ref().unwrap(), + profile, + )?; + Ok(Some(Selection(owner))) +} + +fn changed() -> CandidateError { + CandidateError::new( + "admission_owner_changed", + "Verified live admission ownership changed; no create or boot was admitted.", + ) +} + +impl Selection { + pub(super) fn reverify(&self, candidate: &Candidate) -> Result<(), CandidateError> { + let current = Owner::load(candidate)?; + if current != self.0 || current.phase != "running" { + return Err(changed()); + } + verify_live(candidate, ¤t)?; + verify_disk_allocation( + current.storage.as_ref().ok_or_else(changed)?, + current.overlay.as_ref().ok_or_else(changed)?, + current.profile, + ) + } +} + +pub(super) fn probe( + candidate: &Candidate, + profile: crate::provider::Profile, + selected: Option<&Selection>, +) -> Result { + match selected { + None => admission::probe_for(&candidate.checkout, profile), + Some(selected) => { + selected.reverify(candidate)?; + let observed = admission::probe_owned_development(&candidate.checkout)?; + selected.reverify(candidate)?; + Ok(observed) + } + } +} + +#[cfg(all(test, target_os = "macos"))] +mod tests; diff --git a/packages/runtime-core/src/provider/lifecycle/admission_pool/tests.rs b/packages/runtime-core/src/provider/lifecycle/admission_pool/tests.rs new file mode 100644 index 000000000..cc33363a0 --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/admission_pool/tests.rs @@ -0,0 +1,283 @@ +use super::*; +use crate::provider::{NetworkIntent, Profile, artifact, identity, state}; +use std::{ + fs::{self, OpenOptions}, + io::{Seek, SeekFrom, Write}, + os::unix::fs::{DirBuilderExt, OpenOptionsExt, PermissionsExt}, + path::{Path, PathBuf}, + process::{Child, Command}, + sync::atomic::{AtomicU64, Ordering}, + time::{SystemTime, UNIX_EPOCH}, +}; + +const GIB: u64 = 1024 * 1024 * 1024; + +struct Cleanup { + directory: PathBuf, + alias: Option<(PathBuf, PathBuf)>, + child: Option, +} + +impl Drop for Cleanup { + fn drop(&mut self) { + if let Some(child) = &mut self.child { + let _ = child.kill(); + let _ = child.wait(); + } + if let Some((alias, target)) = &self.alias + && fs::read_link(alias).ok().as_ref() == Some(target) + { + let _ = fs::remove_file(alias); + } + let _ = fs::remove_dir_all(&self.directory); + } +} + +fn private_directory() -> PathBuf { + static NEXT: AtomicU64 = AtomicU64::new(0); + let base = fs::canonicalize(std::env::temp_dir()).unwrap(); + let timestamp = SystemTime::now() + .duration_since(UNIX_EPOCH) + .unwrap() + .as_nanos(); + loop { + let directory = base.join(format!( + "hack-admission-pool-{}-{timestamp}-{}", + std::process::id(), + NEXT.fetch_add(1, Ordering::Relaxed) + )); + match fs::DirBuilder::new().mode(0o700).create(&directory) { + Ok(()) => return directory, + Err(error) if error.kind() == std::io::ErrorKind::AlreadyExists => continue, + Err(error) => panic!("cannot create private admission fixture: {error}"), + } + } +} + +fn sparse_disk(path: &Path, gib: u64, tag: u8) { + let mut file = OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .open(path) + .unwrap(); + file.set_len(gib * GIB).unwrap(); + file.seek(SeekFrom::Start(1080)).unwrap(); + file.write_all(&[0x53, 0xef]).unwrap(); + file.seek(SeekFrom::Start(1128)).unwrap(); + file.write_all(&[tag; 16]).unwrap(); +} + +struct Pool { + candidate: Candidate, + owner: Owner, + cleanup: Cleanup, +} + +impl Pool { + fn new() -> Self { + let directory = private_directory(); + let mut cleanup = Cleanup { + directory: directory.clone(), + alias: None, + child: None, + }; + let candidate = Candidate::discover(&directory).unwrap(); + let provider = artifact::root(&candidate); + state::private_directory(&provider).unwrap(); + let binary = provider.join("smolvm-bin"); + fs::copy("/bin/sleep", &binary).unwrap(); + fs::set_permissions(&binary, fs::Permissions::from_mode(0o755)).unwrap(); + + let mut owner = Owner::create( + &candidate, + Profile::Development, + None, + NetworkIntent::Isolated, + ) + .unwrap(); + cleanup.alias = Some(( + owner.short_home.clone(), + candidate.state_root.join("run/smolvm/home"), + )); + let data = owner.real_data_dir(&candidate).unwrap(); + state::private_directory(&data).unwrap(); + sparse_disk(&data.join("storage.raw"), 32, 1); + sparse_disk(&data.join("overlay.raw"), 10, 2); + fs::write(data.join("name"), &owner.machine).unwrap(); + + cleanup.child = Some(Command::new(&binary).arg("120").spawn().unwrap()); + let process = identity::observe(cleanup.child.as_ref().unwrap().id() as i32).unwrap(); + assert_eq!(process.executable, binary); + fs::write( + data.join("agent.pid"), + format!("{}\n{}\n", process.pid, process.start_micros), + ) + .unwrap(); + owner.process = Some(process); + owner.created = true; + owner.phase = "running".into(); + owner.guest_boot_id = Some("11111111-1111-1111-1111-111111111111".into()); + owner.storage = Some(identity::disk(&data.join("storage.raw")).unwrap()); + owner.overlay = Some(identity::disk(&data.join("overlay.raw")).unwrap()); + owner.save(&candidate).unwrap(); + Self { + candidate, + owner, + cleanup, + } + } + + fn data(&self) -> PathBuf { + self.owner.real_data_dir(&self.candidate).unwrap() + } + + fn save(&self) { + self.owner.save(&self.candidate).unwrap(); + } +} + +#[test] +fn only_a_live_owned_development_pool_selects_the_host_reserve() { + let pool = Pool::new(); + let receipt = fs::read(root(&pool.candidate).join("owner.json")).unwrap(); + let storage = identity::disk(&pool.data().join("storage.raw")).unwrap(); + let overlay = identity::disk(&pool.data().join("overlay.raw")).unwrap(); + + let selection = select(&pool.candidate, Profile::Development) + .unwrap() + .expect("exact live development ownership"); + selection.reverify(&pool.candidate).unwrap(); + assert!( + select(&pool.candidate, Profile::Research) + .unwrap() + .is_none() + ); + assert_eq!( + fs::read(root(&pool.candidate).join("owner.json")).unwrap(), + receipt + ); + assert_eq!( + identity::disk(&pool.data().join("storage.raw")).unwrap(), + storage + ); + assert_eq!( + identity::disk(&pool.data().join("overlay.raw")).unwrap(), + overlay + ); +} + +#[test] +fn absent_or_non_running_owners_do_not_select_the_host_reserve() { + let directory = private_directory(); + let cleanup = Cleanup { + directory: directory.clone(), + alias: None, + child: None, + }; + let candidate = Candidate::discover(&directory).unwrap(); + assert!(select(&candidate, Profile::Development).unwrap().is_none()); + drop(cleanup); + + for phase in ["stopped", "stopped-before-engine"] { + let mut pool = Pool::new(); + pool.owner.phase = phase.into(); + pool.save(); + assert!( + select(&pool.candidate, Profile::Development) + .unwrap() + .is_none() + ); + } + let mut pool = Pool::new(); + pool.owner.profile = Profile::Research; + pool.save(); + assert!( + select(&pool.candidate, Profile::Development) + .unwrap() + .is_none() + ); +} + +#[test] +fn incomplete_or_mis_sized_running_capacity_refuses_selection() { + for missing in ["created", "storage", "overlay", "size"] { + let mut pool = Pool::new(); + match missing { + "created" => pool.owner.created = false, + "storage" => pool.owner.storage = None, + "overlay" => pool.owner.overlay = None, + "size" => { + let path = pool.data().join("overlay.raw"); + OpenOptions::new() + .write(true) + .open(path) + .unwrap() + .set_len(9 * GIB) + .unwrap(); + pool.owner.overlay = + Some(identity::disk(&pool.data().join("overlay.raw")).unwrap()); + } + _ => unreachable!(), + } + pool.save(); + assert!( + select(&pool.candidate, Profile::Development).is_err(), + "{missing}" + ); + } +} + +#[test] +fn stale_or_dead_process_and_replaced_or_missing_disks_refuse_selection() { + let mut pool = Pool::new(); + pool.owner.process.as_mut().unwrap().start_micros += 1; + pool.save(); + assert!(select(&pool.candidate, Profile::Development).is_err()); + + let mut pool = Pool::new(); + let child = pool.cleanup.child.as_mut().unwrap(); + child.kill().unwrap(); + child.wait().unwrap(); + assert!(select(&pool.candidate, Profile::Development).is_err()); + + let pool = Pool::new(); + let path = pool.data().join("storage.raw"); + fs::rename(&path, pool.data().join("storage.retained")).unwrap(); + sparse_disk(&path, 32, 1); + assert!(select(&pool.candidate, Profile::Development).is_err()); + + let pool = Pool::new(); + fs::remove_file(pool.data().join("overlay.raw")).unwrap(); + assert!(select(&pool.candidate, Profile::Development).is_err()); +} + +#[test] +fn selected_owner_or_provider_changes_refuse_reverification() { + let mut pool = Pool::new(); + let selection = select(&pool.candidate, Profile::Development) + .unwrap() + .unwrap(); + pool.owner.guest_boot_id = Some("22222222-2222-2222-2222-222222222222".into()); + pool.save(); + assert_eq!( + selection.reverify(&pool.candidate).err().unwrap().code, + "admission_owner_changed" + ); + + let pool = Pool::new(); + let selection = select(&pool.candidate, Profile::Development) + .unwrap() + .unwrap(); + fs::write(pool.data().join("agent.pid"), b"999999\n1\n").unwrap(); + assert!(selection.reverify(&pool.candidate).is_err()); + + let pool = Pool::new(); + let selection = select(&pool.candidate, Profile::Development) + .unwrap() + .unwrap(); + let path = pool.data().join("overlay.raw"); + fs::rename(&path, pool.data().join("overlay.retained")).unwrap(); + sparse_disk(&path, 10, 2); + assert!(selection.reverify(&pool.candidate).is_err()); +} diff --git a/packages/runtime-core/src/provider/lifecycle/host_filesystem.rs b/packages/runtime-core/src/provider/lifecycle/host_filesystem.rs new file mode 100644 index 000000000..2b7ad0ddb --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/host_filesystem.rs @@ -0,0 +1,297 @@ +//! Explicit legacy device-number migration for an offline stock pool. +//! +//! Old receipts have neither a host boot UUID nor a filesystem volume UUID. Calendar start +//! time, unchanged inode/size/ext4 UUID and exact paths constrain this migration, but cannot +//! establish original volume continuity. Only an explicitly accepted, hash-selected inspection +//! may update the owner. Normal runtime operations retain their strict identity comparisons. +use super::{Owner, binary, identity, lock_absent_disks, root, state}; +use crate::{Candidate, CandidateError}; +use serde::Serialize; +use sha2::{Digest, Sha256}; +use std::{fs, os::unix::fs::MetadataExt, path::Path}; + +#[derive(Debug, Serialize, PartialEq, Eq)] +pub struct Inspection { + pub schema: &'static str, + pub selection_sha256: String, + pub owner_sha256: String, + pub machine: String, + pub host_boot_micros: u64, + pub provider_start_micros: u64, + pub old_device: u64, + pub new_device: u64, + pub pool_inode: u64, + pub storage: identity::DiskIdentity, + pub overlay: identity::DiskIdentity, + pub project_share: Option, + pub qualification: &'static str, +} + +fn refused(detail: &str) -> CandidateError { + CandidateError::new( + "host_filesystem_recovery", + format!("{detail}; no identity was changed."), + ) +} + +#[cfg(target_os = "macos")] +pub(crate) fn host_boot_micros() -> Result { + let mut time = std::mem::MaybeUninit::::zeroed(); + let mut length = std::mem::size_of::(); + // SAFETY: read-only sysctl writes an exactly sized timeval; no input or retained pointers. + if unsafe { + libc::sysctlbyname( + c"kern.boottime".as_ptr(), + time.as_mut_ptr().cast(), + &mut length, + std::ptr::null_mut(), + 0, + ) + } != 0 + || length != std::mem::size_of::() + { + return Err(refused("Native host boot time is unavailable")); + } + // SAFETY: a complete timeval was returned above. + let time = unsafe { time.assume_init() }; + let seconds = u64::try_from(time.tv_sec).map_err(|_| refused("Invalid host boot time"))?; + let micros = u64::try_from(time.tv_usec).map_err(|_| refused("Invalid host boot time"))?; + if micros >= 1_000_000 { + return Err(refused("Invalid host boot time")); + } + seconds + .checked_mul(1_000_000) + .and_then(|v| v.checked_add(micros)) + .filter(|v| *v > 0) + .ok_or_else(|| refused("Invalid host boot time")) +} + +#[cfg(not(target_os = "macos"))] +fn host_boot_micros() -> Result { + Err(CandidateError::new( + "unsupported_host", + "Legacy filesystem recovery requires macOS.", + )) +} + +pub(crate) fn no_auxiliary_update(candidate: &Candidate) -> Result<(), CandidateError> { + crate::provider::network_update::require_complete(candidate)?; + for name in [ + "owner.pending", + "prepared-base.json", + "prepared-base.json.pending", + ] { + match fs::symlink_metadata(root(candidate).join(name)) { + Err(e) if e.kind() == std::io::ErrorKind::NotFound => {} + _ => { + return Err(refused( + "Pending owner or prepared-base state requires separate recovery", + )); + } + } + } + Ok(()) +} + +fn dead_provider(candidate: &Candidate, owner: &Owner, boot: u64) -> Result<(), CandidateError> { + let process = owner + .process + .as_ref() + .ok_or_else(|| refused("No recorded provider identity"))?; + // SAFETY: geteuid has no preconditions. + identity::verify(process, process, &binary(candidate), unsafe { + libc::geteuid() + })?; + if !owner.created + || process.start_micros >= boot + || identity::alive(process.pid)? + || identity::executable_running(&binary(candidate))? + { + return Err(refused( + "Provider must be absent and its recorded start must predate this host boot", + )); + } + Ok(()) +} + +fn same_disk(before: &identity::DiskIdentity, after: &identity::DiskIdentity) -> bool { + before.device != after.device + && before.inode == after.inode + && before.bytes == after.bytes + && before.uuid == after.uuid +} + +/// A retained flock cannot fence a writer that opens a substituted lock pathname. +fn bound_locks( + candidate: &Candidate, + owner: &Owner, + operation: &state::Lock, + vm: &fs::File, +) -> Result<(), CandidateError> { + let check = |path: &Path, expected: (u64, u64)| -> Result<(), CandidateError> { + let metadata = fs::symlink_metadata(path).map_err(state::io)?; + if !metadata.is_file() + || metadata.nlink() != 1 + || (metadata.dev(), metadata.ino()) != expected + { + return Err(refused("Held lock pathname was replaced")); + } + Ok(()) + }; + check( + &root(candidate).join("operation.lock"), + operation.identity()?, + )?; + let metadata = vm.metadata().map_err(state::io)?; + check( + &owner.real_data_dir(candidate)?.join("vm.lock"), + (metadata.dev(), metadata.ino()), + ) +} + +/// Build a candidate owner in memory. This function neither adopts a new disk nor writes state. +fn selection( + candidate: &Candidate, + owner: &Owner, + boot: u64, +) -> Result<(Inspection, Owner), CandidateError> { + no_auxiliary_update(candidate)?; + dead_provider(candidate, owner, boot)?; + let data = owner.real_data_dir(candidate)?; + let storage = identity::disk(&data.join("storage.raw"))?; + let overlay = identity::disk(&data.join("overlay.raw"))?; + let before = owner + .storage + .as_ref() + .ok_or_else(|| refused("Storage identity is missing"))?; + let prior_overlay = owner + .overlay + .as_ref() + .ok_or_else(|| refused("Overlay identity is missing"))?; + let metadata = fs::symlink_metadata(root(candidate)).map_err(state::io)?; + if !same_disk(before, &storage) + || !same_disk(prior_overlay, &overlay) + || before.device != prior_overlay.device + || storage.device != overlay.device + || storage.device != metadata.dev() + { + return Err(refused( + "Only a common device-number change with unchanged disk inode, size and UUID is accepted", + )); + } + let mut next = owner.clone(); + if let Some(share) = &owner.project_share { + let observed = + crate::provider::ProjectShareIntent::approve(&share.project, share.unfiltered_source)?; + let mut expected = share.clone(); + expected.device = storage.device; + if share.device != before.device || observed != expected { + return Err(refused( + "Project share changed beyond the same filesystem device number", + )); + } + next.project_share = Some(observed); + } + next.storage = Some(storage.clone()); + next.overlay = Some(overlay.clone()); + // The canonical read is bounded and validated independently of raw-byte hashing. + let bytes = crate::provider::prepared_base::read_private( + &root(candidate).join("owner.json"), + 1024 * 1024, + )?; + if Owner::load_for_short_home_recovery(candidate)? != *owner { + return Err(refused("Owner selection changed")); + } + let mut inspection = Inspection { + schema: "hack.host-filesystem-recovery/v1", + selection_sha256: String::new(), + owner_sha256: format!("{:x}", Sha256::digest(bytes)), + machine: owner.machine.clone(), + host_boot_micros: boot, + provider_start_micros: owner + .process + .as_ref() + .ok_or_else(|| refused("No process"))? + .start_micros, + old_device: before.device, + new_device: storage.device, + pool_inode: metadata.ino(), + storage, + overlay, + project_share: next.project_share.clone(), + qualification: "explicit-legacy-migration-original-volume-continuity-unproven", + }; + inspection.selection_sha256 = format!( + "{:x}", + Sha256::digest( + serde_json::to_vec(&inspection).map_err(|_| refused("Cannot encode selection"))? + ) + ); + Ok((inspection, next)) +} + +/// Read-only inspection holds both existing locks and proves disks have no open handles. +pub fn inspect(candidate: &Candidate) -> Result { + let operation = state::Lock::acquire_existing(&root(candidate))?; + let owner = Owner::load_for_short_home_recovery(candidate)?; + let (inspection, next) = selection(candidate, &owner, host_boot_micros()?)?; + let vm = lock_absent_disks(candidate, &next)?; + if selection(candidate, &owner, host_boot_micros()?)?.0 != inspection { + return Err(refused("Inspection changed during absence verification")); + } + bound_locks(candidate, &owner, &operation, &vm)?; + Ok(inspection) +} + +/// Publish disk and source device changes in ONE atomic owner replacement. A crash leaves +/// old or new committed metadata; any unfinished owner.pending is preserved and blocks retry. +/// This does not launch, restore the alias, retire historical sockets, or rewrite graph receipts. +pub fn recover(candidate: &Candidate, expected: &str) -> Result { + recover_with_boot(candidate, expected, host_boot_micros) +} + +fn recover_with_boot( + candidate: &Candidate, + expected: &str, + boot: impl Fn() -> Result, +) -> Result { + if expected.len() != 64 || !expected.bytes().all(|b| b.is_ascii_hexdigit()) { + return Err(refused("An exact inspection SHA-256 is required")); + } + let operation = state::Lock::acquire_existing(&root(candidate))?; + let owner = Owner::load_for_short_home_recovery(candidate)?; + let (inspection, next) = selection(candidate, &owner, boot()?)?; + if inspection.selection_sha256 != expected { + return Err(refused("Inspection selection is stale")); + } + let vm = lock_absent_disks(candidate, &next)?; + if selection(candidate, &owner, boot()?)?.0 != inspection { + return Err(refused("Selected identities changed before publication")); + } + bound_locks(candidate, &owner, &operation, &vm)?; + next.save(candidate)?; + if Owner::load_for_short_home_recovery(candidate)? != next { + return Err(CandidateError::new( + "host_filesystem_recovery_incomplete", + "Owner was published but changed during confirmation; retained state requires inspection.", + )); + } + super::verify_disks(candidate, &next).map_err(|e| { + CandidateError::new( + "host_filesystem_recovery_incomplete", + format!( + "Owner was published but disk confirmation failed ({}); inspect retained state.", + e.code + ), + ) + })?; + if let Some(share) = &next.project_share { + share.validate().map_err(|e| CandidateError::new( + "host_filesystem_recovery_incomplete", format!("Owner was published but source confirmation failed ({}); inspect retained state.", e.code) + ))?; + } + Ok(inspection) +} + +#[cfg(all(test, target_os = "macos"))] +mod tests; diff --git a/packages/runtime-core/src/provider/lifecycle/host_filesystem/tests.rs b/packages/runtime-core/src/provider/lifecycle/host_filesystem/tests.rs new file mode 100644 index 000000000..89ebeda4b --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/host_filesystem/tests.rs @@ -0,0 +1,367 @@ +use super::*; +use crate::provider::{NetworkIntent, Profile, ProjectShareIntent}; +use std::{ + fs::{self, File}, + io::ErrorKind, + os::{fd::AsRawFd, unix::fs::DirBuilderExt}, + path::PathBuf, + process::Command, + sync::atomic::{AtomicU64, Ordering}, + time::{SystemTime, UNIX_EPOCH}, +}; + +fn private_pool_directory_at(timestamp_nanos: u128, sequence: &AtomicU64) -> PathBuf { + let base = fs::canonicalize(std::env::temp_dir()).unwrap(); + loop { + let directory = base.join(format!( + "hack-device-recovery-{}-{timestamp_nanos}-{}", + std::process::id(), + sequence.fetch_add(1, Ordering::Relaxed) + )); + match fs::DirBuilder::new().mode(0o700).create(&directory) { + Ok(()) => return directory, + Err(error) if error.kind() == ErrorKind::AlreadyExists => continue, + Err(error) => panic!("cannot create private host-filesystem fixture: {error}"), + } + } +} + +fn private_pool_directory() -> PathBuf { + static NEXT: AtomicU64 = AtomicU64::new(0); + let timestamp_nanos = SystemTime::now() + .duration_since(UNIX_EPOCH) + .unwrap() + .as_nanos(); + private_pool_directory_at(timestamp_nanos, &NEXT) +} + +struct Pool { + candidate: Candidate, + owner: Owner, + directory: PathBuf, + boot: u64, +} +impl Pool { + fn new() -> Self { + let mut child = Command::new("/bin/sleep").arg("30").spawn().unwrap(); + let mut process = identity::observe(child.id() as i32).unwrap(); + child.kill().unwrap(); + child.wait().unwrap(); + let boot = process.start_micros + 1; + let directory = private_pool_directory(); + let candidate = Candidate::discover(&directory).unwrap(); + let operation = state::Lock::acquire(&root(&candidate)).unwrap(); + let project = directory.join("app"); + state::private_directory(&project).unwrap(); + fs::write(project.join("package.json"), b"{}").unwrap(); + let share = ProjectShareIntent::approve(&project, true).unwrap(); + let mut owner = Owner::create_with_project_share( + &candidate, + Profile::Development, + None, + NetworkIntent::Isolated, + None, + Some(share), + ) + .unwrap(); + let data = owner.real_data_dir(&candidate).unwrap(); + state::private_directory(&data).unwrap(); + for (name, tag) in [("storage.raw", 1u8), ("overlay.raw", 2)] { + let mut bytes = vec![0; 4096]; + bytes[1080..1082].copy_from_slice(&[0x53, 0xef]); + bytes[1128..1144].fill(tag); + bytes[2048..2059].copy_from_slice(b"data-marker"); + fs::write(data.join(name), bytes).unwrap(); + } + fs::write(data.join("name"), &owner.machine).unwrap(); + fs::write(data.join("vm.lock"), b"").unwrap(); + process.executable = binary(&candidate); + owner.process = Some(process); + owner.created = true; + owner.phase = "running".into(); + let mut storage = identity::disk(&data.join("storage.raw")).unwrap(); + let mut overlay = identity::disk(&data.join("overlay.raw")).unwrap(); + let old = storage.device + 1; + storage.device = old; + overlay.device = old; + owner.storage = Some(storage); + owner.overlay = Some(overlay); + owner.project_share.as_mut().unwrap().device = old; + owner.save(&candidate).unwrap(); + fs::remove_file(&owner.short_home).unwrap(); + drop(operation); + Self { + candidate, + owner, + directory, + boot, + } + } + fn data(&self) -> PathBuf { + self.owner.real_data_dir(&self.candidate).unwrap() + } + fn selected(&self) -> Inspection { + selection(&self.candidate, &self.owner, self.boot) + .unwrap() + .0 + } + fn receipt(&self) -> Vec { + fs::read(root(&self.candidate).join("owner.json")).unwrap() + } + fn apply(&self, hash: &str) -> Result { + recover_with_boot(&self.candidate, hash, || Ok(self.boot)) + } + fn unchanged(&self, before: &[u8]) { + assert_eq!(self.receipt(), before); + assert!(self.owner.short_home.symlink_metadata().is_err()); + } +} + +#[test] +fn same_timestamp_parallel_roots_preserve_an_existing_directory() { + let timestamp_nanos = SystemTime::now() + .duration_since(UNIX_EPOCH) + .unwrap() + .as_nanos(); + let sequence = AtomicU64::new(0); + let (first, second) = std::thread::scope(|scope| { + let first = scope.spawn(|| private_pool_directory_at(timestamp_nanos, &sequence)); + let second = scope.spawn(|| private_pool_directory_at(timestamp_nanos, &sequence)); + (first.join().unwrap(), second.join().unwrap()) + }); + assert_ne!(first, second); + let sentinel = first.join("sentinel"); + fs::write(&sentinel, b"untouched").unwrap(); + let restarted_sequence = AtomicU64::new(0); + let third = private_pool_directory_at(timestamp_nanos, &restarted_sequence); + assert_ne!(third, first); + assert_ne!(third, second); + assert_eq!(fs::read(sentinel).unwrap(), b"untouched"); + for directory in [first, second, third] { + fs::remove_dir_all(directory).unwrap(); + } +} +impl Drop for Pool { + fn drop(&mut self) { + let _ = fs::remove_file(&self.owner.short_home); + fs::remove_dir_all(&self.directory).unwrap(); + } +} + +#[test] +fn migration_changes_only_devices_and_keeps_alias_graph_history_and_data() { + let pool = Pool::new(); + let old = pool.receipt(); + let storage = fs::read(pool.data().join("storage.raw")).unwrap(); + let overlay = fs::read(pool.data().join("overlay.raw")).unwrap(); + let history = pool.directory.join("historical-graph.json"); + fs::write(&history, b"historical-source-binding").unwrap(); + let inspected = pool.selected(); + pool.unchanged(&old); + let result = pool.apply(&inspected.selection_sha256).unwrap(); + assert_eq!(result, inspected); + let mut expected = pool.owner.clone(); + expected.storage.as_mut().unwrap().device = inspected.new_device; + expected.overlay.as_mut().unwrap().device = inspected.new_device; + expected.project_share.as_mut().unwrap().device = inspected.new_device; + assert_eq!( + Owner::load_for_short_home_recovery(&pool.candidate).unwrap(), + expected + ); + assert!(pool.owner.short_home.symlink_metadata().is_err()); + assert_eq!(fs::read(history).unwrap(), b"historical-source-binding"); + assert_eq!(fs::read(pool.data().join("storage.raw")).unwrap(), storage); + assert_eq!(fs::read(pool.data().join("overlay.raw")).unwrap(), overlay); + assert!(pool.apply(&inspected.selection_sha256).is_err()); + // Ordinary recovery still owns HOME restoration and the recovered phase. + let recovered = crate::provider::recover(&pool.candidate).unwrap(); + assert_eq!(recovered.phase, "recovered-unclean"); + assert_eq!(recovered.process_alive, Some(false)); + assert_eq!(fs::read(pool.data().join("storage.raw")).unwrap(), storage); +} + +#[test] +fn disk_and_share_changes_beyond_common_device_refuse_without_writes() { + for case in 0..7 { + let mut pool = Pool::new(); + let observed_device = identity::disk(&pool.data().join("storage.raw")) + .unwrap() + .device; + let disk = pool.owner.storage.as_mut().unwrap(); + match case { + 0 => disk.inode += 1, + 1 => disk.bytes += 1, + 2 => disk.uuid.push('0'), + 3 => disk.device = observed_device, + 4 => pool.owner.overlay.as_mut().unwrap().device += 1, + 5 => pool.owner.project_share.as_mut().unwrap().inode += 1, + _ => pool.owner.project_share.as_mut().unwrap().device += 1, + } + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert!(pool.apply(&"a".repeat(64)).is_err()); + pool.unchanged(&before); + } +} + +#[test] +fn stale_owner_or_host_boot_selection_refuses() { + let mut pool = Pool::new(); + let selection = pool.selected(); + assert!(pool.apply(&"a".repeat(64)).is_err()); + pool.owner.phase = "stopped".into(); + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert!(pool.apply(&selection.selection_sha256).is_err()); + pool.unchanged(&before); + let selected = pool.selected(); + assert!( + recover_with_boot(&pool.candidate, &selected.selection_sha256, || Ok(pool + .boot + + 1)) + .is_err() + ); + pool.unchanged(&before); +} + +#[test] +fn same_boot_live_reused_pid_and_unknown_identity_refuse() { + for case in 0..4 { + let mut pool = Pool::new(); + if case == 0 { + pool.boot = pool.owner.process.as_ref().unwrap().start_micros; + } + if case == 1 || case == 2 { + let p = pool.owner.process.as_mut().unwrap(); + p.pid = std::process::id() as i32; + if case == 1 { + p.start_micros = identity::observe(p.pid).unwrap().start_micros; + } + } + if case == 3 { + pool.owner.process = None; + } + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert!(pool.apply(&"a".repeat(64)).is_err()); + pool.unchanged(&before); + } +} + +#[test] +fn held_operation_vm_lock_and_disk_handles_refuse() { + for case in 0..3 { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let _operation = + (case == 0).then(|| state::Lock::acquire_existing(&root(&pool.candidate)).unwrap()); + let file = if case == 1 { + Some(File::open(pool.data().join("vm.lock")).unwrap()) + } else if case == 2 { + Some(File::open(pool.data().join("storage.raw")).unwrap()) + } else { + None + }; + if case == 1 { + assert_eq!( + unsafe { + libc::flock( + file.as_ref().unwrap().as_raw_fd(), + libc::LOCK_EX | libc::LOCK_NB, + ) + }, + 0 + ); + } + assert!(pool.apply(&selected.selection_sha256).is_err()); + pool.unchanged(&before); + } +} + +#[test] +fn pending_foreign_updates_and_prepared_pools_are_preserved() { + for name in [ + "owner.pending", + "network-update.json", + "network-update.pending", + "prepared-base.json", + "prepared-base.json.pending", + ] { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let path = root(&pool.candidate).join(name); + fs::write(&path, b"foreign-or-interrupted").unwrap(); + assert!(pool.apply(&selected.selection_sha256).is_err()); + pool.unchanged(&before); + assert_eq!(fs::read(path).unwrap(), b"foreign-or-interrupted"); + } +} + +#[test] +fn foreign_alias_and_source_substitution_refuse() { + for case in 0..2 { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + if case == 0 { + fs::write(&pool.owner.short_home, b"foreign").unwrap(); + } else { + let project = &pool.owner.project_share.as_ref().unwrap().project; + fs::rename(project, pool.directory.join("preserved-app")).unwrap(); + state::private_directory(project).unwrap(); + fs::write(project.join("package.json"), b"{}").unwrap(); + } + assert!(pool.apply(&selected.selection_sha256).is_err()); + assert_eq!(pool.receipt(), before); + if case == 0 { + assert_eq!(fs::read(&pool.owner.short_home).unwrap(), b"foreign"); + } + } +} + +#[test] +fn final_recheck_rejects_identity_change_before_publication() { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let called = std::cell::Cell::new(false); + let result = recover_with_boot(&pool.candidate, &selected.selection_sha256, || { + if called.replace(true) { + fs::write(pool.data().join("storage.raw"), b"changed").unwrap(); + } + Ok(pool.boot) + }); + assert!(result.is_err()); + pool.unchanged(&before); + assert_eq!( + fs::read(pool.data().join("storage.raw")).unwrap(), + b"changed" + ); +} + +#[test] +fn substituted_lock_paths_do_not_authorize_publication_on_an_old_descriptor() { + for name in ["operation.lock", "vm.lock"] { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let called = std::cell::Cell::new(false); + let path = if name == "operation.lock" { + root(&pool.candidate).join(name) + } else { + pool.data().join(name) + }; + let result = recover_with_boot(&pool.candidate, &selected.selection_sha256, || { + if called.replace(true) { + fs::rename(&path, path.with_extension("preserved")).unwrap(); + fs::write(&path, b"replacement lock").unwrap(); + } + Ok(pool.boot) + }); + assert!(matches!(result, Err(e) if e.code == "host_filesystem_recovery")); + pool.unchanged(&before); + assert_eq!(fs::read(path).unwrap(), b"replacement lock"); + } +} diff --git a/packages/runtime-core/src/provider/lifecycle/short_home.rs b/packages/runtime-core/src/provider/lifecycle/short_home.rs new file mode 100644 index 000000000..6678ab602 --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/short_home.rs @@ -0,0 +1,70 @@ +//! Explicit two-phase recovery of the receipt-bound temporary HOME after host cleanup. +//! +//! No provider is launched and no existing alias is replaced. Durable disks and VM absence +//! are audited before restoring the short socket path; socket absence is then checked before +//! committing the recovered phase. A later refusal retains the exact alias and unchanged data. +use super::{ + Owner, RuntimeStatus, binary, finish_absent_locked, identity, lock_absent_disks, root, state, + status, verify_disks, +}; +use crate::{Candidate, CandidateError}; + +fn dead_provider(candidate: &Candidate, owner: &Owner) -> Result<(), CandidateError> { + let process = owner.process.as_ref().ok_or_else(|| { + CandidateError::new( + "recovery_required", + "Missing HOME recovery requires a recorded provider identity; nothing was changed.", + ) + })?; + // SAFETY: geteuid has no preconditions. + identity::verify(process, process, &binary(candidate), unsafe { + libc::geteuid() + })?; + if identity::alive(process.pid)? || identity::executable_running(&binary(candidate))? { + return Err(CandidateError::new( + "recovery_required", + "Missing HOME recovery requires a confirmed dead provider and no active provider command; nothing was changed.", + )); + } + if !owner.created || owner.storage.is_none() || owner.overlay.is_none() { + return Err(CandidateError::new( + "recovery_required", + "Missing HOME recovery requires both previously identified disks; nothing was adopted.", + )); + } + Ok(()) +} + +pub(super) fn recover(candidate: &Candidate) -> Result { + let _operation = state::Lock::acquire_existing(&root(candidate))?; + let mut owner = Owner::load_for_short_home_recovery(candidate)?; + dead_provider(candidate, &owner)?; + let vm_lock = lock_absent_disks(candidate, &owner)?; + // Recheck immediately before the only new external effect. The locks fence cooperative + // writers; no signal, disk adoption, overwrite or receipt update happens in this phase. + dead_provider(candidate, &owner)?; + verify_disks(candidate, &owner)?; + owner.restore_missing_short_home(candidate)?; + let mut finish = || -> Result<(), CandidateError> { + let observed = Owner::load(candidate)?; + if observed != owner { + return Err(CandidateError::new( + "foreign_state", + "Provider receipt changed after HOME restoration.", + )); + } + dead_provider(candidate, &owner)?; + verify_disks(candidate, &owner)?; + finish_absent_locked(candidate, &mut owner, "recovered-unclean", false, &vm_lock) + }; + finish().map_err(|error| { + CandidateError::new( + "provider_home_restored_recovery_incomplete", + format!("Owned HOME alias restored; runtime recovery is incomplete ({}). Data retained; inspect and retry runtime recover.", error.code), + ) + })?; + status(candidate) +} + +#[cfg(all(test, target_os = "macos"))] +mod tests; diff --git a/packages/runtime-core/src/provider/lifecycle/short_home/tests.rs b/packages/runtime-core/src/provider/lifecycle/short_home/tests.rs new file mode 100644 index 000000000..2654463e9 --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/short_home/tests.rs @@ -0,0 +1,368 @@ +use super::*; +use crate::provider::{NetworkIntent, Profile}; +use std::{ + fs, + os::{ + fd::AsRawFd, + unix::{ + fs::{DirBuilderExt, MetadataExt}, + net::UnixListener, + }, + }, + path::PathBuf, + process::Command, + sync::atomic::{AtomicU64, Ordering}, +}; + +static NEXT_POOL: AtomicU64 = AtomicU64::new(0); + +fn pool_directory(timestamp: u128) -> PathBuf { + fs::canonicalize(std::env::temp_dir()) + .unwrap() + .join(format!( + "hack-home-recovery-{}-{timestamp}-{}", + std::process::id(), + NEXT_POOL.fetch_add(1, Ordering::Relaxed) + )) +} + +fn create_pool_directory(path: &std::path::Path) -> std::io::Result<()> { + fs::DirBuilder::new().mode(0o700).create(path) +} + +#[test] +fn parallel_pool_roots_are_unique_at_the_same_timestamp_and_never_adopt() { + let paths: Vec<_> = std::thread::scope(|scope| { + let workers: Vec<_> = (0..16) + .map(|_| { + scope.spawn(|| { + let path = pool_directory(123); + create_pool_directory(&path).unwrap(); + path + }) + }) + .collect(); + workers + .into_iter() + .map(|worker| worker.join().unwrap()) + .collect() + }); + assert_eq!( + paths + .iter() + .collect::>() + .len(), + 16 + ); + for path in paths { + let before = fs::symlink_metadata(&path).unwrap(); + assert_eq!(before.mode() & 0o777, 0o700); + fs::write(path.join("marker"), b"owned fixture").unwrap(); + assert_eq!( + create_pool_directory(&path).unwrap_err().kind(), + std::io::ErrorKind::AlreadyExists + ); + assert_eq!(fs::symlink_metadata(&path).unwrap().ino(), before.ino()); + assert_eq!(fs::read(path.join("marker")).unwrap(), b"owned fixture"); + fs::remove_file(path.join("marker")).unwrap(); + fs::remove_dir(path).unwrap(); + } +} + +struct Pool { + candidate: Candidate, + owner: Owner, + directory: PathBuf, +} +impl Pool { + fn new() -> Self { + let mut child = Command::new("/bin/sleep").arg("30").spawn().unwrap(); + let mut process = identity::observe(child.id() as i32).unwrap(); + child.kill().unwrap(); + child.wait().unwrap(); + let directory = pool_directory( + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .unwrap() + .as_nanos(), + ); + create_pool_directory(&directory).unwrap(); + let candidate = Candidate::discover(&directory).unwrap(); + let operation = state::Lock::acquire(&root(&candidate)).unwrap(); + let mut owner = Owner::create( + &candidate, + Profile::Development, + None, + NetworkIntent::Isolated, + ) + .unwrap(); + let data = owner.real_data_dir(&candidate).unwrap(); + state::private_directory(&data).unwrap(); + for (name, tag) in [("storage.raw", 1u8), ("overlay.raw", 2)] { + let mut bytes = vec![0u8; 4096]; + bytes[1080..1082].copy_from_slice(&[0x53, 0xef]); + bytes[1128..1144].fill(tag); + bytes[2048..2059].copy_from_slice(b"data-marker"); + fs::write(data.join(name), bytes).unwrap(); + } + fs::write(data.join("vm.lock"), b"").unwrap(); + fs::write(data.join("name"), &owner.machine).unwrap(); + process.executable = binary(&candidate); + owner.process = Some(process); + owner.created = true; + owner.phase = "running".into(); + owner.storage = Some(identity::disk(&data.join("storage.raw")).unwrap()); + owner.overlay = Some(identity::disk(&data.join("overlay.raw")).unwrap()); + owner.save(&candidate).unwrap(); + drop(operation); + Self { + candidate, + owner, + directory, + } + } + fn remove_alias(&self) { + fs::remove_file(&self.owner.short_home).unwrap(); + } + fn receipt(&self) -> Vec { + fs::read(root(&self.candidate).join("owner.json")).unwrap() + } + fn data(&self) -> PathBuf { + self.owner.real_data_dir(&self.candidate).unwrap() + } + fn unchanged(&self, before: &[u8]) { + assert_eq!(self.receipt(), before); + assert!(self.owner.short_home.symlink_metadata().is_err()); + } +} +impl Drop for Pool { + fn drop(&mut self) { + let _ = fs::remove_file(&self.owner.short_home); + fs::remove_dir_all(&self.directory).unwrap(); + } +} + +#[test] +fn explicit_recovery_restores_only_missing_alias_and_preserves_disk_bytes() { + let pool = Pool::new(); + let storage = fs::read(pool.data().join("storage.raw")).unwrap(); + let overlay = fs::read(pool.data().join("overlay.raw")).unwrap(); + pool.remove_alias(); + assert_eq!( + status(&pool.candidate).unwrap_err().code, + "provider_home_missing" + ); + let recovered = crate::provider::recover(&pool.candidate).unwrap(); + assert_eq!(recovered.phase, "recovered-unclean"); + assert_eq!(recovered.process_alive, Some(false)); + assert_eq!( + fs::read_link(&pool.owner.short_home).unwrap(), + root(&pool.candidate).join("home") + ); + assert_eq!(fs::read(pool.data().join("storage.raw")).unwrap(), storage); + assert_eq!(fs::read(pool.data().join("overlay.raw")).unwrap(), overlay); + assert_eq!( + Owner::load(&pool.candidate).unwrap().storage, + pool.owner.storage + ); +} + +#[test] +fn live_or_reused_pid_and_missing_disk_identity_refuse_without_alias_effect() { + for case in 0..3 { + let mut pool = Pool::new(); + if case < 2 { + let p = pool.owner.process.as_mut().unwrap(); + p.pid = std::process::id() as i32; + p.start_micros = if case == 0 { + identity::observe(p.pid).unwrap().start_micros + } else { + 1 + }; + } else { + pool.owner.overlay = None; + } + pool.owner.save(&pool.candidate).unwrap(); + pool.remove_alias(); + let before = pool.receipt(); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap_err().code, + "recovery_required" + ); + pool.unchanged(&before); + } +} + +#[test] +fn changed_disk_held_vm_lock_and_open_disk_refuse_before_alias_creation() { + for case in 0..3 { + let pool = Pool::new(); + pool.remove_alias(); + let before = pool.receipt(); + let mut held = None; + if case == 0 { + fs::write(pool.data().join("overlay.raw"), b"changed disk").unwrap(); + } + if case == 1 { + let f = fs::File::open(pool.data().join("vm.lock")).unwrap(); + assert_eq!( + unsafe { libc::flock(f.as_raw_fd(), libc::LOCK_EX | libc::LOCK_NB) }, + 0 + ); + held = Some(f); + } + if case == 2 { + held = Some(fs::File::open(pool.data().join("storage.raw")).unwrap()); + } + assert!(crate::provider::recover(&pool.candidate).is_err()); + pool.unchanged(&before); + drop(held); + } +} + +#[test] +fn existing_file_directory_and_foreign_symlink_are_never_replaced() { + for case in 0..3 { + let pool = Pool::new(); + pool.remove_alias(); + let before = pool.receipt(); + if case == 0 { + fs::write(&pool.owner.short_home, b"foreign marker").unwrap(); + } + if case == 1 { + fs::create_dir(&pool.owner.short_home).unwrap(); + } + if case == 2 { + std::os::unix::fs::symlink(&pool.directory, &pool.owner.short_home).unwrap(); + } + let inode = fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap_err().code, + "foreign_state" + ); + assert_eq!(pool.receipt(), before); + assert_eq!( + fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(), + inode + ); + if case == 1 { + fs::remove_dir(&pool.owner.short_home).unwrap(); + } + } +} + +#[test] +fn active_socket_leaves_explicit_incomplete_recovery_then_retry_finishes() { + let pool = Pool::new(); + let listener = UnixListener::bind(pool.owner.data_dir().join("agent.sock")).unwrap(); + pool.remove_alias(); + let before = pool.receipt(); + let error = crate::provider::recover(&pool.candidate).unwrap_err(); + assert_eq!(error.code, "provider_home_restored_recovery_incomplete"); + assert!(error.message.contains("stop_uncertain")); + assert_eq!(pool.receipt(), before); + assert_eq!( + fs::read_link(&pool.owner.short_home).unwrap(), + root(&pool.candidate).join("home") + ); + drop(listener); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap().phase, + "recovered-unclean" + ); +} + +#[test] +fn receipt_change_and_concurrent_alias_creation_refuse_exclusive_repair() { + let mut pool = Pool::new(); + pool.remove_alias(); + let observed = Owner::load_for_short_home_recovery(&pool.candidate).unwrap(); + pool.owner.phase = "stopped".into(); + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert_eq!( + observed + .restore_missing_short_home(&pool.candidate) + .unwrap_err() + .code, + "foreign_state" + ); + pool.unchanged(&before); + std::os::unix::fs::symlink(root(&pool.candidate).join("home"), &pool.owner.short_home).unwrap(); + let inode = fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(); + assert_eq!( + pool.owner + .restore_missing_short_home(&pool.candidate) + .unwrap_err() + .code, + "socket_alias_collision" + ); + assert_eq!( + fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(), + inode + ); + assert_eq!(pool.receipt(), before); +} + +#[test] +fn pending_owner_update_is_preserved_and_refuses_alias_repair() { + let pool = Pool::new(); + pool.remove_alias(); + let before = pool.receipt(); + let pending = root(&pool.candidate).join("owner.pending"); + fs::write(&pending, b"interrupted update").unwrap(); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap_err().code, + "recovery_required" + ); + pool.unchanged(&before); + assert_eq!(fs::read(&pending).unwrap(), b"interrupted update"); +} + +#[test] +fn vm_lock_releases_on_success_and_error_even_with_an_inherited_descriptor() { + for fail in [false, true] { + let pool = Pool::new(); + let listener = + fail.then(|| UnixListener::bind(pool.owner.data_dir().join("agent.sock")).unwrap()); + let before = pool.receipt(); + let guard = lock_absent_disks(&pool.candidate, &pool.owner).unwrap(); + // A dup shares the open-file description, exactly as an inherited fork FD does. + let inherited = guard.try_clone().unwrap(); + let independent = fs::File::open(pool.data().join("vm.lock")).unwrap(); + assert_ne!( + unsafe { libc::flock(independent.as_raw_fd(), libc::LOCK_EX | libc::LOCK_NB) }, + 0 + ); + let mut owner = pool.owner.clone(); + let result = finish_absent_locked( + &pool.candidate, + &mut owner, + "recovered-unclean", + false, + &guard, + ); + if fail { + assert_eq!(result.unwrap_err().code, "stop_uncertain"); + assert_eq!(pool.receipt(), before); + } else { + result.unwrap(); + assert_eq!( + Owner::load(&pool.candidate).unwrap().phase, + "recovered-unclean" + ); + } + drop(guard); + assert!(inherited.metadata().is_ok()); + assert_eq!( + unsafe { libc::flock(independent.as_raw_fd(), libc::LOCK_EX | libc::LOCK_NB) }, + 0 + ); + assert_eq!( + unsafe { libc::flock(independent.as_raw_fd(), libc::LOCK_UN) }, + 0 + ); + drop(inherited); + drop(listener); + } +} diff --git a/packages/runtime-core/src/provider/mod.rs b/packages/runtime-core/src/provider/mod.rs index 30f7082bf..3f0b265da 100644 --- a/packages/runtime-core/src/provider/mod.rs +++ b/packages/runtime-core/src/provider/mod.rs @@ -6,6 +6,7 @@ mod bridge; mod config_audit; mod dependency_socket; pub use bridge::BridgeIntent; +pub use dependency_socket::quiescent_recovery as quiescent_dependency_socket_recovery; pub use dependency_socket::recovery as dependency_socket_recovery; pub use dependency_socket::{ DependencySocketIntent, DependencySocketObservation, DependencySocketPath, @@ -19,6 +20,8 @@ pub mod environment; mod environment_probe_test; pub mod environment_recovery; pub mod graph; +#[cfg(target_os = "macos")] +mod host_pin; pub use engine::{EngineInfo, info as engine_info}; #[cfg(all(test, target_os = "macos", target_arch = "aarch64"))] mod gateway_probe_test; @@ -26,6 +29,7 @@ pub mod guest_storage; pub mod host_endpoint; pub mod hostname_authority; pub mod http_probe; +pub mod https_recovery; mod identity; pub mod image_ensure; mod image_load; @@ -42,9 +46,12 @@ pub mod relay_loop; #[cfg(target_os = "macos")] pub mod relay_owner; pub mod resources; +#[cfg(target_os = "macos")] +pub mod shared_https_recovery; pub mod storage_usage; pub use image_load::load as load_image; mod lifecycle; +pub use lifecycle::host_filesystem; mod network_intent; mod network_update; pub use network_update::{enable_internet, extend_network}; @@ -92,10 +99,10 @@ mod socket_requests; mod state; pub use lifecycle::{ PreparedBaseBuilt, build_prepared_base, check_project_share, down, prepared_base_status, - recover, remove_prepared_base, status, up, up_with_bridge, up_with_capabilities, - up_with_minimum_bridges, up_with_network_sockets, up_with_prepared_base, up_with_profile, - up_with_project_share, up_with_retained_project_share, up_with_socket_requests, - up_with_sockets, verify_prepared_base, + probe_with_profile, recover, remove_prepared_base, status, up, up_with_bridge, + up_with_capabilities, up_with_minimum_bridges, up_with_network_sockets, up_with_prepared_base, + up_with_profile, up_with_project_share, up_with_retained_project_share, + up_with_socket_requests, up_with_sockets, verify_prepared_base, }; pub use socket_requests::SocketRequests; pub use source_transfer::{ diff --git a/packages/runtime-core/src/provider/publication.rs b/packages/runtime-core/src/provider/publication.rs index da048d7fd..92451b189 100644 --- a/packages/runtime-core/src/provider/publication.rs +++ b/packages/runtime-core/src/provider/publication.rs @@ -511,6 +511,16 @@ pub fn inspect_claims(c: &Candidate) -> Result Result<(), CandidateError> { + if !load(c, owner)?.is_empty() { + return Err(error()); + } + Ok(()) +} + fn observe_publication(e: &Entry) -> Result<(), CandidateError> { let observed = identity::observe(e.process.pid)?; identity::verify(&e.process, &observed, &e.process.executable, unsafe { diff --git a/packages/runtime-core/src/provider/relay_owner/lifecycle_intent.rs b/packages/runtime-core/src/provider/relay_owner/lifecycle_intent.rs index 8b3fd8708..fb0c61745 100644 --- a/packages/runtime-core/src/provider/relay_owner/lifecycle_intent.rs +++ b/packages/runtime-core/src/provider/relay_owner/lifecycle_intent.rs @@ -12,6 +12,7 @@ use crate::{ }, }; use serde::{Deserialize, Serialize}; +use sha2::{Digest, Sha256}; use std::{ fs::{self, File, OpenOptions}, io::{Read, Write}, @@ -70,6 +71,26 @@ struct Record { ack_required: bool, } impl Record { + fn recovery_fingerprint(&self) -> Result { + let bytes = serde_json::to_vec(&( + "hack-relay-lifecycle-recovery-v1", + self.version, + self.parent, + self.runtime, + self.boot, + self.owner, + &self.process, + self.publication, + self.operation, + self.effect, + &self.targets, + self.graph, + self.selected, + self.attempt, + )) + .map_err(|_| refused())?; + Ok(format!("{:x}", Sha256::digest(bytes))) + } fn validate(&self, parent: (u64, u64)) -> Result<(), CandidateError> { if ![1, 2].contains(&self.version) || self.owner == [0; 16] @@ -314,6 +335,10 @@ pub struct Inspection { pub graph: Option, pub selection_observed: bool, pub acknowledgement_pending: bool, + pub owner: [u8; 16], + pub process: ProcessIdentity, + pub publication: [u8; 32], + pub recovery_fingerprint: String, } impl Inspection { /// Requires an existing journal and lock; never initializes absent runtime state. @@ -339,6 +364,10 @@ impl Inspection { .transpose()?, selection_observed: record.selected.is_some(), acknowledgement_pending: record.ack_required, + owner: record.owner, + process: record.process.clone(), + publication: record.publication, + recovery_fingerprint: record.recovery_fingerprint()?, }) } } @@ -492,6 +521,64 @@ impl Coordinator { pub fn acknowledgement_pending(&self) -> bool { self.snapshot.record.ack_required } + /// A separate selected recovery may inspect an interrupted graph effect while + /// holding this coordinator lock. This grants no replay of `execute`. + pub fn verify_selected_effect( + &self, + selection: &Selection, + graph: GraphScope, + owner: [u8; 16], + process: &ProcessIdentity, + publication: [u8; 32], + recovery_fingerprint: &str, + ) -> Result<(), CandidateError> { + self.verify_recovery_identity(selection, graph, recovery_fingerprint)?; + self.store.verify(&self.snapshot)?; + let record = &self.snapshot.record; + if record.phase != Phase::EffectStarted + || !record.ack_required + || record.selected.is_none() + || record.owner != owner + || &record.process != process + || record.publication != publication + || record.recovery_fingerprint()? != recovery_fingerprint + || record.runtime != selection.context.runtime + || record.boot != selection.context.boot + || record.operation != selection.operation + || record.effect != selection.effect + || record.graph != Some(graph.id) + || selection.context != graph.context + { + return Err(refused()); + } + Ok(()) + } + /// Exact selected identity across EffectStarted -> Confirmed -> caller ACK. + /// The fingerprint binds owner/process/publication/attempt and targets while + /// phase, observation and ACK bit may change only under this held store lock. + pub fn verify_recovery_identity( + &self, + selection: &Selection, + graph: GraphScope, + expected_fingerprint: &str, + ) -> Result<(), CandidateError> { + self.store.verify(&self.snapshot)?; + let record = &self.snapshot.record; + if !matches!(record.phase, Phase::EffectStarted | Phase::Confirmed) + || !record.ack_required + || record.selected.is_none() + || record.runtime != selection.context.runtime + || record.boot != selection.context.boot + || record.operation != selection.operation + || record.effect != selection.effect + || record.graph != Some(graph.id) + || selection.context != graph.context + || record.recovery_fingerprint()? != expected_fingerprint + { + return Err(refused()); + } + Ok(()) + } /// Acknowledge durable caller confirmation before allowing the single coordinator /// record to roll over. This does not inspect caller metadata; the caller must /// hold the VM lease and first persist its matching confirmed enrollment marker. diff --git a/packages/runtime-core/src/provider/relay_owner/lifecycle_intent/tests.rs b/packages/runtime-core/src/provider/relay_owner/lifecycle_intent/tests.rs index 96c0b0d5c..b92253946 100644 --- a/packages/runtime-core/src/provider/relay_owner/lifecycle_intent/tests.rs +++ b/packages/runtime-core/src/provider/relay_owner/lifecycle_intent/tests.rs @@ -809,6 +809,98 @@ fn enrolled_confirmation_requires_caller_acknowledgement_before_any_rollover() { .prepare_graph(&server.endpoint, Duration::from_secs(2)) .unwrap(); coordinator.execute(EFFECT, || Ok(())).unwrap(); + let selected_effect = selection(&coordinator); + let selected_owner = server.endpoint.incarnation(); + let selected_process = server.endpoint.process().clone(); + let selected_publication = server.endpoint.fingerprint(); + let selected_identity = coordinator.snapshot.record.recovery_fingerprint().unwrap(); + coordinator + .verify_selected_effect( + &selected_effect, + scope, + selected_owner, + &selected_process, + selected_publication, + &selected_identity, + ) + .unwrap(); + assert!( + coordinator + .verify_selected_effect( + &selected_effect, + GraphScope::new(context(), [7; 32]).unwrap(), + selected_owner, + &selected_process, + selected_publication, + &selected_identity + ) + .is_err() + ); + assert!( + coordinator + .verify_selected_effect( + &Selection { + effect: [8; 32], + ..selected_effect + }, + scope, + selected_owner, + &selected_process, + selected_publication, + &selected_identity + ) + .is_err() + ); + assert!( + coordinator + .verify_selected_effect( + &selection(&coordinator), + scope, + [3; 16], + &selected_process, + selected_publication, + &selected_identity + ) + .is_err() + ); + let mut other_process = selected_process.clone(); + other_process.pid += 1; + assert!( + coordinator + .verify_selected_effect( + &selection(&coordinator), + scope, + selected_owner, + &other_process, + selected_publication, + &selected_identity + ) + .is_err() + ); + assert!( + coordinator + .verify_selected_effect( + &selection(&coordinator), + scope, + selected_owner, + &selected_process, + [4; 32], + &selected_identity + ) + .is_err() + ); + assert!( + coordinator + .verify_selected_effect( + &selection(&coordinator), + scope, + selected_owner, + &selected_process, + selected_publication, + &"0".repeat(64) + ) + .is_err() + ); assert!(coordinator.acknowledge().is_err()); coordinator.confirm(EFFECT, || Ok([7; 32])).unwrap(); assert!(coordinator.acknowledgement_pending()); @@ -817,6 +909,7 @@ fn enrolled_confirmation_requires_caller_acknowledgement_before_any_rollover() { let inspection = Inspection::load(&root.0, context()).unwrap(); assert_eq!(inspection.phase, Phase::Confirmed); assert!(inspection.acknowledgement_pending); + assert_eq!(inspection.recovery_fingerprint, selected_identity); assert!( registration(&root.0).is_ok(), "ordinary registration keeps its existing contract" @@ -834,8 +927,20 @@ fn enrolled_confirmation_requires_caller_acknowledgement_before_any_rollover() { .is_err() ); let mut recovered = Coordinator::resume(&root.0, selected).unwrap(); + recovered + .verify_recovery_identity(&selection(&recovered), scope, &selected_identity) + .unwrap(); recovered.acknowledge().unwrap(); assert!(!recovered.acknowledgement_pending()); + assert_eq!( + recovered.snapshot.record.recovery_fingerprint().unwrap(), + selected_identity + ); + assert!( + recovered + .verify_recovery_identity(&selection(&recovered), scope, &selected_identity) + .is_err() + ); let identity = recovered.snapshot.identity; recovered.acknowledge().unwrap(); assert_eq!( @@ -862,6 +967,41 @@ fn enrolled_confirmation_requires_caller_acknowledgement_before_any_rollover() { assert!(Coordinator::begin(&server.endpoint, server.mutation()).is_ok()); } +#[test] +fn selected_effect_refuses_same_effect_replaced_coordinator_record() { + let root = Root::new(); + let server = Server::scoped(&root, Some(GraphScope::new(context(), [8; 32]).unwrap())); + let scope = GraphScope::new(context(), [9; 32]).unwrap(); + let mut coordinator = + Coordinator::begin_graph_enrolled(&server.endpoint, scope, EFFECT, Some([6; 16]), |_| { + Ok(()) + }) + .unwrap(); + coordinator + .prepare_graph(&server.endpoint, Duration::from_secs(2)) + .unwrap(); + coordinator.execute(EFFECT, || Ok(())).unwrap(); + let selected = selection(&coordinator); + let identity = coordinator.snapshot.record.recovery_fingerprint().unwrap(); + let mut replacement = coordinator.snapshot.record.clone(); + replacement.owner = [7; 16]; + assert_eq!(replacement.effect, coordinator.snapshot.record.effect); + assert_ne!(replacement.recovery_fingerprint().unwrap(), identity); + state::write(&root.0.join("relay-lifecycle/state.json"), &replacement).unwrap(); + assert!( + coordinator + .verify_selected_effect( + &selected, + scope, + server.endpoint.incarnation(), + server.endpoint.process(), + server.endpoint.fingerprint(), + &identity, + ) + .is_err() + ); +} + #[test] fn legacy_record_cannot_claim_enrollment_acknowledgement() { let root = Root::new(); diff --git a/packages/runtime-core/src/provider/relay_owner/publication.rs b/packages/runtime-core/src/provider/relay_owner/publication.rs index 2210b12ec..178e0a97d 100644 --- a/packages/runtime-core/src/provider/relay_owner/publication.rs +++ b/packages/runtime-core/src/provider/relay_owner/publication.rs @@ -7,6 +7,7 @@ use super::{ use crate::{ CandidateError, provider::{ + host_pin::DeviceRebind, identity::{self, ProcessIdentity}, state, }, @@ -102,6 +103,16 @@ pub struct PinnedEndpoint { receipt: Receipt, receipt_id: FileId, bytes: Vec, + device_rebind: Option, +} +#[derive(Clone, Debug, Eq, PartialEq, Deserialize, Serialize)] +#[serde(deny_unknown_fields)] +pub(crate) struct LegacyControl { + pub receipt_sha256: String, + pub receipt_id: FileId, + pub parent: FileId, + pub endpoint: FileId, + pub process: ProcessIdentity, } impl PinnedEndpoint { pub(crate) fn runtime_root(&self) -> &Path { @@ -134,6 +145,13 @@ impl PinnedEndpoint { Self::read(Paths::new(root)?, context) } fn read(paths: Paths, context: Context) -> Result { + Self::read_with_rebind(paths, context, None) + } + fn read_with_rebind( + paths: Paths, + context: Context, + device_rebind: Option, + ) -> Result { let parent = paths.parent()?; let mut file = OpenOptions::new() .read(true) @@ -160,7 +178,9 @@ impl PinnedEndpoint { || context.boot == [0; 16] || receipt.owner == [0; 16] || receipt.socket != paths.socket - || receipt.parent != parent + || !device_rebind.map_or(receipt.parent == parent, |v| { + v.matches(receipt.parent, parent) + }) || receipt.endpoint.1 == 0 || !receipt.process.executable.is_absolute() { @@ -179,8 +199,30 @@ impl PinnedEndpoint { receipt, receipt_id: id(&m), bytes, + device_rebind, }) } + pub(crate) fn load_legacy_recovery( + root: &Path, + context: Context, + device_rebind: DeviceRebind, + host_boot_micros: u64, + ) -> Result { + let pin = Self::read_with_rebind(Paths::new(root)?, context, Some(device_rebind))?; + pin.verify_dead()?; + device_rebind.definitely_dead_before_boot(&pin.receipt.process, host_boot_micros)?; + dead::no_listener(&pin.paths.socket)?; + Ok(pin) + } + pub(crate) fn legacy_summary(&self) -> LegacyControl { + LegacyControl { + receipt_sha256: format!("{:x}", Sha256::digest(&self.bytes)), + receipt_id: self.receipt_id, + parent: self.receipt.parent, + endpoint: self.receipt.endpoint, + process: self.receipt.process.clone(), + } + } /// Read-only predecessor proof: exact private receipt/socket, absent owner. #[cfg(target_os = "macos")] pub(crate) fn verify_dead(&self) -> Result<(), CandidateError> { @@ -198,12 +240,13 @@ impl PinnedEndpoint { Sha256::digest(&self.bytes).into() } fn verify_receipt(&self) -> Result<(), CandidateError> { - let current = Self::read( + let current = Self::read_with_rebind( self.paths.clone(), Context { runtime: self.receipt.runtime, boot: self.receipt.boot, }, + self.device_rebind, )?; if current.receipt_id != self.receipt_id || current.bytes != self.bytes { return Err(refused()); @@ -211,14 +254,24 @@ impl PinnedEndpoint { Ok(()) } fn verify_socket(&self) -> Result<(), CandidateError> { - if self.paths.parent()? != self.receipt.parent { + let parent = self.paths.parent()?; + if !self + .device_rebind + .map_or(parent == self.receipt.parent, |v| { + v.matches(self.receipt.parent, parent) + }) + { return Err(refused()); } let m = fs::symlink_metadata(&self.paths.socket).map_err(|_| refused())?; if !m.file_type().is_socket() || !private(&m) || m.nlink() != 1 - || id(&m) != self.receipt.endpoint + || !self + .device_rebind + .map_or(id(&m) == self.receipt.endpoint, |v| { + v.matches(self.receipt.endpoint, id(&m)) + }) { return Err(refused()); } @@ -353,6 +406,7 @@ impl ControlListener { receipt, receipt_id, bytes, + device_rebind: None, }; endpoint.verify_receipt()?; endpoint.verify_socket()?; diff --git a/packages/runtime-core/src/provider/relay_owner/publication/dead.rs b/packages/runtime-core/src/provider/relay_owner/publication/dead.rs index edec289f5..fe759f0b1 100644 --- a/packages/runtime-core/src/provider/relay_owner/publication/dead.rs +++ b/packages/runtime-core/src/provider/relay_owner/publication/dead.rs @@ -41,6 +41,7 @@ impl Selection { receipt, receipt_id: self.record_id, bytes: self.bytes.clone(), + device_rebind: None, }) } } @@ -49,6 +50,115 @@ pub(crate) struct Witness { pin: PinnedEndpoint, lock: state::Lock, } +/// A clean managed owner exit unlinks both publication paths while leaving the +/// operation lock. Absence is only a current endpoint fact: callers must bind +/// the original owner/process/publication to a selected coordinator record. +pub(crate) enum CleanupWitness { + Present(Witness), + Absent(AbsentWitness), +} + +pub(crate) struct AbsentWitness { + paths: Paths, + lock: state::Lock, + parent: FileId, + process: ProcessIdentity, +} + +#[derive(Serialize)] +struct AbsentSelection<'a> { + kind: &'static str, + parent: FileId, + lock: FileId, + process: &'a ProcessIdentity, +} + +impl CleanupWitness { + pub(crate) fn acquire( + root: &Path, + context: Context, + process: &ProcessIdentity, + ) -> Result { + let paths = Paths::new(root)?; + let lock = state::Lock::acquire_existing(&paths.directory)?; + match (absent(&paths.receipt)?, absent(&paths.socket)?) { + (false, false) => { + let pin = PinnedEndpoint::read(paths, context)?; + if &pin.receipt.process != process { + return Err(refused()); + } + let witness = Witness { pin, lock }; + witness.verify()?; + Ok(Self::Present(witness)) + } + (true, true) if context.runtime != [0; 16] && context.boot != [0; 16] => { + // SAFETY: geteuid takes no arguments. + identity::verify(process, process, &process.executable, unsafe { + libc::geteuid() + })?; + let parent = paths.parent()?; + let witness = AbsentWitness { + paths, + lock, + parent, + process: process.clone(), + }; + witness.verify()?; + Ok(Self::Absent(witness)) + } + _ => Err(refused()), + } + } + + pub(crate) fn selection_sha256(&self) -> Result { + let bytes = match self { + Self::Present(witness) => { + serde_json::to_vec(&witness.selection()).map_err(|_| refused())? + } + Self::Absent(witness) => serde_json::to_vec(&AbsentSelection { + kind: "absent", + parent: witness.parent, + lock: witness.lock.identity()?, + process: &witness.process, + }) + .map_err(|_| refused())?, + }; + Ok(format!("{:x}", Sha256::digest(bytes))) + } + + pub(crate) fn present_identity(&self) -> Option<([u8; 16], [u8; 32])> { + match self { + Self::Present(witness) => Some((witness.owner(), witness.publication())), + Self::Absent(_) => None, + } + } + + pub(crate) fn verify(&self) -> Result<(), CandidateError> { + match self { + Self::Present(witness) => witness.verify(), + Self::Absent(witness) => witness.verify(), + } + } +} + +impl AbsentWitness { + fn verify(&self) -> Result<(), CandidateError> { + let lock = fs::symlink_metadata(self.paths.directory.join("operation.lock")) + .map_err(|_| refused())?; + if self.paths.parent()? != self.parent + || !lock.is_file() + || lock.nlink() != 1 + || !private(&lock) + || id(&lock) != self.lock.identity()? + || !absent(&self.paths.receipt)? + || !absent(&self.paths.socket)? + || identity::alive(self.process.pid)? + { + return Err(refused()); + } + Ok(()) + } +} impl Witness { pub(crate) fn acquire( root: &Path, @@ -71,6 +181,12 @@ impl Witness { record_id: self.pin.receipt_id, } } + pub(crate) fn owner(&self) -> [u8; 16] { + self.pin.incarnation() + } + pub(crate) fn publication(&self) -> [u8; 32] { + self.pin.fingerprint() + } pub(crate) fn verify(&self) -> Result<(), CandidateError> { let m = fs::symlink_metadata(self.pin.paths.directory.join("operation.lock")) .map_err(|_| refused())?; diff --git a/packages/runtime-core/src/provider/relay_owner/publication/tests.rs b/packages/runtime-core/src/provider/relay_owner/publication/tests.rs index 8cf51ccab..067964d1b 100644 --- a/packages/runtime-core/src/provider/relay_owner/publication/tests.rs +++ b/packages/runtime-core/src/provider/relay_owner/publication/tests.rs @@ -97,6 +97,8 @@ fn private_named_endpoint_retires_and_cleans_only_owned_publication() { assert!(absent(&paths(&root).socket).unwrap()); assert!(absent(&paths(&root).receipt).unwrap()); assert!(paths(&root).directory.join("operation.lock").is_file()); + #[cfg(target_os = "macos")] + assert!(dead::CleanupWitness::acquire(&root.0, context(), &pin.receipt.process).is_err()); pin.recover().unwrap(); } #[test] @@ -252,7 +254,7 @@ fn publication_child() { return; } let mut byte = [0]; - std::io::stdin().read_exact(&mut byte).unwrap(); + let _ = std::io::stdin().read(&mut byte).unwrap(); } fn spawn_child(root: &Root, backlog: bool) -> OwnedChild { OwnedChild( @@ -406,3 +408,64 @@ fn dead_witness_rejects_replacement_listener_and_changed_lock() { assert!(dead::Witness::acquire(&root.0, context(), &pin.receipt.process).is_err()); drop(replacement); } + +#[cfg(target_os = "macos")] +#[test] +fn normal_owner_exit_leaves_an_exact_absent_pair_under_its_operation_lock() { + let root = Root::new(); + let mut child = spawn_child(&root, false); + let deadline = Instant::now() + Duration::from_secs(3); + let pin = loop { + if let Ok(pin) = PinnedEndpoint::load(&root.0, context()) { + break pin; + } + assert!(Instant::now() < deadline); + assert!(child.0.try_wait().unwrap().is_none()); + std::thread::sleep(Duration::from_millis(1)); + }; + assert!(dead::CleanupWitness::acquire(&root.0, context(), &pin.receipt.process).is_err()); + drop(child.0.stdin.take()); + assert!(child.0.wait().unwrap().success()); + assert!(absent(&pin.paths.receipt).unwrap()); + assert!(absent(&pin.paths.socket).unwrap()); + assert!(pin.paths.directory.join("operation.lock").is_file()); + + let witness = dead::CleanupWitness::acquire(&root.0, context(), &pin.receipt.process).unwrap(); + assert!(matches!(&witness, dead::CleanupWitness::Absent(_))); + let selected = witness.selection_sha256().unwrap(); + assert!(witness.verify().is_ok()); + let mut wrong_process = pin.receipt.process.clone(); + wrong_process.start_micros += 1; + drop(witness); + let changed = dead::CleanupWitness::acquire(&root.0, context(), &wrong_process).unwrap(); + assert_ne!(changed.selection_sha256().unwrap(), selected); + drop(changed); + + fs::write(&pin.paths.receipt, b"foreign").unwrap(); + assert!(dead::CleanupWitness::acquire(&root.0, context(), &pin.receipt.process).is_err()); + fs::remove_file(&pin.paths.receipt).unwrap(); + let witness = dead::CleanupWitness::acquire(&root.0, context(), &pin.receipt.process).unwrap(); + fs::write(&pin.paths.socket, b"foreign").unwrap(); + assert!(witness.verify().is_err()); + drop(witness); + fs::remove_file(&pin.paths.socket).unwrap(); + + let witness = dead::CleanupWitness::acquire(&root.0, context(), &pin.receipt.process).unwrap(); + let lock = pin.paths.directory.join("operation.lock"); + let linked = root.0.join("linked.lock"); + fs::hard_link(&lock, &linked).unwrap(); + assert!(witness.verify().is_err()); + fs::remove_file(&linked).unwrap(); + assert!(witness.verify().is_ok()); + fs::rename(&lock, pin.paths.directory.join("old.lock")).unwrap(); + let replacement = state::Lock::acquire(&pin.paths.directory).unwrap(); + assert!(witness.verify().is_err()); + drop(replacement); + drop(witness); + + let witness = dead::CleanupWitness::acquire(&root.0, context(), &pin.receipt.process).unwrap(); + fs::rename(&pin.paths.directory, root.0.join("old-control")).unwrap(); + state::private_directory(&pin.paths.directory).unwrap(); + let _replacement = state::Lock::acquire(&pin.paths.directory).unwrap(); + assert!(witness.verify().is_err()); +} diff --git a/packages/runtime-core/src/provider/shared_https_recovery.rs b/packages/runtime-core/src/provider/shared_https_recovery.rs new file mode 100644 index 000000000..620a9649a --- /dev/null +++ b/packages/runtime-core/src/provider/shared_https_recovery.rs @@ -0,0 +1,855 @@ +//! Explicit no-signal archival of one dead, prior-boot shared HTTPS owner. +//! The endpoint v1 record has no PID/birth; complete executable absence and a +//! refused exact socket are quiescence observations, never process adoption. +use super::{graph::HttpsArchiveGuard, https_recovery, identity, state}; +use crate::{Candidate, CandidateError}; +use serde::{Deserialize, Serialize}; +use sha2::{Digest, Sha256}; +use std::{ + fs::{self, File, OpenOptions}, + io::{Read, Write}, + os::unix::{ + fs::{DirBuilderExt, FileTypeExt, MetadataExt, OpenOptionsExt}, + net::UnixStream, + }, + path::{Path, PathBuf}, +}; + +const HEX32: usize = 32; +const MAX_FILE: u64 = 16_384; +fn refused() -> CandidateError { + CandidateError::new( + "shared_https_previous_boot_recovery", + "Exact previous-boot shared HTTPS ownership or quiescence is unproven; retained evidence was preserved.", + ) +} +fn hex(value: &str, len: usize) -> bool { + value.len() == len + && value + .bytes() + .all(|b| b.is_ascii_digit() || (b'a'..=b'f').contains(&b)) +} +fn digest(bytes: &[u8]) -> String { + format!("{:x}", Sha256::digest(bytes)) +} +fn absent(path: &Path) -> Result { + match fs::symlink_metadata(path) { + Ok(_) => Ok(false), + Err(e) if e.kind() == std::io::ErrorKind::NotFound => Ok(true), + Err(_) => Err(refused()), + } +} +#[derive(Clone, Copy, Debug, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct Inode { + dev: u64, + ino: u64, +} +fn inode(m: &fs::Metadata) -> Inode { + Inode { + dev: m.dev(), + ino: m.ino(), + } +} +fn dir(path: &Path) -> Result { + let m = fs::symlink_metadata(path).map_err(|_| refused())?; + if !m.is_dir() + || m.uid() != unsafe { libc::geteuid() } + || m.mode() & 0o777 != 0o700 + || fs::canonicalize(path).map_err(|_| refused())? != path + { + return Err(refused()); + } + Ok(inode(&m)) +} +fn exact_entries(path: &Path, expected: &[&str]) -> Result<(), CandidateError> { + let mut got = fs::read_dir(path) + .map_err(|_| refused())? + .map(|e| { + e.map_err(|_| refused()) + .and_then(|e| e.file_name().into_string().map_err(|_| refused())) + }) + .collect::, _>>()?; + got.sort(); + let mut want = expected.iter().map(|s| (*s).to_owned()).collect::>(); + want.sort(); + if got != want { + return Err(refused()); + } + Ok(()) +} +#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct FilePin { + inode: Inode, + sha256: String, +} +fn read_file(path: &Path, limit: u64, private: bool) -> Result<(Vec, FilePin), CandidateError> { + let mut file = OpenOptions::new() + .read(true) + .custom_flags(libc::O_NOFOLLOW | libc::O_NONBLOCK) + .open(path) + .map_err(|_| refused())?; + let before = file.metadata().map_err(|_| refused())?; + if !before.is_file() + || before.nlink() != 1 + || before.uid() != unsafe { libc::geteuid() } + || (private && before.mode() & 0o777 != 0o600) + || (!private && before.mode() & 0o022 != 0) + || before.len() == 0 + || before.len() > limit + { + return Err(refused()); + } + let mut bytes = Vec::new(); + (&mut file) + .take(limit + 1) + .read_to_end(&mut bytes) + .map_err(|_| refused())?; + let after = file.metadata().map_err(|_| refused())?; + if bytes.len() as u64 != before.len() + || inode(&before) != inode(&after) + || before.mtime() != after.mtime() + || before.mtime_nsec() != after.mtime_nsec() + || before.ctime() != after.ctime() + || before.ctime_nsec() != after.ctime_nsec() + || fs::canonicalize(path).map_err(|_| refused())? != path + { + return Err(refused()); + } + Ok(( + bytes.clone(), + FilePin { + inode: inode(&before), + sha256: digest(&bytes), + }, + )) +} +fn pinned_file(path: &Path, expected: &FilePin, private: bool) -> Result, CandidateError> { + let (bytes, actual) = read_file(path, MAX_FILE, private)?; + if actual != *expected { + return Err(refused()); + } + Ok(bytes) +} +#[derive(Clone, Debug, Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Runtime { + binary: PathBuf, + home: PathBuf, +} +#[derive(Clone, Debug, Deserialize)] +#[serde(deny_unknown_fields)] +struct Binary { + binary: PathBuf, + sha256: String, +} +#[derive(Clone, Debug, Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Pool { + owner: String, + boot_id: String, +} +#[derive(Clone, Debug, Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Binding { + runtime: Runtime, + frontend: Binary, + runtime_sha256: String, + pool: Pool, + caddy_binary: PathBuf, + caddy_sha256: String, + https_port: u16, + certificate_name_limit: u16, +} +#[derive(Clone, Debug, Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Configuration { + version: u8, + owner_generation: String, + binding: Binding, +} +#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Lease { + version: u8, + owner_generation: String, + lease_id: String, + owner: String, + run: String, + attempt: String, + namespace: String, + plan_id: String, +} +#[derive(Clone, Debug, Deserialize)] +#[serde(deny_unknown_fields, rename_all = "camelCase")] +struct Endpoint { + version: u8, + owner_generation: String, + socket: PathBuf, + dev: u64, + ino: u64, +} +#[derive(Clone, Debug, Serialize, Deserialize)] +#[serde(deny_unknown_fields)] +struct Journal { + version: u8, + run: String, + owner_generation: String, + lease_id: String, + old_boot: String, + current_boot: String, + recovery_process: identity::ProcessIdentity, + admission: Inode, + admission_lock: Inode, + root: Inode, + leases: Inode, + configuration: FilePin, + endpoint: FilePin, + lease: FilePin, + socket: SocketEvidence, + ca: FilePin, +} +#[derive(Clone, Debug, Serialize, Deserialize)] +#[serde(deny_unknown_fields, tag = "kind", rename_all = "snake_case")] +enum SocketEvidence { + Present { + path: PathBuf, + parent: Inode, + socket: Inode, + }, + Absent { + path: PathBuf, + }, +} +impl SocketEvidence { + fn path(&self) -> &Path { + match self { + Self::Present { path, .. } | Self::Absent { path } => path, + } + } +} +struct Paths { + storage: PathBuf, + source: PathBuf, + archive: PathBuf, + intent: PathBuf, + complete: PathBuf, + admission: PathBuf, +} +impl Paths { + fn new(home: &Path, generation: &str, lease_id: &str) -> Self { + let storage = home.join("native-https"); + let stem = format!("{generation}-{lease_id}"); + Self { + source: storage.join("shared-owner"), + archive: storage.join(format!("archived-previous-boot-shared-owner-{stem}")), + intent: storage.join(format!("previous-boot-shared-owner-{stem}.intent.json")), + complete: storage.join(format!("previous-boot-shared-owner-{stem}.complete.json")), + admission: storage.join("shared-owner-admission.lock"), + storage, + } + } +} +fn selected<'a>( + source: &'a Path, + archive: &'a Path, + expected: Inode, + directory: bool, +) -> Result<&'a Path, CandidateError> { + let path = match (absent(source)?, absent(archive)?) { + (false, true) => source, + (true, false) => archive, + _ => return Err(refused()), + }; + let m = fs::symlink_metadata(path).map_err(|_| refused())?; + if inode(&m) != expected + || m.uid() != unsafe { libc::geteuid() } + || if directory { + !m.is_dir() || m.mode() & 0o777 != 0o700 + } else { + !m.file_type().is_socket() || m.mode() & 0o777 != 0o600 + } + { + return Err(refused()); + } + Ok(path) +} +fn socket_archive(path: &Path, journal: &Journal) -> Result { + if path.file_name().and_then(|n| n.to_str()) != Some("control.sock") { + return Err(refused()); + } + let archived = path.with_file_name(format!("s{}", &journal.lease_id[..8])); + if archived.as_os_str().len() >= 100 { + return Err(refused()); + } + Ok(archived) +} +fn socket_dead(path: &Path) -> Result<(), CandidateError> { + match UnixStream::connect(path) { + Err(e) if e.kind() == std::io::ErrorKind::ConnectionRefused => Ok(()), + _ => Err(refused()), + } +} +fn socket_parent(path: &Path) -> Result<&Path, CandidateError> { + let parent = path.parent().ok_or_else(refused)?; + if path.file_name().and_then(|s| s.to_str()) != Some("control.sock") + || path.as_os_str().len() >= 100 + || parent.parent() != Some(Path::new("/private/tmp")) + || parent + .file_name() + .and_then(|s| s.to_str()) + .is_none_or(|s| !s.starts_with("hk-https-leases-")) + { + return Err(refused()); + } + Ok(parent) +} +fn sync_dir(path: &Path, expected: Inode) -> Result<(), CandidateError> { + if dir(path)? != expected { + return Err(refused()); + } + File::open(path) + .and_then(|f| f.sync_all()) + .map_err(|_| refused()) +} +#[cfg(target_os = "macos")] +fn rename_exclusive(from: &Path, to: &Path) -> Result<(), CandidateError> { + use std::os::unix::ffi::OsStrExt; + let from = std::ffi::CString::new(from.as_os_str().as_bytes()).map_err(|_| refused())?; + let to = std::ffi::CString::new(to.as_os_str().as_bytes()).map_err(|_| refused())?; + if unsafe { libc::renamex_np(from.as_ptr(), to.as_ptr(), libc::RENAME_EXCL) } != 0 { + return Err(refused()); + } + Ok(()) +} +fn write_new(path: &Path, value: &impl Serialize, parent: Inode) -> Result<(), CandidateError> { + let bytes = serde_json::to_vec(value).map_err(|_| refused())?; + let mut file = OpenOptions::new() + .write(true) + .create_new(true) + .mode(0o600) + .custom_flags(libc::O_NOFOLLOW) + .open(path) + .map_err(|_| refused())?; + file.write_all(&bytes) + .and_then(|_| file.sync_all()) + .map_err(|_| refused())?; + sync_dir(path.parent().ok_or_else(refused)?, parent) +} +fn lease_id(config: &Configuration, lease: &Lease) -> Result { + let parts = serde_json::json!([ + "native-https-lease-v1", + config.owner_generation, + config.binding.pool.owner, + lease.run, + lease.attempt, + lease.namespace, + lease.plan_id, + ]); + let bytes = serde_json::to_vec(&parts).map_err(|_| refused())?; + Ok(digest(&bytes)[..32].into()) +} +fn validate_config( + config: &Configuration, + lease: &Lease, + endpoint: &Endpoint, + home: &Path, + generation: &str, + expected_lease: &str, +) -> Result<(), CandidateError> { + if config.version != 1 + || lease.version != 1 + || endpoint.version != 1 + || config.owner_generation != generation + || lease.owner_generation != generation + || endpoint.owner_generation != generation + || lease.lease_id != expected_lease + || lease_id(config, lease)? != expected_lease + || config.binding.runtime.home != home + || config.binding.pool.owner != lease.owner + || config.binding.https_port == 0 + || !(1..=4096).contains(&config.binding.certificate_name_limit) + || !hex(&lease.run, HEX32) + || !hex(&lease.attempt, HEX32) + || !hex(&lease.owner, HEX32) + || !hex(&lease.namespace, 64) + || !hex(&lease.plan_id, 64) + || !hex(&config.binding.runtime_sha256, 64) + || !hex(&config.binding.frontend.sha256, 64) + || !hex(&config.binding.caddy_sha256, 64) + || [ + &config.binding.runtime.binary, + &config.binding.frontend.binary, + &config.binding.caddy_binary, + ] + .iter() + .any(|p| !p.is_absolute()) + { + return Err(refused()); + } + Ok(()) +} +fn selected_evidence( + paths: &Paths, + journal: &Journal, +) -> Result<(Configuration, Lease, Endpoint, Option), CandidateError> { + let root = selected(&paths.source, &paths.archive, journal.root, true)?; + exact_entries(root, &["configuration.json", "endpoint.json", "leases"])?; + if dir(&root.join("leases"))? != journal.leases { + return Err(refused()); + } + exact_entries( + &root.join("leases"), + &[&format!("{}.json", journal.lease_id)], + )?; + let config: Configuration = serde_json::from_slice(&pinned_file( + &root.join("configuration.json"), + &journal.configuration, + true, + )?) + .map_err(|_| refused())?; + let lease: Lease = serde_json::from_slice(&pinned_file( + &root + .join("leases") + .join(format!("{}.json", journal.lease_id)), + &journal.lease, + true, + )?) + .map_err(|_| refused())?; + let endpoint: Endpoint = serde_json::from_slice(&pinned_file( + &root.join("endpoint.json"), + &journal.endpoint, + true, + )?) + .map_err(|_| refused())?; + validate_config( + &config, + &lease, + &endpoint, + &config.binding.runtime.home, + &journal.owner_generation, + &journal.lease_id, + )?; + if endpoint.socket != journal.socket.path() { + return Err(refused()); + } + socket_parent(&endpoint.socket)?; + let socket_parent = endpoint.socket.parent().ok_or_else(refused)?; + let selected_socket = match &journal.socket { + SocketEvidence::Present { parent, socket, .. } => { + if endpoint.dev != socket.dev + || endpoint.ino != socket.ino + || dir(socket_parent)? != *parent + { + return Err(refused()); + } + let retired_socket = socket_archive(&endpoint.socket, journal)?; + let selected = selected(&endpoint.socket, &retired_socket, *socket, false)?; + exact_entries( + socket_parent, + &[selected + .file_name() + .and_then(|n| n.to_str()) + .ok_or_else(refused)?], + )?; + Some(selected.to_path_buf()) + } + SocketEvidence::Absent { .. } => { + if !absent(socket_parent)? || !absent(&endpoint.socket)? { + return Err(refused()); + } + None + } + }; + let ca = paths + .storage + .join("data/caddy/pki/authorities/local/root.crt"); + pinned_file(&ca, &journal.ca, false)?; + Ok((config, lease, endpoint, selected_socket)) +} +fn prove_quiescence(config: &Configuration, socket: Option<&Path>) -> Result<(), CandidateError> { + let self_binary = + fs::canonicalize(std::env::current_exe().map_err(|_| refused())?).map_err(|_| refused())?; + for (path, hash) in [ + ( + &config.binding.frontend.binary, + &config.binding.frontend.sha256, + ), + ( + &config.binding.runtime.binary, + &config.binding.runtime_sha256, + ), + (&config.binding.caddy_binary, &config.binding.caddy_sha256), + ] { + if https_recovery::executable_hash(path)? != *hash { + return Err(refused()); + } + } + for binary in [ + &config.binding.frontend.binary, + &config.binding.runtime.binary, + &config.binding.caddy_binary, + ] { + if fs::canonicalize(binary).map_err(|_| refused())? == self_binary + || identity::executable_running(binary)? + { + return Err(refused()); + } + } + if let Some(socket) = socket { + socket_dead(socket)?; + } + Ok(()) +} + +pub struct ArchiveSelection<'a> { + pub run: &'a str, + pub owner_generation: &'a str, + pub lease_id: &'a str, + pub attempt: &'a str, + pub owner: &'a str, + pub namespace: &'a str, + pub plan: &'a str, +} + +fn acquire_admission(paths: &Paths) -> Result<(bool, Inode, state::Lock, Inode), CandidateError> { + let fresh = match fs::DirBuilder::new().mode(0o700).create(&paths.admission) { + Ok(()) => true, + Err(e) if e.kind() == std::io::ErrorKind::AlreadyExists => false, + Err(_) => return Err(refused()), + }; + let admission_id = dir(&paths.admission)?; + // An existing empty admission can belong to a normal ensure. Never create + // its lock or infer ownership from an older completed archive. + if !fresh && absent(&paths.intent)? && absent(&paths.complete)? { + return Err(refused()); + } + let admission_lock = if fresh { + state::Lock::acquire(&paths.admission)? + } else { + state::Lock::acquire_existing(&paths.admission)? + }; + let lock_id = fs::symlink_metadata(paths.admission.join("operation.lock")) + .map(|m| inode(&m)) + .map_err(|_| refused())?; + Ok((fresh, admission_id, admission_lock, lock_id)) +} + +fn verify_admission( + paths: &Paths, + admission_id: Inode, + lock_id: Inode, + lock: &state::Lock, +) -> Result<(), CandidateError> { + if dir(&paths.admission)? != admission_id || lock.identity()? != (lock_id.dev, lock_id.ino) { + return Err(refused()); + } + exact_entries(&paths.admission, &["operation.lock"])?; + let m = fs::symlink_metadata(paths.admission.join("operation.lock")).map_err(|_| refused())?; + if !m.is_file() + || m.nlink() != 1 + || m.uid() != unsafe { libc::geteuid() } + || m.mode() & 0o777 != 0o600 + || inode(&m) != lock_id + { + return Err(refused()); + } + Ok(()) +} + +fn verify_recovery_process(journal: &Journal) -> Result<(), CandidateError> { + if journal.recovery_process.pid == std::process::id() as i32 { + if identity::observe(journal.recovery_process.pid)? != journal.recovery_process { + return Err(refused()); + } + } else if identity::alive(journal.recovery_process.pid)? { + return Err(refused()); + } + Ok(()) +} + +fn verify_journal_admission( + journal: &Journal, + fresh: bool, + complete: bool, + admission_id: Inode, + lock_id: Inode, +) -> Result<(), CandidateError> { + // An incomplete intent can resume only behind its original retained barrier. + // A completed exact archive may be replayed behind a newly created barrier. + if (fresh && !complete) + || ((!complete || !fresh) + && (journal.admission != admission_id || journal.admission_lock != lock_id)) + { + return Err(refused()); + } + Ok(()) +} + +/// An explicit current-branch native operation. A stale helper is never signaled, +/// its exact socket is retired by same-filesystem rename, and CA data is not changed. +pub fn archive( + candidate: &Candidate, + selection: ArchiveSelection<'_>, +) -> Result { + let ArchiveSelection { + run, + owner_generation: generation, + lease_id, + attempt, + owner, + namespace, + plan, + } = selection; + if !hex(run, HEX32) + || !hex(generation, HEX32) + || !hex(lease_id, HEX32) + || !hex(attempt, HEX32) + || !hex(owner, HEX32) + || !hex(namespace, 64) + || !hex(plan, 64) + || !cfg!(target_os = "macos") + { + return Err(refused()); + } + let home = &candidate.checkout; + let paths = Paths::new(home, generation, lease_id); + let storage_id = dir(&paths.storage)?; + let (fresh, admission_id, admission_lock, lock_id) = acquire_admission(&paths)?; + verify_admission(&paths, admission_id, lock_id, &admission_lock)?; + let mut intent_written = !absent(&paths.intent)?; + let result = (|| { + let journal: Journal = if intent_written { + let (bytes, _) = read_file(&paths.intent, MAX_FILE, true)?; + let journal: Journal = serde_json::from_slice(&bytes).map_err(|_| refused())?; + if journal.version != 1 + || journal.run != run + || journal.owner_generation != generation + || journal.lease_id != lease_id + { + return Err(refused()); + } + verify_journal_admission( + &journal, + fresh, + !absent(&paths.complete)?, + admission_id, + lock_id, + )?; + // flock excludes a still-running predecessor; matching PID requires + // matching birth, UID, and executable to exclude a reused process. + verify_recovery_process(&journal)?; + journal + } else { + if !fresh || !absent(&paths.complete)? || !absent(&paths.archive)? { + return Err(refused()); + } + let root = dir(&paths.source)?; + exact_entries( + &paths.source, + &["configuration.json", "endpoint.json", "leases"], + )?; + let (config_bytes, configuration) = + read_file(&paths.source.join("configuration.json"), MAX_FILE, true)?; + let config: Configuration = + serde_json::from_slice(&config_bytes).map_err(|_| refused())?; + let leases_path = paths.source.join("leases"); + let leases = dir(&leases_path)?; + exact_entries(&leases_path, &[&format!("{lease_id}.json")])?; + let (lease_bytes, lease_pin) = read_file( + &leases_path.join(format!("{lease_id}.json")), + MAX_FILE, + true, + )?; + let lease: Lease = serde_json::from_slice(&lease_bytes).map_err(|_| refused())?; + let (endpoint_bytes, endpoint_pin) = + read_file(&paths.source.join("endpoint.json"), MAX_FILE, true)?; + let endpoint: Endpoint = + serde_json::from_slice(&endpoint_bytes).map_err(|_| refused())?; + validate_config(&config, &lease, &endpoint, home, generation, lease_id)?; + if lease.run != run { + return Err(refused()); + } + let parent = socket_parent(&endpoint.socket)?; + let socket = match (absent(parent)?, absent(&endpoint.socket)?) { + (true, true) => SocketEvidence::Absent { + path: endpoint.socket, + }, + (false, false) => { + let parent_id = dir(parent)?; + exact_entries(parent, &["control.sock"])?; + let socket_m = fs::symlink_metadata(&endpoint.socket).map_err(|_| refused())?; + if !socket_m.file_type().is_socket() + || socket_m.mode() & 0o777 != 0o600 + || socket_m.uid() != unsafe { libc::geteuid() } + || socket_m.dev() != endpoint.dev + || socket_m.ino() != endpoint.ino + { + return Err(refused()); + } + SocketEvidence::Present { + path: endpoint.socket, + parent: parent_id, + socket: inode(&socket_m), + } + } + _ => return Err(refused()), + }; + let (_, ca) = read_file( + &paths + .storage + .join("data/caddy/pki/authorities/local/root.crt"), + MAX_FILE, + false, + )?; + Journal { + version: 1, + run: run.into(), + owner_generation: generation.into(), + lease_id: lease_id.into(), + old_boot: config.binding.pool.boot_id, + current_boot: String::new(), + recovery_process: identity::observe(std::process::id() as i32)?, + admission: admission_id, + admission_lock: lock_id, + root, + leases, + configuration, + endpoint: endpoint_pin, + lease: lease_pin, + socket, + ca, + } + }; + let (config, lease, _, socket) = selected_evidence(&paths, &journal)?; + if config.binding.runtime.home != *home + || lease.run != run + || lease.owner != owner + || lease.attempt != attempt + || lease.namespace != namespace + || lease.plan_id != plan + || lease.owner != config.binding.pool.owner + || journal.old_boot != config.binding.pool.boot_id + { + return Err(refused()); + } + let graph = HttpsArchiveGuard::acquire( + candidate, + run, + &lease.owner, + &lease.namespace, + &lease.plan_id, + &journal.old_boot, + )?; + if !journal.current_boot.is_empty() && journal.current_boot != graph.current_boot() { + return Err(refused()); + } + prove_quiescence(&config, socket.as_deref())?; + let _ports = https_recovery::port_absent(config.binding.https_port)?; + if !intent_written { + let current = Journal { + current_boot: graph.current_boot().into(), + ..journal + }; + graph.verify(candidate)?; + selected_evidence(&paths, ¤t)?; + verify_admission(&paths, admission_id, lock_id, &admission_lock)?; + // Create-new/fsync may fail after publication. Keep the admission + // barrier until a selected resume can prove the exact journal. + intent_written = true; + write_new(&paths.intent, ¤t, storage_id)?; + run_archive(&paths, ¤t, storage_id, &mut || { + verify_admission(&paths, admission_id, lock_id, &admission_lock)?; + graph.verify(candidate) + })?; + } else { + run_archive(&paths, &journal, storage_id, &mut || { + verify_admission(&paths, admission_id, lock_id, &admission_lock)?; + graph.verify(candidate) + })?; + } + Ok( + serde_json::json!({"archived":true,"run":run,"owner_generation":generation,"lease_id":lease_id,"data_retained":true,"processes_signaled":0}), + ) + })(); + if result.is_ok() || !intent_written { + // A completed archive may unblock normal admission. Before intent, + // exact owned lock cleanup is safe; after uncertain effect it is not. + let lock_path = paths.admission.join("operation.lock"); + if result.is_ok() { + verify_admission(&paths, admission_id, lock_id, &admission_lock)?; + } + if verify_admission(&paths, admission_id, lock_id, &admission_lock).is_ok() { + fs::remove_file(&lock_path).map_err(|_| refused())?; + fs::remove_dir(&paths.admission).map_err(|_| refused())?; + sync_dir(&paths.storage, storage_id)?; + } + } + drop(admission_lock); + result +} +fn run_archive( + paths: &Paths, + journal: &Journal, + storage_id: Inode, + verify_graph: &mut impl FnMut() -> Result<(), CandidateError>, +) -> Result<(), CandidateError> { + let intent_bytes = pinned_intent(paths, journal)?; + if absent(&paths.complete)? { + let (config, _, _, socket) = selected_evidence(paths, journal)?; + prove_quiescence(&config, socket.as_deref())?; + verify_graph()?; + if let (SocketEvidence::Present { path, parent, .. }, Some(selected)) = + (&journal.socket, socket.as_ref()) + { + if selected == path { + rename_exclusive(selected, &socket_archive(selected, journal)?)?; + sync_dir(selected.parent().ok_or_else(refused)?, *parent)?; + } + } + let (config, _, _, socket) = selected_evidence(paths, journal)?; + prove_quiescence(&config, socket.as_deref())?; + verify_graph()?; + if !absent(&paths.source)? { + rename_exclusive(&paths.source, &paths.archive)?; + sync_dir(&paths.storage, storage_id)?; + } + let (config, _, _, socket) = selected_evidence(paths, journal)?; + prove_quiescence(&config, socket.as_deref())?; + verify_graph()?; + let complete = serde_json::json!({"version":1,"intent_sha256":digest(&intent_bytes),"run":journal.run, + "owner_generation":journal.owner_generation,"lease_id":journal.lease_id, + "old_boot":journal.old_boot,"current_boot":journal.current_boot}); + write_new(&paths.complete, &complete, storage_id)?; + } else { + let (bytes, _) = read_file(&paths.complete, MAX_FILE, true)?; + let value: serde_json::Value = serde_json::from_slice(&bytes).map_err(|_| refused())?; + if value + != serde_json::json!({"version":1,"intent_sha256":digest(&intent_bytes),"run":journal.run, + "owner_generation":journal.owner_generation,"lease_id":journal.lease_id, + "old_boot":journal.old_boot,"current_boot":journal.current_boot}) + { + return Err(refused()); + } + } + if !absent(&paths.source)? || !absent(journal.socket.path())? { + return Err(refused()); + } + let (config, _, _, socket) = selected_evidence(paths, journal)?; + prove_quiescence(&config, socket.as_deref())?; + verify_graph() +} +fn pinned_intent(paths: &Paths, journal: &Journal) -> Result, CandidateError> { + let (bytes, _) = read_file(&paths.intent, MAX_FILE, true)?; + let parsed: Journal = serde_json::from_slice(&bytes).map_err(|_| refused())?; + if serde_json::to_vec(&parsed).map_err(|_| refused())? + != serde_json::to_vec(journal).map_err(|_| refused())? + { + return Err(refused()); + } + Ok(bytes) +} + +#[cfg(all(test, target_os = "macos"))] +mod tests; diff --git a/packages/runtime-core/src/provider/shared_https_recovery/tests.rs b/packages/runtime-core/src/provider/shared_https_recovery/tests.rs new file mode 100644 index 000000000..ad760b41d --- /dev/null +++ b/packages/runtime-core/src/provider/shared_https_recovery/tests.rs @@ -0,0 +1,430 @@ +use super::*; +use std::{ + os::unix::fs::PermissionsExt, + process::Command, + sync::atomic::{AtomicU64, Ordering}, + time::{Duration, SystemTime, UNIX_EPOCH}, +}; + +static NEXT_FIXTURE: AtomicU64 = AtomicU64::new(0); + +struct Fixture { + home: PathBuf, + socket_parent: PathBuf, + paths: Paths, + journal: Journal, +} +impl Drop for Fixture { + fn drop(&mut self) { + let _ = fs::remove_dir_all(&self.home); + let _ = fs::remove_dir_all(&self.socket_parent); + } +} +fn private_dir(path: &Path) { + fs::create_dir_all(path).unwrap(); + fs::set_permissions(path, fs::Permissions::from_mode(0o700)).unwrap(); +} +fn private_file(path: &Path, bytes: &[u8]) { + fs::write(path, bytes).unwrap(); + fs::set_permissions(path, fs::Permissions::from_mode(0o600)).unwrap(); +} +fn executable(path: &Path) -> String { + fs::write(path, b"#!/bin/sh\nexit 0\n").unwrap(); + fs::set_permissions(path, fs::Permissions::from_mode(0o700)).unwrap(); + https_recovery::executable_hash(path).unwrap() +} +fn fixture(present_socket: bool) -> Fixture { + let nonce = SystemTime::now() + .duration_since(UNIX_EPOCH) + .unwrap() + .as_nanos(); + let sequence = NEXT_FIXTURE.fetch_add(1, Ordering::Relaxed); + let home = PathBuf::from(format!( + "/private/tmp/hk-shared-recovery-test-{}-{nonce}-{sequence}", + std::process::id() + )); + let socket_parent = PathBuf::from(format!( + "/private/tmp/hk-https-leases-{}-{nonce}-{sequence}", + std::process::id() + )); + private_dir(&home); + let storage = home.join("native-https"); + private_dir(&storage); + let source = storage.join("shared-owner"); + private_dir(&source); + private_dir(&source.join("leases")); + let ca_path = storage.join("data/caddy/pki/authorities/local/root.crt"); + private_dir(ca_path.parent().unwrap()); + private_file(&ca_path, b"fixture-ca-unchanged"); + let frontend = home.join("old-frontend"); + let runtime = home.join("old-runtime"); + let caddy = home.join("old-caddy"); + let frontend_sha = executable(&frontend); + let runtime_sha = executable(&runtime); + let caddy_sha = executable(&caddy); + let reserved = std::net::TcpListener::bind("127.0.0.1:0").unwrap(); + let port = reserved.local_addr().unwrap().port(); + drop(reserved); + let generation = "a".repeat(32); + let run = "b".repeat(32); + let mut lease = Lease { + version: 1, + owner_generation: generation.clone(), + lease_id: String::new(), + owner: "c".repeat(32), + run: run.clone(), + attempt: "d".repeat(32), + namespace: "e".repeat(64), + plan_id: "f".repeat(64), + }; + let config_value = serde_json::json!({ + "version":1,"ownerGeneration":generation, + "binding": { + "runtime":{"binary":runtime,"home":home}, + "frontend":{"binary":frontend,"sha256":frontend_sha}, + "runtimeSha256":runtime_sha, + "pool":{"owner":lease.owner,"bootId":"11111111-1111-1111-1111-111111111111"}, + "caddyBinary":caddy,"caddySha256":caddy_sha, + "httpsPort":port,"certificateNameLimit":256 + } + }); + let config: Configuration = serde_json::from_value(config_value.clone()).unwrap(); + lease.lease_id = lease_id(&config, &lease).unwrap(); + private_file( + &source.join("configuration.json"), + serde_json::to_string(&config_value).unwrap().as_bytes(), + ); + private_file( + &source + .join("leases") + .join(format!("{}.json", lease.lease_id)), + serde_json::to_string(&lease).unwrap().as_bytes(), + ); + let socket_path = socket_parent.join("control.sock"); + let socket = if present_socket { + private_dir(&socket_parent); + let status = Command::new("/usr/bin/python3") + .arg("-c") + .arg("import socket,sys; s=socket.socket(socket.AF_UNIX); s.bind(sys.argv[1]); s.close()") + .arg(&socket_path) + .status().unwrap(); + assert!(status.success()); + fs::set_permissions(&socket_path, fs::Permissions::from_mode(0o600)).unwrap(); + let m = fs::symlink_metadata(&socket_path).unwrap(); + SocketEvidence::Present { + path: socket_path.clone(), + parent: dir(&socket_parent).unwrap(), + socket: inode(&m), + } + } else { + SocketEvidence::Absent { + path: socket_path.clone(), + } + }; + let (dev, ino) = match &socket { + SocketEvidence::Present { socket, .. } => (socket.dev, socket.ino), + SocketEvidence::Absent { .. } => (101, 202), + }; + private_file( + &source.join("endpoint.json"), + serde_json::to_string(&serde_json::json!({ + "version":1,"ownerGeneration":generation,"socket":socket_path,"dev":dev,"ino":ino + })) + .unwrap() + .as_bytes(), + ); + let paths = Paths::new(&home, &generation, &lease.lease_id); + private_dir(&paths.admission); + private_file(&paths.admission.join("operation.lock"), b"lock"); + let journal = Journal { + version: 1, + run, + owner_generation: generation, + lease_id: lease.lease_id.clone(), + old_boot: config.binding.pool.boot_id, + current_boot: "22222222-2222-2222-2222-222222222222".into(), + recovery_process: identity::observe(std::process::id() as i32).unwrap(), + admission: dir(&paths.admission).unwrap(), + admission_lock: inode( + &fs::symlink_metadata(paths.admission.join("operation.lock")).unwrap(), + ), + root: dir(&source).unwrap(), + leases: dir(&source.join("leases")).unwrap(), + configuration: read_file(&source.join("configuration.json"), MAX_FILE, true) + .unwrap() + .1, + endpoint: read_file(&source.join("endpoint.json"), MAX_FILE, true) + .unwrap() + .1, + lease: read_file( + &source + .join("leases") + .join(format!("{}.json", lease.lease_id)), + MAX_FILE, + true, + ) + .unwrap() + .1, + socket, + ca: read_file(&ca_path, MAX_FILE, false).unwrap().1, + }; + write_new(&paths.intent, &journal, dir(&paths.storage).unwrap()).unwrap(); + Fixture { + home, + socket_parent, + paths, + journal, + } +} + +#[test] +fn real_subprocess_socket_and_absent_parent_archive_preserve_exact_bytes_and_inodes() { + for present_socket in [false, true] { + let f = fixture(present_socket); + let old_lease = fs::read( + f.paths + .source + .join("leases") + .join(format!("{}.json", f.journal.lease_id)), + ) + .unwrap(); + let old_ca = fs::read( + f.home + .join("native-https/data/caddy/pki/authorities/local/root.crt"), + ) + .unwrap(); + run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || Ok(()), + ) + .unwrap(); + assert!(absent(&f.paths.source).unwrap()); + assert!(!absent(&f.paths.archive).unwrap()); + assert!(!absent(&f.paths.complete).unwrap()); + let (_, _, _, selected_socket) = selected_evidence(&f.paths, &f.journal).unwrap(); + assert_eq!(selected_socket.is_some(), present_socket); + assert_eq!(dir(&f.paths.archive).unwrap(), f.journal.root); + assert_eq!( + fs::read( + f.paths + .archive + .join("leases") + .join(format!("{}.json", f.journal.lease_id)) + ) + .unwrap(), + old_lease + ); + assert_eq!( + fs::read( + f.home + .join("native-https/data/caddy/pki/authorities/local/root.crt") + ) + .unwrap(), + old_ca + ); + run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || Ok(()), + ) + .unwrap(); + } +} + +#[test] +fn interrupted_owner_move_requires_same_selection_and_retains_barrier_for_resume() { + let f = fixture(false); + let mut checks = 0; + assert!( + run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || { + checks += 1; + if checks == 3 { Err(refused()) } else { Ok(()) } + } + ) + .is_err() + ); + assert!(absent(&f.paths.source).unwrap()); + assert!(!absent(&f.paths.archive).unwrap()); + assert!(absent(&f.paths.complete).unwrap()); + assert!(!absent(&f.paths.admission).unwrap()); + let mut stale = f.journal.clone(); + stale.current_boot = "33333333-3333-3333-3333-333333333333".into(); + assert!( + run_archive( + &f.paths, + &stale, + dir(&f.paths.storage).unwrap(), + &mut || Ok(()) + ) + .is_err() + ); + assert!(absent(&f.paths.complete).unwrap()); + run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || Ok(()), + ) + .unwrap(); + assert!(!absent(&f.paths.complete).unwrap()); +} + +#[test] +fn missing_graph_proof_or_partial_socket_never_archives_owner() { + let f = fixture(false); + assert!( + run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || Err(refused()) + ) + .is_err() + ); + assert!(!absent(&f.paths.source).unwrap()); + assert!(absent(&f.paths.archive).unwrap()); + assert!(absent(&f.paths.complete).unwrap()); + private_dir(&f.socket_parent); + assert!(selected_evidence(&f.paths, &f.journal).is_err()); + assert!(!absent(&f.paths.source).unwrap()); +} + +#[test] +fn admission_is_private_and_completed_history_cannot_claim_a_foreign_empty_lock() { + let f = fixture(false); + fs::remove_file(f.paths.admission.join("operation.lock")).unwrap(); + fs::remove_dir(&f.paths.admission).unwrap(); + let (fresh, admission_id, lock, lock_id) = acquire_admission(&f.paths).unwrap(); + assert!(fresh); + assert_eq!(dir(&f.paths.admission).unwrap(), admission_id); + verify_admission(&f.paths, admission_id, lock_id, &lock).unwrap(); + drop(lock); + + private_file(&f.paths.complete, b"old completed archive"); + fs::remove_file(f.paths.admission.join("operation.lock")).unwrap(); + assert!(acquire_admission(&f.paths).is_err()); + assert!(absent(&f.paths.admission.join("operation.lock")).unwrap()); + assert!(!absent(&f.paths.complete).unwrap()); +} + +#[test] +fn incomplete_intent_cannot_resume_behind_a_recreated_admission() { + let f = fixture(false); + fs::remove_file(f.paths.admission.join("operation.lock")).unwrap(); + fs::remove_dir(&f.paths.admission).unwrap(); + let (fresh, admission_id, lock, lock_id) = acquire_admission(&f.paths).unwrap(); + assert!(fresh); + assert!(verify_journal_admission(&f.journal, fresh, false, admission_id, lock_id).is_err()); + let mut recycled = f.journal.clone(); + recycled.admission = admission_id; + recycled.admission_lock = lock_id; + assert!(verify_journal_admission(&recycled, fresh, false, admission_id, lock_id).is_err()); + assert!(absent(&f.paths.archive).unwrap()); + assert!(!absent(&f.paths.source).unwrap()); + // Completed replay can use a new barrier, but only after its separate + // exact completion, graph, and owner checks in archive(). + verify_journal_admission(&f.journal, fresh, true, admission_id, lock_id).unwrap(); + drop(lock); +} + +#[test] +fn same_pid_with_different_birth_refuses_recovery() { + let f = fixture(false); + verify_recovery_process(&f.journal).unwrap(); + let mut reused = f.journal.clone(); + reused.recovery_process.start_micros += 1; + assert!(verify_recovery_process(&reused).is_err()); +} + +#[test] +fn live_pinned_executable_refuses_quiescence() { + let f = fixture(false); + let (bytes, _) = read_file(&f.paths.source.join("configuration.json"), MAX_FILE, true).unwrap(); + let mut config: Configuration = serde_json::from_slice(&bytes).unwrap(); + config.binding.caddy_binary = f.home.join("live-caddy"); + fs::copy("/bin/sleep", &config.binding.caddy_binary).unwrap(); + fs::set_permissions( + &config.binding.caddy_binary, + fs::Permissions::from_mode(0o700), + ) + .unwrap(); + config.binding.caddy_sha256 = + https_recovery::executable_hash(&config.binding.caddy_binary).unwrap(); + let mut child = Command::new(&config.binding.caddy_binary) + .arg("10") + .spawn() + .unwrap(); + std::thread::sleep(Duration::from_millis(50)); + let refusal = prove_quiescence(&config, None); + child.kill().unwrap(); + child.wait().unwrap(); + assert!(refusal.is_err()); +} + +#[test] +fn replaced_admission_refuses_before_archive_effect() { + let f = fixture(false); + let lock = state::Lock::acquire_existing(&f.paths.admission).unwrap(); + let original = dir(&f.paths.admission).unwrap(); + let lock_id = inode(&fs::symlink_metadata(f.paths.admission.join("operation.lock")).unwrap()); + let moved = f.paths.storage.join("replaced-admission-evidence"); + let mut replaced = false; + let result = run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || { + if !replaced { + fs::rename(&f.paths.admission, &moved).unwrap(); + private_dir(&f.paths.admission); + private_file(&f.paths.admission.join("operation.lock"), b"foreign"); + replaced = true; + } + verify_admission(&f.paths, original, lock_id, &lock) + }, + ); + assert!(result.is_err()); + assert!(!absent(&f.paths.source).unwrap()); + assert!(absent(&f.paths.archive).unwrap()); + assert!(absent(&f.paths.complete).unwrap()); + assert_eq!( + fs::read(f.paths.admission.join("operation.lock")).unwrap(), + b"foreign" + ); +} + +#[test] +fn completed_replay_refuses_when_a_new_shared_owner_has_published() { + let f = fixture(false); + run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || Ok(()), + ) + .unwrap(); + let old_archive = dir(&f.paths.archive).unwrap(); + private_dir(&f.paths.source); + private_file(&f.paths.source.join("new-owner"), b"another app"); + assert!( + run_archive( + &f.paths, + &f.journal, + dir(&f.paths.storage).unwrap(), + &mut || Ok(()) + ) + .is_err() + ); + assert_eq!(dir(&f.paths.archive).unwrap(), old_archive); + assert_eq!( + fs::read(f.paths.source.join("new-owner")).unwrap(), + b"another app" + ); +} diff --git a/packages/runtime-core/src/provider/state.rs b/packages/runtime-core/src/provider/state.rs index 7f3b1a81a..bc4c3bf84 100644 --- a/packages/runtime-core/src/provider/state.rs +++ b/packages/runtime-core/src/provider/state.rs @@ -57,7 +57,6 @@ impl Lock { } /// Observe existing state without initializing a directory or lock file. - #[cfg(target_os = "macos")] pub fn acquire_existing(root: &Path) -> Result { check_private_directory(root)?; let file = OpenOptions::new() @@ -177,7 +176,7 @@ impl Default for ReclamationPolicy { } } -#[derive(Debug, Serialize, Deserialize)] +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] #[serde(deny_unknown_fields)] pub struct Owner { #[serde(default, skip_serializing_if = "Option::is_none")] @@ -211,8 +210,34 @@ pub struct Owner { } impl Owner { pub fn load(candidate: &Candidate) -> Result { + Self::load_with_short_home(candidate, false) + } + + /// Explicit recovery may inspect a missing temporary alias, never a replaced one. + /// Loading here is read-only; the lifecycle boundary still proves provider absence. + pub(super) fn load_for_short_home_recovery( + candidate: &Candidate, + ) -> Result { + Self::load_with_short_home(candidate, true) + } + + fn load_with_short_home( + candidate: &Candidate, + allow_missing: bool, + ) -> Result { let root = candidate.state_root.join("run/smolvm"); reject_aliased_state(&root)?; + if allow_missing { + match fs::symlink_metadata(root.join("owner.pending")) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => {} + _ => { + return Err(CandidateError::new( + "recovery_required", + "Provider owner update is pending or unobservable; HOME alias was not restored.", + )); + } + } + } let owner: Self = read(&root.join("owner.json"))?; owner.network.validate()?; if let Some(share) = &owner.project_share { @@ -250,20 +275,69 @@ impl Owner { "Provider owner identity does not match this checkout.", )); } - // The only permitted alias is an exact, receipt-bound short HOME for Unix sockets. - let alias = fs::symlink_metadata(&owner.short_home).map_err(io)?; + owner.check_short_home(candidate, allow_missing)?; + check_private_directory(&root.join("home"))?; + Ok(owner) + } + /// The only permitted alias is the exact receipt-bound HOME for Unix sockets. + fn check_short_home( + &self, + candidate: &Candidate, + allow_missing: bool, + ) -> Result<(), CandidateError> { + let alias = match fs::symlink_metadata(&self.short_home) { + Ok(alias) => alias, + Err(error) if error.kind() == std::io::ErrorKind::NotFound => { + return if allow_missing { + Ok(()) + } else { + Err(CandidateError::new( + "provider_home_missing", + "Temporary provider HOME alias is absent. Use runtime recover to verify the stopped pool before restoring it.", + )) + }; + } + Err(error) => return Err(io(error)), + }; if !alias.file_type().is_symlink() || alias.uid() != unsafe { libc::geteuid() } - || fs::read_link(&owner.short_home).map_err(io)? != root.join("home") + || fs::read_link(&self.short_home).map_err(io)? + != candidate.state_root.join("run/smolvm/home") { return Err(CandidateError::new( "foreign_state", "Short provider HOME alias was replaced.", )); } - check_private_directory(&root.join("home"))?; - Ok(owner) + Ok(()) } + + /// Called only while the lifecycle owns the pool and VM locks and has audited absence. + /// Recheck the durable selection, then create exclusively; a concurrent alias is retained. + pub(super) fn restore_missing_short_home( + &self, + candidate: &Candidate, + ) -> Result<(), CandidateError> { + let observed = Self::load_for_short_home_recovery(candidate)?; + if observed != *self { + return Err(CandidateError::new( + "foreign_state", + "Provider owner changed during HOME recovery; no alias was restored.", + )); + } + std::os::unix::fs::symlink( + candidate.state_root.join("run/smolvm/home"), + &self.short_home, + ) + .map_err(|_| { + CandidateError::new( + "socket_alias_collision", + "Cannot exclusively restore provider HOME alias; no existing path was replaced.", + ) + })?; + self.check_short_home(candidate, false) + } + #[cfg(all(test, target_os = "macos"))] pub fn create( candidate: &Candidate, diff --git a/packages/runtime-core/tests/absent_recovery_cli_contract.rs b/packages/runtime-core/tests/absent_recovery_cli_contract.rs new file mode 100644 index 000000000..3ea54549d --- /dev/null +++ b/packages/runtime-core/tests/absent_recovery_cli_contract.rs @@ -0,0 +1,107 @@ +//! Recovery acknowledgements must reject at the CLI boundary before any runtime admission. +use serde_json::Value; +use std::{ + fs, + path::{Path, PathBuf}, + process::Command, +}; + +struct Fixture(PathBuf); +impl Fixture { + fn new() -> Self { + let root = std::env::temp_dir().join(format!( + "hack-absence-cli-{}-{}", + std::process::id(), + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .unwrap() + .as_nanos() + )); + fs::create_dir(&root).unwrap(); + Self(root.canonicalize().unwrap()) + } +} +impl Drop for Fixture { + fn drop(&mut self) { + fs::remove_dir_all(&self.0).unwrap(); + } +} + +#[test] +fn absent_recovery_requires_one_exact_selection_and_both_acknowledgements() { + let fixture = Fixture::new(); + // The development executable intentionally resolves its compiled checkout + // before command dispatch. Supply that valid identity so these assertions + // exercise recovery parsing instead of the unrelated checkout gate. + let checkout = Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../..") + .canonicalize() + .unwrap(); + let original = fixture.0.join("original-not-created.json"); + let inspection = fixture.0.join("inspection-not-created.json"); + let complete: Vec = [ + "graph", + "recover-absent-publication-cleanup", + "--run-id", + &"a".repeat(32), + "--original-owner-file", + original.to_str().unwrap(), + "--host-inspection-file", + inspection.to_str().unwrap(), + "--expect-selection", + &"b".repeat(64), + "--retain-data", + "--accept-unpinned-post-reboot", + "--json", + ] + .into_iter() + .map(str::to_owned) + .collect(); + let mut cases = Vec::new(); + for flag in ["--retain-data", "--accept-unpinned-post-reboot"] { + let mut missing = complete.clone(); + missing.retain(|argument| argument != flag); + cases.push(missing); + let mut duplicated = complete.clone(); + duplicated.push(flag.into()); + cases.push(duplicated); + } + for flag in [ + "--run-id", + "--original-owner-file", + "--host-inspection-file", + "--expect-selection", + ] { + let mut missing = complete.clone(); + let position = missing + .iter() + .position(|argument| argument == flag) + .unwrap(); + missing.drain(position..position + 2); + cases.push(missing); + } + for extra in ["--remove-data", "--json", "--force"] { + let mut invalid = complete.clone(); + invalid.push(extra.into()); + cases.push(invalid); + } + for arguments in cases { + let output = Command::new(env!("CARGO_BIN_EXE_hack-runtime-candidate")) + .arg("--candidate-root") + .arg(&checkout) + .args(&arguments) + .env_clear() + .env("PATH", "/nonexistent") + .env("HOME", fixture.0.join("home-not-created")) + .output() + .unwrap(); + assert_eq!(output.status.code(), Some(2), "{arguments:?}"); + assert!( + output.stdout.is_empty(), + "refusal must not emit a success result" + ); + let failure: Value = serde_json::from_slice(&output.stderr).unwrap(); + assert_eq!(failure["code"], "graph_arguments", "{arguments:?}"); + assert_eq!(fs::read_dir(&fixture.0).unwrap().count(), 0); + } +} diff --git a/packages/runtime-core/tests/installed_candidate_contract.rs b/packages/runtime-core/tests/installed_candidate_contract.rs index 0f5ad26b6..cf5dcee3b 100644 --- a/packages/runtime-core/tests/installed_candidate_contract.rs +++ b/packages/runtime-core/tests/installed_candidate_contract.rs @@ -93,6 +93,44 @@ fn copied_installed_binary_uses_only_explicit_home_and_preserves_reexec_identity assert_eq!(fs::read_dir(&cwd).unwrap().count(), 0); } +#[test] +fn quiescent_https_recovery_requires_exact_selectors_without_initializing_state() { + let fixture = Fixture::new(); + let home = fixture.directory("private"); + let cwd = fixture.directory("cwd"); + for (owner, pid) in [ + ("wrong", "999999"), + ( + "0000000000000000000000000000000000000000000000000000000000000000", + "1", + ), + ] { + let output = invoke( + Path::new(env!("CARGO_BIN_EXE_hack-native")), + &home, + &cwd, + &[ + "runtime", + "recover-quiescent-https", + "--expect-owner", + owner, + "--expect-configuration", + "1111111111111111111111111111111111111111111111111111111111111111", + "--expect-frontend-pid", + pid, + "--json", + ], + ); + assert!(!output.status.success()); + assert!( + String::from_utf8(output.stderr) + .unwrap() + .contains("https_recovery_refused") + ); + assert_eq!(fs::read_dir(&home).unwrap().count(), 0); + } +} + #[test] fn installed_entrypoint_refuses_public_missing_relative_and_aliased_homes() { let fixture = Fixture::new(); diff --git a/packages/runtime-core/tests/quiescent_dependency_recovery_cli_contract.rs b/packages/runtime-core/tests/quiescent_dependency_recovery_cli_contract.rs new file mode 100644 index 000000000..fa1a831ee --- /dev/null +++ b/packages/runtime-core/tests/quiescent_dependency_recovery_cli_contract.rs @@ -0,0 +1,60 @@ +//! An explicit live recovery selection is required before provider access. +use serde_json::Value; +use std::{fs, path::Path, process::Command}; + +#[test] +fn quiescent_recovery_rejects_missing_duplicate_and_invalid_selection() { + let checkout = Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../..") + .canonicalize() + .unwrap(); + let home = std::env::temp_dir().join(format!("hack-quiescent-cli-{}", std::process::id())); + fs::create_dir(&home).unwrap(); + let cases = [ + vec!["runtime", "recover-quiescent-dependency-sockets", "--json"], + vec![ + "runtime", + "recover-quiescent-dependency-sockets", + "--expect-sha256", + "invalid", + "--json", + ], + vec![ + "runtime", + "recover-quiescent-dependency-sockets", + "--expect-sha256", + "invalid", + "--json", + "--json", + ], + vec![ + "runtime", + "quiescent-dependency-socket-recovery", + "--json", + "--json", + ], + ]; + for arguments in cases { + let output = Command::new(env!("CARGO_BIN_EXE_hack-runtime-candidate")) + .arg("--candidate-root") + .arg(&checkout) + .args(&arguments) + .env_clear() + .env("PATH", "/nonexistent") + .env("HOME", &home) + .output() + .unwrap(); + assert_eq!(output.status.code(), Some(2), "{arguments:?}"); + assert!(output.stdout.is_empty()); + let failure: Value = serde_json::from_slice(&output.stderr).unwrap(); + assert!( + matches!( + failure["code"].as_str(), + Some("unsupported_command" | "quiescent_dependency_socket_recovery") + ), + "{failure:?}" + ); + assert_eq!(fs::read_dir(&home).unwrap().count(), 0); + } + fs::remove_dir(home).unwrap(); +} diff --git a/packages/runtime-core/tests/relay_fence_contract.rs b/packages/runtime-core/tests/relay_fence_contract.rs index 013064b48..fe8253c8c 100644 --- a/packages/runtime-core/tests/relay_fence_contract.rs +++ b/packages/runtime-core/tests/relay_fence_contract.rs @@ -28,7 +28,7 @@ fn interrupted_fence_publication_requires_canonical_valid_transition() { ("closing", "stopped"), ("stopped", "closing"), ] { - for action in ["stop", "remove", "start", "inspect"] { + for action in ["stop", "remove", "start", "inspect", "inspect-retirement"] { let before = format!("7 {allocation} {phase}\n"); let pending = format!("7 {allocation} {next}\n"); fs::write(root.join("state"), &before).unwrap(); @@ -85,7 +85,7 @@ fn interrupted_fence_publication_requires_canonical_valid_transition() { "closing", ] { for next in ["preparing", "cancelled", "launching", "stopped"] { - for action in ["stop", "remove", "start", "inspect"] { + for action in ["stop", "remove", "start", "inspect", "inspect-retirement"] { let before = format!("6 {allocation} {phase}\n"); let pending = format!("7 {allocation} {next}\n"); fs::write(root.join("state"), &before).unwrap(); diff --git a/packages/runtime-core/tests/source_device_rebind_cli_contract.rs b/packages/runtime-core/tests/source_device_rebind_cli_contract.rs new file mode 100644 index 000000000..d50bd340a --- /dev/null +++ b/packages/runtime-core/tests/source_device_rebind_cli_contract.rs @@ -0,0 +1,62 @@ +//! The explicit continuity acknowledgement must be exact before provider access. +use serde_json::Value; +use std::{fs, path::Path, process::Command}; + +#[test] +fn source_device_rebind_requires_exact_selection_and_acknowledgement() { + let checkout = Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../..") + .canonicalize() + .unwrap(); + let home = std::env::temp_dir().join(format!( + "hack-source-rebind-cli-{}-{}", + std::process::id(), + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .unwrap() + .as_nanos() + )); + fs::create_dir(&home).unwrap(); + let full = [ + "graph".to_owned(), + "recover-source-device-rebind".to_owned(), + "--run-id".to_owned(), + "a".repeat(32), + "--expect-selection".to_owned(), + "b".repeat(64), + "--accept-legacy-device-rebind".to_owned(), + "--json".to_owned(), + ]; + let mut cases = Vec::new(); + for flag in ["--accept-legacy-device-rebind", "--json"] { + let mut duplicate = full.to_vec(); + duplicate.push(flag.to_owned()); + cases.push(duplicate); + } + for flag in ["--run-id", "--expect-selection"] { + let mut missing = full.to_vec(); + let index = missing.iter().position(|value| value == flag).unwrap(); + missing.drain(index..index + 2); + cases.push(missing); + } + let mut unacknowledged = full.to_vec(); + unacknowledged.retain(|value| value != "--accept-legacy-device-rebind"); + cases.push(unacknowledged); + for arguments in cases { + let output = Command::new(env!("CARGO_BIN_EXE_hack-runtime-candidate")) + .arg("--candidate-root") + .arg(&checkout) + .args(&arguments) + .env_clear() + .env("PATH", "/nonexistent") + .env("HOME", &home) + .output() + .unwrap(); + assert_eq!(output.status.code(), Some(2), "{arguments:?}"); + assert!(output.stdout.is_empty()); + let failure: Value = serde_json::from_slice(&output.stderr).unwrap(); + assert_eq!(failure["code"], "graph_arguments", "{arguments:?}"); + assert_eq!(fs::read_dir(&home).unwrap().count(), 0); + } + fs::remove_dir(&home).unwrap(); +} diff --git a/packages/runtime-core/tests/stream_relay_contract.rs b/packages/runtime-core/tests/stream_relay_contract.rs index 8ff435c2a..06bc03779 100644 --- a/packages/runtime-core/tests/stream_relay_contract.rs +++ b/packages/runtime-core/tests/stream_relay_contract.rs @@ -773,15 +773,16 @@ fn publisher_refuses_port_conflicts_and_stale_reused_upstream() { #[test] fn publisher_requires_ack_before_forwarding_and_preserves_upstream_socket() { - use std::os::unix::net::UnixListener; + use std::os::unix::{fs::MetadataExt, net::UnixListener}; let root = PathBuf::from(format!("/tmp/hkp-ack-{}", std::process::id())); fs::create_dir(&root).unwrap(); let path = root.join("socket"); let upstream = UnixListener::bind(&path).unwrap(); fs::set_permissions(&path, fs::Permissions::from_mode(0o700)).unwrap(); let token = "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"; - let port = free_loopback_port(); - let mut publication = publisher(&path, token, port); + let owned = fs::symlink_metadata(&path).unwrap(); + // The ack contract is transport-independent; avoid releasing a TCP port before spawn. + let mut publication = frontend_publisher(&path, token, 0, None, true); let worker = thread::spawn(move || { let (mut socket, _) = upstream.accept().unwrap(); socket @@ -795,7 +796,7 @@ fn publisher_requires_ack_before_forwarding_and_preserves_upstream_socket() { socket.read_to_end(&mut extra).unwrap(); assert!(extra.is_empty()); }); - let mut client = std::net::TcpStream::connect(("127.0.0.1", port)).unwrap(); + let mut client = publication.connect(); client .set_read_timeout(Some(Duration::from_secs(3))) .unwrap(); @@ -808,7 +809,9 @@ fn publisher_requires_ack_before_forwarding_and_preserves_upstream_socket() { } worker.join().unwrap(); publication.stop(); - assert!(path.exists()); + let retained = fs::symlink_metadata(&path).unwrap(); + assert_eq!((retained.dev(), retained.ino()), (owned.dev(), owned.ino())); + assert_eq!(retained.mode(), owned.mode()); fs::remove_dir_all(root).unwrap(); } diff --git a/scripts/lib/tla-runtime-models.ts b/scripts/lib/tla-runtime-models.ts index c0222485e..3ad78cf80 100644 --- a/scripts/lib/tla-runtime-models.ts +++ b/scripts/lib/tla-runtime-models.ts @@ -26,6 +26,61 @@ type ModelContract = { // Bounds and witnesses are reviewed contracts, not learned from each run. const contracts: readonly ModelContract[] = [ + { + name: "absent-publication-recovery", + module: "Absent", + states: 247, + invariant: "NoPrematurePublication", + action: "Publish", + fields: [ + "intent = TRUE", + "complete = FALSE", + "retired = FALSE", + "published = TRUE", + "unsafePublication = TRUE", + ], + additionalControls: [ + { + name: "unwitnessed-cleanup", + negative: true, + invariant: "NoUnwitnessedCleanup", + action: "CleanupOne", + fields: [ + "intent = FALSE", + "progress = 1", + 'engine = "recovery"', + 'foreground = "recovery"', + "unsafeCleanup = TRUE", + ], + }, + { + name: "stale-selection", + negative: true, + invariant: "NoUnwitnessedCleanup", + action: "CleanupOne", + fields: [ + "intent = TRUE", + "selected = 1", + "version = 2", + "progress = 1", + "unsafeCleanup = TRUE", + ], + }, + { + name: "unconfirmed-retirement", + negative: true, + invariant: "NoUnconfirmedRetirement", + action: "RetireAbsentPublisher", + fields: [ + "intent = FALSE", + "complete = FALSE", + "progress = 0", + "retired = TRUE", + "unsafeRetirement = TRUE", + ], + }, + ], + }, { name: "shared-https-lifetime", module: "SharedHttps", @@ -34,6 +89,49 @@ const contracts: readonly ModelContract[] = [ action: "FinishClose", fields: ['mode = "closed"', "leases = {2}", "unsafeShutdown = TRUE"], }, + { + name: "previous-boot-shared-https", + module: "ArchiveHttps", + states: 92, + invariant: "NoPrematurePublication", + action: "Publish", + fields: [ + "intent = TRUE", + "complete = FALSE", + 'archived = {"owner"}', + "crashed = TRUE", + "published = TRUE", + "unsafePublication = TRUE", + ], + additionalControls: [ + { + name: "stale-selection", + negative: true, + invariant: "NoUnprovedArchive", + action: "ArchiveOwner", + fields: [ + "intent = TRUE", + "selected = 1", + "version = 2", + "engine = TRUE", + 'admission = "recovery"', + "unsafeArchive = TRUE", + ], + }, + { + name: "unproved-archive", + negative: true, + invariant: "NoUnprovedArchive", + action: "ArchiveOwner", + fields: [ + "intent = TRUE", + "eligible = FALSE", + "engine = TRUE", + "unsafeArchive = TRUE", + ], + }, + ], + }, { name: "stopped-pool-startup", module: "Resume", diff --git a/src/backends/native-https-empty-owner-recovery.ts b/src/backends/native-https-empty-owner-recovery.ts new file mode 100644 index 000000000..a83114d5c --- /dev/null +++ b/src/backends/native-https-empty-owner-recovery.ts @@ -0,0 +1,562 @@ +/** Explicit, conservative archival of a legacy shared HTTPS owner that never published an endpoint. + * + * The old creator PID is operator evidence: configuration.json did not record it. Current-source + * ensure calls share an admission lock, while a legacy binary does not. The operator must keep + * legacy startup quiescent; a pinned executable inventory refuses any live other caller. + * This API never signals a process or removes Caddy data. An interrupted intent is retained. + */ +import { spawn } from "node:child_process"; +import { createHash } from "node:crypto"; +import { constants } from "node:fs"; +import { + lstat, + mkdir, + open, + readdir, + realpath, + rename, + rmdir, +} from "node:fs/promises"; +import { join } from "node:path"; +import { isRecord } from "../lib/guards.ts"; +import { acquireNativeHttpsOwnerAdmission } from "./native-https-owner-admission.ts"; +import { + isNativeHttpsOwnerConfiguration, + NATIVE_HTTPS_OWNER_ARGUMENT, + type NativeHttpsOwnerConfiguration, + nativeHttpsOwnerRefused, +} from "./native-https-owner-protocol.ts"; +import { + nativeHttpsExecutableSha256, + nativeHttpsOwnerRoot, + nativeHttpsPrivateDirectory, + nativeHttpsReadFile, + nativeHttpsWriteNew, +} from "./native-https-owner-storage.ts"; +import { checkNativeHttpsPort } from "./native-https-port.ts"; +import { + invokeNativeRuntime, + type NativeRuntimeSelection, +} from "./native-runtime-client.ts"; + +interface Identity { + readonly dev: number; + readonly ino: number; +} +const HEX32 = /^[a-f0-9]{32}$/; +const HEX64 = /^[a-f0-9]{64}$/; +const PS_LINE = /^\s*(\d+)\s+(.+)$/; +interface SelectedOwner { + readonly root: Identity; + readonly configuration: Identity & { readonly sha256: string }; + readonly leases: Identity; + readonly owner: NativeHttpsOwnerConfiguration; +} +export interface EmptyNativeHttpsOwnerSelection { + readonly ownerGeneration: string; + readonly configurationSha256: string; + readonly runtime: NativeRuntimeSelection; + readonly runtimeSha256: string; + /** Observed when the legacy configuration was first written, not inferred from it. */ + readonly originalSpawnerPid: number; +} + +/** Only effect-free external observations are replaceable in isolated tests. */ +export interface EmptyNativeHttpsOwnerInspection { + originalSpawnerAbsent(pid: number): Promise; + selectedOwnerProcessAbsent(opts: { + readonly frontendBinary: string; + readonly configurationPath: string; + }): Promise; + selectedPortAbsent(port: number): Promise; + poolAndPublications(opts: { + readonly runtime: NativeRuntimeSelection; + readonly expectedOwner: string; + readonly expectedBootId: string; + }): Promise; +} + +function refused(): Error { + return nativeHttpsOwnerRefused(); +} +function isAbsent(error: unknown): boolean { + return isRecord(error) && error.code === "ENOENT"; +} +async function absent(path: string): Promise { + try { + await lstat(path); + return false; + } catch (error) { + if (isAbsent(error)) { + return true; + } + throw refused(); + } +} +async function requireAbsent(path: string): Promise { + if (!(await absent(path))) { + throw refused(); + } +} +async function directoryIdentity(path: string): Promise { + const stat = await lstat(path); + if ( + !stat.isDirectory() || + stat.uid !== process.getuid?.() || + (stat.mode & 0o777) !== 0o700 || + (await realpath(path)) !== path + ) { + throw refused(); + } + const after = await lstat(path); + if (after.dev !== stat.dev || after.ino !== stat.ino) { + throw refused(); + } + return { dev: stat.dev, ino: stat.ino }; +} +function sameIdentity(a: Identity, b: Identity): boolean { + return a.dev === b.dev && a.ino === b.ino; +} +async function exactEntries( + path: string, + expected: readonly string[] +): Promise { + const entries = (await readdir(path)).sort(); + if (entries.join("\0") !== [...expected].sort().join("\0")) { + throw refused(); + } +} +async function selectOwner( + selection: EmptyNativeHttpsOwnerSelection +): Promise { + const home = selection.runtime.home; + if ( + !Number.isSafeInteger(selection.originalSpawnerPid) || + selection.originalSpawnerPid <= 1 || + !HEX32.test(selection.ownerGeneration) || + !HEX64.test(selection.configurationSha256) || + !HEX64.test(selection.runtimeSha256) || + (await realpath(home)) !== home + ) { + throw refused(); + } + await nativeHttpsPrivateDirectory(home); + await nativeHttpsPrivateDirectory(join(home, "native-https")); + const rootPath = nativeHttpsOwnerRoot(home); + const root = await directoryIdentity(rootPath); + await exactEntries(rootPath, ["configuration.json", "leases"]); + const leases = await directoryIdentity(join(rootPath, "leases")); + await exactEntries(join(rootPath, "leases"), []); + const { bytes, identity } = await nativeHttpsReadFile( + join(rootPath, "configuration.json") + ); + const parsed: unknown = JSON.parse(bytes.toString("utf8")); + if (!isNativeHttpsOwnerConfiguration(parsed)) { + throw refused(); + } + if ( + parsed.ownerGeneration !== selection.ownerGeneration || + identity.sha256 !== selection.configurationSha256 || + parsed.binding.runtime.home !== home || + parsed.binding.runtime.binary !== selection.runtime.binary || + parsed.binding.runtimeSha256 !== selection.runtimeSha256 || + (await nativeHttpsExecutableSha256(selection.runtime.binary)) !== + selection.runtimeSha256 || + (await nativeHttpsExecutableSha256(parsed.binding.frontend.binary)) !== + parsed.binding.frontend.sha256 || + (await nativeHttpsExecutableSha256(parsed.binding.caddyBinary)) !== + parsed.binding.caddySha256 + ) { + throw refused(); + } + return { root, leases, configuration: identity, owner: parsed }; +} +async function sameSelected( + selection: EmptyNativeHttpsOwnerSelection, + before: SelectedOwner +): Promise { + const after = await selectOwner(selection); + if ( + !( + sameIdentity(before.root, after.root) && + sameIdentity(before.leases, after.leases) && + sameIdentity(before.configuration, after.configuration) + ) || + before.configuration.sha256 !== after.configuration.sha256 || + JSON.stringify(before.owner) !== JSON.stringify(after.owner) + ) { + throw refused(); + } +} + +async function boundedCommand( + binary: string, + args: readonly string[], + acceptedExit: readonly number[] +): Promise { + const child = spawn(binary, args, { + env: { PATH: "/usr/bin:/bin" }, + stdio: ["ignore", "pipe", "pipe"], + }); + let output = ""; + let diagnostics = ""; + let tooLarge = false; + const take = (target: "output" | "diagnostics", chunk: Buffer) => { + if (target === "output") { + output += chunk.toString("utf8"); + tooLarge ||= Buffer.byteLength(output) > 1_048_576; + } else { + diagnostics += chunk.toString("utf8"); + tooLarge ||= Buffer.byteLength(diagnostics) > 4096; + } + if (tooLarge) { + child.kill("SIGKILL"); + } + }; + child.stdout.on("data", (chunk: Buffer) => take("output", chunk)); + child.stderr.on("data", (chunk: Buffer) => take("diagnostics", chunk)); + const timer = setTimeout(() => child.kill("SIGKILL"), 5000); + try { + const code = await new Promise((resolve, reject) => { + child.once("error", reject); + child.once("close", (value) => resolve(value ?? -1)); + }); + if (tooLarge || diagnostics !== "" || !acceptedExit.includes(code)) { + throw refused(); + } + return output; + } finally { + clearTimeout(timer); + } +} + +const productionInspection: EmptyNativeHttpsOwnerInspection = { + originalSpawnerAbsent(pid) { + try { + process.kill(pid, 0); + return Promise.resolve(false); + } catch (error) { + return Promise.resolve(isRecord(error) && error.code === "ESRCH"); + } + }, + async selectedOwnerProcessAbsent(opts) { + // Do not return raw process arguments: they can contain unrelated private data. + const lines = ( + await boundedCommand("/bin/ps", ["-axo", "pid=,command=", "-ww"], [0]) + ).split("\n"); + for (const line of lines) { + if (line.trim() === "") { + continue; + } + const match = PS_LINE.exec(line); + if (!match) { + return false; + } + const pid = Number(match[1]); + const command = match[2] ?? ""; + if (pid === process.pid) { + continue; + } + // A legacy CLI ignores admission. Its exact executable may be live without + // the internal-owner argument; HOME is not reliably present in ps output. + if ( + command === opts.frontendBinary || + command.startsWith(`${opts.frontendBinary} `) || + (command.includes(NATIVE_HTTPS_OWNER_ARGUMENT) && + command.includes(opts.configurationPath)) + ) { + return false; + } + } + return true; + }, + async selectedPortAbsent(port) { + try { + // The local bind probe catches listeners invisible to process inventory. + await checkNativeHttpsPort(port); + } catch { + return false; + } + const lines = await boundedCommand( + "/usr/sbin/lsof", + ["-nP", `-iTCP:${port}`, "-sTCP:LISTEN"], + [0, 1] + ); + return lines.trim() === ""; + }, + async poolAndPublications(opts) { + const { runtime, expectedOwner, expectedBootId } = opts; + for (const dir of [ + join(runtime.home, ".hack-local"), + join(runtime.home, ".hack-local", "run"), + join(runtime.home, ".hack-local", "run", "smolvm"), + ]) { + await nativeHttpsPrivateDirectory(dir); + } + const ownerPath = join(runtime.home, ".hack-local/run/smolvm/owner.json"); + const { bytes: before, identity: beforeIdentity } = + await nativeHttpsReadFile(ownerPath, 1_048_576); + const owner: unknown = JSON.parse(before.toString("utf8")); + if ( + !isRecord(owner) || + owner.token !== expectedOwner || + owner.guest_boot_id !== expectedBootId || + owner.checkout !== runtime.home || + (await absent( + join(runtime.home, ".hack-local/run/smolvm/owner.pending") + )) === false + ) { + return false; + } + const [status, authority, publications] = await Promise.all([ + invokeNativeRuntime({ + runtime, + cwd: runtime.home, + args: ["runtime", "status", "--json"], + timeoutMs: 5000, + }), + invokeNativeRuntime({ + runtime, + cwd: runtime.home, + args: ["runtime", "managed-hostname-authority", "--json"], + timeoutMs: 5000, + }), + invokeNativeRuntime({ + runtime, + cwd: runtime.home, + args: ["runtime", "publication-hostnames", "--json"], + timeoutMs: 5000, + }), + ]); + const { bytes: after, identity: afterIdentity } = await nativeHttpsReadFile( + ownerPath, + 1_048_576 + ); + return ( + sameIdentity(beforeIdentity, afterIdentity) && + beforeIdentity.sha256 === afterIdentity.sha256 && + before.equals(after) && + isRecord(status) && + status.phase === "running" && + status.process_alive === true && + status.guest_boot_id === expectedBootId && + isRecord(authority) && + isRecord(authority.authority) && + authority.authority.present === false && + isRecord(publications) && + publications.scope === "durable-ownership-only" && + Array.isArray(publications.claims) && + publications.claims.length === 0 + ); + }, +}; + +async function observe( + selection: EmptyNativeHttpsOwnerSelection, + owner: NativeHttpsOwnerConfiguration, + inspection: EmptyNativeHttpsOwnerInspection, + heldOwnerLock: Identity +): Promise { + const storage = join(selection.runtime.home, "native-https"); + for (const name of ["owner.sock", "active-owner.json"]) { + if (!(await absent(join(storage, name)))) { + throw refused(); + } + } + if ( + !sameIdentity( + await directoryIdentity(join(storage, "owner.lock")), + heldOwnerLock + ) + ) { + throw refused(); + } + await exactEntries(join(storage, "owner.lock"), []); + const checks = await Promise.all([ + inspection.originalSpawnerAbsent(selection.originalSpawnerPid), + inspection.selectedOwnerProcessAbsent({ + frontendBinary: owner.binding.frontend.binary, + configurationPath: join( + nativeHttpsOwnerRoot(selection.runtime.home), + "configuration.json" + ), + }), + inspection.selectedPortAbsent(owner.binding.httpsPort), + inspection.poolAndPublications({ + runtime: selection.runtime, + expectedOwner: owner.binding.pool.owner, + expectedBootId: owner.binding.pool.bootId, + }), + ]); + if (checks.some((value) => value !== true)) { + throw refused(); + } +} + +async function noRecoveryClaims( + storage: string, + ownIntent?: string +): Promise { + const entries = await readdir(storage); + for (const entry of entries) { + if ( + (entry.startsWith("empty-owner-recovery-") || + entry.startsWith("archived-empty-owner-")) && + entry !== ownIntent + ) { + throw refused(); + } + } +} +async function syncDirectory(path: string): Promise { + const file = await open(path, constants.O_RDONLY); + try { + await file.sync(); + } finally { + await file.close(); + } +} +async function removeExactEmptyLock( + path: string, + identity: Identity +): Promise { + if (!sameIdentity(await directoryIdentity(path), identity)) { + throw refused(); + } + await exactEntries(path, []); + await rmdir(path); +} + +/** Archive only a selected empty legacy owner. An existing intent or lock is never replayed. */ +export async function archiveEmptyNativeHttpsOwner(opts: { + readonly selection: EmptyNativeHttpsOwnerSelection; + readonly acceptLegacyOwnerWithoutPid: true; + /** Pure-test seam; production callers omit it and use native and OS observations. */ + readonly inspection?: EmptyNativeHttpsOwnerInspection; + /** Pure-test fault seam; production uses the durable private-file publisher. */ + readonly writeJournal?: typeof nativeHttpsWriteNew; +}): Promise<{ + readonly archive: string; + readonly configurationSha256: string; +}> { + if (opts.acceptLegacyOwnerWithoutPid !== true) { + throw refused(); + } + const { selection } = opts; + const storage = join(selection.runtime.home, "native-https"); + const root = nativeHttpsOwnerRoot(selection.runtime.home); + await nativeHttpsPrivateDirectory(storage); + const storageId = await directoryIdentity(storage); + const admission = await acquireNativeHttpsOwnerAdmission({ + home: selection.runtime.home, + waitMs: 0, + }); + let ownerLockId: Identity | undefined; + let intentAttempted = false; + try { + await noRecoveryClaims(storage); + const selected = await selectOwner(selection); + const ownerLock = join(storage, "owner.lock"); + if (!(await absent(ownerLock))) { + throw refused(); + } + // This is also the normal HTTPS startup exclusion lock. Losing its mkdir race refuses. + await mkdir(ownerLock, { mode: 0o700 }); + ownerLockId = await directoryIdentity(ownerLock); + const inspection = opts.inspection ?? productionInspection; + await observe(selection, selected.owner, inspection, ownerLockId); + await sameSelected(selection, selected); + if (!sameIdentity(await directoryIdentity(storage), storageId)) { + throw refused(); + } + const stem = `${selection.ownerGeneration}-${selection.configurationSha256}`; + const archive = join(storage, `archived-empty-owner-${stem}`); + const intent = join(storage, `empty-owner-recovery-${stem}.intent.json`); + const complete = join( + storage, + `empty-owner-recovery-${stem}.complete.json` + ); + if ( + !( + (await absent(archive)) && + (await absent(intent)) && + (await absent(complete)) + ) + ) { + throw refused(); + } + // Recheck directly at the effect boundary, after the operation and HTTPS locks are held. + await observe(selection, selected.owner, inspection, ownerLockId); + await sameSelected(selection, selected); + await noRecoveryClaims(storage); + const proof = { + version: 1, + action: "archive-empty-native-https-owner", + home: selection.runtime.home, + ownerGeneration: selection.ownerGeneration, + configurationSha256: selection.configurationSha256, + originalSpawnerPid: selection.originalSpawnerPid, + root: selected.root, + configuration: selected.configuration, + leases: selected.leases, + poolOwner: selected.owner.binding.pool.owner, + poolBootId: selected.owner.binding.pool.bootId, + runtimeSha256: selection.runtimeSha256, + }; + // Publication can succeed before its writer fails during fsync/readback. + intentAttempted = true; + await (opts.writeJournal ?? nativeHttpsWriteNew)(intent, proof); + await observe(selection, selected.owner, inspection, ownerLockId); + await sameSelected(selection, selected); + await noRecoveryClaims(storage, `empty-owner-recovery-${stem}.intent.json`); + if (!((await absent(archive)) && (await absent(complete)))) { + throw refused(); + } + await rename(root, archive); + await syncDirectory(storage); + await requireAbsent(root); + if (!sameIdentity(await directoryIdentity(archive), selected.root)) { + throw refused(); + } + await exactEntries(archive, ["configuration.json", "leases"]); + await exactEntries(join(archive, "leases"), []); + const archived = await nativeHttpsReadFile( + join(archive, "configuration.json") + ); + if ( + !sameIdentity(archived.identity, selected.configuration) || + archived.identity.sha256 !== selected.configuration.sha256 || + !sameIdentity( + await directoryIdentity(join(archive, "leases")), + selected.leases + ) + ) { + throw refused(); + } + await (opts.writeJournal ?? nativeHttpsWriteNew)(complete, { + ...proof, + archive, + intentSha256: createHash("sha256") + .update((await nativeHttpsReadFile(intent)).bytes) + .digest("hex"), + }); + await requireAbsent(root); + await removeExactEmptyLock(ownerLock, ownerLockId); + ownerLockId = undefined; + await admission.release(); + await syncDirectory(storage); + return { archive, configurationSha256: selection.configurationSha256 }; + } catch (error) { + // Before durable intent, owned empty locks can be released. After it, uncertainty stays. + if (!intentAttempted) { + if (ownerLockId) { + await removeExactEmptyLock( + join(storage, "owner.lock"), + ownerLockId + ).catch(() => undefined); + } + await admission.release().catch(() => undefined); + } + throw error; + } +} diff --git a/src/backends/native-https-owner-admission.ts b/src/backends/native-https-owner-admission.ts new file mode 100644 index 000000000..498347dde --- /dev/null +++ b/src/backends/native-https-owner-admission.ts @@ -0,0 +1,83 @@ +/** Short shared-owner admission lock. It covers config creation/read/spawn and explicit archival, + * never a lease or long-lived HTTPS session. An interrupted owner retains the lock for review. + */ +import { lstat, mkdir, readdir, realpath, rmdir } from "node:fs/promises"; +import { join } from "node:path"; +import { isRecord } from "../lib/guards.ts"; +import { nativeHttpsOwnerRefused } from "./native-https-owner-protocol.ts"; +import { nativeHttpsPrivateDirectory } from "./native-https-owner-storage.ts"; + +export const NATIVE_HTTPS_OWNER_ADMISSION = "shared-owner-admission.lock"; + +async function identity( + path: string +): Promise<{ readonly dev: number; readonly ino: number }> { + const stat = await lstat(path); + if ( + !stat.isDirectory() || + stat.uid !== process.getuid?.() || + (stat.mode & 0o777) !== 0o700 || + (await realpath(path)) !== path + ) { + throw nativeHttpsOwnerRefused(); + } + const after = await lstat(path); + if (after.dev !== stat.dev || after.ino !== stat.ino) { + throw nativeHttpsOwnerRefused(); + } + return { dev: stat.dev, ino: stat.ino }; +} + +/** Existing private locks are waited on only by normal admission; recovery fails closed. */ +export async function acquireNativeHttpsOwnerAdmission(opts: { + readonly home: string; + readonly waitMs: 0 | 15_000; +}): Promise<{ readonly path: string; release(): Promise }> { + await nativeHttpsPrivateDirectory(opts.home); + const storage = join(opts.home, "native-https"); + await nativeHttpsPrivateDirectory(storage); + const path = join(storage, NATIVE_HTTPS_OWNER_ADMISSION); + const deadline = Date.now() + opts.waitMs; + for (;;) { + try { + await mkdir(path, { mode: 0o700 }); + const mine = await identity(path); + let released = false; + return { + path, + async release() { + if (released) { + throw nativeHttpsOwnerRefused(); + } + const current = await identity(path); + if ( + current.dev !== mine.dev || + current.ino !== mine.ino || + (await readdir(path)).length !== 0 + ) { + throw nativeHttpsOwnerRefused(); + } + await rmdir(path); + released = true; + }, + }; + } catch (error) { + if (!(isRecord(error) && error.code === "EEXIST")) { + throw error; + } + // A foreign or malformed entry is never waited on, adopted, or removed. + try { + await identity(path); + } catch (identityError) { + if (isRecord(identityError) && identityError.code === "ENOENT") { + continue; + } + throw identityError; + } + if (Date.now() >= deadline) { + throw nativeHttpsOwnerRefused(); + } + await Bun.sleep(25); + } + } +} diff --git a/src/backends/native-https-owner.ts b/src/backends/native-https-owner.ts index bd4eb6d9b..4b0e1e4b0 100644 --- a/src/backends/native-https-owner.ts +++ b/src/backends/native-https-owner.ts @@ -4,6 +4,7 @@ import { lstat, mkdir, realpath } from "node:fs/promises"; import { connect, type Socket } from "node:net"; import { dirname, join } from "node:path"; import { isRecord } from "../lib/guards.ts"; +import { acquireNativeHttpsOwnerAdmission } from "./native-https-owner-admission.ts"; import { decodeNativeHttpsOwnerFrame, encodeNativeHttpsOwnerFrame, @@ -40,7 +41,10 @@ import { verifyActiveNativeHttpsConnection, verifyNativeHttpsHostname, } from "./native-project-https.ts"; -import type { NativeRuntimeSelection } from "./native-runtime-client.ts"; +import { + invokeNativeRuntime, + type NativeRuntimeSelection, +} from "./native-runtime-client.ts"; export type NativeHttpsLeaseIdentity = LeaseIdentity; export function isNativeHttpsLeaseIdentity( @@ -136,24 +140,35 @@ function sameBinding( export async function readNativeHttpsOwnerConfiguration( runtime: NativeRuntimeSelection ): Promise { - if ((await realpath(runtime.home)) !== runtime.home) { + const configuration = await readNativeHttpsOwnerConfigurationAtHome( + runtime.home + ); + if (configuration.binding.runtime.binary !== runtime.binary) { + throw nativeHttpsOwnerRefused(); + } + return configuration; +} + +async function readNativeHttpsOwnerConfigurationAtHome( + home: string +): Promise { + if ((await realpath(home)) !== home) { throw nativeHttpsOwnerRefused(); } for (const directory of [ - runtime.home, - join(runtime.home, "native-https"), - nativeHttpsOwnerRoot(runtime.home), + home, + join(home, "native-https"), + nativeHttpsOwnerRoot(home), ]) { await nativeHttpsPrivateDirectory(directory); } const { bytes } = await nativeHttpsReadFile( - join(nativeHttpsOwnerRoot(runtime.home), "configuration.json") + join(nativeHttpsOwnerRoot(home), "configuration.json") ); const configuration: unknown = JSON.parse(bytes.toString("utf8")); if ( !isNativeHttpsOwnerConfiguration(configuration) || - configuration.binding.runtime.home !== runtime.home || - configuration.binding.runtime.binary !== runtime.binary + configuration.binding.runtime.home !== home ) { throw nativeHttpsOwnerRefused(); } @@ -177,43 +192,53 @@ export async function ensureNativeHttpsOwner(opts: { } await nativeHttpsPrivateDirectory(binding.runtime.home); await nativeHttpsPrivateDirectory(dirname(root), true); - let created = false; + const admission = await acquireNativeHttpsOwnerAdmission({ + home: binding.runtime.home, + waitMs: 15_000, + }); try { - await mkdir(root, { mode: 0o700 }); - created = true; - } catch (error) { - if (!isRecord(error) || error.code !== "EEXIST") { - throw nativeHttpsOwnerRefused(); - } - } - if (created) { - await mkdir(join(root, "leases"), { mode: 0o700 }); - const configurationPath = join(root, "configuration.json"); - await nativeHttpsWriteNew(configurationPath, configuration); - await (opts.spawnOwner ?? spawnNativeHttpsOwner)({ - frontend: binding.frontend, - home: binding.runtime.home, - configurationPath, - }); - return configuration; - } - // Only startup publication can be retried; no acquire/release request is replayed. - const deadline = Date.now() + 15_000; - while (Date.now() < deadline) { + let created = false; try { - const existing = await readNativeHttpsOwnerConfiguration(binding.runtime); - if (!sameBinding(existing.binding, binding)) { + await mkdir(root, { mode: 0o700 }); + created = true; + } catch (error) { + if (!isRecord(error) || error.code !== "EEXIST") { throw nativeHttpsOwnerRefused(); } - return existing; - } catch (error) { - if (!isRecord(error) || error.code !== "ENOENT") { - throw error; + } + if (created) { + await mkdir(join(root, "leases"), { mode: 0o700 }); + const configurationPath = join(root, "configuration.json"); + await nativeHttpsWriteNew(configurationPath, configuration); + await (opts.spawnOwner ?? spawnNativeHttpsOwner)({ + frontend: binding.frontend, + home: binding.runtime.home, + configurationPath, + }); + return configuration; + } + // Only startup publication can be retried; no acquire/release request is replayed. + const deadline = Date.now() + 15_000; + while (Date.now() < deadline) { + try { + const existing = await readNativeHttpsOwnerConfiguration( + binding.runtime + ); + if (!sameBinding(existing.binding, binding)) { + throw nativeHttpsOwnerRefused(); + } + return existing; + } catch (error) { + if (!isRecord(error) || error.code !== "ENOENT") { + throw error; + } + await Bun.sleep(25); } - await Bun.sleep(25); } + throw nativeHttpsOwnerRefused(); + } finally { + await admission.release(); } - throw nativeHttpsOwnerRefused(); } async function endpointFor( @@ -470,10 +495,17 @@ async function hasActiveNativeHttpsLease( identity: NativeHttpsLeaseIdentity ): Promise { try { - const configuration = await readNativeHttpsOwnerConfiguration(runtime); + // A newer owner may pin a different bundle. Its validated generation can + // exclude this historical lease without authorizing access to that owner. + const configuration = await readNativeHttpsOwnerConfigurationAtHome( + runtime.home + ); if (configuration.ownerGeneration !== identity.ownerGeneration) { return false; } + if (configuration.binding.runtime.binary !== runtime.binary) { + throw nativeHttpsOwnerRefused(); + } await lstat( join( nativeHttpsOwnerRoot(runtime.home), @@ -490,11 +522,86 @@ async function hasActiveNativeHttpsLease( } } +async function archivePreviousBootSharedHttps(opts: { + readonly runtime: NativeRuntimeSelection; + readonly identity: NativeHttpsLeaseIdentity; +}): Promise { + let configuration: NativeHttpsOwnerConfiguration | undefined; + try { + configuration = await readNativeHttpsOwnerConfigurationAtHome( + opts.runtime.home + ); + } catch (error) { + if (!isRecord(error) || error.code !== "ENOENT") { + throw error; + } + } + const status = await invokeNativeRuntime({ + runtime: opts.runtime, + cwd: opts.runtime.home, + args: ["runtime", "status", "--json"], + timeoutMs: 30_000, + }); + if ( + !isRecord(status) || + status.phase !== "running" || + status.process_alive !== true || + typeof status.guest_boot_id !== "string" + ) { + throw nativeHttpsOwnerRefused(); + } + if ( + !configuration || + (configuration.ownerGeneration === opts.identity.ownerGeneration && + configuration.binding.pool.bootId !== status.guest_boot_id) + ) { + const archived = await invokeNativeRuntime({ + runtime: opts.runtime, + cwd: opts.runtime.home, + args: [ + "runtime", + "archive-previous-boot-shared-https", + "--run-id", + opts.identity.run, + "--expect-owner-generation", + opts.identity.ownerGeneration, + "--expect-lease-id", + opts.identity.leaseId, + "--expect-attempt", + opts.identity.attempt, + "--expect-owner", + opts.identity.owner, + "--expect-namespace", + opts.identity.namespace, + "--expect-plan", + opts.identity.planId, + "--json", + ], + timeoutMs: 60_000, + }); + if ( + !isRecord(archived) || + archived.archived !== true || + archived.run !== opts.identity.run || + archived.owner_generation !== opts.identity.ownerGeneration || + archived.lease_id !== opts.identity.leaseId || + archived.data_retained !== true || + archived.processes_signaled !== 0 + ) { + throw nativeHttpsOwnerRefused(); + } + return true; + } + return false; +} + /** Recovery never treats absent/dead helper state as permission to adopt its children. */ export async function recoverNativeHttpsLease(opts: { readonly runtime: NativeRuntimeSelection; readonly identity: NativeHttpsLeaseIdentity; readonly verifyReleased?: typeof verifyNativeHttpsLeaseGraph; + /** Only an explicit v3 dead-frontend recovery may select prior-boot archival. */ + readonly archivePreviousBoot?: true; }): Promise { if (!isNativeHttpsLeaseIdentity(opts.identity)) { throw nativeHttpsOwnerRefused(); @@ -549,6 +656,12 @@ export async function recoverNativeHttpsLease(opts: { throw error; } } + if ( + opts.archivePreviousBoot === true && + (await archivePreviousBootSharedHttps(opts)) + ) { + return; + } const configuration = await readNativeHttpsOwnerConfiguration(opts.runtime); if ( configuration.ownerGeneration !== opts.identity.ownerGeneration || diff --git a/src/backends/native-project-down.ts b/src/backends/native-project-down.ts index f11a0e316..ecbf08378 100644 --- a/src/backends/native-project-down.ts +++ b/src/backends/native-project-down.ts @@ -1,10 +1,17 @@ import { basename } from "node:path"; +import { isDeepStrictEqual } from "node:util"; import { isRecord } from "../lib/guards.ts"; import { resolveProjectEnvConfig, selectProjectEnvValuesForExecutionTarget, } from "../lib/project-env-config.ts"; +import { + captureNativeProjectFinalization, + type NativeProjectFinalizationToken, + waitNativeProjectFinalization, +} from "./native-project-finalization.ts"; import { inspectNativeProjectGraph } from "./native-project-inspect.ts"; +import { recoverNativeInterruptedStartupCleanup } from "./native-project-interrupted-cleanup.ts"; import { loadNativeProjectRun, type NativeProjectRun, @@ -20,6 +27,21 @@ function refused(): Error { "Native down cleanup is unconfirmed; retained mapping and runtime state require inspection. No cleanup was replayed." ); } +function finalizationRefused(): Error { + return new Error( + "Native down frontend finalization is unconfirmed or its ownership changed; retained mapping and runtime state require inspection. No cleanup was replayed." + ); +} +async function captureFinalization(opts: { + readonly scope: NativeProjectRunScope; + readonly run: NativeProjectRun; +}): Promise { + try { + return await captureNativeProjectFinalization(opts); + } catch { + throw finalizationRefused(); + } +} function verify(value: unknown, run: NativeProjectRun, stopped: boolean): void { if ( !isRecord(value) || @@ -57,8 +79,8 @@ function verify(value: unknown, run: NativeProjectRun, stopped: boolean): void { } } } -/** Retaining cleanup only. Retire owned host processes after confirmed compute - * absence, including a previously recovered stop; user hooks are not replayed. +/** Retaining cleanup only. Success follows confirmed compute absence and the + * original frontend's finalizers, including a previously recovered stop. */ export async function nativeProjectDown(opts: { readonly runtime: NativeRuntimeSelection; @@ -67,6 +89,9 @@ export async function nativeProjectDown(opts: { readonly retireHostProcesses?: (run: NativeProjectRun) => Promise; readonly after?: (run: NativeProjectRun) => Promise; readonly invoke?: typeof invokeNativeRuntime; + /** Only the restart coordinator may defer to its already captured exact wait/recovery. */ + readonly deferFinalization?: boolean; + readonly finalizationTimeoutMs?: number; }) { const run = await loadNativeProjectRun(opts.scope); if (!run) { @@ -86,6 +111,42 @@ export async function nativeProjectDown(opts: { }); const initial = await inspect(); verify(initial, run, false); + const finalization = opts.deferFinalization + ? undefined + : await captureFinalization({ scope: opts.scope, run }); + const wait = async () => { + if (finalization === undefined) { + return; + } + try { + await waitNativeProjectFinalization({ + scope: opts.scope, + run, + token: finalization, + timeoutMs: opts.finalizationTimeoutMs, + }); + } catch { + throw finalizationRefused(); + } + }; + const recovered = await recoverNativeInterruptedStartupCleanup({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + run, + snapshot: initial, + invoke, + }); + if (recovered !== null) { + verify(recovered, run, true); + await opts.retireHostProcesses?.(run); + await wait(); + return { + backend: "native", + status: "stopped", + run: run.run, + dataPreserved: true, + } as const; + } if ( isRecord(initial) && isRecord(initial.receipt) && @@ -94,6 +155,7 @@ export async function nativeProjectDown(opts: { // Recovery already stopped compute. User hooks and guest cleanup must not replay. verify(initial, run, true); await opts.retireHostProcesses?.(run); + await wait(); return { backend: "native", status: "stopped", @@ -104,6 +166,15 @@ export async function nativeProjectDown(opts: { await opts.before?.(run); // Recheck after hooks, which may run arbitrary user-authorized commands. verify(await inspect(), run, false); + if ( + finalization !== undefined && + !isDeepStrictEqual( + finalization, + await captureFinalization({ scope: opts.scope, run }) + ) + ) { + throw finalizationRefused(); + } await invoke({ runtime: opts.runtime, cwd: opts.scope.projectRoot, @@ -113,6 +184,7 @@ export async function nativeProjectDown(opts: { verify(await inspect(), run, true); await opts.retireHostProcesses?.(run); await opts.after?.(run); + await wait(); return { backend: "native", status: "stopped", diff --git a/src/backends/native-project-interrupted-cleanup.ts b/src/backends/native-project-interrupted-cleanup.ts new file mode 100644 index 000000000..bd008cd06 --- /dev/null +++ b/src/backends/native-project-interrupted-cleanup.ts @@ -0,0 +1,153 @@ +import { isDeepStrictEqual } from "node:util"; +import { isRecord } from "../lib/guards.ts"; +import { inspectNativeProjectGraph } from "./native-project-inspect.ts"; +import { confirmedNativeRetainedGraph } from "./native-project-retained.ts"; +import type { NativeProjectRun } from "./native-project-run.ts"; +import { + invokeNativeRuntime, + type NativeRuntimeSelection, +} from "./native-runtime-client.ts"; + +const SHA256 = /^[a-f0-9]{64}$/; + +function refused(): Error { + return new Error( + "Native interrupted startup cleanup is unconfirmed; inspect owned state before retrying." + ); +} + +function sameResources(before: unknown, after: unknown): boolean { + if ( + !(isRecord(before) && isRecord(after)) || + Object.keys(before).length === 0 || + Object.keys(before).length !== Object.keys(after).length + ) { + return false; + } + return Object.entries(before).every(([key, resource]) => { + const final = after[key]; + return ( + isRecord(resource) && + isRecord(final) && + isDeepStrictEqual( + { ...resource, phase: undefined }, + { ...final, phase: undefined } + ) + ); + }); +} + +/** Recover a selected failed startup after its native foreground has exited. + * Native inspection must prove the dead owner, unchanged boot and exact pending + * effect before issuing a separate retaining recovery. A bounded native journal + * hint also admits inspection of unfinished post-ACK retirement. Neither hint + * grants authority: native selection is revalidated before recovery. The old + * effect is never replayed, and acknowledgement alone cannot establish cleanup. + */ +export async function recoverNativeInterruptedStartupCleanup(opts: { + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly run: NativeProjectRun; + readonly snapshot: unknown; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + const value = opts.snapshot; + if (!(isRecord(value) && isRecord(value.receipt))) { + return null; + } + const { receipt } = value; + const pending = + receipt.phase === "cleanup-intent" && + isRecord(receipt.relay_cleanup) && + receipt.relay_cleanup.phase === "pending"; + if (!(pending || value.interrupted_start_cleanup_incomplete === true)) { + return null; + } + if ( + !["cleanup-intent", "stopped-data-retained"].includes( + String(receipt.phase) + ) || + value.journal_incomplete !== false || + receipt.run !== opts.run.run || + receipt.owner !== opts.run.owner || + receipt.namespace !== opts.run.namespace || + receipt.plan_id !== opts.run.planId + ) { + throw refused(); + } + const invoke = opts.invoke ?? invokeNativeRuntime; + try { + const selection = await invoke({ + runtime: opts.runtime, + cwd: opts.projectRoot, + args: [ + "graph", + "inspect-interrupted-start-cleanup", + "--run-id", + opts.run.run, + "--json", + ], + timeoutMs: 30_000, + }); + if ( + !isRecord(selection) || + selection.run !== opts.run.run || + selection.phase !== receipt.phase || + selection.eligible !== true || + selection.data_retained !== true || + selection.same_boot !== true || + typeof selection.selection_sha256 !== "string" || + !SHA256.test(selection.selection_sha256) + ) { + throw refused(); + } + const recovered = await invoke({ + runtime: opts.runtime, + cwd: opts.projectRoot, + args: [ + "graph", + "recover-interrupted-start-cleanup", + "--run-id", + opts.run.run, + "--expect-selection", + selection.selection_sha256, + "--retain-data", + "--json", + ], + timeoutMs: 590_000, + }); + if ( + !isRecord(recovered) || + recovered.run !== opts.run.run || + recovered.phase !== "stopped-data-retained" || + recovered.recovered !== true || + recovered.data_retained !== true || + recovered.same_boot !== true || + recovered.publisher_retired !== true || + recovered.reservation_released !== true + ) { + throw refused(); + } + const final = await inspectNativeProjectGraph({ + runtime: opts.runtime, + projectRoot: opts.projectRoot, + run: opts.run.run, + invoke, + }); + if ( + !( + confirmedNativeRetainedGraph(final, opts.run) && + isRecord(final) && + final.interrupted_start_cleanup_incomplete !== true && + isRecord(final.receipt) && + sameResources(receipt.resources, final.receipt.resources) + ) + ) { + throw refused(); + } + return final; + } catch { + // Native errors and peer replies may contain arbitrary private diagnostics. + throw refused(); + } +} diff --git a/src/backends/native-project-mapping-command.ts b/src/backends/native-project-mapping-command.ts new file mode 100644 index 000000000..995d5b476 --- /dev/null +++ b/src/backends/native-project-mapping-command.ts @@ -0,0 +1,131 @@ +import { CliUsageError } from "../cli/command.ts"; +import { resolveEffectiveBranch } from "../lib/branches.ts"; +import { emitCliResult, okResult } from "../lib/cli-result.ts"; +import { + findProjectContext, + readProjectConfig, + resolveWorktreeAutoBranch, + sanitizeBranchSlug, +} from "../lib/project.ts"; +import { + inspectNativeProjectRunFilesystemRecovery, + recoverNativeProjectRunFilesystem, +} from "./native-project-run.ts"; +import type { NativeRuntimeSelection } from "./native-runtime-client.ts"; + +const SELECTION = /^[a-f0-9]{64}$/; + +type Selection = + | { readonly action: "inspect" } + | { readonly action: "repair"; readonly expectSelection: string }; + +/** A device rebind is an explicit selected metadata migration, never an automatic Doctor fix. */ +export function parseNativeRunMappingRecoveryOptions(opts: { + readonly action?: string; + readonly expectSelection?: string; + readonly acceptLegacyDeviceRebind?: boolean; + readonly branch?: string; + readonly otherOptions: boolean; +}): Selection | null { + if (opts.action === undefined) { + if ( + opts.expectSelection !== undefined || + opts.acceptLegacyDeviceRebind || + opts.branch !== undefined + ) { + throw new CliUsageError( + "Run-mapping selection and branch options require --native-run-mapping inspect|repair." + ); + } + return null; + } + if (opts.otherOptions) { + throw new CliUsageError( + "--native-run-mapping cannot be combined with other Doctor repair or browser options." + ); + } + if (opts.action === "inspect") { + if (opts.expectSelection !== undefined || opts.acceptLegacyDeviceRebind) { + throw new CliUsageError( + "Run-mapping inspection is read-only; omit repair selection and acceptance flags." + ); + } + return { action: "inspect" }; + } + if ( + opts.action !== "repair" || + !opts.acceptLegacyDeviceRebind || + !opts.expectSelection || + opts.expectSelection.length !== 64 || + !SELECTION.test(opts.expectSelection) + ) { + throw new CliUsageError( + "Run-mapping repair requires --native-run-mapping repair --expect-selection <64-hex> --accept-legacy-device-rebind." + ); + } + return { action: "repair", expectSelection: opts.expectSelection }; +} + +/** Repair only mapping scope metadata. Native graph, socket and data authority remain unchanged. */ +export async function runNativeProjectMappingCommand(opts: { + readonly selection: Selection; + readonly startDir: string; + readonly branch?: string; + readonly runtime: NativeRuntimeSelection; + readonly json: boolean; +}): Promise { + const project = await findProjectContext(opts.startDir); + if (!project) { + throw new CliUsageError( + "Native run-mapping recovery requires a Hack project." + ); + } + const cfg = await readProjectConfig(project); + if (cfg.parseError) { + throw new CliUsageError( + "Native run-mapping recovery requires valid project configuration." + ); + } + const explicit = opts.branch?.trim(); + if (explicit !== undefined && explicit.length === 0) { + throw new CliUsageError("--branch requires a nonempty instance selection."); + } + const selected = await resolveEffectiveBranch({ + explicitBranch: explicit ? sanitizeBranchSlug(explicit) || "branch" : null, + projectRoot: project.projectRoot, + autoBranchEnabled: resolveWorktreeAutoBranch(cfg), + }); + if (selected.source === "detached-worktree") { + throw new CliUsageError( + "Detached worktree recovery requires --branch ." + ); + } + const scope = { + projectRoot: project.projectRoot, + projectDir: project.projectDir, + nativeHome: opts.runtime.home, + branch: selected.branch, + }; + const data = + opts.selection.action === "inspect" + ? await inspectNativeProjectRunFilesystemRecovery({ + scope, + runtime: opts.runtime, + }) + : await recoverNativeProjectRunFilesystem({ + scope, + runtime: opts.runtime, + expectSelection: opts.selection.expectSelection, + acceptLegacyDeviceRebind: true, + }); + if (opts.json) { + emitCliResult({ result: okResult({ data }) }); + } else { + process.stdout.write( + `Native run mapping ${opts.selection.action === "repair" ? "repaired" : "inspected"}: branch=${selected.branch ?? "base"}, run=${data.run.run}\n` + + `selection=${data.selectionSha256}\nqualification=${data.qualification}\n` + + "Graph recovery, retained-data readback and application readiness require separate verification.\n" + ); + } + return 0; +} diff --git a/src/backends/native-project-recovery.ts b/src/backends/native-project-recovery.ts index 45d48701c..f1d686215 100644 --- a/src/backends/native-project-recovery.ts +++ b/src/backends/native-project-recovery.ts @@ -216,6 +216,7 @@ export async function verifyNativeFrontendRecovery( await (opts.recoverLease ?? recoverNativeHttpsLease)({ runtime: opts.runtime, identity: opts.httpsLease, + archivePreviousBoot: true, }); await opts.cleanupLifecycle?.(); await noLifecycleEntries(opts.scope.projectDir); diff --git a/src/backends/native-project-restart-preflight.ts b/src/backends/native-project-restart-preflight.ts index eee54851b..cbc0075ac 100644 --- a/src/backends/native-project-restart-preflight.ts +++ b/src/backends/native-project-restart-preflight.ts @@ -5,7 +5,6 @@ import { dirname, join } from "node:path"; import { isRecord } from "../lib/guards.ts"; import { adaptNativeAwsEnvironment } from "./native-aws-environment.ts"; import { prepareNativeProjectAdaptation } from "./native-project-adaptation.ts"; -import { prepareNativeProjectBranch } from "./native-project-branch.ts"; import { discoverNativeHostDependency, type NativeHostDependency, @@ -21,7 +20,13 @@ import { verifyNativeSourceCompatibility, } from "./native-project-restore.ts"; import { confirmedNativeRetainedGraph } from "./native-project-retained.ts"; -import { withNativeProjectReview } from "./native-project-review.ts"; +import { selectNativeRetainedImages } from "./native-project-retained-images.ts"; +import { + prepareNativeReviewBranch, + selectNativeProjectReviewIdentity, + verifyNativeActiveReview, + withNativeProjectReview, +} from "./native-project-review.ts"; import { type NativeProjectRun, type NativeProjectRunScope, @@ -41,6 +46,8 @@ const DEFAULTS = { dependencies: readNativeHostDependencies, invoke: invokeNativeRuntime, review: withNativeProjectReview, + selectReview: selectNativeProjectReviewIdentity, + retainedImages: selectNativeRetainedImages, }; const IMAGE = /^sha256:[a-f0-9]{64}$/; const SHA = /^[a-f0-9]{64}$/; @@ -207,6 +214,88 @@ export function nativeRestartSelection(opts: { }; } +async function inspectRestartGraph(opts: { + readonly runtime: NativeRuntimeSelection; + readonly scope: NativeProjectRunScope; + readonly run: NativeProjectRun; + readonly invoke: typeof invokeNativeRuntime; +}): Promise { + try { + return await inspectNativeProjectGraph({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + run: opts.run.run, + invoke: opts.invoke, + }); + } catch { + throw new Error( + "Native restart cannot confirm the stopped retained graph; no listener discovery or cleanup was requested." + ); + } +} + +async function stoppedRestartMode(opts: { + readonly runtime: NativeRuntimeSelection; + readonly scope: NativeProjectRunScope; + readonly run: NativeProjectRun; + readonly cleanedRetry?: boolean; + readonly recoverStopped?: boolean; + readonly invoke: typeof invokeNativeRuntime; +}): Promise { + if (!(opts.cleanedRetry || opts.recoverStopped)) { + return false; + } + const observed = await inspectRestartGraph(opts); + const confirmed = confirmedNativeRetainedGraph(observed, opts.run); + if ( + !confirmed && + (opts.cleanedRetry || + (isRecord(observed) && + isRecord(observed.receipt) && + observed.receipt.phase === "stopped-data-retained")) + ) { + throw new Error( + "Native restart cannot confirm the stopped retained graph; no listener discovery or cleanup was requested." + ); + } + // Unconfirmed graphs retain active selection and live-listener checks. + return confirmed; +} + +async function pinRestartImages(opts: { + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly specs: Record>; + readonly retainedImages: ReadonlyMap; + readonly invoke: typeof invokeNativeRuntime; +}): Promise { + for (const [name, spec] of Object.entries(opts.specs)) { + const retainedImage = opts.retainedImages.get(name); + if (retainedImage) { + spec.image = retainedImage; + continue; + } + const image = String(spec.image); + if (IMAGE.test(image)) { + continue; + } + // Owned cache acquisition grants no authority to stop the current graph. + const pinned = await opts.invoke({ + runtime: opts.runtime, + cwd: opts.projectRoot, + args: ["runtime", "ensure-image", "--reference", image, "--json"], + }); + if ( + !isRecord(pinned) || + typeof pinned.image_id !== "string" || + !IMAGE.test(pinned.image_id) + ) { + throw new Error("Native restart image selection failed before cleanup."); + } + spec.image = pinned.image_id; + } +} + /** Review the same public graph before stopping anything; never run startup hooks here. */ export async function preflightNativeRestart(opts: { readonly runtime: NativeRuntimeSelection; @@ -214,6 +303,8 @@ export async function preflightNativeRestart(opts: { readonly composeFile: string; readonly run: NativeProjectRun; readonly cleanedRetry?: boolean; + /** Explicit frontend recovery may encounter a graph stopped before an intent was saved. */ + readonly recoverStopped?: boolean; readonly adaptationFile?: string; readonly dependencyFile?: string; readonly allowedHosts?: readonly string[]; @@ -221,26 +312,10 @@ export async function preflightNativeRestart(opts: { }) { const deps = { ...DEFAULTS, ...opts.dependencies }; const selected = nativeRestartSelection({ run: opts.run }); - if (opts.cleanedRetry) { - let observed: unknown; - try { - observed = await inspectNativeProjectGraph({ - runtime: opts.runtime, - projectRoot: opts.scope.projectRoot, - run: opts.run.run, - invoke: deps.invoke, - }); - } catch { - throw new Error( - "Native restart cannot confirm the stopped retained graph; no listener discovery or cleanup was requested." - ); - } - if (!confirmedNativeRetainedGraph(observed, opts.run)) { - throw new Error( - "Native restart cannot confirm the stopped retained graph; no listener discovery or cleanup was requested." - ); - } - } + const stoppedRetry = await stoppedRestartMode({ + ...opts, + invoke: deps.invoke, + }); requireNativeRestartNetwork( await deps.invoke({ runtime: opts.runtime, @@ -268,11 +343,19 @@ export async function preflightNativeRestart(opts: { input, path: opts.adaptationFile, }); - input = await prepareNativeProjectBranch({ - input, + const reviewedBranch = await prepareNativeReviewBranch({ + runtime: opts.runtime, scope: opts.scope, composeFile: opts.composeFile, + profiles: selected.profiles, + input, + retained: opts.run, + retainedMode: stoppedRetry ? "stopped" : "active", + invoke: deps.invoke, + select: deps.selectReview, }); + input = reviewedBranch.input; + const reviewIdentity = reviewedBranch.identity; if (selected.aws) { input = (await deps.aws({ input, ...selected.aws })).input; } @@ -280,7 +363,7 @@ export async function preflightNativeRestart(opts: { input, opts.dependencyFile !== undefined ); - const dependencies = opts.cleanedRetry + const dependencies = stoppedRetry ? await readNativeRestartDependencyIntent({ path: opts.dependencyFile, services: Object.keys(specs), @@ -298,26 +381,22 @@ export async function preflightNativeRestart(opts: { }), }); prepareNativeDependencyServices({ dependencies, services: specs }); - for (const spec of Object.values(specs)) { - const image = String(spec.image); - if (IMAGE.test(image)) { - continue; - } - // Image acquisition may populate an owned cache, but cannot stop the current graph. - const pinned = await deps.invoke({ - runtime: opts.runtime, - cwd: opts.scope.projectRoot, - args: ["runtime", "ensure-image", "--reference", image, "--json"], - }); - if ( - !isRecord(pinned) || - typeof pinned.image_id !== "string" || - !IMAGE.test(pinned.image_id) - ) { - throw new Error("Native restart image selection failed before cleanup."); - } - spec.image = pinned.image_id; - } + const retainedImages = stoppedRetry + ? await deps.retainedImages({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + originalSha256: input.originalSha256, + restore: opts.run, + invoke: deps.invoke, + }) + : new Map(); + await pinRestartImages({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + specs, + retainedImages, + invoke: deps.invoke, + }); const compose = JSON.parse(input.normalizedComposeJson); compose.services = specs; await deps.review({ @@ -326,6 +405,10 @@ export async function preflightNativeRestart(opts: { composeFile: opts.composeFile, profiles: selected.profiles, branch: opts.scope.branch, + retained: opts.run, + retainedMode: stoppedRetry ? "stopped" : "active", + reviewIdentity, + invoke: deps.invoke, input: { ...input, normalizedComposeJson: JSON.stringify(compose) }, run: async (review) => { if (review.namespace !== opts.run.namespace) { @@ -348,8 +431,26 @@ export async function preflightNativeRestart(opts: { planId: review.planId, dependencies, invoke: deps.invoke, - cleanedRetry: opts.cleanedRetry === true, + cleanedRetry: stoppedRetry, + }); + await verifyNativeActiveReview({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + retained: opts.run, + proof: reviewIdentity?.activeProof, + invoke: deps.invoke, }); + if (stoppedRetry && opts.recoverStopped && !opts.cleanedRetry) { + const observed = await inspectRestartGraph({ + ...opts, + invoke: deps.invoke, + }); + if (!confirmedNativeRetainedGraph(observed, opts.run)) { + throw new Error( + "Native restart stopped recovery changed during review; no cleanup was requested." + ); + } + } }, }); } diff --git a/src/backends/native-project-restore.ts b/src/backends/native-project-restore.ts index 1f9885d4f..da477f80b 100644 --- a/src/backends/native-project-restore.ts +++ b/src/backends/native-project-restore.ts @@ -9,6 +9,30 @@ import { } from "./native-runtime-client.ts"; const SHA = /^[a-f0-9]{64}$/; +/** Recheck native authority without deriving or translating any ownership identity. */ +export function validateNativeRestoreSelection(opts: { + readonly selected: unknown; + readonly saved: NativeProjectRun; + readonly namespace: string; +}): { readonly generation: string } { + const { selected, saved } = opts; + if ( + !isRecord(selected) || + selected.run !== saved.run || + selected.owner !== saved.owner || + selected.namespace !== opts.namespace || + selected.namespace !== saved.namespace || + selected.plan !== saved.planId || + typeof selected.generation !== "string" || + !SHA.test(selected.generation) + ) { + throw new Error( + "Native restore selection changed; retained data was not adopted." + ); + } + return { generation: selected.generation }; +} + /** Eligible routes consult current native provenance even when reverting to the ownership plan. */ export function nativeReviewNeedsCompatibility(opts: { readonly saved: NativeProjectRun; @@ -94,15 +118,14 @@ export async function selectNativeProjectRestore(opts: { cwd: opts.projectRoot, args: ["graph", "restore-selection", "--run-id", saved.run, "--json"], }); + const selection = validateNativeRestoreSelection({ + selected, + saved, + namespace: opts.review.namespace, + }); if ( - !isRecord(selected) || - selected.run !== saved.run || - selected.owner !== saved.owner || - selected.namespace !== opts.review.namespace || - selected.namespace !== saved.namespace || - selected.plan !== saved.planId || - typeof selected.generation !== "string" || - !SHA.test(selected.generation) + opts.review.retainedGeneration !== undefined && + opts.review.retainedGeneration !== selection.generation ) { throw new Error( "Native restore selection changed; retained data was not adopted." @@ -122,7 +145,7 @@ export async function selectNativeProjectRestore(opts: { : null; return { run: saved.run, - flags: ["--expect-generation", selected.generation], + flags: ["--expect-generation", selection.generation], restoring: true, ...(sourceRevision ? { sourceRevision } : {}), }; diff --git a/src/backends/native-project-retained-images.ts b/src/backends/native-project-retained-images.ts new file mode 100644 index 000000000..f84b2336b --- /dev/null +++ b/src/backends/native-project-retained-images.ts @@ -0,0 +1,84 @@ +import { isRecord } from "../lib/guards.ts"; +import { inspectNativeProjectGraph } from "./native-project-inspect.ts"; +import { confirmedNativeRetainedGraph } from "./native-project-retained.ts"; +import type { NativeProjectRun } from "./native-project-run.ts"; +import type { + invokeNativeRuntime, + NativeRuntimeSelection, +} from "./native-runtime-client.ts"; + +const SHA = /^[a-f0-9]{64}$/; +const IMAGE = /^sha256:[a-f0-9]{64}$/; + +/** + * Restore unchanged declarations with their authenticated content IDs. Resolving + * a mutable tag again could change compute around retained data. This observation + * grants no restore authority: generation, input, source and resource checks still + * run in native code at admission. Changed original input keeps normal resolution. + */ +export async function selectNativeRetainedImages(opts: { + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly originalSha256: string; + readonly restore?: NativeProjectRun; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise> { + if (!opts.restore) { + return new Map(); + } + const observed = await inspectNativeProjectGraph({ + ...opts, + run: opts.restore.run, + }); + if ( + !( + confirmedNativeRetainedGraph(observed, opts.restore) && + isRecord(observed) && + isRecord(observed.receipt) + ) + ) { + throw new Error( + "Native retained image selection changed; no images were resolved." + ); + } + const { receipt } = observed; + if (receipt.normalized_input === undefined) { + return new Map(); + } + const normalized = receipt.normalized_input; + if ( + !isRecord(normalized) || + normalized.namespace !== opts.restore.namespace || + typeof normalized.original_compose_sha256 !== "string" || + !SHA.test(normalized.original_compose_sha256) || + typeof normalized.normalized_compose_sha256 !== "string" || + !SHA.test(normalized.normalized_compose_sha256) || + !SHA.test(opts.originalSha256) || + !isRecord(receipt.resources) + ) { + throw new Error( + "Native retained image provenance is invalid; no images were resolved." + ); + } + if (normalized.original_compose_sha256 !== opts.originalSha256) { + return new Map(); + } + const images = new Map(); + for (const [key, resource] of Object.entries(receipt.resources)) { + if (!isRecord(resource) || resource.kind !== "container") { + continue; + } + if ( + typeof resource.key !== "string" || + key !== `container:${resource.key}` || + typeof resource.image !== "string" || + !IMAGE.test(resource.image) + ) { + throw new Error( + "Native retained image identity is invalid; no images were resolved." + ); + } + images.set(resource.key, resource.image); + } + return images; +} diff --git a/src/backends/native-project-review.ts b/src/backends/native-project-review.ts index 384be4738..89ae4ef2c 100644 --- a/src/backends/native-project-review.ts +++ b/src/backends/native-project-review.ts @@ -1,39 +1,163 @@ -import { chmod, mkdtemp, rm, writeFile } from "node:fs/promises"; +import { chmod, mkdtemp, realpath, rm, writeFile } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join, relative } from "node:path"; import { isRecord } from "../lib/guards.ts"; -import { nativeProjectBranchArgs } from "./native-project-branch.ts"; +import { + nativeProjectBranchArgs, + prepareNativeProjectBranch, +} from "./native-project-branch.ts"; import type { NativeProjectInput } from "./native-project-input.ts"; +import { validateNativeRestoreSelection } from "./native-project-restore.ts"; +import type { + NativeProjectRun, + NativeProjectRunScope, +} from "./native-project-run.ts"; import { invokeNativeRuntime, type NativeRuntimeSelection, } from "./native-runtime-client.ts"; const DIGEST = /^[a-f0-9]{64}$/; +const CONTROL = /[\x00-\x1f\x7f]/; +const SERVICE = /^[a-zA-Z0-9_.-]{1,128}$/; export interface NativeProjectReview { readonly planId: string; readonly namespace: string; readonly report: Readonly>; + /** Exact legacy fallback selection; restore must observe it again after review. */ + readonly retainedGeneration?: string; /** Reuse unchanged for source publication and admission within the callback. */ readonly projectArgs: readonly string[]; } -/** - * Native code owns namespace and plan identities. Hold only public normalized - * bytes in a private temporary directory while the caller publishes/starts the - * exact plan. The native side rechecks original input identity at every effect. - * Managed values remain in the caller's in-memory input, outside this file. - */ -export async function withNativeProjectReview(opts: { +/** Authenticated active service selection; never a stopped restore generation. */ +export interface NativeActiveReviewProof { + readonly run: string; + readonly owner: string; + readonly namespace: string; + readonly plan: string; + readonly service: string; + readonly container: string; + readonly boot: string; + readonly generation: string; +} + +async function selectActiveReview(opts: { + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly retained: NativeProjectRun; + readonly service: string; + readonly invoke: typeof invokeNativeRuntime; +}): Promise { + const selected = await opts.invoke({ + runtime: opts.runtime, + cwd: opts.projectRoot, + args: [ + "graph", + "run-selection", + "--run-id", + opts.retained.run, + "--service", + opts.service, + "--json", + ], + }); + const { generation } = validateNativeRestoreSelection({ + selected, + saved: opts.retained, + namespace: opts.retained.namespace, + }); + if ( + !isRecord(selected) || + selected.ok !== true || + selected.service !== opts.service || + typeof selected.container !== "string" || + !DIGEST.test(selected.container) || + typeof selected.boot !== "string" || + selected.boot.length < 1 || + selected.boot.length > 128 || + CONTROL.test(selected.boot) + ) { + throw new Error( + "Native active review identity changed; the current graph was not stopped." + ); + } + return { + run: opts.retained.run, + owner: opts.retained.owner, + namespace: opts.retained.namespace, + plan: opts.retained.planId, + service: opts.service, + container: selected.container, + boot: selected.boot, + generation, + }; +} + +function activeReviewService(plan: Readonly>): string { + // `active` means profile-enabled. Native run-selection also verifies admitted + // completed initializers; selecting their identity never executes them again. + const services = isRecord(plan.services) ? plan.services : {}; + const service = Object.keys(services) + .sort() + .find( + (key) => + SERVICE.test(key) && + isRecord(services[key]) && + services[key].active === true + ); + if (!service) { + throw new Error( + "Native active review has no supported service authority; the current graph was not stopped." + ); + } + return service; +} + +/** Repeat the exact active selection after compatibility review, before cleanup eligibility. */ +export async function verifyNativeActiveReview(opts: { + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly retained: NativeProjectRun; + readonly proof?: NativeActiveReviewProof; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + if (!opts.proof) { + return; + } + const current = await selectActiveReview({ + ...opts, + service: opts.proof.service, + invoke: opts.invoke ?? invokeNativeRuntime, + }); + if (JSON.stringify(current) !== JSON.stringify(opts.proof)) { + throw new Error( + "Native active review identity changed; the current graph was not stopped." + ); + } +} + +/** Original-input identity chosen before any branch route normalization. */ +export interface NativeProjectReviewIdentity { + readonly namespace: string; + readonly branch: string | null; + readonly projectArgs: readonly string[]; + readonly retainedGeneration?: string; + readonly activeProof?: NativeActiveReviewProof; +} + +export async function selectNativeProjectReviewIdentity(opts: { readonly runtime: NativeRuntimeSelection; readonly projectRoot: string; readonly composeFile: string; readonly profiles?: readonly string[]; readonly branch?: string | null; readonly input: NativeProjectInput; - readonly run: (review: NativeProjectReview) => Promise; -}): Promise { + readonly retained?: NativeProjectRun; + readonly invoke?: typeof invokeNativeRuntime; + readonly retainedMode?: "stopped" | "active"; +}): Promise { const file = relative(opts.projectRoot, opts.composeFile); if ( !file || @@ -44,7 +168,7 @@ export async function withNativeProjectReview(opts: { "Native project review requires an owned original Compose file." ); } - const projectArgs = [ + let projectArgs = [ "--project", opts.projectRoot, "--file", @@ -52,7 +176,8 @@ export async function withNativeProjectReview(opts: { ...nativeProjectBranchArgs(opts.branch), ...(opts.profiles ?? []).flatMap((profile) => ["--profile", profile]), ]; - const original = await invokeNativeRuntime({ + const invoke = opts.invoke ?? invokeNativeRuntime; + const original = await invoke({ runtime: opts.runtime, cwd: opts.projectRoot, args: ["project", "plan", ...projectArgs, "--json"], @@ -67,7 +192,104 @@ export async function withNativeProjectReview(opts: { "Native project input changed before review; prepare it again." ); } - const namespace = original.plan.namespace; + let namespace = original.plan.namespace; + let branch = opts.branch ?? null; + let retainedGeneration: string | undefined; + let activeProof: NativeActiveReviewProof | undefined; + if (opts.branch && opts.retained && namespace !== opts.retained.namespace) { + if (opts.retainedMode === "active") { + activeProof = await selectActiveReview({ + runtime: opts.runtime, + projectRoot: opts.projectRoot, + retained: opts.retained, + service: activeReviewService(original.plan), + invoke, + }); + } else { + const selected = await invoke({ + runtime: opts.runtime, + cwd: opts.projectRoot, + args: [ + "graph", + "restore-selection", + "--run-id", + opts.retained.run, + "--json", + ], + }); + retainedGeneration = validateNativeRestoreSelection({ + selected, + saved: opts.retained, + namespace: opts.retained.namespace, + }).generation; + } + const unbranchedArgs = [ + "--project", + opts.projectRoot, + "--file", + file, + ...(opts.profiles ?? []).flatMap((profile) => ["--profile", profile]), + ]; + const unbranched = await invoke({ + runtime: opts.runtime, + cwd: opts.projectRoot, + args: ["project", "plan", ...unbranchedArgs, "--json"], + }); + const canonicalProject = await realpath(opts.projectRoot); + if ( + !(isRecord(unbranched) && isRecord(unbranched.plan)) || + unbranched.plan.namespace !== opts.retained.namespace || + original.plan.source !== canonicalProject || + unbranched.plan.source !== canonicalProject || + unbranched.plan.compose_sha256 !== opts.input.originalSha256 + ) { + throw new Error( + "Native retained project review changed; retained data was not adopted." + ); + } + projectArgs = unbranchedArgs; + namespace = opts.retained.namespace; + branch = null; + } + return { + namespace, + branch, + projectArgs, + ...(retainedGeneration ? { retainedGeneration } : {}), + ...(activeProof ? { activeProof } : {}), + }; +} + +/** + * Native code owns namespace and plan identities. Hold only public normalized + * bytes in a private temporary directory while the caller publishes/starts the + * exact plan. The native side rechecks original input identity at every effect. + * Managed values remain in the caller's in-memory input, outside this file. + */ +export async function withNativeProjectReview(opts: { + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly composeFile: string; + readonly profiles?: readonly string[]; + readonly branch?: string | null; + readonly input: NativeProjectInput; + readonly retained?: NativeProjectRun; + readonly invoke?: typeof invokeNativeRuntime; + readonly retainedMode?: "stopped" | "active"; + readonly reviewIdentity?: NativeProjectReviewIdentity; + readonly run: (review: NativeProjectReview) => Promise; +}): Promise { + const identity = await selectNativeProjectReviewIdentity(opts); + if ( + opts.reviewIdentity && + JSON.stringify(opts.reviewIdentity) !== JSON.stringify(identity) + ) { + throw new Error( + "Native project review selection changed; retained data was not adopted." + ); + } + const { namespace, projectArgs, retainedGeneration } = identity; + const invoke = opts.invoke ?? invokeNativeRuntime; const directory = await mkdtemp(join(tmpdir(), "hack-native-review-")); try { await chmod(directory, 0o700); @@ -85,7 +307,7 @@ export async function withNativeProjectReview(opts: { "--expect-namespace", namespace, ]; - const report = await invokeNativeRuntime({ + const report = await invoke({ runtime: opts.runtime, cwd: opts.projectRoot, args: ["project", "plan", ...normalizedArgs, "--json"], @@ -102,9 +324,56 @@ export async function withNativeProjectReview(opts: { planId: report.plan_id, namespace, report, + ...(retainedGeneration ? { retainedGeneration } : {}), projectArgs: normalizedArgs, }); } finally { await rm(directory, { recursive: true, force: true }); } } + +/** Choose retained native identity before rewriting any adapted route labels. */ +export async function prepareNativeReviewBranch(opts: { + readonly runtime: NativeRuntimeSelection; + readonly phase?: "before-runtime" | "after-runtime"; + readonly scope: NativeProjectRunScope; + readonly composeFile: string; + readonly profiles?: readonly string[]; + readonly input: NativeProjectInput; + readonly retained?: NativeProjectRun; + readonly retainedMode?: "stopped" | "active"; + readonly invoke?: typeof invokeNativeRuntime; + readonly select?: typeof selectNativeProjectReviewIdentity; +}): Promise<{ + input: NativeProjectInput; + identity?: NativeProjectReviewIdentity; +}> { + nativeProjectBranchArgs(opts.scope.branch); + const deferred = Boolean(opts.retained && opts.scope.branch); + if ( + (opts.phase === "before-runtime" && deferred) || + (opts.phase === "after-runtime" && !deferred) + ) { + return { input: opts.input }; + } + const identity = + opts.retained && opts.scope.branch + ? await (opts.select ?? selectNativeProjectReviewIdentity)({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + composeFile: opts.composeFile, + profiles: opts.profiles, + branch: opts.scope.branch, + input: opts.input, + retained: opts.retained, + retainedMode: opts.retainedMode, + invoke: opts.invoke, + }) + : undefined; + const input = await prepareNativeProjectBranch({ + input: opts.input, + composeFile: opts.composeFile, + scope: identity ? { ...opts.scope, branch: identity.branch } : opts.scope, + }); + return { input, ...(identity ? { identity } : {}) }; +} diff --git a/src/backends/native-project-run.ts b/src/backends/native-project-run.ts index e48916f04..c2825cbe7 100644 --- a/src/backends/native-project-run.ts +++ b/src/backends/native-project-run.ts @@ -16,10 +16,16 @@ import { isNativeProjectFinalizationToken, type NativeProjectFinalizationToken, } from "./native-project-finalization.ts"; +import { inspectNativeProjectGraph } from "./native-project-inspect.ts"; +import { + invokeNativeRuntime, + type NativeRuntimeSelection, +} from "./native-runtime-client.ts"; const HEX32 = /^[a-f0-9]{32}$/; const HEX64 = /^[a-f0-9]{64}$/; const LIMIT = 8192; +const RECOVERY_SCHEMA = "hack.native-project-run-filesystem-recovery/v1"; // Equivalent to normalizeEnvConfigName(value) === value, with a bounded length. const ENV_NAME = /^[a-z0-9]+(?:-[a-z0-9]+)*$/; const AWS_PROFILE = /^[A-Za-z0-9_+=,.@-]{1,128}$/; @@ -390,6 +396,548 @@ function record(value: unknown, identity: unknown): NativeProjectRun { } return value.run; } + +type DirectoryIdentity = { readonly dev: number; readonly ino: number }; +type RunScopeIdentity = Awaited>; +type RecoveryPaths = Awaited>; + +export type NativeProjectRunFilesystemInspection = { + readonly schema: typeof RECOVERY_SCHEMA; + readonly selectionSha256: string; + readonly mappingSha256: string; + readonly run: NativeProjectRun; + readonly oldDevice: number; + readonly newDevice: number; + readonly qualification: "explicit-legacy-rebind-original-volume-continuity-unproven"; +}; + +export type NativeProjectRunFilesystemRecovery = + NativeProjectRunFilesystemInspection & { + readonly repaired: true; + readonly auditPath: string; + }; + +function sameIdentity( + value: unknown, + expected: DirectoryIdentity, + oldDevice: number +): boolean { + return ( + isRecord(value) && + Object.keys(value).sort().join() === "dev,ino" && + typeof value.dev === "number" && + Number.isSafeInteger(value.dev) && + value.dev === oldDevice && + value.ino === expected.ino + ); +} + +async function readRecoveryFile(path: string): Promise<{ + bytes: Buffer; + value: unknown; + identity: DirectoryIdentity; +}> { + const fd = await open( + path, + constants.O_RDONLY | constants.O_NONBLOCK | constants.O_NOFOLLOW + ); + try { + const before = await fd.stat(); + if ( + !before.isFile() || + before.nlink !== 1 || + before.size < 1 || + before.size > LIMIT || + before.uid !== process.getuid?.() || + (before.mode & 0o777) !== 0o600 + ) { + throw refused(); + } + const bytes = Buffer.alloc(before.size); + const result = await fd.read(bytes, 0, bytes.length, 0); + const after = await fd.stat(); + const named = await lstat(path); + if ( + result.bytesRead !== bytes.length || + before.dev !== after.dev || + before.ino !== after.ino || + before.size !== after.size || + before.mtimeMs !== after.mtimeMs || + before.ctimeMs !== after.ctimeMs || + named.dev !== after.dev || + named.ino !== after.ino || + !named.isFile() || + named.isSymbolicLink() || + after.nlink !== 1 || + after.uid !== process.getuid?.() || + (after.mode & 0o777) !== 0o600 + ) { + throw refused(); + } + const value: unknown = JSON.parse( + new TextDecoder("utf-8", { fatal: true }).decode(bytes) + ); + return { bytes, value, identity: { dev: after.dev, ino: after.ino } }; + } finally { + await fd.close(); + } +} + +async function preserveRecoveryAudit( + path: string, + expected: string +): Promise { + const limit = LIMIT * 2 + 4096; + if (Buffer.byteLength(expected) > limit) { + throw refused(); + } + try { + await write(path, expected); + return; + } catch (error) { + if (!isRecord(error) || error.code !== "EEXIST") { + throw error; + } + } + const fd = await open( + path, + constants.O_RDONLY | constants.O_NONBLOCK | constants.O_NOFOLLOW + ); + try { + const before = await fd.stat(); + if ( + !before.isFile() || + before.nlink !== 1 || + before.uid !== process.getuid?.() || + (before.mode & 0o777) !== 0o600 || + before.size !== Buffer.byteLength(expected) || + before.size > limit + ) { + throw refused(); + } + const bytes = Buffer.alloc(before.size); + const read = await fd.read(bytes, 0, bytes.length, 0); + const after = await fd.stat(); + const named = await lstat(path); + if ( + read.bytesRead !== bytes.length || + !bytes.equals(Buffer.from(expected)) || + before.dev !== after.dev || + before.ino !== after.ino || + before.mtimeMs !== after.mtimeMs || + before.ctimeMs !== after.ctimeMs || + named.dev !== after.dev || + named.ino !== after.ino || + !named.isFile() || + named.nlink !== 1 + ) { + throw refused(); + } + } finally { + await fd.close(); + } +} + +function oldMapping( + value: unknown, + current: RunScopeIdentity +): { + readonly oldDevice: number; + readonly run: NativeProjectRun; + readonly next: string; +} { + if ( + !isRecord(value) || + Object.keys(value).sort().join() !== "run,scope,version" || + value.version !== 1 || + !valid(value.run) || + !isRecord(value.scope) || + Object.keys(value.scope).sort().join() !== + "branch,dirIdentity,homeIdentity,nativeHome,projectDir,projectRoot,rootIdentity" + ) { + throw refused(); + } + const stored = value.scope; + const oldRoot = stored.rootIdentity; + if ( + !isRecord(oldRoot) || + typeof oldRoot.dev !== "number" || + !Number.isSafeInteger(oldRoot.dev) || + oldRoot.dev <= 0 + ) { + throw refused(); + } + const oldDevice = oldRoot.dev; + if ( + oldDevice === current.rootIdentity.dev || + current.rootIdentity.dev !== current.dirIdentity.dev || + current.rootIdentity.dev !== current.homeIdentity.dev || + stored.projectRoot !== current.projectRoot || + stored.projectDir !== current.projectDir || + stored.nativeHome !== current.nativeHome || + stored.branch !== current.branch || + !sameIdentity(oldRoot, current.rootIdentity, oldDevice) || + !sameIdentity(stored.dirIdentity, current.dirIdentity, oldDevice) || + !sameIdentity(stored.homeIdentity, current.homeIdentity, oldDevice) + ) { + throw refused(); + } + const next = JSON.stringify({ ...value, scope: current }); + if ( + Buffer.byteLength(next) > LIMIT || + JSON.stringify(value.run) !== JSON.stringify(JSON.parse(next).run) + ) { + throw refused(); + } + return { oldDevice, run: value.run, next }; +} + +async function absentPath(path: string): Promise { + try { + await lstat(path); + } catch (error) { + if (isRecord(error) && error.code === "ENOENT") { + return; + } + throw refused(); + } + throw refused(); +} + +async function recoveryLocksAbsent(p: RecoveryPaths): Promise { + const restartFile = `${p.file.slice(0, -5)}.restart.json`; + const restartLock = `${p.lock.slice(0, -5)}.restart.lock`; + await Promise.all([ + absentPath(p.lock), + absentPath(`${p.lock}.operation`), + absentPath(restartFile), + absentPath(restartLock), + absentPath(`${restartLock}.operation`), + ]); +} + +async function recoveryDirectories(p: RecoveryPaths): Promise { + for (const path of [ + p.identity.projectRoot, + p.identity.projectDir, + p.identity.nativeHome, + join(p.identity.projectDir, ".internal"), + p.root, + ]) { + const metadata = await lstat(path); + if ( + !metadata.isDirectory() || + metadata.isSymbolicLink() || + metadata.uid !== process.getuid?.() || + (metadata.mode & 0o022) !== 0 || + (path === p.root && (metadata.mode & 0o777) !== 0o700) + ) { + throw refused(); + } + } +} + +async function heldDirectory(path: string) { + const fd = await open(path, constants.O_RDONLY | constants.O_NOFOLLOW); + try { + const metadata = await fd.stat(); + if ( + !metadata.isDirectory() || + metadata.uid !== process.getuid?.() || + (metadata.mode & 0o777) !== 0o700 + ) { + throw refused(); + } + const verify = async () => { + const named = await lstat(path); + const held = await fd.stat(); + if ( + !named.isDirectory() || + named.isSymbolicLink() || + named.dev !== metadata.dev || + named.ino !== metadata.ino || + held.dev !== metadata.dev || + held.ino !== metadata.ino || + named.uid !== metadata.uid || + (named.mode & 0o777) !== 0o700 + ) { + throw refused(); + } + }; + await verify(); + return { verify, close: () => fd.close() }; + } catch (error) { + await fd.close(); + throw error; + } +} + +async function releaseHeldDirectory( + path: string, + held: Awaited> +) { + try { + await held.verify(); + await rmdir(path); + } finally { + await held.close(); + } +} + +async function verifySelectedMappingPath( + path: string, + selected: Buffer, + identity: DirectoryIdentity +): Promise<() => Promise> { + const fd = await open( + path, + constants.O_RDONLY | constants.O_NONBLOCK | constants.O_NOFOLLOW + ); + try { + const before = await fd.stat(); + if ( + !before.isFile() || + before.nlink !== 1 || + before.uid !== process.getuid?.() || + (before.mode & 0o777) !== 0o600 || + before.size !== selected.length || + before.dev !== identity.dev || + before.ino !== identity.ino + ) { + throw refused(); + } + const bytes = Buffer.alloc(selected.length); + const read = await fd.read(bytes, 0, bytes.length, 0); + const after = await fd.stat(); + const named = await lstat(path); + if ( + read.bytesRead !== bytes.length || + !bytes.equals(selected) || + before.dev !== after.dev || + before.ino !== after.ino || + before.mtimeMs !== after.mtimeMs || + before.ctimeMs !== after.ctimeMs || + named.dev !== after.dev || + named.ino !== after.ino || + named.nlink !== 1 || + named.uid !== process.getuid?.() || + (named.mode & 0o777) !== 0o600 + ) { + throw refused(); + } + return () => fd.close(); + } catch (error) { + await fd.close(); + throw error; + } +} + +async function nativeRecoveryAuthority(opts: { + readonly p: RecoveryPaths; + readonly run: NativeProjectRun; + readonly oldDevice: number; + readonly runtime: NativeRuntimeSelection; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + const invoke = opts.invoke ?? invokeNativeRuntime; + const [graph, status] = await Promise.all([ + inspectNativeProjectGraph({ + runtime: opts.runtime, + projectRoot: opts.p.identity.projectRoot, + run: opts.run.run, + invoke, + }), + invoke({ + runtime: opts.runtime, + cwd: opts.p.identity.projectRoot, + args: ["runtime", "status", "--json"], + timeoutMs: 30_000, + }), + ]); + const receipt = isRecord(graph) ? graph.receipt : undefined; + const source = isRecord(receipt) ? receipt.source : undefined; + const shared = isRecord(source) ? source.shared : undefined; + const share = isRecord(status) ? status.project_share : undefined; + if ( + !isRecord(graph) || + graph.journal_incomplete !== false || + !isRecord(receipt) || + receipt.run !== opts.run.run || + receipt.owner !== opts.run.owner || + receipt.namespace !== opts.run.namespace || + receipt.plan_id !== opts.run.planId || + !["ready-observed", "stopped-data-retained"].includes( + String(receipt.phase) + ) || + !isRecord(status) || + status.phase !== "running" || + status.process_alive !== true || + status.persistent_disks_identified !== true || + !isRecord(share) || + share.project !== opts.p.identity.projectRoot || + share.device !== opts.p.identity.rootIdentity.dev || + share.inode !== opts.p.identity.rootIdentity.ino || + share.unfiltered_source !== true || + !isRecord(shared) || + shared.project !== share.project || + shared.guest_path !== share.guest_path || + shared.inode !== share.inode || + shared.device !== opts.oldDevice || + shared.unfiltered_source !== true + ) { + throw refused(); + } +} + +async function recoverySelection(opts: { + readonly scope: NativeProjectRunScope; + readonly runtime: NativeRuntimeSelection; + readonly invoke?: typeof invokeNativeRuntime; + readonly heldLocks?: boolean; +}): Promise<{ + readonly inspection: NativeProjectRunFilesystemInspection; + readonly p: RecoveryPaths; + readonly bytes: Buffer; + readonly fileIdentity: DirectoryIdentity; + readonly next: string; +}> { + const p = await paths(opts.scope, false); + await recoveryDirectories(p); + if (!opts.heldLocks) { + await recoveryLocksAbsent(p); + } + const { bytes, value, identity } = await readRecoveryFile(p.file); + const mapping = oldMapping(value, p.identity); + await nativeRecoveryAuthority({ + p, + run: mapping.run, + oldDevice: mapping.oldDevice, + runtime: opts.runtime, + invoke: opts.invoke, + }); + const mappingSha256 = createHash("sha256").update(bytes).digest("hex"); + const selected = { + schema: RECOVERY_SCHEMA, + mappingSha256, + fileIdentity: identity, + scope: p.identity, + oldDevice: mapping.oldDevice, + run: mapping.run, + }; + const inspection: NativeProjectRunFilesystemInspection = { + schema: RECOVERY_SCHEMA, + selectionSha256: createHash("sha256") + .update(JSON.stringify(selected)) + .digest("hex"), + mappingSha256, + run: mapping.run, + oldDevice: mapping.oldDevice, + newDevice: p.identity.rootIdentity.dev, + qualification: "explicit-legacy-rebind-original-volume-continuity-unproven", + }; + return { inspection, p, bytes, fileIdentity: identity, next: mapping.next }; +} + +/** Inspect an exact legacy device renumbering without changing frontend or native state. */ +export async function inspectNativeProjectRunFilesystemRecovery(opts: { + readonly scope: NativeProjectRunScope; + readonly runtime: NativeRuntimeSelection; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + return (await recoverySelection(opts)).inspection; +} + +/** Explicitly rebind only three mapping device numbers; graph history is untouched. */ +export async function recoverNativeProjectRunFilesystem(opts: { + readonly scope: NativeProjectRunScope; + readonly runtime: NativeRuntimeSelection; + readonly expectSelection: string; + readonly acceptLegacyDeviceRebind: true; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + if ( + !HEX64.test(opts.expectSelection) || + opts.acceptLegacyDeviceRebind !== true + ) { + throw refused(); + } + const initial = await recoverySelection(opts); + if (initial.inspection.selectionSha256 !== opts.expectSelection) { + throw refused(); + } + const restartLock = `${initial.p.lock.slice(0, -5)}.restart.lock.operation`; + await mkdir(restartLock, { mode: 0o700 }); + const restartHeld = await heldDirectory(restartLock); + try { + await mkdir(initial.p.lock, { mode: 0o700 }); + const mappingHeld = await heldDirectory(initial.p.lock); + try { + const selected = await recoverySelection({ ...opts, heldLocks: true }); + if ( + selected.inspection.selectionSha256 !== opts.expectSelection || + selected.p.root !== initial.p.root || + selected.p.file !== initial.p.file || + !selected.bytes.equals(initial.bytes) || + selected.fileIdentity.dev !== initial.fileIdentity.dev || + selected.fileIdentity.ino !== initial.fileIdentity.ino + ) { + throw refused(); + } + await absentPath(`${selected.p.file.slice(0, -5)}.restart.json`); + const audit = `${selected.p.file.slice(0, -5)}.filesystem-recovery-${selected.inspection.mappingSha256}.json`; + const auditText = JSON.stringify({ + version: 1, + selectionSha256: selected.inspection.selectionSha256, + mappingSha256: selected.inspection.mappingSha256, + mapping: selected.bytes.toString("utf8"), + }); + await preserveRecoveryAudit(audit, auditText); + await sync(selected.p.root); + const final = await recoverySelection({ ...opts, heldLocks: true }); + if ( + final.inspection.selectionSha256 !== opts.expectSelection || + !final.bytes.equals(selected.bytes) || + final.fileIdentity.dev !== selected.fileIdentity.dev || + final.fileIdentity.ino !== selected.fileIdentity.ino + ) { + throw refused(); + } + const temporary = join(selected.p.root, `${randomUUID()}.tmp`); + try { + await write(temporary, selected.next); + await Promise.all([restartHeld.verify(), mappingHeld.verify()]); + const closeMapping = await verifySelectedMappingPath( + selected.p.file, + selected.bytes, + selected.fileIdentity + ); + try { + await rename(temporary, selected.p.file); + } finally { + await closeMapping(); + } + await sync(selected.p.root); + } finally { + await unlink(temporary).catch((error: unknown) => { + if (!isRecord(error) || error.code !== "ENOENT") { + throw error; + } + }); + } + if ( + JSON.stringify(await loadNativeProjectRun(opts.scope)) !== + JSON.stringify(selected.inspection.run) + ) { + throw refused(); + } + return { ...selected.inspection, repaired: true, auditPath: audit }; + } finally { + await releaseHeldDirectory(initial.p.lock, mappingHeld); + } + } finally { + await releaseHeldDirectory(restartLock, restartHeld); + } +} /** Establish excluded metadata before source identity is reviewed. */ export async function prepareNativeProjectRunStorage( opts: NativeProjectRunScope diff --git a/src/backends/native-project-start.ts b/src/backends/native-project-start.ts index d900650a1..8b5d6a1e2 100644 --- a/src/backends/native-project-start.ts +++ b/src/backends/native-project-start.ts @@ -14,7 +14,6 @@ import { preparedBaseArguments, } from "./native-prepared-base.ts"; import { prepareNativeProjectAdaptation } from "./native-project-adaptation.ts"; -import { prepareNativeProjectBranch } from "./native-project-branch.ts"; import { nativeSharedSourceFlags, nativeStartCacheSource, @@ -24,23 +23,33 @@ import { prepareNativeDependencyServices, readNativeHostDependencies, } from "./native-project-dependencies.ts"; -import { beginNativeProjectFinalization } from "./native-project-finalization.ts"; +import { + beginNativeProjectFinalization, + captureNativeProjectFinalization, + waitNativeProjectFinalization, +} from "./native-project-finalization.ts"; import { isNativeHttpsProbePath } from "./native-project-https.ts"; import { type NativeProjectInput, prepareNativeProjectInput, } from "./native-project-input.ts"; import { inspectNativeProjectGraph } from "./native-project-inspect.ts"; +import { recoverNativeInterruptedStartupCleanup } from "./native-project-interrupted-cleanup.ts"; import { validateNativeAllowedHosts } from "./native-project-network.ts"; import { serveNativeProjectGraph } from "./native-project-process.ts"; import { selectNativeProjectRestore } from "./native-project-restore.ts"; import { confirmedNativeRetainedGraph } from "./native-project-retained.ts"; +import { selectNativeRetainedImages } from "./native-project-retained-images.ts"; import { preflightNativeRetainedStartup, verifyNativeResumedRetainedGraph, verifyNativeRetainedMapping, } from "./native-project-retained-startup.ts"; -import { withNativeProjectReview } from "./native-project-review.ts"; +import { + prepareNativeReviewBranch, + selectNativeProjectReviewIdentity, + withNativeProjectReview, +} from "./native-project-review.ts"; import { hasOnlyNativeSupportedLabels, nativeBridgeCapacity, @@ -211,6 +220,8 @@ type Dependencies = { prepareStorage: typeof prepareNativeProjectRunStorage; adaptAws: typeof adaptNativeAwsEnvironment; review: typeof withNativeProjectReview; + selectReview: typeof selectNativeProjectReviewIdentity; + retainedImages: typeof selectNativeRetainedImages; serve: typeof serveNativeProjectGraph; https: typeof acquireNativeHttpsLease; recoverHttps: typeof recoverNativeHttpsLease; @@ -220,12 +231,16 @@ type Dependencies = { save: typeof saveNativeProjectRun; remove: typeof removeNativeProjectRun; finalization: typeof beginNativeProjectFinalization; + captureFinalization: typeof captureNativeProjectFinalization; + waitFinalization: typeof waitNativeProjectFinalization; }; const DEFAULTS: Dependencies = { prepare: prepareNativeProjectInput, prepareStorage: prepareNativeProjectRunStorage, adaptAws: adaptNativeAwsEnvironment, review: withNativeProjectReview, + selectReview: selectNativeProjectReviewIdentity, + retainedImages: selectNativeRetainedImages, serve: serveNativeProjectGraph, https: acquireNativeHttpsLease, recoverHttps: recoverNativeHttpsLease, @@ -235,6 +250,8 @@ const DEFAULTS: Dependencies = { save: saveNativeProjectRun, remove: removeNativeProjectRun, finalization: beginNativeProjectFinalization, + captureFinalization: captureNativeProjectFinalization, + waitFinalization: waitNativeProjectFinalization, }; function refused(): Error { return new Error( @@ -273,6 +290,43 @@ export function prepareNativeProjectServices( } return result; } +/** Keep image acquisition separate from graph admission and preserve cancellation between requests. */ +async function pinNativeProjectImages(opts: { + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly specs: Record>; + readonly retainedImages: ReadonlyMap; + readonly signal: AbortSignal; + readonly invoke: typeof invokeNativeRuntime; +}): Promise { + for (const [name, spec] of Object.entries(opts.specs)) { + if (opts.signal.aborted) { + throw refused(); + } + const retainedImage = opts.retainedImages.get(name); + if (retainedImage) { + spec.image = retainedImage; + continue; + } + const image = String(spec.image); + if (IMAGE.test(image)) { + continue; + } + const ensured = await opts.invoke({ + runtime: opts.runtime, + cwd: opts.projectRoot, + args: ["runtime", "ensure-image", "--reference", image, "--json"], + }); + if ( + !isRecord(ensured) || + typeof ensured.image_id !== "string" || + !IMAGE.test(ensured.image_id) + ) { + throw refused(); + } + spec.image = ensured.image_id; + } +} function readiness( specs: Record>, initializers: ReadonlySet, @@ -511,6 +565,39 @@ function requireConfirmedCleanup( throw cleanupUnconfirmed(startupFailure, nativeCode); } } + +async function confirmForegroundCleanup(opts: { + readonly final: unknown; + readonly runtime: NativeRuntimeSelection; + readonly projectRoot: string; + readonly run: string; + readonly namespace: string; + readonly planId: string; + readonly invoke: typeof invokeNativeRuntime; + readonly startupFailure: unknown; + readonly nativeCode: string | undefined; +}): Promise<{ final: unknown; owned: NativeProjectRun }> { + let final = opts.final; + let owned = authoritative(final, opts.run, opts.namespace, opts.planId); + try { + const recovered = await recoverNativeInterruptedStartupCleanup({ + runtime: opts.runtime, + projectRoot: opts.projectRoot, + run: owned, + snapshot: final, + invoke: opts.invoke, + }); + if (recovered !== null) { + final = recovered; + owned = authoritative(final, opts.run, opts.namespace, opts.planId); + } + } catch { + throw cleanupUnconfirmed(opts.startupFailure, opts.nativeCode); + } + requireConfirmedCleanup(final, opts.startupFailure, opts.nativeCode); + return { final, owned }; +} + async function refusePendingFreshStart( restore: NativeProjectRun | undefined, scope: NativeProjectRunScope, @@ -664,6 +751,26 @@ async function acknowledgeFinalization( await finalization?.complete(); } } +async function requireAcknowledgedRetainedFinalization(opts: { + readonly scope: NativeProjectRunScope; + readonly run: NativeProjectRun; + readonly capture: typeof captureNativeProjectFinalization; + readonly wait: typeof waitNativeProjectFinalization; +}): Promise { + try { + const token = await opts.capture({ scope: opts.scope, run: opts.run }); + await opts.wait({ + scope: opts.scope, + run: opts.run, + token, + timeoutMs: 1, + }); + } catch { + throw new Error( + "Native retained frontend finalization is unconfirmed or its ownership changed; no hooks, runtime start or HTTPS owner were requested." + ); + } +} /** Explicit unfiltered project sharing; foreground only. Native refusal never falls back to Compose. */ export async function startNativeProject(opts: { readonly runtime: NativeRuntimeSelection; @@ -709,6 +816,20 @@ export async function startNativeProject(opts: { signal: opts.signal, }); requireActiveStartup(opts.signal); + if (retained) { + await requireAcknowledgedRetainedFinalization({ + scope: opts.scope, + run: retained, + capture: deps.captureFinalization, + wait: deps.waitFinalization, + }); + await verifyNativeRetainedMapping({ + scope: opts.scope, + run: retained, + load: deps.load, + }); + requireActiveStartup(opts.signal); + } const profiles = selection.profiles; await preflightNativeProjectSource({ runtime: opts.runtime, @@ -735,11 +856,16 @@ export async function startNativeProject(opts: { input, path: opts.adaptationFile, }); - input = await prepareNativeProjectBranch({ - input, - scope: opts.scope, - composeFile: opts.composeFile, - }); + input = ( + await prepareNativeReviewBranch({ + runtime: opts.runtime, + scope: opts.scope, + composeFile: opts.composeFile, + input, + retained: restore, + phase: "before-runtime", + }) + ).input; let specs = prepareNativeProjectServices( input, opts.dependencyFile !== undefined @@ -854,28 +980,43 @@ export async function startNativeProject(opts: { run: retained, invoke: deps.invoke, }); - for (const spec of Object.values(specs)) { - if (controller.signal.aborted) { - throw refused(); - } - const image = String(spec.image); - if (IMAGE.test(image)) { - continue; - } - const ensured = await deps.invoke({ - runtime: opts.runtime, - cwd: opts.scope.projectRoot, - args: ["runtime", "ensure-image", "--reference", image, "--json"], - }); - if ( - !isRecord(ensured) || - typeof ensured.image_id !== "string" || - !IMAGE.test(ensured.image_id) - ) { - throw refused(); - } - spec.image = ensured.image_id; - } + const reviewedBranch = await prepareNativeReviewBranch({ + runtime: opts.runtime, + scope: opts.scope, + composeFile: opts.composeFile, + profiles, + input, + retained: restore, + invoke: deps.invoke, + select: deps.selectReview, + phase: "after-runtime", + }); + input = reviewedBranch.input; + const reviewIdentity = reviewedBranch.identity; + requireActiveStartup(controller.signal); + specs = prepareNativeProjectServices( + input, + opts.dependencyFile !== undefined + ); + prepareNativeDependencyServices({ + dependencies: hostDependencies, + services: specs, + }); + const retainedImages = await deps.retainedImages({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + originalSha256: input.originalSha256, + restore, + invoke: deps.invoke, + }); + await pinNativeProjectImages({ + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, + specs, + retainedImages, + signal: controller.signal, + invoke: deps.invoke, + }); const compose = JSON.parse(input.normalizedComposeJson); compose.services = specs; const pinned = { ...input, normalizedComposeJson: JSON.stringify(compose) }; @@ -886,6 +1027,8 @@ export async function startNativeProject(opts: { composeFile: opts.composeFile, profiles, branch: opts.scope.branch, + retained: restore, + reviewIdentity, input: pinned, run: async (review) => { requireEnrollmentCompatible(review.report.plan); @@ -1087,18 +1230,23 @@ export async function startNativeProject(opts: { nativeExitCode ); } - const owned = authoritative( + const confirmed = await confirmForegroundCleanup({ final, + runtime: opts.runtime, + projectRoot: opts.scope.projectRoot, run, - review.namespace, - restore?.planId ?? review.planId - ); - requireConfirmedCleanup(final, serveFailure, nativeExitCode); + namespace: review.namespace, + planId: restore?.planId ?? review.planId, + invoke: deps.invoke, + startupFailure: serveFailure, + nativeCode: nativeExitCode, + }); + final = confirmed.final; graphCleanupConfirmed = true; await retireRemovedMapping({ final, mapping, - owned, + owned: confirmed.owned, scope: opts.scope, remove: deps.remove, }); diff --git a/src/backends/native-runtime-client.ts b/src/backends/native-runtime-client.ts index 3d5aa8dde..35fd3833b 100644 --- a/src/backends/native-runtime-client.ts +++ b/src/backends/native-runtime-client.ts @@ -10,13 +10,16 @@ interface NativeFailure { /** Structured native codes remain safe to inspect without exposing subprocess diagnostics. */ export class NativeRuntimeRequestError extends Error { readonly nativeCode: string | undefined; + readonly nativeCauseCode: string | undefined; constructor(opts: { readonly message: string; readonly nativeCode?: string; + readonly nativeCauseCode?: string; }) { super(opts.message); this.nativeCode = opts.nativeCode; + this.nativeCauseCode = opts.nativeCauseCode; } } @@ -163,6 +166,7 @@ function completionResponse( ? "Native runtime request timed out; inspect owned state before retrying." : `Native runtime request failed${failure ? ` (${failure.code}${failure.causeCode ? `: ${failure.causeCode}` : ""})` : ""}; inspect owned state before retrying.`, nativeCode: timedOut ? undefined : failure?.code, + nativeCauseCode: timedOut ? undefined : failure?.causeCode, }); } try { @@ -282,7 +286,9 @@ async function readNativeFailure( ) { return { code: value.code, - ...(value.code === "graph_one_off_failed" && + ...(["graph_one_off_failed", "graph_owner_recovery"].includes( + value.code + ) && typeof value.cause_code === "string" && ERROR_CODE.test(value.cause_code) ? { causeCode: value.cause_code } diff --git a/src/commands/doctor.ts b/src/commands/doctor.ts index 7b3355051..883b609ed 100644 --- a/src/commands/doctor.ts +++ b/src/commands/doctor.ts @@ -8,6 +8,10 @@ import { checkLegacyProjectAgentArtifacts, checkLegacyUserAgentArtifacts, } from "../agents/legacy-artifacts.ts"; +import { + parseNativeRunMappingRecoveryOptions, + runNativeProjectMappingCommand, +} from "../backends/native-project-mapping-command.ts"; import { type NativeRuntimeSelection, resolveNativeRuntimeSelection, @@ -19,7 +23,7 @@ import { defineOption, withHandler, } from "../cli/command.ts"; -import { optJson, optPath } from "../cli/options.ts"; +import { optBranch, optJson, optPath } from "../cli/options.ts"; import { DEFAULT_CADDY_IP, DEFAULT_HOST_DNS_IP, @@ -208,6 +212,29 @@ const doctorOptions = [ optJson, optBrowserUrl, optBrowserResult, + optBranch, + defineOption({ + name: "nativeRunMapping", + type: "string", + long: "--native-run-mapping", + valueHint: "inspect|repair", + description: + "Inspect or explicitly repair a native run mapping after filesystem device renumbering", + } as const), + defineOption({ + name: "expectSelection", + type: "string", + long: "--expect-selection", + valueHint: "<64-hex>", + description: "Require the exact run-mapping recovery inspection selection", + } as const), + defineOption({ + name: "acceptLegacyDeviceRebind", + type: "boolean", + long: "--accept-legacy-device-rebind", + description: + "Explicitly accept legacy migration without proof of original filesystem volume continuity", + } as const), ] as const; const doctorPositionals = [] as const; @@ -332,9 +359,49 @@ async function maybeRunDomainMigration( return null; } +async function maybeRunNativeMappingRecovery( + args: Parameters>[0]["args"] +): Promise { + const mapping = parseNativeRunMappingRecoveryOptions({ + action: args.options.nativeRunMapping, + expectSelection: args.options.expectSelection, + acceptLegacyDeviceRebind: args.options.acceptLegacyDeviceRebind, + branch: args.options.branch, + otherOptions: Boolean( + args.options.fix || + args.options.migrateEnvConfig || + args.options.domainMigration || + args.options.browserUrl || + args.options.browserResult + ), + }); + if (mapping) { + const runtime = resolveNativeRuntimeSelection(); + if (!runtime) { + throw new CliUsageError( + "Native run-mapping recovery requires an explicitly selected native runtime." + ); + } + return await runNativeProjectMappingCommand({ + selection: mapping, + runtime, + startDir: args.options.path + ? resolve(process.cwd(), args.options.path) + : process.cwd(), + branch: args.options.branch, + json: args.options.json === true, + }); + } + return null; +} + const handleDoctor: CommandHandlerFor = async ({ args, }): Promise => { + const mappingResult = await maybeRunNativeMappingRecovery(args); + if (mappingResult !== null) { + return mappingResult; + } const domainResult = await maybeRunDomainMigration(args); if (domainResult !== null) { return domainResult; diff --git a/src/commands/project.ts b/src/commands/project.ts index da34b6541..1c29586e8 100644 --- a/src/commands/project.ts +++ b/src/commands/project.ts @@ -6052,11 +6052,13 @@ async function handleNativeUp({ allowedHosts: startup.allowedHosts, run, cleanedRetry, + recoverStopped: recovery !== undefined, }); }, down: async () => { const code = await handleDown({ ctx, + deferFinalization: true, args: { options: { path: args.options.path, @@ -6609,9 +6611,11 @@ function writeDownNotice(opts: { async function handleDown({ ctx, args, + deferFinalization, }: { readonly ctx: CliContext; readonly args: DownArgs; + readonly deferFinalization?: boolean; }): Promise { const native = resolveNativeRuntimeSelection(); if (native) { @@ -6684,6 +6688,7 @@ async function handleDown({ const result = await nativeProjectDown({ runtime: native, scope, + deferFinalization, retireHostProcesses: async () => { await stopLifecycleProcesses({ project, diff --git a/tests/models/tla/README.md b/tests/models/tla/README.md index cd08ad461..2d0a118ee 100644 --- a/tests/models/tla/README.md +++ b/tests/models/tla/README.md @@ -16,6 +16,47 @@ pinned artifact. Each check has a 120-second timeout, 512 MiB Java heap, two wor and bounded output; temporary TLC metadata is removed after success or failure. No credentials or running VM are needed. +## Missing post-reboot publications + +`absent-publication-recovery/Absent.tla` checks one explicitly selected cleanup +racing one foreground publisher. The positive configuration explores 247 distinct +states, including one recovery-process crash, two non-volume cleanup effects and +one input-version change. Both participants acquire the foreground lock before the +Engine lease. Ordinary publication occurs under the foreground lock before Engine +admission, matching the implementation. The durable absence intent survives a +crash; a restart cannot publish until cleanup and separate absence retirement are +confirmed. A changed selection +cannot resume cleanup. Completion and retirement are separate steps, so a crash +between them remains visible. + +Four guard-removal controls require TLC exit 12 and a named same-state witness: + +| Control | Required failure | +| --- | --- | +| `negative` | `NoPrematurePublication` in `Publish`, with a durable intent, incomplete cleanup, no retirement and a newly published owner | +| `unwitnessed-cleanup` | `NoUnwitnessedCleanup` in `CleanupOne`, with both locks held but no intent | +| `stale-selection` | `NoUnwitnessedCleanup` in `CleanupOne`, with a durable intent but selected version 1 and current version 2 | +| `unconfirmed-retirement` | `NoUnconfirmedRetirement` in `RetireAbsentPublisher`, before cleanup or intent | + +| Model action | Implementation boundary under `packages/runtime-core/src/provider/` | +| --- | --- | +| `AcquireRecoveryForeground` / `AcquireRecoveryEngine` | `graph/absent_publication_cleanup.rs`: deterministic lock-only reservation before the Engine lease; matches foreground startup's lock order | +| `WriteIntent` | Private durable run-scoped absence intent before any cleanup effects | +| `CleanupOne` / `CommitCleanup` | Exact selected, data-retaining `cleanup_owned(..., false)` and stopped receipt confirmation; tagged absent bridge authority remains distinct from pinned predecessors | +| `RetireAbsentPublisher` | Separate durable absent-publication retirement proof tied to completed cleanup | +| `AcquirePublisherForeground` / `AcquirePublisherEngine` / `Publish` | `graph/foreground/transport.rs` admission barrier plus `graph/foreground.rs` Engine acquisition; explicit retired binding needs completed proof | +| `CrashRecovery` / `ChangeInputs` | Retained intent and exact retry; reinspection refuses changed host/guest boot, input bytes or resources | + +The model assumes validated legacy inputs, absent roots and resource observations +are summarized by one version. It represents cooperating writers under two kernel +locks and durable writes as atomic commits. It does not prove filesystem fsync, +FD/path identity, raw hashes, PID reuse, private permissions, actual guest or +physical host reboot, bridge cleanup, volume preservation, source-device migration, +later restore generations, browser readiness or performance. Concrete refusal, +interruption and data-marker tests remain required. The recorded old Owner is +legacy corroboration, never a reconstructed foreground Pin or proof of original +physical-volume continuity. No fairness or eventual recovery claim is made. + ## Shared HTTPS lifetime `shared-https-lifetime/SharedHttps.tla` checks the last-lease release racing @@ -47,6 +88,61 @@ shared-owner process tests and startup/finalization/recovery regressions provide separate implementation evidence; native multi-application acceptance remains required before claiming complete concurrent application support. +## Previous-boot shared HTTPS archival + +`previous-boot-shared-https/ArchiveHttps.tla` checks one explicit recovery racing +another application's owner startup. It models the owner directory and control +socket as separate atomic renames in either order. Native archival moves the +socket first; the owner-first interleaving is an additional barrier control. +The admission barrier survives one recovery +crash while the provider cleanup lease does not; resume reacquires that lease and +rechecks the exact selection before another move. Startup can observe an absent +owner directory between the two moves, so retaining admission is essential. + +The positive configuration explores 92 distinct states from two initial inputs +(eligible or refused), with one crash/resume and one input-generation change. +Original artifacts remain in exactly one original or archived location. Completed +archival and explicit frontend finalization are separate facts. Another +application may create the next owner after archival commits; the selected +application must still finish its own exact frontend recovery. + +Three guard-removal controls require TLC exit 12 and the named same-state witness: + +| Control | Required failure | +| --- | --- | +| `negative` | `NoPrematurePublication` in `Publish`, after an owner-only move and recovery crash, before socket archival or completion | +| `stale-selection` | `NoUnprovedArchive` in `ArchiveOwner`, with selected generation 1 and current generation 2 after reacquiring the provider lease | +| `unproved-archive` | `NoUnprovedArchive` in `ArchiveOwner`, with an ineligible selection despite held locks and a recorded intent | + +| Model action | Implementation boundary | +| --- | --- | +| `AcquireAdmission` / `AcquireEngine` | Explicit previous-boot shared-owner recovery takes shared HTTPS admission before native provider/graph cleanup exclusion | +| `Prepare` / `eligible` / `selected` | Exact current completed dead-owner cleanup, retired publisher, immediate boot succession and unchanged owner/lease/socket/executable/data observations | +| `ArchiveOwner` / `ArchiveSocket` | Journaled no-signal, same-filesystem moves preserve selected owner/lease and socket inodes; each effect rechecks the selected state | +| `Crash` / `Resume` / `ChangeInputs` | Durable admission barrier and intent survive interruption; exact dead recovery ownership and unchanged inputs are required to resume | +| `Commit` / `Finalize` | Completed archive proof remains distinct from ordinary lease-release acknowledgement; explicit v3 frontend recovery consumes only its exact selected proof | +| `AcquireStartup` / `Publish` | Normal shared-owner admission prevents a new owner during incomplete archival | + +Validated graph, process, filesystem and boot observations are summarized by one +eligibility flag and input version. Admission identity and recovery-process death +are assumed validated before resume. Cooperating provider writers cannot change +the selection under the provider lease; legacy startup must separately remain +quiescent. Atomic renames and durable intent/completion writes are abstractions, +not fsync proofs. The native implementation is in +`packages/runtime-core/src/provider/shared_https_recovery.rs`; explicit v3 +frontend selection is in `src/backends/native-project-recovery.ts` and +`src/backends/native-https-owner.ts`. This model covers the present-socket path +with two moves. The both-absent socket/parent variant has one owner move and +requires separate concrete absence, process and port regression controls. +The implementation also refuses completed replay while a newer shared owner is +active; the model's post-commit startup interleaving is a safety bound, not proof +that this interrupted frontend can finish without first restoring quiescence. +This model does not establish executable absence, PID reuse, +socket/FD identity, permissions, hash integrity, physical reboot, retained volume +continuity, multiple-lease recovery, browser readiness or performance. Concrete +refusal, interrupted-archive, process/socket and native application checks remain +required. No fairness or eventual recovery claim is made. + ## Active dependency rebinding `dependency-rebind/Rebind.tla` checks one physical slot shared by two logical diff --git a/tests/models/tla/absent-publication-recovery/Absent.tla b/tests/models/tla/absent-publication-recovery/Absent.tla new file mode 100644 index 000000000..a736a3182 --- /dev/null +++ b/tests/models/tla/absent-publication-recovery/Absent.tla @@ -0,0 +1,136 @@ +----------------------------- MODULE Absent ----------------------------- +EXTENDS Naturals, TLC +CONSTANTS EnforceIntent, EnforceBarrier, EnforceSelection, EnforceCompletion +VARIABLES foreground, engine, recovery, publisher, intent, progress, + complete, retired, published, version, selected, crashed, + unsafeCleanup, unsafePublication, unsafeRetirement +durable == <> +observations == <> +effects == <> +vars == <> +LocksHeld == foreground = "recovery" /\ engine = "recovery" +SelectionMatches == selected = version +\* Exact legacy inputs and resource/guest observations are one abstract version. +\* No old Pin is synthesized. The selection starts without any live publication. +Init == /\ foreground = "none" /\ engine = "none" + /\ recovery = "idle" /\ publisher = "idle" + /\ intent = FALSE /\ progress = 0 /\ complete = FALSE + /\ retired = FALSE /\ published = FALSE + /\ version = 1 /\ selected = 0 /\ crashed = FALSE + /\ unsafeCleanup = FALSE /\ unsafePublication = FALSE + /\ unsafeRetirement = FALSE +AcquireRecoveryForeground == + /\ recovery \in {"idle", "crashed"} /\ foreground = "none" + /\ foreground' = "recovery" /\ recovery' = "foreground" + /\ selected' = IF intent THEN selected ELSE version + /\ UNCHANGED <> +AcquireRecoveryEngine == + /\ recovery = "foreground" /\ foreground = "recovery" /\ engine = "none" + /\ engine' = "recovery" + /\ recovery' = IF retired THEN "retired" + ELSE IF complete THEN "complete" + ELSE IF intent THEN "cleaning" ELSE "leased" + /\ UNCHANGED <> +WriteIntent == + /\ recovery = "leased" /\ LocksHeld /\ SelectionMatches /\ ~published + /\ intent' = TRUE /\ recovery' = "cleaning" + /\ UNCHANGED <> +\* One effect represents one exact non-volume resource cleanup. Commit is separate. +CleanupOne == + /\ recovery \in {"leased", "cleaning"} /\ LocksHeld /\ ~published + /\ progress < 2 /\ (EnforceIntent => intent) + /\ (EnforceSelection => SelectionMatches) + /\ progress' = progress + 1 /\ recovery' = "cleaning" + /\ unsafeCleanup' = (unsafeCleanup \/ ~intent \/ ~SelectionMatches) + /\ UNCHANGED <> +CommitCleanup == + /\ recovery = "cleaning" /\ LocksHeld /\ intent + /\ SelectionMatches /\ progress = 2 /\ ~published + /\ complete' = TRUE /\ recovery' = "complete" + /\ UNCHANGED <> +RetireAbsentPublisher == + /\ recovery \in {"leased", "cleaning", "complete"} /\ LocksHeld + /\ SelectionMatches /\ ~published + /\ (EnforceCompletion => (intent /\ complete /\ progress = 2)) + /\ retired' = TRUE /\ recovery' = "retired" + /\ unsafeRetirement' = (unsafeRetirement \/ ~intent \/ ~complete \/ progress # 2) + /\ UNCHANGED <> +ReleaseRecovery == + /\ recovery = "retired" /\ LocksHeld + /\ foreground' = "none" /\ engine' = "none" /\ recovery' = "done" + /\ UNCHANGED <> +\* A crash releases kernel locks while retaining every committed durable field. +CrashRecovery == + /\ recovery \in {"foreground", "leased", "cleaning", "complete", "retired"} + /\ ~crashed /\ recovery' = "crashed" /\ crashed' = TRUE + /\ foreground' = "none" /\ engine' = "none" + /\ UNCHANGED <> +\* A cooperating identity/guest-boot change cannot occur under the Engine lease. +ChangeInputs == + /\ engine = "none" /\ version = 1 /\ version' = 2 + /\ UNCHANGED <> +RefuseRecovery == + /\ recovery \in {"leased", "cleaning", "complete"} /\ LocksHeld + /\ (~SelectionMatches \/ published) + /\ recovery' = "refused" /\ foreground' = "none" /\ engine' = "none" + /\ UNCHANGED <> +AcquirePublisherForeground == + /\ publisher = "idle" /\ foreground = "none" + /\ publisher' = "foreground" /\ foreground' = "publisher" + /\ UNCHANGED <> +AcquirePublisherEngine == + /\ publisher = "published" /\ foreground = "publisher" /\ engine = "none" + /\ publisher' = "active" /\ engine' = "publisher" + /\ UNCHANGED <> +ReleasePublisherEngine == + /\ publisher = "active" /\ engine = "publisher" + /\ publisher' = "running" /\ engine' = "none" + /\ UNCHANGED <> +FreshAllowed == ~intent /\ ~published +RetiredAllowed == intent /\ complete /\ retired /\ ~published +\* Ordinary bind publishes while holding the foreground lock, BEFORE Engine admission. +\* Recovery also needs that foreground lock, so intent and publication cannot race. +Publish == + /\ publisher = "foreground" /\ foreground = "publisher" /\ engine = "none" + /\ (FreshAllowed \/ RetiredAllowed \/ (~EnforceBarrier /\ ~published)) + /\ published' = TRUE /\ publisher' = "published" + /\ unsafePublication' = (unsafePublication \/ (intent /\ ~retired)) + /\ UNCHANGED <> +RefusePublisher == + /\ publisher = "foreground" /\ ~FreshAllowed /\ ~RetiredAllowed + /\ publisher' = "refused" /\ foreground' = "none" + /\ UNCHANGED <> +PublisherExit == + /\ publisher \in {"published", "active", "running"} /\ foreground = "publisher" + /\ publisher' = "exited" /\ foreground' = "none" /\ engine' = "none" + /\ UNCHANGED <> +Next == AcquireRecoveryForeground \/ AcquireRecoveryEngine \/ WriteIntent \/ CleanupOne + \/ CommitCleanup \/ RetireAbsentPublisher \/ ReleaseRecovery \/ CrashRecovery + \/ ChangeInputs \/ RefuseRecovery \/ AcquirePublisherForeground + \/ AcquirePublisherEngine \/ ReleasePublisherEngine \/ Publish \/ RefusePublisher + \/ PublisherExit +Spec == Init /\ [][Next]_vars +TypeOK == /\ foreground \in {"none", "recovery", "publisher"} + /\ engine \in {"none", "recovery", "publisher"} + /\ recovery \in {"idle", "foreground", "leased", "cleaning", "complete", + "retired", "crashed", "done", "refused"} + /\ publisher \in {"idle", "foreground", "published", "active", "running", + "refused", "exited"} + /\ intent \in BOOLEAN /\ progress \in 0..2 /\ complete \in BOOLEAN + /\ retired \in BOOLEAN /\ published \in BOOLEAN + /\ version \in 1..2 /\ selected \in 0..2 /\ crashed \in BOOLEAN + /\ unsafeCleanup \in BOOLEAN /\ unsafePublication \in BOOLEAN + /\ unsafeRetirement \in BOOLEAN +ForegroundBeforeEngine == engine # "none" => engine = foreground +NoUnwitnessedCleanup == ~unsafeCleanup +NoPrematurePublication == ~unsafePublication +NoUnconfirmedRetirement == ~unsafeRetirement +CompletionRequiresCleanup == complete => (intent /\ progress = 2) +======================================================================= diff --git a/tests/models/tla/absent-publication-recovery/negative.cfg b/tests/models/tla/absent-publication-recovery/negative.cfg new file mode 100644 index 000000000..16ebac56b --- /dev/null +++ b/tests/models/tla/absent-publication-recovery/negative.cfg @@ -0,0 +1,12 @@ +CONSTANTS EnforceIntent = TRUE + EnforceBarrier = FALSE + EnforceSelection = TRUE + EnforceCompletion = TRUE +SPECIFICATION Spec +INVARIANTS TypeOK + ForegroundBeforeEngine + NoUnwitnessedCleanup + NoPrematurePublication + NoUnconfirmedRetirement + CompletionRequiresCleanup +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/absent-publication-recovery/positive.cfg b/tests/models/tla/absent-publication-recovery/positive.cfg new file mode 100644 index 000000000..db5e420e8 --- /dev/null +++ b/tests/models/tla/absent-publication-recovery/positive.cfg @@ -0,0 +1,12 @@ +CONSTANTS EnforceIntent = TRUE + EnforceBarrier = TRUE + EnforceSelection = TRUE + EnforceCompletion = TRUE +SPECIFICATION Spec +INVARIANTS TypeOK + ForegroundBeforeEngine + NoUnwitnessedCleanup + NoPrematurePublication + NoUnconfirmedRetirement + CompletionRequiresCleanup +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/absent-publication-recovery/stale-selection.cfg b/tests/models/tla/absent-publication-recovery/stale-selection.cfg new file mode 100644 index 000000000..47cec1236 --- /dev/null +++ b/tests/models/tla/absent-publication-recovery/stale-selection.cfg @@ -0,0 +1,12 @@ +CONSTANTS EnforceIntent = TRUE + EnforceBarrier = TRUE + EnforceSelection = FALSE + EnforceCompletion = TRUE +SPECIFICATION Spec +INVARIANTS TypeOK + ForegroundBeforeEngine + NoUnwitnessedCleanup + NoPrematurePublication + NoUnconfirmedRetirement + CompletionRequiresCleanup +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/absent-publication-recovery/unconfirmed-retirement.cfg b/tests/models/tla/absent-publication-recovery/unconfirmed-retirement.cfg new file mode 100644 index 000000000..7e0438c70 --- /dev/null +++ b/tests/models/tla/absent-publication-recovery/unconfirmed-retirement.cfg @@ -0,0 +1,12 @@ +CONSTANTS EnforceIntent = TRUE + EnforceBarrier = TRUE + EnforceSelection = TRUE + EnforceCompletion = FALSE +SPECIFICATION Spec +INVARIANTS TypeOK + ForegroundBeforeEngine + NoUnwitnessedCleanup + NoPrematurePublication + NoUnconfirmedRetirement + CompletionRequiresCleanup +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/absent-publication-recovery/unwitnessed-cleanup.cfg b/tests/models/tla/absent-publication-recovery/unwitnessed-cleanup.cfg new file mode 100644 index 000000000..07f3f71d4 --- /dev/null +++ b/tests/models/tla/absent-publication-recovery/unwitnessed-cleanup.cfg @@ -0,0 +1,12 @@ +CONSTANTS EnforceIntent = FALSE + EnforceBarrier = TRUE + EnforceSelection = TRUE + EnforceCompletion = TRUE +SPECIFICATION Spec +INVARIANTS TypeOK + ForegroundBeforeEngine + NoUnwitnessedCleanup + NoPrematurePublication + NoUnconfirmedRetirement + CompletionRequiresCleanup +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/previous-boot-shared-https/ArchiveHttps.tla b/tests/models/tla/previous-boot-shared-https/ArchiveHttps.tla new file mode 100644 index 000000000..b87ad0350 --- /dev/null +++ b/tests/models/tla/previous-boot-shared-https/ArchiveHttps.tla @@ -0,0 +1,103 @@ +---------------------------- MODULE ArchiveHttps ---------------------------- +EXTENDS TLC, FiniteSets +CONSTANTS KeepCrashBarrier, CheckSelection, CheckProof +VARIABLES admission, engine, recovery, crashed, intent, selected, version, + eligible, originals, archived, complete, finalized, published, + unsafeArchive, unsafePublication +vars == <> +Artifacts == {"owner", "socket"} +Init == /\ admission = "free" /\ engine = FALSE /\ recovery = FALSE + /\ crashed = FALSE /\ intent = FALSE /\ selected = 0 /\ version = 1 + /\ eligible \in BOOLEAN /\ originals = Artifacts /\ archived = {} + /\ complete = FALSE /\ finalized = FALSE /\ published = FALSE + /\ unsafeArchive = FALSE /\ unsafePublication = FALSE +AcquireAdmission == /\ admission = "free" /\ ~intent /\ ~published + /\ admission' = "recovery" /\ recovery' = TRUE + /\ UNCHANGED <> +AcquireEngine == /\ admission = "recovery" /\ recovery /\ ~engine + /\ engine' = TRUE + /\ UNCHANGED <> +Prepare == /\ admission = "recovery" /\ recovery /\ engine /\ ~intent + /\ (eligible \/ ~CheckProof) + /\ intent' = TRUE /\ selected' = version + /\ UNCHANGED <> +ArchiveOwner == /\ admission = "recovery" /\ recovery /\ engine /\ intent + /\ "owner" \in originals + /\ (selected = version \/ ~CheckSelection) + /\ originals' = originals \ {"owner"} + /\ archived' = archived \cup {"owner"} + /\ unsafeArchive' = (~eligible \/ selected # version) + /\ UNCHANGED <> +ArchiveSocket == /\ admission = "recovery" /\ recovery /\ engine /\ intent + /\ "socket" \in originals + /\ (selected = version \/ ~CheckSelection) + /\ originals' = originals \ {"socket"} + /\ archived' = archived \cup {"socket"} + /\ unsafeArchive' = (~eligible \/ selected # version) + /\ UNCHANGED <> +Commit == /\ admission = "recovery" /\ recovery /\ engine /\ intent + /\ archived = Artifacts /\ eligible /\ selected = version + /\ complete' = TRUE /\ engine' = FALSE /\ recovery' = FALSE + /\ admission' = "free" + /\ UNCHANGED <> +Crash == /\ admission = "recovery" /\ recovery /\ ~crashed + /\ crashed' = TRUE /\ engine' = FALSE /\ recovery' = FALSE + /\ admission' = IF KeepCrashBarrier THEN "recovery" ELSE "free" + /\ UNCHANGED <> +Resume == /\ admission = "recovery" /\ ~recovery /\ crashed /\ ~complete + /\ recovery' = TRUE + /\ UNCHANGED <> +ChangeInputs == /\ version = 1 /\ ~engine /\ ~complete + /\ version' = 2 + /\ UNCHANGED <> +Finalize == /\ complete /\ selected = version /\ ~finalized + /\ finalized' = TRUE + /\ UNCHANGED <> +AcquireStartup == /\ admission = "free" /\ "owner" \notin originals + /\ ~published + /\ admission' = "startup" + /\ UNCHANGED <> +Publish == /\ admission = "startup" /\ ~published + /\ published' = TRUE /\ admission' = "free" + /\ unsafePublication' = (~complete \/ archived # Artifacts) + /\ UNCHANGED <> +Next == AcquireAdmission \/ AcquireEngine \/ Prepare \/ ArchiveOwner + \/ ArchiveSocket \/ Commit \/ Crash \/ Resume \/ ChangeInputs + \/ Finalize \/ AcquireStartup \/ Publish +Spec == Init /\ [][Next]_vars +TypeOK == /\ admission \in {"free", "recovery", "startup"} + /\ {engine, recovery, crashed, intent, eligible, complete, finalized, + published, unsafeArchive, unsafePublication} \subseteq BOOLEAN + /\ selected \in {0, 1, 2} /\ version \in {1, 2} + /\ originals \subseteq Artifacts /\ archived \subseteq Artifacts +PreserveArtifacts == /\ originals \cup archived = Artifacts + /\ originals \cap archived = {} +NoUnprovedArchive == ~unsafeArchive +NoPrematurePublication == ~unsafePublication +============================================================================= diff --git a/tests/models/tla/previous-boot-shared-https/negative.cfg b/tests/models/tla/previous-boot-shared-https/negative.cfg new file mode 100644 index 000000000..8b76810ff --- /dev/null +++ b/tests/models/tla/previous-boot-shared-https/negative.cfg @@ -0,0 +1,6 @@ +SPECIFICATION Spec +CONSTANTS KeepCrashBarrier = FALSE + CheckSelection = TRUE + CheckProof = TRUE +INVARIANTS TypeOK PreserveArtifacts NoUnprovedArchive NoPrematurePublication +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/previous-boot-shared-https/positive.cfg b/tests/models/tla/previous-boot-shared-https/positive.cfg new file mode 100644 index 000000000..2e523eef3 --- /dev/null +++ b/tests/models/tla/previous-boot-shared-https/positive.cfg @@ -0,0 +1,6 @@ +SPECIFICATION Spec +CONSTANTS KeepCrashBarrier = TRUE + CheckSelection = TRUE + CheckProof = TRUE +INVARIANTS TypeOK PreserveArtifacts NoUnprovedArchive NoPrematurePublication +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/previous-boot-shared-https/stale-selection.cfg b/tests/models/tla/previous-boot-shared-https/stale-selection.cfg new file mode 100644 index 000000000..44ea8566e --- /dev/null +++ b/tests/models/tla/previous-boot-shared-https/stale-selection.cfg @@ -0,0 +1,6 @@ +SPECIFICATION Spec +CONSTANTS KeepCrashBarrier = TRUE + CheckSelection = FALSE + CheckProof = TRUE +INVARIANTS TypeOK PreserveArtifacts NoUnprovedArchive NoPrematurePublication +CHECK_DEADLOCK FALSE diff --git a/tests/models/tla/previous-boot-shared-https/unproved-archive.cfg b/tests/models/tla/previous-boot-shared-https/unproved-archive.cfg new file mode 100644 index 000000000..80b2763c9 --- /dev/null +++ b/tests/models/tla/previous-boot-shared-https/unproved-archive.cfg @@ -0,0 +1,6 @@ +SPECIFICATION Spec +CONSTANTS KeepCrashBarrier = TRUE + CheckSelection = TRUE + CheckProof = FALSE +INVARIANTS TypeOK PreserveArtifacts NoUnprovedArchive NoPrematurePublication +CHECK_DEADLOCK FALSE diff --git a/tests/native-https-empty-owner-recovery.test.ts b/tests/native-https-empty-owner-recovery.test.ts new file mode 100644 index 000000000..71e778af8 --- /dev/null +++ b/tests/native-https-empty-owner-recovery.test.ts @@ -0,0 +1,380 @@ +import { afterEach, expect, test } from "bun:test"; +import { createHash } from "node:crypto"; +import { + chmod, + lstat, + mkdir, + mkdtemp, + readdir, + readFile, + realpath, + rm, + writeFile, +} from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { + archiveEmptyNativeHttpsOwner, + type EmptyNativeHttpsOwnerInspection, + type EmptyNativeHttpsOwnerSelection, +} from "../src/backends/native-https-empty-owner-recovery.ts"; +import { ensureNativeHttpsOwner } from "../src/backends/native-https-owner.ts"; +import { + acquireNativeHttpsOwnerAdmission, + NATIVE_HTTPS_OWNER_ADMISSION, +} from "../src/backends/native-https-owner-admission.ts"; +import type { NativeHttpsOwnerConfiguration } from "../src/backends/native-https-owner-protocol.ts"; +import { nativeHttpsWriteNew } from "../src/backends/native-https-owner-storage.ts"; + +const homes: string[] = []; +afterEach(async () => { + for (const home of homes.splice(0)) { + await rm(home, { recursive: true, force: true }); + } +}); +const sha = (bytes: string | Buffer) => + createHash("sha256").update(bytes).digest("hex"); +const absent = async (path: string) => + await lstat(path) + .then(() => false) + .catch((error: unknown) => { + if ( + typeof error === "object" && + error !== null && + "code" in error && + error.code === "ENOENT" + ) { + return true; + } + throw error; + }); + +async function fixture() { + const home = await mkdtemp(join(await realpath(tmpdir()), "hk-empty-owner-")); + homes.push(home); + await chmod(home, 0o700); + const storage = join(home, "native-https"); + const root = join(storage, "shared-owner"); + await mkdir(join(root, "leases"), { recursive: true, mode: 0o700 }); + await chmod(storage, 0o700); + await chmod(root, 0o700); + await chmod(join(root, "leases"), 0o700); + const sibling = join(storage, "data", "caddy", "marker"); + await mkdir(join(storage, "data", "caddy"), { recursive: true, mode: 0o700 }); + await writeFile(sibling, "retained sibling data", { mode: 0o600 }); + const binary = join(home, "reviewed-binary"); + await writeFile(binary, "#!/bin/sh\nexit 0\n", { mode: 0o700 }); + const binarySha = sha(await readFile(binary)); + const config: NativeHttpsOwnerConfiguration = { + version: 1, + ownerGeneration: "a".repeat(32), + binding: { + runtime: { home, binary }, + runtimeSha256: binarySha, + frontend: { binary, sha256: binarySha }, + pool: { + owner: "b".repeat(32), + bootId: "12345678-1234-1234-1234-123456789abc", + }, + caddyBinary: binary, + caddySha256: binarySha, + httpsPort: 18_443, + certificateNameLimit: 16, + }, + }; + const configPath = join(root, "configuration.json"); + const bytes = `${JSON.stringify(config)}\n`; + await writeFile(configPath, bytes, { mode: 0o600 }); + const selection: EmptyNativeHttpsOwnerSelection = { + ownerGeneration: config.ownerGeneration, + configurationSha256: sha(bytes), + runtime: config.binding.runtime, + runtimeSha256: config.binding.runtimeSha256, + originalSpawnerPid: 47_321, + }; + const inspection: EmptyNativeHttpsOwnerInspection = { + originalSpawnerAbsent: async () => true, + selectedOwnerProcessAbsent: async () => true, + selectedPortAbsent: async () => true, + poolAndPublications: async () => true, + }; + return { + home, + storage, + root, + configPath, + config, + bytes, + selection, + inspection, + sibling, + }; +} +async function archive(f: Awaited>) { + return archiveEmptyNativeHttpsOwner({ + selection: f.selection, + acceptLegacyOwnerWithoutPid: true, + inspection: f.inspection, + }); +} + +test("archives exact empty owner, retaining its inode, bytes, and sibling data", async () => { + const f = await fixture(); + const before = await lstat(f.root); + const result = await archive(f); + const after = await lstat(result.archive); + expect(after.dev).toBe(before.dev); + expect(after.ino).toBe(before.ino); + expect( + await readFile(join(result.archive, "configuration.json"), "utf8") + ).toBe(f.bytes); + expect(await readdir(join(result.archive, "leases"))).toEqual([]); + expect(await readFile(f.sibling, "utf8")).toBe("retained sibling data"); + expect(await absent(f.root)).toBe(true); + expect(await absent(join(f.storage, "owner.lock"))).toBe(true); + expect(await absent(join(f.storage, NATIVE_HTTPS_OWNER_ADMISSION))).toBe( + true + ); + const entries = await readdir(f.storage); + expect( + entries.filter((value) => value.includes("empty-owner-recovery-")) + ).toHaveLength(2); + expect( + await archive(f) + .then(() => true) + .catch(() => false) + ).toBe(false); +}); + +test.each([ + "originalSpawnerAbsent", + "selectedOwnerProcessAbsent", + "selectedPortAbsent", + "poolAndPublications", +] as const)("refuses uncertain or live %s without moving selected state", async (method) => { + const f = await fixture(); + f.inspection[method] = async () => false; + await expect(archive(f)).rejects.toThrow(); + expect(await readFile(f.configPath, "utf8")).toBe(f.bytes); + expect(await absent(join(f.storage, "owner.lock"))).toBe(true); + expect(await absent(join(f.storage, NATIVE_HTTPS_OWNER_ADMISSION))).toBe( + true + ); +}); + +test.each([ + ["endpoint.json", "root"], + ["configuration.pending", "root"], + ["unknown", "root"], + ["lease.json", "leases"], + ["owner.sock", "storage"], + ["active-owner.json", "storage"], + ["owner.lock", "storage"], + ["empty-owner-recovery-foreign.pending", "storage"], + ["archived-empty-owner-foreign", "storage"], +] as const)("refuses %s at %s while retaining exact selected state", async (name, where) => { + const f = await fixture(); + const parent = + where === "root" + ? f.root + : where === "leases" + ? join(f.root, "leases") + : f.storage; + if (name === "owner.lock") { + await mkdir(join(parent, name), { mode: 0o700 }); + } else { + await writeFile(join(parent, name), "foreign", { mode: 0o600 }); + } + await expect(archive(f)).rejects.toThrow(); + expect(await readFile(f.configPath, "utf8")).toBe(f.bytes); + expect( + await readFile(join(parent, name), "utf8").catch(() => "directory") + ).toBe(name === "owner.lock" ? "directory" : "foreign"); +}); + +test("requires exact acknowledgement, selected digest, generation, and runtime", async () => { + const f = await fixture(); + await expect( + archiveEmptyNativeHttpsOwner({ + selection: f.selection, + acceptLegacyOwnerWithoutPid: false as true, + inspection: f.inspection, + }) + ).rejects.toThrow(); + for (const selection of [ + { ...f.selection, configurationSha256: "0".repeat(64) }, + { ...f.selection, ownerGeneration: "0".repeat(32) }, + { ...f.selection, runtimeSha256: "0".repeat(64) }, + { ...f.selection, runtime: { ...f.selection.runtime, home: f.storage } }, + { ...f.selection, originalSpawnerPid: 0 }, + ]) { + await expect( + archiveEmptyNativeHttpsOwner({ + selection, + acceptLegacyOwnerWithoutPid: true, + inspection: f.inspection, + }) + ).rejects.toThrow(); + } + expect(await readFile(f.configPath, "utf8")).toBe(f.bytes); +}); + +test("identity drift immediately before rename refuses and preserves both names", async () => { + const f = await fixture(); + let inspections = 0; + f.inspection.selectedPortAbsent = async () => { + inspections++; + if (inspections === 2) { + await writeFile(f.configPath, "replacement", { mode: 0o600 }); + } + return true; + }; + await expect(archive(f)).rejects.toThrow(); + expect(await readFile(f.configPath, "utf8")).toBe("replacement"); + expect(await absent(f.root)).toBe(false); + expect(await absent(join(f.storage, "owner.lock"))).toBe(true); +}); + +test("failure after durable intent preserves pending proof and exclusion locks", async () => { + const f = await fixture(); + let inspections = 0; + f.inspection.selectedPortAbsent = async () => { + inspections++; + return inspections < 3; + }; + await expect(archive(f)).rejects.toThrow(); + expect(await readFile(f.configPath, "utf8")).toBe(f.bytes); + expect(await absent(join(f.storage, "owner.lock"))).toBe(false); + expect(await absent(join(f.storage, NATIVE_HTTPS_OWNER_ADMISSION))).toBe( + false + ); + expect( + (await readdir(f.storage)).some((name) => name.endsWith(".intent.json")) + ).toBe(true); + await expect(archive(f)).rejects.toThrow(); +}); + +test("normal concurrent ensures publish and share one generation", async () => { + const f = await fixture(); + await rm(f.root, { recursive: true }); + let spawns = 0; + const options = { + binding: f.config.binding, + spawnOwner: async () => { + spawns++; + await Bun.sleep(30); + }, + }; + const [first, second] = await Promise.all([ + ensureNativeHttpsOwner(options), + ensureNativeHttpsOwner(options), + ]); + expect(first).toEqual(second); + expect(spawns).toBe(1); + expect(await absent(join(f.storage, NATIVE_HTTPS_OWNER_ADMISSION))).toBe( + true + ); +}); + +test("archival admission blocks a concurrent ensure until complete proof and lock retirement", async () => { + const f = await fixture(); + let finishComplete: (() => void) | undefined; + let enteredComplete: (() => void) | undefined; + const atComplete = new Promise((resolve) => { + enteredComplete = resolve; + }); + const holdComplete = new Promise((resolve) => { + finishComplete = resolve; + }); + const recovering = archiveEmptyNativeHttpsOwner({ + selection: f.selection, + acceptLegacyOwnerWithoutPid: true, + inspection: f.inspection, + writeJournal: async (path, value) => { + if (path.endsWith(".complete.json")) { + enteredComplete?.(); + await holdComplete; + } + return nativeHttpsWriteNew(path, value); + }, + }); + await atComplete; + let spawns = 0; + const ensuring = ensureNativeHttpsOwner({ + binding: f.config.binding, + spawnOwner: async () => { + spawns++; + }, + }); + await Bun.sleep(50); + expect(spawns).toBe(0); + expect(await absent(f.root)).toBe(true); + finishComplete?.(); + const archived = await recovering; + const replacement = await ensuring; + expect(spawns).toBe(1); + expect(replacement.ownerGeneration).not.toBe(f.selection.ownerGeneration); + expect( + await readFile(join(archived.archive, "configuration.json"), "utf8") + ).toBe(f.bytes); + expect(await absent(join(f.storage, NATIVE_HTTPS_OWNER_ADMISSION))).toBe( + true + ); +}); + +test("uncertain failure after intent publication retains admission and HTTPS barriers", async () => { + const f = await fixture(); + await expect( + archiveEmptyNativeHttpsOwner({ + selection: f.selection, + acceptLegacyOwnerWithoutPid: true, + inspection: f.inspection, + writeJournal: async (path, value) => { + const published = await nativeHttpsWriteNew(path, value); + if (path.endsWith(".intent.json")) { + throw new Error("injected post-publication fsync failure"); + } + return published; + }, + }) + ).rejects.toThrow("injected post-publication fsync failure"); + expect(await absent(f.root)).toBe(false); + expect(await absent(join(f.storage, "owner.lock"))).toBe(false); + expect(await absent(join(f.storage, NATIVE_HTTPS_OWNER_ADMISSION))).toBe( + false + ); + expect( + (await readdir(f.storage)).some((name) => name.endsWith(".intent.json")) + ).toBe(true); + await expect( + acquireNativeHttpsOwnerAdmission({ home: f.home, waitMs: 0 }) + ).rejects.toThrow(); +}); + +test("a replacement appearing after rename is retained and success is refused", async () => { + const f = await fixture(); + const foreign = "foreign replacement"; + await expect( + archiveEmptyNativeHttpsOwner({ + selection: f.selection, + acceptLegacyOwnerWithoutPid: true, + inspection: f.inspection, + writeJournal: async (path, value) => { + if (path.endsWith(".complete.json")) { + await mkdir(f.root, { mode: 0o700 }); + await writeFile(join(f.root, "foreign"), foreign, { mode: 0o600 }); + } + return nativeHttpsWriteNew(path, value); + }, + }) + ).rejects.toThrow(); + expect(await readFile(join(f.root, "foreign"), "utf8")).toBe(foreign); + expect( + (await readdir(f.storage)).some((name) => + name.startsWith("archived-empty-owner-") + ) + ).toBe(true); + expect(await absent(join(f.storage, NATIVE_HTTPS_OWNER_ADMISSION))).toBe( + false + ); +}); diff --git a/tests/native-https-owner.test.ts b/tests/native-https-owner.test.ts index 339a944fa..abf76c301 100644 --- a/tests/native-https-owner.test.ts +++ b/tests/native-https-owner.test.ts @@ -7,6 +7,7 @@ import { realpath, rename, rm, + symlink, writeFile, } from "node:fs/promises"; import type { Socket } from "node:net"; @@ -27,6 +28,7 @@ import { import { verifyNativeHttpsLeaseGraph } from "../src/backends/native-https-owner-server.ts"; import { nativeHttpsExecutableSha256, + nativeHttpsLeaseReleasePath, nativeHttpsOwnerRoot, nativeHttpsReadRelease, nativeHttpsRecordRelease, @@ -85,7 +87,7 @@ async function fixture( frontend, `#!${process.execPath} import { serveNativeHttpsOwner } from ${JSON.stringify(serverModule)}; -import { appendFile, readFile, unlink, writeFile } from "node:fs/promises"; +import { appendFile, chmod, readFile, unlink, writeFile } from "node:fs/promises"; const home = ${JSON.stringify(home)}; const fault = ${JSON.stringify(options.afterPublish ?? null)}; await writeFile(home + "/helper-pid", String(process.pid)); @@ -105,7 +107,8 @@ try { afterPublish: async (path) => { await writeFile(home + "/published-path", path); if (fault === "fail") { throw new Error("injected failure after publication"); } - if (fault === "replace") { await unlink(path); await writeFile(path, "foreign replacement", { mode: 0o644 }); } + // Set the foreign mode exactly even when the helper inherits a private umask. + if (fault === "replace") { await unlink(path); await writeFile(path, "foreign replacement", { mode: 0o644 }); await chmod(path, 0o644); } }, verifyIdle: async () => { ${ @@ -217,6 +220,218 @@ try { return { home, binding, sockets, start, acquire, clean, release }; } +/** Historical recovery exercises only private files; neither owner is spawned. */ +async function historicalReleaseFixture(sameGeneration = false) { + const f = await fixture(); + const current = await ensureNativeHttpsOwner({ + binding: { + ...f.binding, + runtime: { ...f.binding.runtime, binary: join(f.home, "new-runtime") }, + frontend: { + binary: join(f.home, "new-frontend"), + sha256: "3".repeat(64), + }, + runtimeSha256: "4".repeat(64), + }, + spawnOwner: async () => {}, + }); + const historical = { + ...current, + ownerGeneration: sameGeneration ? current.ownerGeneration : "5".repeat(32), + binding: f.binding, + }; + const identity = nativeHttpsLeaseIdentity(historical, lease("b")); + await nativeHttpsRecordRelease({ + version: 1, + identity, + binding: f.binding, + finalOwner: true, + }); + const configurationPath = join( + nativeHttpsOwnerRoot(f.home), + "configuration.json" + ); + const releasePath = nativeHttpsLeaseReleasePath(f.home, identity); + const siblingPath = join( + nativeHttpsOwnerRoot(f.home), + "leases", + "sibling.json" + ); + await writeFile(siblingPath, "new owner's retained evidence", { + mode: 0o600, + }); + const snapshot = () => + Promise.all( + [configurationPath, releasePath, siblingPath].map((path) => + readFile(path) + ) + ); + return { ...f, current, identity, configurationPath, snapshot }; +} + +test("historical HTTPS release revalidates across a new bundle generation without changing its owner", async () => { + const f = await historicalReleaseFixture(); + const before = await f.snapshot(); + let verified = 0; + await recoverNativeHttpsLease({ + runtime: f.binding.runtime, + identity: f.identity, + verifyReleased: async (binding, identity, phase) => { + expect(binding).toEqual(f.binding); + expect(identity).toEqual(f.identity); + expect(phase).toBe("release"); + verified += 1; + }, + }); + expect(verified).toBe(1); + expect(await f.snapshot()).toEqual(before); + // Public owner access still requires the current exact executable binding. + await expect( + readNativeHttpsOwnerConfiguration(f.binding.runtime) + ).rejects.toThrow(); + expect( + await readNativeHttpsOwnerConfiguration(f.current.binding.runtime) + ).toEqual(f.current); +}); + +test("explicit prior-boot frontend recovery selects the exact native archive contract", async () => { + const f = await fixture(); + const old = await ensureNativeHttpsOwner({ + binding: f.binding, + spawnOwner: async () => {}, + }); + const identity = nativeHttpsLeaseIdentity(old, lease("b")); + const selectedBinary = join(f.home, "selected-native"); + const callsPath = join(f.home, "native-calls"); + const script = (bootId: string) => `#!/bin/sh +printf '%s\\n' "$*" >> '${callsPath}' +case "$4" in + status) printf '%s\\n' '${JSON.stringify({ phase: "running", process_alive: true, guest_boot_id: bootId })}' ;; + archive-previous-boot-shared-https) printf '%s\\n' '${JSON.stringify({ archived: true, run: identity.run, owner_generation: identity.ownerGeneration, lease_id: identity.leaseId, data_retained: true, processes_signaled: 0 })}' ;; + *) exit 1 ;; +esac +`; + await writeFile(selectedBinary, script(f.binding.pool.bootId), { + mode: 0o700, + }); + const runtime = { home: f.home, binary: selectedBinary }; + await expect( + recoverNativeHttpsLease({ runtime, identity }) + ).rejects.toThrow(); + await expect( + recoverNativeHttpsLease({ runtime, identity, archivePreviousBoot: true }) + ).rejects.toThrow(); + expect((await readFile(callsPath, "utf8")).trim().split("\n")).toHaveLength( + 1 + ); + + await writeFile( + selectedBinary, + script("22222222-2222-2222-2222-222222222222") + ); + await recoverNativeHttpsLease({ + runtime, + identity, + archivePreviousBoot: true, + }); + const calls = (await readFile(callsPath, "utf8")).trim().split("\n"); + expect(calls).toHaveLength(3); + expect(calls.at(-1)).toContain(`--run-id ${identity.run}`); + expect(calls.at(-1)).toContain( + `--expect-owner-generation ${identity.ownerGeneration}` + ); + expect(calls.at(-1)).toContain(`--expect-lease-id ${identity.leaseId}`); + expect(calls.at(-1)).toContain(`--expect-attempt ${identity.attempt}`); + expect(calls.at(-1)).toContain(`--expect-owner ${identity.owner}`); + expect(calls.at(-1)).toContain(`--expect-namespace ${identity.namespace}`); + expect(calls.at(-1)).toContain(`--expect-plan ${identity.planId}`); +}); + +test("historical HTTPS release refuses a same-generation runtime mismatch before graph verification", async () => { + const f = await historicalReleaseFixture(true); + const before = await f.snapshot(); + let verified = false; + await expect( + recoverNativeHttpsLease({ + runtime: f.binding.runtime, + identity: f.identity, + verifyReleased: async () => { + verified = true; + }, + }) + ).rejects.toThrow(); + expect(verified).toBe(false); + expect(await f.snapshot()).toEqual(before); +}); + +test.each([ + "malformed", + "symlink", + "foreign-home", +] as const)("historical HTTPS release refuses %s current configuration before graph verification", async (fault) => { + const f = await historicalReleaseFixture(); + if (fault === "symlink") { + const target = join(f.home, "configuration-target.json"); + await rename(f.configurationPath, target); + await symlink(target, f.configurationPath); + } else { + await writeFile( + f.configurationPath, + JSON.stringify( + fault === "malformed" + ? { + ownerGeneration: f.current.ownerGeneration, + } + : { + ...f.current, + binding: { + ...f.current.binding, + runtime: { + ...f.current.binding.runtime, + home: join(f.home, "foreign"), + }, + }, + } + ) + ); + } + const before = await f.snapshot(); + let verified = false; + await expect( + recoverNativeHttpsLease({ + runtime: f.binding.runtime, + identity: f.identity, + verifyReleased: async () => { + verified = true; + }, + }) + ).rejects.toThrow(); + expect(verified).toBe(false); + expect(await f.snapshot()).toEqual(before); + expect((await lstat(f.configurationPath)).isSymbolicLink()).toBe( + fault === "symlink" + ); +}); + +test("historical HTTPS release still refuses a dirty graph after a bundle generation change", async () => { + const f = await historicalReleaseFixture(); + const before = await f.snapshot(); + let verified = false; + await expect( + recoverNativeHttpsLease({ + runtime: f.binding.runtime, + identity: f.identity, + verifyReleased: async (binding) => { + expect(binding).toEqual(f.binding); + verified = true; + throw new Error("graph restarted"); + }, + }) + ).rejects.toThrow("graph restarted"); + expect(verified).toBe(true); + expect(await f.snapshot()).toEqual(before); +}); + test("detached helper outlives its spawning CLI and first release preserves the second lease", async () => { const f = await fixture(); const first = await f.start(true); diff --git a/tests/native-project-down.test.ts b/tests/native-project-down.test.ts index a68a3b70f..99b3e3594 100644 --- a/tests/native-project-down.test.ts +++ b/tests/native-project-down.test.ts @@ -1,4 +1,5 @@ import { afterEach, expect, test } from "bun:test"; +import { createHash } from "node:crypto"; import { mkdir, mkdtemp, realpath, rm, writeFile } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; @@ -6,8 +7,13 @@ import { nativeDownEnvironment, nativeProjectDown, } from "../src/backends/native-project-down.ts"; +import { + beginNativeProjectFinalization, + captureNativeProjectFinalization, +} from "../src/backends/native-project-finalization.ts"; import { loadNativeProjectRun, + type NativeProjectRunScope, removeNativeProjectRun, saveNativeProjectRun, } from "../src/backends/native-project-run.ts"; @@ -25,12 +31,12 @@ const run = { planId: "d".repeat(64), }; const id = "e".repeat(64); -function snapshot() { +function snapshot(stopped = false) { return { receipt: { ...run, plan_id: run.planId, - phase: "ready-observed", + phase: stopped ? "stopped-data-retained" : "ready-observed", resources: { "container:web": { kind: "container", @@ -41,7 +47,12 @@ function snapshot() { }, }, journal_incomplete: false, - observations: { "container:web": { state: "running", health: "healthy" } }, + observations: { + "container:web": { + state: stopped ? "absent" : "running", + health: "healthy", + }, + }, }; } async function fixture(saved = true) { @@ -52,14 +63,265 @@ async function fixture(saved = true) { const projectDir = join(projectRoot, ".hack"), nativeHome = join(projectRoot, "candidate"); await mkdir(projectDir); - await mkdir(nativeHome); + await mkdir(nativeHome, { mode: 0o700 }); const scope = { projectRoot, projectDir, nativeHome, branch: null }; if (saved) { await saveNativeProjectRun({ ...scope, run }); } + const finalization = await beginNativeProjectFinalization({ scope, run }); + await finalization.complete(); return { scope, runtime: { binary: "/not-invoked", home: nativeHome } }; } +function finalizationRoot(scope: NativeProjectRunScope): string { + const hash = createHash("sha256") + .update( + JSON.stringify({ + home: scope.nativeHome, + projectRoot: scope.projectRoot, + projectDir: scope.projectDir, + branch: scope.branch, + }) + ) + .digest("hex"); + return join( + scope.nativeHome, + ".hack-local", + "frontend-finalization", + hash, + run.run + ); +} + +test("down waits for the captured frontend after cleanup and lifecycle retirement", async () => { + const opts = await fixture(); + const lifetime = await beginNativeProjectFinalization({ + scope: opts.scope, + run, + }); + const hooksFinished = Promise.withResolvers(); + const calls: string[] = []; + let cleaned = false; + let stopped = false; + const result = nativeProjectDown({ + ...opts, + retireHostProcesses: async () => { + calls.push("retire-host"); + }, + after: async () => { + calls.push("after"); + hooksFinished.resolve(); + }, + invoke: async (request) => { + calls.push(request.args[1]!); + if (request.args[1] === "cleanup") { + cleaned = true; + return {}; + } + return snapshot(cleaned); + }, + }).then((value) => { + stopped = true; + return value; + }); + await hooksFinished.promise; + await Bun.sleep(20); + expect(stopped).toBe(false); + expect(calls).toEqual([ + "inspect", + "inspect", + "cleanup", + "inspect", + "retire-host", + "after", + ]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); + await lifetime.complete(); + expect((await result).status).toBe("stopped"); + expect(stopped).toBe(true); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); +}); + +test("missing frontend identity refuses before hooks or cleanup without leaking paths", async () => { + const opts = await fixture(); + await rm(finalizationRoot(opts.scope), { recursive: true }); + const calls: string[] = []; + const result = nativeProjectDown({ + ...opts, + before: async () => { + calls.push("before"); + }, + invoke: async (request) => { + calls.push(request.args[1]!); + return snapshot(); + }, + }); + await expect(result).rejects.toThrow("frontend finalization is unconfirmed"); + try { + await result; + } catch (error) { + expect(String(error)).not.toContain(opts.scope.nativeHome); + } + expect(calls).toEqual(["inspect"]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); +}); + +test("frontend identity changed by a before hook refuses before compute cleanup", async () => { + const opts = await fixture(); + const calls: string[] = []; + await expect( + nativeProjectDown({ + ...opts, + before: async () => { + await beginNativeProjectFinalization({ scope: opts.scope, run }); + calls.push("before"); + }, + invoke: async (request) => { + calls.push(request.args[1]!); + return snapshot(); + }, + }) + ).rejects.toThrow("ownership changed"); + expect(calls).toEqual(["inspect", "before", "inspect"]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); +}); + +test("stopped compute cannot borrow a wrong or missing frontend completion", async () => { + for (const fault of [ + "changed", + "missing-active", + "missing-completion", + ] as const) { + const opts = await fixture(); + const root = finalizationRoot(opts.scope); + const calls: string[] = []; + let cleaned = false; + await expect( + nativeProjectDown({ + ...opts, + finalizationTimeoutMs: 1, + invoke: async (request) => { + calls.push(request.args[1]!); + if (request.args[1] === "cleanup") { + cleaned = true; + if (fault === "changed") { + const replacement = await beginNativeProjectFinalization({ + scope: opts.scope, + run, + }); + await replacement.complete(); + } else { + await rm( + join( + root, + fault === "missing-active" ? "active.json" : "completed.json" + ) + ); + } + return {}; + } + return snapshot(cleaned); + }, + }) + ).rejects.toThrow("frontend finalization is unconfirmed"); + expect(calls).toEqual(["inspect", "inspect", "cleanup", "inspect"]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); + } +}); + +test("already-stopped down waits for finalization without guest cleanup or hooks", async () => { + const opts = await fixture(); + const lifetime = await beginNativeProjectFinalization({ + scope: opts.scope, + run, + }); + const retired = Promise.withResolvers(); + const calls: string[] = []; + let stopped = false; + const result = nativeProjectDown({ + ...opts, + before: async () => { + throw new Error("replayed before"); + }, + after: async () => { + throw new Error("replayed after"); + }, + retireHostProcesses: async () => { + calls.push("retire-host"); + retired.resolve(); + }, + invoke: async (request) => { + calls.push(request.args[1]!); + return snapshot(true); + }, + }).then((value) => { + stopped = true; + return value; + }); + await retired.promise; + await Bun.sleep(20); + expect(stopped).toBe(false); + await lifetime.complete(); + expect((await result).status).toBe("stopped"); + expect(calls).toEqual(["inspect", "retire-host"]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); +}); + +test("already-stopped unacknowledged frontend refuses without replaying compute", async () => { + const opts = await fixture(); + const root = finalizationRoot(opts.scope); + const token = await captureNativeProjectFinalization({ + scope: opts.scope, + run, + }); + await writeFile( + join(root, "active.json"), + JSON.stringify({ ...token, pid: 2_147_483_647 }) + ); + await rm(join(root, "completed.json")); + const calls: string[] = []; + await expect( + nativeProjectDown({ + ...opts, + finalizationTimeoutMs: 1, + invoke: async (request) => { + calls.push(request.args[1]!); + return snapshot(true); + }, + }) + ).rejects.toThrow("frontend finalization is unconfirmed"); + expect(calls).toEqual(["inspect"]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); +}); + +test("restart alone may defer its finalization barrier until explicit recovery", async () => { + const opts = await fixture(); + const lifetime = await beginNativeProjectFinalization({ + scope: opts.scope, + run, + }); + expect( + await captureNativeProjectFinalization({ scope: opts.scope, run }) + ).toEqual(lifetime.token); + let cleaned = false; + const calls: string[] = []; + const result = await nativeProjectDown({ + ...opts, + deferFinalization: true, + invoke: async (request) => { + calls.push(request.args[1]!); + if (request.args[1] === "cleanup") { + cleaned = true; + return {}; + } + return snapshot(cleaned); + }, + }); + expect(result.status).toBe("stopped"); + expect(calls).toEqual(["inspect", "inspect", "cleanup", "inspect"]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); +}); + test("down invokes cleanup once and retains the stopped mapping for later up", async () => { const opts = await fixture(); const events: string[] = []; @@ -440,6 +702,67 @@ test("recovered stopped graph retires host processes without replaying hooks or expect(await loadNativeProjectRun(opts.scope)).toEqual(run); }); +test("down recovers a selected interrupted startup without replaying hooks or ordinary cleanup", async () => { + const opts = await fixture(); + const calls: string[] = []; + let recovered = false; + const result = await nativeProjectDown({ + ...opts, + before: () => { + throw new Error("unexpected before-hook replay"); + }, + after: () => { + throw new Error("unexpected after-hook replay"); + }, + retireHostProcesses: async () => { + calls.push("retire-host"); + }, + invoke: async ({ args }) => { + calls.push(args[1]!); + if (args[1] === "inspect-interrupted-start-cleanup") { + return { + run: run.run, + phase: "cleanup-intent", + eligible: true, + same_boot: true, + data_retained: true, + selection_sha256: "f".repeat(64), + }; + } + if (args[1] === "recover-interrupted-start-cleanup") { + recovered = true; + return { + run: run.run, + phase: "stopped-data-retained", + recovered: true, + data_retained: true, + same_boot: true, + publisher_retired: true, + reservation_released: true, + }; + } + const value = snapshot(recovered); + if (!recovered) { + value.receipt.phase = "cleanup-intent"; + return { + ...value, + receipt: { ...value.receipt, relay_cleanup: { phase: "pending" } }, + }; + } + return value; + }, + }); + expect(result.status).toBe("stopped"); + expect(calls).toEqual([ + "inspect", + "inspect-interrupted-start-cleanup", + "recover-interrupted-start-cleanup", + "inspect", + "retire-host", + ]); + expect(await loadNativeProjectRun(opts.scope)).toEqual(run); +}); + test("stopped receipt with a remaining container preserves mapping", async () => { const opts = await fixture(); await expect( diff --git a/tests/native-project-interrupted-cleanup.test.ts b/tests/native-project-interrupted-cleanup.test.ts new file mode 100644 index 000000000..ed6b42383 --- /dev/null +++ b/tests/native-project-interrupted-cleanup.test.ts @@ -0,0 +1,341 @@ +import { expect, test } from "bun:test"; +import { recoverNativeInterruptedStartupCleanup } from "../src/backends/native-project-interrupted-cleanup.ts"; +import type { invokeNativeRuntime } from "../src/backends/native-runtime-client.ts"; + +const run = { + run: "a".repeat(32), + owner: "b".repeat(32), + namespace: "c".repeat(64), + planId: "d".repeat(64), +}; +const hash = "e".repeat(64); +const runtime = { binary: "/not-executed", home: "/not-used" }; + +function snapshot(stopped = false) { + return { + journal_incomplete: false, + interrupted_start_cleanup_incomplete: false, + receipt: { + ...run, + plan_id: run.planId, + phase: stopped ? "stopped-data-retained" : "cleanup-intent", + relay_cleanup: { phase: stopped ? "confirmed" : "pending" }, + resources: { + "container:web": { + kind: "container", + key: "web", + name: "selected-web", + id: "f".repeat(64), + phase: stopped ? "absent" : "uncertain", + }, + "volume:data": { + kind: "volume", + key: "data", + name: "selected-data", + phase: "created", + }, + }, + }, + observations: { + "container:web": { state: stopped ? "absent" : "present" }, + "volume:data": { state: "present" }, + }, + }; +} + +function selection() { + return { + run: run.run, + phase: "cleanup-intent", + eligible: true, + data_retained: true, + same_boot: true, + selection_sha256: hash, + }; +} + +function recovered() { + return { + run: run.run, + phase: "stopped-data-retained", + recovered: true, + data_retained: true, + same_boot: true, + publisher_retired: true, + reservation_released: true, + }; +} + +test("interrupted cleanup uses the exact native selection once and independently verifies retained data", async () => { + const calls: string[][] = []; + const result = await recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: snapshot(), + invoke: async ({ args }) => { + calls.push([...args]); + if (args[1] === "inspect-interrupted-start-cleanup") { + return selection(); + } + if (args[1] === "recover-interrupted-start-cleanup") { + return recovered(); + } + return snapshot(true); + }, + }); + expect(result).toEqual(snapshot(true)); + expect(calls).toEqual([ + [ + "graph", + "inspect-interrupted-start-cleanup", + "--run-id", + run.run, + "--json", + ], + [ + "graph", + "recover-interrupted-start-cleanup", + "--run-id", + run.run, + "--expect-selection", + hash, + "--retain-data", + "--json", + ], + ["graph", "inspect", "--run-id", run.run, "--json"], + ]); +}); + +test("ordinary, completed and unenrolled graph states do not request recovery", async () => { + for (const phase of [ + "ready-observed", + "failed-retained", + "stopped-data-retained", + ]) { + const value = snapshot(); + value.receipt.phase = phase; + expect( + await recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: value, + invoke: () => { + throw new Error("unexpected recovery"); + }, + }) + ).toBeNull(); + } +}); + +test("a post-ACK journal hint revalidates selection and finishes exact retirement", async () => { + const partial = snapshot(true); + partial.interrupted_start_cleanup_incomplete = true; + const selected = { ...selection(), phase: "stopped-data-retained" }; + const calls: string[][] = []; + const result = await recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: partial, + invoke: async ({ args }) => { + calls.push([...args]); + if (args[1] === "inspect-interrupted-start-cleanup") { + return selected; + } + if (args[1] === "recover-interrupted-start-cleanup") { + return recovered(); + } + return snapshot(true); + }, + }); + expect(result).toEqual(snapshot(true)); + expect(calls.length).toBe(3); + expect(calls[1]).toContain(hash); +}); + +test("a native journal hint cannot grant authority to a different graph phase", async () => { + const partial = snapshot(true); + partial.interrupted_start_cleanup_incomplete = true; + partial.receipt.phase = "ready-observed"; + let calls = 0; + await expect( + recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: partial, + invoke: () => { + calls++; + throw new Error("unexpected recovery"); + }, + }) + ).rejects.toThrow("cleanup is unconfirmed"); + expect(calls).toBe(0); +}); + +test("foreign or incomplete graph selection refuses before native requests", async () => { + for (const field of [ + "run", + "owner", + "namespace", + "plan_id", + "journal", + ] as const) { + const value = snapshot(); + if (field === "journal") { + value.journal_incomplete = true; + } else { + value.receipt[field] = "invalid"; + } + let calls = 0; + await expect( + recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: value, + invoke: () => { + calls++; + throw new Error("synthetic-private-canary"); + }, + }) + ).rejects.toThrow("cleanup is unconfirmed"); + expect(calls).toBe(0); + } +}); + +test("malformed native selection never admits a recovery effect", async () => { + for (const field of [ + "run", + "phase", + "eligible", + "data_retained", + "same_boot", + "selection_sha256", + ]) { + const selected: Record = selection(); + selected[field] = "synthetic-private-canary"; + let calls = 0; + const failure = await recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: snapshot(), + invoke: async () => { + calls++; + return selected; + }, + }).catch((error: unknown) => error); + expect(failure).toBeInstanceOf(Error); + expect(String(failure)).not.toContain("synthetic-private-canary"); + expect(calls).toBe(1); + } +}); + +test("lost recovery replies are never replayed and private diagnostics are omitted", async () => { + let calls = 0; + const failure = await recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: snapshot(), + invoke: async () => { + if (++calls === 1) { + return selection(); + } + throw new Error("synthetic-private-canary"); + }, + }).catch((error: unknown) => error); + expect(calls).toBe(2); + expect(String(failure)).toContain("cleanup is unconfirmed"); + expect(Bun.inspect(failure)).not.toContain("synthetic-private-canary"); +}); + +test("graph-only confirmation cannot claim publisher or dependency retirement", async () => { + for (const field of [ + "run", + "phase", + "recovered", + "data_retained", + "same_boot", + "publisher_retired", + "reservation_released", + ]) { + const reply: Record = recovered(); + reply[field] = "synthetic-private-canary"; + let calls = 0; + const failure = await recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: snapshot(), + invoke: async () => (++calls === 1 ? selection() : reply), + }).catch((error: unknown) => error); + expect(calls).toBe(2); + expect(String(failure)).toContain("cleanup is unconfirmed"); + expect(Bun.inspect(failure)).not.toContain("synthetic-private-canary"); + } +}); + +test("recovery acknowledgement does not replace fresh exact retained-data observations", async () => { + for (const fault of [ + "container", + "volume", + "name", + "id", + "missing", + "foreign", + "journal", + "phase", + "unfinished", + ]) { + const final = snapshot(true); + if (fault === "container") { + final.observations["container:web"].state = "running"; + } + if (fault === "volume") { + final.observations["volume:data"].state = "absent"; + } + if (fault === "name") { + final.receipt.resources["volume:data"].name = "changed"; + } + if (fault === "id") { + final.receipt.resources["container:web"].id = "1".repeat(64); + } + if (fault === "missing") { + Reflect.deleteProperty(final.receipt.resources, "volume:data"); + } + if (fault === "foreign") { + final.receipt.owner = "1".repeat(32); + } + if (fault === "journal") { + final.journal_incomplete = true; + } + if (fault === "phase") { + final.receipt.phase = "cleanup-intent"; + } + if (fault === "unfinished") { + final.interrupted_start_cleanup_incomplete = true; + } + const invoke: typeof invokeNativeRuntime = async ({ args }) => { + if (args[1] === "inspect-interrupted-start-cleanup") { + return selection(); + } + if (args[1] === "recover-interrupted-start-cleanup") { + return recovered(); + } + return final; + }; + await expect( + recoverNativeInterruptedStartupCleanup({ + runtime, + projectRoot: "/not-used", + run, + snapshot: snapshot(), + invoke, + }) + ).rejects.toThrow("cleanup is unconfirmed"); + } +}); diff --git a/tests/native-project-mapping-command.test.ts b/tests/native-project-mapping-command.test.ts new file mode 100644 index 000000000..8e707d586 --- /dev/null +++ b/tests/native-project-mapping-command.test.ts @@ -0,0 +1,73 @@ +import { expect, test } from "bun:test"; +import { parseNativeRunMappingRecoveryOptions as parse } from "../src/backends/native-project-mapping-command.ts"; + +const selected = "a".repeat(64); + +test("ordinary Doctor does not select a mapping migration", () => { + expect(parse({ otherOptions: false })).toBeNull(); +}); + +test("inspection refuses every mutation selector", () => { + expect( + parse({ action: "inspect", branch: "feature-api", otherOptions: false }) + ).toEqual({ action: "inspect" }); + for (const options of [ + { expectSelection: selected }, + { acceptLegacyDeviceRebind: true }, + ]) { + expect(() => + parse({ action: "inspect", ...options, otherOptions: false }) + ).toThrow("read-only"); + } +}); + +test("repair requires both an exact inspection selection and explicit legacy acceptance", () => { + expect( + parse({ + action: "repair", + expectSelection: selected, + acceptLegacyDeviceRebind: true, + otherOptions: false, + }) + ).toEqual({ action: "repair", expectSelection: selected }); + for (const options of [ + {}, + { expectSelection: selected }, + { acceptLegacyDeviceRebind: true }, + { expectSelection: "a".repeat(63), acceptLegacyDeviceRebind: true }, + { expectSelection: "A".repeat(64), acceptLegacyDeviceRebind: true }, + { expectSelection: `${selected}\n`, acceptLegacyDeviceRebind: true }, + ]) { + expect(() => + parse({ action: "repair", ...options, otherOptions: false }) + ).toThrow("requires"); + } +}); + +test("unknown actions and orphaned selectors refuse before project lookup", () => { + expect(() => parse({ action: "force", otherOptions: false })).toThrow( + "requires" + ); + for (const options of [ + { expectSelection: selected }, + { acceptLegacyDeviceRebind: true }, + { branch: "feature-api" }, + ]) { + expect(() => parse({ ...options, otherOptions: false })).toThrow( + "require --native-run-mapping" + ); + } +}); + +test("mapping migration cannot combine with ordinary repair, domain or browser flows", () => { + for (const action of ["inspect", "repair"]) { + expect(() => + parse({ + action, + expectSelection: selected, + acceptLegacyDeviceRebind: true, + otherOptions: true, + }) + ).toThrow("cannot be combined"); + } +}); diff --git a/tests/native-project-restart.test.ts b/tests/native-project-restart.test.ts index c7f12d496..a523837a7 100644 --- a/tests/native-project-restart.test.ts +++ b/tests/native-project-restart.test.ts @@ -16,6 +16,7 @@ import { nativeRestartSelection, preflightNativeRestart, } from "../src/backends/native-project-restart-preflight.ts"; +import { selectNativeRetainedImages } from "../src/backends/native-project-retained-images.ts"; import { completeNativeRestartCleanup, loadNativeProjectRun, @@ -534,6 +535,7 @@ async function listenerIntentFixture() { run, dependencyFile: path, dependencies: { + retainedImages: async () => new Map(), prepare: async ({ envName }) => { expect(envName).toBe("qa"); return input; @@ -633,6 +635,302 @@ test("cleaned retry with absent listeners reaches startup hooks before actual id } }); +test("explicit stopped recovery without an intent retains the normal cleanup and startup sequence", async () => { + const listener = await listenerIntentFixture(); + try { + const f = fixture(); + expect( + await restartNativeProject({ + ...f.options, + preflight: async (selected, { cleanedRetry }) => { + expect(cleanedRetry).toBe(false); + await preflightNativeRestart({ + ...listener.options, + run: selected, + cleanedRetry, + recoverStopped: true, + }); + }, + }) + ).toBe(0); + expect(listener.calls).toEqual([ + "graph inspect", + "runtime status", + "runtime probe", + "review", + "graph inspect", + ]); + expect(f.events).toEqual([ + "capture", + "intent", + "down", + "finalized", + "start", + "remove", + "unlock", + "serving", + ]); + expect(JSON.parse(await readFile(listener.path, "utf8"))).toEqual( + listener.selection + ); + } finally { + await rm(listener.directory, { recursive: true, force: true }); + } +}); + +test("explicit stopped recovery refuses incomplete, foreign or missing retained observations before cleanup", async () => { + const listener = await listenerIntentFixture(); + try { + const complete = stoppedRetainedGraph(); + for (const observed of [ + { ...complete, journal_incomplete: true }, + { ...complete, receipt: { ...complete.receipt, run: "9".repeat(32) } }, + { ...complete, receipt: { ...complete.receipt, owner: "9".repeat(32) } }, + { + ...complete, + observations: { + ...complete.observations, + "volume:data": { state: "absent" }, + }, + }, + { + ...complete, + observations: { + ...complete.observations, + "container:app": { state: "unknown" }, + }, + }, + ]) { + const f = fixture(); + const calls: string[] = []; + await expect( + restartNativeProject({ + ...f.options, + preflight: async () => + await preflightNativeRestart({ + ...listener.options, + recoverStopped: true, + dependencies: { + ...listener.options.dependencies, + invoke: async ({ args }) => { + calls.push(String(args[1])); + return observed; + }, + }, + }), + }) + ).rejects.toThrow("cannot confirm the stopped retained graph"); + expect(calls).toEqual(["inspect"]); + expect(f.events).toEqual([]); + expect(f.state.pending).toBeNull(); + } + } finally { + await rm(listener.directory, { recursive: true, force: true }); + } +}); + +test("explicit stopped recovery rechecks compute and retained data after review", async () => { + const listener = await listenerIntentFixture(); + try { + for (const changed of ["compute", "data"]) { + const f = fixture(); + let inspections = 0; + await expect( + restartNativeProject({ + ...f.options, + preflight: async () => + await preflightNativeRestart({ + ...listener.options, + recoverStopped: true, + dependencies: { + ...listener.options.dependencies, + invoke: async (options) => { + if (options.args[1] !== "inspect") { + return await listener.options.dependencies!.invoke!( + options + ); + } + inspections += 1; + const observed = stoppedRetainedGraph(); + if (inspections === 2) { + observed.observations[ + changed === "compute" ? "container:app" : "volume:data" + ].state = changed === "compute" ? "present" : "absent"; + } + return observed; + }, + }, + }), + }) + ).rejects.toThrow("stopped recovery changed during review"); + expect(inspections).toBe(2); + expect(f.events).toEqual([]); + expect(f.state.pending).toBeNull(); + } + } finally { + await rm(listener.directory, { recursive: true, force: true }); + } +}); + +test("frontend recovery on an active graph still requires live host listeners", async () => { + const listener = await listenerIntentFixture(); + try { + const f = fixture(); + await expect( + restartNativeProject({ + ...f.options, + preflight: async () => + await preflightNativeRestart({ + ...listener.options, + recoverStopped: true, + dependencies: { + ...listener.options.dependencies, + invoke: async (options) => + options.args[1] === "inspect" + ? { + ...stoppedRetainedGraph(), + receipt: { + ...stoppedRetainedGraph().receipt, + phase: "ready", + }, + } + : await listener.options.dependencies!.invoke!(options), + }, + }), + }) + ).rejects.toThrow("bounded explicit listener selection; values omitted"); + expect(listener.calls).toEqual([ + "runtime status", + "runtime probe", + "graph dependency-discover", + ]); + expect(f.events).toEqual([]); + expect(f.state.pending).toBeNull(); + } finally { + await rm(listener.directory, { recursive: true, force: true }); + } +}); + +test("stopped recovery reviews retained image IDs and resolves tags only after original input changes", async () => { + const listener = await listenerIntentFixture(); + try { + for (const changedOriginal of [false, true]) { + let imageResolutions = 0; + let reviewedImage: unknown; + const oldImage = `sha256:${"7".repeat(64)}`; + const newImage = `sha256:${"8".repeat(64)}`; + await preflightNativeRestart({ + ...listener.options, + recoverStopped: true, + dependencies: { + ...listener.options.dependencies, + retainedImages: selectNativeRetainedImages, + prepare: async () => ({ + originalSha256: changedOriginal ? "2".repeat(64) : "1".repeat(64), + environmentFiles: [], + serviceNames: ["app"], + normalizedComposeJson: JSON.stringify({ + services: { app: { image: "example/app:latest" } }, + }), + managedEnvironment: {}, + lifecycleHostEnvironment: {}, + effectiveEnvName: "qa", + }), + invoke: async (options) => { + if (options.args[1] === "inspect") { + const observed = stoppedRetainedGraph(); + return { + ...observed, + receipt: { + ...observed.receipt, + normalized_input: { + namespace: run.namespace, + original_compose_sha256: "1".repeat(64), + normalized_compose_sha256: "3".repeat(64), + }, + resources: { + default: observed.receipt.resources.default, + data: observed.receipt.resources.data, + "container:app": { + kind: "container", + key: "app", + image: oldImage, + }, + }, + }, + }; + } + if (options.args[1] === "ensure-image") { + imageResolutions += 1; + return { image_id: newImage }; + } + return await listener.options.dependencies!.invoke!(options); + }, + review: async (review) => { + reviewedImage = JSON.parse(review.input.normalizedComposeJson) + .services.app.image; + return await review.run({ + planId: run.planId, + namespace: run.namespace, + report: {}, + projectArgs: [], + }); + }, + }, + }); + expect(reviewedImage).toBe(changedOriginal ? newImage : oldImage); + expect(imageResolutions).toBe(changedOriginal ? 1 : 0); + } + } finally { + await rm(listener.directory, { recursive: true, force: true }); + } +}); + +test("invalid retained image provenance refuses stopped preflight before tag resolution", async () => { + const listener = await listenerIntentFixture(); + try { + const f = fixture(); + const calls: string[] = []; + await expect( + restartNativeProject({ + ...f.options, + preflight: async () => + await preflightNativeRestart({ + ...listener.options, + recoverStopped: true, + dependencies: { + ...listener.options.dependencies, + retainedImages: selectNativeRetainedImages, + invoke: async (options) => { + calls.push(String(options.args[1])); + if (options.args[1] === "inspect") { + const observed = stoppedRetainedGraph(); + return { + ...observed, + receipt: { + ...observed.receipt, + normalized_input: { + namespace: "9".repeat(64), + original_compose_sha256: "1".repeat(64), + normalized_compose_sha256: "3".repeat(64), + }, + }, + }; + } + return await listener.options.dependencies!.invoke!(options); + }, + }, + }), + }) + ).rejects.toThrow("retained image provenance is invalid"); + expect(calls).not.toContain("ensure-image"); + expect(f.events).toEqual([]); + expect(f.state.pending).toBeNull(); + } finally { + await rm(listener.directory, { recursive: true, force: true }); + } +}); + test("persisted retaining cleanup keeps its mapping and retry captures listeners after hooks", async () => { const listener = await listenerIntentFixture(); try { @@ -678,6 +976,7 @@ test("persisted retaining cleanup keeps its mapping and retry captures listeners await nativeProjectDown({ runtime, scope: retainedScope, + deferFinalization: true, invoke: listener.options.dependencies?.invoke, before: async () => { throw new Error("down hooks must not replay"); @@ -915,3 +1214,132 @@ test("malformed dependency intent refuses cleaned retry without capture or start await rm(listener.directory, { recursive: true, force: true }); } }); + +test("active legacy restart preserves adapted labels and rechecks authenticated selection before cleanup", async () => { + const root = await mkdtemp(join(tmpdir(), "native-active-review-")); + try { + const projectRoot = await realpath(root); + for (const drift of [false, true]) { + const f = fixture(); + const scoped = { + ...scope, + projectRoot, + projectDir: join(projectRoot, ".hack"), + branch: "feature-a", + }; + const calls: string[] = []; + let compatible = false; + const input = { + originalSha256: "1".repeat(64), + environmentFiles: [], + serviceNames: ["app"], + normalizedComposeJson: JSON.stringify({ + services: { + app: { + image: `sha256:${"a".repeat(64)}`, + labels: { + caddy: "app.hack, app.hack.gy", + "caddy.tls": "internal", + "caddy.reverse_proxy": "{{upstreams 3000}}", + }, + }, + }, + }), + managedEnvironment: {}, + lifecycleHostEnvironment: {}, + effectiveEnvName: "qa", + }; + const result = restartNativeProject({ + ...f.options, + scope: scoped, + preflight: async () => + await preflightNativeRestart({ + runtime: { binary: "/unused", home: "/candidate" }, + scope: scoped, + composeFile: join(projectRoot, "compose.yml"), + run, + dependencies: { + prepare: async () => input, + adapt: async ({ input: value }) => value, + dependencies: async () => [], + invoke: async ({ args }) => { + calls.push(args[1] ?? "unknown"); + if (args[1] === "status") { + return { network: "internet" }; + } + if (args[1] === "probe") { + return { admitted: true }; + } + if (args[0] === "project") { + if (args.includes("--normalized-file")) { + const value = JSON.parse( + await readFile( + args[args.indexOf("--normalized-file") + 1] ?? "", + "utf8" + ) + ); + expect(value.services.app.labels.caddy).toBe( + "app.hack, app.hack.gy" + ); + expect(args).not.toContain("--branch"); + } + return { + plan_id: "e".repeat(64), + plan: { + source: projectRoot, + namespace: args.includes("--branch") + ? "f".repeat(64) + : run.namespace, + compose_sha256: input.originalSha256, + services: { app: { active: true } }, + }, + }; + } + if (args[1] === "run-selection") { + expect(args[args.indexOf("--service") + 1]).toBe("app"); + return { + ok: true, + run: run.run, + owner: run.owner, + namespace: run.namespace, + plan: run.planId, + service: "app", + container: "2".repeat(64), + boot: compatible && drift ? "changed-boot" : "owned-boot", + generation: "3".repeat(64), + }; + } + if (args[1] === "source-compatibility") { + compatible = true; + return { + run: run.run, + owner: run.owner, + namespace: run.namespace, + plan: run.planId, + reviewed_plan: "e".repeat(64), + source_revision: null, + }; + } + throw new Error("Unexpected effect"); + }, + }, + }), + }); + if (drift) { + await expect(result).rejects.toThrow("active review identity changed"); + expect(f.events).toEqual([]); + } else { + expect(await result).toBe(0); + expect(f.events).toContain("down"); + } + expect(calls.at(-1)).toBe("run-selection"); + expect(calls.filter((action) => action === "run-selection")).toHaveLength( + 3 + ); + expect(calls).not.toContain("restore-selection"); + expect(calls).not.toContain("run-service"); + } + } finally { + await rm(root, { recursive: true, force: true }); + } +}); diff --git a/tests/native-project-restore.test.ts b/tests/native-project-restore.test.ts index acd673a1e..7e539d6ba 100644 --- a/tests/native-project-restore.test.ts +++ b/tests/native-project-restore.test.ts @@ -173,3 +173,14 @@ test("eligible hostname rollback checks current native provenance even at the ow }) ).rejects.toThrow("route change refused"); }); + +test("legacy review refuses a different valid restore generation", async () => { + await expect( + selectNativeProjectRestore({ + ...options, + restore: saved, + review: { ...options.review, retainedGeneration: "1".repeat(64) }, + invoke: async () => ({ ...selected, generation: "2".repeat(64) }), + }) + ).rejects.toThrow("restore selection changed"); +}); diff --git a/tests/native-project-retained-images.test.ts b/tests/native-project-retained-images.test.ts new file mode 100644 index 000000000..0aef126b5 --- /dev/null +++ b/tests/native-project-retained-images.test.ts @@ -0,0 +1,165 @@ +import { expect, test } from "bun:test"; +import { selectNativeRetainedImages } from "../src/backends/native-project-retained-images.ts"; + +const saved = { + run: "a".repeat(32), + owner: "b".repeat(32), + namespace: "c".repeat(64), + planId: "d".repeat(64), +}; +const image = `sha256:${"e".repeat(64)}`; +const originalSha256 = "f".repeat(64); +const options = { + runtime: { binary: "/not-invoked", home: "/fixture" }, + projectRoot: "/fixture/project", + originalSha256, + restore: saved, +}; +function observation() { + return { + journal_incomplete: false, + receipt: { + run: saved.run, + owner: saved.owner, + namespace: saved.namespace, + plan_id: saved.planId, + phase: "stopped-data-retained", + normalized_input: { + namespace: saved.namespace, + original_compose_sha256: originalSha256, + normalized_compose_sha256: "1".repeat(64), + }, + resources: { + "container:web": { kind: "container", key: "web", image }, + "volume:data": { kind: "volume", key: "data" }, + }, + }, + observations: { + "container:web": { state: "absent" }, + "volume:data": { state: "present" }, + }, + }; +} +test("unchanged retained input selects verified content IDs without resolving tags", async () => { + const calls: string[][] = []; + const selected = await selectNativeRetainedImages({ + ...options, + invoke: async ({ args }) => { + calls.push([...args]); + return observation(); + }, + }); + expect([...selected]).toEqual([["web", image]]); + expect(calls).toEqual([ + ["graph", "inspect", "--run-id", saved.run, "--json"], + ]); +}); +test("fresh input never observes retained state", async () => { + expect( + ( + await selectNativeRetainedImages({ + ...options, + restore: undefined, + invoke: async () => { + throw new Error("unexpected access"); + }, + }) + ).size + ).toBe(0); +}); +test("edited original input and legacy absence keep ordinary resolution", async () => { + expect( + ( + await selectNativeRetainedImages({ + ...options, + originalSha256: "2".repeat(64), + invoke: async () => observation(), + }) + ).size + ).toBe(0); + const { normalized_input: _normalized, ...legacy } = observation().receipt; + expect( + ( + await selectNativeRetainedImages({ + ...options, + invoke: async () => ({ ...observation(), receipt: legacy }), + }) + ).size + ).toBe(0); +}); +test("stale ownership and uncertain compute or data refuse image reuse", async () => { + const baseline = observation(); + for (const changed of [ + { ...baseline, journal_incomplete: true }, + ...["run", "owner", "namespace", "plan_id", "phase"].map((key) => ({ + ...baseline, + receipt: { ...baseline.receipt, [key]: "changed" }, + })), + { + ...baseline, + observations: { + ...baseline.observations, + "container:web": { state: "running" }, + }, + }, + { + ...baseline, + observations: { + ...baseline.observations, + "volume:data": { state: "absent" }, + }, + }, + ]) { + await expect( + selectNativeRetainedImages({ ...options, invoke: async () => changed }) + ).rejects.toThrow("image selection changed"); + } +}); +test("malformed provenance or image and mismatched resource keys refuse reuse", async () => { + const baseline = observation(); + for (const field of [ + "namespace", + "original_compose_sha256", + "normalized_compose_sha256", + ]) { + await expect( + selectNativeRetainedImages({ + ...options, + invoke: async () => ({ + ...baseline, + receipt: { + ...baseline.receipt, + normalized_input: { + ...baseline.receipt.normalized_input, + [field]: "invalid", + }, + }, + }), + }) + ).rejects.toThrow("provenance is invalid"); + } + for (const resource of [ + { ...baseline.receipt.resources["container:web"], image: "redis:latest" }, + { ...baseline.receipt.resources["container:web"], key: "worker" }, + ]) { + await expect( + selectNativeRetainedImages({ + ...options, + invoke: async () => ({ + ...baseline, + receipt: { + ...baseline.receipt, + resources: { + ...baseline.receipt.resources, + "container:web": resource, + }, + }, + }), + }) + ).rejects.toThrow( + resource.key === "web" + ? "image identity is invalid" + : "image selection changed" + ); + } +}); diff --git a/tests/native-project-review.test.ts b/tests/native-project-review.test.ts index e47f7b814..4e5d52d0e 100644 --- a/tests/native-project-review.test.ts +++ b/tests/native-project-review.test.ts @@ -1,9 +1,16 @@ import { afterEach, expect, test } from "bun:test"; -import { chmod, mkdtemp, readFile, rm, stat } from "node:fs/promises"; +import { chmod, mkdtemp, readFile, realpath, rm, stat } from "node:fs/promises"; import { tmpdir } from "node:os"; import { dirname, join } from "node:path"; import { prepareNativeProjectInput } from "../src/backends/native-project-input.ts"; -import { withNativeProjectReview } from "../src/backends/native-project-review.ts"; +import { selectNativeProjectRestore } from "../src/backends/native-project-restore.ts"; +import { + prepareNativeReviewBranch, + selectNativeProjectReviewIdentity, + verifyNativeActiveReview, + withNativeProjectReview, +} from "../src/backends/native-project-review.ts"; +import type { invokeNativeRuntime } from "../src/backends/native-runtime-client.ts"; const roots: string[] = []; afterEach(async () => { @@ -131,3 +138,377 @@ test("noncanonical branch review refuses before invoking the executor", async () }) ).rejects.toThrow("canonical branch"); }); + +const retained = { + run: "d".repeat(32), + owner: "e".repeat(32), + namespace: "a".repeat(64), + planId: "b".repeat(64), +}; + +async function legacyFixture() { + const options = await fixture(); + const calls: string[][] = []; + const selected: Record = { + run: retained.run, + owner: retained.owner, + namespace: retained.namespace, + plan: retained.planId, + generation: "f".repeat(64), + ok: true, + service: "web", + container: "1".repeat(64), + boot: "owned-boot", + }; + const unbranched: Record = { + namespace: retained.namespace, + source: await realpath(options.projectRoot), + compose_sha256: options.input.originalSha256, + services: { web: { active: true } }, + }; + const invoke: typeof invokeNativeRuntime = async ({ args }) => { + calls.push([...args]); + if (args[0] === "graph") { + return selected; + } + if (!args.includes("--branch")) { + return { plan_id: retained.planId, plan: unbranched }; + } + return { + plan_id: retained.planId, + plan: { ...unbranched, namespace: "c".repeat(64) }, + }; + }; + return { options, calls, selected, unbranched, invoke }; +} + +test("legacy retained review uses native unbranched authority and rechecks restore selection", async () => { + const { options, calls, invoke } = await legacyFixture(); + let temporary = ""; + await withNativeProjectReview({ + ...options, + branch: "feature-a", + retained, + invoke, + run: async (review) => { + expect(review.namespace).toBe(retained.namespace); + expect(review.projectArgs).not.toContain("--branch"); + temporary = + review.projectArgs[ + review.projectArgs.indexOf("--normalized-file") + 1 + ] ?? ""; + expect(await Bun.file(temporary).text()).toBe( + options.input.normalizedComposeJson + ); + const selection = await selectNativeProjectRestore({ + runtime: options.runtime, + projectRoot: options.projectRoot, + restore: retained, + review, + invoke, + }); + expect(selection.run).toBe(retained.run); + expect(selection.flags).toEqual(["--expect-generation", "f".repeat(64)]); + }, + }); + expect(calls.map((args) => args.slice(0, 2).join(" "))).toEqual([ + "project plan", + "graph restore-selection", + "project plan", + "project plan", + "graph restore-selection", + ]); + expect(await Bun.file(temporary).exists()).toBe(false); +}); + +test("fresh and current branch-native review never consult legacy restore authority", async () => { + for (const saved of [undefined, { ...retained, namespace: "c".repeat(64) }]) { + const { options, calls, invoke } = await legacyFixture(); + await withNativeProjectReview({ + ...options, + branch: "feature-a", + retained: saved, + invoke, + run: async (review) => { + expect(review.namespace).toBe("c".repeat(64)); + expect(review.projectArgs).toContain("--branch"); + }, + }); + expect(calls).toHaveLength(2); + expect(calls.every((args) => args.includes("--branch"))).toBe(true); + } +}); + +test("legacy review rejects changed native selection before unbranched review", async () => { + for (const key of ["run", "owner", "namespace", "plan", "generation"]) { + const { options, calls, selected, invoke } = await legacyFixture(); + selected[key] = "invalid"; + await expect( + withNativeProjectReview({ + ...options, + branch: "feature-a", + retained, + invoke, + run: async () => { + throw new Error("unexpected callback"); + }, + }) + ).rejects.toThrow("restore selection changed"); + expect(calls).toHaveLength(2); + } +}); + +test("legacy review rejects a different project, namespace or edited original without admission", async () => { + for (const key of ["source", "namespace", "compose_sha256"]) { + const { options, calls, unbranched, invoke } = await legacyFixture(); + const originalInvoke: typeof invokeNativeRuntime = async (request) => { + if (request.args[0] === "project" && !request.args.includes("--branch")) { + unbranched[key] = "foreign"; + } + return await invoke(request); + }; + await expect( + withNativeProjectReview({ + ...options, + branch: "feature-a", + retained, + invoke: originalInvoke, + run: async () => { + throw new Error("unexpected callback"); + }, + }) + ).rejects.toThrow("retained project review changed"); + expect(calls).toHaveLength(3); + } +}); + +test("legacy review cleans temporary input and refuses selection drift after review", async () => { + const { options, selected, invoke } = await legacyFixture(); + let temporary = ""; + await expect( + withNativeProjectReview({ + ...options, + branch: "feature-a", + retained, + invoke, + run: async (review) => { + temporary = + review.projectArgs[ + review.projectArgs.indexOf("--normalized-file") + 1 + ] ?? ""; + selected.owner = "0".repeat(32); + await selectNativeProjectRestore({ + runtime: options.runtime, + projectRoot: options.projectRoot, + restore: retained, + review, + invoke, + }); + }, + }) + ).rejects.toThrow("restore selection changed"); + expect(await Bun.file(temporary).exists()).toBe(false); +}); + +test("legacy identity is selected before route normalization and original adapted labels survive", async () => { + const { options, invoke } = await legacyFixture(); + const compose = JSON.parse(options.input.normalizedComposeJson); + compose.services.web.labels = { + caddy: "app.hack, app.hack.gy", + "caddy.reverse_proxy": "{{upstreams 3000}}", + }; + const input = { + ...options.input, + normalizedComposeJson: JSON.stringify(compose), + }; + const scope = { + projectRoot: options.projectRoot, + projectDir: join(options.projectRoot, ".hack"), + nativeHome: options.runtime.home, + branch: "feature-a", + }; + const prepared = await prepareNativeReviewBranch({ + ...options, + scope, + input, + retained, + invoke, + }); + expect(prepared.input.normalizedComposeJson).toBe( + input.normalizedComposeJson + ); + expect(prepared.identity?.branch).toBeNull(); + await withNativeProjectReview({ + ...options, + input: prepared.input, + branch: scope.branch, + retained, + invoke, + reviewIdentity: prepared.identity, + run: async (review) => { + expect(review.namespace).toBe(retained.namespace); + expect(review.projectArgs).not.toContain("--branch"); + }, + }); +}); + +test("early identity proof is rechecked after route preparation before normalized publication", async () => { + for (const key of ["generation", "owner"]) { + const { options, selected, calls, invoke } = await legacyFixture(); + const identity = await selectNativeProjectReviewIdentity({ + ...options, + branch: "feature-a", + retained, + invoke, + }); + selected[key] = "1".repeat(key === "generation" ? 64 : 32); + calls.splice(0); + await expect( + withNativeProjectReview({ + ...options, + branch: "feature-a", + retained, + invoke, + reviewIdentity: identity, + run: async () => { + throw new Error("unexpected admission"); + }, + }) + ).rejects.toThrow("selection changed"); + expect(calls.some((args) => args.includes("--normalized-file"))).toBe( + false + ); + } +}); + +test("active branch-native needs no stopped selection and legacy selects authenticated service authority", async () => { + const { options, calls, invoke } = await legacyFixture(); + const scoped = await selectNativeProjectReviewIdentity({ + ...options, + branch: "feature-a", + retained: { ...retained, namespace: "c".repeat(64) }, + retainedMode: "active", + invoke, + }); + expect(scoped.branch).toBe("feature-a"); + expect(calls).toHaveLength(1); + calls.splice(0); + const identity = await selectNativeProjectReviewIdentity({ + ...options, + branch: "feature-a", + retained, + retainedMode: "active", + invoke, + }); + expect(identity.branch).toBeNull(); + expect(identity.retainedGeneration).toBeUndefined(); + expect(identity.activeProof?.service).toBe("web"); + await verifyNativeActiveReview({ + ...options, + retained, + proof: identity.activeProof, + invoke, + }); + expect( + calls.filter((args) => args[0] === "graph").map((args) => args[1]) + ).toEqual(["run-selection", "run-selection"]); +}); + +test("deferred retained review rejects malformed branch before hooks or runtime selection", async () => { + const { options, calls, invoke } = await legacyFixture(); + await expect( + prepareNativeReviewBranch({ + ...options, + retained, + invoke, + phase: "before-runtime", + scope: { + projectRoot: options.projectRoot, + projectDir: join(options.projectRoot, ".hack"), + nativeHome: options.runtime.home, + branch: "foreign/branch", + }, + }) + ).rejects.toThrow("canonical branch"); + expect(calls).toHaveLength(0); +}); + +test("active proof refuses every tuple drift after compatibility and unavailable service authority", async () => { + for (const key of [ + "run", + "owner", + "namespace", + "plan", + "service", + "container", + "boot", + "generation", + ]) { + const { options, selected, invoke } = await legacyFixture(); + const identity = await selectNativeProjectReviewIdentity({ + ...options, + branch: "feature-a", + retained, + retainedMode: "active", + invoke, + }); + selected[key] = + key === "boot" + ? "another-boot" + : "2".repeat(key === "run" || key === "owner" ? 32 : 64); + await expect( + verifyNativeActiveReview({ + ...options, + retained, + proof: identity.activeProof, + invoke, + }) + ).rejects.toThrow("changed"); + } + for (const services of [{}, { web: { active: false } }]) { + const { options, calls, unbranched, invoke } = await legacyFixture(); + unbranched.services = services; + await expect( + selectNativeProjectReviewIdentity({ + ...options, + branch: "feature-a", + retained, + retainedMode: "active", + invoke, + }) + ).rejects.toThrow("no supported service authority"); + expect(calls).toHaveLength(1); + } +}); + +test("active legacy proof never grants an unrelated source and propagates stale dependency refusal", async () => { + for (const failure of ["source", "dependencies"]) { + const { options, invoke, unbranched, calls } = await legacyFixture(); + const checked: typeof invokeNativeRuntime = async (request) => { + if (failure === "dependencies" && request.args[1] === "run-selection") { + throw new Error("stale dependency"); + } + if (failure === "source") { + unbranched.source = "/foreign/project"; + } + return await invoke(request); + }; + await expect( + selectNativeProjectReviewIdentity({ + ...options, + branch: "feature-a", + retained, + retainedMode: "active", + invoke: checked, + }) + ).rejects.toThrow( + failure === "source" ? "review changed" : "stale dependency" + ); + expect( + calls.some( + (args) => + args[1] === "run-service" || args.includes("--normalized-file") + ) + ).toBe(false); + } +}); diff --git a/tests/native-project-run-filesystem.test.ts b/tests/native-project-run-filesystem.test.ts new file mode 100644 index 000000000..742a1c949 --- /dev/null +++ b/tests/native-project-run-filesystem.test.ts @@ -0,0 +1,331 @@ +import { afterEach, expect, test } from "bun:test"; +import { createHash } from "node:crypto"; +import { + chmod, + lstat, + mkdir, + mkdtemp, + readFile, + realpath, + rename, + rm, + writeFile, +} from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { + inspectNativeProjectRunFilesystemRecovery as inspect, + loadNativeProjectRun as load, + type NativeProjectRunScope, + recoverNativeProjectRunFilesystem as recover, + saveNativeProjectRun as save, +} from "../src/backends/native-project-run.ts"; +import type { NativeRuntimeSelection } from "../src/backends/native-runtime-client.ts"; + +const roots: string[] = []; +afterEach(async () => { + await Promise.all( + roots.splice(0).map((root) => rm(root, { recursive: true, force: true })) + ); +}); + +const run = { + run: "a".repeat(32), + owner: "b".repeat(32), + namespace: "c".repeat(64), + planId: "d".repeat(64), + effectiveEnvName: "qa", + profiles: ["worker"], + aws: { profile: "livenation_qa", region: "us-east-1" }, +}; +const runtime: NativeRuntimeSelection = { + binary: "/unused/hack-native", + home: "/unused/home", +}; + +async function fixture() { + const projectRoot = await realpath( + await mkdtemp(join(tmpdir(), "native-mapping-reboot-")) + ); + roots.push(projectRoot); + const projectDir = join(projectRoot, ".hack"); + const nativeHome = join(projectRoot, "candidate"); + await mkdir(projectDir); + await mkdir(nativeHome); + const scope: NativeProjectRunScope = { + projectRoot, + projectDir, + nativeHome, + branch: "event-agent", + }; + await save({ ...scope, run }); + const key = createHash("sha256") + .update(JSON.stringify(scope.branch)) + .digest("hex"); + const directory = join(projectDir, ".internal/native-runs"); + const file = join(directory, `${key}.json`); + const original = JSON.parse(await readFile(file, "utf8")); + const oldDevice = original.scope.rootIdentity.dev + 1; + for (const name of ["rootIdentity", "dirIdentity", "homeIdentity"]) { + original.scope[name].dev = oldDevice; + } + const oldBytes = JSON.stringify(original); + await writeFile(file, oldBytes); + return { scope, directory, file, key, oldDevice, oldBytes, original }; +} + +function authority( + pool: Awaited>, + opts: { + onCall?: (call: number) => Promise; + graph?: Record; + status?: Record; + } = {} +) { + let call = 0; + return async (input: { + readonly args: readonly string[]; + }): Promise => { + call++; + await opts.onCall?.(call); + if (input.args[0] === "runtime") { + return ( + opts.status ?? { + phase: "running", + process_alive: true, + persistent_disks_identified: true, + project_share: { + project: pool.scope.projectRoot, + guest_path: "/mnt/hack-projects/fixture", + device: (await lstat(pool.scope.projectRoot)).dev, + inode: (await lstat(pool.scope.projectRoot)).ino, + unfiltered_source: true, + }, + } + ); + } + return ( + opts.graph ?? { + journal_incomplete: false, + receipt: { + run: run.run, + owner: run.owner, + namespace: run.namespace, + plan_id: run.planId, + phase: "ready-observed", + source: { + shared: { + project: pool.scope.projectRoot, + guest_path: "/mnt/hack-projects/fixture", + device: pool.oldDevice, + inode: (await lstat(pool.scope.projectRoot)).ino, + unfiltered_source: true, + }, + }, + }, + } + ); + }; +} + +test("explicit mapping repair changes only three devices and retains byte-exact audit", async () => { + const pool = await fixture(); + const invoke = authority(pool); + await expect(load(pool.scope)).rejects.toThrow("Native project run mapping"); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + expect(selected.qualification).toContain("unproven"); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); + const repaired = await recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }); + expect(repaired.repaired).toBe(true); + expect(repaired.selectionSha256).toBe(selected.selectionSha256); + expect(await load(pool.scope)).toEqual(run); + const after = JSON.parse(await readFile(pool.file, "utf8")); + expect(after.run).toEqual(pool.original.run); + for (const name of ["rootIdentity", "dirIdentity", "homeIdentity"]) { + expect(after.scope[name].ino).toBe(pool.original.scope[name].ino); + expect(after.scope[name].dev).toBe(selected.newDevice); + } + const audit = JSON.parse(await readFile(repaired.auditPath, "utf8")); + expect(audit.mapping).toBe(pool.oldBytes); + expect(audit.mappingSha256).toBe(selected.mappingSha256); + await expect( + inspect({ scope: pool.scope, runtime, invoke }) + ).rejects.toThrow(); +}); + +test("partial device changes, changed inode, paths, branch and malformed mapping refuse", async () => { + for (const change of [ + "partial", + "inode", + "path", + "branch", + "schema", + "mode", + "link", + ]) { + const pool = await fixture(); + const invoke = authority(pool); + if (change === "mode") { + await chmod(pool.file, 0o644); + } else if (change === "link") { + await writeFile(`${pool.file}.link`, "foreign"); + await rm(`${pool.file}.link`); + const { link } = await import("node:fs/promises"); + await link(pool.file, `${pool.file}.link`); + } else { + const value = JSON.parse(pool.oldBytes); + if (change === "partial") { + value.scope.dirIdentity.dev++; + } + if (change === "inode") { + value.scope.homeIdentity.ino++; + } + if (change === "path") { + value.scope.projectRoot = "/foreign"; + } + if (change === "branch") { + value.scope.branch = "foreign"; + } + if (change === "schema") { + value.extra = true; + } + await writeFile(pool.file, JSON.stringify(value)); + } + await expect( + inspect({ scope: pool.scope, runtime, invoke }) + ).rejects.toThrow(); + } +}); + +test("native graph and current share must match the exact mapped run", async () => { + const pool = await fixture(); + for (const override of [ + { graph: { journal_incomplete: true, receipt: {} } }, + { graph: { journal_incomplete: false, receipt: { run: "f".repeat(32) } } }, + { status: { phase: "running", project_share: null } }, + ]) { + await expect( + inspect({ scope: pool.scope, runtime, invoke: authority(pool, override) }) + ).rejects.toThrow(); + } +}); + +test("stale selection, pending restart and foreign locks never change the mapping", async () => { + const pool = await fixture(); + const invoke = authority(pool); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: "f".repeat(64), + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + const lock = join(pool.directory, `${pool.key}.restart.lock.operation`); + await mkdir(lock); + await expect( + inspect({ scope: pool.scope, runtime, invoke }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); + await rm(lock, { recursive: true }); + await writeFile(join(pool.directory, `${pool.key}.restart.json`), "{}"); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); +}); + +test("replacement mapping inode during final authority check refuses", async () => { + const pool = await fixture(); + const invoke = authority(pool, { + onCall: async (call) => { + if (call === 7) { + await rename(pool.file, `${pool.file}.preserved`); + await writeFile(pool.file, pool.oldBytes, { mode: 0o600 }); + } + }, + }); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); +}); + +test("substituted lock paths are preserved and cannot authorize publication", async () => { + for (const suffix of [".restart.lock.operation", ".lock"]) { + const pool = await fixture(); + const lock = join(pool.directory, `${pool.key}${suffix}`); + const invoke = authority(pool, { + onCall: async (call) => { + if (call === 7) { + await rename(lock, `${lock}.preserved`); + await mkdir(lock, { mode: 0o700 }); + } + }, + }); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); + expect((await lstat(lock)).isDirectory()).toBe(true); + } +}); + +test("an exact published audit supports retry after prepublication refusal", async () => { + const pool = await fixture(); + const invoke = authority(pool, { + onCall: async (call) => { + if (call === 7) { + await writeFile(pool.file, `${pool.oldBytes} `); + } + }, + }); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + await writeFile(pool.file, pool.oldBytes); + const repaired = await recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke: authority(pool), + }); + expect(repaired.repaired).toBe(true); + expect(await load(pool.scope)).toEqual(run); +}); diff --git a/tests/native-project-start.test.ts b/tests/native-project-start.test.ts index b6afbe41f..fd440f61a 100644 --- a/tests/native-project-start.test.ts +++ b/tests/native-project-start.test.ts @@ -1,11 +1,23 @@ import { afterEach, expect, test } from "bun:test"; -import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; +import { + mkdir, + mkdtemp, + readFile, + realpath, + rm, + writeFile, +} from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; import type { acquireNativeHttpsLease } from "../src/backends/native-https-owner.ts"; -import { beginNativeProjectFinalization } from "../src/backends/native-project-finalization.ts"; +import { + beginNativeProjectFinalization, + captureNativeProjectFinalization, + waitNativeProjectFinalization, +} from "../src/backends/native-project-finalization.ts"; import type { NativeProjectInput } from "../src/backends/native-project-input.ts"; import { preflightNativeRestart } from "../src/backends/native-project-restart-preflight.ts"; +import { selectNativeRetainedImages } from "../src/backends/native-project-retained-images.ts"; import type { NativeProjectRun } from "../src/backends/native-project-run.ts"; import { parseNativeHttpsSelection, @@ -43,6 +55,7 @@ async function fixture(withEnvironment = true) { Parameters[0]["dependencies"] > = { prepareStorage: async () => {}, + retainedImages: async () => new Map(), load: async () => null, loadRestart: async () => null, save: async () => { @@ -192,6 +205,53 @@ function retainedPreflight(run: NativeProjectRun, phase = "running") { live_resources_verified: false, }; } +async function acknowledgeRetainedFrontend( + scope: Awaited>["opts"]["scope"], + run: NativeProjectRun +) { + const lifetime = await beginNativeProjectFinalization({ scope, run }); + await lifetime.complete(); + return lifetime; +} +async function retainedFinalizationFixture() { + const { opts, events } = await fixture(false); + const saved: NativeProjectRun = { + run: "1".repeat(32), + owner: "c".repeat(32), + namespace: "b".repeat(64), + planId: "a".repeat(64), + effectiveEnvName: null, + profiles: [], + aws: null, + }; + opts.dependencies.load = async () => saved; + const invoke = opts.dependencies.invoke!; + opts.dependencies.invoke = async (request) => { + if (request.args[1] === "retained-preflight") { + events.push("graph retained-preflight"); + return retainedPreflight(saved); + } + if ( + request.args[1] === "inspect" && + !events.includes("graph dependency-plan") + ) { + events.push("graph inspect"); + return stoppedGraph(saved); + } + return request.args[1] === "restore-selection" + ? { ...saved, plan: saved.planId, generation: "2".repeat(64) } + : await invoke(request); + }; + opts.dependencies.prepareStorage = async () => { + events.push("storage"); + }; + const prepare = opts.dependencies.prepare!; + opts.dependencies.prepare = async (request) => { + events.push("prepare"); + return await prepare(request); + }; + return { opts, events, saved }; +} test("source pool mismatch refuses before hooks and storage preparation", async () => { const { opts, events } = await fixture(); opts.dependencies.prepareStorage = async () => { @@ -251,6 +311,77 @@ test("uncertain foreground failure retains published mapping and cleans lifecycl expect(process.listenerCount("SIGINT")).toBe(int); expect(process.listenerCount("SIGTERM")).toBe(term); }); + +test("failed startup recovers exact pending cleanup and still reports the startup failure", async () => { + const { opts, events } = await fixture(); + const invoke = opts.dependencies.invoke!; + let selectedRun = ""; + let recovered = false; + const requests: string[] = []; + opts.dependencies.serve = async (request) => { + selectedRun = request.run; + throw new Error( + "Native graph startup failed (engine_rejected); inspect owned state before retrying." + ); + }; + opts.dependencies.invoke = async (request) => { + if (request.args[1] === "inspect-interrupted-start-cleanup") { + requests.push(request.args[1]); + return { + run: selectedRun, + phase: "cleanup-intent", + eligible: true, + same_boot: true, + data_retained: true, + selection_sha256: "1".repeat(64), + }; + } + if (request.args[1] === "recover-interrupted-start-cleanup") { + requests.push(request.args[1]); + recovered = true; + return { + run: selectedRun, + phase: "stopped-data-retained", + recovered: true, + data_retained: true, + same_boot: true, + publisher_retired: true, + reservation_released: true, + }; + } + if (request.args[1] === "inspect") { + requests.push(request.args[1]); + const value = stoppedGraph({ + run: selectedRun, + owner: "c".repeat(32), + namespace: "b".repeat(64), + planId: "a".repeat(64), + }); + if (!recovered) { + value.receipt.phase = "cleanup-intent"; + value.observations["container:web"].state = "present"; + return { + ...value, + receipt: { ...value.receipt, relay_cleanup: { phase: "pending" } }, + }; + } + return value; + } + return invoke(request); + }; + await expect(startNativeProject(opts)).rejects.toThrow( + "Native graph startup failed (engine_rejected)" + ); + expect(requests).toEqual([ + "inspect", + "inspect-interrupted-start-cleanup", + "recover-interrupted-start-cleanup", + "inspect", + ]); + expect(events).not.toContain("ready"); + expect(events).not.toContain("save"); + expect(events.at(-1)).toBe("cleanup"); +}); test("source opt-in and occupied mapping refuse before lifecycle or runtime effects", async () => { const { opts, events } = await fixture(); await expect( @@ -1411,6 +1542,101 @@ test("invalid profile selection refuses before lifecycle or runtime effects", as ).rejects.toThrow("profiles"); }); +test("retained startup refuses missing, unacknowledged or changed finalization before effects", async () => { + for (const fault of [ + "missing", + "unacknowledged", + "stale-completion", + "changed", + ] as const) { + const { opts, events, saved } = await retainedFinalizationFixture(); + if (fault === "unacknowledged") { + await beginNativeProjectFinalization({ scope: opts.scope, run: saved }); + } else if (fault !== "missing") { + await acknowledgeRetainedFrontend(opts.scope, saved); + if (fault === "stale-completion") { + await beginNativeProjectFinalization({ scope: opts.scope, run: saved }); + } else { + opts.dependencies.captureFinalization = async (selection) => { + const original = await captureNativeProjectFinalization(selection); + await acknowledgeRetainedFrontend(opts.scope, saved); + return original; + }; + } + } + opts.dependencies.https = async () => { + throw new Error("unexpected HTTPS owner effect"); + }; + const attempt = startNativeProject({ + ...opts, + https: httpsSelection, + adaptationFile: join(opts.scope.projectRoot, "unread-adaptation.json"), + }); + await expect(attempt).rejects.toThrow( + "frontend finalization is unconfirmed" + ); + try { + await attempt; + } catch (error) { + expect(String(error)).not.toContain(opts.scope.nativeHome); + } + expect(events).toEqual(["graph retained-preflight", "graph inspect"]); + } +}); + +test("retained startup admits only the acknowledged original frontend before effects", async () => { + const { opts, events, saved } = await retainedFinalizationFixture(); + const lifetime = await acknowledgeRetainedFrontend(opts.scope, saved); + opts.dependencies.waitFinalization = async (selection) => { + expect(selection.token).toEqual(lifetime.token); + expect(selection.run).toEqual(saved); + expect(selection.timeoutMs).toBe(1); + await waitNativeProjectFinalization(selection); + events.push("frontend-acknowledged"); + }; + expect(await startNativeProject(opts)).toBe(0); + expect(events.indexOf("frontend-acknowledged")).toBeLessThan( + events.indexOf("storage") + ); + expect(events.indexOf("frontend-acknowledged")).toBeLessThan( + events.indexOf("prepare") + ); + expect(events.indexOf("frontend-acknowledged")).toBeLessThan( + events.indexOf("before") + ); + expect(events.indexOf("frontend-acknowledged")).toBeLessThan( + events.indexOf("runtime up") + ); + expect(events).toContain("save"); +}); + +test("retained acknowledgement cannot authorize a newer or concurrently changed mapping", async () => { + for (const fault of ["newer-run", "changed-mapping"] as const) { + const { opts, events, saved } = await retainedFinalizationFixture(); + await acknowledgeRetainedFrontend(opts.scope, saved); + const newer = { ...saved, run: "2".repeat(32) }; + let loaded = 0; + opts.dependencies.load = async () => { + loaded++; + return fault === "newer-run" || loaded > 1 ? newer : saved; + }; + if (fault === "newer-run") { + opts.dependencies.invoke = async (request) => { + events.push(request.args.slice(0, 2).join(" ")); + return request.args[1] === "retained-preflight" + ? retainedPreflight(newer) + : stoppedGraph(newer); + }; + } + await expect(startNativeProject(opts)).rejects.toThrow( + fault === "newer-run" + ? "frontend finalization is unconfirmed" + : "mapping changed" + ); + expect(events).toEqual(["graph retained-preflight", "graph inspect"]); + } +}); + test("ordinary up reuses a stopped owned run and publishes its mapping by comparison", async () => { const { opts, events } = await fixture(); const saved: NativeProjectRun = { @@ -1422,6 +1648,7 @@ test("ordinary up reuses a stopped owned run and publishes its mapping by compar profiles: [], aws: null, }; + await acknowledgeRetainedFrontend(opts.scope, saved); opts.dependencies.load = async () => saved; const invoke = opts.dependencies.invoke!; opts.dependencies.invoke = async (request) => { @@ -1485,6 +1712,11 @@ test("restore startup passes the retained run and selected generation to its own namespace: "b".repeat(64), planId: "a".repeat(64), }; + const review = opts.dependencies.review!; + opts.dependencies.review = async (request) => { + expect(request.retained).toEqual(saved); + return await review(request); + }; const invoke = opts.dependencies.invoke!; const serve = opts.dependencies.serve!; opts.dependencies.invoke = async (request) => @@ -1520,6 +1752,7 @@ test("ordinary up resumes a stopped owned VM once and requires live proof before profiles: [], aws: null, }; + await acknowledgeRetainedFrontend(opts.scope, saved); const controller = new AbortController(); let up = 0, served = 0, @@ -1937,13 +2170,21 @@ test("startup and restart review use identical branch routes after adaptation", }; return { ...input, normalizedComposeJson: JSON.stringify(compose) }; }; + const before = async (input: NativeProjectInput) => { + expect( + JSON.parse(input.normalizedComposeJson).services.web.labels.caddy + ).toBe("feature-a.app.hack.local, feature-a.app.hack.gy"); + return await opts.before(); + }; const observed: NativeProjectInput[] = []; opts.dependencies.review = async (request) => { expect(request.branch).toBe("feature-a"); observed.push(request.input); return await review(request); }; - expect(await startNativeProject({ ...opts, scope, adaptationFile })).toBe(0); + expect( + await startNativeProject({ ...opts, scope, adaptationFile, before }) + ).toBe(0); await preflightNativeRestart({ runtime: opts.runtime, scope, @@ -1963,6 +2204,12 @@ test("startup and restart review use identical branch routes after adaptation", review: opts.dependencies.review, dependencies: async () => [], invoke: async ({ args }) => { + if (args[0] === "project") { + expect(args).toContain("--branch"); + return { + plan: { namespace: "b".repeat(64), compose_sha256: "d".repeat(64) }, + }; + } if (args[1] === "status") { return { network: "internet" }; } @@ -1982,3 +2229,139 @@ test("startup and restart review use identical branch routes after adaptation", .caddy ).toBe("feature-a.app.hack.local, feature-a.app.hack.gy"); }); + +test("retained startup selects legacy labels only after runtime admission and before review", async () => { + for (const changed of [false, true]) { + const { opts } = await fixture(false); + const saved = { + run: "1".repeat(32), + owner: "c".repeat(32), + namespace: "b".repeat(64), + planId: "a".repeat(64), + }; + const prepare = opts.dependencies.prepare!; + opts.dependencies.prepare = async (request) => { + const input = await prepare(request); + const compose = JSON.parse(input.normalizedComposeJson); + compose.services.web.labels = { + caddy: "app.hack.local", + "caddy.tls": "internal", + "caddy.reverse_proxy": "{{upstreams 3000}}", + }; + return { ...input, normalizedComposeJson: JSON.stringify(compose) }; + }; + const invoke = opts.dependencies.invoke!; + let admitted = false; + let reviewed = false; + opts.dependencies.invoke = async (request) => { + if (request.args[0] === "runtime" && request.args[1] === "up") { + admitted = true; + } + if (request.args[0] === "project") { + expect(admitted).toBe(true); + return { + plan: { + namespace: request.args.includes("--branch") + ? "d".repeat(64) + : saved.namespace, + compose_sha256: "d".repeat(64), + source: await realpath(opts.scope.projectRoot), + }, + }; + } + if (request.args[1] === "restore-selection") { + expect(admitted).toBe(true); + return { + ...saved, + owner: changed ? "0".repeat(32) : saved.owner, + plan: saved.planId, + generation: "2".repeat(64), + }; + } + return await invoke(request); + }; + const review = opts.dependencies.review!; + opts.dependencies.review = async (request) => { + reviewed = true; + expect(request.reviewIdentity?.branch).toBeNull(); + expect( + JSON.parse(request.input.normalizedComposeJson).services.web.labels + .caddy + ).toBe("app.hack.local"); + return await review(request); + }; + const running = startNativeProject({ + ...opts, + scope: { ...opts.scope, branch: "feature-a" }, + restore: saved, + }); + if (changed) { + await expect(running).rejects.toThrow("restore selection changed"); + expect(reviewed).toBe(false); + } else { + expect(await running).toBe(0); + expect(reviewed).toBe(true); + } + } +}); + +test("retained startup reviews the saved content ID instead of a newly resolved mutable tag", async () => { + const { opts, events } = await fixture(false); + const saved = { + run: "1".repeat(32), + owner: "c".repeat(32), + namespace: "b".repeat(64), + planId: "a".repeat(64), + }; + const image = `sha256:${"e".repeat(64)}`; + const stopped = stoppedGraph(saved); + const observed = { + ...stopped, + receipt: { + ...stopped.receipt, + normalized_input: { + namespace: saved.namespace, + original_compose_sha256: "d".repeat(64), + normalized_compose_sha256: "f".repeat(64), + }, + resources: { + "container:web": { + ...stopped.receipt.resources["container:web"], + image, + }, + }, + }, + }; + opts.dependencies.retainedImages = selectNativeRetainedImages; + const prepare = opts.dependencies.prepare!; + opts.dependencies.prepare = async (request) => ({ + ...(await prepare(request)), + normalizedComposeJson: JSON.stringify({ + services: { web: { image: "redis:latest" } }, + }), + }); + const invoke = opts.dependencies.invoke!; + let beforeReview = true; + opts.dependencies.invoke = async (request) => { + if (request.args[1] === "ensure-image") { + throw new Error("mutable tag was unexpectedly resolved"); + } + if (request.args[1] === "inspect" && beforeReview) { + expect(events).toContain("runtime up"); + return observed; + } + if (request.args[1] === "restore-selection") { + return { ...saved, plan: saved.planId, generation: "2".repeat(64) }; + } + return await invoke(request); + }; + const review = opts.dependencies.review!; + opts.dependencies.review = async (request) => { + beforeReview = false; + expect( + JSON.parse(request.input.normalizedComposeJson).services.web.image + ).toBe(image); + return await review(request); + }; + expect(await startNativeProject({ ...opts, restore: saved })).toBe(0); +}); diff --git a/tests/native-runtime-client.test.ts b/tests/native-runtime-client.test.ts index dc827520a..6df6a9aaa 100644 --- a/tests/native-runtime-client.test.ts +++ b/tests/native-runtime-client.test.ts @@ -26,6 +26,8 @@ if (process.argv.includes("hang")) await Bun.sleep(60_000); if (process.argv.includes("structured")) { console.error(JSON.stringify({code:"source_conflict",message:"synthetic-secret-diagnostic"})); process.exit(23); } if (process.argv.includes("one-off-fail")) { console.error(JSON.stringify({code:"graph_one_off_failed",cause_code:"graph_one_off_cancelled",message:"synthetic-secret-diagnostic"})); process.exit(23); } if (process.argv.includes("one-off-unsafe")) { console.error(JSON.stringify({code:"graph_one_off_failed",cause_code:"synthetic-secret-diagnostic",message:"synthetic-secret-diagnostic"})); process.exit(23); } +if (process.argv.includes("owner-fail")) { const {appendFileSync}=await import("node:fs"); appendFileSync("attempts","attempt"); console.error(JSON.stringify({code:"graph_owner_recovery",cause_code:"graph_relay_identity",message:"synthetic-secret-diagnostic"})); process.exit(23); } +if (process.argv.includes("owner-unsafe")) { console.error(JSON.stringify({code:"graph_owner_recovery",cause_code:"synthetic-secret-diagnostic",message:"synthetic-secret-diagnostic"})); process.exit(23); } if (process.argv.includes("fail")) { console.error("synthetic-secret-diagnostic"); process.exit(23); } if (process.argv.includes("overflow")) { process.stdout.write("x".repeat(17 * 1024 * 1024)); } else { @@ -207,6 +209,35 @@ test("one-off failures expose only a validated owner cause code", async () => { ); }); +test("owner cleanup failures preserve a bounded cause without replay or raw diagnostics", async () => { + const runtime = await fixture(); + const error = await invokeNativeRuntime({ + runtime, + cwd: runtime.home, + args: ["owner-fail"], + }).catch((failure: unknown) => failure); + expect(error).toBeInstanceOf(NativeRuntimeRequestError); + expect(error).toMatchObject({ + nativeCode: "graph_owner_recovery", + nativeCauseCode: "graph_relay_identity", + }); + expect(String(error)).toContain("graph_owner_recovery: graph_relay_identity"); + expect(String(error)).not.toContain("synthetic-secret-diagnostic"); + expect(await Bun.file(join(runtime.home, "attempts")).text()).toBe("attempt"); + + const unsafe = await invokeNativeRuntime({ + runtime, + cwd: runtime.home, + args: ["owner-unsafe"], + }).catch((failure: unknown) => failure); + expect(unsafe).toBeInstanceOf(NativeRuntimeRequestError); + expect(unsafe).toMatchObject({ + nativeCode: "graph_owner_recovery", + nativeCauseCode: undefined, + }); + expect(String(unsafe)).not.toContain("synthetic-secret-diagnostic"); +}); + test("early structured rejection survives a full private-input pipe without replay", async () => { const runtime = await fixture(); const marker = join(runtime.home, "attempts");