diff --git a/docs/guides/native-candidate.md b/docs/guides/native-candidate.md index 7a31e89b0..2d21f9d08 100644 --- a/docs/guides/native-candidate.md +++ b/docs/guides/native-candidate.md @@ -118,6 +118,77 @@ complete bundle to change the candidate. binary refusing newer state is an unsupported downgrade, not a successful rollback. Preserve that state; do not rewrite receipts or adopt another home's data. +Host cleanup or a reboot can remove the temporary `/private/tmp/hkl-` HOME +alias while the private candidate home and disks remain intact. Normal commands +report `provider_home_missing`; they do not initialize another pool or recreate the +alias during observation. Explicit `runtime recover --json` can restore only that +absent exact alias, after verifying the receipt-bound dead provider, both recorded +disks, free VM lock, closed disk handles, and no active provider command. Existing +files, directories, foreign links, a live or reused PID, and uncertain ownership +refuse without replacement. Recovery rechecks the receipt before exclusive creation, +then validates socket absence and flushes the disks before recording +`recovered-unclean`. If those final checks fail, it reports +`provider_home_restored_recovery_incomplete`, keeps the exact owned alias and retained +data, and requires inspection before retry. No guest is started by alias recovery. +Interrupted starts without a recorded process or both identified disks remain +outside this repair path. + +A physical macOS reboot may also renumber the mounted filesystem device. Strict +disk and source checks still refuse a changed device number; missing-HOME recovery +does not waive them. For an offline **stock pool**, inspect the separate migration: + +```sh +./hack-native --candidate-root /absolute/private/candidate-home runtime host-filesystem-recovery --json +./hack-native --candidate-root /absolute/private/candidate-home runtime recover-host-filesystem --expect-sha256 --accept-legacy-device-rebind --json +./hack-native --candidate-root /absolute/private/candidate-home runtime recover --json +``` + +Review the inspection before selecting its hash. This explicit legacy migration +requires an absent recorded provider whose start predates the current host boot, +no active provider commands, exclusive existing operation/VM locks and closed disk +handles. Both disk inodes, sizes and ext4 UUIDs, the exact source path/inode and its +ownership must be unchanged; only one common old-to-new device-number change is +allowed. The inspection is read-only. Publication atomically changes only the +owner's disk and source device numbers, retaining its phase and process record. +Normal commands keep their strict identity checks. + +Legacy receipts have no original host boot UUID or filesystem volume UUID. Calendar +timestamps corroborate a reboot, and matching retained file identities constrain +the migration, but neither proves original volume continuity. The opt-in explicitly +accepts that limitation; copied or relocated pools are outside this procedure. +Prepared-base pools and pending owner/network/activation updates require separate +recovery and are refused. A torn owner publication preserves `owner.pending` and +blocks another migration; do not delete or adopt that file manually. + +This does not start the VM, restore HOME, retire stale sockets or rewrite historical +graph receipts. Use ordinary recovery afterward. Historical shared-source graphs +retain the prior identity and may require verified cleanup plus a new generation; +successful metadata migration alone does not establish application recovery. + +The frontend project-run mapping also records directory device numbers. If native +inspection succeeds after provider recovery but ordinary project commands refuse +the mapping, inspect the exact instance with the current bundle: + +```sh +./hack-v5 doctor --path /absolute/project --branch my-instance --native-run-mapping inspect --json +./hack-v5 doctor --path /absolute/project --branch my-instance --native-run-mapping repair --expect-selection --accept-legacy-device-rebind --json +``` + +Omitting `--branch` uses the same linked-worktree default as project commands; +detached linked worktrees require an explicit instance. Inspection writes nothing. +Repair requires the exact selection and explicit acceptance of unproven original +filesystem volume continuity. It changes only the three scope directory device +numbers, preserving canonical paths, inodes, branch, run, owner, plan, environment +selection, profiles and AWS selector. The selected graph and current project share +must still match native authority. An audit copy retains the original private +mapping. Held locks, pending restart state, substituted directories, stale hashes +and changed native identities refuse publication. + +This is a metadata repair, not an app restart: it does not retire sockets, modify +graph history, remove data or establish browser readiness. Doctor's ordinary +`--fix` never applies it. Run the normal project command separately after reviewing +the result; its existing native graph and ownership checks still apply. + Compatibility is qualified for specific bundle hashes and state formats. The displayed version alone does not establish frontend/executor or downgrade compatibility. V4 and candidate homes remain separate; this procedure does not diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 164130b76..ccf439c3f 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -1263,6 +1263,10 @@ hack doctor [options] | `--json` | Output JSON (machine-readable) | | `--browser-url https://app.hack` | HTTPS origin manually tested in the browser (no path or credentials) | | `--browser-result unknown|works|fails|permission-denied` | Your manual browser observation for --browser-url (default: unknown) | +| `--branch ` | Run against a branch-specific instance (compose name + hostnames) | +| `--native-run-mapping inspect|repair` | Inspect or explicitly repair a native run mapping after filesystem device renumbering | +| `--expect-selection <64-hex>` | Require the exact run-mapping recovery inspection selection | +| `--accept-legacy-device-rebind` | Explicitly accept legacy migration without proof of original filesystem volume continuity | | `--no-interactive` | Never prompt: apply documented defaults or fail with E_INTERACTIVE_REQUIRED (also via HACK_NO_INTERACTIVE=1) | | `--help, -h` | Show help | | `--version, -v` | Show version | diff --git a/packages/runtime-core/src/main.rs b/packages/runtime-core/src/main.rs index 68bdc05ed..be6ee57a7 100644 --- a/packages/runtime-core/src/main.rs +++ b/packages/runtime-core/src/main.rs @@ -86,6 +86,8 @@ Usage: hack-local runtime up --profile development --internet --json hack-local runtime network extend --allow-host [--allow-host ] --json hack-local runtime up|status|down|recover [--json] + hack-local runtime host-filesystem-recovery [--json] + hack-local runtime recover-host-filesystem --expect-sha256 --accept-legacy-device-rebind [--json] hack-local node serve|status|inspect hack-local node request hack-local --version @@ -857,6 +859,32 @@ fn run() -> Result<(), CandidateError> { image, )?)?; } + ["runtime", "host-filesystem-recovery"] + | ["runtime", "host-filesystem-recovery", "--json"] => { + print_json(&hack_runtime_core::provider::host_filesystem::inspect( + &discover_candidate(&requested)?, + )?)?; + } + [ + "runtime", + "recover-host-filesystem", + "--expect-sha256", + hash, + "--accept-legacy-device-rebind", + ] + | [ + "runtime", + "recover-host-filesystem", + "--expect-sha256", + hash, + "--accept-legacy-device-rebind", + "--json", + ] => { + print_json(&hack_runtime_core::provider::host_filesystem::recover( + &discover_candidate(&requested)?, + hash, + )?)?; + } ["runtime", action] | ["runtime", action, "--json"] => { let candidate = discover_candidate(&requested)?; let result = match *action { diff --git a/packages/runtime-core/src/provider/lifecycle.rs b/packages/runtime-core/src/provider/lifecycle.rs index 60e7109d3..36decce95 100644 --- a/packages/runtime-core/src/provider/lifecycle.rs +++ b/packages/runtime-core/src/provider/lifecycle.rs @@ -1,8 +1,10 @@ +pub mod host_filesystem; mod interrupted; mod prepared_boot; #[cfg(any(target_os = "macos", test))] mod private_child; mod relay_process; +mod short_home; use super::{ admission, agent, artifact, identity, process, state::{self, Owner, io}, @@ -1812,6 +1814,12 @@ fn finish_absent( value: &str, record_disks: bool, ) -> Result<(), CandidateError> { + let vm_lock = lock_absent_disks(candidate, owner)?; + finish_absent_locked(candidate, owner, value, record_disks, &vm_lock) +} + +/// Keep this descriptor through alias restoration and the stopped receipt. +fn lock_absent_disks(candidate: &Candidate, owner: &Owner) -> Result { let directory = owner.real_data_dir(candidate)?; let vm_lock = OpenOptions::new() .read(true) @@ -1847,6 +1855,16 @@ fn finish_absent( )); } verify_disks(candidate, owner)?; + Ok(vm_lock) +} + +fn finish_absent_locked( + candidate: &Candidate, + owner: &mut Owner, + value: &str, + record_disks: bool, + _vm_lock: &File, +) -> Result<(), CandidateError> { if record_disks && owner.storage.is_none() { owner.storage = Some(identity::disk( &owner.real_data_dir(candidate)?.join("storage.raw"), @@ -1896,7 +1914,13 @@ fn finish_absent( } pub fn recover(candidate: &Candidate) -> Result { - let initial = status(candidate)?; + let initial = match status(candidate) { + Ok(initial) => initial, + Err(error) if error.code == "provider_home_missing" => { + return short_home::recover(candidate); + } + Err(error) => return Err(error), + }; if initial.phase == "uninitialized" { return Ok(initial); } diff --git a/packages/runtime-core/src/provider/lifecycle/host_filesystem.rs b/packages/runtime-core/src/provider/lifecycle/host_filesystem.rs new file mode 100644 index 000000000..6c9e10634 --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/host_filesystem.rs @@ -0,0 +1,297 @@ +//! Explicit legacy device-number migration for an offline stock pool. +//! +//! Old receipts have neither a host boot UUID nor a filesystem volume UUID. Calendar start +//! time, unchanged inode/size/ext4 UUID and exact paths constrain this migration, but cannot +//! establish original volume continuity. Only an explicitly accepted, hash-selected inspection +//! may update the owner. Normal runtime operations retain their strict identity comparisons. +use super::{Owner, binary, identity, lock_absent_disks, root, state}; +use crate::{Candidate, CandidateError}; +use serde::Serialize; +use sha2::{Digest, Sha256}; +use std::{fs, os::unix::fs::MetadataExt, path::Path}; + +#[derive(Debug, Serialize, PartialEq, Eq)] +pub struct Inspection { + pub schema: &'static str, + pub selection_sha256: String, + pub owner_sha256: String, + pub machine: String, + pub host_boot_micros: u64, + pub provider_start_micros: u64, + pub old_device: u64, + pub new_device: u64, + pub pool_inode: u64, + pub storage: identity::DiskIdentity, + pub overlay: identity::DiskIdentity, + pub project_share: Option, + pub qualification: &'static str, +} + +fn refused(detail: &str) -> CandidateError { + CandidateError::new( + "host_filesystem_recovery", + format!("{detail}; no identity was changed."), + ) +} + +#[cfg(target_os = "macos")] +fn host_boot_micros() -> Result { + let mut time = std::mem::MaybeUninit::::zeroed(); + let mut length = std::mem::size_of::(); + // SAFETY: read-only sysctl writes an exactly sized timeval; no input or retained pointers. + if unsafe { + libc::sysctlbyname( + c"kern.boottime".as_ptr(), + time.as_mut_ptr().cast(), + &mut length, + std::ptr::null_mut(), + 0, + ) + } != 0 + || length != std::mem::size_of::() + { + return Err(refused("Native host boot time is unavailable")); + } + // SAFETY: a complete timeval was returned above. + let time = unsafe { time.assume_init() }; + let seconds = u64::try_from(time.tv_sec).map_err(|_| refused("Invalid host boot time"))?; + let micros = u64::try_from(time.tv_usec).map_err(|_| refused("Invalid host boot time"))?; + if micros >= 1_000_000 { + return Err(refused("Invalid host boot time")); + } + seconds + .checked_mul(1_000_000) + .and_then(|v| v.checked_add(micros)) + .filter(|v| *v > 0) + .ok_or_else(|| refused("Invalid host boot time")) +} + +#[cfg(not(target_os = "macos"))] +fn host_boot_micros() -> Result { + Err(CandidateError::new( + "unsupported_host", + "Legacy filesystem recovery requires macOS.", + )) +} + +fn no_auxiliary_update(candidate: &Candidate) -> Result<(), CandidateError> { + crate::provider::network_update::require_complete(candidate)?; + for name in [ + "owner.pending", + "prepared-base.json", + "prepared-base.json.pending", + ] { + match fs::symlink_metadata(root(candidate).join(name)) { + Err(e) if e.kind() == std::io::ErrorKind::NotFound => {} + _ => { + return Err(refused( + "Pending owner or prepared-base state requires separate recovery", + )); + } + } + } + Ok(()) +} + +fn dead_provider(candidate: &Candidate, owner: &Owner, boot: u64) -> Result<(), CandidateError> { + let process = owner + .process + .as_ref() + .ok_or_else(|| refused("No recorded provider identity"))?; + // SAFETY: geteuid has no preconditions. + identity::verify(process, process, &binary(candidate), unsafe { + libc::geteuid() + })?; + if !owner.created + || process.start_micros >= boot + || identity::alive(process.pid)? + || identity::executable_running(&binary(candidate))? + { + return Err(refused( + "Provider must be absent and its recorded start must predate this host boot", + )); + } + Ok(()) +} + +fn same_disk(before: &identity::DiskIdentity, after: &identity::DiskIdentity) -> bool { + before.device != after.device + && before.inode == after.inode + && before.bytes == after.bytes + && before.uuid == after.uuid +} + +/// A retained flock cannot fence a writer that opens a substituted lock pathname. +fn bound_locks( + candidate: &Candidate, + owner: &Owner, + operation: &state::Lock, + vm: &fs::File, +) -> Result<(), CandidateError> { + let check = |path: &Path, expected: (u64, u64)| -> Result<(), CandidateError> { + let metadata = fs::symlink_metadata(path).map_err(state::io)?; + if !metadata.is_file() + || metadata.nlink() != 1 + || (metadata.dev(), metadata.ino()) != expected + { + return Err(refused("Held lock pathname was replaced")); + } + Ok(()) + }; + check( + &root(candidate).join("operation.lock"), + operation.identity()?, + )?; + let metadata = vm.metadata().map_err(state::io)?; + check( + &owner.real_data_dir(candidate)?.join("vm.lock"), + (metadata.dev(), metadata.ino()), + ) +} + +/// Build a candidate owner in memory. This function neither adopts a new disk nor writes state. +fn selection( + candidate: &Candidate, + owner: &Owner, + boot: u64, +) -> Result<(Inspection, Owner), CandidateError> { + no_auxiliary_update(candidate)?; + dead_provider(candidate, owner, boot)?; + let data = owner.real_data_dir(candidate)?; + let storage = identity::disk(&data.join("storage.raw"))?; + let overlay = identity::disk(&data.join("overlay.raw"))?; + let before = owner + .storage + .as_ref() + .ok_or_else(|| refused("Storage identity is missing"))?; + let prior_overlay = owner + .overlay + .as_ref() + .ok_or_else(|| refused("Overlay identity is missing"))?; + let metadata = fs::symlink_metadata(root(candidate)).map_err(state::io)?; + if !same_disk(before, &storage) + || !same_disk(prior_overlay, &overlay) + || before.device != prior_overlay.device + || storage.device != overlay.device + || storage.device != metadata.dev() + { + return Err(refused( + "Only a common device-number change with unchanged disk inode, size and UUID is accepted", + )); + } + let mut next = owner.clone(); + if let Some(share) = &owner.project_share { + let observed = + crate::provider::ProjectShareIntent::approve(&share.project, share.unfiltered_source)?; + let mut expected = share.clone(); + expected.device = storage.device; + if share.device != before.device || observed != expected { + return Err(refused( + "Project share changed beyond the same filesystem device number", + )); + } + next.project_share = Some(observed); + } + next.storage = Some(storage.clone()); + next.overlay = Some(overlay.clone()); + // The canonical read is bounded and validated independently of raw-byte hashing. + let bytes = crate::provider::prepared_base::read_private( + &root(candidate).join("owner.json"), + 1024 * 1024, + )?; + if Owner::load_for_short_home_recovery(candidate)? != *owner { + return Err(refused("Owner selection changed")); + } + let mut inspection = Inspection { + schema: "hack.host-filesystem-recovery/v1", + selection_sha256: String::new(), + owner_sha256: format!("{:x}", Sha256::digest(bytes)), + machine: owner.machine.clone(), + host_boot_micros: boot, + provider_start_micros: owner + .process + .as_ref() + .ok_or_else(|| refused("No process"))? + .start_micros, + old_device: before.device, + new_device: storage.device, + pool_inode: metadata.ino(), + storage, + overlay, + project_share: next.project_share.clone(), + qualification: "explicit-legacy-migration-original-volume-continuity-unproven", + }; + inspection.selection_sha256 = format!( + "{:x}", + Sha256::digest( + serde_json::to_vec(&inspection).map_err(|_| refused("Cannot encode selection"))? + ) + ); + Ok((inspection, next)) +} + +/// Read-only inspection holds both existing locks and proves disks have no open handles. +pub fn inspect(candidate: &Candidate) -> Result { + let operation = state::Lock::acquire_existing(&root(candidate))?; + let owner = Owner::load_for_short_home_recovery(candidate)?; + let (inspection, next) = selection(candidate, &owner, host_boot_micros()?)?; + let vm = lock_absent_disks(candidate, &next)?; + if selection(candidate, &owner, host_boot_micros()?)?.0 != inspection { + return Err(refused("Inspection changed during absence verification")); + } + bound_locks(candidate, &owner, &operation, &vm)?; + Ok(inspection) +} + +/// Publish disk and source device changes in ONE atomic owner replacement. A crash leaves +/// old or new committed metadata; any unfinished owner.pending is preserved and blocks retry. +/// This does not launch, restore the alias, retire historical sockets, or rewrite graph receipts. +pub fn recover(candidate: &Candidate, expected: &str) -> Result { + recover_with_boot(candidate, expected, host_boot_micros) +} + +fn recover_with_boot( + candidate: &Candidate, + expected: &str, + boot: impl Fn() -> Result, +) -> Result { + if expected.len() != 64 || !expected.bytes().all(|b| b.is_ascii_hexdigit()) { + return Err(refused("An exact inspection SHA-256 is required")); + } + let operation = state::Lock::acquire_existing(&root(candidate))?; + let owner = Owner::load_for_short_home_recovery(candidate)?; + let (inspection, next) = selection(candidate, &owner, boot()?)?; + if inspection.selection_sha256 != expected { + return Err(refused("Inspection selection is stale")); + } + let vm = lock_absent_disks(candidate, &next)?; + if selection(candidate, &owner, boot()?)?.0 != inspection { + return Err(refused("Selected identities changed before publication")); + } + bound_locks(candidate, &owner, &operation, &vm)?; + next.save(candidate)?; + if Owner::load_for_short_home_recovery(candidate)? != next { + return Err(CandidateError::new( + "host_filesystem_recovery_incomplete", + "Owner was published but changed during confirmation; retained state requires inspection.", + )); + } + super::verify_disks(candidate, &next).map_err(|e| { + CandidateError::new( + "host_filesystem_recovery_incomplete", + format!( + "Owner was published but disk confirmation failed ({}); inspect retained state.", + e.code + ), + ) + })?; + if let Some(share) = &next.project_share { + share.validate().map_err(|e| CandidateError::new( + "host_filesystem_recovery_incomplete", format!("Owner was published but source confirmation failed ({}); inspect retained state.", e.code) + ))?; + } + Ok(inspection) +} + +#[cfg(all(test, target_os = "macos"))] +mod tests; diff --git a/packages/runtime-core/src/provider/lifecycle/host_filesystem/tests.rs b/packages/runtime-core/src/provider/lifecycle/host_filesystem/tests.rs new file mode 100644 index 000000000..4c6cc502f --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/host_filesystem/tests.rs @@ -0,0 +1,319 @@ +use super::*; +use crate::provider::{NetworkIntent, Profile, ProjectShareIntent}; +use std::{fs::File, os::fd::AsRawFd, path::PathBuf, process::Command}; + +struct Pool { + candidate: Candidate, + owner: Owner, + directory: PathBuf, + boot: u64, +} +impl Pool { + fn new() -> Self { + let mut child = Command::new("/bin/sleep").arg("30").spawn().unwrap(); + let mut process = identity::observe(child.id() as i32).unwrap(); + child.kill().unwrap(); + child.wait().unwrap(); + let boot = process.start_micros + 1; + let directory = fs::canonicalize(std::env::temp_dir()) + .unwrap() + .join(format!( + "hack-device-recovery-{}-{}", + std::process::id(), + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .unwrap() + .as_nanos() + )); + state::private_directory(&directory).unwrap(); + let candidate = Candidate::discover(&directory).unwrap(); + let operation = state::Lock::acquire(&root(&candidate)).unwrap(); + let project = directory.join("app"); + state::private_directory(&project).unwrap(); + fs::write(project.join("package.json"), b"{}").unwrap(); + let share = ProjectShareIntent::approve(&project, true).unwrap(); + let mut owner = Owner::create_with_project_share( + &candidate, + Profile::Development, + None, + NetworkIntent::Isolated, + None, + Some(share), + ) + .unwrap(); + let data = owner.real_data_dir(&candidate).unwrap(); + state::private_directory(&data).unwrap(); + for (name, tag) in [("storage.raw", 1u8), ("overlay.raw", 2)] { + let mut bytes = vec![0; 4096]; + bytes[1080..1082].copy_from_slice(&[0x53, 0xef]); + bytes[1128..1144].fill(tag); + bytes[2048..2059].copy_from_slice(b"data-marker"); + fs::write(data.join(name), bytes).unwrap(); + } + fs::write(data.join("name"), &owner.machine).unwrap(); + fs::write(data.join("vm.lock"), b"").unwrap(); + process.executable = binary(&candidate); + owner.process = Some(process); + owner.created = true; + owner.phase = "running".into(); + let mut storage = identity::disk(&data.join("storage.raw")).unwrap(); + let mut overlay = identity::disk(&data.join("overlay.raw")).unwrap(); + let old = storage.device + 1; + storage.device = old; + overlay.device = old; + owner.storage = Some(storage); + owner.overlay = Some(overlay); + owner.project_share.as_mut().unwrap().device = old; + owner.save(&candidate).unwrap(); + fs::remove_file(&owner.short_home).unwrap(); + drop(operation); + Self { + candidate, + owner, + directory, + boot, + } + } + fn data(&self) -> PathBuf { + self.owner.real_data_dir(&self.candidate).unwrap() + } + fn selected(&self) -> Inspection { + selection(&self.candidate, &self.owner, self.boot) + .unwrap() + .0 + } + fn receipt(&self) -> Vec { + fs::read(root(&self.candidate).join("owner.json")).unwrap() + } + fn apply(&self, hash: &str) -> Result { + recover_with_boot(&self.candidate, hash, || Ok(self.boot)) + } + fn unchanged(&self, before: &[u8]) { + assert_eq!(self.receipt(), before); + assert!(self.owner.short_home.symlink_metadata().is_err()); + } +} +impl Drop for Pool { + fn drop(&mut self) { + let _ = fs::remove_file(&self.owner.short_home); + fs::remove_dir_all(&self.directory).unwrap(); + } +} + +#[test] +fn migration_changes_only_devices_and_keeps_alias_graph_history_and_data() { + let pool = Pool::new(); + let old = pool.receipt(); + let storage = fs::read(pool.data().join("storage.raw")).unwrap(); + let overlay = fs::read(pool.data().join("overlay.raw")).unwrap(); + let history = pool.directory.join("historical-graph.json"); + fs::write(&history, b"historical-source-binding").unwrap(); + let inspected = pool.selected(); + pool.unchanged(&old); + let result = pool.apply(&inspected.selection_sha256).unwrap(); + assert_eq!(result, inspected); + let mut expected = pool.owner.clone(); + expected.storage.as_mut().unwrap().device = inspected.new_device; + expected.overlay.as_mut().unwrap().device = inspected.new_device; + expected.project_share.as_mut().unwrap().device = inspected.new_device; + assert_eq!( + Owner::load_for_short_home_recovery(&pool.candidate).unwrap(), + expected + ); + assert!(pool.owner.short_home.symlink_metadata().is_err()); + assert_eq!(fs::read(history).unwrap(), b"historical-source-binding"); + assert_eq!(fs::read(pool.data().join("storage.raw")).unwrap(), storage); + assert_eq!(fs::read(pool.data().join("overlay.raw")).unwrap(), overlay); + assert!(pool.apply(&inspected.selection_sha256).is_err()); + // Ordinary recovery still owns HOME restoration and the recovered phase. + let recovered = crate::provider::recover(&pool.candidate).unwrap(); + assert_eq!(recovered.phase, "recovered-unclean"); + assert_eq!(recovered.process_alive, Some(false)); + assert_eq!(fs::read(pool.data().join("storage.raw")).unwrap(), storage); +} + +#[test] +fn disk_and_share_changes_beyond_common_device_refuse_without_writes() { + for case in 0..7 { + let mut pool = Pool::new(); + let observed_device = identity::disk(&pool.data().join("storage.raw")) + .unwrap() + .device; + let disk = pool.owner.storage.as_mut().unwrap(); + match case { + 0 => disk.inode += 1, + 1 => disk.bytes += 1, + 2 => disk.uuid.push('0'), + 3 => disk.device = observed_device, + 4 => pool.owner.overlay.as_mut().unwrap().device += 1, + 5 => pool.owner.project_share.as_mut().unwrap().inode += 1, + _ => pool.owner.project_share.as_mut().unwrap().device += 1, + } + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert!(pool.apply(&"a".repeat(64)).is_err()); + pool.unchanged(&before); + } +} + +#[test] +fn stale_owner_or_host_boot_selection_refuses() { + let mut pool = Pool::new(); + let selection = pool.selected(); + assert!(pool.apply(&"a".repeat(64)).is_err()); + pool.owner.phase = "stopped".into(); + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert!(pool.apply(&selection.selection_sha256).is_err()); + pool.unchanged(&before); + let selected = pool.selected(); + assert!( + recover_with_boot(&pool.candidate, &selected.selection_sha256, || Ok(pool + .boot + + 1)) + .is_err() + ); + pool.unchanged(&before); +} + +#[test] +fn same_boot_live_reused_pid_and_unknown_identity_refuse() { + for case in 0..4 { + let mut pool = Pool::new(); + if case == 0 { + pool.boot = pool.owner.process.as_ref().unwrap().start_micros; + } + if case == 1 || case == 2 { + let p = pool.owner.process.as_mut().unwrap(); + p.pid = std::process::id() as i32; + if case == 1 { + p.start_micros = identity::observe(p.pid).unwrap().start_micros; + } + } + if case == 3 { + pool.owner.process = None; + } + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert!(pool.apply(&"a".repeat(64)).is_err()); + pool.unchanged(&before); + } +} + +#[test] +fn held_operation_vm_lock_and_disk_handles_refuse() { + for case in 0..3 { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let _operation = + (case == 0).then(|| state::Lock::acquire_existing(&root(&pool.candidate)).unwrap()); + let file = if case == 1 { + Some(File::open(pool.data().join("vm.lock")).unwrap()) + } else if case == 2 { + Some(File::open(pool.data().join("storage.raw")).unwrap()) + } else { + None + }; + if case == 1 { + assert_eq!( + unsafe { + libc::flock( + file.as_ref().unwrap().as_raw_fd(), + libc::LOCK_EX | libc::LOCK_NB, + ) + }, + 0 + ); + } + assert!(pool.apply(&selected.selection_sha256).is_err()); + pool.unchanged(&before); + } +} + +#[test] +fn pending_foreign_updates_and_prepared_pools_are_preserved() { + for name in [ + "owner.pending", + "network-update.json", + "network-update.pending", + "prepared-base.json", + "prepared-base.json.pending", + ] { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let path = root(&pool.candidate).join(name); + fs::write(&path, b"foreign-or-interrupted").unwrap(); + assert!(pool.apply(&selected.selection_sha256).is_err()); + pool.unchanged(&before); + assert_eq!(fs::read(path).unwrap(), b"foreign-or-interrupted"); + } +} + +#[test] +fn foreign_alias_and_source_substitution_refuse() { + for case in 0..2 { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + if case == 0 { + fs::write(&pool.owner.short_home, b"foreign").unwrap(); + } else { + let project = &pool.owner.project_share.as_ref().unwrap().project; + fs::rename(project, pool.directory.join("preserved-app")).unwrap(); + state::private_directory(project).unwrap(); + fs::write(project.join("package.json"), b"{}").unwrap(); + } + assert!(pool.apply(&selected.selection_sha256).is_err()); + assert_eq!(pool.receipt(), before); + if case == 0 { + assert_eq!(fs::read(&pool.owner.short_home).unwrap(), b"foreign"); + } + } +} + +#[test] +fn final_recheck_rejects_identity_change_before_publication() { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let called = std::cell::Cell::new(false); + let result = recover_with_boot(&pool.candidate, &selected.selection_sha256, || { + if called.replace(true) { + fs::write(pool.data().join("storage.raw"), b"changed").unwrap(); + } + Ok(pool.boot) + }); + assert!(result.is_err()); + pool.unchanged(&before); + assert_eq!( + fs::read(pool.data().join("storage.raw")).unwrap(), + b"changed" + ); +} + +#[test] +fn substituted_lock_paths_do_not_authorize_publication_on_an_old_descriptor() { + for name in ["operation.lock", "vm.lock"] { + let pool = Pool::new(); + let selected = pool.selected(); + let before = pool.receipt(); + let called = std::cell::Cell::new(false); + let path = if name == "operation.lock" { + root(&pool.candidate).join(name) + } else { + pool.data().join(name) + }; + let result = recover_with_boot(&pool.candidate, &selected.selection_sha256, || { + if called.replace(true) { + fs::rename(&path, path.with_extension("preserved")).unwrap(); + fs::write(&path, b"replacement lock").unwrap(); + } + Ok(pool.boot) + }); + assert!(matches!(result, Err(e) if e.code == "host_filesystem_recovery")); + pool.unchanged(&before); + assert_eq!(fs::read(path).unwrap(), b"replacement lock"); + } +} diff --git a/packages/runtime-core/src/provider/lifecycle/short_home.rs b/packages/runtime-core/src/provider/lifecycle/short_home.rs new file mode 100644 index 000000000..6678ab602 --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/short_home.rs @@ -0,0 +1,70 @@ +//! Explicit two-phase recovery of the receipt-bound temporary HOME after host cleanup. +//! +//! No provider is launched and no existing alias is replaced. Durable disks and VM absence +//! are audited before restoring the short socket path; socket absence is then checked before +//! committing the recovered phase. A later refusal retains the exact alias and unchanged data. +use super::{ + Owner, RuntimeStatus, binary, finish_absent_locked, identity, lock_absent_disks, root, state, + status, verify_disks, +}; +use crate::{Candidate, CandidateError}; + +fn dead_provider(candidate: &Candidate, owner: &Owner) -> Result<(), CandidateError> { + let process = owner.process.as_ref().ok_or_else(|| { + CandidateError::new( + "recovery_required", + "Missing HOME recovery requires a recorded provider identity; nothing was changed.", + ) + })?; + // SAFETY: geteuid has no preconditions. + identity::verify(process, process, &binary(candidate), unsafe { + libc::geteuid() + })?; + if identity::alive(process.pid)? || identity::executable_running(&binary(candidate))? { + return Err(CandidateError::new( + "recovery_required", + "Missing HOME recovery requires a confirmed dead provider and no active provider command; nothing was changed.", + )); + } + if !owner.created || owner.storage.is_none() || owner.overlay.is_none() { + return Err(CandidateError::new( + "recovery_required", + "Missing HOME recovery requires both previously identified disks; nothing was adopted.", + )); + } + Ok(()) +} + +pub(super) fn recover(candidate: &Candidate) -> Result { + let _operation = state::Lock::acquire_existing(&root(candidate))?; + let mut owner = Owner::load_for_short_home_recovery(candidate)?; + dead_provider(candidate, &owner)?; + let vm_lock = lock_absent_disks(candidate, &owner)?; + // Recheck immediately before the only new external effect. The locks fence cooperative + // writers; no signal, disk adoption, overwrite or receipt update happens in this phase. + dead_provider(candidate, &owner)?; + verify_disks(candidate, &owner)?; + owner.restore_missing_short_home(candidate)?; + let mut finish = || -> Result<(), CandidateError> { + let observed = Owner::load(candidate)?; + if observed != owner { + return Err(CandidateError::new( + "foreign_state", + "Provider receipt changed after HOME restoration.", + )); + } + dead_provider(candidate, &owner)?; + verify_disks(candidate, &owner)?; + finish_absent_locked(candidate, &mut owner, "recovered-unclean", false, &vm_lock) + }; + finish().map_err(|error| { + CandidateError::new( + "provider_home_restored_recovery_incomplete", + format!("Owned HOME alias restored; runtime recovery is incomplete ({}). Data retained; inspect and retry runtime recover.", error.code), + ) + })?; + status(candidate) +} + +#[cfg(all(test, target_os = "macos"))] +mod tests; diff --git a/packages/runtime-core/src/provider/lifecycle/short_home/tests.rs b/packages/runtime-core/src/provider/lifecycle/short_home/tests.rs new file mode 100644 index 000000000..f22b62192 --- /dev/null +++ b/packages/runtime-core/src/provider/lifecycle/short_home/tests.rs @@ -0,0 +1,265 @@ +use super::*; +use crate::provider::{NetworkIntent, Profile}; +use std::{ + fs, + os::{ + fd::AsRawFd, + unix::{fs::MetadataExt, net::UnixListener}, + }, + path::PathBuf, + process::Command, +}; + +struct Pool { + candidate: Candidate, + owner: Owner, + directory: PathBuf, +} +impl Pool { + fn new() -> Self { + let mut child = Command::new("/bin/sleep").arg("30").spawn().unwrap(); + let mut process = identity::observe(child.id() as i32).unwrap(); + child.kill().unwrap(); + child.wait().unwrap(); + let directory = fs::canonicalize(std::env::temp_dir()) + .unwrap() + .join(format!( + "hack-home-recovery-{}-{}", + std::process::id(), + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .unwrap() + .as_nanos() + )); + state::private_directory(&directory).unwrap(); + let candidate = Candidate::discover(&directory).unwrap(); + let operation = state::Lock::acquire(&root(&candidate)).unwrap(); + let mut owner = Owner::create( + &candidate, + Profile::Development, + None, + NetworkIntent::Isolated, + ) + .unwrap(); + let data = owner.real_data_dir(&candidate).unwrap(); + state::private_directory(&data).unwrap(); + for (name, tag) in [("storage.raw", 1u8), ("overlay.raw", 2)] { + let mut bytes = vec![0u8; 4096]; + bytes[1080..1082].copy_from_slice(&[0x53, 0xef]); + bytes[1128..1144].fill(tag); + bytes[2048..2059].copy_from_slice(b"data-marker"); + fs::write(data.join(name), bytes).unwrap(); + } + fs::write(data.join("vm.lock"), b"").unwrap(); + fs::write(data.join("name"), &owner.machine).unwrap(); + process.executable = binary(&candidate); + owner.process = Some(process); + owner.created = true; + owner.phase = "running".into(); + owner.storage = Some(identity::disk(&data.join("storage.raw")).unwrap()); + owner.overlay = Some(identity::disk(&data.join("overlay.raw")).unwrap()); + owner.save(&candidate).unwrap(); + drop(operation); + Self { + candidate, + owner, + directory, + } + } + fn remove_alias(&self) { + fs::remove_file(&self.owner.short_home).unwrap(); + } + fn receipt(&self) -> Vec { + fs::read(root(&self.candidate).join("owner.json")).unwrap() + } + fn data(&self) -> PathBuf { + self.owner.real_data_dir(&self.candidate).unwrap() + } + fn unchanged(&self, before: &[u8]) { + assert_eq!(self.receipt(), before); + assert!(self.owner.short_home.symlink_metadata().is_err()); + } +} +impl Drop for Pool { + fn drop(&mut self) { + let _ = fs::remove_file(&self.owner.short_home); + fs::remove_dir_all(&self.directory).unwrap(); + } +} + +#[test] +fn explicit_recovery_restores_only_missing_alias_and_preserves_disk_bytes() { + let pool = Pool::new(); + let storage = fs::read(pool.data().join("storage.raw")).unwrap(); + let overlay = fs::read(pool.data().join("overlay.raw")).unwrap(); + pool.remove_alias(); + assert_eq!( + status(&pool.candidate).unwrap_err().code, + "provider_home_missing" + ); + let recovered = crate::provider::recover(&pool.candidate).unwrap(); + assert_eq!(recovered.phase, "recovered-unclean"); + assert_eq!(recovered.process_alive, Some(false)); + assert_eq!( + fs::read_link(&pool.owner.short_home).unwrap(), + root(&pool.candidate).join("home") + ); + assert_eq!(fs::read(pool.data().join("storage.raw")).unwrap(), storage); + assert_eq!(fs::read(pool.data().join("overlay.raw")).unwrap(), overlay); + assert_eq!( + Owner::load(&pool.candidate).unwrap().storage, + pool.owner.storage + ); +} + +#[test] +fn live_or_reused_pid_and_missing_disk_identity_refuse_without_alias_effect() { + for case in 0..3 { + let mut pool = Pool::new(); + if case < 2 { + let p = pool.owner.process.as_mut().unwrap(); + p.pid = std::process::id() as i32; + p.start_micros = if case == 0 { + identity::observe(p.pid).unwrap().start_micros + } else { + 1 + }; + } else { + pool.owner.overlay = None; + } + pool.owner.save(&pool.candidate).unwrap(); + pool.remove_alias(); + let before = pool.receipt(); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap_err().code, + "recovery_required" + ); + pool.unchanged(&before); + } +} + +#[test] +fn changed_disk_held_vm_lock_and_open_disk_refuse_before_alias_creation() { + for case in 0..3 { + let pool = Pool::new(); + pool.remove_alias(); + let before = pool.receipt(); + let mut held = None; + if case == 0 { + fs::write(pool.data().join("overlay.raw"), b"changed disk").unwrap(); + } + if case == 1 { + let f = fs::File::open(pool.data().join("vm.lock")).unwrap(); + assert_eq!( + unsafe { libc::flock(f.as_raw_fd(), libc::LOCK_EX | libc::LOCK_NB) }, + 0 + ); + held = Some(f); + } + if case == 2 { + held = Some(fs::File::open(pool.data().join("storage.raw")).unwrap()); + } + assert!(crate::provider::recover(&pool.candidate).is_err()); + pool.unchanged(&before); + drop(held); + } +} + +#[test] +fn existing_file_directory_and_foreign_symlink_are_never_replaced() { + for case in 0..3 { + let pool = Pool::new(); + pool.remove_alias(); + let before = pool.receipt(); + if case == 0 { + fs::write(&pool.owner.short_home, b"foreign marker").unwrap(); + } + if case == 1 { + fs::create_dir(&pool.owner.short_home).unwrap(); + } + if case == 2 { + std::os::unix::fs::symlink(&pool.directory, &pool.owner.short_home).unwrap(); + } + let inode = fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap_err().code, + "foreign_state" + ); + assert_eq!(pool.receipt(), before); + assert_eq!( + fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(), + inode + ); + if case == 1 { + fs::remove_dir(&pool.owner.short_home).unwrap(); + } + } +} + +#[test] +fn active_socket_leaves_explicit_incomplete_recovery_then_retry_finishes() { + let pool = Pool::new(); + let listener = UnixListener::bind(pool.owner.data_dir().join("agent.sock")).unwrap(); + pool.remove_alias(); + let before = pool.receipt(); + let error = crate::provider::recover(&pool.candidate).unwrap_err(); + assert_eq!(error.code, "provider_home_restored_recovery_incomplete"); + assert!(error.message.contains("stop_uncertain")); + assert_eq!(pool.receipt(), before); + assert_eq!( + fs::read_link(&pool.owner.short_home).unwrap(), + root(&pool.candidate).join("home") + ); + drop(listener); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap().phase, + "recovered-unclean" + ); +} + +#[test] +fn receipt_change_and_concurrent_alias_creation_refuse_exclusive_repair() { + let mut pool = Pool::new(); + pool.remove_alias(); + let observed = Owner::load_for_short_home_recovery(&pool.candidate).unwrap(); + pool.owner.phase = "stopped".into(); + pool.owner.save(&pool.candidate).unwrap(); + let before = pool.receipt(); + assert_eq!( + observed + .restore_missing_short_home(&pool.candidate) + .unwrap_err() + .code, + "foreign_state" + ); + pool.unchanged(&before); + std::os::unix::fs::symlink(root(&pool.candidate).join("home"), &pool.owner.short_home).unwrap(); + let inode = fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(); + assert_eq!( + pool.owner + .restore_missing_short_home(&pool.candidate) + .unwrap_err() + .code, + "socket_alias_collision" + ); + assert_eq!( + fs::symlink_metadata(&pool.owner.short_home).unwrap().ino(), + inode + ); + assert_eq!(pool.receipt(), before); +} + +#[test] +fn pending_owner_update_is_preserved_and_refuses_alias_repair() { + let pool = Pool::new(); + pool.remove_alias(); + let before = pool.receipt(); + let pending = root(&pool.candidate).join("owner.pending"); + fs::write(&pending, b"interrupted update").unwrap(); + assert_eq!( + crate::provider::recover(&pool.candidate).unwrap_err().code, + "recovery_required" + ); + pool.unchanged(&before); + assert_eq!(fs::read(&pending).unwrap(), b"interrupted update"); +} diff --git a/packages/runtime-core/src/provider/mod.rs b/packages/runtime-core/src/provider/mod.rs index 30f7082bf..e851f1890 100644 --- a/packages/runtime-core/src/provider/mod.rs +++ b/packages/runtime-core/src/provider/mod.rs @@ -45,6 +45,7 @@ pub mod resources; pub mod storage_usage; pub use image_load::load as load_image; mod lifecycle; +pub use lifecycle::host_filesystem; mod network_intent; mod network_update; pub use network_update::{enable_internet, extend_network}; diff --git a/packages/runtime-core/src/provider/state.rs b/packages/runtime-core/src/provider/state.rs index 7f3b1a81a..bc4c3bf84 100644 --- a/packages/runtime-core/src/provider/state.rs +++ b/packages/runtime-core/src/provider/state.rs @@ -57,7 +57,6 @@ impl Lock { } /// Observe existing state without initializing a directory or lock file. - #[cfg(target_os = "macos")] pub fn acquire_existing(root: &Path) -> Result { check_private_directory(root)?; let file = OpenOptions::new() @@ -177,7 +176,7 @@ impl Default for ReclamationPolicy { } } -#[derive(Debug, Serialize, Deserialize)] +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] #[serde(deny_unknown_fields)] pub struct Owner { #[serde(default, skip_serializing_if = "Option::is_none")] @@ -211,8 +210,34 @@ pub struct Owner { } impl Owner { pub fn load(candidate: &Candidate) -> Result { + Self::load_with_short_home(candidate, false) + } + + /// Explicit recovery may inspect a missing temporary alias, never a replaced one. + /// Loading here is read-only; the lifecycle boundary still proves provider absence. + pub(super) fn load_for_short_home_recovery( + candidate: &Candidate, + ) -> Result { + Self::load_with_short_home(candidate, true) + } + + fn load_with_short_home( + candidate: &Candidate, + allow_missing: bool, + ) -> Result { let root = candidate.state_root.join("run/smolvm"); reject_aliased_state(&root)?; + if allow_missing { + match fs::symlink_metadata(root.join("owner.pending")) { + Err(error) if error.kind() == std::io::ErrorKind::NotFound => {} + _ => { + return Err(CandidateError::new( + "recovery_required", + "Provider owner update is pending or unobservable; HOME alias was not restored.", + )); + } + } + } let owner: Self = read(&root.join("owner.json"))?; owner.network.validate()?; if let Some(share) = &owner.project_share { @@ -250,20 +275,69 @@ impl Owner { "Provider owner identity does not match this checkout.", )); } - // The only permitted alias is an exact, receipt-bound short HOME for Unix sockets. - let alias = fs::symlink_metadata(&owner.short_home).map_err(io)?; + owner.check_short_home(candidate, allow_missing)?; + check_private_directory(&root.join("home"))?; + Ok(owner) + } + /// The only permitted alias is the exact receipt-bound HOME for Unix sockets. + fn check_short_home( + &self, + candidate: &Candidate, + allow_missing: bool, + ) -> Result<(), CandidateError> { + let alias = match fs::symlink_metadata(&self.short_home) { + Ok(alias) => alias, + Err(error) if error.kind() == std::io::ErrorKind::NotFound => { + return if allow_missing { + Ok(()) + } else { + Err(CandidateError::new( + "provider_home_missing", + "Temporary provider HOME alias is absent. Use runtime recover to verify the stopped pool before restoring it.", + )) + }; + } + Err(error) => return Err(io(error)), + }; if !alias.file_type().is_symlink() || alias.uid() != unsafe { libc::geteuid() } - || fs::read_link(&owner.short_home).map_err(io)? != root.join("home") + || fs::read_link(&self.short_home).map_err(io)? + != candidate.state_root.join("run/smolvm/home") { return Err(CandidateError::new( "foreign_state", "Short provider HOME alias was replaced.", )); } - check_private_directory(&root.join("home"))?; - Ok(owner) + Ok(()) } + + /// Called only while the lifecycle owns the pool and VM locks and has audited absence. + /// Recheck the durable selection, then create exclusively; a concurrent alias is retained. + pub(super) fn restore_missing_short_home( + &self, + candidate: &Candidate, + ) -> Result<(), CandidateError> { + let observed = Self::load_for_short_home_recovery(candidate)?; + if observed != *self { + return Err(CandidateError::new( + "foreign_state", + "Provider owner changed during HOME recovery; no alias was restored.", + )); + } + std::os::unix::fs::symlink( + candidate.state_root.join("run/smolvm/home"), + &self.short_home, + ) + .map_err(|_| { + CandidateError::new( + "socket_alias_collision", + "Cannot exclusively restore provider HOME alias; no existing path was replaced.", + ) + })?; + self.check_short_home(candidate, false) + } + #[cfg(all(test, target_os = "macos"))] pub fn create( candidate: &Candidate, diff --git a/src/backends/native-project-mapping-command.ts b/src/backends/native-project-mapping-command.ts new file mode 100644 index 000000000..995d5b476 --- /dev/null +++ b/src/backends/native-project-mapping-command.ts @@ -0,0 +1,131 @@ +import { CliUsageError } from "../cli/command.ts"; +import { resolveEffectiveBranch } from "../lib/branches.ts"; +import { emitCliResult, okResult } from "../lib/cli-result.ts"; +import { + findProjectContext, + readProjectConfig, + resolveWorktreeAutoBranch, + sanitizeBranchSlug, +} from "../lib/project.ts"; +import { + inspectNativeProjectRunFilesystemRecovery, + recoverNativeProjectRunFilesystem, +} from "./native-project-run.ts"; +import type { NativeRuntimeSelection } from "./native-runtime-client.ts"; + +const SELECTION = /^[a-f0-9]{64}$/; + +type Selection = + | { readonly action: "inspect" } + | { readonly action: "repair"; readonly expectSelection: string }; + +/** A device rebind is an explicit selected metadata migration, never an automatic Doctor fix. */ +export function parseNativeRunMappingRecoveryOptions(opts: { + readonly action?: string; + readonly expectSelection?: string; + readonly acceptLegacyDeviceRebind?: boolean; + readonly branch?: string; + readonly otherOptions: boolean; +}): Selection | null { + if (opts.action === undefined) { + if ( + opts.expectSelection !== undefined || + opts.acceptLegacyDeviceRebind || + opts.branch !== undefined + ) { + throw new CliUsageError( + "Run-mapping selection and branch options require --native-run-mapping inspect|repair." + ); + } + return null; + } + if (opts.otherOptions) { + throw new CliUsageError( + "--native-run-mapping cannot be combined with other Doctor repair or browser options." + ); + } + if (opts.action === "inspect") { + if (opts.expectSelection !== undefined || opts.acceptLegacyDeviceRebind) { + throw new CliUsageError( + "Run-mapping inspection is read-only; omit repair selection and acceptance flags." + ); + } + return { action: "inspect" }; + } + if ( + opts.action !== "repair" || + !opts.acceptLegacyDeviceRebind || + !opts.expectSelection || + opts.expectSelection.length !== 64 || + !SELECTION.test(opts.expectSelection) + ) { + throw new CliUsageError( + "Run-mapping repair requires --native-run-mapping repair --expect-selection <64-hex> --accept-legacy-device-rebind." + ); + } + return { action: "repair", expectSelection: opts.expectSelection }; +} + +/** Repair only mapping scope metadata. Native graph, socket and data authority remain unchanged. */ +export async function runNativeProjectMappingCommand(opts: { + readonly selection: Selection; + readonly startDir: string; + readonly branch?: string; + readonly runtime: NativeRuntimeSelection; + readonly json: boolean; +}): Promise { + const project = await findProjectContext(opts.startDir); + if (!project) { + throw new CliUsageError( + "Native run-mapping recovery requires a Hack project." + ); + } + const cfg = await readProjectConfig(project); + if (cfg.parseError) { + throw new CliUsageError( + "Native run-mapping recovery requires valid project configuration." + ); + } + const explicit = opts.branch?.trim(); + if (explicit !== undefined && explicit.length === 0) { + throw new CliUsageError("--branch requires a nonempty instance selection."); + } + const selected = await resolveEffectiveBranch({ + explicitBranch: explicit ? sanitizeBranchSlug(explicit) || "branch" : null, + projectRoot: project.projectRoot, + autoBranchEnabled: resolveWorktreeAutoBranch(cfg), + }); + if (selected.source === "detached-worktree") { + throw new CliUsageError( + "Detached worktree recovery requires --branch ." + ); + } + const scope = { + projectRoot: project.projectRoot, + projectDir: project.projectDir, + nativeHome: opts.runtime.home, + branch: selected.branch, + }; + const data = + opts.selection.action === "inspect" + ? await inspectNativeProjectRunFilesystemRecovery({ + scope, + runtime: opts.runtime, + }) + : await recoverNativeProjectRunFilesystem({ + scope, + runtime: opts.runtime, + expectSelection: opts.selection.expectSelection, + acceptLegacyDeviceRebind: true, + }); + if (opts.json) { + emitCliResult({ result: okResult({ data }) }); + } else { + process.stdout.write( + `Native run mapping ${opts.selection.action === "repair" ? "repaired" : "inspected"}: branch=${selected.branch ?? "base"}, run=${data.run.run}\n` + + `selection=${data.selectionSha256}\nqualification=${data.qualification}\n` + + "Graph recovery, retained-data readback and application readiness require separate verification.\n" + ); + } + return 0; +} diff --git a/src/backends/native-project-run.ts b/src/backends/native-project-run.ts index e48916f04..c2825cbe7 100644 --- a/src/backends/native-project-run.ts +++ b/src/backends/native-project-run.ts @@ -16,10 +16,16 @@ import { isNativeProjectFinalizationToken, type NativeProjectFinalizationToken, } from "./native-project-finalization.ts"; +import { inspectNativeProjectGraph } from "./native-project-inspect.ts"; +import { + invokeNativeRuntime, + type NativeRuntimeSelection, +} from "./native-runtime-client.ts"; const HEX32 = /^[a-f0-9]{32}$/; const HEX64 = /^[a-f0-9]{64}$/; const LIMIT = 8192; +const RECOVERY_SCHEMA = "hack.native-project-run-filesystem-recovery/v1"; // Equivalent to normalizeEnvConfigName(value) === value, with a bounded length. const ENV_NAME = /^[a-z0-9]+(?:-[a-z0-9]+)*$/; const AWS_PROFILE = /^[A-Za-z0-9_+=,.@-]{1,128}$/; @@ -390,6 +396,548 @@ function record(value: unknown, identity: unknown): NativeProjectRun { } return value.run; } + +type DirectoryIdentity = { readonly dev: number; readonly ino: number }; +type RunScopeIdentity = Awaited>; +type RecoveryPaths = Awaited>; + +export type NativeProjectRunFilesystemInspection = { + readonly schema: typeof RECOVERY_SCHEMA; + readonly selectionSha256: string; + readonly mappingSha256: string; + readonly run: NativeProjectRun; + readonly oldDevice: number; + readonly newDevice: number; + readonly qualification: "explicit-legacy-rebind-original-volume-continuity-unproven"; +}; + +export type NativeProjectRunFilesystemRecovery = + NativeProjectRunFilesystemInspection & { + readonly repaired: true; + readonly auditPath: string; + }; + +function sameIdentity( + value: unknown, + expected: DirectoryIdentity, + oldDevice: number +): boolean { + return ( + isRecord(value) && + Object.keys(value).sort().join() === "dev,ino" && + typeof value.dev === "number" && + Number.isSafeInteger(value.dev) && + value.dev === oldDevice && + value.ino === expected.ino + ); +} + +async function readRecoveryFile(path: string): Promise<{ + bytes: Buffer; + value: unknown; + identity: DirectoryIdentity; +}> { + const fd = await open( + path, + constants.O_RDONLY | constants.O_NONBLOCK | constants.O_NOFOLLOW + ); + try { + const before = await fd.stat(); + if ( + !before.isFile() || + before.nlink !== 1 || + before.size < 1 || + before.size > LIMIT || + before.uid !== process.getuid?.() || + (before.mode & 0o777) !== 0o600 + ) { + throw refused(); + } + const bytes = Buffer.alloc(before.size); + const result = await fd.read(bytes, 0, bytes.length, 0); + const after = await fd.stat(); + const named = await lstat(path); + if ( + result.bytesRead !== bytes.length || + before.dev !== after.dev || + before.ino !== after.ino || + before.size !== after.size || + before.mtimeMs !== after.mtimeMs || + before.ctimeMs !== after.ctimeMs || + named.dev !== after.dev || + named.ino !== after.ino || + !named.isFile() || + named.isSymbolicLink() || + after.nlink !== 1 || + after.uid !== process.getuid?.() || + (after.mode & 0o777) !== 0o600 + ) { + throw refused(); + } + const value: unknown = JSON.parse( + new TextDecoder("utf-8", { fatal: true }).decode(bytes) + ); + return { bytes, value, identity: { dev: after.dev, ino: after.ino } }; + } finally { + await fd.close(); + } +} + +async function preserveRecoveryAudit( + path: string, + expected: string +): Promise { + const limit = LIMIT * 2 + 4096; + if (Buffer.byteLength(expected) > limit) { + throw refused(); + } + try { + await write(path, expected); + return; + } catch (error) { + if (!isRecord(error) || error.code !== "EEXIST") { + throw error; + } + } + const fd = await open( + path, + constants.O_RDONLY | constants.O_NONBLOCK | constants.O_NOFOLLOW + ); + try { + const before = await fd.stat(); + if ( + !before.isFile() || + before.nlink !== 1 || + before.uid !== process.getuid?.() || + (before.mode & 0o777) !== 0o600 || + before.size !== Buffer.byteLength(expected) || + before.size > limit + ) { + throw refused(); + } + const bytes = Buffer.alloc(before.size); + const read = await fd.read(bytes, 0, bytes.length, 0); + const after = await fd.stat(); + const named = await lstat(path); + if ( + read.bytesRead !== bytes.length || + !bytes.equals(Buffer.from(expected)) || + before.dev !== after.dev || + before.ino !== after.ino || + before.mtimeMs !== after.mtimeMs || + before.ctimeMs !== after.ctimeMs || + named.dev !== after.dev || + named.ino !== after.ino || + !named.isFile() || + named.nlink !== 1 + ) { + throw refused(); + } + } finally { + await fd.close(); + } +} + +function oldMapping( + value: unknown, + current: RunScopeIdentity +): { + readonly oldDevice: number; + readonly run: NativeProjectRun; + readonly next: string; +} { + if ( + !isRecord(value) || + Object.keys(value).sort().join() !== "run,scope,version" || + value.version !== 1 || + !valid(value.run) || + !isRecord(value.scope) || + Object.keys(value.scope).sort().join() !== + "branch,dirIdentity,homeIdentity,nativeHome,projectDir,projectRoot,rootIdentity" + ) { + throw refused(); + } + const stored = value.scope; + const oldRoot = stored.rootIdentity; + if ( + !isRecord(oldRoot) || + typeof oldRoot.dev !== "number" || + !Number.isSafeInteger(oldRoot.dev) || + oldRoot.dev <= 0 + ) { + throw refused(); + } + const oldDevice = oldRoot.dev; + if ( + oldDevice === current.rootIdentity.dev || + current.rootIdentity.dev !== current.dirIdentity.dev || + current.rootIdentity.dev !== current.homeIdentity.dev || + stored.projectRoot !== current.projectRoot || + stored.projectDir !== current.projectDir || + stored.nativeHome !== current.nativeHome || + stored.branch !== current.branch || + !sameIdentity(oldRoot, current.rootIdentity, oldDevice) || + !sameIdentity(stored.dirIdentity, current.dirIdentity, oldDevice) || + !sameIdentity(stored.homeIdentity, current.homeIdentity, oldDevice) + ) { + throw refused(); + } + const next = JSON.stringify({ ...value, scope: current }); + if ( + Buffer.byteLength(next) > LIMIT || + JSON.stringify(value.run) !== JSON.stringify(JSON.parse(next).run) + ) { + throw refused(); + } + return { oldDevice, run: value.run, next }; +} + +async function absentPath(path: string): Promise { + try { + await lstat(path); + } catch (error) { + if (isRecord(error) && error.code === "ENOENT") { + return; + } + throw refused(); + } + throw refused(); +} + +async function recoveryLocksAbsent(p: RecoveryPaths): Promise { + const restartFile = `${p.file.slice(0, -5)}.restart.json`; + const restartLock = `${p.lock.slice(0, -5)}.restart.lock`; + await Promise.all([ + absentPath(p.lock), + absentPath(`${p.lock}.operation`), + absentPath(restartFile), + absentPath(restartLock), + absentPath(`${restartLock}.operation`), + ]); +} + +async function recoveryDirectories(p: RecoveryPaths): Promise { + for (const path of [ + p.identity.projectRoot, + p.identity.projectDir, + p.identity.nativeHome, + join(p.identity.projectDir, ".internal"), + p.root, + ]) { + const metadata = await lstat(path); + if ( + !metadata.isDirectory() || + metadata.isSymbolicLink() || + metadata.uid !== process.getuid?.() || + (metadata.mode & 0o022) !== 0 || + (path === p.root && (metadata.mode & 0o777) !== 0o700) + ) { + throw refused(); + } + } +} + +async function heldDirectory(path: string) { + const fd = await open(path, constants.O_RDONLY | constants.O_NOFOLLOW); + try { + const metadata = await fd.stat(); + if ( + !metadata.isDirectory() || + metadata.uid !== process.getuid?.() || + (metadata.mode & 0o777) !== 0o700 + ) { + throw refused(); + } + const verify = async () => { + const named = await lstat(path); + const held = await fd.stat(); + if ( + !named.isDirectory() || + named.isSymbolicLink() || + named.dev !== metadata.dev || + named.ino !== metadata.ino || + held.dev !== metadata.dev || + held.ino !== metadata.ino || + named.uid !== metadata.uid || + (named.mode & 0o777) !== 0o700 + ) { + throw refused(); + } + }; + await verify(); + return { verify, close: () => fd.close() }; + } catch (error) { + await fd.close(); + throw error; + } +} + +async function releaseHeldDirectory( + path: string, + held: Awaited> +) { + try { + await held.verify(); + await rmdir(path); + } finally { + await held.close(); + } +} + +async function verifySelectedMappingPath( + path: string, + selected: Buffer, + identity: DirectoryIdentity +): Promise<() => Promise> { + const fd = await open( + path, + constants.O_RDONLY | constants.O_NONBLOCK | constants.O_NOFOLLOW + ); + try { + const before = await fd.stat(); + if ( + !before.isFile() || + before.nlink !== 1 || + before.uid !== process.getuid?.() || + (before.mode & 0o777) !== 0o600 || + before.size !== selected.length || + before.dev !== identity.dev || + before.ino !== identity.ino + ) { + throw refused(); + } + const bytes = Buffer.alloc(selected.length); + const read = await fd.read(bytes, 0, bytes.length, 0); + const after = await fd.stat(); + const named = await lstat(path); + if ( + read.bytesRead !== bytes.length || + !bytes.equals(selected) || + before.dev !== after.dev || + before.ino !== after.ino || + before.mtimeMs !== after.mtimeMs || + before.ctimeMs !== after.ctimeMs || + named.dev !== after.dev || + named.ino !== after.ino || + named.nlink !== 1 || + named.uid !== process.getuid?.() || + (named.mode & 0o777) !== 0o600 + ) { + throw refused(); + } + return () => fd.close(); + } catch (error) { + await fd.close(); + throw error; + } +} + +async function nativeRecoveryAuthority(opts: { + readonly p: RecoveryPaths; + readonly run: NativeProjectRun; + readonly oldDevice: number; + readonly runtime: NativeRuntimeSelection; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + const invoke = opts.invoke ?? invokeNativeRuntime; + const [graph, status] = await Promise.all([ + inspectNativeProjectGraph({ + runtime: opts.runtime, + projectRoot: opts.p.identity.projectRoot, + run: opts.run.run, + invoke, + }), + invoke({ + runtime: opts.runtime, + cwd: opts.p.identity.projectRoot, + args: ["runtime", "status", "--json"], + timeoutMs: 30_000, + }), + ]); + const receipt = isRecord(graph) ? graph.receipt : undefined; + const source = isRecord(receipt) ? receipt.source : undefined; + const shared = isRecord(source) ? source.shared : undefined; + const share = isRecord(status) ? status.project_share : undefined; + if ( + !isRecord(graph) || + graph.journal_incomplete !== false || + !isRecord(receipt) || + receipt.run !== opts.run.run || + receipt.owner !== opts.run.owner || + receipt.namespace !== opts.run.namespace || + receipt.plan_id !== opts.run.planId || + !["ready-observed", "stopped-data-retained"].includes( + String(receipt.phase) + ) || + !isRecord(status) || + status.phase !== "running" || + status.process_alive !== true || + status.persistent_disks_identified !== true || + !isRecord(share) || + share.project !== opts.p.identity.projectRoot || + share.device !== opts.p.identity.rootIdentity.dev || + share.inode !== opts.p.identity.rootIdentity.ino || + share.unfiltered_source !== true || + !isRecord(shared) || + shared.project !== share.project || + shared.guest_path !== share.guest_path || + shared.inode !== share.inode || + shared.device !== opts.oldDevice || + shared.unfiltered_source !== true + ) { + throw refused(); + } +} + +async function recoverySelection(opts: { + readonly scope: NativeProjectRunScope; + readonly runtime: NativeRuntimeSelection; + readonly invoke?: typeof invokeNativeRuntime; + readonly heldLocks?: boolean; +}): Promise<{ + readonly inspection: NativeProjectRunFilesystemInspection; + readonly p: RecoveryPaths; + readonly bytes: Buffer; + readonly fileIdentity: DirectoryIdentity; + readonly next: string; +}> { + const p = await paths(opts.scope, false); + await recoveryDirectories(p); + if (!opts.heldLocks) { + await recoveryLocksAbsent(p); + } + const { bytes, value, identity } = await readRecoveryFile(p.file); + const mapping = oldMapping(value, p.identity); + await nativeRecoveryAuthority({ + p, + run: mapping.run, + oldDevice: mapping.oldDevice, + runtime: opts.runtime, + invoke: opts.invoke, + }); + const mappingSha256 = createHash("sha256").update(bytes).digest("hex"); + const selected = { + schema: RECOVERY_SCHEMA, + mappingSha256, + fileIdentity: identity, + scope: p.identity, + oldDevice: mapping.oldDevice, + run: mapping.run, + }; + const inspection: NativeProjectRunFilesystemInspection = { + schema: RECOVERY_SCHEMA, + selectionSha256: createHash("sha256") + .update(JSON.stringify(selected)) + .digest("hex"), + mappingSha256, + run: mapping.run, + oldDevice: mapping.oldDevice, + newDevice: p.identity.rootIdentity.dev, + qualification: "explicit-legacy-rebind-original-volume-continuity-unproven", + }; + return { inspection, p, bytes, fileIdentity: identity, next: mapping.next }; +} + +/** Inspect an exact legacy device renumbering without changing frontend or native state. */ +export async function inspectNativeProjectRunFilesystemRecovery(opts: { + readonly scope: NativeProjectRunScope; + readonly runtime: NativeRuntimeSelection; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + return (await recoverySelection(opts)).inspection; +} + +/** Explicitly rebind only three mapping device numbers; graph history is untouched. */ +export async function recoverNativeProjectRunFilesystem(opts: { + readonly scope: NativeProjectRunScope; + readonly runtime: NativeRuntimeSelection; + readonly expectSelection: string; + readonly acceptLegacyDeviceRebind: true; + readonly invoke?: typeof invokeNativeRuntime; +}): Promise { + if ( + !HEX64.test(opts.expectSelection) || + opts.acceptLegacyDeviceRebind !== true + ) { + throw refused(); + } + const initial = await recoverySelection(opts); + if (initial.inspection.selectionSha256 !== opts.expectSelection) { + throw refused(); + } + const restartLock = `${initial.p.lock.slice(0, -5)}.restart.lock.operation`; + await mkdir(restartLock, { mode: 0o700 }); + const restartHeld = await heldDirectory(restartLock); + try { + await mkdir(initial.p.lock, { mode: 0o700 }); + const mappingHeld = await heldDirectory(initial.p.lock); + try { + const selected = await recoverySelection({ ...opts, heldLocks: true }); + if ( + selected.inspection.selectionSha256 !== opts.expectSelection || + selected.p.root !== initial.p.root || + selected.p.file !== initial.p.file || + !selected.bytes.equals(initial.bytes) || + selected.fileIdentity.dev !== initial.fileIdentity.dev || + selected.fileIdentity.ino !== initial.fileIdentity.ino + ) { + throw refused(); + } + await absentPath(`${selected.p.file.slice(0, -5)}.restart.json`); + const audit = `${selected.p.file.slice(0, -5)}.filesystem-recovery-${selected.inspection.mappingSha256}.json`; + const auditText = JSON.stringify({ + version: 1, + selectionSha256: selected.inspection.selectionSha256, + mappingSha256: selected.inspection.mappingSha256, + mapping: selected.bytes.toString("utf8"), + }); + await preserveRecoveryAudit(audit, auditText); + await sync(selected.p.root); + const final = await recoverySelection({ ...opts, heldLocks: true }); + if ( + final.inspection.selectionSha256 !== opts.expectSelection || + !final.bytes.equals(selected.bytes) || + final.fileIdentity.dev !== selected.fileIdentity.dev || + final.fileIdentity.ino !== selected.fileIdentity.ino + ) { + throw refused(); + } + const temporary = join(selected.p.root, `${randomUUID()}.tmp`); + try { + await write(temporary, selected.next); + await Promise.all([restartHeld.verify(), mappingHeld.verify()]); + const closeMapping = await verifySelectedMappingPath( + selected.p.file, + selected.bytes, + selected.fileIdentity + ); + try { + await rename(temporary, selected.p.file); + } finally { + await closeMapping(); + } + await sync(selected.p.root); + } finally { + await unlink(temporary).catch((error: unknown) => { + if (!isRecord(error) || error.code !== "ENOENT") { + throw error; + } + }); + } + if ( + JSON.stringify(await loadNativeProjectRun(opts.scope)) !== + JSON.stringify(selected.inspection.run) + ) { + throw refused(); + } + return { ...selected.inspection, repaired: true, auditPath: audit }; + } finally { + await releaseHeldDirectory(initial.p.lock, mappingHeld); + } + } finally { + await releaseHeldDirectory(restartLock, restartHeld); + } +} /** Establish excluded metadata before source identity is reviewed. */ export async function prepareNativeProjectRunStorage( opts: NativeProjectRunScope diff --git a/src/commands/doctor.ts b/src/commands/doctor.ts index 7b3355051..883b609ed 100644 --- a/src/commands/doctor.ts +++ b/src/commands/doctor.ts @@ -8,6 +8,10 @@ import { checkLegacyProjectAgentArtifacts, checkLegacyUserAgentArtifacts, } from "../agents/legacy-artifacts.ts"; +import { + parseNativeRunMappingRecoveryOptions, + runNativeProjectMappingCommand, +} from "../backends/native-project-mapping-command.ts"; import { type NativeRuntimeSelection, resolveNativeRuntimeSelection, @@ -19,7 +23,7 @@ import { defineOption, withHandler, } from "../cli/command.ts"; -import { optJson, optPath } from "../cli/options.ts"; +import { optBranch, optJson, optPath } from "../cli/options.ts"; import { DEFAULT_CADDY_IP, DEFAULT_HOST_DNS_IP, @@ -208,6 +212,29 @@ const doctorOptions = [ optJson, optBrowserUrl, optBrowserResult, + optBranch, + defineOption({ + name: "nativeRunMapping", + type: "string", + long: "--native-run-mapping", + valueHint: "inspect|repair", + description: + "Inspect or explicitly repair a native run mapping after filesystem device renumbering", + } as const), + defineOption({ + name: "expectSelection", + type: "string", + long: "--expect-selection", + valueHint: "<64-hex>", + description: "Require the exact run-mapping recovery inspection selection", + } as const), + defineOption({ + name: "acceptLegacyDeviceRebind", + type: "boolean", + long: "--accept-legacy-device-rebind", + description: + "Explicitly accept legacy migration without proof of original filesystem volume continuity", + } as const), ] as const; const doctorPositionals = [] as const; @@ -332,9 +359,49 @@ async function maybeRunDomainMigration( return null; } +async function maybeRunNativeMappingRecovery( + args: Parameters>[0]["args"] +): Promise { + const mapping = parseNativeRunMappingRecoveryOptions({ + action: args.options.nativeRunMapping, + expectSelection: args.options.expectSelection, + acceptLegacyDeviceRebind: args.options.acceptLegacyDeviceRebind, + branch: args.options.branch, + otherOptions: Boolean( + args.options.fix || + args.options.migrateEnvConfig || + args.options.domainMigration || + args.options.browserUrl || + args.options.browserResult + ), + }); + if (mapping) { + const runtime = resolveNativeRuntimeSelection(); + if (!runtime) { + throw new CliUsageError( + "Native run-mapping recovery requires an explicitly selected native runtime." + ); + } + return await runNativeProjectMappingCommand({ + selection: mapping, + runtime, + startDir: args.options.path + ? resolve(process.cwd(), args.options.path) + : process.cwd(), + branch: args.options.branch, + json: args.options.json === true, + }); + } + return null; +} + const handleDoctor: CommandHandlerFor = async ({ args, }): Promise => { + const mappingResult = await maybeRunNativeMappingRecovery(args); + if (mappingResult !== null) { + return mappingResult; + } const domainResult = await maybeRunDomainMigration(args); if (domainResult !== null) { return domainResult; diff --git a/tests/native-project-mapping-command.test.ts b/tests/native-project-mapping-command.test.ts new file mode 100644 index 000000000..8e707d586 --- /dev/null +++ b/tests/native-project-mapping-command.test.ts @@ -0,0 +1,73 @@ +import { expect, test } from "bun:test"; +import { parseNativeRunMappingRecoveryOptions as parse } from "../src/backends/native-project-mapping-command.ts"; + +const selected = "a".repeat(64); + +test("ordinary Doctor does not select a mapping migration", () => { + expect(parse({ otherOptions: false })).toBeNull(); +}); + +test("inspection refuses every mutation selector", () => { + expect( + parse({ action: "inspect", branch: "feature-api", otherOptions: false }) + ).toEqual({ action: "inspect" }); + for (const options of [ + { expectSelection: selected }, + { acceptLegacyDeviceRebind: true }, + ]) { + expect(() => + parse({ action: "inspect", ...options, otherOptions: false }) + ).toThrow("read-only"); + } +}); + +test("repair requires both an exact inspection selection and explicit legacy acceptance", () => { + expect( + parse({ + action: "repair", + expectSelection: selected, + acceptLegacyDeviceRebind: true, + otherOptions: false, + }) + ).toEqual({ action: "repair", expectSelection: selected }); + for (const options of [ + {}, + { expectSelection: selected }, + { acceptLegacyDeviceRebind: true }, + { expectSelection: "a".repeat(63), acceptLegacyDeviceRebind: true }, + { expectSelection: "A".repeat(64), acceptLegacyDeviceRebind: true }, + { expectSelection: `${selected}\n`, acceptLegacyDeviceRebind: true }, + ]) { + expect(() => + parse({ action: "repair", ...options, otherOptions: false }) + ).toThrow("requires"); + } +}); + +test("unknown actions and orphaned selectors refuse before project lookup", () => { + expect(() => parse({ action: "force", otherOptions: false })).toThrow( + "requires" + ); + for (const options of [ + { expectSelection: selected }, + { acceptLegacyDeviceRebind: true }, + { branch: "feature-api" }, + ]) { + expect(() => parse({ ...options, otherOptions: false })).toThrow( + "require --native-run-mapping" + ); + } +}); + +test("mapping migration cannot combine with ordinary repair, domain or browser flows", () => { + for (const action of ["inspect", "repair"]) { + expect(() => + parse({ + action, + expectSelection: selected, + acceptLegacyDeviceRebind: true, + otherOptions: true, + }) + ).toThrow("cannot be combined"); + } +}); diff --git a/tests/native-project-run-filesystem.test.ts b/tests/native-project-run-filesystem.test.ts new file mode 100644 index 000000000..742a1c949 --- /dev/null +++ b/tests/native-project-run-filesystem.test.ts @@ -0,0 +1,331 @@ +import { afterEach, expect, test } from "bun:test"; +import { createHash } from "node:crypto"; +import { + chmod, + lstat, + mkdir, + mkdtemp, + readFile, + realpath, + rename, + rm, + writeFile, +} from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { + inspectNativeProjectRunFilesystemRecovery as inspect, + loadNativeProjectRun as load, + type NativeProjectRunScope, + recoverNativeProjectRunFilesystem as recover, + saveNativeProjectRun as save, +} from "../src/backends/native-project-run.ts"; +import type { NativeRuntimeSelection } from "../src/backends/native-runtime-client.ts"; + +const roots: string[] = []; +afterEach(async () => { + await Promise.all( + roots.splice(0).map((root) => rm(root, { recursive: true, force: true })) + ); +}); + +const run = { + run: "a".repeat(32), + owner: "b".repeat(32), + namespace: "c".repeat(64), + planId: "d".repeat(64), + effectiveEnvName: "qa", + profiles: ["worker"], + aws: { profile: "livenation_qa", region: "us-east-1" }, +}; +const runtime: NativeRuntimeSelection = { + binary: "/unused/hack-native", + home: "/unused/home", +}; + +async function fixture() { + const projectRoot = await realpath( + await mkdtemp(join(tmpdir(), "native-mapping-reboot-")) + ); + roots.push(projectRoot); + const projectDir = join(projectRoot, ".hack"); + const nativeHome = join(projectRoot, "candidate"); + await mkdir(projectDir); + await mkdir(nativeHome); + const scope: NativeProjectRunScope = { + projectRoot, + projectDir, + nativeHome, + branch: "event-agent", + }; + await save({ ...scope, run }); + const key = createHash("sha256") + .update(JSON.stringify(scope.branch)) + .digest("hex"); + const directory = join(projectDir, ".internal/native-runs"); + const file = join(directory, `${key}.json`); + const original = JSON.parse(await readFile(file, "utf8")); + const oldDevice = original.scope.rootIdentity.dev + 1; + for (const name of ["rootIdentity", "dirIdentity", "homeIdentity"]) { + original.scope[name].dev = oldDevice; + } + const oldBytes = JSON.stringify(original); + await writeFile(file, oldBytes); + return { scope, directory, file, key, oldDevice, oldBytes, original }; +} + +function authority( + pool: Awaited>, + opts: { + onCall?: (call: number) => Promise; + graph?: Record; + status?: Record; + } = {} +) { + let call = 0; + return async (input: { + readonly args: readonly string[]; + }): Promise => { + call++; + await opts.onCall?.(call); + if (input.args[0] === "runtime") { + return ( + opts.status ?? { + phase: "running", + process_alive: true, + persistent_disks_identified: true, + project_share: { + project: pool.scope.projectRoot, + guest_path: "/mnt/hack-projects/fixture", + device: (await lstat(pool.scope.projectRoot)).dev, + inode: (await lstat(pool.scope.projectRoot)).ino, + unfiltered_source: true, + }, + } + ); + } + return ( + opts.graph ?? { + journal_incomplete: false, + receipt: { + run: run.run, + owner: run.owner, + namespace: run.namespace, + plan_id: run.planId, + phase: "ready-observed", + source: { + shared: { + project: pool.scope.projectRoot, + guest_path: "/mnt/hack-projects/fixture", + device: pool.oldDevice, + inode: (await lstat(pool.scope.projectRoot)).ino, + unfiltered_source: true, + }, + }, + }, + } + ); + }; +} + +test("explicit mapping repair changes only three devices and retains byte-exact audit", async () => { + const pool = await fixture(); + const invoke = authority(pool); + await expect(load(pool.scope)).rejects.toThrow("Native project run mapping"); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + expect(selected.qualification).toContain("unproven"); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); + const repaired = await recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }); + expect(repaired.repaired).toBe(true); + expect(repaired.selectionSha256).toBe(selected.selectionSha256); + expect(await load(pool.scope)).toEqual(run); + const after = JSON.parse(await readFile(pool.file, "utf8")); + expect(after.run).toEqual(pool.original.run); + for (const name of ["rootIdentity", "dirIdentity", "homeIdentity"]) { + expect(after.scope[name].ino).toBe(pool.original.scope[name].ino); + expect(after.scope[name].dev).toBe(selected.newDevice); + } + const audit = JSON.parse(await readFile(repaired.auditPath, "utf8")); + expect(audit.mapping).toBe(pool.oldBytes); + expect(audit.mappingSha256).toBe(selected.mappingSha256); + await expect( + inspect({ scope: pool.scope, runtime, invoke }) + ).rejects.toThrow(); +}); + +test("partial device changes, changed inode, paths, branch and malformed mapping refuse", async () => { + for (const change of [ + "partial", + "inode", + "path", + "branch", + "schema", + "mode", + "link", + ]) { + const pool = await fixture(); + const invoke = authority(pool); + if (change === "mode") { + await chmod(pool.file, 0o644); + } else if (change === "link") { + await writeFile(`${pool.file}.link`, "foreign"); + await rm(`${pool.file}.link`); + const { link } = await import("node:fs/promises"); + await link(pool.file, `${pool.file}.link`); + } else { + const value = JSON.parse(pool.oldBytes); + if (change === "partial") { + value.scope.dirIdentity.dev++; + } + if (change === "inode") { + value.scope.homeIdentity.ino++; + } + if (change === "path") { + value.scope.projectRoot = "/foreign"; + } + if (change === "branch") { + value.scope.branch = "foreign"; + } + if (change === "schema") { + value.extra = true; + } + await writeFile(pool.file, JSON.stringify(value)); + } + await expect( + inspect({ scope: pool.scope, runtime, invoke }) + ).rejects.toThrow(); + } +}); + +test("native graph and current share must match the exact mapped run", async () => { + const pool = await fixture(); + for (const override of [ + { graph: { journal_incomplete: true, receipt: {} } }, + { graph: { journal_incomplete: false, receipt: { run: "f".repeat(32) } } }, + { status: { phase: "running", project_share: null } }, + ]) { + await expect( + inspect({ scope: pool.scope, runtime, invoke: authority(pool, override) }) + ).rejects.toThrow(); + } +}); + +test("stale selection, pending restart and foreign locks never change the mapping", async () => { + const pool = await fixture(); + const invoke = authority(pool); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: "f".repeat(64), + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + const lock = join(pool.directory, `${pool.key}.restart.lock.operation`); + await mkdir(lock); + await expect( + inspect({ scope: pool.scope, runtime, invoke }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); + await rm(lock, { recursive: true }); + await writeFile(join(pool.directory, `${pool.key}.restart.json`), "{}"); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); +}); + +test("replacement mapping inode during final authority check refuses", async () => { + const pool = await fixture(); + const invoke = authority(pool, { + onCall: async (call) => { + if (call === 7) { + await rename(pool.file, `${pool.file}.preserved`); + await writeFile(pool.file, pool.oldBytes, { mode: 0o600 }); + } + }, + }); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); +}); + +test("substituted lock paths are preserved and cannot authorize publication", async () => { + for (const suffix of [".restart.lock.operation", ".lock"]) { + const pool = await fixture(); + const lock = join(pool.directory, `${pool.key}${suffix}`); + const invoke = authority(pool, { + onCall: async (call) => { + if (call === 7) { + await rename(lock, `${lock}.preserved`); + await mkdir(lock, { mode: 0o700 }); + } + }, + }); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + expect(await readFile(pool.file, "utf8")).toBe(pool.oldBytes); + expect((await lstat(lock)).isDirectory()).toBe(true); + } +}); + +test("an exact published audit supports retry after prepublication refusal", async () => { + const pool = await fixture(); + const invoke = authority(pool, { + onCall: async (call) => { + if (call === 7) { + await writeFile(pool.file, `${pool.oldBytes} `); + } + }, + }); + const selected = await inspect({ scope: pool.scope, runtime, invoke }); + await expect( + recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke, + }) + ).rejects.toThrow(); + await writeFile(pool.file, pool.oldBytes); + const repaired = await recover({ + scope: pool.scope, + runtime, + expectSelection: selected.selectionSha256, + acceptLegacyDeviceRebind: true, + invoke: authority(pool), + }); + expect(repaired.repaired).toBe(true); + expect(await load(pool.scope)).toEqual(run); +});