diff --git a/CHANGELOG.md b/CHANGELOG.md index 7a84944..da80ddb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +### Added + +- **Runtime range access control: the `expensive-slice-access-control` feature (implies `alloc` + `set`; Rust only; off by default).** A per-stack point table assigns an access mode to every payload offset, and each mutating/reading entry point (`set`/`get`/`push`/`pop`/`extend`/`resize`/`ensure`/`swap`/`splice`/`atrunc`/`cas`/`copy`/`process`/`process_gen`/`inplace_gen` and the batched/sparse/`try_*` variants) checks the touched range's required authority before committing, failing with `io::ErrorKind::PermissionDenied` on denial. Protection is armed through `BStackOwnedSlice::protect`/`protect_as` and carried by two one-shot, `!Clone`, pointer-identity-checked capability tokens — `BStackProtection` (guard authority, minted once via `take_protection`) and `BStackAllocAuthority` (allocator authority, `take_alloc_authority`) — which are mutually incomparable, so a token from a different stack or the wrong axis grants nothing. A holder reaches a protected range through the `_as` sibling of any checked method (`set_as`, `get_as`, `swap_as`, `resize_as`, `push_as`, …); a `BStackOwnedSlice`/`BStackSlice` granted an authority via `authorize` routes its region I/O through them automatically, and `merge`/`merge_adjacent` refuse when two views carry different authorities. New public types: `BStackAccess`, `AccessOp`, `BStackAccessRequirement`, `BStackAccessAuthorities`, `BStackProtection`, `BStackAllocAuthority`, `BStackAuthority`. The whole feature compiles away when off — the check macro folds to nothing and a build without the feature is byte-for-byte unchanged. +- **Allocators reserve their own metadata and refuse to free protected regions (`expensive-slice-access-control`; Rust only).** Every built-in allocator burns the alloc-authority mint on construction and marks its fixed header `Alloc`, routing its own metadata I/O through that authority so a protected header cannot be corrupted. `dealloc`/`dealloc_bulk` refuse to free any range still carrying a caller mode outside `{All, Alloc}`, returning `PermissionDenied` and handing the handle(s) back intact rather than silently dropping the caller's protection into the region's next owner; the caller must lift its own `protect` first. + ### Changed - **`SegregatedBStackAllocator` (Rust) / `segregated_bstack_allocator_*` (C) shrink/split heuristics tuned (`alloc` + `set` / `BSTACK_FEATURE_SET`; the tail-shrink reclaim additionally `atomic` / `BSTACK_FEATURE_ATOMIC`).** `SPLIT_MIN` / `ALSG_SPLIT_MIN` raised from `LINEAR_MAX` (256) to `MAX_CLASS` (4096): a smaller excess is now retained as slack rather than carved into the class free lists, where the pieces rarely reuse and strand as dead arena. Separately, a non-tail (interior) shrink now always retains its freed excess in place instead of carving it — only a *tail* shrink still reclaims, and only under `atomic` / `BSTACK_FEATURE_ATOMIC`, via the existing `Len` + `Atrunc` (`BSTACK_GEN_LEN` + `BSTACK_GEN_SPLICE`). Both cut durable syncs and file growth on skewed and realloc-heavy workloads. Heuristics only and the on-disk format is unchanged. diff --git a/Cargo.toml b/Cargo.toml index 74826bb..29e6f7e 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -22,6 +22,9 @@ alloc = [] atomic = [] # Enables BStackGuardedSlice and related hook-based slice abstractions. guarded = ["alloc"] +# Enables runtime range access control on BStack / BStackOwnedSlice (the +# BStackAccess point table). Off by default; See the `acl_core` module. +expensive-slice-access-control = ["alloc", "set"] # Enables deterministic I/O-fault injection at the BStack API level, for testing # allocator and downstream error-handling paths. Only takes effect in builds with # `debug_assertions` on (dev/test); release builds are completely unaffected even diff --git a/PLANNED.md b/PLANNED.md index d040f09..2a5ce29 100644 --- a/PLANNED.md +++ b/PLANNED.md @@ -354,75 +354,6 @@ Cost is `n` staged bytes and `2k` syncs. Staging is independent of `k`, so rotat --- -## Range access control on `BStack` and `BStackOwnedSlice` - -**Feature flag:** `expensive-slice-access-control` (implies `alloc` + `set`). Off by default. -**Breaking change:** No. A build without the flag compiles to exactly today's code. - -### Motivation - -`bstack` has one enforcement mechanism for "these bytes must not change": `lock_up_to`. It is shaped for the stack case — a consumer whose bottom `n` bytes are settled and whose later pushes build on them — and is right there. For anything else it is a prefix, it is all-or-nothing, and it conflates writing with truncating. - -What nothing encodes is *who is asking*. Ownership and borrowing settle aliasing, but aliasing is not authority: two callers holding the same range are indistinguishable, and nothing can say that one may write it and the other may not. It is the axis an OS gives every page, where `rwx` belongs to the mapping rather than to any pointer into it. Three things follow: - -- **Allocator metadata is protected only as far as the slice API.** No handle spans a block header, but an allocator hands out its stack, and `allocator.stack().set(..)` reaches any byte in the arena. The rule being broken — *only the allocator may write here* — is about the caller, not about aliasing. -- **Truncating cannot be separated from writing.** A tail region may be freely writable yet must not be discarded. -- **Reads cannot be denied.** Atomicity guarantees a read is never torn, not that the bytes are still *yours*: a slice held across a `dealloc` reads whatever the block was reused for. - -None of this is a correction — used as documented, the APIs keep a stack intact. Access control is an **additional layer** for callers who would rather have an invariant checked at runtime. - -### Design - -#### Modes - -```rust -pub enum BStackAccess { All, Rw, RwStrict, Prot, RwProt, Alloc, ReadOnly, Locked } -``` - -| Mode | Read | Write | Truncate | -|------------|-----------|-----------|--------------------| -| `All` | any | any | any | -| `Rw` | any | any | allocator or guard | -| `RwStrict` | any | any | none | -| `Prot` | guard | guard | guard | -| `RwProt` | guard | guard | none | -| `Alloc` | allocator | allocator | allocator | -| `ReadOnly` | any | none | none | -| `Locked` | none | none | none | - -Each cell lists the authorities that satisfy it; `any` means no token is needed. The two tokens are **incomparable** — neither outranks the other — so `Prot` and `Alloc` are each private to their own holder on all three axes: a range marked `Alloc` cannot be read, written, or truncated by a guard holder, which is what makes metadata inviolable even to the policy owner. `All` is the default everywhere and is what an unprotected stack reports. - -#### Authority - -Two capability tokens, `BStackProtection<'a>` and `BStackAllocAuthority<'a>`, neither `Clone` nor `Copy`, each minted at most once per handle (`take_protection() -> Option<_>`, `None` thereafter). One-shot minting is what makes them mean anything. Allocator constructors claim the second naturally, since they already consume a `BStack` exclusively. Checked entry points gain a token-carrying sibling, `set_as(&self, auth, offset, data)`. - -#### The point table - -A sorted `Vec<(u64, BStackAccess)>` of change points: `(16, Alloc)` means `Alloc` from offset 16 until the next point, with an absent leading point implying `All` from 0. Adjacent equal modes coalesce, so the table is proportional to the number of distinct regions and one protected header is two entries. Deliberately **not** a `BTreeMap`: the table is read far more often than written, a read is a `partition_point` over a contiguous array. - -- **Point lookup** — `partition_point(|p| p.0 <= off) - 1`. -- **Range check** for `[a, b)` — one `partition_point`, then a forward scan while `points[j].0 < b`, folding to the most restrictive mode. The common case spans one point and the scan does not run. -- **Setting** `[a, b)` to `M` — record the mode in effect at `b` as a point at `b`, insert `(a, M)`, drop points strictly inside, coalesce. - -The table takes its own `RwLock`, separate from the stack lock, because the locked-region read fast path bypasses the stack lock and must still reject a `Locked` read. Mutation takes both, stack lock first, so in-flight writers drain before the new policy is published — the ordering and the reasoning of `lock_up_to`. An `AtomicBool` short-circuits every check on a stack that has never been protected. - -#### Checks - -Writes check their target range, truncations `[new_len, old_len)`, reads their read range. Batched paths check every block before the journal is armed. Points beyond `len` are retained rather than trimmed, so a range can be armed before its bytes arrive. Denials return `PermissionDenied`. - -The locked prefix is checked first and stays out of the table, which can only further restrict it; folding the two would cost `lock_up_to` its lock-free read path. Protection is set through an owned handle — `BStackOwnedSlice::protect(mode)`, forwarded to the stack's table — never through `BStack` with an arbitrary range. A caller may set any range whose current mode already admits its token; one without a token may only tighten a range currently at `All`. Nothing is persisted, so reopening clears the table. - -#### Cost - -The flag is named for it. On a protected stack every checked call pays a relaxed load, an `RwLock` read acquisition, and a binary search before any I/O — a real fraction of a small `set`, whose fast path is one write and one sync. Batched ops pay per block, and every checked entry point grows a token-carrying sibling. - -### Open questions - -- **Named modes or an axis triple.** A `{ read, write, truncate }` triple of authorities is more expressive and no larger, at the cost of admitting nonsense (`read: none, write: any`). The enum is proposed because the curated eight are what callers want and a one-byte discriminant keeps the table compact. -- **What an allocator may do inside a `Prot` range.** Incomparability settles one direction — a guard holder cannot reach an `Alloc` range — but not the other. A caller may mark its own allocation `Prot` and then free it, leaving a mode over bytes the allocator is about to hand to someone else. Either `dealloc` resets the reclaimed range to `All`, which means an allocator overriding a mode it otherwise cannot touch, or the protection outlives the allocation and poisons the block for its next owner. - ---- - ## `BStackGuardedUnit` and `BStackGuardedBuilder` — composable transform units for `guarded` (0.5.0) **Feature flag:** `guarded`. diff --git a/src/acl.rs b/src/acl.rs new file mode 100644 index 0000000..f81393c --- /dev/null +++ b/src/acl.rs @@ -0,0 +1,3744 @@ +//! Range access control on [`BStack`] (the `expensive-slice-access-control` +//! feature): the capability tokens, the authority resolution, and the checked +//! policy entry points. +//! +//! The lock-free policy machinery — the mode enum, the point table, and its +//! lookup / range-check / set operations — lives in [`acl_core`](crate::acl_core). +//! This module bolts it onto a live stack: the two one-shot tokens, the check +//! that every guarded I/O path consults, and [`protect`](BStack::protect) / +//! [`protect_as`](BStack::protect_as) that arm ranges. +//! +//! The [`acl_check!`] macro lives at the module root (always compiled) so the +//! guarded I/O paths in `lib.rs` can invoke it unconditionally; it folds to +//! nothing without the feature. Everything else is in the gated `inner` module. + +/// Enforce the access-control policy over `[$a, $b)` for one axis on `$stack`, +/// then `?` on denial — expanding to nothing without the +/// `expensive-slice-access-control` feature. The op is a bare +/// [`AccessOp`](crate::AccessOp) variant (`Read`/`Write`/`Truncate`); authorities +/// default to [`NONE`](crate::BStackAccessAuthorities::NONE), overridable by a +/// trailing expression. Takes the stack explicitly (like [`fault_point!`], since +/// `self` cannot cross the macro's hygiene boundary). +#[allow(unused_macros)] +macro_rules! acl_check { + ($stack:expr, $a:expr, $b:expr, $op:ident $(,)?) => { + acl_check!($stack, $a, $b, $op, $crate::BStackAccessAuthorities::NONE) + }; + ($stack:expr, $a:expr, $b:expr, $op:ident, $held:expr $(,)?) => {{ + #[cfg(feature = "expensive-slice-access-control")] + $stack.acl_check($a, $b, $crate::AccessOp::$op, $held)?; + }}; +} +pub(crate) use acl_check; + +/// Dispatch one allocator metadata op through its held authority. +/// +/// `alloc_meta!(self, op, op_as, args...)` expands to +/// `self.stack.op_as(&self.alloc_auth, args...)` with the +/// `expensive-slice-access-control` feature — presenting the allocator's *held* +/// real [`BStackAllocAuthority`] to the token op, so its own permanently +/// `Alloc`-marked metadata is admitted — or to the plain `self.stack.op(args...)` +/// without it. There are no `meta_*` methods; this expands inline at the call +/// site, so an op whose call site is `atomic`-gated needs no gating here. +/// +/// The allocator (`self`) is passed explicitly — `self` does not cross +/// `macro_rules` hygiene, the same reason [`acl_check!`] takes its stack. The +/// allocator must have a `stack` field and, under the feature, an `alloc_auth` +/// field holding its token. It is pure dispatch: a caller that needs to decode +/// (e.g. a little-endian `u64`) reads into a buffer and decodes at the call site. +#[cfg(all(feature = "alloc", feature = "set"))] +macro_rules! alloc_meta { + ($self:expr, $op:ident, $op_as:ident $(, $arg:expr)* $(,)?) => {{ + #[cfg(feature = "expensive-slice-access-control")] + let __r = $self.stack.$op_as(&$self.alloc_auth $(, $arg)*); + #[cfg(not(feature = "expensive-slice-access-control"))] + let __r = $self.stack.$op($($arg),*); + __r + }}; +} +#[cfg(all(feature = "alloc", feature = "set"))] +pub(crate) use alloc_meta; + +#[cfg(feature = "expensive-slice-access-control")] +mod inner { + use crate::fault::fault_point; + use crate::io_core::{HEADER_SIZE, commit_shrink, repeat_fill, set_in_place}; + use crate::{ + AccessOp, BStack, BStackAccess, BStackAccessAuthorities, check_offset_unlocked, checked_end, + }; + use std::io::{self, Seek, SeekFrom}; + use std::sync::atomic::Ordering; + + #[cfg(unix)] + use crate::io_core::pread_exact_raw; + #[cfg(windows)] + use crate::io_core::pread_exact_raw_handle; + #[cfg(any(unix, windows))] + use crate::io_core::{pread_exact, pread_exact_into}; + #[cfg(not(any(unix, windows)))] + use std::io::Read; + + #[cfg(feature = "atomic")] + use crate::io_core::{ + durable_sync, is_atomic_write, journaled_copy, journaled_exchange, journaled_move, write_at, + }; + // Frontier/tail commit helpers shared by the append/shrink/read `_as` siblings; + // these mirror non-atomic `BStack` methods, so they cannot be `atomic`-gated. + #[cfg(feature = "atomic")] + use crate::BStackGenOp; + #[cfg(feature = "atomic")] + use crate::io_core::commit_tail_replace; + use crate::io_core::{commit_grow, commit_sparse_extend, read_at}; + use crate::validate_sparse_blocks; + use std::io::Write; + // Helpers for the duplicated `inplace_gen_as` engine. + #[cfg(feature = "atomic")] + use crate::fault::fault_probe; + #[cfg(feature = "atomic")] + use crate::io_core::{ + OverlayData, inplace_overlay_insert, inplace_overlay_read, inplace_validate_read, + inplace_validate_repeat, inplace_validate_write, journaled_multi_overlay, + journaled_multi_set, + }; + + /// A one-shot capability token authorizing guard-level access to a stack's + /// protected ranges. + /// + /// Minted at most once per stack via [`BStack::take_protection`] (there is only + /// ever one guard authority per stack); neither `Clone` nor `Copy`, so it + /// cannot be duplicated. It is owned (it records its origin stack's identity + /// rather than borrowing it), so a holder may store it — e.g. a wrapper that + /// owns the stack by value — and hand it back with + /// [`BStack::return_protection`]. Present it to a `*_as` entry point to act on a + /// [`Prot`](BStackAccess::Prot)/[`RwProt`](BStackAccess::RwProt) range or to + /// re-arm a range it governs. Incomparable with [`BStackAllocAuthority`]: + /// neither reaches the other's private ranges. + pub struct BStackProtection { + origin: u64, + } + + /// A one-shot capability token authorizing allocator-level access to a stack's + /// [`Alloc`](BStackAccess::Alloc) ranges. + /// + /// The allocator-axis counterpart of [`BStackProtection`], with the same + /// one-per-stack owned move-out shape (mint via [`BStack::take_alloc_authority`], + /// hand back via [`BStack::return_alloc_authority`]); neither `Clone` nor + /// `Copy`. Incomparable with [`BStackProtection`]. + pub struct BStackAllocAuthority { + origin: u64, + } + + /// A presented access token, resolved to the authorities it carries for the + /// stack it was minted from. Implemented for references to the two token types + /// and for `()` (no token). + pub trait BStackAuthority { + /// The authorities this token grants when acting on `stack`. A token minted + /// from a *different* stack grants nothing. + fn authorities_for(&self, stack: &BStack) -> BStackAccessAuthorities; + } + + impl BStackAuthority for () { + #[inline] + fn authorities_for(&self, _stack: &BStack) -> BStackAccessAuthorities { + BStackAccessAuthorities::NONE + } + } + + // Resolved authorities carry themselves — the form a slice stores after a + // token grant, so its I/O can present the authority it was given. + impl BStackAuthority for BStackAccessAuthorities { + #[inline] + fn authorities_for(&self, _stack: &BStack) -> BStackAccessAuthorities { + *self + } + } + + impl BStackAuthority for &BStackProtection { + #[inline] + fn authorities_for(&self, stack: &BStack) -> BStackAccessAuthorities { + // The token records its origin stack's identity, so this rejects a + // token minted from any other stack. + if self.origin == stack.acl_identity() { + BStackAccessAuthorities::GUARD + } else { + BStackAccessAuthorities::NONE + } + } + } + + impl BStackAuthority for &BStackAllocAuthority { + #[inline] + fn authorities_for(&self, stack: &BStack) -> BStackAccessAuthorities { + if self.origin == stack.acl_identity() { + BStackAccessAuthorities::ALLOC + } else { + BStackAccessAuthorities::NONE + } + } + } + + impl BStack { + /// This stack's identity, recorded in a minted token and re-checked when it + /// is presented, so a token only ever authorizes the stack it came from. + /// The fd (Unix) / handle (Windows) — unique among live stacks, matching + /// the [`Hash`]/[`PartialEq`] identity — or the instance address elsewhere. + #[inline] + pub(crate) fn acl_identity(&self) -> u64 { + #[cfg(unix)] + { + self.fd as u64 + } + #[cfg(windows)] + { + self.handle as u64 + } + #[cfg(not(any(unix, windows)))] + { + self as *const BStack as u64 + } + } + + /// Move out the guard capability [token](BStackProtection), or `None` if it + /// has already been taken (until [`return_protection`](Self::return_protection) + /// hands it back). The one-shot move-out is what makes the token mean + /// anything. + #[inline] + pub fn take_protection(&self) -> Option { + self.protection + .lock() + .unwrap() + .take() + .map(|()| BStackProtection { + origin: self.acl_identity(), + }) + } + + /// Hand a guard token back to the stack it came from, so a later + /// [`take_protection`](Self::take_protection) can mint again. A token from a + /// different stack is dropped without re-arming this one. + #[inline] + pub fn return_protection(&self, token: BStackProtection) { + if token.origin == self.acl_identity() { + *self.protection.lock().unwrap() = Some(()); + } + } + + /// Move out the allocator capability [token](BStackAllocAuthority), or + /// `None` if it has already been taken. See + /// [`take_protection`](Self::take_protection). + #[inline] + pub fn take_alloc_authority(&self) -> Option { + self.alloc_authority + .lock() + .unwrap() + .take() + .map(|()| BStackAllocAuthority { + origin: self.acl_identity(), + }) + } + + /// Hand an allocator token back to the stack it came from. See + /// [`return_protection`](Self::return_protection). + #[inline] + pub fn return_alloc_authority(&self, token: BStackAllocAuthority) { + if token.origin == self.acl_identity() { + *self.alloc_authority.lock().unwrap() = Some(()); + } + } + + /// Check that `[a, b)` permits `op` under the authorities `held`, returning + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) otherwise. An + /// unprotected stack has an empty table, so the check is a cheap miss. The + /// `acl` lock is separate from the stack lock, so this works + /// on the lock-free read fast path too. + pub(crate) fn acl_check( + &self, + a: u64, + b: u64, + op: AccessOp, + held: BStackAccessAuthorities, + ) -> io::Result<()> { + let table = self.acl.read().unwrap(); + if table.check(a, b, op, held) { + Ok(()) + } else { + Err(io_error!( + PermissionDenied, + format!("{op:?} on [{a}, {b}) denied by access control") + )) + } + } + + /// [`process_gen`](BStack::process_gen) run under an access token: every + /// per-op check is evaluated as if `auth` were presented, so an allocator + /// can drive its own [`Alloc`](BStackAccess)-marked metadata (the free-list + /// head, say) through a journalled sequence without ever lifting the mark — + /// leaving no window in which a concurrent op could see it unprotected. + /// + /// Present [`ALLOC`](BStackAccessAuthorities::ALLOC) or an alloc token. The + /// body duplicates `process_gen` (kept in `lib.rs` as the tokenless engine) + /// rather than threading authority through the crash-atomic core; the two + /// must stay in sync. + #[cfg(feature = "atomic")] + pub fn process_gen_as<'a, A, F>(&self, auth: A, mut f: F) -> io::Result<()> + where + A: super::BStackAuthority, + F: FnMut() -> Option>, + { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + let locked = self.locked.load(Ordering::Acquire); + fault_point!(self, "process_gen"); + loop { + match f() { + Some(BStackGenOp::Read { offset, buf }) => { + let end = checked_end( + offset, + buf.len() as u64, + "process_gen: read offset + buf.len() overflows u64", + )?; + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: read range [{offset}, {end}) exceeds payload size ({data_size})" + ) + )); + } + // Per-step read fault: stands in for this `Read`'s I/O, and + // like a genuine read failure here it ends the whole call — + // `process_gen` has no channel to report a step failure on. + // Consulted for every `Read` op, including ones the fast + // paths below serve without touching the disk, so the + // schedule does not shift with cache state. + fault_point!(self, "process_gen:read"); + acl_check!(self, offset, end, Read, held); + // Fast path: locked bytes are immutable, so they can be + // served from the cache or via a lock-free pread instead + // of going through the held file handle — mirroring how + // `get_into` treats reads of the locked region. + #[cfg(any(unix, windows))] + { + if end <= locked { + if self.cache_enabled { + let cache = self.cache.lock().unwrap(); + buf.copy_from_slice(&cache[offset as usize..end as usize]); + } else { + #[cfg(unix)] + pread_exact_raw(self.fd, HEADER_SIZE + offset, buf)?; + #[cfg(windows)] + pread_exact_raw_handle(self.handle, HEADER_SIZE + offset, buf)?; + } + } else { + pread_exact_into(file, HEADER_SIZE + offset, buf)?; + } + } + #[cfg(not(any(unix, windows)))] + { + if end <= locked && self.cache_enabled { + let cache = self.cache.lock().unwrap(); + buf.copy_from_slice(&cache[offset as usize..end as usize]); + } else { + read_at(file, offset, buf)?; + } + } + } + Some(BStackGenOp::Write { offset, data }) => { + let end = checked_end( + offset, + data.len() as u64, + "process_gen: write offset + data.len() overflows u64", + )?; + if offset < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: write range [{offset}, {end}) overlaps locked region [0, {locked})" + ) + )); + } + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: write range [{offset}, {end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, offset, end, Write, held); + if !data.is_empty() { + Self::mark_replay(replay, set_in_place(file, data_size, offset, data))?; + } + return Ok(()); + } + Some(BStackGenOp::Repeat { + offset, + pattern, + count, + }) => { + // Empty pattern or zero count is a no-op, matching `zero(_, 0)`. + if pattern.is_empty() || count == 0 { + return Ok(()); + } + let total = (pattern.len() as u64).checked_mul(count).ok_or_else(|| { + io_error!(InvalidInput, "process_gen: repeat length overflows u64") + })?; + let end = checked_end( + offset, + total, + "process_gen: repeat offset + length overflows u64", + )?; + if offset < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: repeat range [{offset}, {end}) overlaps locked region [0, {locked})" + ) + )); + } + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: repeat range [{offset}, {end}) exceeds payload size ({data_size})" + ) + )); + } + Self::mark_replay( + replay, + repeat_fill(file, data_size, offset, pattern, count), + )?; + return Ok(()); + } + Some(BStackGenOp::Swap { + a_offset, + b_offset, + len, + }) => { + let a_end = checked_end( + a_offset, + len, + "process_gen: a_offset + len overflows u64", + )?; + let b_end = checked_end( + b_offset, + len, + "process_gen: b_offset + len overflows u64", + )?; + if len > 0 { + let (lo, hi) = if a_offset < b_offset { + (a_offset, b_offset) + } else { + (b_offset, a_offset) + }; + if lo + len > hi { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: swap regions [{a_offset}, {a_end}) and [{b_offset}, {b_end}) overlap" + ) + )); + } + } + if a_offset < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: swap region [{a_offset}, {a_end}) overlaps locked region [0, {locked})" + ) + )); + } + if b_offset < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: swap region [{b_offset}, {b_end}) overlaps locked region [0, {locked})" + ) + )); + } + if a_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: swap region [{a_offset}, {a_end}) exceeds payload size ({data_size})" + ) + )); + } + if b_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: swap region [{b_offset}, {b_end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, a_offset, a_end, Write, held); + acl_check!(self, b_offset, b_end, Write, held); + if len > 0 { + Self::mark_replay( + replay, + journaled_exchange(file, data_size, a_offset, b_offset, len), + )?; + } + return Ok(()); + } + Some(BStackGenOp::Push { data }) => { + if !data.is_empty() { + let file_end = file.seek(SeekFrom::End(0))?; + let logical_offset = file_end - HEADER_SIZE; + acl_check!( + self, + logical_offset, + logical_offset + data.len() as u64, + Write, + held + ); + if let Err(e) = file.write_all(data) { + // A failed rollback leaves a stale tail past the committed length: + // defer it to the next write's replay. + if file.set_len(file_end).is_err() { + *replay = true; + } + return Err(e); + } + let new_len = logical_offset + data.len() as u64; + Self::mark_replay( + replay, + commit_grow(file, clen, new_len, logical_offset, file_end), + )?; + } + return Ok(()); + } + Some(BStackGenOp::Pop { buf }) => { + let n = buf.len() as u64; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("process_gen: pop({n}) exceeds payload size ({data_size})") + )); + } + let new_data_len = data_size - n; + if new_data_len < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: pop({n}) would shrink payload below locked length ({locked})" + ) + )); + } + acl_check!(self, new_data_len, data_size, Truncate, held); + if n > 0 { + read_at(file, new_data_len, buf)?; + Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; + } + return Ok(()); + } + Some(BStackGenOp::Discard { len }) => { + if len > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: discard({len}) exceeds payload size ({data_size})" + ) + )); + } + let new_data_len = data_size - len; + if new_data_len < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: discard({len}) would shrink payload below locked length ({locked})" + ) + )); + } + acl_check!(self, new_data_len, data_size, Truncate, held); + if len > 0 { + Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; + } + return Ok(()); + } + Some(BStackGenOp::Atrunc { n, data }) => { + if n > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: atrunc n ({n}) exceeds payload size ({data_size})" + ) + )); + } + let new_tail_start = data_size - n; + if new_tail_start < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: atrunc would modify locked region [0, {locked})" + ) + )); + } + acl_check!(self, new_tail_start, data_size, Truncate, held); + if n != 0 || !data.is_empty() { + let file_end = HEADER_SIZE + data_size; + Self::mark_replay( + replay, + commit_tail_replace(file, clen, new_tail_start, n, data, file_end), + )?; + } + return Ok(()); + } + Some(BStackGenOp::Splice { old, new }) => { + let n = old.len() as u64; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: splice n ({n}) exceeds payload size ({data_size})" + ) + )); + } + let new_tail_start = data_size - n; + if new_tail_start < locked { + return Err(io_error!( + InvalidInput, + format!( + "process_gen: splice would modify locked region [0, {locked})" + ) + )); + } + acl_check!(self, new_tail_start, data_size, Truncate, held); + if n != 0 || !new.is_empty() { + // Read the removed bytes before any mutation. + read_at(file, new_tail_start, old)?; + let file_end = HEADER_SIZE + data_size; + Self::mark_replay( + replay, + commit_tail_replace(file, clen, new_tail_start, n, new, file_end), + )?; + } + return Ok(()); + } + Some(BStackGenOp::Sparse { writes, length }) => { + // Copy the borrowed blocks into a local list so they can be + // filtered and sorted for validation (the source slice is `&'a`). + let mut blocks: Vec<(u64, &[u8])> = writes + .iter() + .map(|(off, d)| (*off, *d)) + .filter(|(_, d)| !d.is_empty()) + .collect(); + validate_sparse_blocks(&mut blocks, length, "process_gen: sparse")?; + if length != 0 { + let file_end = HEADER_SIZE + data_size; + let new_len = checked_end( + data_size, + length, + "process_gen: sparse data_size + length overflows u64", + )?; + acl_check!(self, data_size, new_len, Write, held); + Self::mark_replay( + replay, + commit_sparse_extend( + file, clen, data_size, file_end, new_len, &blocks, + ), + )?; + } + return Ok(()); + } + Some(BStackGenOp::Len { out }) => { + *out = data_size; + } + Some(BStackGenOp::Abort { source }) => { + // Nothing has been mutated: every mutating op ends the + // sequence, so reaching here means only reads have run. + return source.map_or(Ok(()), Err); + } + None => return Ok(()), + } + } + } + + /// [`inplace_gen`](BStack::inplace_gen) run under an access token — the + /// generator counterpart to [`process_gen_as`](Self::process_gen_as), for + /// the overlay-resolving in-place engine. Duplicates `inplace_gen` (kept in + /// `lib.rs`) rather than threading authority through it; the two must stay + /// in sync. + #[cfg(feature = "atomic")] + pub fn inplace_gen_as<'a, A, F>(&self, auth: A, mut f: F) -> io::Result<()> + where + A: super::BStackAuthority, + F: FnMut(io::Result<()>) -> Option>, + { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + let locked = self.locked.load(Ordering::Acquire); + fault_point!(self, "inplace_gen"); + // Sorted, pairwise-non-overlapping set of pending in-place edits, each + // borrowing the caller's `Write` data (or `Repeat` pattern) for the + // lifetime of the call. + let mut overlay: Vec<(u64, OverlayData<'a>)> = Vec::new(); + let mut feedback: io::Result<()> = Ok(()); + loop { + match f(feedback) { + Some(BStackGenOp::Read { offset, buf }) => { + // Validate first so a bad range still beats an injected + // fault, then let the policy stand in for the read itself. + // Unlike `process_gen`, a failed `Read` here does not end + // the call: it is reported to the generator through its + // `feedback` argument, so an injected fault must take the + // same route as a genuine one. + feedback = match inplace_validate_read(offset, buf.len() as u64, data_size) + { + Err(e) => Err(e), + Ok(()) => { + // A denied read is reported through `feedback`, the same + // route a genuine read failure takes. Validation passed, + // so `offset + len` cannot overflow. + #[cfg(feature = "expensive-slice-access-control")] + let gate = self.acl_check( + offset, + offset + buf.len() as u64, + AccessOp::Read, + held, + ); + #[cfg(not(feature = "expensive-slice-access-control"))] + let gate: io::Result<()> = Ok(()); + gate.and_then(|()| { + fault_probe!(self, "inplace_gen:read").map_or_else( + || { + inplace_overlay_read( + file, data_size, offset, buf, &overlay, + ) + }, + Err, + ) + }) + } + }; + } + Some(BStackGenOp::Write { offset, data }) => { + feedback = inplace_validate_write(offset, data, data_size, locked); + // A denial is routed to the generator like a validation error, + // not returned from the call. `is_ok` implies the range already + // passed `inplace_validate_write`'s `checked_end`, so `offset + + // len` cannot overflow here. + #[cfg(feature = "expensive-slice-access-control")] + if feedback.is_ok() { + let end = offset + data.len() as u64; + feedback = self.acl_check(offset, end, AccessOp::Write, held); + } + if feedback.is_ok() && !data.is_empty() { + inplace_overlay_insert( + &mut overlay, + offset, + OverlayData::Literal(data), + ); + } + } + Some(BStackGenOp::Repeat { + offset, + pattern, + count, + }) => { + feedback = + inplace_validate_repeat(offset, pattern, count, data_size, locked); + if feedback.is_ok() && !pattern.is_empty() && count > 0 { + // Non-overflowing after validation. + let len = pattern.len() as u64 * count; + inplace_overlay_insert( + &mut overlay, + offset, + OverlayData::Repeat { + pattern, + phase: 0, + len, + }, + ); + } + } + Some(BStackGenOp::Len { out }) => { + *out = data_size; + feedback = Ok(()); + } + Some(BStackGenOp::Abort { source }) => { + // Drop the overlay without committing: the pending writes + // only ever existed in memory, so the file is untouched. + return source.map_or(Ok(()), Err); + } + Some(BStackGenOp::Swap { .. }) => { + feedback = Err(io_error!( + InvalidInput, + "inplace_gen: Swap is not permitted (Read/Write/Len only)" + )); + } + Some(BStackGenOp::Push { .. }) => { + feedback = Err(io_error!( + InvalidInput, + "inplace_gen: Push is not permitted (in-place writes only)" + )); + } + Some(BStackGenOp::Pop { .. }) => { + feedback = Err(io_error!( + InvalidInput, + "inplace_gen: Pop is not permitted (in-place writes only)" + )); + } + Some(BStackGenOp::Discard { .. }) => { + feedback = Err(io_error!( + InvalidInput, + "inplace_gen: Discard is not permitted (in-place writes only)" + )); + } + Some(BStackGenOp::Atrunc { .. }) => { + feedback = Err(io_error!( + InvalidInput, + "inplace_gen: Atrunc is not permitted (in-place writes only)" + )); + } + Some(BStackGenOp::Splice { .. }) => { + feedback = Err(io_error!( + InvalidInput, + "inplace_gen: Splice is not permitted (in-place writes only)" + )); + } + Some(BStackGenOp::Sparse { .. }) => { + feedback = Err(io_error!( + InvalidInput, + "inplace_gen: Sparse is not permitted (in-place writes only)" + )); + } + None => break, + } + } + // Commit the accumulated edits. Zero → nothing to do; a lone literal takes + // the ordinary single-write path and a lone repeat the compact repeat-fill + // journal; several edits go through the multi-write journal (which streams + // any repeat block rather than materialising it). + match overlay.len() { + 0 => Ok(()), + 1 => match overlay[0] { + (offset, OverlayData::Literal(data)) => { + Self::mark_replay(replay, set_in_place(file, data_size, offset, data)) + } + // A lone repeat is never sliced (slicing needs an overlapping edit, + // which would leave it non-lone), so `phase == 0` and `len` is a + // whole number of periods — exactly what `repeat_fill` expects. + ( + offset, + OverlayData::Repeat { + pattern, + phase: 0, + len, + }, + ) if len % pattern.len() as u64 == 0 => Self::mark_replay( + replay, + repeat_fill(file, data_size, offset, pattern, len / pattern.len() as u64), + ), + _ => Self::mark_replay( + replay, + journaled_multi_overlay(file, data_size, &overlay), + ), + }, + _ => Self::mark_replay(replay, journaled_multi_overlay(file, data_size, &overlay)), + } + } + + /// Arm `[offset, offset + len)` with `mode`, presenting `auth`. Crate-internal. + /// + /// All public protection goes through + /// [`BStackOwnedSlice::protect`](crate::BStackOwnedSlice::protect), which + /// bounds the range to a genuine allocation rather than an arbitrary span. + /// Nothing is persisted, so reopening clears the policy. + /// + /// A caller may re-mode any range whose + /// current mode its token can already write (which, by the incomparability + /// of the two tokens, keeps a guard out of [`Alloc`](BStackAccess::Alloc) + /// ranges and an allocator out of [`Prot`](BStackAccess::Prot) ranges). A + /// tokenless `auth` of `()` falls back to the tighten-`All`-only rule. + /// + /// # Errors + /// + /// [`InvalidInput`](io::ErrorKind::InvalidInput) if `offset + len` overflows; + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the current policy + /// does not admit the token over the whole range. + pub(crate) fn protect_as( + &self, + auth: impl BStackAuthority, + offset: u64, + len: u64, + mode: BStackAccess, + ) -> io::Result<()> { + if len == 0 { + return Ok(()); + } + let end = checked_end(offset, len, "protect: offset + len overflows u64")?; + let held = auth.authorities_for(self); + // Stack lock first, then the acl lock, so in-flight writers drain before + // the new policy is published (the ordering of `lock_up_to`). + let _guard = self.write_lock()?; + let mut table = self.acl.write().unwrap(); + let admitted = if held == BStackAccessAuthorities::NONE { + table.all_over(offset, end) + } else { + table.check(offset, end, AccessOp::Write, held) + }; + if !admitted { + return Err(io_error!( + PermissionDenied, + format!("protect: [{offset}, {end}) not admitted by current policy") + )); + } + table.set(offset, end, mode); + Ok(()) + } + + /// Mark `[offset, offset + len)` as allocator-owned + /// [`Alloc`](BStackAccess::Alloc) metadata. + /// + /// Presents synthetic [`ALLOC`](BStackAccessAuthorities::ALLOC) authority: + /// the crate-internal caller *is* the allocator, so it needs no minted + /// token to arm its own metadata. Re-marking a range already `Alloc` (a + /// reopened arena, a reused block) is admitted. + /// + // The metadata-marking hook. Marking a region `Alloc` also requires the + // allocator to route its own reads/writes of that region through the + // `_as(ALLOC)` siblings, so per-allocator adoption is staged separately + // from the (universal) mint-burn and dealloc-reclaim wiring. + #[allow(dead_code)] + pub(crate) fn acl_mark_alloc(&self, offset: u64, len: u64) -> io::Result<()> { + self.protect_as( + BStackAccessAuthorities::ALLOC, + offset, + len, + BStackAccess::Alloc, + ) + } + + /// Whether `[offset, offset + len)` may be reclaimed on `dealloc` — i.e. + /// carries no caller-set policy (nothing but + /// [`All`](BStackAccess::All)/[`Alloc`](BStackAccess::Alloc)). + /// + /// The read-only half of [`acl_reclaim`](Self::acl_reclaim), split out so a + /// bulk free can validate every handle *before* clearing any — keeping the + /// batch atomic. Returns [`PermissionDenied`](io::ErrorKind::PermissionDenied) + /// otherwise. + pub(crate) fn acl_reclaimable(&self, offset: u64, len: u64) -> io::Result<()> { + if len == 0 { + return Ok(()); + } + let end = checked_end(offset, len, "acl_reclaimable: offset + len overflows u64")?; + if self.acl.read().unwrap().reclaimable_by_alloc(offset, end) { + Ok(()) + } else { + Err(io_error!( + PermissionDenied, + format!( + "dealloc: [{offset}, {end}) carries access-control policy; \ + unprotect it before freeing" + ) + )) + } + } + + /// Reclaim `[offset, offset + len)` on `dealloc`: refuse (see + /// [`acl_reclaimable`](Self::acl_reclaimable)) if it carries caller-set + /// policy, so a stated policy is never silently dropped nor left to poison + /// the block's next owner; otherwise reset the range to + /// [`All`](BStackAccess::All), clearing the allocator's own + /// [`Alloc`](BStackAccess::Alloc) metadata marks over it. + pub(crate) fn acl_reclaim(&self, offset: u64, len: u64) -> io::Result<()> { + if len == 0 { + return Ok(()); + } + let end = checked_end(offset, len, "acl_reclaim: offset + len overflows u64")?; + // Stack lock first, then the acl lock, as in `protect_as`. + let _guard = self.write_lock()?; + let mut table = self.acl.write().unwrap(); + if !table.reclaimable_by_alloc(offset, end) { + return Err(io_error!( + PermissionDenied, + format!( + "dealloc: [{offset}, {end}) carries access-control policy; \ + unprotect it before freeing" + ) + )); + } + table.set(offset, end, BStackAccess::All); + Ok(()) + } + + /// The mode currently governing logical `offset` (for inspection/testing). + #[inline] + #[must_use] + pub fn access_at(&self, offset: u64) -> BStackAccess { + self.acl.read().unwrap().mode_at(offset) + } + + /// [`set`](BStack::set) presenting an access token: writing a + /// [`Prot`](BStackAccess::Prot)/[`Alloc`](BStackAccess::Alloc) range requires + /// the matching capability, checked against the range's mode before any I/O. + /// + /// The body mirrors [`set`](BStack::set) with the token check spliced in after + /// the locked-prefix check. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's access + /// mode denies the write under `auth`, plus every error [`set`](BStack::set) + /// itself can return. + pub fn set_as( + &self, + auth: impl BStackAuthority, + offset: u64, + data: impl AsRef<[u8]>, + ) -> io::Result<()> { + let data = data.as_ref(); + if data.is_empty() { + return Ok(()); + } + let held = auth.authorities_for(self); + let end = checked_end(offset, data.len() as u64, "set: offset + len overflows u64")?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + check_offset_unlocked("set", offset, end, locked)?; + acl_check!(self, offset, end, Write, held); + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("set: write end ({end}) exceeds payload size ({data_size})") + )); + } + fault_point!(self, "set"); + Self::mark_replay(replay, set_in_place(file, data_size, offset, data)) + } + + /// [`discard`](BStack::discard) presenting an access token: truncating a range + /// whose mode restricts the truncate axis requires the matching capability, + /// checked over the discarded tail `[new_len, old_len)` before any I/O. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the tail's access + /// mode denies truncation under `auth`, plus every error + /// [`discard`](BStack::discard) itself can return. + pub fn discard_as(&self, auth: impl BStackAuthority, n: u64) -> io::Result<()> { + if n == 0 { + return Ok(()); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let raw_size = file.seek(SeekFrom::End(0))?; + let data_size = raw_size - HEADER_SIZE; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("discard({n}) exceeds payload size ({data_size})") + )); + } + let new_data_len = data_size - n; + let locked = self.locked.load(Ordering::Acquire); + if new_data_len < locked { + return Err(io_error!( + InvalidInput, + format!("discard({n}) would shrink payload below locked length ({locked})") + )); + } + acl_check!(self, new_data_len, data_size, Truncate, held); + fault_point!(self, "discard"); + Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; + Ok(()) + } + + /// [`get`](BStack::get) presenting an access token: reading a + /// [`Prot`](BStackAccess::Prot)/[`Alloc`](BStackAccess::Alloc) range (or any + /// range that denies tokenless reads) requires the matching capability, + /// checked over `[start, end)` before any I/O — ahead of the locked-region + /// fast path, so a [`Locked`](BStackAccess::Locked) range is still denied. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's access + /// mode denies the read under `auth`, plus every error [`get`](BStack::get) + /// itself can return. + pub fn get_as( + &self, + auth: impl BStackAuthority, + start: u64, + end: u64, + ) -> io::Result> { + acl_check!(self, start, end, Read, auth.authorities_for(self)); + if end < start { + return Err(io_error!( + InvalidInput, + format!("get: end ({end}) < start ({start})") + )); + } + // Fast-path: if the range lies entirely within the locked region, serve + // from the in-memory cache (if enabled) or a lock-free pread. + #[cfg(any(unix, windows))] + { + let locked = self.locked.load(Ordering::Acquire); + if end <= locked { + if self.cache_enabled { + let len = (end - start) as usize; + let mut buf = vec![0u8; len]; + let cache = self.cache.lock().unwrap(); + buf.copy_from_slice(&cache[start as usize..end as usize]); + return Ok(buf); + } + #[cfg(unix)] + { + let mut buf = vec![0u8; (end - start) as usize]; + pread_exact_raw(self.fd, HEADER_SIZE + start, &mut buf)?; + return Ok(buf); + } + #[cfg(windows)] + { + let mut buf = vec![0u8; (end - start) as usize]; + pread_exact_raw_handle(self.handle, HEADER_SIZE + start, &mut buf)?; + return Ok(buf); + } + } + } + #[cfg(any(unix, windows))] + { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("get: end ({end}) exceeds payload size ({data_size})") + )); + } + fault_point!(self, "get"); + pread_exact(file, HEADER_SIZE + start, (end - start) as usize) + } + #[cfg(not(any(unix, windows)))] + { + let locked = self.locked.load(Ordering::Acquire); + if end <= locked && self.cache_enabled { + let cache = self.cache.lock().unwrap(); + return Ok(cache[start as usize..end as usize].to_vec()); + } + let mut guard = self.write_lock_read()?; + let file = &mut guard.0; + let raw_size = file.seek(SeekFrom::End(0))?; + let data_size = raw_size.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("get: end ({end}) exceeds payload size ({data_size})") + )); + } + fault_point!(self, "get"); + file.seek(SeekFrom::Start(HEADER_SIZE + start))?; + let mut buf = vec![0u8; (end - start) as usize]; + file.read_exact(&mut buf)?; + Ok(buf) + } + } + + /// [`get_into`](BStack::get_into) presenting an access token, mirroring the + /// tokenless body but checking with `auth`. + pub fn get_into_as( + &self, + auth: impl BStackAuthority, + start: u64, + buf: &mut [u8], + ) -> io::Result<()> { + if buf.is_empty() { + return Ok(()); + } + let len = buf.len() as u64; + let end = start + .checked_add(len) + .ok_or_else(|| io_error!(InvalidInput, "get_into: start + len overflows u64"))?; + acl_check!(self, start, end, Read, auth.authorities_for(self)); + #[cfg(any(unix, windows))] + { + let locked = self.locked.load(Ordering::Acquire); + if end <= locked { + if self.cache_enabled { + let cache = self.cache.lock().unwrap(); + buf.copy_from_slice(&cache[start as usize..end as usize]); + return Ok(()); + } + #[cfg(unix)] + return pread_exact_raw(self.fd, HEADER_SIZE + start, buf); + #[cfg(windows)] + return pread_exact_raw_handle(self.handle, HEADER_SIZE + start, buf); + } + } + #[cfg(any(unix, windows))] + { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("get_into: end ({end}) exceeds payload size ({data_size})") + )); + } + fault_point!(self, "get_into"); + pread_exact_into(file, HEADER_SIZE + start, buf) + } + #[cfg(not(any(unix, windows)))] + { + let locked = self.locked.load(Ordering::Acquire); + if end <= locked && self.cache_enabled { + let cache = self.cache.lock().unwrap(); + buf.copy_from_slice(&cache[start as usize..end as usize]); + return Ok(()); + } + let mut guard = self.write_lock_read()?; + let file = &mut guard.0; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("get_into: end ({end}) exceeds payload size ({data_size})") + )); + } + fault_point!(self, "get_into"); + file.seek(SeekFrom::Start(HEADER_SIZE + start))?; + file.read_exact(buf) + } + } + + /// [`zero`](BStack::zero) presenting an access token: zeroing a + /// [`Prot`](BStackAccess::Prot)/[`Alloc`](BStackAccess::Alloc) range requires the + /// matching capability, checked before any I/O. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's mode denies + /// the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `offset + n` overflows `u64` or the range exceeds the payload size; plus any I/O + /// error from the write. + pub fn zero_as(&self, auth: impl BStackAuthority, offset: u64, n: u64) -> io::Result<()> { + if n == 0 { + return Ok(()); + } + let held = auth.authorities_for(self); + let end = checked_end(offset, n, "zero: offset + n overflows u64")?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + check_offset_unlocked("zero", offset, end, locked)?; + acl_check!(self, offset, end, Write, held); + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("zero: write end ({end}) exceeds payload size ({data_size})") + )); + } + fault_point!(self, "zero"); + Self::mark_replay(replay, repeat_fill(file, data_size, offset, &[0u8], n)) + } + + /// [`repeat`](BStack::repeat) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's mode denies + /// the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if the + /// pattern is empty with a non-zero count, or the filled range overflows `u64` or + /// exceeds the payload size; plus any I/O error from the write. + pub fn repeat_as( + &self, + auth: impl BStackAuthority, + offset: u64, + pattern: impl AsRef<[u8]>, + count: u64, + ) -> io::Result<()> { + let pattern = pattern.as_ref(); + if pattern.is_empty() || count == 0 { + return Ok(()); + } + let held = auth.authorities_for(self); + let total = (pattern.len() as u64).checked_mul(count).ok_or_else(|| { + io_error!(InvalidInput, "repeat: count * pattern.len() overflows u64") + })?; + let end = checked_end( + offset, + total, + "repeat: offset + count*pattern.len() overflows u64", + )?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + check_offset_unlocked("repeat", offset, end, locked)?; + acl_check!(self, offset, end, Write, held); + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("repeat: write end ({end}) exceeds payload size ({data_size})") + )); + } + fault_point!(self, "repeat"); + Self::mark_replay(replay, repeat_fill(file, data_size, offset, pattern, count)) + } + + /// [`cas`](BStack::cas) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's mode denies + /// the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `offset + old.len()` overflows `u64` or exceeds the payload size; plus any I/O + /// error. A length mismatch or a failed compare returns `Ok(false)`, not an error. + #[cfg(feature = "atomic")] + pub fn cas_as( + &self, + auth: impl BStackAuthority, + offset: u64, + old: impl AsRef<[u8]>, + new: impl AsRef<[u8]>, + ) -> io::Result { + let old = old.as_ref(); + let new = new.as_ref(); + if old.len() != new.len() { + return Ok(false); + } + if old.is_empty() { + return Ok(true); + } + let held = auth.authorities_for(self); + let end = checked_end(offset, old.len() as u64, "cas: offset + len overflows u64")?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + check_offset_unlocked("cas", offset, end, locked)?; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("cas: range [{offset}, {end}) exceeds payload size ({data_size})") + )); + } + acl_check!(self, offset, end, Write, held); + fault_point!(self, "cas"); + let mut current = vec![0u8; old.len()]; + read_at(file, offset, &mut current)?; + if current != old { + return Ok(false); + } + Self::mark_replay(replay, set_in_place(file, data_size, offset, new))?; + Ok(true) + } + + /// [`cross_exchange`](BStack::cross_exchange) presenting an access token: both + /// regions are checked for write under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if either region's mode + /// denies the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a + /// region overflows `u64`, exceeds the payload size, overlaps the other, or overlaps + /// the locked prefix; plus any I/O error. + #[cfg(feature = "atomic")] + pub fn cross_exchange_as( + &self, + auth: impl BStackAuthority, + a: u64, + b: u64, + n: u64, + ) -> io::Result<()> { + let held = auth.authorities_for(self); + let a_end = checked_end(a, n, "cross_exchange: a + n overflows u64")?; + let b_end = checked_end(b, n, "cross_exchange: b + n overflows u64")?; + if n > 0 { + let (lo, hi) = if a < b { (a, b) } else { (b, a) }; + if lo + n > hi { + return Err(io_error!( + InvalidInput, + format!( + "cross_exchange: regions [{a}, {a_end}) and [{b}, {b_end}) overlap" + ) + )); + } + } + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + if a < locked { + return Err(io_error!( + InvalidInput, + format!( + "cross_exchange: region [{a}, {a_end}) overlaps locked region [0, {locked})" + ) + )); + } + if b < locked { + return Err(io_error!( + InvalidInput, + format!( + "cross_exchange: region [{b}, {b_end}) overlaps locked region [0, {locked})" + ) + )); + } + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if a_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "cross_exchange: region [{a}, {a_end}) exceeds payload size ({data_size})" + ) + )); + } + if b_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "cross_exchange: region [{b}, {b_end}) exceeds payload size ({data_size})" + ) + )); + } + if n == 0 { + return Ok(()); + } + acl_check!(self, a, a_end, Write, held); + acl_check!(self, b, b_end, Write, held); + fault_point!(self, "cross_exchange"); + Self::mark_replay(replay, journaled_exchange(file, data_size, a, b, n)) + } + + /// [`copy`](BStack::copy) presenting an access token: the source range is checked + /// for read and the destination for write under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the source denies the + /// read or the destination denies the write under `auth`; + /// [`InvalidInput`](io::ErrorKind::InvalidInput) if either range overflows `u64`, + /// exceeds the payload size, or the destination overlaps the locked prefix; plus + /// any I/O error. + #[cfg(feature = "atomic")] + pub fn copy_as( + &self, + auth: impl BStackAuthority, + from: u64, + to: u64, + n: u64, + ) -> io::Result<()> { + let held = auth.authorities_for(self); + let from_end = checked_end(from, n, "copy: from + n overflows u64")?; + let to_end = checked_end(to, n, "copy: to + n overflows u64")?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + if to < locked { + return Err(io_error!( + InvalidInput, + format!( + "copy: destination [{to}, {to_end}) overlaps locked region [0, {locked})" + ) + )); + } + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if from_end > data_size { + return Err(io_error!( + InvalidInput, + format!("copy: source [{from}, {from_end}) exceeds payload size ({data_size})") + )); + } + if to_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "copy: destination [{to}, {to_end}) exceeds payload size ({data_size})" + ) + )); + } + if n == 0 { + return Ok(()); + } + if from == to { + return Ok(()); + } + acl_check!(self, from, from_end, Read, held); + acl_check!(self, to, to_end, Write, held); + fault_point!(self, "copy"); + if is_atomic_write(to, n) { + let mut buf = vec![0u8; n as usize]; + read_at(file, from, &mut buf)?; + Self::mark_replay(replay, write_at(file, to, &buf))?; + Self::mark_replay(replay, durable_sync(file)) + } else if from < to_end && to < from_end { + Self::mark_replay(replay, journaled_move(file, data_size, from, to, n)) + } else { + Self::mark_replay(replay, journaled_copy(file, data_size, from, to, n)) + } + } + + /// [`swap`](BStack::swap) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's mode denies + /// the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `offset + buf.len()` overflows `u64` or exceeds the payload size; plus any I/O + /// error from the read-back or write. + #[cfg(feature = "atomic")] + pub fn swap_as( + &self, + auth: impl BStackAuthority, + offset: u64, + buf: impl AsRef<[u8]>, + ) -> io::Result> { + let buf = buf.as_ref(); + if buf.is_empty() { + return Ok(Vec::new()); + } + let held = auth.authorities_for(self); + let end = checked_end(offset, buf.len() as u64, "swap: offset + len overflows u64")?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + check_offset_unlocked("swap", offset, end, locked)?; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("swap: range [{offset}, {end}) exceeds payload size ({data_size})") + )); + } + acl_check!(self, offset, end, Write, held); + fault_point!(self, "swap"); + let mut old = vec![0u8; buf.len()]; + read_at(file, offset, &mut old)?; + Self::mark_replay(replay, set_in_place(file, data_size, offset, buf))?; + Ok(old) + } + + /// [`swap_into`](BStack::swap_into) presenting an access token. + /// + /// # Errors + /// + /// As [`swap_as`](Self::swap_as): [`PermissionDenied`](io::ErrorKind::PermissionDenied) + /// if the range's mode denies the write under `auth`; + /// [`InvalidInput`](io::ErrorKind::InvalidInput) if `offset + buf.len()` overflows + /// `u64` or exceeds the payload size; plus any I/O error. + #[cfg(feature = "atomic")] + pub fn swap_into_as( + &self, + auth: impl BStackAuthority, + offset: u64, + buf: &mut [u8], + ) -> io::Result<()> { + if buf.is_empty() { + return Ok(()); + } + let held = auth.authorities_for(self); + let end = checked_end( + offset, + buf.len() as u64, + "swap_into: offset + len overflows u64", + )?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + check_offset_unlocked("swap_into", offset, end, locked)?; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "swap_into: range [{offset}, {end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, offset, end, Write, held); + fault_point!(self, "swap_into"); + let mut tmp = vec![0u8; buf.len()]; + read_at(file, offset, &mut tmp)?; + Self::mark_replay(replay, set_in_place(file, data_size, offset, buf))?; + buf.copy_from_slice(&tmp); + Ok(()) + } + + /// [`splice`](BStack::splice) presenting an access token: the truncated tail range + /// is checked under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the tail range's mode + /// denies the truncate under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `n` exceeds the payload size or the cut reaches into the locked prefix; plus any + /// I/O error. + #[cfg(feature = "atomic")] + pub fn splice_as( + &self, + auth: impl BStackAuthority, + n: u64, + buf: impl AsRef<[u8]>, + ) -> io::Result> { + let buf = buf.as_ref(); + let buf_len = buf.len() as u64; + if n == 0 && buf_len == 0 { + return Ok(Vec::new()); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("splice: n ({n}) exceeds payload size ({data_size})") + )); + } + let locked = self.locked.load(Ordering::Acquire); + let new_tail_start = data_size - n; + if new_tail_start < locked { + return Err(io_error!( + InvalidInput, + format!("splice: operation would modify locked region [0, {locked})") + )); + } + acl_check!(self, new_tail_start, data_size, Truncate, held); + fault_point!(self, "splice"); + let mut removed = vec![0u8; n as usize]; + read_at(file, new_tail_start, &mut removed)?; + Self::mark_replay( + replay, + commit_tail_replace(file, clen, new_tail_start, n, buf, file_end), + )?; + Ok(removed) + } + + /// [`splice_into`](BStack::splice_into) presenting an access token. + /// + /// # Errors + /// + /// As [`splice_as`](Self::splice_as): [`PermissionDenied`](io::ErrorKind::PermissionDenied) + /// if the tail range's mode denies the truncate under `auth`; + /// [`InvalidInput`](io::ErrorKind::InvalidInput) if `old.len()` exceeds the payload + /// size or the cut reaches the locked prefix; plus any I/O error. + #[cfg(feature = "atomic")] + pub fn splice_into_as( + &self, + auth: impl BStackAuthority, + old: &mut [u8], + new: impl AsRef<[u8]>, + ) -> io::Result<()> { + let new = new.as_ref(); + let n = old.len() as u64; + let new_len = new.len() as u64; + if n == 0 && new_len == 0 { + return Ok(()); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("splice_into: n ({n}) exceeds payload size ({data_size})") + )); + } + let locked = self.locked.load(Ordering::Acquire); + let new_tail_start = data_size - n; + if new_tail_start < locked { + return Err(io_error!( + InvalidInput, + format!("splice_into: operation would modify locked region [0, {locked})") + )); + } + acl_check!(self, new_tail_start, data_size, Truncate, held); + fault_point!(self, "splice_into"); + read_at(file, new_tail_start, old)?; + Self::mark_replay( + replay, + commit_tail_replace(file, clen, new_tail_start, n, new, file_end), + ) + } + + /// [`replace`](BStack::replace) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the tail range's mode + /// denies the truncate under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `n` exceeds the payload size or the cut reaches the locked prefix; plus any I/O + /// error. + #[cfg(feature = "atomic")] + pub fn replace_as(&self, auth: impl BStackAuthority, n: u64, f: F) -> io::Result<()> + where + F: FnOnce(&[u8]) -> Vec, + { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("replace: n ({n}) exceeds payload size ({data_size})") + )); + } + let locked = self.locked.load(Ordering::Acquire); + let new_tail_start = data_size - n; + if new_tail_start < locked { + return Err(io_error!( + InvalidInput, + format!("replace: operation would modify locked region [0, {locked})") + )); + } + acl_check!(self, new_tail_start, data_size, Truncate, held); + fault_point!(self, "replace"); + let mut old_tail = vec![0u8; n as usize]; + read_at(file, new_tail_start, &mut old_tail)?; + let new_tail = f(&old_tail); + Self::mark_replay( + replay, + commit_tail_replace(file, clen, new_tail_start, n, &new_tail, file_end), + ) + } + + /// [`set_batched`](BStack::set_batched) presenting an access token: every write + /// block is checked under `auth` before the batch is journaled. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if any block's mode denies + /// the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a block + /// overflows `u64`, exceeds the payload size, overlaps the locked prefix, or two + /// blocks overlap; plus any I/O error. + #[cfg(feature = "atomic")] + pub fn set_batched_as(&self, auth: impl BStackAuthority, writes: I) -> io::Result<()> + where + I: IntoIterator, + D: AsRef<[u8]>, + { + let owned: Vec<(u64, D)> = writes.into_iter().collect(); + let mut blocks: Vec<(u64, &[u8])> = owned + .iter() + .map(|(off, d)| (*off, d.as_ref())) + .filter(|(_, d)| !d.is_empty()) + .collect(); + if blocks.is_empty() { + return Ok(()); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + for (off, data) in &blocks { + let end = checked_end( + *off, + data.len() as u64, + "set_batched: offset + len overflows u64", + )?; + if *off < locked { + return Err(io_error!( + InvalidInput, + format!( + "set_batched: write range [{off}, {end}) overlaps locked region [0, {locked})" + ) + )); + } + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "set_batched: write range [{off}, {end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, *off, end, Write, held); + } + fault_point!(self, "set_batched"); + if blocks.len() == 1 { + let (off, data) = blocks[0]; + return Self::mark_replay(replay, set_in_place(file, data_size, off, data)); + } + blocks.sort_by_key(|(off, _)| *off); + for pair in blocks.windows(2) { + let (a_off, a_data) = pair[0]; + let (b_off, _) = pair[1]; + let a_end = a_off + a_data.len() as u64; + if a_end > b_off { + return Err(io_error!( + InvalidInput, + format!( + "set_batched: write range [{a_off}, {a_end}) overlaps [{b_off}, ...)" + ) + )); + } + } + Self::mark_replay(replay, journaled_multi_set(file, data_size, &blocks)) + } + + /// [`process`](BStack::process) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's mode denies + /// the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if the range + /// overflows `u64` or exceeds the payload size; plus any I/O error. The callback must + /// not change the buffer length. + #[cfg(feature = "atomic")] + pub fn process_as( + &self, + auth: impl BStackAuthority, + start: u64, + end: u64, + f: F, + ) -> io::Result<()> + where + F: FnOnce(&mut [u8]), + { + if end < start { + return Err(io_error!( + InvalidInput, + format!("process: end ({end}) < start ({start})") + )); + } + let held = auth.authorities_for(self); + let n = end - start; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!("process: end ({end}) exceeds payload size ({data_size})") + )); + } + let locked = self.locked.load(Ordering::Acquire); + if start < locked { + return Err(io_error!( + InvalidInput, + format!("process: range [{start}, {end}) overlaps locked region [0, {locked})") + )); + } + acl_check!(self, start, end, Write, held); + fault_point!(self, "process"); + let mut buf = vec![0u8; n as usize]; + if n > 0 { + read_at(file, start, &mut buf)?; + } + f(&mut buf); + if n > 0 { + Self::mark_replay(replay, set_in_place(file, data_size, start, &buf))?; + } + Ok(()) + } + + /// [`eq_crds`](BStack::eq_crds) presenting an access token: the compared range and the + /// write range are checked under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if a touched range's mode + /// denies the access under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a + /// range overflows `u64` or exceeds the payload size, or the lengths mismatch; plus + /// any I/O error. A failed compare returns `Ok(None)`, not an error. + #[cfg(feature = "atomic")] + pub fn eq_crds_as( + &self, + auth: impl BStackAuthority, + a_offset: u64, + a_expected: impl AsRef<[u8]>, + b_offset: u64, + b_buf: impl AsRef<[u8]>, + ) -> io::Result>> { + let a_expected = a_expected.as_ref(); + let b_buf = b_buf.as_ref(); + let held = auth.authorities_for(self); + let a_end = checked_end( + a_offset, + a_expected.len() as u64, + "eq_crds: a_offset + a_len overflows u64", + )?; + let b_end = checked_end( + b_offset, + b_buf.len() as u64, + "eq_crds: b_offset + b_len overflows u64", + )?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + if !b_buf.is_empty() && b_offset < locked { + return Err(io_error!( + InvalidInput, + format!( + "eq_crds: B range [{b_offset}, {b_end}) overlaps locked region [0, {locked})" + ) + )); + } + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if !a_expected.is_empty() && a_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "eq_crds: A range [{a_offset}, {a_end}) exceeds payload size ({data_size})" + ) + )); + } + if !b_buf.is_empty() && b_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "eq_crds: B range [{b_offset}, {b_end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, a_offset, a_end, Read, held); + acl_check!(self, b_offset, b_end, Write, held); + fault_point!(self, "eq_crds"); + let mut a_current = vec![0u8; a_expected.len()]; + if !a_expected.is_empty() { + read_at(file, a_offset, &mut a_current)?; + } + if a_current != a_expected { + return Ok(None); + } + if b_buf.is_empty() { + return Ok(Some(Vec::new())); + } + let mut old_b = vec![0u8; b_buf.len()]; + read_at(file, b_offset, &mut old_b)?; + Self::mark_replay(replay, set_in_place(file, data_size, b_offset, b_buf))?; + Ok(Some(old_b)) + } + + /// [`ne_crds`](BStack::ne_crds) presenting an access token: the compared range and the + /// write range are checked under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if a touched range's mode + /// denies the access under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a + /// range overflows `u64` or exceeds the payload size, or the lengths mismatch; plus + /// any I/O error. A failed compare returns `Ok(None)`, not an error. + #[cfg(feature = "atomic")] + pub fn ne_crds_as( + &self, + auth: impl BStackAuthority, + a_offset: u64, + a_expected: impl AsRef<[u8]>, + b_offset: u64, + b_buf: impl AsRef<[u8]>, + ) -> io::Result>> { + let a_expected = a_expected.as_ref(); + let b_buf = b_buf.as_ref(); + let held = auth.authorities_for(self); + let a_end = checked_end( + a_offset, + a_expected.len() as u64, + "ne_crds: a_offset + a_len overflows u64", + )?; + let b_end = checked_end( + b_offset, + b_buf.len() as u64, + "ne_crds: b_offset + b_len overflows u64", + )?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + if !b_buf.is_empty() && b_offset < locked { + return Err(io_error!( + InvalidInput, + format!( + "ne_crds: B range [{b_offset}, {b_end}) overlaps locked region [0, {locked})" + ) + )); + } + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if !a_expected.is_empty() && a_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "ne_crds: A range [{a_offset}, {a_end}) exceeds payload size ({data_size})" + ) + )); + } + if !b_buf.is_empty() && b_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "ne_crds: B range [{b_offset}, {b_end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, a_offset, a_end, Read, held); + acl_check!(self, b_offset, b_end, Write, held); + fault_point!(self, "ne_crds"); + let mut a_current = vec![0u8; a_expected.len()]; + if !a_expected.is_empty() { + read_at(file, a_offset, &mut a_current)?; + } + if a_current == a_expected { + return Ok(None); + } + if b_buf.is_empty() { + return Ok(Some(Vec::new())); + } + let mut old_b = vec![0u8; b_buf.len()]; + read_at(file, b_offset, &mut old_b)?; + Self::mark_replay(replay, set_in_place(file, data_size, b_offset, b_buf))?; + Ok(Some(old_b)) + } + + /// [`masked_eq_crds`](BStack::masked_eq_crds) presenting an access token: the compared range and the + /// write range are checked under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if a touched range's mode + /// denies the access under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a + /// range overflows `u64` or exceeds the payload size, or the lengths mismatch; plus + /// any I/O error. A failed compare returns `Ok(None)`, not an error. + #[cfg(feature = "atomic")] + pub fn masked_eq_crds_as( + &self, + auth: impl BStackAuthority, + a_offset: u64, + mask: impl AsRef<[u8]>, + a_expected: impl AsRef<[u8]>, + b_offset: u64, + b_buf: impl AsRef<[u8]>, + ) -> io::Result>> { + let mask = mask.as_ref(); + let a_expected = a_expected.as_ref(); + let b_buf = b_buf.as_ref(); + if mask.len() != a_expected.len() { + return Err(io_error!( + InvalidInput, + "masked_eq_crds: mask length ({}) != a_expected length ({})", + mask.len(), + a_expected.len() + )); + } + let held = auth.authorities_for(self); + let a_end = checked_end( + a_offset, + a_expected.len() as u64, + "masked_eq_crds: a_offset + a_len overflows u64", + )?; + let b_end = checked_end( + b_offset, + b_buf.len() as u64, + "masked_eq_crds: b_offset + b_len overflows u64", + )?; + let mut guard = self.write_lock()?; + let (file, _, replay) = &mut *guard; + let locked = self.locked.load(Ordering::Acquire); + if !b_buf.is_empty() && b_offset < locked { + return Err(io_error!( + InvalidInput, + format!( + "masked_eq_crds: B range [{b_offset}, {b_end}) overlaps locked region [0, {locked})" + ) + )); + } + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if !a_expected.is_empty() && a_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "masked_eq_crds: A range [{a_offset}, {a_end}) exceeds payload size ({data_size})" + ) + )); + } + if !b_buf.is_empty() && b_end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "masked_eq_crds: B range [{b_offset}, {b_end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, a_offset, a_end, Read, held); + acl_check!(self, b_offset, b_end, Write, held); + fault_point!(self, "masked_eq_crds"); + let mut a_current = vec![0u8; a_expected.len()]; + if !a_expected.is_empty() { + read_at(file, a_offset, &mut a_current)?; + } + let masked_match = a_current + .iter() + .zip(mask.iter()) + .zip(a_expected.iter()) + .all(|((&a, &m), &e)| (a & m) == (e & m)); + if !masked_match { + return Ok(None); + } + if b_buf.is_empty() { + return Ok(Some(Vec::new())); + } + let mut old_b = vec![0u8; b_buf.len()]; + read_at(file, b_offset, &mut old_b)?; + Self::mark_replay(replay, set_in_place(file, data_size, b_offset, b_buf))?; + Ok(Some(old_b)) + } + + /// [`push`](BStack::push) presenting an access token: the appended range + /// `[len, len + data.len())` is checked for write under `auth` (it need not be + /// `All` — a range may be armed before its bytes arrive). + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the appended range's mode + /// denies the write under `auth`; plus any I/O error from the append. + pub fn push_as( + &self, + auth: impl BStackAuthority, + data: impl AsRef<[u8]>, + ) -> io::Result { + let data = data.as_ref(); + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let logical_offset = file_end - HEADER_SIZE; + if data.is_empty() { + return Ok(logical_offset); + } + acl_check!( + self, + logical_offset, + logical_offset + data.len() as u64, + Write, + held + ); + fault_point!(self, "push"); + if let Err(e) = file.write_all(data) { + if file.set_len(file_end).is_err() { + *replay = true; + } + return Err(e); + } + let new_len = logical_offset + data.len() as u64; + Self::mark_replay( + replay, + commit_grow(file, clen, new_len, logical_offset, file_end), + )?; + Ok(logical_offset) + } + + /// [`extend`](BStack::extend) presenting an access token: the grown range + /// `[len, len + n)` is checked for write under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the grown range's mode + /// denies the write under `auth`; plus any I/O error. + pub fn extend_as(&self, auth: impl BStackAuthority, n: u64) -> io::Result { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let logical_offset = file_end - HEADER_SIZE; + if n == 0 { + return Ok(logical_offset); + } + acl_check!(self, logical_offset, logical_offset + n, Write, held); + fault_point!(self, "extend"); + let new_file_end = file_end + n; + Self::mark_replay(replay, file.set_len(new_file_end))?; + let new_len = logical_offset + n; + Self::mark_replay( + replay, + commit_grow(file, clen, new_len, logical_offset, file_end), + )?; + Ok(logical_offset) + } + + /// [`resize`](BStack::resize) presenting an access token: a grow checks the new tail + /// for write, a shrink checks the discarded tail for truncate, under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the affected range's mode + /// denies the operation under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// a shrink would cut into the locked prefix; plus any I/O error. + pub fn resize_as(&self, auth: impl BStackAuthority, target: u64) -> io::Result { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if target == data_size { + return Ok(data_size); + } + if target < data_size { + let locked = self.locked.load(Ordering::Acquire); + if target < locked { + return Err(io_error!( + InvalidInput, + format!( + "resize({target}) would shrink payload below locked length ({locked})" + ) + )); + } + acl_check!(self, target, data_size, Truncate, held); + fault_point!(self, "resize"); + Self::mark_replay(replay, commit_shrink(file, clen, target))?; + return Ok(data_size); + } + acl_check!(self, data_size, target, Write, held); + fault_point!(self, "resize"); + Self::mark_replay(replay, file.set_len(HEADER_SIZE + target))?; + Self::mark_replay(replay, commit_grow(file, clen, target, data_size, file_end))?; + Ok(data_size) + } + + /// [`ensure`](BStack::ensure) presenting an access token: when it grows, the new + /// range is checked for write under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the grown range's mode + /// denies the write under `auth`; plus any I/O error. + pub fn ensure_as(&self, auth: impl BStackAuthority, target: u64) -> io::Result { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if target <= data_size { + return Ok(data_size); + } + acl_check!(self, data_size, target, Write, held); + fault_point!(self, "ensure"); + Self::mark_replay(replay, file.set_len(HEADER_SIZE + target))?; + Self::mark_replay(replay, commit_grow(file, clen, target, data_size, file_end))?; + Ok(data_size) + } + + /// [`extend_sparse`](BStack::extend_sparse) presenting an access token: the grown + /// range `[len, len + length)` is checked for write under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the grown range's mode + /// denies the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `buf.len()` exceeds `length` or the payload size plus `length` overflows `u64`; + /// plus any I/O error. + pub fn extend_sparse_as( + &self, + auth: impl BStackAuthority, + buf: impl AsRef<[u8]>, + length: u64, + ) -> io::Result { + let buf = buf.as_ref(); + if buf.len() as u64 > length { + return Err(io_error!( + InvalidInput, + "extend_sparse: buffer length ({}) exceeds extension length ({length})", + buf.len() + )); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let logical_offset = file_end - HEADER_SIZE; + if length == 0 { + return Ok(logical_offset); + } + let new_len = logical_offset.checked_add(length).ok_or_else(|| { + io_error!( + InvalidInput, + "extend_sparse: payload size + length overflows u64" + ) + })?; + acl_check!(self, logical_offset, new_len, Write, held); + fault_point!(self, "extend_sparse"); + let one = [(0u64, buf)]; + let blocks: &[(u64, &[u8])] = if buf.is_empty() { &[] } else { &one }; + Self::mark_replay( + replay, + commit_sparse_extend(file, clen, logical_offset, file_end, new_len, blocks), + )?; + Ok(logical_offset) + } + + /// [`extend_sparse_batched`](BStack::extend_sparse_batched) presenting an access + /// token: the grown range is checked for write under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the grown range's mode + /// denies the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a + /// write overflows `u64` or exceeds `length`, two writes overlap, or the payload size + /// plus `length` overflows `u64`; plus any I/O error. + pub fn extend_sparse_batched_as( + &self, + auth: impl BStackAuthority, + writes: I, + length: u64, + ) -> io::Result + where + I: IntoIterator, + D: AsRef<[u8]>, + { + let owned: Vec<(u64, D)> = writes.into_iter().collect(); + let mut blocks: Vec<(u64, &[u8])> = owned + .iter() + .map(|(off, d)| (*off, d.as_ref())) + .filter(|(_, d)| !d.is_empty()) + .collect(); + validate_sparse_blocks(&mut blocks, length, "extend_sparse_batched")?; + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let logical_offset = file_end - HEADER_SIZE; + if length == 0 { + return Ok(logical_offset); + } + let new_len = logical_offset.checked_add(length).ok_or_else(|| { + io_error!( + InvalidInput, + "extend_sparse_batched: payload size + length overflows u64" + ) + })?; + acl_check!(self, logical_offset, new_len, Write, held); + fault_point!(self, "extend_sparse_batched"); + Self::mark_replay( + replay, + commit_sparse_extend(file, clen, logical_offset, file_end, new_len, &blocks), + )?; + Ok(logical_offset) + } + + /// [`pop`](BStack::pop) presenting an access token: the removed tail range + /// `[len - n, len)` is checked for truncate under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the tail range's mode + /// denies the truncate under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `n` exceeds the payload size or shrinks below the locked length; plus any I/O + /// error. + pub fn pop_as(&self, auth: impl BStackAuthority, n: u64) -> io::Result> { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let raw_size = file.seek(SeekFrom::End(0))?; + let data_size = raw_size - HEADER_SIZE; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("pop({n}) exceeds payload size ({data_size})") + )); + } + let new_data_len = data_size - n; + let locked = self.locked.load(Ordering::Acquire); + if new_data_len < locked { + return Err(io_error!( + InvalidInput, + format!("pop({n}) would shrink payload below locked length ({locked})") + )); + } + acl_check!(self, new_data_len, data_size, Truncate, held); + let mut buf = vec![0u8; n as usize]; + fault_point!(self, "pop"); + read_at(file, new_data_len, &mut buf)?; + Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; + Ok(buf) + } + + /// [`pop_into`](BStack::pop_into) presenting an access token. + /// + /// # Errors + /// + /// As [`pop_as`](Self::pop_as): [`PermissionDenied`](io::ErrorKind::PermissionDenied) + /// if the tail range's mode denies the truncate under `auth`; + /// [`InvalidInput`](io::ErrorKind::InvalidInput) if `buf.len()` exceeds the payload + /// size or shrinks below the locked length; plus any I/O error. + pub fn pop_into_as(&self, auth: impl BStackAuthority, buf: &mut [u8]) -> io::Result<()> { + if buf.is_empty() { + return Ok(()); + } + let n = buf.len() as u64; + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let raw_size = file.seek(SeekFrom::End(0))?; + let data_size = raw_size - HEADER_SIZE; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("pop_into({n}) exceeds payload size ({data_size})") + )); + } + let new_data_len = data_size - n; + let locked = self.locked.load(Ordering::Acquire); + if new_data_len < locked { + return Err(io_error!( + InvalidInput, + format!("pop_into({n}) would shrink payload below locked length ({locked})") + )); + } + acl_check!(self, new_data_len, data_size, Truncate, held); + fault_point!(self, "pop_into"); + read_at(file, new_data_len, buf)?; + Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; + Ok(()) + } + + /// [`peek`](BStack::peek) presenting an access token: the read range `[offset, len)` + /// is checked for read under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's mode denies + /// the read under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if `offset` + /// exceeds the payload size; plus any I/O error. + pub fn peek_as(&self, auth: impl BStackAuthority, offset: u64) -> io::Result> { + let held = auth.authorities_for(self); + #[cfg(any(unix, windows))] + { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + if offset > data_size { + return Err(io_error!( + InvalidInput, + format!("peek offset ({offset}) exceeds payload size ({data_size})") + )); + } + acl_check!(self, offset, data_size, Read, held); + fault_point!(self, "peek"); + pread_exact(file, HEADER_SIZE + offset, (data_size - offset) as usize) + } + #[cfg(not(any(unix, windows)))] + { + let mut guard = self.write_lock_read()?; + let file = &mut guard.0; + let raw_size = file.seek(SeekFrom::End(0))?; + let data_size = raw_size.saturating_sub(HEADER_SIZE); + if offset > data_size { + return Err(io_error!( + InvalidInput, + format!("peek offset ({offset}) exceeds payload size ({data_size})") + )); + } + acl_check!(self, offset, data_size, Read, held); + fault_point!(self, "peek"); + file.seek(SeekFrom::Start(HEADER_SIZE + offset))?; + let mut buf = vec![0u8; (data_size - offset) as usize]; + file.read_exact(&mut buf)?; + Ok(buf) + } + } + + /// [`peek_into`](BStack::peek_into) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the range's mode denies + /// the read under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `offset + buf.len()` overflows `u64` or exceeds the payload size; plus any I/O + /// error. + pub fn peek_into_as( + &self, + auth: impl BStackAuthority, + offset: u64, + buf: &mut [u8], + ) -> io::Result<()> { + if buf.is_empty() { + return Ok(()); + } + let len = buf.len() as u64; + let end = offset + .checked_add(len) + .ok_or_else(|| io_error!(InvalidInput, "peek_into: offset + len overflows u64"))?; + let held = auth.authorities_for(self); + acl_check!(self, offset, end, Read, held); + #[cfg(any(unix, windows))] + { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "peek_into: range [{offset}, {end}) exceeds payload size ({data_size})" + ) + )); + } + fault_point!(self, "peek_into"); + pread_exact_into(file, HEADER_SIZE + offset, buf) + } + #[cfg(not(any(unix, windows)))] + { + let mut guard = self.write_lock_read()?; + let file = &mut guard.0; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "peek_into: range [{offset}, {end}) exceeds payload size ({data_size})" + ) + )); + } + fault_point!(self, "peek_into"); + file.seek(SeekFrom::Start(HEADER_SIZE + offset))?; + file.read_exact(buf) + } + } + + /// [`atrunc`](BStack::atrunc) presenting an access token: the rewritten tail range + /// `[data_size - n, data_size)` is checked under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the tail range's mode + /// denies the truncate under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `n` exceeds the payload size or the cut reaches the locked prefix; plus any I/O + /// error. + #[cfg(feature = "atomic")] + pub fn atrunc_as( + &self, + auth: impl BStackAuthority, + n: u64, + buf: impl AsRef<[u8]>, + ) -> io::Result<()> { + let buf = buf.as_ref(); + let buf_len = buf.len() as u64; + if n == 0 && buf_len == 0 { + return Ok(()); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("atrunc: n ({n}) exceeds payload size ({data_size})") + )); + } + let locked = self.locked.load(Ordering::Acquire); + let new_tail_start = data_size - n; + if new_tail_start < locked { + return Err(io_error!( + InvalidInput, + format!("atrunc: operation would modify locked region [0, {locked})") + )); + } + acl_check!(self, new_tail_start, data_size, Truncate, held); + fault_point!(self, "atrunc"); + Self::mark_replay( + replay, + commit_tail_replace(file, clen, new_tail_start, n, buf, file_end), + ) + } + + /// [`try_extend`](BStack::try_extend) presenting an access token: when the size + /// matches, the appended range is checked for write under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the appended range's mode + /// denies the write under `auth`; plus any I/O error. A size mismatch returns + /// `Ok(false)`, not an error. + #[cfg(feature = "atomic")] + pub fn try_extend_as( + &self, + auth: impl BStackAuthority, + s: u64, + buf: impl AsRef<[u8]>, + ) -> io::Result { + let buf = buf.as_ref(); + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if data_size != s { + return Ok(false); + } + if buf.is_empty() { + return Ok(true); + } + acl_check!(self, data_size, data_size + buf.len() as u64, Write, held); + fault_point!(self, "try_extend"); + if let Err(e) = file.write_all(buf) { + if file.set_len(file_end).is_err() { + *replay = true; + } + return Err(e); + } + let new_len = data_size + buf.len() as u64; + Self::mark_replay( + replay, + commit_grow(file, clen, new_len, data_size, file_end), + )?; + Ok(true) + } + + /// [`try_extend_zeros`](BStack::try_extend_zeros) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the grown range's mode + /// denies the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if the + /// payload size plus `n` overflows `u64`; plus any I/O error. A size mismatch returns + /// `Ok(false)`. + #[cfg(feature = "atomic")] + pub fn try_extend_zeros_as( + &self, + auth: impl BStackAuthority, + s: u64, + n: u64, + ) -> io::Result { + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if data_size != s { + return Ok(false); + } + if n == 0 { + return Ok(true); + } + let new_len = checked_end( + data_size, + n, + "try_extend_zeros: data_size + n overflows u64", + )?; + acl_check!(self, data_size, new_len, Write, held); + fault_point!(self, "try_extend_zeros"); + Self::mark_replay(replay, file.set_len(HEADER_SIZE + new_len))?; + Self::mark_replay( + replay, + commit_grow(file, clen, new_len, data_size, file_end), + )?; + Ok(true) + } + + /// [`try_extend_sparse`](BStack::try_extend_sparse) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the grown range's mode + /// denies the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `buf.len()` exceeds `length` or the payload size plus `length` overflows `u64`; + /// plus any I/O error. A size mismatch returns `Ok(false)`. + #[cfg(feature = "atomic")] + pub fn try_extend_sparse_as( + &self, + auth: impl BStackAuthority, + s: u64, + buf: impl AsRef<[u8]>, + length: u64, + ) -> io::Result { + let buf = buf.as_ref(); + if buf.len() as u64 > length { + return Err(io_error!( + InvalidInput, + "try_extend_sparse: buffer length ({}) exceeds extension length ({length})", + buf.len() + )); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if data_size != s { + return Ok(false); + } + if length == 0 { + return Ok(true); + } + let new_len = checked_end( + data_size, + length, + "try_extend_sparse: data_size + length overflows u64", + )?; + acl_check!(self, data_size, new_len, Write, held); + fault_point!(self, "try_extend_sparse"); + let one = [(0u64, buf)]; + let blocks: &[(u64, &[u8])] = if buf.is_empty() { &[] } else { &one }; + Self::mark_replay( + replay, + commit_sparse_extend(file, clen, data_size, file_end, new_len, blocks), + )?; + Ok(true) + } + + /// [`try_extend_sparse_batched`](BStack::try_extend_sparse_batched) presenting an + /// access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the grown range's mode + /// denies the write under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a + /// write overflows `u64` or exceeds `length`, two writes overlap, or the payload size + /// plus `length` overflows `u64`; plus any I/O error. A size mismatch returns + /// `Ok(false)`. + #[cfg(feature = "atomic")] + pub fn try_extend_sparse_batched_as( + &self, + auth: impl BStackAuthority, + s: u64, + writes: I, + length: u64, + ) -> io::Result + where + I: IntoIterator, + D: AsRef<[u8]>, + { + let owned: Vec<(u64, D)> = writes.into_iter().collect(); + let mut blocks: Vec<(u64, &[u8])> = owned + .iter() + .map(|(off, d)| (*off, d.as_ref())) + .filter(|(_, d)| !d.is_empty()) + .collect(); + validate_sparse_blocks(&mut blocks, length, "try_extend_sparse_batched")?; + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let file_end = file.seek(SeekFrom::End(0))?; + let data_size = file_end - HEADER_SIZE; + if data_size != s { + return Ok(false); + } + if length == 0 { + return Ok(true); + } + let new_len = checked_end( + data_size, + length, + "try_extend_sparse_batched: data_size + length overflows u64", + )?; + acl_check!(self, data_size, new_len, Write, held); + fault_point!(self, "try_extend_sparse_batched"); + Self::mark_replay( + replay, + commit_sparse_extend(file, clen, data_size, file_end, new_len, &blocks), + )?; + Ok(true) + } + + /// [`try_discard`](BStack::try_discard) presenting an access token: when the size + /// matches, the removed tail range is checked for truncate under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if the tail range's mode + /// denies the truncate under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if + /// `n` exceeds the payload size or shrinks below the locked length; plus any I/O + /// error. A size mismatch returns `Ok(false)`. + #[cfg(feature = "atomic")] + pub fn try_discard_as( + &self, + auth: impl BStackAuthority, + s: u64, + n: u64, + ) -> io::Result { + if n == 0 { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + return Ok(data_size == s); + } + let held = auth.authorities_for(self); + let mut guard = self.write_lock()?; + let (file, clen, replay) = &mut *guard; + let raw_size = file.seek(SeekFrom::End(0))?; + let data_size = raw_size - HEADER_SIZE; + if data_size != s { + return Ok(false); + } + if n > data_size { + return Err(io_error!( + InvalidInput, + format!("try_discard: n ({n}) exceeds payload size ({data_size})") + )); + } + let new_data_len = data_size - n; + let locked = self.locked.load(Ordering::Acquire); + if new_data_len < locked { + return Err(io_error!( + InvalidInput, + format!("try_discard: would shrink payload below locked length ({locked})") + )); + } + acl_check!(self, new_data_len, data_size, Truncate, held); + fault_point!(self, "try_discard"); + Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; + Ok(true) + } + + /// [`get_batched`](BStack::get_batched) presenting an access token: every range is + /// checked for read under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if any range's mode denies + /// the read under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if a range has + /// `end < start` or exceeds the payload size; plus any I/O error. + #[cfg(feature = "atomic")] + pub fn get_batched_as( + &self, + auth: impl BStackAuthority, + ranges: I, + ) -> io::Result>> + where + I: IntoIterator>, + { + let held = auth.authorities_for(self); + let ranges: Vec> = ranges.into_iter().collect(); + if ranges.is_empty() { + return Ok(Vec::new()); + } + for r in &ranges { + if r.end < r.start { + return Err(io_error!( + InvalidInput, + "get_batched: end ({}) < start ({})", + r.end, + r.start + )); + } + acl_check!(self, r.start, r.end, Read, held); + } + #[cfg(any(unix, windows))] + { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + fault_point!(self, "get_batched"); + let mut results = Vec::with_capacity(ranges.len()); + for r in &ranges { + if r.end > data_size { + return Err(io_error!( + InvalidInput, + "get_batched: end ({}) exceeds payload size ({data_size})", + r.end + )); + } + results.push(pread_exact( + file, + HEADER_SIZE + r.start, + (r.end - r.start) as usize, + )?); + } + Ok(results) + } + #[cfg(not(any(unix, windows)))] + { + let mut guard = self.write_lock_read()?; + let file = &mut guard.0; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + fault_point!(self, "get_batched"); + let mut results = Vec::with_capacity(ranges.len()); + for r in &ranges { + if r.end > data_size { + return Err(io_error!( + InvalidInput, + "get_batched: end ({}) exceeds payload size ({data_size})", + r.end + )); + } + file.seek(SeekFrom::Start(HEADER_SIZE + r.start))?; + let mut buf = vec![0u8; (r.end - r.start) as usize]; + file.read_exact(&mut buf)?; + results.push(buf); + } + Ok(results) + } + } + + /// [`get_batched_into`](BStack::get_batched_into) presenting an access token. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if any range's mode denies + /// the read under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if an + /// `offset + buf.len()` overflows `u64` or exceeds the payload size; plus any I/O + /// error. + #[cfg(feature = "atomic")] + pub fn get_batched_into_as<'a, I>( + &self, + auth: impl BStackAuthority, + bufs: I, + ) -> io::Result<()> + where + I: IntoIterator, + { + let held = auth.authorities_for(self); + let bufs: Vec<(u64, &'a mut [u8])> = bufs.into_iter().collect(); + if bufs.is_empty() { + return Ok(()); + } + #[cfg(any(unix, windows))] + { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + fault_point!(self, "get_batched_into"); + for (ptr, buf) in bufs { + let end = ptr.checked_add(buf.len() as u64).ok_or_else(|| { + io_error!( + InvalidInput, + "get_batched_into: offset + buf.len() overflows u64" + ) + })?; + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "get_batched_into: end ({end}) exceeds payload size ({data_size})", + ) + )); + } + acl_check!(self, ptr, end, Read, held); + pread_exact_into(file, HEADER_SIZE + ptr, buf)?; + } + Ok(()) + } + #[cfg(not(any(unix, windows)))] + { + let mut guard = self.write_lock_read()?; + let file = &mut guard.0; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + fault_point!(self, "get_batched_into"); + for (ptr, buf) in bufs { + let end = ptr.checked_add(buf.len() as u64).ok_or_else(|| { + io_error!( + InvalidInput, + "get_batched_into: offset + buf.len() overflows u64" + ) + })?; + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "get_batched_into: end ({end}) exceeds payload size ({data_size})", + ) + )); + } + acl_check!(self, ptr, end, Read, held); + file.seek(SeekFrom::Start(HEADER_SIZE + ptr))?; + file.read_exact(buf)?; + } + Ok(()) + } + } + + /// [`get_batched_gen`](BStack::get_batched_gen) presenting an access token: each + /// requested read range is checked under `auth`. + /// + /// # Errors + /// + /// [`PermissionDenied`](io::ErrorKind::PermissionDenied) if a requested range's mode + /// denies the read under `auth`; [`InvalidInput`](io::ErrorKind::InvalidInput) if an + /// `offset + buf.len()` overflows `u64` or exceeds the payload size; plus any I/O + /// error. + #[cfg(feature = "atomic")] + pub fn get_batched_gen_as<'a, F>( + &self, + auth: impl BStackAuthority, + mut f: F, + ) -> io::Result<()> + where + F: FnMut() -> Option<(u64, &'a mut [u8])>, + { + let held = auth.authorities_for(self); + #[cfg(any(unix, windows))] + { + let guard = self.read_lock()?; + let file = &guard.0; + let data_size = file.metadata()?.len().saturating_sub(HEADER_SIZE); + fault_point!(self, "get_batched_gen"); + while let Some((offset, buf)) = f() { + let end = offset.checked_add(buf.len() as u64).ok_or_else(|| { + io_error!( + InvalidInput, + "get_batched_gen: offset + buf.len() overflows u64" + ) + })?; + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "get_batched_gen: end ({end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, offset, end, Read, held); + fault_point!(self, "get_batched_gen:read"); + pread_exact_into(file, HEADER_SIZE + offset, buf)?; + } + Ok(()) + } + #[cfg(not(any(unix, windows)))] + { + let mut guard = self.write_lock_read()?; + let file = &mut guard.0; + let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); + fault_point!(self, "get_batched_gen"); + while let Some((offset, buf)) = f() { + let end = offset.checked_add(buf.len() as u64).ok_or_else(|| { + io_error!( + InvalidInput, + "get_batched_gen: offset + buf.len() overflows u64" + ) + })?; + if end > data_size { + return Err(io_error!( + InvalidInput, + format!( + "get_batched_gen: end ({end}) exceeds payload size ({data_size})" + ) + )); + } + acl_check!(self, offset, end, Read, held); + fault_point!(self, "get_batched_gen:read"); + file.seek(SeekFrom::Start(HEADER_SIZE + offset))?; + file.read_exact(buf)?; + } + Ok(()) + } + } + } +} + +#[cfg(feature = "expensive-slice-access-control")] +pub use inner::*; + +// Without the feature the allocator wiring still compiles: these fold to nothing +// and inline away, keeping the build byte-identical to one that never mentioned +// access control. Only the allocators (the `alloc` feature) call them. +#[cfg(all(feature = "alloc", not(feature = "expensive-slice-access-control")))] +impl crate::BStack { + #[inline] + #[allow(dead_code)] // metadata-marking hook; see the gated sibling + pub(crate) fn acl_mark_alloc(&self, _offset: u64, _len: u64) -> std::io::Result<()> { + Ok(()) + } + + #[inline] + pub(crate) fn acl_reclaimable(&self, _offset: u64, _len: u64) -> std::io::Result<()> { + Ok(()) + } + + #[inline] + pub(crate) fn acl_reclaim(&self, _offset: u64, _len: u64) -> std::io::Result<()> { + Ok(()) + } +} + +#[cfg(all(test, feature = "expensive-slice-access-control"))] +mod acl_tests { + use crate::*; + use std::io; + use std::path::PathBuf; + use std::sync::atomic::{AtomicU64, Ordering}; + + fn mk() -> (BStack, PathBuf) { + static COUNTER: AtomicU64 = AtomicU64::new(0); + let id = COUNTER.fetch_add(1, Ordering::Relaxed); + let pid = std::process::id(); + let path = std::env::temp_dir().join(format!("bstack_acl_{pid}_{id}.bin")); + let _ = std::fs::remove_file(&path); + (BStack::open(&path).unwrap(), path) + } + + struct Guard(PathBuf); + impl Drop for Guard { + fn drop(&mut self) { + let _ = std::fs::remove_file(&self.0); + } + } + + fn seed(s: &BStack, n: usize) { + s.push(vec![0xAAu8; n]).unwrap(); + } + + #[test] + fn unprotected_stack_is_transparent() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + assert_eq!(s.access_at(10), BStackAccess::All); + s.set(0, [1, 2, 3]).unwrap(); + assert_eq!(s.get(0, 3).unwrap(), [1, 2, 3]); + s.discard(8).unwrap(); + assert_eq!(s.len().unwrap(), 56); + } + + #[test] + fn prot_range_needs_guard_on_every_axis() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 16, 16, BStackAccess::Prot).unwrap(); + assert_eq!(s.access_at(20), BStackAccess::Prot); + + // Tokenless is denied on read, write, and (via the tail) truncate. + assert_eq!( + s.set(16, [0; 4]).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + assert_eq!( + s.get(16, 20).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + assert_eq!( + s.discard(56).unwrap_err().kind(), // would cut into [16,32) + io::ErrorKind::PermissionDenied + ); + + // The guard token satisfies all three. + s.set_as(&prot, 16, [7u8; 4]).unwrap(); + assert_eq!(s.get_as(&prot, 16, 20).unwrap(), [7, 7, 7, 7]); + s.discard_as(&prot, 56).unwrap(); + assert_eq!(s.len().unwrap(), 8); + } + + #[test] + fn alloc_and_guard_are_incomparable() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let alloc = s.take_alloc_authority().unwrap(); + let prot = s.take_protection().unwrap(); + s.protect_as(&alloc, 0, 16, BStackAccess::Alloc).unwrap(); + + // The guard cannot touch an Alloc range, nor re-mode it. + assert_eq!( + s.get_as(&prot, 0, 4).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + assert_eq!( + s.protect_as(&prot, 0, 16, BStackAccess::Prot) + .unwrap_err() + .kind(), + io::ErrorKind::PermissionDenied + ); + // The allocator can. + s.set_as(&alloc, 0, [1u8; 4]).unwrap(); + assert_eq!(s.get_as(&alloc, 0, 4).unwrap(), [1, 1, 1, 1]); + } + + #[test] + fn readonly_denies_writes_only() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 32); + s.protect_as((), 0, 16, BStackAccess::ReadOnly).unwrap(); + assert_eq!(s.get(0, 4).unwrap().len(), 4); // reads pass + assert_eq!( + s.set(0, [0u8; 4]).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + assert_eq!( + s.zero(0, 4).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + } + + #[test] + fn locked_denies_reads() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 32); + s.protect_as((), 0, 16, BStackAccess::Locked).unwrap(); + assert_eq!( + s.get(0, 4).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + assert_eq!( + s.peek(0).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + } + + #[test] + fn tokenless_protect_only_tightens_all() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 32); + s.protect_as((), 0, 16, BStackAccess::ReadOnly).unwrap(); + // Re-arming a now-non-All range without a token is denied. + assert_eq!( + s.protect_as((), 0, 16, BStackAccess::All) + .unwrap_err() + .kind(), + io::ErrorKind::PermissionDenied + ); + } + + #[test] + fn tokens_are_one_shot() { + let (s, p) = mk(); + let _g = Guard(p); + assert!(s.take_protection().is_some()); + assert!(s.take_protection().is_none()); + assert!(s.take_alloc_authority().is_some()); + assert!(s.take_alloc_authority().is_none()); + } + + #[test] + fn returned_token_can_be_reminted() { + let (s, p) = mk(); + let _g = Guard(p); + let prot = s.take_protection().expect("first mint"); + assert!(s.take_protection().is_none(), "one-shot while held"); + s.return_protection(prot); + assert!(s.take_protection().is_some(), "re-mintable after return"); + + let auth = s.take_alloc_authority().expect("first alloc mint"); + assert!(s.take_alloc_authority().is_none()); + s.return_alloc_authority(auth); + assert!(s.take_alloc_authority().is_some()); + } + + #[test] + fn returning_a_foreign_token_does_not_rearm() { + let (s1, p1) = mk(); + let _g1 = Guard(p1); + let (s2, p2) = mk(); + let _g2 = Guard(p2); + let from_s1 = s1.take_protection().unwrap(); + let from_s2 = s2.take_protection().unwrap(); + // s1's token carries s1's identity, so handing it to s2 is a no-op. + s2.return_protection(from_s1); + assert!( + s2.take_protection().is_none(), + "a foreign token must not re-arm the mint" + ); + // s2's own token hands back and re-mints normally. + s2.return_protection(from_s2); + assert!(s2.take_protection().is_some()); + } + + #[test] + fn foreign_token_grants_nothing() { + let (s1, p1) = mk(); + let _g1 = Guard(p1); + let (s2, p2) = mk(); + let _g2 = Guard(p2); + seed(&s1, 32); + let foreign = s2.take_protection().unwrap(); + s1.protect_as((), 0, 16, BStackAccess::Prot).unwrap(); + // A token minted from s2 is treated as tokenless on s1. + assert_eq!( + s1.get_as(&foreign, 0, 4).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + } + + #[test] + fn append_checks_its_target_region() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 16); + // Arm the region a push would land in, before its bytes exist. + s.protect_as((), 16, 16, BStackAccess::Locked).unwrap(); + assert_eq!( + s.push(vec![0u8; 8]).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + assert_eq!( + s.extend(8).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + assert_eq!(s.len().unwrap(), 16); // nothing appended + } + + #[cfg(feature = "atomic")] + #[test] + fn atomic_ops_respect_protection() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 16, 16, BStackAccess::Prot).unwrap(); + let denied = |k: io::ErrorKind| assert_eq!(k, io::ErrorKind::PermissionDenied); + + // In-place mutators over the protected range are denied tokenless. + denied(s.swap(16, [0u8; 4]).unwrap_err().kind()); + denied(s.cas(16, [0xAAu8; 4], [0u8; 4]).unwrap_err().kind()); + denied(s.process(16, 20, |b| b.fill(0)).unwrap_err().kind()); + // copy: source read denied. + denied(s.copy(16, 40, 4).unwrap_err().kind()); + // copy: destination write denied. + denied(s.copy(40, 16, 4).unwrap_err().kind()); + // cross_exchange touching the range denied. + denied(s.cross_exchange(16, 40, 4).unwrap_err().kind()); + // A tail replace that reaches into the range is denied. + denied(s.atrunc(56, []).unwrap_err().kind()); // truncates [8, 64) ⊇ [16,32) + // Batched read of the range denied. + denied(s.get_batched(std::iter::once(16..20)).unwrap_err().kind()); + + // Outside the protected range, the same ops succeed. + s.swap(40, [1u8; 4]).unwrap(); + assert_eq!(s.get(40, 44).unwrap(), [1, 1, 1, 1]); + } + + #[test] + fn authorized_slice_reaches_prot_region() { + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + let mut slice = alloc.alloc(32).unwrap(); + let prot = alloc.stack().take_protection().unwrap(); + slice.protect_as(&prot, BStackAccess::Prot).unwrap(); + // Without authority, the slice's own I/O is denied. + assert_eq!( + slice.read().unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + // Grant the slice the guard authority; its I/O now reaches the region. + slice.authorize(&prot); + slice.write([7u8; 4]).unwrap(); + let bytes = slice.read().unwrap(); + assert_eq!(bytes.len(), 32); + assert_eq!(&bytes[..4], &[7, 7, 7, 7]); + } + + #[test] + fn owned_slice_protect_forwards_to_stack() { + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + let mut slice = alloc.alloc(32).unwrap(); + // Arm the allocation as ReadOnly through the owned handle. + slice.protect(BStackAccess::ReadOnly).unwrap(); + assert_eq!( + alloc.stack().access_at(slice.start()), + BStackAccess::ReadOnly + ); + // Reads through the slice pass; writes are denied by the forwarded policy. + assert!(slice.read().is_ok()); + assert_eq!( + slice.write([0u8; 4]).unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + } + + #[test] + fn owned_slice_protect_as_with_guard() { + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + let slice = alloc.alloc(32).unwrap(); + let prot = alloc.stack().take_protection().unwrap(); + // Arm as Prot; only the guard token may then read or write it. + slice.protect_as(&prot, BStackAccess::Prot).unwrap(); + assert_eq!( + slice.read().unwrap_err().kind(), + io::ErrorKind::PermissionDenied + ); + // The stack's token-carrying entry points still reach it. + let stack = alloc.stack(); + stack.set_as(&prot, slice.start(), [9u8; 4]).unwrap(); + assert_eq!( + stack + .get_as(&prot, slice.start(), slice.start() + 4) + .unwrap(), + [9, 9, 9, 9] + ); + } + + #[cfg(feature = "atomic")] + #[test] + fn set_batched_checks_every_block() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + s.protect_as((), 32, 8, BStackAccess::ReadOnly).unwrap(); + // One block lands in the ReadOnly region → the whole batch is refused. + assert_eq!( + s.set_batched([(0u64, vec![1u8; 4]), (32u64, vec![2u8; 4])]) + .unwrap_err() + .kind(), + io::ErrorKind::PermissionDenied + ); + // The permitted block was not applied (checks run before the journal). + assert_eq!(s.get(0, 4).unwrap(), [0xAA, 0xAA, 0xAA, 0xAA]); + } + + #[cfg(feature = "atomic")] + #[test] + fn authorized_slice_atomic_ops_reach_prot() { + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + let mut slice = alloc.alloc(32).unwrap(); + let prot = alloc.stack().take_protection().unwrap(); + slice.protect_as(&prot, BStackAccess::Prot).unwrap(); + // Tokenless slice process (a write) is denied. + assert_eq!( + slice + .as_slice_mut() + .process(|b| b.fill(9)) + .unwrap_err() + .kind(), + io::ErrorKind::PermissionDenied + ); + // With authority the atomic slice ops reach the Prot region. + slice.authorize(&prot); + slice.as_slice_mut().process(|b| b.fill(1)).unwrap(); + assert_eq!(&slice.read().unwrap()[..4], &[1, 1, 1, 1]); + } + + #[test] + fn authorized_stack_compound_ops_reach_prot() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + // Interior range for the in-place writes; tail range for the replacements. + s.protect_as(&prot, 16, 16, BStackAccess::Prot).unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + let denied = |k: io::ErrorKind| assert_eq!(k, io::ErrorKind::PermissionDenied); + + // swap over [16,20): tokenless denied, `swap_as` reaches it. + denied(s.swap(16, [0u8; 4]).unwrap_err().kind()); + s.swap_as(&prot, 16, [1u8; 4]).unwrap(); + assert_eq!(s.get_as(&prot, 16, 20).unwrap(), [1, 1, 1, 1]); + + // swap_into over [16,20): reads back the bytes just written. + let mut buf = [2u8; 4]; + denied(s.swap_into(16, &mut [0u8; 4]).unwrap_err().kind()); + s.swap_into_as(&prot, 16, &mut buf).unwrap(); + assert_eq!(buf, [1, 1, 1, 1]); + assert_eq!(s.get_as(&prot, 16, 20).unwrap(), [2, 2, 2, 2]); + + // set_batched touching [16,20): tokenless denied, `set_batched_as` succeeds. + denied( + s.set_batched(std::iter::once((16u64, [3u8; 4]))) + .unwrap_err() + .kind(), + ); + s.set_batched_as(&prot, std::iter::once((16u64, [3u8; 4]))) + .unwrap(); + assert_eq!(s.get_as(&prot, 16, 20).unwrap(), [3, 3, 3, 3]); + + // Tail replacements reaching [48,64) via a 20-byte tail (touches [44,64)). + denied(s.splice(20, []).unwrap_err().kind()); + let removed = s.splice_as(&prot, 20, [7u8; 20]).unwrap(); + assert_eq!(removed.len(), 20); + assert_eq!(s.get_as(&prot, 44, 64).unwrap(), [7u8; 20]); + + let mut old = [0u8; 20]; + denied(s.splice_into(&mut [0u8; 20], []).unwrap_err().kind()); + s.splice_into_as(&prot, &mut old, [8u8; 20]).unwrap(); + assert_eq!(old, [7u8; 20]); + assert_eq!(s.get_as(&prot, 44, 64).unwrap(), [8u8; 20]); + + denied(s.replace(20, |b| b.to_vec()).unwrap_err().kind()); + s.replace_as(&prot, 20, |b| b.iter().map(|x| x + 1).collect()) + .unwrap(); + assert_eq!(s.get_as(&prot, 44, 64).unwrap(), [9u8; 20]); + } + + #[cfg(feature = "atomic")] + #[test] + fn authorized_stack_append_ops_reach_prot() { + // Appends land in a Prot window armed just past the tail: tokenless denied, + // `_as` reaches it. Chained so each op fills the next slice of the window. + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 48); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 32, BStackAccess::Prot).unwrap(); + let denied = |k: io::ErrorKind| assert_eq!(k, io::ErrorKind::PermissionDenied); + + denied(s.push([1u8; 4]).unwrap_err().kind()); + assert_eq!(s.push_as(&prot, [1u8; 4]).unwrap(), 48); + + denied(s.extend(4).unwrap_err().kind()); + assert_eq!(s.extend_as(&prot, 4).unwrap(), 52); + + denied(s.extend_sparse([2u8; 2], 4).unwrap_err().kind()); + assert_eq!(s.extend_sparse_as(&prot, [2u8; 2], 4).unwrap(), 56); + + denied( + s.extend_sparse_batched(std::iter::once((0u64, [3u8; 2])), 4) + .unwrap_err() + .kind(), + ); + assert_eq!( + s.extend_sparse_batched_as(&prot, std::iter::once((0u64, [3u8; 2])), 4) + .unwrap(), + 60 + ); + + denied(s.try_extend(64, [4u8; 4]).unwrap_err().kind()); + assert!(s.try_extend_as(&prot, 64, [4u8; 4]).unwrap()); + + denied(s.try_extend_zeros(68, 4).unwrap_err().kind()); + assert!(s.try_extend_zeros_as(&prot, 68, 4).unwrap()); + + denied(s.try_extend_sparse(72, [5u8; 2], 4).unwrap_err().kind()); + assert!(s.try_extend_sparse_as(&prot, 72, [5u8; 2], 4).unwrap()); + + denied( + s.try_extend_sparse_batched(76, std::iter::once((0u64, [6u8; 2])), 4) + .unwrap_err() + .kind(), + ); + assert!( + s.try_extend_sparse_batched_as(&prot, 76, std::iter::once((0u64, [6u8; 2])), 4) + .unwrap() + ); + assert_eq!(s.len().unwrap(), 80); + } + + #[cfg(feature = "atomic")] + #[test] + fn authorized_stack_removal_ops_reach_prot() { + // Tail removals reaching a Prot tail are denied tokenless and succeed with + // the token. Fresh stack per op keeps the protected region well-defined. + let denied = |k: io::ErrorKind| assert_eq!(k, io::ErrorKind::PermissionDenied); + { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + denied(s.pop(20).unwrap_err().kind()); + assert_eq!(s.pop_as(&prot, 20).unwrap().len(), 20); + assert_eq!(s.len().unwrap(), 44); + } + { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + let mut buf = [0u8; 20]; + denied(s.pop_into(&mut [0u8; 20]).unwrap_err().kind()); + s.pop_into_as(&prot, &mut buf).unwrap(); + assert_eq!(s.len().unwrap(), 44); + } + { + // atrunc actually rewrites the tail, so its check matters most here. + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + denied(s.atrunc(20, [1u8; 4]).unwrap_err().kind()); + s.atrunc_as(&prot, 20, [1u8; 4]).unwrap(); + assert_eq!(s.len().unwrap(), 48); + } + { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + denied(s.try_discard(64, 20).unwrap_err().kind()); + assert!(s.try_discard_as(&prot, 64, 20).unwrap()); + assert_eq!(s.len().unwrap(), 44); + } + } + + #[cfg(feature = "atomic")] + #[test] + fn authorized_stack_read_ops_reach_prot() { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 16, 16, BStackAccess::Prot).unwrap(); + s.set_as(&prot, 16, [5u8; 16]).unwrap(); + let denied = |k: io::ErrorKind| assert_eq!(k, io::ErrorKind::PermissionDenied); + + // peek reads from offset to end; starting inside the Prot region is denied. + denied(s.peek(16).unwrap_err().kind()); + assert_eq!(&s.peek_as(&prot, 16).unwrap()[..4], &[5, 5, 5, 5]); + + let mut b = [0u8; 4]; + denied(s.peek_into(16, &mut [0u8; 4]).unwrap_err().kind()); + s.peek_into_as(&prot, 16, &mut b).unwrap(); + assert_eq!(b, [5, 5, 5, 5]); + + denied(s.get_batched(std::iter::once(16..20)).unwrap_err().kind()); + assert_eq!( + s.get_batched_as(&prot, std::iter::once(16..20)).unwrap()[0], + vec![5, 5, 5, 5] + ); + + let mut b2 = [0u8; 4]; + { + let mut dbuf = [0u8; 4]; + denied( + s.get_batched_into(std::iter::once((16u64, &mut dbuf[..]))) + .unwrap_err() + .kind(), + ); + } + s.get_batched_into_as(&prot, std::iter::once((16u64, &mut b2[..]))) + .unwrap(); + assert_eq!(b2, [5, 5, 5, 5]); + + // Lending closure: yield a raw-pointer slice, as the other gen tests do. + let mut gbuf = [0u8; 4]; + let ptr = gbuf.as_mut_ptr(); + let mut called = false; + denied( + s.get_batched_gen(|| { + if called { + None + } else { + called = true; + Some((16u64, unsafe { std::slice::from_raw_parts_mut(ptr, 4) })) + } + }) + .unwrap_err() + .kind(), + ); + let mut called2 = false; + s.get_batched_gen_as(&prot, || { + if called2 { + None + } else { + called2 = true; + Some((16u64, unsafe { std::slice::from_raw_parts_mut(ptr, 4) })) + } + }) + .unwrap(); + assert_eq!(gbuf, [5, 5, 5, 5]); + } + + #[test] + fn authorized_stack_resize_ops_reach_prot() { + let denied = |k: io::ErrorKind| assert_eq!(k, io::ErrorKind::PermissionDenied); + // resize shrink reaching a Prot tail. + { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 64); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + denied(s.resize(44).unwrap_err().kind()); + assert_eq!(s.resize_as(&prot, 44).unwrap(), 64); + assert_eq!(s.len().unwrap(), 44); + } + // resize grow into a Prot region armed past the tail. + { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 48); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + denied(s.resize(60).unwrap_err().kind()); + assert_eq!(s.resize_as(&prot, 60).unwrap(), 48); + assert_eq!(s.len().unwrap(), 60); + } + // ensure grow into a Prot region. + { + let (s, p) = mk(); + let _g = Guard(p); + seed(&s, 48); + let prot = s.take_protection().unwrap(); + s.protect_as(&prot, 48, 16, BStackAccess::Prot).unwrap(); + denied(s.ensure(60).unwrap_err().kind()); + assert_eq!(s.ensure_as(&prot, 60).unwrap(), 48); + assert_eq!(s.len().unwrap(), 60); + } + } + + #[test] + fn merge_requires_matching_authority() { + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + let owned = alloc.alloc(32).unwrap(); + let full = owned.as_slice(); + let mut left = full.subslice(0, 16); + let right = full.subslice(16, 32); + // Same (NONE) authority: the adjacent subslices merge. + assert!(left.merge_adjacent(&right).is_some()); + // Grant `left` an authority; now the authorities differ and merge refuses. + let prot = alloc.stack().take_protection().unwrap(); + left.authorize(&prot); + assert!(left.merge(&right).is_none()); + assert!(left.merge_adjacent(&right).is_none()); + } + + #[test] + fn allocator_constructor_burns_alloc_authority() { + // Every allocator claims the alloc-authority mint on construction, so no + // external caller can obtain `Alloc` authority over its arena. + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + assert!(alloc.stack().take_alloc_authority().is_none()); + // The guard mint is independent and still available. + assert!(alloc.stack().take_protection().is_some()); + + let (s2, p2) = mk(); + let _g2 = Guard(p2); + let ff = crate::FirstFitBStackAllocator::new(s2).unwrap(); + assert!(ff.stack().take_alloc_authority().is_none()); + } + + #[test] + fn dealloc_refuses_protected_region_then_frees_once_cleared() { + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + let slice = alloc.alloc(32).unwrap(); + let start = slice.start(); + let len = slice.len(); + let prot = alloc.stack().take_protection().unwrap(); + slice.protect_as(&prot, BStackAccess::Prot).unwrap(); + + // Freeing a region that carries caller policy is refused, and the handle + // comes back intact rather than being consumed. + let err = alloc.dealloc(slice).unwrap_err(); + assert_eq!(err.source.kind(), io::ErrorKind::PermissionDenied); + let slice = err.into_handle().expect("region survives a refused free"); + + // The guard holder clears the policy, and the free now succeeds. + alloc + .stack() + .protect_as(&prot, start, len, BStackAccess::All) + .unwrap(); + alloc.dealloc(slice).unwrap(); + assert_eq!(alloc.stack().len().unwrap(), 0); + } + + #[test] + fn dealloc_bulk_refuses_batch_with_any_protected_handle() { + use crate::BStackBulkAllocator; + let (s, p) = mk(); + let _g = Guard(p); + let alloc = crate::LinearBStackAllocator::new(s); + let a = alloc.alloc(16).unwrap(); + let b = alloc.alloc(16).unwrap(); + let prot = alloc.stack().take_protection().unwrap(); + // Arm just one of the two allocations. + b.protect_as(&prot, BStackAccess::ReadOnly).unwrap(); + + let err = alloc.dealloc_bulk([a, b]).unwrap_err(); + assert_eq!(err.source.kind(), io::ErrorKind::PermissionDenied); + // The whole batch is refused up front, so both handles are returned. + assert_eq!(err.into_handles().len(), 2); + } +} diff --git a/src/acl_core.rs b/src/acl_core.rs new file mode 100644 index 0000000..9f68531 --- /dev/null +++ b/src/acl_core.rs @@ -0,0 +1,581 @@ +//! Core primitives for range access control (the `expensive-slice-access-control` +//! feature). +//! +//! This module is the lock-free heart of the range access-control policy layer: +//! the [`BStackAccess`] mode enum, the sorted change-point table +//! ([`PointTable`]) with its lookup / range-check / set operations, and the +//! authority model that decides whether a caller may act on a range. +//! +//! It carries **no** locking, no `BStack` wiring, and no capability-token +//! lifetimes — those live in the `BStack` integration. Here an authority is just +//! the set of tokens a caller *presents* ([`BStackAccessAuthorities`]); the table answers +//! whether that presentation satisfies a mode. +//! +//! # Model +//! +//! Every offset in the payload has a mode. The table stores only the offsets +//! where the mode *changes* (`(offset, mode)`, sorted, coalesced); an absent +//! leading point implies [`All`](BStackAccess::All) from 0, so an unprotected +//! stack is the empty table. Each mode fixes, per axis (read / write / truncate), +//! which authorities satisfy it. The two tokens — a guard and an allocator — are +//! **incomparable**: neither implies the other, so [`Prot`](BStackAccess::Prot) +//! and [`Alloc`](BStackAccess::Alloc) are each private to their own holder. + +/// The three axes a mode governs independently. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum AccessOp { + /// Reading bytes in the range. + Read, + /// Writing bytes in the range (same-length, in place). + Write, + /// Discarding the range by truncation (checked over `[new_len, old_len)`). + Truncate, +} + +/// Which authorities satisfy one cell of the mode table. +/// +/// `Any` needs no token; `None` is satisfiable by no one. The two tokens are +/// incomparable, so `Guard` and `Allocator` are distinct, and +/// `GuardOrAllocator` is the only cell either token satisfies. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum BStackAccessRequirement { + /// Anyone, with or without a token. + Any, + /// The guard token only. + Guard, + /// The allocator token only. + Allocator, + /// Either token (used by [`Rw`](BStackAccess::Rw) truncate). + GuardOrAllocator, + /// No authority satisfies it. + None, +} + +impl BStackAccessRequirement { + /// Whether a caller presenting `held` satisfies this requirement. + #[inline] + #[must_use] + pub const fn satisfied_by(self, held: BStackAccessAuthorities) -> bool { + match self { + Self::Any => true, + Self::Guard => held.guard(), + Self::Allocator => held.alloc(), + Self::GuardOrAllocator => held.guard() || held.alloc(), + Self::None => false, + } + } +} + +/// The tokens a caller presents when acting on a range. Pure input to the +/// checks; the one-shot minting that makes tokens meaningful lives in the +/// `BStack` integration, not here. +/// +/// A bit set: bit 0 is the guard capability, bit 1 the allocator capability. +/// [`Default`] is [`NONE`](Self::NONE). +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub struct BStackAccessAuthorities(u8); + +impl BStackAccessAuthorities { + const GUARD_BIT: u8 = 1 << 0; + const ALLOC_BIT: u8 = 1 << 1; + + /// No tokens — an ordinary caller. + pub const NONE: Self = Self(0); + /// The guard token only. + pub const GUARD: Self = Self(Self::GUARD_BIT); + /// The allocator token only. + pub const ALLOC: Self = Self(Self::ALLOC_BIT); + + /// Whether the guard capability (the protection token) is held. + #[inline] + #[must_use] + pub const fn guard(self) -> bool { + self.0 & Self::GUARD_BIT != 0 + } + + /// Whether the allocator capability (the alloc-authority token) is held. + #[inline] + #[must_use] + pub const fn alloc(self) -> bool { + self.0 & Self::ALLOC_BIT != 0 + } +} + +/// Access mode for a range, one per distinct region of the point table. +/// +/// [`requirement`](BStackAccess::requirement) encodes each mode's per-axis +/// authority. [`All`](BStackAccess::All) is the default everywhere and what an +/// unprotected stack reports. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +#[repr(u8)] +pub enum BStackAccess { + /// Read / write / truncate all open to anyone. The default. + #[default] + All, + /// Read / write open to anyone; truncate needs a guard or the allocator. + Rw, + /// Read / write open to anyone; truncate denied to everyone. + RwStrict, + /// Read / write / truncate all need the guard token. + Prot, + /// Read / write need the guard token; truncate denied to everyone. + RwProt, + /// Read / write / truncate all need the allocator token. + Alloc, + /// Read open to anyone; write and truncate denied. + ReadOnly, + /// Read / write / truncate all denied. + Locked, +} + +impl BStackAccess { + /// The authority that satisfies `op` under this mode. + #[inline] + #[must_use] + pub const fn requirement(self, op: AccessOp) -> BStackAccessRequirement { + use AccessOp::{Read, Truncate, Write}; + use BStackAccessRequirement::{Allocator, Any, Guard, GuardOrAllocator, None}; + match (self, op) { + (Self::All, _) => Any, + + (Self::Rw, Read | Write) => Any, + (Self::Rw, Truncate) => GuardOrAllocator, + + (Self::RwStrict, Read | Write) => Any, + (Self::RwStrict, Truncate) => None, + + (Self::Prot, _) => Guard, + + (Self::RwProt, Read | Write) => Guard, + (Self::RwProt, Truncate) => None, + + (Self::Alloc, _) => Allocator, + + (Self::ReadOnly, Read) => Any, + (Self::ReadOnly, Write | Truncate) => None, + + (Self::Locked, _) => None, + } + } + + /// Whether a caller presenting `held` may perform `op` under this mode. + #[inline] + #[must_use] + pub const fn permits(self, op: AccessOp, held: BStackAccessAuthorities) -> bool { + self.requirement(op).satisfied_by(held) + } +} + +/// A sorted, coalesced table of change points. +/// +/// `(off, mode)` means `mode` applies from `off` until the next point; an absent +/// leading point implies [`All`](BStackAccess::All) from 0. Adjacent equal modes +/// are coalesced, so the table is proportional to the number of distinct regions +/// and one protected header is two entries. The vec is read far more often than +/// written, so a read is a `partition_point` over a contiguous array rather than +/// a tree walk. +// Scaffolding for the not-yet-wired BStack integration; unused until then. +#[allow(dead_code)] +#[derive(Clone, Debug, Default)] +pub(crate) struct PointTable { + /// Change points, strictly increasing in offset, no adjacent equal modes, + /// no leading `All`. + points: Vec<(u64, BStackAccess)>, +} + +#[allow(dead_code)] +impl PointTable { + /// An empty table: [`All`](BStackAccess::All) everywhere. + #[inline] + #[must_use] + pub const fn new() -> Self { + Self { points: Vec::new() } + } + + /// Whether the table is [`All`](BStackAccess::All) everywhere. The + /// `AtomicBool` short-circuit in the `BStack` integration mirrors this. + #[inline] + #[must_use] + pub fn is_unprotected(&self) -> bool { + self.points.is_empty() + } + + /// The stored change points, for inspection and testing. + #[inline] + #[must_use] + pub fn points(&self) -> &[(u64, BStackAccess)] { + &self.points + } + + /// The mode in effect at `off`. + /// + /// `partition_point(|p| p.0 <= off) - 1`, with `All` when no point starts at + /// or before `off`. Points beyond `len` are retained, so a mode set before + /// its bytes arrive still reports here. + #[inline] + #[must_use] + pub fn mode_at(&self, off: u64) -> BStackAccess { + let i = self.points.partition_point(|p| p.0 <= off); + if i == 0 { + BStackAccess::All + } else { + self.points[i - 1].1 + } + } + + /// Whether every offset in `[a, b)` is currently [`All`](BStackAccess::All). + /// + /// The tokenless authorization rule for arming a range: a caller with no + /// token may only tighten a range that is entirely unprotected. An empty + /// range is trivially uniform. + #[must_use] + pub fn all_over(&self, a: u64, b: u64) -> bool { + if a >= b { + return true; + } + if self.mode_at(a) != BStackAccess::All { + return false; + } + // Any change point inside `[a, b)` breaks uniformity. + let i = self.points.partition_point(|p| p.0 <= a); + i >= self.points.len() || self.points[i].0 >= b + } + + /// Whether every offset in `[a, b)` permits `op` under `held`. + /// + /// One `partition_point` locates the mode at `a`; a forward scan folds in + /// each point strictly inside `(a, b)`, stopping at the first denial. The + /// common case spans one region and the scan never runs. An empty range + /// (`a >= b`) is trivially permitted. + #[must_use] + pub fn check(&self, a: u64, b: u64, op: AccessOp, held: BStackAccessAuthorities) -> bool { + if a >= b { + return true; + } + // Mode governing the start of the range. + let start = self.points.partition_point(|p| p.0 <= a); + let mode = if start == 0 { + BStackAccess::All + } else { + self.points[start - 1].1 + }; + if !mode.permits(op, held) { + return false; + } + // Every point strictly inside (a, b) opens a new region to check. + let mut j = start; + while j < self.points.len() && self.points[j].0 < b { + if !self.points[j].1.permits(op, held) { + return false; + } + j += 1; + } + true + } + + /// Whether every offset in `[a, b)` is [`All`](BStackAccess::All) or + /// [`Alloc`](BStackAccess::Alloc) — the modes an allocator may clear when it + /// reclaims the range on `dealloc`. + /// + /// Any caller-set mode (`Rw`/`RwStrict`/`Prot`/`RwProt`/`ReadOnly`/`Locked`) + /// makes the range non-reclaimable: the free must be refused so the policy is + /// never silently discarded. Folds the range exactly like [`check`](Self::check); + /// an empty range is trivially reclaimable. + #[must_use] + pub fn reclaimable_by_alloc(&self, a: u64, b: u64) -> bool { + if a >= b { + return true; + } + let is_ok = |m: BStackAccess| matches!(m, BStackAccess::All | BStackAccess::Alloc); + let start = self.points.partition_point(|p| p.0 <= a); + let mode = if start == 0 { + BStackAccess::All + } else { + self.points[start - 1].1 + }; + if !is_ok(mode) { + return false; + } + let mut j = start; + while j < self.points.len() && self.points[j].0 < b { + if !is_ok(self.points[j].1) { + return false; + } + j += 1; + } + true + } + + /// Set `[a, b)` to `mode`, leaving everything at or after `b` unchanged. + /// + /// Splices in the up-to-two boundaries the range needs — `(a, mode)` and a + /// `(b, resume)` restoring the mode that governed `b` — dropping each when it + /// would duplicate the region it opens against, and dropping the following + /// point when the tail now resumes its mode. Two binary searches locate the + /// affected span; the rest is proportional to the points in `[a, b]` plus the + /// array shift, with no full-table pass. An empty range is a no-op. This is + /// the raw table mutation: it does **not** consult authorities — the `BStack` + /// integration gates who may call it. + pub fn set(&mut self, a: u64, b: u64, mode: BStackAccess) { + if a >= b { + return; + } + // Mode governing b now; restored as the tail resumes after the set. + let hi = self.points.partition_point(|p| p.0 <= b); + let resume = if hi == 0 { + BStackAccess::All + } else { + self.points[hi - 1].1 + }; + // First point at or after a, and the mode of the region entering a. + let lo = self.points.partition_point(|p| p.0 < a); + let before = if lo == 0 { + BStackAccess::All + } else { + self.points[lo - 1].1 + }; + + // Boundaries for [a, b): each omitted when it repeats the mode before it. + // The point after the range (index `hi`) never needs dropping: it already + // differed from `resume` in the coalesced table, and `resume` is the mode + // the tail now resumes with, so the join stays distinct. + let mut mid = [(0u64, BStackAccess::All); 2]; + let mut n = 0; + if mode != before { + mid[n] = (a, mode); + n += 1; + } + if resume != mode { + mid[n] = (b, resume); + n += 1; + } + + self.points.splice(lo..hi, mid[..n].iter().copied()); + } +} + +#[cfg(test)] +mod tests { + use super::*; + use AccessOp::{Read, Truncate, Write}; + use BStackAccess::{All, Alloc, Locked, Prot, ReadOnly, Rw, RwProt, RwStrict}; + + const NONE: BStackAccessAuthorities = BStackAccessAuthorities::NONE; + const GUARD: BStackAccessAuthorities = BStackAccessAuthorities::GUARD; + const ALLOC: BStackAccessAuthorities = BStackAccessAuthorities::ALLOC; + + #[test] + fn requirement_matrix_is_correct() { + // All: everything to anyone. + for op in [Read, Write, Truncate] { + assert!(All.permits(op, NONE)); + } + // Rw: rw to anyone, truncate to either token but not the tokenless. + assert!(Rw.permits(Write, NONE)); + assert!(!Rw.permits(Truncate, NONE)); + assert!(Rw.permits(Truncate, GUARD)); + assert!(Rw.permits(Truncate, ALLOC)); + // RwStrict: rw to anyone, truncate to no one. + assert!(RwStrict.permits(Write, NONE)); + assert!(!RwStrict.permits(Truncate, ALLOC)); + // Prot vs Alloc are incomparable on every axis. + for op in [Read, Write, Truncate] { + assert!(Prot.permits(op, GUARD)); + assert!(!Prot.permits(op, ALLOC)); + assert!(Alloc.permits(op, ALLOC)); + assert!(!Alloc.permits(op, GUARD)); + } + // RwProt: rw to the guard, truncate to no one. + assert!(RwProt.permits(Write, GUARD)); + assert!(!RwProt.permits(Truncate, GUARD)); + // ReadOnly / Locked. + assert!(ReadOnly.permits(Read, NONE)); + assert!(!ReadOnly.permits(Write, GUARD)); + assert!(!Locked.permits(Read, GUARD)); + } + + #[test] + fn empty_table_is_all() { + let t = PointTable::new(); + assert!(t.is_unprotected()); + assert_eq!(t.mode_at(0), All); + assert_eq!(t.mode_at(1_000_000), All); + assert!(t.check(0, 100, Write, NONE)); + } + + #[test] + fn mode_at_boundaries() { + let mut t = PointTable::new(); + t.set(16, 32, Alloc); + // [0,16) All, [16,32) Alloc, [32,..) All. + assert_eq!(t.mode_at(15), All); + assert_eq!(t.mode_at(16), Alloc); + assert_eq!(t.mode_at(31), Alloc); + assert_eq!(t.mode_at(32), All); + // One protected header is exactly two entries. + assert_eq!(t.points(), &[(16, Alloc), (32, All)]); + } + + #[test] + fn range_check_folds_most_restrictive() { + let mut t = PointTable::new(); + t.set(16, 32, Alloc); + // A read fully inside the Alloc region needs the alloc token. + assert!(t.check(20, 24, Read, ALLOC)); + assert!(!t.check(20, 24, Read, GUARD)); + // A range spanning All -> Alloc -> All is denied without the token, + // because the Alloc slice in the middle dominates. + assert!(!t.check(0, 64, Read, NONE)); + assert!(t.check(0, 64, Read, ALLOC)); + // A range entirely outside the protected slice is open. + assert!(t.check(0, 16, Write, NONE)); + assert!(t.check(32, 64, Write, NONE)); + } + + #[test] + fn set_coalesces_adjacent_equal_modes() { + let mut t = PointTable::new(); + t.set(10, 20, Prot); + t.set(20, 30, Prot); + // Two adjacent Prot regions collapse to one. + assert_eq!(t.points(), &[(10, Prot), (30, All)]); + } + + #[test] + fn set_over_all_leaves_no_leading_point() { + let mut t = PointTable::new(); + t.set(0, 50, All); + assert!(t.is_unprotected()); + } + + #[test] + fn set_overwrites_inner_points() { + let mut t = PointTable::new(); + t.set(10, 20, Prot); + t.set(30, 40, Alloc); + // Overwrite a span that swallows the first region and part of the gap. + t.set(5, 35, Rw); + // [0,5) All, [5,35) Rw, [35,40) Alloc, [40,..) All. + assert_eq!(t.points(), &[(5, Rw), (35, Alloc), (40, All)]); + assert_eq!(t.mode_at(4), All); + assert_eq!(t.mode_at(5), Rw); + assert_eq!(t.mode_at(35), Alloc); + assert_eq!(t.mode_at(40), All); + } + + #[test] + fn set_restores_resume_mode_at_b() { + let mut t = PointTable::new(); + t.set(10, 40, Alloc); + // Carve a Prot window inside the Alloc region; Alloc must resume after. + t.set(20, 30, Prot); + assert_eq!(t.mode_at(19), Alloc); + assert_eq!(t.mode_at(20), Prot); + assert_eq!(t.mode_at(30), Alloc); + assert_eq!(t.mode_at(40), All); + } + + #[test] + fn empty_range_is_noop_and_permitted() { + let mut t = PointTable::new(); + t.set(10, 20, Locked); + assert!(t.check(15, 15, Read, NONE)); + t.set(15, 15, All); + assert_eq!(t.mode_at(15), Locked); + } + + #[test] + fn all_over_detects_uniform_all() { + let mut t = PointTable::new(); + assert!(t.all_over(0, 100)); + assert!(t.all_over(50, 50)); // empty + t.set(20, 40, Prot); + assert!(t.all_over(0, 20)); // touches boundary but stays All + assert!(!t.all_over(0, 21)); // reaches into Prot + assert!(!t.all_over(25, 30)); // inside Prot + assert!(t.all_over(40, 100)); // All resumes after + assert!(!t.all_over(30, 60)); // spans Prot -> All + } + + #[test] + fn reclaimable_by_alloc_accepts_all_and_alloc_only() { + let mut t = PointTable::new(); + assert!(t.reclaimable_by_alloc(0, 100)); // empty table = All + assert!(t.reclaimable_by_alloc(50, 50)); // empty range + t.set(16, 32, Alloc); + assert!(t.reclaimable_by_alloc(0, 64)); // All -> Alloc -> All, all clearable + t.set(40, 48, Prot); + assert!(!t.reclaimable_by_alloc(0, 64)); // Prot region blocks reclaim + assert!(t.reclaimable_by_alloc(0, 40)); // stops before the Prot region + assert!(!t.reclaimable_by_alloc(40, 48)); // exactly the Prot region + // Every non-All/non-Alloc mode blocks a reclaim. + for m in [Rw, RwStrict, RwProt, ReadOnly, Locked] { + let mut u = PointTable::new(); + u.set(10, 20, m); + assert!(!u.reclaimable_by_alloc(0, 30), "{m:?} should block reclaim"); + } + } + + #[test] + fn points_beyond_len_are_retained() { + let mut t = PointTable::new(); + // Arm a region before its bytes exist. + t.set(1000, 2000, Prot); + assert_eq!(t.mode_at(1500), Prot); + assert_eq!(t.points(), &[(1000, Prot), (2000, All)]); + } + + // Naive reference: materialize per-offset modes and re-derive the change + // points. Deliberately O(span) and obviously correct. + fn reference_set(cells: &mut [BStackAccess], a: u64, b: u64, mode: BStackAccess) { + for c in &mut cells[a as usize..b as usize] { + *c = mode; + } + } + fn cells_to_points(cells: &[BStackAccess]) -> Vec<(u64, BStackAccess)> { + let mut out = Vec::new(); + let mut prev = All; + for (i, &m) in cells.iter().enumerate() { + if m != prev { + out.push((i as u64, m)); + prev = m; + } + } + out + } + + #[test] + fn set_matches_naive_reference_randomized() { + use rand::RngExt; + use rand::SeedableRng; + let modes = [All, Rw, RwStrict, Prot, RwProt, Alloc, ReadOnly, Locked]; + let mut rng = rand::rngs::StdRng::seed_from_u64(0xB57A_C0DE); + const SPAN: u64 = 32; + for _ in 0..2000 { + let mut t = PointTable::new(); + // One extra sentinel cell (always All) captures the closing boundary + // a set with `b == SPAN` records at offset SPAN. + let mut cells = [All; SPAN as usize + 1]; + for _ in 0..8 { + let x = rng.random_range(0..=SPAN); + let y = rng.random_range(0..=SPAN); + let (a, b) = (x.min(y), x.max(y)); + let m = modes[rng.random_range(0..modes.len())]; + t.set(a, b, m); + reference_set(&mut cells, a, b, m); + // Points always match the coalesced reference derivation. + assert_eq!(t.points(), cells_to_points(&cells).as_slice()); + // Invariant: strictly increasing offsets, no adjacent equal modes, + // no leading All. + let pts = t.points(); + for w in pts.windows(2) { + assert!(w[0].0 < w[1].0); + assert_ne!(w[0].1, w[1].1); + } + if let Some(first) = pts.first() { + assert_ne!(first.1, All); + } + } + } + } +} diff --git a/src/alloc/checked_slab.rs b/src/alloc/checked_slab.rs index 3c7e64a..bfcb6fa 100644 --- a/src/alloc/checked_slab.rs +++ b/src/alloc/checked_slab.rs @@ -15,6 +15,7 @@ use super::{BStackBulkAllocError, BStackBulkAllocator, ensure_own_handles}; use crate::BStack; #[cfg(feature = "atomic")] use crate::BStackGenOp; +use crate::acl::alloc_meta; #[cfg(feature = "atomic")] use crate::{bstack_unsafe_reborrow, bstack_unsafe_reborrow_mut}; #[cfg(not(feature = "atomic"))] @@ -175,6 +176,10 @@ const ALCK_MAGIC_PREFIX: [u8; 6] = *b"ALCK\x00\x01"; #[cfg(feature = "set")] pub struct CheckedSlabBStackAllocator { stack: BStack, + /// The real allocator capability, minted from `stack` at construction and + /// presented to the `_as` ops by the `alloc_meta!` forwarders. + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth: crate::BStackAllocAuthority, /// Cached from the on-disk header; fixed for the lifetime of the allocator. /// Covers the full block including the 8-byte overhead; must be `≥ 16`. block_size: u64, @@ -263,6 +268,16 @@ impl CheckedSlabBStackAllocator { /// file). /// * Any [`io::Error`] propagated from the underlying [`BStack`] operations. pub fn new(stack: BStack, data_size: u64) -> io::Result { + // Acquire the allocator authority up front, so every metadata write below + // is made as the allocator. Refused (rather than panicking) if the stack's + // permit has already been taken. + #[cfg(feature = "expensive-slice-access-control")] + let alloc_auth = stack.take_alloc_authority().ok_or_else(|| { + io_error!( + PermissionDenied, + "CheckedSlabBStackAllocator: alloc authority already taken from this stack" + ) + })?; if !stack.is_empty()? { return Err(io_error!( InvalidInput, @@ -291,7 +306,14 @@ impl CheckedSlabBStackAllocator { write_buf!(block_size => hdr, off + 8); // free_head at off+16 remains 0 (SENTINEL) stack.push(hdr)?; + // The whole header stays `Alloc` for the allocator's lifetime; its own + // metadata I/O goes through the `meta_*` helpers, which present allocator + // authority. No external `stack().set(..)` can corrupt magic, geometry, or + // the free-list head. + stack.acl_mark_alloc(Self::OFFSET_SIZE, Self::HEADER_SIZE)?; Ok(Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, block_size, #[cfg(feature = "atomic")] @@ -321,6 +343,16 @@ impl CheckedSlabBStackAllocator { /// `block_size`, misaligned arena, or an invalid `free_head`. /// * Any [`io::Error`] propagated from the underlying [`BStack`] operations. pub fn open(stack: BStack) -> io::Result { + // Acquire the allocator authority up front and read the header through it, + // so a reopen sees its own `Alloc`-marked header (marks are not persisted, + // but a same-object reopen still carries them). + #[cfg(feature = "expensive-slice-access-control")] + let alloc_auth = stack.take_alloc_authority().ok_or_else(|| { + io_error!( + PermissionDenied, + "CheckedSlabBStackAllocator: alloc authority already taken from this stack" + ) + })?; if stack.is_empty()? { return Err(io_error!( InvalidInput, @@ -337,6 +369,9 @@ impl CheckedSlabBStackAllocator { } let mut header = [0u8; Self::HEADER_SIZE as usize]; + #[cfg(feature = "expensive-slice-access-control")] + stack.get_into_as(&alloc_auth, Self::OFFSET_SIZE, &mut header)?; + #[cfg(not(feature = "expensive-slice-access-control"))] stack.get_into(Self::OFFSET_SIZE, &mut header)?; if header[..ALCK_MAGIC_PREFIX.len()] != ALCK_MAGIC_PREFIX { @@ -390,6 +425,8 @@ impl CheckedSlabBStackAllocator { } let allocator = Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, block_size: stored_block_size, #[cfg(feature = "atomic")] @@ -397,6 +434,11 @@ impl CheckedSlabBStackAllocator { #[cfg(not(feature = "atomic"))] _not_sync: PhantomData, }; + // Re-arm the header mark on every open (policy is not persisted); the + // free-list I/O in `recover` below goes through the `meta_*` helpers. + allocator + .stack + .acl_mark_alloc(Self::OFFSET_SIZE, Self::HEADER_SIZE)?; // Reclaim leaks and repair a failed tail truncation left by an unclean // shutdown. Best-effort: the residual unsure-block count is discarded // here; call [`recover`](Self::recover) explicitly to inspect it. @@ -533,7 +575,7 @@ impl CheckedSlabBStackAllocator { let mut node_buf = [0u8; 16]; let mut oh_buf = [0u8; 8]; - self.stack.process_gen(|| { + alloc_meta!(self, process_gen, process_gen_as, || { loop { match st { // Read the authoritative payload size first. @@ -824,7 +866,17 @@ impl CheckedSlabBStackAllocator { fn scan_free_list(&self, stack_len: u64) -> io::Result<(Vec, bool)> { let mut free = Vec::new(); let mut seen: HashSet = HashSet::new(); - let mut head = u64::from_le_bytes(read_bstack!(self.stack, Self::FREE_HEAD_OFFSET => u64)); + let mut head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + }; let mut corrupt = false; while head != Self::SENTINEL { if head < Self::ARENA_START @@ -987,7 +1039,7 @@ impl CheckedSlabBStackAllocator { let mut head_opt: Option = None; let mut corrupt: Option = None; - self.stack.process_gen(|| { + alloc_meta!(self, process_gen, process_gen_as, || { let op = match step { // Step 0: read the current free-list head. 0 => Some(BStackGenOp::Read { @@ -1070,7 +1122,17 @@ impl CheckedSlabBStackAllocator { /// two writes merely leaks the detached block. #[cfg(not(feature = "atomic"))] fn pop_and_claim_block(&self, num_blocks: u64, init: bool) -> io::Result> { - let head = u64::from_le_bytes(read_bstack!(self.stack, Self::FREE_HEAD_OFFSET => u64)); + let head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + }; if head == Self::SENTINEL { return Ok(None); } @@ -1086,7 +1148,7 @@ impl CheckedSlabBStackAllocator { )); } // Advance free_head to the next block (stored in data[0..8]). - self.stack.set(Self::FREE_HEAD_OFFSET, &prefix[8..16])?; + alloc_meta!(self, set, set_as, Self::FREE_HEAD_OFFSET, &prefix[8..16])?; // Mark in-use and, when the caller is owed zeroes, scrub the data in the // same write. `false` writes just the overhead word — still one // `set`, but `block_size - OVERHEAD` fewer bytes — and leaves the data @@ -1126,7 +1188,18 @@ impl CheckedSlabBStackAllocator { fn write_free_run(&self, first_block: u64, count: u64) -> io::Result<()> { debug_assert!(count > 0); #[cfg(not(feature = "atomic"))] - let old_head = read_bstack!(self.stack, Self::FREE_HEAD_OFFSET => u64); + let old_head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + } + .to_le_bytes(); let total = count .checked_mul(self.block_size) .ok_or_else(|| io_error!(InvalidInput, "freed region size overflows u64"))?; @@ -1194,13 +1267,24 @@ impl CheckedSlabBStackAllocator { io_error!(InvalidInput, "last free-list offset overflows u64") })?) .ok_or_else(|| io_error!(InvalidInput, "last block offset overflows u64"))?; - self.stack - .cross_exchange(last_block + Self::OVERHEAD, Self::FREE_HEAD_OFFSET, 8) + alloc_meta!( + self, + cross_exchange, + cross_exchange_as, + last_block + Self::OVERHEAD, + Self::FREE_HEAD_OFFSET, + 8 + ) } #[cfg(not(feature = "atomic"))] { - self.stack - .set(Self::FREE_HEAD_OFFSET, first_block.to_le_bytes()) + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + first_block.to_le_bytes() + ) } } } @@ -1491,7 +1575,10 @@ impl CheckedSlabBStackAllocator { // the pre-built run onto free_head with one cross_exchange. if !self.stack.try_discard(sentinel, excess_backing)? { let last_block = excess_start + (excess_count - 1) * self.block_size; - self.stack.cross_exchange( + alloc_meta!( + self, + cross_exchange, + cross_exchange_as, last_block + Self::OVERHEAD, Self::FREE_HEAD_OFFSET, 8, @@ -1502,8 +1589,13 @@ impl CheckedSlabBStackAllocator { { // Non-tail (tail handled above): the run already points at the old // head, so publishing its first block as the new head splices it. - self.stack - .set(Self::FREE_HEAD_OFFSET, excess_start.to_le_bytes())?; + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + excess_start.to_le_bytes() + )?; } // SAFETY: // 1. No overflow: slice.start() + new_len ≤ block_start + OVERHEAD + new_n * block_size − OVERHEAD ≤ u64::MAX @@ -1543,6 +1635,10 @@ impl BStackAllocator for CheckedSlabBStackAllocator { #[inline] fn into_stack(self) -> BStack { + // Hand the allocator capability back so a caller that re-wraps the + // reclaimed stack can mint it again. + #[cfg(feature = "expensive-slice-access-control")] + self.stack.return_alloc_authority(self.alloc_auth); self.stack } @@ -1600,6 +1696,9 @@ impl BStackAllocator for CheckedSlabBStackAllocator { let slice = ensure_own_handle(self, slice, "CheckedSlabBStackAllocator::dealloc")?; let start = slice.start(); let len = slice.len(); + if let Err(source) = self.stack.acl_reclaim(start, len) { + return Err(BStackAllocError::with_handle(source, slice)); + } // Set once the caller's blocks may have been partially freed, after // which returning the handle for retry would risk a double-free. let mut lost = false; @@ -1750,7 +1849,7 @@ impl CheckedSlabBStackAllocator { } let mut st = St::ReadHead; - self.stack.process_gen(|| { + alloc_meta!(self, process_gen, process_gen_as, || { loop { match st { St::ReadHead => { @@ -1885,7 +1984,7 @@ impl CheckedSlabBStackAllocator { } let mut st = St::ReadHead; - self.stack.inplace_gen(|prev| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |prev| { // A failed read reports its error here, not as a return value, and // leaves the buffer unfilled: bail before consuming stale bytes. if let Err(e) = prev { @@ -2042,7 +2141,7 @@ impl CheckedSlabBStackAllocator { b }; let mut i = 0usize; - self.stack.inplace_gen(|_prev| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |_prev| { // `i` walks `jobs` (collected before this call), so `get` is in bounds // until it runs off the end, where `?` ends the sequence. let (off, overhead) = jobs.get(i)?; @@ -2083,8 +2182,14 @@ impl CheckedSlabBStackAllocator { batch.push((blocks[i], buf)); } self.stack.set_batched(batch)?; - self.stack - .cross_exchange(blocks[k - 1] + Self::OVERHEAD, Self::FREE_HEAD_OFFSET, 8) + alloc_meta!( + self, + cross_exchange, + cross_exchange_as, + blocks[k - 1] + Self::OVERHEAD, + Self::FREE_HEAD_OFFSET, + 8 + ) } /// Best-effort cleanup after an `alloc_bulk` extend path fails partway. @@ -2310,6 +2415,11 @@ impl BStackBulkAllocator for CheckedSlabBStackAllocator { ) -> Result<(), BStackBulkAllocError<'a, Self>> { let slices: Vec> = handles.into_iter().collect(); let slices = ensure_own_handles(self, slices, "CheckedSlabBStackAllocator::dealloc_bulk")?; + for s in &slices { + if let Err(source) = self.stack.acl_reclaimable(s.start(), s.len()) { + return Err(BStackBulkAllocError::with_handles(source, slices)); + } + } let bs = self.block_size; let mut freeing = false; @@ -2920,7 +3030,7 @@ mod tests { let _c = alloc.alloc(8).unwrap(); // block 80 alloc.dealloc(b).unwrap(); // free_head = 64 // Simulate a pop crash: advance free_head past b without claiming it. - alloc.stack().set(40, 0u64.to_le_bytes()).unwrap(); + crate::acl::alloc_meta!(alloc, set, set_as, 40, 0u64.to_le_bytes()).unwrap(); // Block 64 is now leaked (overhead 0, not in the free list). assert_eq!(alloc.recover().unwrap(), 0); // The reclaimed block is handed back out on the next allocation. @@ -2983,7 +3093,7 @@ mod tests { let b = alloc.alloc(8).unwrap(); let _c = alloc.alloc(8).unwrap(); alloc.dealloc(b).unwrap(); - alloc.stack().set(40, 0u64.to_le_bytes()).unwrap(); // leak block 64 + crate::acl::alloc_meta!(alloc, set, set_as, 40, 0u64.to_le_bytes()).unwrap(); // leak block 64 assert_eq!(alloc.recover().unwrap(), 0); let len_after = alloc.stack().len().unwrap(); // A second run finds nothing further and changes nothing. diff --git a/src/alloc/first_fit.rs b/src/alloc/first_fit.rs index 01cc594..59567a3 100644 --- a/src/alloc/first_fit.rs +++ b/src/alloc/first_fit.rs @@ -5,6 +5,7 @@ use super::{ use crate::BStack; #[cfg(feature = "atomic")] use crate::BStackGenOp; +use crate::acl::alloc_meta; #[cfg(feature = "atomic")] use crate::{bstack_unsafe_reborrow, bstack_unsafe_reborrow_mut}; #[cfg(not(feature = "atomic"))] @@ -178,6 +179,10 @@ const ALFF_MAGIC_PREFIX: [u8; 6] = *b"ALFF\x00\x01"; #[cfg(feature = "set")] pub struct FirstFitBStackAllocator { stack: BStack, + /// The real allocator capability, minted from `stack` at construction and + /// presented to the `_as` ops by the `meta_*` forwarders. + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth: crate::BStackAllocAuthority, /// Serialises the operations that span multiple [`BStack`] calls and are /// therefore not made atomic by `BStack`'s own locking: free-list mutation /// and stack extension/discard. It is **not** taken for general writes @@ -241,6 +246,16 @@ impl FirstFitBStackAllocator { /// type). /// * Any [`io::Error`] propagated from the underlying [`BStack`] operations. pub fn new(stack: BStack) -> Result { + // Acquire the allocator authority up front and use it for the header I/O + // below, so a same-object reopen reads its own `Alloc`-marked header. + // Refused (not panic) if the stack's permit has already been taken. + #[cfg(feature = "expensive-slice-access-control")] + let alloc_auth = stack.take_alloc_authority().ok_or_else(|| { + io_error!( + PermissionDenied, + "FirstFitBStackAllocator: alloc authority already taken from this stack" + ) + })?; // Initialize empty stack with allocator header if stack.is_empty()? { let mut hdr = [0u8; (Self::OFFSET_SIZE + Self::HEADER_SIZE) as usize]; @@ -248,7 +263,12 @@ impl FirstFitBStackAllocator { .copy_from_slice(&ALFF_MAGIC); // flags, _reserved, free_head remain zero stack.push(hdr)?; + // Header (magic, recovery flag, free_head) stays `Alloc` for the + // allocator's lifetime; own I/O via the `meta_*` helpers. + stack.acl_mark_alloc(Self::OFFSET_SIZE, Self::HEADER_SIZE)?; return Ok(Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, #[cfg(feature = "atomic")] lock: Mutex::new(()), @@ -267,6 +287,9 @@ impl FirstFitBStackAllocator { )); } let mut header = [0u8; Self::HEADER_SIZE as usize]; + #[cfg(feature = "expensive-slice-access-control")] + stack.get_into_as(&alloc_auth, Self::OFFSET_SIZE, &mut header)?; + #[cfg(not(feature = "expensive-slice-access-control"))] stack.get_into(Self::OFFSET_SIZE, &mut header)?; // Check magic prefix for compatibility with 0.1.x files. if header[..ALFF_MAGIC_PREFIX.len()] != ALFF_MAGIC_PREFIX { @@ -292,6 +315,8 @@ impl FirstFitBStackAllocator { } } let alloc = Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, #[cfg(feature = "atomic")] lock: Mutex::new(()), @@ -300,6 +325,11 @@ impl FirstFitBStackAllocator { #[cfg(not(feature = "atomic"))] _not_sync: PhantomData, }; + // Re-arm the header mark on reopen (policy is not persisted); the free-list + // I/O in `recovery` below goes through the `meta_*` helpers. + alloc + .stack + .acl_mark_alloc(Self::OFFSET_SIZE, Self::HEADER_SIZE)?; if recovery_needed { alloc.recovery()?; } @@ -309,8 +339,13 @@ impl FirstFitBStackAllocator { #[cfg(not(feature = "atomic"))] #[inline] fn set_recovery_needed(&self) -> io::Result<()> { - self.stack - .set(Self::OFFSET_SIZE + 8, 1u32.to_le_bytes().as_slice()) + alloc_meta!( + self, + set, + set_as, + Self::OFFSET_SIZE + 8, + 1u32.to_le_bytes().as_slice() + ) } #[cfg(feature = "atomic")] @@ -322,7 +357,10 @@ impl FirstFitBStackAllocator { // failure means the flag was left set by a previously crashed or failed // operation, so the stack needs recovery (reopen) before it is safe to // mutate, and we surface that as an error rather than proceeding. - if !self.stack.cas( + if !alloc_meta!( + self, + cas, + cas_as, Self::OFFSET_SIZE + 8, [0u8; 4].as_slice(), 1u32.to_le_bytes().as_slice(), @@ -373,7 +411,13 @@ impl FirstFitBStackAllocator { #[cfg(not(feature = "atomic"))] #[inline] fn clear_recovery_needed(&self) -> io::Result<()> { - self.stack.set(Self::OFFSET_SIZE + 8, [0u8; 4].as_slice()) + alloc_meta!( + self, + set, + set_as, + Self::OFFSET_SIZE + 8, + [0u8; 4].as_slice() + ) } #[cfg(feature = "atomic")] @@ -384,7 +428,10 @@ impl FirstFitBStackAllocator { // this CAS is a no-cost check over the disk write. A failure means the // flag was not set when we expected it to be, indicating the paired set // was lost or the flag was disturbed out of band. - if !self.stack.cas( + if !alloc_meta!( + self, + cas, + cas_as, Self::OFFSET_SIZE + 8, 1u32.to_le_bytes().as_slice(), [0u8; 4].as_slice(), @@ -523,7 +570,13 @@ impl FirstFitBStackAllocator { if prev != 0 { self.stack.set(prev, next.to_le_bytes())?; } else { - self.stack.set(Self::FREE_HEAD_OFFSET, next.to_le_bytes())?; + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + next.to_le_bytes() + )?; } if next != 0 { self.stack.set(next + 8, prev.to_le_bytes())?; @@ -618,9 +671,17 @@ impl FirstFitBStackAllocator { // free_head <- result_block -> next // free_head --------------------> next -> ... // free_head <------------------- next <- ... - let mut head_buf = [0u8; 8]; - self.stack.get_into(Self::FREE_HEAD_OFFSET, &mut head_buf)?; - let old_head = u64::from_le_bytes(head_buf); + let old_head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + }; // The old head becomes our next_free and takes a back-link at old_head + 8, // so validate it like any other free-list link before writing through it. if !Self::is_valid_link_ptr(old_head, stack_len) { @@ -641,8 +702,13 @@ impl FirstFitBStackAllocator { // free_head -> result_block -> next -> ... // free_head <------------------ next <- ... // If this step fails, the free list is still consistent but the result block is orphaned - self.stack - .set(Self::FREE_HEAD_OFFSET, result_start.to_le_bytes())?; + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + result_start.to_le_bytes() + )?; // After adding result block: // free_head -> result_block -> next -> ... @@ -742,7 +808,7 @@ impl FirstFitBStackAllocator { let mut head_buf = [0u8; 8]; let mut back_buf = [0u8; 8]; - self.stack.inplace_gen(|feedback| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |feedback| { // A failed read or rejected write must tear the batch down, not commit // a partial one if let Err(e) = feedback { @@ -1098,10 +1164,17 @@ impl FirstFitBStackAllocator { / (Self::MIN_BLOCK_PAYLOAD_SIZE + Self::BLOCK_OVERHEAD_SIZE) + 1; let mut walk_count = 0u64; - let mut free_head_buf = [0u8; 8]; - self.stack - .get_into(Self::FREE_HEAD_OFFSET, &mut free_head_buf)?; - let mut head = u64::from_le_bytes(free_head_buf); + let mut head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + }; while head != 0 { walk_count += 1; if walk_count > max_walk { @@ -1258,7 +1331,13 @@ impl FirstFitBStackAllocator { if prev != 0 { self.stack.set(prev, next.to_le_bytes())?; } else { - self.stack.set(Self::FREE_HEAD_OFFSET, next.to_le_bytes())?; + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + next.to_le_bytes() + )?; } // Then commit forward pointer @@ -1368,7 +1447,7 @@ impl FirstFitBStackAllocator { let mut rd16 = [0u8; 16]; let zeros = [0u8; 8]; - self.stack.inplace_gen(|feedback| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |feedback| { if let Err(e) = feedback { return Some(BStackGenOp::Abort { source: Some(e) }); } @@ -1651,13 +1730,24 @@ impl FirstFitBStackAllocator { // Update free_head to the first free block found, or 0 if none. let new_free_head = free_blocks.first().copied().unwrap_or(0); - self.stack - .set(Self::FREE_HEAD_OFFSET, new_free_head.to_le_bytes())?; + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + new_free_head.to_le_bytes() + )?; // Authoritative reset: recovery may have been triggered with the on-disk flag already // clear (e.g. an out-of-range free_head in `new`), so write 0 directly rather than via // the CAS clear, which under the `atomic` feature would fail when the flag is not 1. - self.stack.set(Self::OFFSET_SIZE + 8, [0u8; 4].as_slice()) + alloc_meta!( + self, + set, + set_as, + Self::OFFSET_SIZE + 8, + [0u8; 4].as_slice() + ) } } @@ -1673,6 +1763,10 @@ impl BStackAllocator for FirstFitBStackAllocator { #[inline] fn into_stack(self) -> BStack { + // Hand the allocator capability back so a caller that re-wraps the + // reclaimed stack can mint it again. + #[cfg(feature = "expensive-slice-access-control")] + self.stack.return_alloc_authority(self.alloc_auth); self.stack } @@ -1688,6 +1782,9 @@ impl BStackAllocator for FirstFitBStackAllocator { let slice = ensure_own_handle(self, slice, "FirstFitBStackAllocator::dealloc")?; let start = slice.start(); let len = slice.len(); + if let Err(source) = self.stack.acl_reclaim(start, len) { + return Err(BStackAllocError::with_handle(source, slice)); + } // Set to true once the block is being physically reclaimed and can no // longer be safely handed back to the caller. let mut lost = false; @@ -2024,9 +2121,17 @@ impl FirstFitBStackAllocator { self.stack.get_into(next_block, &mut link_buf)?; let nnext = read_buf_le!(link_buf, 0 => u64); // next.next_free let nprev = read_buf_le!(link_buf, 8 => u64); // next.prev_free - let mut head_buf = [0u8; 8]; - self.stack.get_into(Self::FREE_HEAD_OFFSET, &mut head_buf)?; - let free_head = u64::from_le_bytes(head_buf); + let free_head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + }; if !Self::is_valid_link_ptr(nnext, stack_len) || !Self::is_valid_link_ptr(nprev, stack_len) || !Self::is_valid_link_ptr(free_head, stack_len) @@ -2107,7 +2212,7 @@ impl FirstFitBStackAllocator { (0, NONE) }, ]; - self.stack.set_batched(writes)?; + alloc_meta!(self, set_batched, set_batched_as, writes)?; } else { // No split: absorb `next` entirely; just remove it from the list. let header_le = merged_size.to_le_bytes(); @@ -2141,7 +2246,7 @@ impl FirstFitBStackAllocator { (0, NONE) }, ]; - self.stack.set_batched(writes)?; + alloc_meta!(self, set_batched, set_batched_as, writes)?; } // SAFETY: slice resized by merging with adjacent free block return Ok(true); @@ -2174,9 +2279,17 @@ impl FirstFitBStackAllocator { let remainder_size = merged_size - aligned_new_len - Self::BLOCK_OVERHEAD_SIZE; let new_free_start = start + aligned_new_len + Self::BLOCK_OVERHEAD_SIZE; - let mut head_buf = [0u8; 8]; - self.stack.get_into(Self::FREE_HEAD_OFFSET, &mut head_buf)?; - let old_head = u64::from_le_bytes(head_buf); + let old_head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + }; // All offsets are relative to zero_buff[0] = start + block_size. let alloc_footer_off = (aligned_new_len - block_size) as usize; @@ -2217,8 +2330,13 @@ impl FirstFitBStackAllocator { )?; // Link forward: free_head → new free block // Failure cause: orphaned block - self.stack - .set(Self::FREE_HEAD_OFFSET, new_free_start.to_le_bytes())?; + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + new_free_start.to_le_bytes() + )?; // Link backward: old head's prev_free → new free block // Failure cause: orphaned block with stale forward link from old head (detectable in recovery) but no backward link if old_head != 0 { @@ -2751,7 +2869,7 @@ impl FirstFitBStackAllocator { let off_retftr = start + block_size; let mut committed = false; - let r = self.stack.inplace_gen(|feedback| { + let r = alloc_meta!(self, inplace_gen, inplace_gen_as, |feedback| { // A failed read or rejected write tears the batch down — never // `None`, which would commit the partial state. if let Err(e) = feedback { @@ -3192,7 +3310,7 @@ impl FirstFitBStackAllocator { let mut back_le = [0u8; 8]; let mut committed = false; - let r = self.stack.inplace_gen(|feedback| { + let r = alloc_meta!(self, inplace_gen, inplace_gen_as, |feedback| { if let Err(e) = feedback { return Some(BStackGenOp::Abort { source: Some(e) }); } @@ -3922,13 +4040,15 @@ mod fault_tests { // Corrupt `free_head` to an out-of-bounds offset. `b` has no free neighbour to // coalesce, so `add_to_free_list` reaches the head read and rejects it. - alloc - .stack() - .set( - FirstFitBStackAllocator::FREE_HEAD_OFFSET, - u64::MAX.to_le_bytes(), - ) - .unwrap(); + // The head is allocator metadata (ACL-protected), so poke it as the allocator. + crate::acl::alloc_meta!( + alloc, + set, + set_as, + FirstFitBStackAllocator::FREE_HEAD_OFFSET, + u64::MAX.to_le_bytes(), + ) + .unwrap(); let err = alloc .dealloc(b) diff --git a/src/alloc/ghost_tree.rs b/src/alloc/ghost_tree.rs index 8006947..114c887 100644 --- a/src/alloc/ghost_tree.rs +++ b/src/alloc/ghost_tree.rs @@ -6,6 +6,7 @@ use super::{ use crate::BStack; #[cfg(feature = "atomic")] use crate::BStackGenOp; +use crate::acl::alloc_meta; #[cfg(not(feature = "atomic"))] use std::cell::Cell; use std::fmt; @@ -189,6 +190,10 @@ struct PathEntry { /// ``` pub struct GhostTreeBstackAllocator { stack: BStack, + /// The real allocator capability, minted from `stack` at construction and + /// presented to the `_as` ops by the `meta_*` forwarders. + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth: crate::BStackAllocAuthority, #[cfg(feature = "atomic")] lock: Mutex<()>, #[cfg(not(feature = "atomic"))] @@ -224,13 +229,28 @@ impl GhostTreeBstackAllocator { /// Returns [`io::ErrorKind::InvalidData`] if the payload size falls in the /// unrecoverable range, or if the magic prefix does not match `ALGT`. pub fn new(stack: BStack) -> io::Result { + // Acquire the allocator authority up front and use it for the header I/O + // below, so a same-object reopen reads its own `Alloc`-marked header. + // Refused (not panic) if the stack's permit has already been taken. + #[cfg(feature = "expensive-slice-access-control")] + let alloc_auth = stack.take_alloc_authority().ok_or_else(|| { + io_error!( + PermissionDenied, + "GhostTreeBstackAllocator: alloc authority already taken from this stack" + ) + })?; let size = stack.len()?; if size == 0 { stack.extend(ARENA_START)?; stack.set(MAGIC_OFFSET, ALGT_MAGIC)?; // ROOT_OFFSET is zeroed by extend — null root pointer. + // Header (magic + root pointer) stays `Alloc` for the allocator's + // lifetime; own I/O via the `meta_*` helpers. + stack.acl_mark_alloc(MAGIC_OFFSET, ARENA_START - MAGIC_OFFSET)?; return Ok(Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, #[cfg(feature = "atomic")] lock: Mutex::new(()), @@ -251,6 +271,9 @@ impl GhostTreeBstackAllocator { // Verify magic prefix. let mut magic_buf = [0u8; 6]; + #[cfg(feature = "expensive-slice-access-control")] + stack.get_into_as(&alloc_auth, MAGIC_OFFSET, &mut magic_buf)?; + #[cfg(not(feature = "expensive-slice-access-control"))] stack.get_into(MAGIC_OFFSET, &mut magic_buf)?; if magic_buf != ALGT_MAGIC_PREFIX { return Err(io_error!( @@ -267,12 +290,18 @@ impl GhostTreeBstackAllocator { } let this = Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, #[cfg(feature = "atomic")] lock: Mutex::new(()), #[cfg(not(feature = "atomic"))] _not_sync: PhantomData, }; + // Re-arm the header mark on reopen (policy is not persisted); the root I/O + // in `coalesce_and_rebalance` below goes through the `meta_*` helpers. + this.stack + .acl_mark_alloc(MAGIC_OFFSET, ARENA_START - MAGIC_OFFSET)?; this.coalesce_and_rebalance()?; Ok(this) } @@ -282,9 +311,10 @@ impl GhostTreeBstackAllocator { /// Read the AVL root pointer from the header. #[inline] fn read_root(&self) -> io::Result { - let buf = &mut [0u8; 8]; - self.stack.get_into(ROOT_OFFSET, buf)?; - Ok(read_buf_le!(buf, 0 => u64)) + // Header stays `Alloc`; the root pointer is read as the allocator. + let mut buf = [0u8; 8]; + alloc_meta!(self, get_into, get_into_as, ROOT_OFFSET, &mut buf)?; + Ok(u64::from_le_bytes(buf)) } /// Write the AVL root pointer to the header. @@ -292,7 +322,7 @@ impl GhostTreeBstackAllocator { fn write_root(&self, ptr: u64) -> io::Result<()> { let mut buf = [0u8; 8]; write_buf!(ptr => buf, 0); - self.stack.set(ROOT_OFFSET, buf)?; + alloc_meta!(self, set, set_as, ROOT_OFFSET, buf)?; Ok(()) } @@ -1404,6 +1434,10 @@ impl BStackAllocator for GhostTreeBstackAllocator { #[inline] fn into_stack(self) -> BStack { + // Hand the allocator capability back so a caller that re-wraps the + // reclaimed stack can mint it again. + #[cfg(feature = "expensive-slice-access-control")] + self.stack.return_alloc_authority(self.alloc_auth); self.stack } @@ -1461,6 +1495,9 @@ impl BStackAllocator for GhostTreeBstackAllocator { let slice = ensure_own_handle(self, slice, "GhostTreeBstackAllocator::dealloc")?; let start = slice.start(); let len = slice.len(); + if let Err(source) = self.stack.acl_reclaim(start, len) { + return Err(BStackAllocError::with_handle(source, slice)); + } // Set once the (non-atomic) AVL insert has begun. A torn insert leaves // the tree inconsistent, and GhostTree has no is_free flag to repair it // in-process, so the block can no longer be returned for a retry — @@ -1681,6 +1718,11 @@ impl BStackBulkAllocator for GhostTreeBstackAllocator { ) -> Result<(), BStackBulkAllocError<'a, Self>> { let slices: Vec> = slices.into_iter().collect(); let slices = ensure_own_handles(self, slices, "GhostTreeBstackAllocator::dealloc_bulk")?; + for s in &slices { + if let Err(source) = self.stack.acl_reclaimable(s.start(), s.len()) { + return Err(BStackBulkAllocError::with_handles(source, slices)); + } + } // Set once any block has begun to be freed. This free is progressive // (tail discard, then per-block zero + AVL insert), so once it starts a diff --git a/src/alloc/guarded.rs b/src/alloc/guarded.rs index 1cba377..e5c66f4 100644 --- a/src/alloc/guarded.rs +++ b/src/alloc/guarded.rs @@ -30,7 +30,10 @@ //! [`as_slice`]: BStackGuardedSlice::as_slice //! [`raw_block`]: BStackGuardedSlice::raw_block -use super::{BStackAllocator, BStackOwnedSlice, BStackRange, BStackSlice}; +use super::{BStackAllocator, BStackRange, BStackSlice}; +// Used only by the `set`-gated `to_owned_in`/`to_owned_uninit_in`. +#[cfg(feature = "set")] +use super::BStackOwnedSlice; use std::{borrow::Cow, io, ops::Range}; /// A [`BStackSlice`] abstraction with lifecycle hooks for transparent I/O @@ -1115,7 +1118,6 @@ mod tests { #[cfg(feature = "set")] #[test] fn to_owned_in_copies_decoded_bytes() { - use crate::BStackAllocator; let (stack, _c) = mk_stack(); let key = 0x5A; let plain = b"secret payload!!"; diff --git a/src/alloc/linear.rs b/src/alloc/linear.rs index d67a340..03ebcbb 100644 --- a/src/alloc/linear.rs +++ b/src/alloc/linear.rs @@ -92,6 +92,12 @@ use std::{fmt, io}; /// ``` pub struct LinearBStackAllocator { stack: BStack, + /// The allocator capability, held so no outside caller can mint alloc + /// authority over the arena. `LinearBStackAllocator` marks no metadata, so it + /// never presents the token — it just holds it, and hands it back on + /// [`into_stack`](crate::BStackAllocator::into_stack). + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth: crate::BStackAllocAuthority, #[cfg(not(feature = "atomic"))] _not_sync: PhantomData>, } @@ -106,6 +112,10 @@ impl LinearBStackAllocator { #[must_use] pub fn new(stack: BStack) -> Self { Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth: stack + .take_alloc_authority() + .expect("fresh stack owns its alloc permit"), stack, #[cfg(not(feature = "atomic"))] _not_sync: PhantomData, @@ -145,6 +155,10 @@ impl BStackAllocator for LinearBStackAllocator { #[inline] fn into_stack(self) -> BStack { + // Hand the allocator capability back so a caller that re-wraps the + // reclaimed stack can mint it again. + #[cfg(feature = "expensive-slice-access-control")] + self.stack.return_alloc_authority(self.alloc_auth); self.stack } @@ -259,6 +273,9 @@ impl BStackAllocator for LinearBStackAllocator { let start = handle.start(); let end = handle.end(); let len = handle.len(); + if let Err(source) = self.stack.acl_reclaim(start, len) { + return Err(BStackAllocError::with_handle(source, handle)); + } (|| -> io::Result<()> { let current_tail = self.stack.len()?; if end == current_tail { @@ -283,6 +300,9 @@ impl BStackAllocator for LinearBStackAllocator { let start = handle.start(); let end = handle.end(); let len = handle.len(); + if let Err(source) = self.stack.acl_reclaim(start, len) { + return Err(BStackAllocError::with_handle(source, handle)); + } // try_discard is a no-op when the tail has moved, matching non-tail dealloc semantics. self.stack .try_discard(end, len) @@ -422,6 +442,11 @@ impl BStackBulkAllocator for LinearBStackAllocator { return Ok(()); } let mut sorted = ensure_own_handles(self, owned, "LinearBStackAllocator::dealloc_bulk")?; + for h in &sorted { + if let Err(source) = self.stack.acl_reclaimable(h.start(), h.len()) { + return Err(BStackBulkAllocError::with_handles(source, sorted)); + } + } sorted.sort_by_key(|s| std::cmp::Reverse(s.end())); let result = (|| -> io::Result<()> { let current_tail = self.stack.len()?; @@ -459,6 +484,11 @@ impl BStackBulkAllocator for LinearBStackAllocator { return Ok(()); } let mut sorted = ensure_own_handles(self, owned, "LinearBStackAllocator::dealloc_bulk")?; + for h in &sorted { + if let Err(source) = self.stack.acl_reclaimable(h.start(), h.len()) { + return Err(BStackBulkAllocError::with_handles(source, sorted)); + } + } sorted.sort_by_key(|s| std::cmp::Reverse(s.end())); let result = (|| -> io::Result<()> { let current_tail = self.stack.len()?; diff --git a/src/alloc/segregated.rs b/src/alloc/segregated.rs index 088e5f7..1f77504 100644 --- a/src/alloc/segregated.rs +++ b/src/alloc/segregated.rs @@ -49,6 +49,7 @@ use super::{BStackBulkAllocError, BStackBulkAllocator, ensure_own_handles}; use crate::BStack; #[cfg(feature = "atomic")] use crate::BStackGenOp; +use crate::acl::alloc_meta; #[cfg(feature = "atomic")] use crate::{bstack_unsafe_reborrow, bstack_unsafe_reborrow_mut}; #[cfg(not(feature = "atomic"))] @@ -119,6 +120,10 @@ const ALSG_MAGIC_PREFIX: [u8; 6] = *b"ALSG\x00\x02"; #[cfg(feature = "set")] pub struct SegregatedBStackAllocator { stack: BStack, + /// The real allocator capability, minted from `stack` at construction and + /// presented to the `_as` ops by the `alloc_meta!` forwarders. + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth: crate::BStackAllocAuthority, #[cfg(not(feature = "atomic"))] _not_sync: PhantomData>, } @@ -278,6 +283,16 @@ impl SegregatedBStackAllocator { /// allocator of the expected version). /// * Any [`io::Error`] from the underlying [`BStack`] operations. pub fn new(stack: BStack) -> io::Result { + // Acquire the allocator authority up front and use it for the header I/O + // below, so a same-object reopen reads its own `Alloc`-marked header. + // Refused (not panic) if the stack's permit has already been taken. + #[cfg(feature = "expensive-slice-access-control")] + let alloc_auth = stack.take_alloc_authority().ok_or_else(|| { + io_error!( + PermissionDenied, + "SegregatedBStackAllocator: alloc authority already taken from this stack" + ) + })?; if stack.is_empty()? { // Initialize a new stack: write the header and return a fresh allocator. const OFFSET_OFFSET: usize = SegregatedBStackAllocator::OFFSET_SIZE as usize; @@ -285,7 +300,12 @@ impl SegregatedBStackAllocator { hdr[OFFSET_OFFSET..].copy_from_slice(&ALSG_MAGIC); // the reserved words and every free_head remain 0. let _ = stack.extend_sparse(hdr, Self::ARENA_START)?; + // Header (magic + the 33 free-list heads) stays `Alloc` for the + // allocator's lifetime; own I/O via the `meta_*` helpers. + stack.acl_mark_alloc(Self::OFFSET_SIZE, Self::ARENA_START - Self::OFFSET_SIZE)?; return Ok(Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, #[cfg(not(feature = "atomic"))] _not_sync: PhantomData, @@ -302,6 +322,9 @@ impl SegregatedBStackAllocator { } let mut magic = [0u8; 8]; + #[cfg(feature = "expensive-slice-access-control")] + stack.get_into_as(&alloc_auth, Self::OFFSET_SIZE, &mut magic)?; + #[cfg(not(feature = "expensive-slice-access-control"))] stack.get_into(Self::OFFSET_SIZE, &mut magic)?; if magic[..ALSG_MAGIC_PREFIX.len()] != ALSG_MAGIC_PREFIX { return Err(io_error!( @@ -321,10 +344,17 @@ impl SegregatedBStackAllocator { )); } let allocator = Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, #[cfg(not(feature = "atomic"))] _not_sync: PhantomData, }; + // Re-arm the header mark on reopen (policy is not persisted); the head I/O + // in `recover` below goes through the `meta_*` helpers. + allocator + .stack + .acl_mark_alloc(Self::OFFSET_SIZE, Self::ARENA_START - Self::OFFSET_SIZE)?; // SAFETY: `allocator` was just constructed and has not yet escaped this // function, so no other thread can hold it — it is trivially quiescent. unsafe { allocator.recover()? }; @@ -419,7 +449,7 @@ impl SegregatedBStackAllocator { for (c, &h) in heads.iter().enumerate() { write_buf!(h => head_bytes, c * 8); } - self.stack.set(Self::FREE_HEAD_BASE, head_bytes)?; + alloc_meta!(self, set, set_as, Self::FREE_HEAD_BASE, head_bytes)?; Ok(unsure) } @@ -519,7 +549,7 @@ impl SegregatedBStackAllocator { let mut wi = 0usize; let mut state = Scan::NeedLen; - self.stack.inplace_gen(|feedback| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |feedback| { if let Err(e) = feedback { return Some(BStackGenOp::Abort { source: Some(e) }); } @@ -664,12 +694,16 @@ impl SegregatedBStackAllocator { #[cfg(not(feature = "atomic"))] fn pop_class(&self, class: u64) -> io::Result> { let head_off = Self::head_off(class); - let head = u64::from_le_bytes(read_bstack!(self.stack, head_off => u64)); + let head = { + let mut buf = [0u8; 8]; + alloc_meta!(self, get_into, get_into_as, head_off, &mut buf)?; + u64::from_le_bytes(buf) + }; if head == Self::SENTINEL { return Ok(None); } let next = u64::from_le_bytes(read_bstack!(self.stack, head + Self::OVERHEAD => u64)); - self.stack.set(head_off, next.to_le_bytes())?; + alloc_meta!(self, set, set_as, head_off, next.to_le_bytes())?; Ok(Some(head)) } @@ -686,7 +720,7 @@ impl SegregatedBStackAllocator { let mut next_buf = [0u8; 8]; let mut step = 0u32; let mut popped: Option = None; - self.stack.process_gen(|| { + alloc_meta!(self, process_gen, process_gen_as, || { let op = match step { 0 => Some(BStackGenOp::Read { offset: head_off, @@ -730,7 +764,11 @@ impl SegregatedBStackAllocator { #[cfg(not(feature = "atomic"))] fn pop_oversized(&self, need: u64) -> io::Result> { let head_off = Self::head_off(Self::OVERSIZED_CLASS); - let head = u64::from_le_bytes(read_bstack!(self.stack, head_off => u64)); + let head = { + let mut buf = [0u8; 8]; + alloc_meta!(self, get_into, get_into_as, head_off, &mut buf)?; + u64::from_le_bytes(buf) + }; if head == Self::SENTINEL { return Ok(None); } @@ -743,7 +781,7 @@ impl SegregatedBStackAllocator { return Ok(None); } let next = read_buf_le!(buf, 8 => u64); - self.stack.set(head_off, next.to_le_bytes())?; + alloc_meta!(self, set, set_as, head_off, next.to_le_bytes())?; Ok(Some((head, size))) } @@ -766,7 +804,7 @@ impl SegregatedBStackAllocator { let mut head = 0u64; let mut size = 0u64; let mut popped: Option<(u64, u64)> = None; - self.stack.process_gen(|| { + alloc_meta!(self, process_gen, process_gen_as, || { let op = match step { 0 => Some(BStackGenOp::Read { offset: head_off, @@ -926,18 +964,22 @@ impl SegregatedBStackAllocator { #[cfg(not(feature = "atomic"))] { // Non-atomic path: read head, write overhead+next_free, write head. - let head = u64::from_le_bytes(read_bstack!(self.stack, head_off => u64)); + let head = { + let mut buf = [0u8; 8]; + alloc_meta!(self, get_into, get_into_as, head_off, &mut buf)?; + u64::from_le_bytes(buf) + }; write_buf!(head => overhead_buf, 8); self.stack.set(block_start, overhead_buf)?; // A crash between these two writes leaves the block free-tagged so it is // recoverable by `recover`. - self.stack.set(head_off, start_bytes) + alloc_meta!(self, set, set_as, head_off, start_bytes) } #[cfg(feature = "atomic")] { let mut step = 0u32; let mut read_err: Option = None; - self.stack.inplace_gen(|res| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |res| { if let Err(e) = res { read_err = Some(e); return None; @@ -1040,11 +1082,16 @@ impl SegregatedBStackAllocator { // Copy the prefilled overhead into a local buffer, then read the // old head directly into the latter half before writing both. let mut shared = overhead_next[i]; - // next_free ← current head of this class (read straight in). - self.stack.get_into(head_offs[i], &mut shared[8..])?; + // next_free ← current head of this class (read as the allocator). + let head = { + let mut buf = [0u8; 8]; + alloc_meta!(self, get_into, get_into_as, head_offs[i], &mut buf)?; + u64::from_le_bytes(buf) + }; + shared[8..16].copy_from_slice(&head.to_le_bytes()); // overhead || next_free, then head ← this block. self.stack.set(block_offs[i], shared)?; - self.stack.set(head_offs[i], blockoff_bytes[i])?; + alloc_meta!(self, set, set_as, head_offs[i], blockoff_bytes[i])?; } // Commit: expose the pieces (and, for a claim, mark the block in use). self.stack.set(prefix_off, prefix)?; @@ -1062,7 +1109,7 @@ impl SegregatedBStackAllocator { // `overhead_next` already holds the per-piece overhead in its first // 8 bytes; we will read each head directly into its second half. let mut read_err: Option = None; - self.stack.inplace_gen(|res| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |res| { if let Err(e) = res { read_err = Some(e); return None; @@ -1146,7 +1193,7 @@ impl SegregatedBStackAllocator { write_buf!(old_size >> 4 => old_buf, 0); // free tag: high bit clear let mut step = 0u32; let mut read_err: Option = None; - self.stack.inplace_gen(|res| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |res| { if let Err(e) = res { read_err = Some(e); return None; @@ -1210,6 +1257,10 @@ impl BStackAllocator for SegregatedBStackAllocator { #[inline] fn into_stack(self) -> BStack { + // Hand the allocator capability back so a caller that re-wraps the + // reclaimed stack can mint it again. + #[cfg(feature = "expensive-slice-access-control")] + self.stack.return_alloc_authority(self.alloc_auth); self.stack } @@ -1263,6 +1314,9 @@ impl BStackAllocator for SegregatedBStackAllocator { let slice = ensure_own_handle(self, slice, "SegregatedBStackAllocator::dealloc")?; let start = slice.start(); let len = slice.len(); + if let Err(source) = self.stack.acl_reclaim(start, len) { + return Err(BStackAllocError::with_handle(source, slice)); + } if slice.is_empty() { return Ok(()); } @@ -1367,7 +1421,7 @@ impl SegregatedBStackAllocator { let mut st = St::Chase(0); let mut in_class = 0usize; // blocks popped so far from the current class - self.stack.inplace_gen(|res| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |res| { // A prior read failed. All writes live in the write phase, entered // only once every read has completed, so nothing is staged yet: // ending with `None` commits the empty batch (pops nothing). @@ -1516,7 +1570,10 @@ impl SegregatedBStackAllocator { // not merged with the staging above; re-`chunk_by` needs no `splices` array. for run in blocks.chunk_by(same_class) { let last = run[run.len() - 1].0; - let _ = self.stack.cross_exchange( + let _ = alloc_meta!( + self, + cross_exchange, + cross_exchange_as, last + Self::OVERHEAD, Self::head_off(Self::classify(run[0].1)), 8, @@ -1576,7 +1633,7 @@ impl SegregatedBStackAllocator { } let mut st = St::ReadLen; - self.stack.inplace_gen(|res| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |res| { // A prior read failed. Read-phase reads never stage a write and the // write phase issues no reads, so nothing is staged when a read error // can arrive: ending with `None` commits nothing. @@ -2030,6 +2087,11 @@ impl BStackBulkAllocator for SegregatedBStackAllocator { ) -> Result<(), BStackBulkAllocError<'a, Self>> { let slices: Vec> = handles.into_iter().collect(); let slices = ensure_own_handles(self, slices, "SegregatedBStackAllocator::dealloc_bulk")?; + for s in &slices { + if let Err(source) = self.stack.acl_reclaimable(s.start(), s.len()) { + return Err(BStackBulkAllocError::with_handles(source, slices)); + } + } let mut freeing = false; let result = (|| -> io::Result<()> { @@ -2124,10 +2186,14 @@ impl BStackBulkAllocator for SegregatedBStackAllocator { for run in blocks.chunk_by(same_class) { let last = run[run.len() - 1].0; let head_off = Self::head_off(Self::classify(run[0].1)); - if let Err(e) = self - .stack - .cross_exchange(last + Self::OVERHEAD, head_off, 8) - { + if let Err(e) = alloc_meta!( + self, + cross_exchange, + cross_exchange_as, + last + Self::OVERHEAD, + head_off, + 8 + ) { first_err.get_or_insert(e); } } @@ -2385,7 +2451,7 @@ impl SegregatedBStackAllocator { let cur_ptr: *mut u64 = &mut cur_len; let mut phase = 0u8; let mut truncated = false; - self.stack.process_gen(|| match phase { + alloc_meta!(self, process_gen, process_gen_as, || match phase { 0 => { phase = 1; // SAFETY: `process_gen` invokes this closure strictly @@ -3352,7 +3418,9 @@ mod tests { let _b = a.alloc(100).unwrap(); // pins a second class-6 block a.dealloc(a1).unwrap(); // head[6] → a1 // Simulate a leak: clear head[6] so a1 is free-tagged but unreachable. - a.stack().set(Seg::head_off(6), 0u64.to_le_bytes()).unwrap(); + // A free-list head is allocator metadata (protected under the ACL + // feature), so this simulated write goes through the authorized path. + crate::acl::alloc_meta!(a, set, set_as, Seg::head_off(6), 0u64.to_le_bytes()).unwrap(); assert_eq!( unsafe { a.recover() }.unwrap(), 0, @@ -3586,13 +3654,11 @@ mod bulk_tests { a.stack.set(xb + 272, (48u64 >> 4).to_le_bytes()).unwrap(); assert_eq!(unsafe { a.recover() }.unwrap(), 0); // largest_class_le(272) == 256, so it lands on class 15. - let head = u64::from_le_bytes( - a.stack - .get(Seg::head_off(15), Seg::head_off(15) + 8) - .unwrap() - .try_into() - .unwrap(), - ); + let head = { + let mut buf = [0u8; 8]; + crate::acl::alloc_meta!(a, get_into, get_into_as, Seg::head_off(15), &mut buf).unwrap(); + u64::from_le_bytes(buf) + }; assert_eq!(head, xb); // A class-15 request (block 256) reuses it instead of failing the batch. let r = a.alloc_bulk([248u64]).unwrap(); diff --git a/src/alloc/slab.rs b/src/alloc/slab.rs index 0a8d41a..e564bb9 100644 --- a/src/alloc/slab.rs +++ b/src/alloc/slab.rs @@ -12,6 +12,7 @@ use super::{BStackBulkAllocError, BStackBulkAllocator, ensure_own_handles}; use crate::BStack; #[cfg(feature = "atomic")] use crate::BStackGenOp; +use crate::acl::alloc_meta; #[cfg(feature = "atomic")] use crate::{bstack_unsafe_reborrow, bstack_unsafe_reborrow_mut}; #[cfg(not(feature = "atomic"))] @@ -146,6 +147,10 @@ const ALSL_MAGIC_PREFIX: [u8; 6] = *b"ALSL\x00\x01"; #[cfg(feature = "set")] pub struct SlabBStackAllocator { stack: BStack, + /// The real allocator capability, minted from `stack` at construction and + /// presented to the `_as` ops by the `alloc_meta!` forwarders. + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth: crate::BStackAllocAuthority, /// Cached from the on-disk header; fixed for the lifetime of the allocator. block_size: u64, #[cfg(not(feature = "atomic"))] @@ -183,6 +188,15 @@ impl SlabBStackAllocator { /// empty (use [`SlabBStackAllocator::open`] to reopen an existing file). /// * Any [`io::Error`] propagated from the underlying [`BStack`] operations. pub fn new(stack: BStack, block_size: u64) -> io::Result { + // Acquire the allocator authority up front; refused (not panic) if the + // stack's permit has already been taken. + #[cfg(feature = "expensive-slice-access-control")] + let alloc_auth = stack.take_alloc_authority().ok_or_else(|| { + io_error!( + PermissionDenied, + "SlabBStackAllocator: alloc authority already taken from this stack" + ) + })?; if !stack.is_empty()? { return Err(io_error!( InvalidInput, @@ -208,7 +222,11 @@ impl SlabBStackAllocator { write_buf!(block_size => hdr, off + 8); // free_head at off+16 remains 0 (SENTINEL) stack.push(hdr)?; + // Header stays `Alloc` for the allocator's lifetime; own I/O via `meta_*`. + stack.acl_mark_alloc(Self::OFFSET_SIZE, Self::HEADER_SIZE)?; Ok(Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, block_size, #[cfg(not(feature = "atomic"))] @@ -229,6 +247,15 @@ impl SlabBStackAllocator { /// `block_size`, or invalid `free_head`. /// * Any [`io::Error`] propagated from the underlying [`BStack`] operations. pub fn open(stack: BStack) -> io::Result { + // Acquire the allocator authority up front and read the header through it, + // so a same-object reopen sees its own `Alloc`-marked header. + #[cfg(feature = "expensive-slice-access-control")] + let alloc_auth = stack.take_alloc_authority().ok_or_else(|| { + io_error!( + PermissionDenied, + "SlabBStackAllocator: alloc authority already taken from this stack" + ) + })?; if stack.is_empty()? { return Err(io_error!( InvalidInput, @@ -245,6 +272,9 @@ impl SlabBStackAllocator { } let mut header = [0u8; Self::HEADER_SIZE as usize]; + #[cfg(feature = "expensive-slice-access-control")] + stack.get_into_as(&alloc_auth, Self::OFFSET_SIZE, &mut header)?; + #[cfg(not(feature = "expensive-slice-access-control"))] stack.get_into(Self::OFFSET_SIZE, &mut header)?; if header[..ALSL_MAGIC_PREFIX.len()] != ALSL_MAGIC_PREFIX { @@ -288,7 +318,11 @@ impl SlabBStackAllocator { )); } + // Re-arm the header mark on reopen (policy is not persisted). + stack.acl_mark_alloc(Self::OFFSET_SIZE, Self::HEADER_SIZE)?; Ok(Self { + #[cfg(feature = "expensive-slice-access-control")] + alloc_auth, stack, block_size: stored_block_size, #[cfg(not(feature = "atomic"))] @@ -325,7 +359,7 @@ impl SlabBStackAllocator { let mut step = 0usize; let mut popped: Option = None; - self.stack.process_gen(|| { + alloc_meta!(self, process_gen, process_gen_as, || { let op = match step { // Step 0: read the current free-list head. 0 => Some(BStackGenOp::Read { @@ -376,11 +410,24 @@ impl SlabBStackAllocator { /// See the `atomic` variant for the meaning of `init`. #[cfg(not(feature = "atomic"))] fn pop_free_block(&self, init: bool) -> io::Result> { - let head = u64::from_le_bytes(read_bstack!(self.stack, Self::FREE_HEAD_OFFSET => u64)); + let head = { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + }; if head == Self::SENTINEL { return Ok(None); } - self.stack.set( + alloc_meta!( + self, + set, + set_as, Self::FREE_HEAD_OFFSET, read_bstack!(self.stack, head => u64), )?; @@ -404,8 +451,14 @@ impl SlabBStackAllocator { #[cfg(feature = "atomic")] fn push_free_block(&self, block_start: u64) -> io::Result<()> { self.stack.set(block_start, block_start.to_le_bytes())?; - self.stack - .cross_exchange(block_start, Self::FREE_HEAD_OFFSET, 8) + alloc_meta!( + self, + cross_exchange, + cross_exchange_as, + block_start, + Self::FREE_HEAD_OFFSET, + 8 + ) } /// Prepend the block at `block_start` to the free list. @@ -416,10 +469,26 @@ impl SlabBStackAllocator { // rather than corrupting the list. self.stack.set( block_start, - read_bstack!(self.stack, Self::FREE_HEAD_OFFSET => u64), + { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + } + .to_le_bytes(), )?; - self.stack - .set(Self::FREE_HEAD_OFFSET, block_start.to_le_bytes()) + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + block_start.to_le_bytes() + ) } /// Prepend `count` contiguous blocks starting at `first_block` to the free list. @@ -475,7 +544,17 @@ impl SlabBStackAllocator { } #[cfg(not(feature = "atomic"))] { - u64::from_le_bytes(read_bstack!(self.stack, Self::FREE_HEAD_OFFSET => u64)) + { + let mut buf = [0u8; 8]; + alloc_meta!( + self, + get_into, + get_into_as, + Self::FREE_HEAD_OFFSET, + &mut buf + )?; + u64::from_le_bytes(buf) + } } }; let off = usize::try_from( @@ -494,13 +573,24 @@ impl SlabBStackAllocator { io_error!(InvalidInput, "last free-list offset overflows u64") })?) .ok_or_else(|| io_error!(InvalidInput, "last block offset overflows u64"))?; - self.stack - .cross_exchange(last_block, Self::FREE_HEAD_OFFSET, 8) + alloc_meta!( + self, + cross_exchange, + cross_exchange_as, + last_block, + Self::FREE_HEAD_OFFSET, + 8 + ) } #[cfg(not(feature = "atomic"))] { - self.stack - .set(Self::FREE_HEAD_OFFSET, first_block.to_le_bytes()) + alloc_meta!( + self, + set, + set_as, + Self::FREE_HEAD_OFFSET, + first_block.to_le_bytes() + ) } } @@ -749,6 +839,10 @@ impl BStackAllocator for SlabBStackAllocator { #[inline] fn into_stack(self) -> BStack { + // Hand the allocator capability back so a caller that re-wraps the + // reclaimed stack can mint it again. + #[cfg(feature = "expensive-slice-access-control")] + self.stack.return_alloc_authority(self.alloc_auth); self.stack } @@ -785,6 +879,9 @@ impl BStackAllocator for SlabBStackAllocator { let slice = ensure_own_handle(self, slice, "SlabBStackAllocator::dealloc")?; let start = slice.start(); let len = slice.len(); + if let Err(source) = self.stack.acl_reclaim(start, len) { + return Err(BStackAllocError::with_handle(source, slice)); + } // Set once the caller's blocks may have been partially freed, after // which returning the handle for retry would risk a double-free. let mut lost = false; @@ -932,7 +1029,7 @@ impl SlabBStackAllocator { } let mut st = St::ReadHead; - self.stack.process_gen(|| { + alloc_meta!(self, process_gen, process_gen_as, || { loop { match st { // Read the current free-list head. @@ -1052,7 +1149,7 @@ impl SlabBStackAllocator { } let mut st = St::ReadHead; - self.stack.inplace_gen(|prev| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |prev| { // A failed read reports its error here (not as a return value); the // buffer was not filled, so bail out before consuming it. During the // write phase `prev` is always the Ok validation of an in-range write. @@ -1186,7 +1283,7 @@ impl SlabBStackAllocator { } let zero_block = vec![0u8; self.block_size as usize]; let mut i = 0usize; - self.stack.inplace_gen(|_prev| { + alloc_meta!(self, inplace_gen, inplace_gen_as, |_prev| { // `i` walks `blocks` (collected before this call), so `get` is in // bounds until it runs off the end, where `?` ends the sequence. let off = *blocks.get(i)?; @@ -1222,8 +1319,14 @@ impl SlabBStackAllocator { batch.push((blocks[i], next.to_le_bytes())); } self.stack.set_batched(batch)?; - self.stack - .cross_exchange(blocks[k - 1], Self::FREE_HEAD_OFFSET, 8) + alloc_meta!( + self, + cross_exchange, + cross_exchange_as, + blocks[k - 1], + Self::FREE_HEAD_OFFSET, + 8 + ) } /// Best-effort cleanup after an `alloc_bulk` extend path fails partway. @@ -1426,6 +1529,11 @@ impl BStackBulkAllocator for SlabBStackAllocator { ) -> Result<(), BStackBulkAllocError<'a, Self>> { let slices: Vec> = handles.into_iter().collect(); let slices = ensure_own_handles(self, slices, "SlabBStackAllocator::dealloc_bulk")?; + for s in &slices { + if let Err(source) = self.stack.acl_reclaimable(s.start(), s.len()) { + return Err(BStackBulkAllocError::with_handles(source, slices)); + } + } let bs = self.block_size; // Set once the chain build has begun. Before this point every handle is diff --git a/src/alloc/slice.rs b/src/alloc/slice.rs index b4eb678..b7173e5 100644 --- a/src/alloc/slice.rs +++ b/src/alloc/slice.rs @@ -2,12 +2,30 @@ use super::BStackAllocator; #[cfg(all(feature = "set", feature = "atomic"))] use super::BStackOwnedSliceAllocator; use crate::BStack; +#[cfg(feature = "expensive-slice-access-control")] +use crate::{BStackAccess, BStackAccessAuthorities, BStackAuthority}; use std::borrow::Borrow; use std::fmt; use std::hash::{Hash, Hasher}; use std::io; use std::ops::{Deref, Range}; +/// Dispatch one `BStackSlice` I/O op through the slice's stored authority: +/// `slice_meta!(self, op, op_as, args...)` presents `self.auth` to the +/// token-carrying `_as` entry point with the `expensive-slice-access-control` +/// feature, or calls the plain tokenless op without it. Expands inline at the +/// call site — no per-op wrapper methods; `self` is passed explicitly (`self` +/// does not cross `macro_rules` hygiene). All offsets are absolute. +macro_rules! slice_meta { + ($self:expr, $op:ident, $op_as:ident $(, $arg:expr)* $(,)?) => {{ + #[cfg(feature = "expensive-slice-access-control")] + let __r = $self.stack.$op_as($self.auth $(, $arg)*); + #[cfg(not(feature = "expensive-slice-access-control"))] + let __r = $self.stack.$op($($arg),*); + __r + }}; +} + /// A raw `(offset, len)` coordinate pair with no backing reference. /// /// `BStackRange` is the serialization and persistence representation: store it on @@ -288,6 +306,11 @@ impl fmt::Display for BStackRange { pub struct BStackSlice<'a> { stack: &'a BStack, range: BStackRange, + /// Authority this slice's I/O presents to the stack's access-control table. + /// `NONE` unless granted via [`authorize`](BStackSlice::authorize); inherited + /// by derived slices. + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities, } impl<'a> Clone for BStackSlice<'a> { @@ -296,6 +319,8 @@ impl<'a> Clone for BStackSlice<'a> { BStackSlice { stack: self.stack, range: self.range, + #[cfg(feature = "expensive-slice-access-control")] + auth: self.auth, } } } @@ -332,6 +357,8 @@ impl<'a> BStackSlice<'a> { Self { stack, range: unsafe { BStackRange::from_raw_parts(offset, len) }, + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, } } @@ -345,7 +372,12 @@ impl<'a> BStackSlice<'a> { #[inline] #[must_use] pub unsafe fn from_raw_range(stack: &'a BStack, range: BStackRange) -> Self { - Self { stack, range } + Self { + stack, + range, + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, + } } /// Construct a zero-length slice anchored at offset 0. @@ -357,6 +389,8 @@ impl<'a> BStackSlice<'a> { Self { stack, range: BStackRange::empty(), + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, } } @@ -427,17 +461,27 @@ impl<'a> BStackSlice<'a> { /// Merge this slice with `other` into a single slice covering both. /// /// Delegates to [`BStackRange::merge`] on the underlying coordinates. - /// Returns `None` if the ranges are non-empty and disjoint, or if `self` - /// and `other` are backed by different [`BStack`]s. + /// Returns `None` if the ranges are non-empty and disjoint, if `self` and + /// `other` are backed by different [`BStack`]s, or (with the + /// `expensive-slice-access-control` feature) if they carry different + /// authorities. #[inline] #[must_use] pub fn merge(&self, other: &Self) -> Option { if !std::ptr::eq(self.stack, other.stack) { return None; } + // Merging slices carrying different authorities would silently widen or + // narrow the access one of them was granted; refuse it. + #[cfg(feature = "expensive-slice-access-control")] + if self.auth != other.auth { + return None; + } self.range.merge(&other.range).map(|range| Self { stack: self.stack, range, + #[cfg(feature = "expensive-slice-access-control")] + auth: self.auth, }) } @@ -446,16 +490,24 @@ impl<'a> BStackSlice<'a> { /// /// Delegates to [`BStackRange::merge_adjacent`] on the underlying /// coordinates. Returns `None` if the slices are not adjacent, either is - /// empty, or `self` and `other` are backed by different [`BStack`]s. + /// empty, `self` and `other` are backed by different [`BStack`]s, or (with + /// the `expensive-slice-access-control` feature) they carry different + /// authorities. #[inline] #[must_use] pub fn merge_adjacent(&self, other: &Self) -> Option { if !std::ptr::eq(self.stack, other.stack) { return None; } + #[cfg(feature = "expensive-slice-access-control")] + if self.auth != other.auth { + return None; + } self.range.merge_adjacent(&other.range).map(|range| Self { stack: self.stack, range, + #[cfg(feature = "expensive-slice-access-control")] + auth: self.auth, }) } @@ -481,6 +533,8 @@ impl<'a> BStackSlice<'a> { Self { stack, range: BStackRange::from_bytes(bytes), + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, } } @@ -523,6 +577,8 @@ impl<'a> BStackSlice<'a> { range: unsafe { BStackRange::from_raw_parts(self.start() + range.start, range.end - range.start) }, + #[cfg(feature = "expensive-slice-access-control")] + auth: self.auth, } } @@ -590,7 +646,7 @@ impl<'a> BStackSlice<'a> { return Ok(None); } let mut buf = [0u8; 1]; - self.stack.get_into(self.start() + index, &mut buf)?; + slice_meta!(self, get_into, get_into_as, self.start() + index, &mut buf)?; Ok(Some(buf[0])) } @@ -663,14 +719,14 @@ impl<'a> BStackSlice<'a> { /// Read the entire slice into a new `Vec`. #[inline] pub fn read(&self) -> io::Result> { - self.stack.get(self.start(), self.end()) + slice_meta!(self, get, get_as, self.start(), self.end()) } /// Read bytes into `buf`, up to `min(buf.len(), self.len())` bytes. #[inline] pub fn read_into(&self, buf: &mut [u8]) -> io::Result<()> { let n = (buf.len() as u64).min(self.len()) as usize; - self.stack.get_into(self.start(), &mut buf[..n]) + slice_meta!(self, get_into, get_into_as, self.start(), &mut buf[..n]) } /// Read `[start, end)` relative to this slice into a new `Vec`. @@ -682,7 +738,7 @@ impl<'a> BStackSlice<'a> { self.len() )); } - self.stack.get(self.start() + start, self.start() + end) + slice_meta!(self, get, get_as, self.start() + start, self.start() + end) } /// Read `[start, start + buf.len())` relative to this slice into `buf`. @@ -695,7 +751,7 @@ impl<'a> BStackSlice<'a> { self.len() )); } - self.stack.get_into(self.start() + start, buf) + slice_meta!(self, get_into, get_into_as, self.start() + start, buf) } /// Overwrite the beginning of this slice with `data` (up to `self.len()` bytes). @@ -705,7 +761,7 @@ impl<'a> BStackSlice<'a> { pub fn write(&mut self, data: impl AsRef<[u8]>) -> io::Result<()> { let data = data.as_ref(); let n = (data.len() as u64).min(self.len()) as usize; - self.stack.set(self.start(), &data[..n]) + slice_meta!(self, set, set_as, self.start(), &data[..n]) } /// Overwrite `[start, start + data.len())` relative to this slice. @@ -722,7 +778,19 @@ impl<'a> BStackSlice<'a> { self.len() )); } - self.stack.set(self.start() + start, data) + slice_meta!(self, set, set_as, self.start() + start, data) + } + + /// Grant this slice the authority carried by `auth`, so its subsequent I/O + /// (and any view derived from it) may reach a [`Prot`](BStackAccess::Prot)/ + /// [`Alloc`](BStackAccess::Alloc) region it was authorized for. A token minted + /// from a different stack grants nothing. + /// + /// Requires the `expensive-slice-access-control` feature. + #[cfg(feature = "expensive-slice-access-control")] + #[inline] + pub fn authorize(&mut self, auth: impl BStackAuthority) { + self.auth = auth.authorities_for(self.stack); } /// Zero out the entire slice. @@ -731,7 +799,7 @@ impl<'a> BStackSlice<'a> { #[cfg(feature = "set")] #[inline] pub fn zero(&mut self) -> io::Result<()> { - self.stack.zero(self.start(), self.len()) + slice_meta!(self, zero, zero_as, self.start(), self.len()) } /// Zero `[start, start + n)` within this slice. @@ -747,7 +815,7 @@ impl<'a> BStackSlice<'a> { self.len() )); } - self.stack.zero(self.start() + start, n) + slice_meta!(self, zero, zero_as, self.start() + start, n) } /// Fill the entire slice with `value`. @@ -758,7 +826,7 @@ impl<'a> BStackSlice<'a> { #[cfg(feature = "set")] #[inline] pub fn fill(&mut self, value: u8) -> io::Result<()> { - self.stack.repeat(self.start(), [value], self.len()) + slice_meta!(self, repeat, repeat_as, self.start(), [value], self.len()) } /// Fill the slice by calling `f` once per byte. @@ -792,7 +860,7 @@ impl<'a> BStackSlice<'a> { self.len(), "copy_from_slice: length mismatch" ); - self.stack.set(self.start(), src) + slice_meta!(self, set, set_as, self.start(), src) } /// Copy the contents of `src` into this slice. @@ -827,7 +895,7 @@ impl<'a> BStackSlice<'a> { if self.is_empty() { return Ok(()); } - self.stack.copy(src.start(), self.start(), self.len()) + slice_meta!(self, copy, copy_as, src.start(), self.start(), self.len()) } /// Copy this view's contents into a fresh allocation from `allocator`. @@ -927,8 +995,14 @@ impl<'a> BStackSlice<'a> { if n == 0 { return Ok(()); } - self.stack - .copy(self.start() + src_range.start, self.start() + dest, n) + slice_meta!( + self, + copy, + copy_as, + self.start() + src_range.start, + self.start() + dest, + n + ) } /// Overwrite this slice with `new_bytes` if `guard`'s current contents @@ -982,8 +1056,15 @@ impl<'a> BStackSlice<'a> { self.len() )); } - self.stack - .eq_crds(guard.start(), expected, self.start(), new_bytes) + slice_meta!( + self, + eq_crds, + eq_crds_as, + guard.start(), + expected, + self.start(), + new_bytes + ) } /// Overwrite this slice with `new_bytes` if `guard`'s current contents do @@ -1028,8 +1109,15 @@ impl<'a> BStackSlice<'a> { self.len() )); } - self.stack - .ne_crds(guard.start(), expected, self.start(), new_bytes) + slice_meta!( + self, + ne_crds, + ne_crds_as, + guard.start(), + expected, + self.start(), + new_bytes + ) } /// Overwrite this slice with `new_bytes` if `guard`'s current contents @@ -1080,8 +1168,16 @@ impl<'a> BStackSlice<'a> { self.len() )); } - self.stack - .masked_eq_crds(guard.start(), mask, expected, self.start(), new_bytes) + slice_meta!( + self, + masked_eq_crds, + masked_eq_crds_as, + guard.start(), + mask, + expected, + self.start(), + new_bytes + ) } /// Swap the contents of this slice with `other`. @@ -1111,8 +1207,14 @@ impl<'a> BStackSlice<'a> { if self.is_empty() || self.start() == other.start() { return Ok(()); } - self.stack - .cross_exchange(self.start(), other.start(), self.len()) + slice_meta!( + self, + cross_exchange, + cross_exchange_as, + self.start(), + other.start(), + self.len() + ) } /// Run a length-preserving transform over this slice's bytes in place. @@ -1129,7 +1231,7 @@ impl<'a> BStackSlice<'a> { #[cfg(all(feature = "set", feature = "atomic"))] #[inline] pub fn process(&mut self, f: F) -> io::Result<()> { - self.stack.process(self.start(), self.end(), f) + slice_meta!(self, process, process_as, self.start(), self.end(), f) } /// Reverse the byte order of this slice in place. @@ -1141,8 +1243,9 @@ impl<'a> BStackSlice<'a> { #[cfg(all(feature = "set", feature = "atomic"))] #[inline] pub fn reverse(&mut self) -> io::Result<()> { - self.stack - .process(self.start(), self.end(), |buf| buf.reverse()) + slice_meta!(self, process, process_as, self.start(), self.end(), |buf| { + buf.reverse() + }) } /// Rotate the slice in place such that the bytes at `[mid, len)` move to @@ -1162,7 +1265,7 @@ impl<'a> BStackSlice<'a> { mid <= self.len(), "rotate_left: mid must be <= slice length" ); - self.stack.process(self.start(), self.end(), |buf| { + slice_meta!(self, process, process_as, self.start(), self.end(), |buf| { buf.rotate_left(mid as usize) }) } @@ -1181,8 +1284,9 @@ impl<'a> BStackSlice<'a> { #[track_caller] pub fn rotate_right(&mut self, k: u64) -> io::Result<()> { assert!(k <= self.len(), "rotate_right: k must be <= slice length"); - self.stack - .process(self.start(), self.end(), |buf| buf.rotate_right(k as usize)) + slice_meta!(self, process, process_as, self.start(), self.end(), |buf| { + buf.rotate_right(k as usize) + }) } /// Create a cursor-based reader positioned at the start of this slice. @@ -1378,6 +1482,11 @@ impl<'a> From> for BStackSliceWriter<'a> { pub struct BStackOwnedSlice<'a, A: BStackAllocator> { allocator: &'a A, range: BStackRange, + /// Authority the views borrowed from this handle carry into the stack's + /// access-control table. `NONE` until granted via + /// [`authorize`](BStackOwnedSlice::authorize). + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities, } impl<'a, A: BStackAllocator> fmt::Debug for BStackOwnedSlice<'a, A> { @@ -1413,6 +1522,8 @@ impl<'a, A: BStackAllocator> BStackOwnedSlice<'a, A> { Self { allocator, range: unsafe { BStackRange::from_raw_parts(offset, len) }, + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, } } @@ -1427,7 +1538,12 @@ impl<'a, A: BStackAllocator> BStackOwnedSlice<'a, A> { #[inline] #[must_use] pub unsafe fn from_raw_range(allocator: &'a A, range: BStackRange) -> Self { - Self { allocator, range } + Self { + allocator, + range, + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, + } } /// Construct an empty (zero-length) owned handle. @@ -1440,6 +1556,8 @@ impl<'a, A: BStackAllocator> BStackOwnedSlice<'a, A> { Self { allocator, range: BStackRange::empty(), + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, } } @@ -1509,6 +1627,8 @@ impl<'a, A: BStackAllocator> BStackOwnedSlice<'a, A> { Self { allocator, range: BStackRange::from_bytes(bytes), + #[cfg(feature = "expensive-slice-access-control")] + auth: BStackAccessAuthorities::NONE, } } @@ -1543,6 +1663,8 @@ impl<'a, A: BStackAllocator> BStackOwnedSlice<'a, A> { BStackSlice { stack: self.allocator.stack(), range: self.range, + #[cfg(feature = "expensive-slice-access-control")] + auth: self.auth, } } @@ -1557,9 +1679,54 @@ impl<'a, A: BStackAllocator> BStackOwnedSlice<'a, A> { BStackSlice { stack: self.allocator.stack(), range: self.range, + #[cfg(feature = "expensive-slice-access-control")] + auth: self.auth, } } + /// Arm this allocation's range with `mode` in the stack's access-control + /// table. + /// + /// The public entry point for protection: an owned handle proves the range is + /// genuinely this caller's allocation, rather than an arbitrary span. A + /// tokenless caller may only tighten a range currently at + /// [`All`](BStackAccess::All); see [`protect_as`](Self::protect_as) to present + /// a capability token. + /// + /// Requires the `expensive-slice-access-control` feature. + #[cfg(feature = "expensive-slice-access-control")] + #[inline] + pub fn protect(&self, mode: BStackAccess) -> io::Result<()> { + self.allocator + .stack() + .protect_as((), self.range.start(), self.range.len(), mode) + } + + /// [`protect`](Self::protect) presenting an access token — the token sibling + /// for arming [`Prot`](BStackAccess::Prot)/[`Alloc`](BStackAccess::Alloc) + /// ranges or re-moding a range this token governs. + /// + /// Requires the `expensive-slice-access-control` feature. + #[cfg(feature = "expensive-slice-access-control")] + #[inline] + pub fn protect_as(&self, auth: impl BStackAuthority, mode: BStackAccess) -> io::Result<()> { + self.allocator + .stack() + .protect_as(auth, self.range.start(), self.range.len(), mode) + } + + /// Grant this handle the authority carried by `auth`, so every view borrowed + /// from it ([`as_slice`](Self::as_slice) / [`as_slice_mut`](Self::as_slice_mut)) + /// and their I/O may reach a range it was authorized for. A token minted from + /// a different stack grants nothing. + /// + /// Requires the `expensive-slice-access-control` feature. + #[cfg(feature = "expensive-slice-access-control")] + #[inline] + pub fn authorize(&mut self, auth: impl BStackAuthority) { + self.auth = auth.authorities_for(self.allocator.stack()); + } + /// Read the entire allocation into a new `Vec`. /// /// Internally borrows a [`BStackSlice`] via [`as_slice`](Self::as_slice) @@ -2557,7 +2724,7 @@ impl<'a> io::Read for BStackSliceReader<'a> { let available = (self.slice.len() - self.cursor) as usize; let n = buf.len().min(available); let abs_start = self.slice.start() + self.cursor; - self.slice.stack.get_into(abs_start, &mut buf[..n])?; + slice_meta!(self.slice, get_into, get_into_as, abs_start, &mut buf[..n])?; self.cursor += n as u64; Ok(n) } @@ -2744,7 +2911,7 @@ impl<'a> io::Write for BStackSliceWriter<'a> { let available = (self.slice.len() - self.cursor) as usize; let n = buf.len().min(available); let abs_start = self.slice.start() + self.cursor; - self.slice.stack.set(abs_start, &buf[..n])?; + slice_meta!(self.slice, set, set_as, abs_start, &buf[..n])?; self.cursor += n as u64; Ok(n) } diff --git a/src/lib.rs b/src/lib.rs index 73ce572..46b3a4f 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -707,6 +707,16 @@ pub use fault::{FaultPolicy, FaultState}; mod alloc_fuzz; mod test; +#[cfg(feature = "expensive-slice-access-control")] +mod acl_core; +#[cfg(feature = "expensive-slice-access-control")] +pub use acl_core::{AccessOp, BStackAccess, BStackAccessAuthorities, BStackAccessRequirement}; + +mod acl; +use acl::acl_check; +#[cfg(feature = "expensive-slice-access-control")] +pub use acl::{BStackAllocAuthority, BStackAuthority, BStackProtection}; + #[cfg(feature = "alloc")] mod alloc; #[cfg(feature = "alloc")] @@ -944,6 +954,21 @@ pub struct BStack { /// enabled; release builds carry neither the field nor its per-call branch. #[cfg(all(debug_assertions, feature = "fault-injection"))] fault: fault::FaultState, + /// Range access-control point table, under its own lock (separate from the + /// stack lock so the locked-region read fast path can still consult it). Not + /// persisted — reopening clears it. + #[cfg(feature = "expensive-slice-access-control")] + acl: RwLock, + /// One-shot guard-token permit: `Some(())` until [`take_protection`] moves it + /// out to mint the token, `None` after (until [`return_protection`] hands it + /// back). The move-out is what makes the token mean anything — there is only + /// ever one guard authority per stack. + #[cfg(feature = "expensive-slice-access-control")] + protection: Mutex>, + /// One-shot allocator-token permit, the [`take_alloc_authority`] / + /// [`return_alloc_authority`] counterpart of [`protection`](Self::protection). + #[cfg(feature = "expensive-slice-access-control")] + alloc_authority: Mutex>, } // `BStack` is auto-`Send + Sync` on every platform: all fields @@ -1156,6 +1181,12 @@ impl BStack { cache: Mutex::new(Vec::new()), #[cfg(all(debug_assertions, feature = "fault-injection"))] fault: fault::FaultState::new(), + #[cfg(feature = "expensive-slice-access-control")] + acl: RwLock::new(acl_core::PointTable::new()), + #[cfg(feature = "expensive-slice-access-control")] + protection: Mutex::new(Some(())), + #[cfg(feature = "expensive-slice-access-control")] + alloc_authority: Mutex::new(Some(())), }) } @@ -1295,6 +1326,12 @@ impl BStack { return Ok(logical_offset); } + acl_check!( + self, + logical_offset, + logical_offset + data.len() as u64, + Write + ); fault_point!(self, "push"); if let Err(e) = file.write_all(data) { // A failed rollback leaves a stale tail past the committed length: @@ -1339,6 +1376,7 @@ impl BStack { return Ok(logical_offset); } + acl_check!(self, logical_offset, logical_offset + n, Write); fault_point!(self, "extend"); let new_file_end = file_end + n; Self::mark_replay(replay, file.set_len(new_file_end))?; @@ -1405,6 +1443,7 @@ impl BStack { "extend_sparse: payload size + length overflows u64" ) })?; + acl_check!(self, logical_offset, new_len, Write); fault_point!(self, "extend_sparse"); let one = [(0u64, buf)]; let blocks: &[(u64, &[u8])] = if buf.is_empty() { &[] } else { &one }; @@ -1477,6 +1516,7 @@ impl BStack { "extend_sparse_batched: payload size + length overflows u64" ) })?; + acl_check!(self, logical_offset, new_len, Write); fault_point!(self, "extend_sparse_batched"); Self::mark_replay( replay, @@ -1521,11 +1561,13 @@ impl BStack { format!("resize({target}) would shrink payload below locked length ({locked})") )); } + acl_check!(self, target, data_size, Truncate); fault_point!(self, "resize"); Self::mark_replay(replay, commit_shrink(file, clen, target))?; return Ok(data_size); } + acl_check!(self, data_size, target, Write); fault_point!(self, "resize"); Self::mark_replay(replay, file.set_len(HEADER_SIZE + target))?; Self::mark_replay(replay, commit_grow(file, clen, target, data_size, file_end))?; @@ -1557,6 +1599,7 @@ impl BStack { return Ok(data_size); } + acl_check!(self, data_size, target, Write); fault_point!(self, "ensure"); Self::mark_replay(replay, file.set_len(HEADER_SIZE + target))?; Self::mark_replay(replay, commit_grow(file, clen, target, data_size, file_end))?; @@ -1669,6 +1712,7 @@ impl BStack { format!("pop({n}) would shrink payload below locked length ({locked})") )); } + acl_check!(self, new_data_len, data_size, Truncate); let mut buf = vec![0u8; n as usize]; fault_point!(self, "pop"); read_at(file, new_data_len, &mut buf)?; @@ -1711,6 +1755,7 @@ impl BStack { format!("peek offset ({offset}) exceeds payload size ({data_size})") )); } + acl_check!(self, offset, data_size, Read); fault_point!(self, "peek"); pread_exact(file, HEADER_SIZE + offset, (data_size - offset) as usize) } @@ -1726,6 +1771,7 @@ impl BStack { format!("peek offset ({offset}) exceeds payload size ({data_size})") )); } + acl_check!(self, offset, data_size, Read); fault_point!(self, "peek"); file.seek(SeekFrom::Start(HEADER_SIZE + offset))?; let mut buf = vec![0u8; (data_size - offset) as usize]; @@ -1752,6 +1798,9 @@ impl BStack { /// /// Fails with [`InterruptedWrite`] while an earlier write is pending replay. pub fn get(&self, start: u64, end: u64) -> io::Result> { + // The read check runs before the locked-region fast path, so a `Locked` + // range is rejected even though that path bypasses the stack lock. + acl_check!(self, start, end, Read); if end < start { return Err(io_error!( InvalidInput, @@ -1852,6 +1901,7 @@ impl BStack { let end = offset .checked_add(len) .ok_or_else(|| io_error!(InvalidInput, "peek_into: offset + len overflows u64"))?; + acl_check!(self, offset, end, Read); #[cfg(any(unix, windows))] { let guard = self.read_lock()?; @@ -1914,6 +1964,8 @@ impl BStack { let end = start .checked_add(len) .ok_or_else(|| io_error!(InvalidInput, "get_into: start + len overflows u64"))?; + // Before the locked-region fast path, so a `Locked` range is still denied. + acl_check!(self, start, end, Read); // Fast-path: locked region is immutable — serve from cache or pread. #[cfg(any(unix, windows))] { @@ -2006,6 +2058,7 @@ impl BStack { format!("pop_into({n}) would shrink payload below locked length ({locked})") )); } + acl_check!(self, new_data_len, data_size, Truncate); fault_point!(self, "pop_into"); read_at(file, new_data_len, buf)?; Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; @@ -2048,6 +2101,7 @@ impl BStack { format!("discard({n}) would shrink payload below locked length ({locked})") )); } + acl_check!(self, new_data_len, data_size, Truncate); fault_point!(self, "discard"); Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; Ok(()) @@ -2091,6 +2145,7 @@ impl BStack { // our write, letting us mutate a now-immutable byte. let locked = self.locked.load(Ordering::Acquire); check_offset_unlocked("set", offset, end, locked)?; + acl_check!(self, offset, end, Write); let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); if end > data_size { return Err(io_error!( @@ -2136,6 +2191,7 @@ impl BStack { // Load `locked` under the write lock (see `set` for rationale). let locked = self.locked.load(Ordering::Acquire); check_offset_unlocked("zero", offset, end, locked)?; + acl_check!(self, offset, end, Write); let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); if end > data_size { return Err(io_error!( @@ -2201,6 +2257,7 @@ impl BStack { // Load `locked` under the write lock (see `set` for rationale). let locked = self.locked.load(Ordering::Acquire); check_offset_unlocked("repeat", offset, end, locked)?; + acl_check!(self, offset, end, Write); let data_size = file.seek(SeekFrom::End(0))?.saturating_sub(HEADER_SIZE); if end > data_size { return Err(io_error!( @@ -2273,6 +2330,7 @@ impl BStack { format!("atrunc: operation would modify locked region [0, {locked})") )); } + acl_check!(self, new_tail_start, data_size, Truncate); fault_point!(self, "atrunc"); Self::mark_replay( replay, @@ -2323,6 +2381,7 @@ impl BStack { format!("splice: operation would modify locked region [0, {locked})") )); } + acl_check!(self, new_tail_start, data_size, Truncate); fault_point!(self, "splice"); // Read the bytes to remove before any mutation. let mut removed = vec![0u8; n as usize]; @@ -2379,6 +2438,7 @@ impl BStack { format!("splice_into: operation would modify locked region [0, {locked})") )); } + acl_check!(self, new_tail_start, data_size, Truncate); fault_point!(self, "splice_into"); // Read the bytes to remove before any mutation. read_at(file, new_tail_start, old)?; @@ -2416,6 +2476,7 @@ impl BStack { if buf.is_empty() { return Ok(true); } + acl_check!(self, data_size, data_size + buf.len() as u64, Write); fault_point!(self, "try_extend"); if let Err(e) = file.write_all(buf) { // A failed rollback leaves a stale tail past the committed length: @@ -2465,6 +2526,7 @@ impl BStack { n, "try_extend_zeros: data_size + n overflows u64", )?; + acl_check!(self, data_size, new_len, Write); fault_point!(self, "try_extend_zeros"); Self::mark_replay(replay, file.set_len(HEADER_SIZE + new_len))?; Self::mark_replay( @@ -2526,6 +2588,7 @@ impl BStack { length, "try_extend_sparse: data_size + length overflows u64", )?; + acl_check!(self, data_size, new_len, Write); fault_point!(self, "try_extend_sparse"); let one = [(0u64, buf)]; let blocks: &[(u64, &[u8])] = if buf.is_empty() { &[] } else { &one }; @@ -2594,6 +2657,7 @@ impl BStack { length, "try_extend_sparse_batched: data_size + length overflows u64", )?; + acl_check!(self, data_size, new_len, Write); fault_point!(self, "try_extend_sparse_batched"); Self::mark_replay( replay, @@ -2648,6 +2712,7 @@ impl BStack { format!("try_discard: would shrink payload below locked length ({locked})") )); } + acl_check!(self, new_data_len, data_size, Truncate); fault_point!(self, "try_discard"); Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; Ok(true) @@ -2693,6 +2758,7 @@ impl BStack { r.start )); } + acl_check!(self, r.start, r.end, Read); } #[cfg(any(unix, windows))] { @@ -2790,6 +2856,7 @@ impl BStack { format!("get_batched_into: end ({end}) exceeds payload size ({data_size})",) )); } + acl_check!(self, ptr, end, Read); pread_exact_into(file, HEADER_SIZE + ptr, buf)?; } Ok(()) @@ -2813,6 +2880,7 @@ impl BStack { format!("get_batched_into: end ({end}) exceeds payload size ({data_size})",) )); } + acl_check!(self, ptr, end, Read); file.seek(SeekFrom::Start(HEADER_SIZE + ptr))?; file.read_exact(buf)?; } @@ -2865,6 +2933,7 @@ impl BStack { format!("get_batched_gen: end ({end}) exceeds payload size ({data_size})") )); } + acl_check!(self, offset, end, Read); // Per-step read fault: stands in for this read's I/O and ends // the whole call, as a genuine read failure here does. fault_point!(self, "get_batched_gen:read"); @@ -2891,6 +2960,7 @@ impl BStack { format!("get_batched_gen: end ({end}) exceeds payload size ({data_size})") )); } + acl_check!(self, offset, end, Read); // Per-step read fault: stands in for this read's I/O and ends // the whole call, as a genuine read failure here does. fault_point!(self, "get_batched_gen:read"); @@ -2945,6 +3015,7 @@ impl BStack { format!("replace: operation would modify locked region [0, {locked})") )); } + acl_check!(self, new_tail_start, data_size, Truncate); fault_point!(self, "replace"); let mut old_tail = vec![0u8; n as usize]; read_at(file, new_tail_start, &mut old_tail)?; @@ -3194,6 +3265,7 @@ impl BStack { format!("swap: range [{offset}, {end}) exceeds payload size ({data_size})") )); } + acl_check!(self, offset, end, Write); fault_point!(self, "swap"); let mut old = vec![0u8; buf.len()]; read_at(file, offset, &mut old)?; @@ -3246,6 +3318,7 @@ impl BStack { format!("swap_into: range [{offset}, {end}) exceeds payload size ({data_size})") )); } + acl_check!(self, offset, end, Write); fault_point!(self, "swap_into"); let mut tmp = vec![0u8; buf.len()]; read_at(file, offset, &mut tmp)?; @@ -3307,6 +3380,7 @@ impl BStack { format!("cas: range [{offset}, {end}) exceeds payload size ({data_size})") )); } + acl_check!(self, offset, end, Write); fault_point!(self, "cas"); let mut current = vec![0u8; old.len()]; read_at(file, offset, &mut current)?; @@ -3394,6 +3468,8 @@ impl BStack { if n == 0 { return Ok(()); } + acl_check!(self, a, a_end, Write); + acl_check!(self, b, b_end, Write); fault_point!(self, "cross_exchange"); Self::mark_replay(replay, journaled_exchange(file, data_size, a, b, n)) } @@ -3461,6 +3537,8 @@ impl BStack { if from == to { return Ok(()); } + acl_check!(self, from, from_end, Read); + acl_check!(self, to, to_end, Write); fault_point!(self, "copy"); // Write-strategy hierarchy (see `algos/WIP.md`): // * destination within one aligned block → single-block atomic write @@ -3539,6 +3617,7 @@ impl BStack { format!("process: range [{start}, {end}) overlaps locked region [0, {locked})") )); } + acl_check!(self, start, end, Write); fault_point!(self, "process"); let mut buf = vec![0u8; n as usize]; if n > 0 { @@ -3674,6 +3753,7 @@ impl BStack { // paths below serve without touching the disk, so the // schedule does not shift with cache state. fault_point!(self, "process_gen:read"); + acl_check!(self, offset, end, Read); // Fast path: locked bytes are immutable, so they can be // served from the cache or via a lock-free pread instead // of going through the held file handle — mirroring how @@ -3726,6 +3806,7 @@ impl BStack { ) )); } + acl_check!(self, offset, end, Write); if !data.is_empty() { Self::mark_replay(replay, set_in_place(file, data_size, offset, data))?; } @@ -3826,6 +3907,8 @@ impl BStack { ) )); } + acl_check!(self, a_offset, a_end, Write); + acl_check!(self, b_offset, b_end, Write); if len > 0 { Self::mark_replay( replay, @@ -3838,6 +3921,12 @@ impl BStack { if !data.is_empty() { let file_end = file.seek(SeekFrom::End(0))?; let logical_offset = file_end - HEADER_SIZE; + acl_check!( + self, + logical_offset, + logical_offset + data.len() as u64, + Write + ); if let Err(e) = file.write_all(data) { // A failed rollback leaves a stale tail past the committed length: // defer it to the next write's replay. @@ -3871,6 +3960,7 @@ impl BStack { ) )); } + acl_check!(self, new_data_len, data_size, Truncate); if n > 0 { read_at(file, new_data_len, buf)?; Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; @@ -3895,6 +3985,7 @@ impl BStack { ) )); } + acl_check!(self, new_data_len, data_size, Truncate); if len > 0 { Self::mark_replay(replay, commit_shrink(file, clen, new_data_len))?; } @@ -3916,6 +4007,7 @@ impl BStack { format!("process_gen: atrunc would modify locked region [0, {locked})") )); } + acl_check!(self, new_tail_start, data_size, Truncate); if n != 0 || !data.is_empty() { let file_end = HEADER_SIZE + data_size; Self::mark_replay( @@ -3942,6 +4034,7 @@ impl BStack { format!("process_gen: splice would modify locked region [0, {locked})") )); } + acl_check!(self, new_tail_start, data_size, Truncate); if n != 0 || !new.is_empty() { // Read the removed bytes before any mutation. read_at(file, new_tail_start, old)?; @@ -3969,6 +4062,7 @@ impl BStack { length, "process_gen: sparse data_size + length overflows u64", )?; + acl_check!(self, data_size, new_len, Write); Self::mark_replay( replay, commit_sparse_extend(file, clen, data_size, file_end, new_len, &blocks), @@ -4068,6 +4162,7 @@ impl BStack { ) )); } + acl_check!(self, *off, end, Write); } fault_point!(self, "set_batched"); // A lone write cannot overlap anything and is already atomic on its own, @@ -4187,14 +4282,44 @@ impl BStack { // same route as a genuine one. feedback = match inplace_validate_read(offset, buf.len() as u64, data_size) { Err(e) => Err(e), - Ok(()) => match fault_probe!(self, "inplace_gen:read") { - Some(e) => Err(e), - None => inplace_overlay_read(file, data_size, offset, buf, &overlay), - }, + Ok(()) => { + // A denied read is reported through `feedback`, the same + // route a genuine read failure takes. Validation passed, + // so `offset + len` cannot overflow. + #[cfg(feature = "expensive-slice-access-control")] + let gate = self.acl_check( + offset, + offset + buf.len() as u64, + AccessOp::Read, + BStackAccessAuthorities::NONE, + ); + #[cfg(not(feature = "expensive-slice-access-control"))] + let gate: io::Result<()> = Ok(()); + gate.and_then(|()| { + fault_probe!(self, "inplace_gen:read").map_or_else( + || inplace_overlay_read(file, data_size, offset, buf, &overlay), + Err, + ) + }) + } }; } Some(BStackGenOp::Write { offset, data }) => { feedback = inplace_validate_write(offset, data, data_size, locked); + // A denial is routed to the generator like a validation error, + // not returned from the call. `is_ok` implies the range already + // passed `inplace_validate_write`'s `checked_end`, so `offset + + // len` cannot overflow here. + #[cfg(feature = "expensive-slice-access-control")] + if feedback.is_ok() { + let end = offset + data.len() as u64; + feedback = self.acl_check( + offset, + end, + AccessOp::Write, + BStackAccessAuthorities::NONE, + ); + } if feedback.is_ok() && !data.is_empty() { inplace_overlay_insert(&mut overlay, offset, OverlayData::Literal(data)); } @@ -4374,6 +4499,8 @@ impl BStack { ) )); } + acl_check!(self, a_offset, a_end, Read); + acl_check!(self, b_offset, b_end, Write); fault_point!(self, "eq_crds"); let mut a_current = vec![0u8; a_expected.len()]; if !a_expected.is_empty() { @@ -4451,6 +4578,8 @@ impl BStack { ) )); } + acl_check!(self, a_offset, a_end, Read); + acl_check!(self, b_offset, b_end, Write); fault_point!(self, "ne_crds"); let mut a_current = vec![0u8; a_expected.len()]; if !a_expected.is_empty() { @@ -4549,6 +4678,8 @@ impl BStack { ) )); } + acl_check!(self, a_offset, a_end, Read); + acl_check!(self, b_offset, b_end, Write); fault_point!(self, "masked_eq_crds"); let mut a_current = vec![0u8; a_expected.len()]; if !a_expected.is_empty() { diff --git a/src/test.rs b/src/test.rs index 9c294f5..e245d30 100644 --- a/src/test.rs +++ b/src/test.rs @@ -6410,11 +6410,24 @@ mod first_fit_tests { alloc.dealloc(b).unwrap(); let stack = alloc.into_stack(); - // Corrupt: set recovery_needed=1 and scramble free_head to garbage - stack.set(24, 1u32.to_le_bytes()).unwrap(); // flags byte → recovery_needed=1 - stack - .set(FREE_HEAD_OFFSET, 0xDEADBEEFu64.to_le_bytes()) - .unwrap(); + // Corrupt: set recovery_needed=1 and scramble free_head to garbage. These + // are allocator-metadata writes (the header is protected under the ACL + // feature), so present ALLOC authority. + #[cfg(not(feature = "expensive-slice-access-control"))] + { + stack.set(24, 1u32.to_le_bytes()).unwrap(); // flags byte → recovery_needed=1 + stack + .set(FREE_HEAD_OFFSET, 0xDEADBEEFu64.to_le_bytes()) + .unwrap(); + } + #[cfg(feature = "expensive-slice-access-control")] + { + let auth = crate::BStackAccessAuthorities::ALLOC; + stack.set_as(auth, 24, 1u32.to_le_bytes()).unwrap(); + stack + .set_as(auth, FREE_HEAD_OFFSET, 0xDEADBEEFu64.to_le_bytes()) + .unwrap(); + } drop(stack); // Re-open: recovery should run and rebuild the free list from is_free flags @@ -6757,10 +6770,20 @@ mod first_fit_tests { // and the in-memory poison set. The disk flag drives reopen recovery; the // poison is what the single atomic-call paths (which no longer arm the flag // themselves) check to refuse. + #[cfg(not(feature = "expensive-slice-access-control"))] alloc .stack() .set(24u64, 1u32.to_le_bytes().as_slice()) .unwrap(); + #[cfg(feature = "expensive-slice-access-control")] + alloc + .stack() + .set_as( + crate::BStackAccessAuthorities::ALLOC, + 24u64, + 1u32.to_le_bytes().as_slice(), + ) + .unwrap(); alloc.poison_for_test(); // A free-list-touching op must now refuse. `alloc(64)` reuses the freed