Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 35 additions & 0 deletions docs/sandboxing.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,23 @@ Your global `git config user.name` and `user.email` (if configured on the host)

Agent login/config state (e.g. Claude Code's `~/.claude`, Codex's `~/.codex`) is kept in a Docker-managed named volume, not a bind mount of your real home directory, so it survives across `agent run` invocations without exposing anything else on the host. This volume is shared across every profile, project, *and image* using the Docker backend on this machine — logging in once covers all of them, even after switching to a completely different custom image or Dockerfile, since the volume is mounted at the same container path (`/home/agent`) regardless of which image runs.

### Admin shell and installing tools

```bash
stashbase agent docker shell
```

This opens an interactive `bash` in the default sandbox image (`--image <ref>` for another one) with the persistent home volume mounted and `/home/agent` as the working directory. No profile is needed. None of a run's sandboxing applies: there's no per-run network or firewall and no working-directory mount. Use it to log in, edit agent config, install tools, or inspect what the sandbox sees. It runs as the same user agent runs do (your uid on Linux, root on macOS), so files it creates stay writable for them. `--root` forces root on Linux.

Only `$HOME` persists after the shell exits. The image points user-level installs there, so these are kept and available to every sandboxed run, in every repo and profile:

```bash
npm i -g bun # → ~/.npm-global/bin (on PATH)
pip install --user x # → ~/.local/bin (on PATH)
```

These home directories come after the system paths on `PATH`, so a tool installed there never overrides the image's own binaries. `rm -rf ~/.npm-global` resets the global npm installs. System packages (`apt install …`) go into the container's own filesystem and are lost on exit. Add them with a custom `sandbox.dockerfile` instead (see [Custom images](#custom-images)).

### Codex and subscription login

Codex's normal OAuth login flow opens a browser that redirects to a local HTTP callback server. That callback listens inside the container's own network namespace, which the host browser cannot reach — the container's `localhost` is not your machine's `localhost`. Use Codex's device-code flow instead, which doesn't depend on a local callback at all:
Expand All @@ -136,6 +153,24 @@ cpus = "1.5" # docker run --cpus

Or per invocation, without editing the file: `--docker-memory <value>` / `--docker-cpus <value>` (same override precedence as `--docker-image`/`--docker-dockerfile`, but these don't imply the Docker backend on their own — they're only meaningful once Docker is already selected). `agent validate` checks the value looks like something Docker would accept before you ever try to run it.

### Isolated paths (e.g. `node_modules`)

The container is Linux, but it sees your working directory exactly as it is on the host. On macOS or Windows that includes a `node_modules` installed for the host OS. Packages that ship native binaries (Nx, esbuild, swc, rollup, …) only have the host's binary there, so they fail inside the sandbox. It works the other way round too: an `npm ci` inside the sandbox would replace the host's install with Linux binaries.

List such directories under `isolated_paths` to give the container its own copy:

```toml
[sandbox]
backend = "docker"
isolated_paths = ["node_modules"]
```

Each entry (relative to the working directory) gets a Docker volume for that repo, mounted over the directory inside the container. The host's copy is never touched. The volume starts empty, so install once from inside the sandbox (`npm ci`); it persists across runs of that repo. Nested workspace folders (`apps/web/node_modules`) are listed explicitly if needed. The same works for `.venv`, `target/`, etc. On a Linux host this is usually unnecessary, since the host's binaries already match the container.

Or per invocation, without editing the file: `--docker-isolated-paths node_modules,.venv` (comma-separated). It adds to the profile's list rather than replacing it, and like `--docker-cpus` it only applies once Docker is the selected backend.

`stashbase agent docker cleanup --isolated-paths` lists these volumes and removes them after confirmation (`--yes` to skip).

### Cleaning up after a crash

Every per-run Docker network (and its two containers) is named after that run's own session id — the same `ags_...` id shown in "Agent session"/"Audit session" and used for the audit log filename — specifically so leftovers can be traced back to the run that created them. Normal exit paths, including Ctrl+C, tear both containers and the network down as part of `agent run` itself; only a crash or a forceful `SIGKILL` of the `stashbase` process can leave them behind.
Expand Down
11 changes: 11 additions & 0 deletions src/cmd/agent.rs
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,10 @@ pub struct AgentDockerCleanupCommand {
/// Remove every leftover resource without prompting for confirmation
#[arg(long)]
pub yes: bool,

/// Instead of leftover networks, remove the per-repo volumes backing profiles' `sandbox.isolated_paths` (e.g. the container's own node_modules; reinstalled on next use)
#[arg(long)]
pub isolated_paths: bool,
}

#[derive(Debug, Args)]
Expand Down Expand Up @@ -271,6 +275,13 @@ pub struct AgentRunCommand {
#[arg(long)]
pub docker_cpus: Option<String>,

/// Add to the profile's `[sandbox] isolated_paths` for this run only:
/// a directory relative to the working directory (e.g. "node_modules")
/// that gets its own per-repo volume inside the container.
/// Comma-separated, e.g. `--docker-isolated-paths node_modules,.venv`.
#[arg(long, value_name = "PATHS", value_delimiter = ',')]
pub docker_isolated_paths: Vec<String>,

/// Store metadata-only proxy audit events locally
#[arg(
long,
Expand Down
92 changes: 92 additions & 0 deletions src/handlers/agent/docker.rs
Original file line number Diff line number Diff line change
Expand Up @@ -235,6 +235,9 @@ pub async fn handle_docker_cleanup_command(
if let Some(error) = crate::handlers::run::docker_sandbox::docker_enforcement_error() {
anyhow::bail!("Docker sandbox backend unavailable: {error}");
}
if command.isolated_paths {
return cleanup_isolated_path_volumes(command.yes, raw_output, silent);
}

let candidates: Vec<_> = list_networks_with_liveness()?
.into_iter()
Expand Down Expand Up @@ -369,6 +372,95 @@ pub async fn handle_docker_cleanup_command(
Ok(())
}

/// `cleanup --isolated-paths`: unlike leftover networks, these volumes are
/// deliberate persistent state (a repo's container-side node_modules etc.),
/// so they're only ever removed on this explicit request. Removing one just
/// means the next run starts with an empty directory there to reinstall into.
fn cleanup_isolated_path_volumes(yes: bool, raw_output: bool, silent: bool) -> Result<()> {
use crate::handlers::run::docker_sandbox::{list_isolated_path_volumes, remove_volume};

let volumes = list_isolated_path_volumes()
.map_err(|error| anyhow::anyhow!("failed to list isolated-path volumes: {error}"))?;
if volumes.is_empty() {
if raw_output {
println!(
"{}",
get_formatted_json_string(
&serde_json::json!({ "removed": [], "failed": [] }),
true
)?
);
} else if !silent {
println!("No isolated-path volumes found.");
}
return Ok(());
}

if !silent && !raw_output {
println!("Found {} isolated-path volume(s):\n", volumes.len());
for volume in &volumes {
println!(" {} {}/{}", volume.name, volume.repo, volume.path);
}
println!();
}

let should_remove = if yes {
true
} else if silent {
anyhow::bail!(
"{} isolated-path volume(s) found; re-run with --yes to remove them, or without --silent to be prompted",
volumes.len()
);
} else {
let confirmed = crate::utils::interaction::confirm_opt(&format!(
"Remove all {} volume(s) listed above?",
volumes.len()
))
.unwrap_or(false);
let _ = dialoguer::console::Term::stdout().show_cursor();
confirmed
};
if !should_remove {
if !silent && !raw_output {
println!("Nothing removed.");
}
return Ok(());
}

let mut removed = Vec::new();
let mut failures = Vec::new();
for volume in &volumes {
match remove_volume(&volume.name) {
Ok(()) => {
if !silent && !raw_output {
println!("{} {}", "Removed".green_if_tty(), volume.name);
}
removed.push(&volume.name);
}
// Most likely still mounted by a running sandbox.
Err(error) => failures.push(format!("{}: {error}", volume.name)),
}
}

if raw_output {
println!(
"{}",
get_formatted_json_string(
&serde_json::json!({ "removed": removed, "failed": failures }),
true,
)?
);
}
if !failures.is_empty() {
anyhow::bail!(
"failed to remove {} volume(s):\n{}",
failures.len(),
failures.join("\n")
);
}
Ok(())
}

/// Checks whether the Docker sandbox backend can actually run on this
/// machine, without starting a real sandboxed run to find out. Reports
/// each check independently (the `docker` CLI on PATH, the daemon
Expand Down
2 changes: 2 additions & 0 deletions src/handlers/agent/mcp.rs
Original file line number Diff line number Diff line change
Expand Up @@ -681,6 +681,7 @@ async fn proxied_client(
sandbox_dockerfile: profile.sandbox.dockerfile.clone(),
sandbox_memory: profile.sandbox.memory.clone(),
sandbox_cpus: profile.sandbox.cpus.clone(),
sandbox_isolated_paths: profile.sandbox.isolated_paths.clone(),
};
let proxy = Proxy::start_with_port(secrets, policy, None, None).await?;
let proxy_url = proxy.child_env()["HTTPS_PROXY"].clone();
Expand Down Expand Up @@ -834,6 +835,7 @@ async fn remote_proxied_client(
sandbox_dockerfile: profile.sandbox.dockerfile.clone(),
sandbox_memory: profile.sandbox.memory.clone(),
sandbox_cpus: profile.sandbox.cpus.clone(),
sandbox_isolated_paths: profile.sandbox.isolated_paths.clone(),
};
let proxy = Proxy::start_remote_with_port(
RemoteProxyConfig {
Expand Down
9 changes: 9 additions & 0 deletions src/handlers/agent/validate.rs
Original file line number Diff line number Diff line change
Expand Up @@ -433,6 +433,11 @@ fn validate_profile(profile: &AgentProfile) -> Vec<Check> {
));
}
}
for path in &profile.sandbox.isolated_paths {
if let Err(error) = crate::handlers::run::docker_sandbox::normalize_isolated_path(path) {
checks.push(fail("Sandbox isolated path", error));
}
}

let mut bindings: HashMap<&str, Vec<&str>> = HashMap::new();
let mut child_envs: HashMap<&str, Vec<&str>> = HashMap::new();
Expand Down Expand Up @@ -1223,6 +1228,7 @@ mod tests {
};
profile.sandbox.memory = Some("not-a-memory-value".to_owned());
profile.sandbox.cpus = Some("not-a-number".to_owned());
profile.sandbox.isolated_paths = vec!["../outside".to_owned()];

let checks = validate_profile(&profile);
assert!(checks
Expand All @@ -1231,6 +1237,9 @@ mod tests {
assert!(checks
.iter()
.any(|check| check.status == Status::Fail && check.name == "Sandbox CPU limit"));
assert!(checks
.iter()
.any(|check| check.status == Status::Fail && check.name == "Sandbox isolated path"));
}

#[test]
Expand Down
6 changes: 6 additions & 0 deletions src/handlers/entry/root.rs
Original file line number Diff line number Diff line change
Expand Up @@ -802,6 +802,11 @@ pub async fn handle_cli(args: Cli) {
if let Some(cpus) = &agent_run.docker_cpus {
profile.sandbox.cpus = Some(cpus.clone());
}
for path in &agent_run.docker_isolated_paths {
if !profile.sandbox.isolated_paths.contains(path) {
profile.sandbox.isolated_paths.push(path.clone());
}
}
crate::handlers::agent::validate::ensure_profile_is_valid_for_run(&profile)?;
// Egress policy is meaningful only when the child cannot opt out of
// its proxy environment. Contain every session to the loopback
Expand Down Expand Up @@ -1001,6 +1006,7 @@ pub async fn handle_cli(args: Cli) {
sandbox_dockerfile: profile.sandbox.dockerfile.clone(),
sandbox_memory: profile.sandbox.memory.clone(),
sandbox_cpus: profile.sandbox.cpus.clone(),
sandbox_isolated_paths: profile.sandbox.isolated_paths.clone(),
};
let policy_fingerprint = policy.fingerprint();
let profile_source = directory_source
Expand Down
Loading
Loading