Bounded context. Bounded cost. Unbounded codebases.
An AI software-engineering platform for long-running work on large repositories, with bounded context, deterministic verification, persistent task state and optional frontier escalation. It runs a local model by default, or a cloud model API you choose.
Quick start · How it works · Validation · Verification · Security · Limitations · Docs
A real run with the local Qwen3.6-35B-A3B on an RTX 4060 laptop. Waits while the model works are shortened; the task took 1m43s.
The existing tests pass on the buggy code, so a green build proves nothing: BoundedCode asks for a test that fails without the fix.
Important
Public Alpha. BoundedCode is usable and has been validated on a small held-out sample of real software-engineering tasks (tasks not used during development and never shown to the agent). It has been tested on one machine with one model. Commands, configuration and APIs may change, and it is not production-ready. Issue reports, compatibility reports and contributions are welcome.
Coding agents working on large codebases tend to fail in two ways:
- They overflow context. Pouring the repository into the prompt exceeds what a local model can hold and drives up frontier cost.
- They mistake green tests for done. An agent reports success because the tests pass, even when no test exercises the change.
BoundedCode is built around the opposite defaults:
| Principle | What it means |
|---|---|
| Bounded context | The model starts from a small task-specific pack drawn from repository intelligence and reads further code through tools as needed, instead of receiving the whole repository. |
| Local by default | Inference runs on your machine through llama.cpp unless you choose a cloud model API (OpenAI, Anthropic, Gemini or an OpenAI-compatible service; unvalidated, see Cloud models). A frontier model is an optional, policy-triggered exception: enabled but not triggered in the second validation; in the first validation 4 frontier calls were sent and no task was accepted. |
| Evidence, not just green tests | A task is TASK_VERIFIED only when a test it adds fails on the base commit and passes with the change (fail-before/pass-after evidence, not proof of correctness). In a multi-repository task, every gRPC, protobuf or OpenAPI link that the change affects must also be shown compatible by the repositories' own checks (experimental). |
| Durable tasks | A persistent ledger lets long tasks resume after Ctrl-C, a crash or a reboot. |
| Contained agent | The agent runs in a network-less container on its own git worktree. Nothing is pushed or merged for you. |
| Stage | Result | Report |
|---|---|---|
| Initial validation: 8 real public tasks, frozen build | 0/8, then 1/8 after the first defect fixes | report |
| Engineering: the failures used as a development corpus (development evidence, not a validation) | fixes to verification, agent tooling, resource handling, retrieval and execution control | failure-driven · targeted |
| Second validation (held out): 6 tasks not used during development and never shown to the agent, screened for issue-derivable acceptance tests, run once | 5 of 6 strict TASK_VERIFIED and passing the datasets' hidden acceptance tests (hidden from the agent) · 6 of 6 hidden tests pass · all 5 successes local-only: no frontier calls (escalation enabled, not triggered) · 0 false verification passes among the 5 TASK_VERIFIED tasks |
report |
In a fresh small validation on 6 public engineering tasks not used during development and never shown to the agent, whose hidden acceptance criteria were screened for consistency with the issue before execution, 5 of the 6 tasks succeeded, and all 5 successes used only the local model; no task made a frontier call.
The tasks were selected before execution from the SWE-bench Multilingual and Multi-SWE-bench benchmark datasets. They cover Go, JavaScript, TypeScript, an infrastructure tool and a repository of 1.8 M estimated source tokens. Screening rejected 10 of 24 candidates whose hidden tests could not be derived from their issue. There was no human code intervention.
Context use was roughly 18–35 K tokens per task across repositories of 0.12–1.83 M estimated source tokens, so the share of the repository that entered context depends on repository size: at most about 1.6 % on the largest repository, 29.1 % on the smallest (gin). These figures include all tool output (files read, search and test output) and are therefore upper bounds; source tokens are estimated as bytes × 10/32.
The 0 false verification passes covers the 5 TASK_VERIFIED tasks, on a task
set screened for issue-derivable tests. In development runs with the same
gate design, tasks were TASK_VERIFIED but failed hidden tests: 3 in the
final failure-driven run, where the issue allowed another reading or the
hidden test required details the issue did not state
(§C),
and 2 in the targeted pass's development checks, one on another valid reading
and one whose only evidence was a test that does not compile on the base
(verification honesty).
Why 5/6 and not 6/6?
The sixth task (Prometheus) was implemented so that its hidden acceptance
test passed. BoundedCode still classified it UNVERIFIED: its
evidence checker did not associate the modified data-driven test file
(promql/testdata/functions.test) with the Go test function that reads it.
The official score stays 5 of 6; 6 of 6 hidden acceptance tests passed.
A verifier that withholds TASK_VERIFIED when it cannot show fail-before/
pass-after evidence is behaving as intended. After this validation, the
checker was changed to attribute changed test data to the Go package whose
tests read it (unreleased; covered by unit tests, not yet re-validated).
This is a small practical validation sample, not a statistically comprehensive evaluation.
The second set was screened for acceptance tests derivable from the issue;
the first set was screened for environment validity only, and three of its
tasks failed on identifiers that only the reference solution introduces. The
first set also had a "difficult" slot (a multi-file reference patch); the
second did not. The second validation ran with task.ambiguity: proceed, not
the default ask. On the development tasks, the candidate build's checks
before the freeze passed 0 of 2 (0 of 3 runs), per the
targeted engineering pass.
Therefore 0/8 → 5/6 does not measure system improvement; each result stands
on its own, with its own scope.
flowchart LR
U([Task request<br/>bcode chat or CLI]) --> CP[BoundedCode<br/>Go control plane]
CP <--> L[(Task ledger<br/>SQLite)]
CP --> C{Task contract:<br/>ambiguous?}
C -->|yes| Q([Asks you to clarify])
C -->|no| P[Context planner<br/>bounded pack]
CM[codebase-memory-mcp<br/>repository breadth] --- P
SE[Serena + LSP, optional<br/>semantic depth] --- P
P --> A[OpenHands agent<br/>network-less container]
A <-->|model calls over stdio| G[Model gateway<br/>in the control plane]
G <--> M[llama.cpp<br/>local model, default]
G <-.-> K[Cloud API, optional<br/>OpenAI · Anthropic · Gemini]
A --> W[Git worktree<br/>agent/task-id]
W --> V{Verification<br/>targeted, then full,<br/>in the sandbox}
V -->|failed: retry pack| P
V -->|budget exhausted| B([Blocked<br/>task resume])
V -->|passed| X{Affected cross-repo<br/>links compatible?<br/>experimental}
X -->|broken: retry pack| P
X -->|untested, after one request<br/>for a test that exercises it| R2
X -->|compatible or none| E{A test demonstrates<br/>the change?}
E -->|no: ask once for one| P
E -->|yes| R([task_verified<br/>branch ready for review])
E -->|still no| R2([tests_green<br/>UNVERIFIED, review first])
CP -.->|policy: Z1 design risk,<br/>Z2 repeated failures,<br/>Z3 pre-merge review| F[Frontier advisor<br/>Codex CLI or manual]
F -.->|advice in the next pack| P
Each attempt gets a bounded context pack; the agent has only terminal,
file-edit and task-tracker tools, and its model calls go back over stdio to
the gateway, so the container needs no network. Before verification the
control plane checks the worktree's integrity and commits a checkpoint. A
passing change that touches a cross-service contract without updating the
other side gets one more round to check it. In multi-repository tasks, the
affected gRPC, protobuf and OpenAPI links must also be shown compatible by
the repositories' own checks against each other's candidate commits
(experimental), and frontier escalation also
runs when you ask for it (Z4). A task that runs out of attempts, tokens or
time is blocked, not failed: task resume continues it.
| Layer | Role |
|---|---|
bcode chat / CLI |
Interactive chat and views (experimental) or plain commands, over the same operations |
| Go control plane | Orchestrates tasks, budgets, retries and escalation policy |
| codebase-memory-mcp | Repository breadth: code graph, impact, search |
| Serena / LSP (optional) | Semantic depth: definitions, references, implementations |
| Context planner | Builds small task-specific packs |
| OpenHands SDK | Agent runtime, in a sandboxed container |
| Model gateway | Carries the agent's model calls to the local model or a cloud API; metering, budgets, API keys |
| llama.cpp + local model, or a cloud API | Reasoning and editing (local by default) |
| Verification engine | Build, lint and tests in the sandbox, a secret scan of the diff on the host, and behavioural evidence |
| Task ledger | Persistent state and audit log; resume anywhere |
| Frontier gate | Optional escalation when the policy triggers or you ask; advice goes into the next pack |
BoundedCode is the control plane around existing tools. The model, the agent loop, code indexing and language servers come from upstream projects, used unmodified as separate processes or pinned dependencies (no forks, no vendored source).
Implemented in this repository (Go, plus a small Python adapter):
| Component | Where |
|---|---|
| Task orchestration: attempts, retries, budgets, resume after a crash | internal/orchestrator, internal/task |
| Task ledger and audit log (SQLite) | internal/store, internal/telemetry |
| Context planner and ranked retrieval seeds | internal/contextplan |
| Strategy governor (runaway control) and task contract (ambiguity handling) | internal/orchestrator, internal/task |
| Verification engine and behavioural-evidence gate | internal/verify |
| Sandbox setup, secret masking, command and path policy | internal/sandbox, internal/policy |
| Git worktree management and tamper checks | internal/gitops |
| Cross-service contract analysis (HTTP, OpenAPI, gRPC, protobuf, SQL, topics, env, Terraform) | internal/xservice |
| Model gateway (metering, tunnelled agent calls) and llama.cpp supervision | internal/inference |
| Frontier escalation policy, packet building and sanitization | internal/frontier |
| Process management for codebase-memory-mcp and Serena | internal/repointel |
| Terminal interface and chat; guided set-up; applying results to a checkout | internal/tui, internal/cli |
| OpenHands adapter: JSON-RPC bridge to the agent SDK | adapters/openhands/python |
CLI, doctor, benchmark harness |
internal/cli, internal/benchmark |
Integrated from upstream (pinned; licenses in THIRD_PARTY_NOTICES.md):
| Component | Role | How it is used |
|---|---|---|
| llama.cpp v0.5.0 | Local inference | External llama-server process |
| Local model (validated: Qwen3.6-35B-A3B) | Reasoning and editing | Weights downloaded by you; not redistributed |
| OpenHands Software Agent SDK 1.51.0 | Agent loop and tools | Python dependency of the adapter, inside the sandbox container |
| codebase-memory-mcp v0.11.0 | Code graph, impact, search | External binary |
| Serena v1.7.0 (optional) | LSP-backed symbol navigation | External MCP processes, one per task worktree |
Language servers (gopls, typescript-language-server) |
Used by Serena | Installed by you |
| gitleaks v8.30.1 | Secret scanning of task diffs | External binary |
| Docker (or Podman) | Container sandbox | Container engine |
| Codex CLI (optional) | Frontier escalation | External codex exec process, run in its own container |
| Go libraries: cobra, yaml, modernc.org/sqlite | CLI, config, embedded database | Go module dependencies |
Version pins and update policy: upstream components.
More detail: system architecture · product spec · implementation status · ADRs
Linux and macOS:
curl -fsSL https://raw.githubusercontent.com/akynte/boundedcode/main/scripts/install.sh | bash
cd ~/src/my-service # any git repository
bcodeWindows (PowerShell):
irm https://raw.githubusercontent.com/akynte/boundedcode/main/scripts/install.ps1 | iex
cd $HOME\src\my-service
bcodeinstall.sh puts boundedcode and its short name bcode in ~/.local/bin.
It uses a checksum-verified release binary when the release ships one, and
otherwise builds from source with Go. install.ps1 installs the
checksum-verified release binary into
%LOCALAPPDATA%\Programs\BoundedCode\bin and adds it to your user PATH.
Release binaries for macOS and Windows ship from the next release on; until
then, build from source there. bcode opens a chat for the repository
you are in. On the first run it checks the prerequisites and offers to install
what is missing: tools, llama.cpp, the model weights and the Docker sandbox.
It asks before every download or build. After that, describe a change and it
runs as a task. See the terminal interface.
You need git and a container engine (Docker, Podman, or Docker Desktop on
macOS and Windows). For local inference, set-up recommends a model that fits
this machine's memory and GPU (NVIDIA with CUDA, Apple Silicon with Metal,
or CPU only), and downloads a prebuilt llama.cpp where it does not build one
(it builds from source on Linux when a compiler and CMake are present).
Without a capable machine, choose a cloud model API instead. bcode setup --check shows what is missing from the shell. Linux is the validated
platform; macOS and Windows are experimental (see
Known limitations and the
platform table).
This is the Linux path with a CUDA build of llama.cpp, as on the reference
machine. On macOS and Windows, build the CLI with go build and let
bcode setup install the rest (it downloads the pinned prebuilt tools).
Prerequisites:
- Linux (x86-64 or arm64)
- Go (see
go.mod) - Git, ripgrep and Docker
- for GPU inference, an NVIDIA GPU and the CUDA toolkit, to build llama.cpp
uv, only for Serena
The getting-started guide lists tested versions.
git clone https://github.com/akynte/boundedcode.git
cd boundedcode
make build # ./bin/boundedcode
./scripts/install-deps.sh ~/.local/bin # gitleaks + codebase-memory-mcp (checksum-pinned)
./scripts/build-llama-cpp.sh # pinned llama.cpp v0.5.0 with CUDA
# Download a model (see "Models"), pinned to a commit and sha256-verified:
./bin/boundedcode model recommend # the model that suits this machine
./bin/boundedcode model fetch qwen3.6-35b-a3b # into ~/.local/share/boundedcode/models
L=~/.local/share/boundedcode/runtimes/llama.cpp/v0.5.0/bin
./bin/boundedcode init \
--llama-server $L/llama-server --llama-bench $L/llama-bench \
--adapter-dir $PWD/adapters/openhands/python
./bin/boundedcode sandbox build --dir adapters/openhands # agent sandbox image
./bin/boundedcode doctor # checks everything aboveRun your first task:
./bin/boundedcode workspace create demo
./bin/boundedcode workspace add ~/src/my-service
./bin/boundedcode index
./bin/boundedcode task create "Return 404 instead of 500 for unknown users" \
-c "go test ./... passes" --run
./bin/boundedcode task diff <id> # review the agent/<id> branch like a pull request
./bin/boundedcode task resume <id> # after Ctrl-C, a crash or a rebootPrefer a full-screen interface? ./bin/boundedcode tui covers all of the above
and the rest of the CLI: live task activity, diffs, verification, workspaces,
repository intelligence, the runtime, frontier escalations, stats and
doctor. See docs/usage/tui.md.
Tip
Already running an OpenAI-compatible server? Use
init --external-url http://127.0.0.1:8080 instead of the llama.cpp flags.
Run boundedcode doctor whenever something fails: it reports missing
dependencies with an install hint.
This exact path was tested from a clean clone with an empty home directory (record).
Model weights are not distributed with this project. You download them from their publisher and are responsible for complying with each model's license.
- Profiles: each file in
configs/models/pins the upstream source and revision, the file, the license and the llama.cpp settings. - Choice:
bcode model recommendrates every profile against this machine's RAM and GPU (a rule of thumb, not a measurement) and proposes one;bcode model listshows them all. Only the default is validated; the others are marked experimental until they are benchmarked here. - Download:
bcode model fetch NAMEdownloads at the pinned commit, resumes interrupted downloads, and checks the file against the sha256 in its profile. - Tooling:
bcode model use NAMEmakes a model the default;boundedcode bench infra --applytunes one for your machine.
| Model profiles | Status |
|---|---|
| Validated configuration | Qwen3.6-35B-A3B, UD-Q4_K_M (Apache-2.0) on llama.cpp v0.5.0 |
| Other profiles | Present, but not part of the validation |
Instead of a local model, the agent can use a cloud model API:
bcode provider use anthropic --model claude-opus-5-5 # or openai, gemini, openai-compatible
bcode provider key set anthropic # prompts without echo
bcode provider test # one short request- Keys are stored in the OS credential store (Secret Service, macOS Keychain, Windows Credential Manager), or an owner-only file where none exists, never in the configuration file. Only the host-side gateway uses them; the agent's sandbox has no network and never sees a key.
- Your code goes to the provider. With a cloud provider, everything the agent reads (context packs, file contents, command and test output) is sent to that provider, under its data policy. Secrets are masked as with a local model, but repository code is not. Use the local model for code that must not leave the machine.
- Cost is per token.
bcode statsshows usage per provider, and an estimate when you enter prices (--input-price,--output-price). - Status: experimental. The translations are tested against the providers' documented request and response shapes with fake servers, not on real tasks; the validation results above are for the local model only.
See ADR-0010 and the configuration reference.
BoundedCode runs fully local without any frontier account. Escalation is off by default. If you enable it:
local model first ──> escalation policy (Z1-Z4) ──> frontier only when the policy triggers
The policy triggers on repeated failures, rejected strategies, architectural risk or a high-risk review.
- Routes: the Codex CLI with a ChatGPT subscription sign-in, or a manual mode that writes the packet to disk for you to answer.
- No API keys for the frontier: escalation uses only the Codex subscription sign-in (the agent's cloud provider is a separate setting).
- Packets: sanitized (host paths, secrets), and each one needs approval unless pre-approved.
Escalation was enabled but not triggered in the second validation. In the first validation, 4 frontier calls were sent and no task was accepted (report). See ADR-0009.
| State | Meaning |
|---|---|
| builds | It compiles and lints. |
tests_green |
The repository's checks pass. Not enough on its own: in the initial validation, patches that changed nothing passed existing tests. |
TASK_VERIFIED |
Checks pass and there is behavioural evidence: a test the change adds or modifies (test code or test data) fails on the base commit with the changed tests, does not fail there without them, and passes with the change. In a multi-repository task, every gRPC, protobuf or OpenAPI link that the change affects must also be compatible (see below). |
- Missing evidence: the agent is asked once for a reproduction test.
Without one, the task ends
tests_green(UNVERIFIED) and is never presented as a verified merge candidate. - Cross-repository compatibility (experimental): in a multi-repository
task, each gRPC, protobuf or OpenAPI link that the change affects gets a
result:
compatible,brokenoruntested, tied to exact commits. The dependent repository's own checks run in the sandbox against the other repositories' candidate commits. Coverage, or a run with the OpenAPI operation removed, must show that the checks execute the link. A broken link fails verification and is retried. An untested one withholdsTASK_VERIFIED. The report is shown bytask status,verify --fulland the TUI's Verification tab. Supported for Go sides only (design and limits). - Gate integrity: the verification config is read from the base commit, so the agent cannot change which stages run. It can still edit tests and build scripts in its worktree; only review catches an adversarial change there (see the sandbox's residual risks).
A changed test that does not compile or load on the base (it calls code the change adds) shows that the API exists, not that it behaves as asked, so it is not evidence; nor is a stage that times out on the base.
This is not formal verification. Known limits:
- Go failures are compared per test function; other languages per stage (the stage must fail with the changed tests and pass without them);
- a change that only adds new API needs a test that also runs on the
original code, or it ends
tests_green; - a test can only demonstrate the reading of a request that the agent chose.
The agent is treated as potentially wrong or adversarially steered, for example by prompt injection in repository content.
| Control | Mechanism |
|---|---|
| Sandbox | Agent tools run in a container with no network. Only the task worktree is writable; there is no host home, SSH agent or credentials. |
| Secrets | .env*, keys, cloud credentials and kubeconfigs are masked in the sandbox and denied by path policy. |
| Protected paths | Changes to .boundedcode/, CI workflows or CODEOWNERS fail verification. |
| Git integrity | Worktree pointers, admin dirs and HEAD are verified before host git touches them. Nothing is pushed. |
| Command policy | A deterministic policy blocks push, destructive and deploy commands in verification. |
| Host reads | Context building never follows symlinks out of a worktree. |
| Frontier | Packets are sanitized, and the gate fails closed. |
Warning
A container is not a perfect boundary. Running an autonomous coding agent on untrusted repositories still carries risk. Never run BoundedCode where production credentials are reachable. See SECURITY.md and the sandbox design and residual risks.
- Small validation sample. 8 + 6 public tasks, each run once. The second set was screened for issue-derivable tests; real requests are not.
- Ambiguity detection is imperfect. The task contract is derived by the
local model. It has flagged a clear request as ambiguous and misnamed real
alternatives. The second validation ran with
task.ambiguity: proceed, not the defaultask; under the default, one clear task (vue) would have stopped on a false ambiguity flag. Since then, a material ambiguity is checked against the request text before it can stop a task: a second local-model call quotes the request, and the ambiguity is dropped when a quote that settles it occurs in the request, or when fewer than two of its readings have a supporting quote there. This check, and the retry of a contract that names nothing required, have unit tests with scripted model replies only; they have not been run on real tasks. - New models, cloud providers and platforms are unmeasured. Only Qwen3.6-35B-A3B on the reference machine (Linux) is benchmarked. The other model profiles and the cloud providers work but their quality on BoundedCode tasks is unknown, and model fit is a rule of thumb. macOS and Windows build and vet in CI, but their unit tests do not pass there yet (at a8fce66: 5 of 31 test packages fail on macOS, 14 of 31 on Windows), and the full flow has not been run on a Mac or a Windows machine.
- Recent evidence-check changes are not yet validated. Data-driven test files are now attributed to the Go package that reads them, and tests that only fail to compile on the base no longer count. Both changes are unreleased and covered by unit tests only; non-Go stages are compared per stage, not per test.
- Frontier escalation is unproven. It was enabled but not triggered in the second validation; in the first, 4 frontier calls were sent and no task was accepted.
- The strategy governor bounded runaway generation in development runs, but did not trigger during the held-out validation.
- Cross-repository compatibility covers a narrow set of cases. The gate
checks gRPC/protobuf sides written in Go (single module at the
repository root, a
go teststage, generated code committed in a task repository). It checks OpenAPI sides whose tests read the specification. Every other shape is reporteduntested, which withholdsTASK_VERIFIED. It never runs a client against the real server. It is tested on fixtures only, not on real tasks. - No baseline advantage shown. On the two-task baseline in the first validation, BoundedCode did not improve the same local model's result and was slower on those tasks; no baseline was run on the held-out set (baseline comparison).
Smaller limitations
- Terminal only (the CLI and the full-screen
bcodeinterface); there is no daemon or GUI. - Verification runs offline, so a project's dependencies must already be
installed: in the checkout (
node_modules,.venv,vendor) or in this machine's package caches (Go, Cargo, Maven, Gradle). Built-in presets cover Go, JavaScript/TypeScript, Python, Rust, Java (Maven, Gradle), C/C++ (CMake, Meson, Autotools, Make), Ruby and PHP without configuration; other languages run the Makefile'stest/checktarget, or say that no test runner was found (reference).
I designed the architecture, the threat model, the verification model and the evaluation protocol, and made the release and scope decisions. Implementation, test runs and first drafts of the reports were produced with heavy use of AI coding agents, under that design and review. The validation reports, errata and failures are published with their results unedited, including results that did not support a release.
Tested configuration, not a minimum requirement:
| Component | Tested configuration |
|---|---|
| Machine | Lenovo LOQ 15IRH8 laptop |
| CPU | Intel Core i7-13620H |
| GPU | NVIDIA RTX 4060 Laptop, 8 GB VRAM |
| RAM | 64 GB DDR5 |
| OS | Debian 13 |
| Model | Qwen3.6-35B-A3B, UD-Q4_K_M, 131 K context, MoE experts partly on CPU |
During validation the model server used up to 29.3 GiB of RAM, and BoundedCode itself under 70 MiB. No minimum requirement has been measured. Smaller machines may work with smaller models or contexts, but this is untested.
Earlier development benchmarks (decode speed, an 11-task synthetic suite, kill/resume and cross-service ablations) are in benchmarks/reports/ and the model evaluation.
Bug reports, model compatibility reports and pull requests are welcome.
- CONTRIBUTING.md: development setup, tests, and the
DCO sign-off on every commit (
git commit -s). - SECURITY.md: report vulnerabilities privately, not in issues.
- CODE_OF_CONDUCT.md and GOVERNANCE.md.
Licensed under Apache-2.0 (see also NOTICE). Third-party components and their licenses are listed in THIRD_PARTY_NOTICES.md and the upstream license matrix.
BoundedCode is independent and is not affiliated with, sponsored by, or endorsed by the upstream projects or vendors it integrates with.