Legend: implemented (tested, usable), experimental (works, but has limited validation or unstable interfaces), planned.
| Area | Status | Evidence / notes |
|---|---|---|
| Repo bootstrap, licensing, DCO, CI | implemented | make check; license matrix verified 2026-10-03 |
| Typed config + model profiles | implemented | strict YAML, validation tests |
| SQLite state + redacted audit log | implemented | internal/store, internal/telemetry |
doctor, init |
implemented | |
bcode chat, setup, one-command install |
experimental | The chat maps messages to new, clarified and follow-up tasks, with queueing and interrupt (model tests). setup steps (backend tests, plus a pty first run with declined downloads). task apply and follow-ups from a branch (backend tests with real git). The installer was tested on its build-from-source path; no release binary is published yet |
Terminal interface (tui) |
experimental | internal/tui. It covers every CLI command: structured views read the ledger and audit log, and actions run the CLI commands in-process, so policy and audit are unchanged. Model-level tests drive keys and forms against a fake backend at 80×24 to 200×60. Backend tests drive the real CLI (init, workspace, create, cancel, cleanup). The real binary was smoke-tested in a pty |
| llama.cpp supervision (start/reuse/stop, health, OOM classification, external server, idle sleep) | implemented | internal/inference/llamacpp; idle sleep measured 2026-10-04 (RSS 19.3 → 0.8 GiB, VRAM 7.1 → 0.2 GiB, wake 1.7 s) |
| Infrastructure benchmark (feasibility search, prompt/decode, cache reuse, thermals) | implemented | bench infra; reports in benchmarks/reports/ |
| OpenHands adapter (JSON-RPC/stdio, LLM tunnel, resume, condense, protocol v1 handshake) | implemented | real SDK through the adapter, scripted LLM, host and container tests; malformed/partial/crash/deadline/fuzz tests on the protocol |
| Agent runs against the local model end to end | implemented | milestone 1 (kill/resume, condensation, local-only); 11-task suite on 3 models |
| Workspaces (add/remove/disable, per-workspace config) + codebase-memory-mcp (persistent MCP session, version gate, CLI fallback, per-worktree impact) | implemented | Phase 3 gap report; 5-repo lifecycle test; detect_changes limitation and workaround recorded 2026-10-04 |
| Cross-service contract analysis (HTTP, topics, env, Terraform) | implemented | docs/design/cross-service-analysis.md: 7/7 fixture links, idiom tests, fuzzing, contract check in the run loop |
| Serena v1.7.0 symbol navigation (optional, MIT-pinned) | implemented | ADR-0008: per-worktree MCP instances, read-only, graph-for-breadth routing with per-symbol fallback; fake and real-Serena tests (Go, TS, worktree freshness, malicious repo config, multi-repo, crash, cancellation, resume); bench intel overlap study; Stage 2 edit experiment; CI pin guard. Optional, off by default: A/B 9/9 vs 9/9, efficiency trend only (benchmarks/reports/*-ab-serena-summary.md) |
| Task ledger + context planner (resume packs) | implemented | run lease and crash recovery (TestSIGKILLResume, real SIGKILL), changed symbols, step progress, confined and redacted packs, lexical (ripgrep) stage |
| Verification engine (task_verified vs tests_green via behavioural evidence, Go build-tag variants, Go/TS/Terraform-fmt/Helm presets, impact-selected Go packages, diff scope, protected paths, gitleaks) | implemented | config and presets read from the base commit (TestAgentCannotWeakenVerification); JS/TS stages use the checkout's installed node_modules (mounted read-only) and fail when a declared script has no installed dependencies; Terraform/Helm stages are optional and skipped when the tool is absent from the sandbox image |
| Runaway control (reasoning budget, per-response output cap, progress-aware strategy budget) | implemented | governor_test.go; calibrated by replaying archived runs (docs/benchmarks/targeted-engineering-pass-2026-10.md); no live stop observed yet |
Task contract and ambiguity handling (task.ambiguity: ask blocks for task run --clarify; proceed records SPEC_AMBIGUOUS) |
experimental | model-derived; in the second validation 2/6 derivations were empty and 1/6 flagged a false material ambiguity |
| Ranked retrieval seeds (diagnostics first, prose names, config keys, samples last; common names demoted) | implemented | seeds_test.go; replay on the development issues |
| Sandbox + policy (container, masks, read-only git, host git hardening) | implemented | adversarial tests incl. admin-dir tampering, planted symlinks, prompt-injected agent, policy bypasses; residual risks in docs/design/sandbox.md |
| Frontier escalation Z1–Z4 (Codex subscription, manual) | implemented | live Codex run (Phase 7) fixed a task the local model failed; the contained provider is verified live (TestCodexContainedLive) |
| Engineering benchmark harness (14 tasks, hidden checks) | implemented | every task fails on its base (TestHiddenChecksFailOnBase, 14/14 on 2026-10-04); the model evaluation used the original 11 |
| Model evaluation (Qwen3.6, Laguna XS 2.1, Qwen3-Coder-Next) | implemented | docs/design/model-evaluation.md. Default: Qwen3.6 (ADR-0007, accepted). |
Success metrics from real task history (stats, frontier list) |
implemented | local-only rate, escalation rate, token share, verified tasks/hour; escalations record model and whether the code changed after the advice |
| Daemon, GUI | planned | |
| Release packaging, SBOM | implemented (unpublished) | make dist, scripts/sbom. Publishing needs maintainer approval. |
| Local escalation tier (switch to a stronger local model before frontier) | planned | motivated by the 4/4 vs 1/4 result in ADR-0007 |
An implementation audit against both specifications (the original project specification and the Serena integration) was done on 2026-10-04; see audit-2026-10-04.md.
A small real-world validation (8 public SWE-bench-style tasks) was run on
2026-10-04: 0/8 on the frozen build, 1/8 after fixing five defects it found;
see the report.
A failure-driven engineering pass on 2026-10-05 used those failures as a
development corpus (development evidence, not a validation): 1/4 development tasks pass on the
final build (prometheus, previously failing); see
the engineering report.
A targeted pass (runaway control, ambiguity, retrieval seeds) followed, then a
second validation on 6 tasks not used during development and never shown to
the agent, screened for issue-derivable acceptance tests: 5 of 6
TASK_VERIFIED and passing the datasets' hidden acceptance tests (hidden from
the agent), all 5 local-only: no frontier calls (escalation enabled, not
triggered), and 0 false verification passes among those 5 on the screened
set (development runs with the same gate design did have such passes); see
the targeted pass and
the validation report.