Run AI coding agents unattended without letting them go loose: one implements, another reviews, a third tests, a guard blocks what they must never touch, and you only see what needs you.
Zero runtime dependencies · tests run against real git repositories · every claim marked enforced, checked or instructed
Quick start · How it works · Harnesses · Lean mode · FAQ
Give an agent a task and walk away. Agent Flow runs it through separate implement, review and QA steps in its own git worktree, blocks the edits an agent must never make, lets it fix its own failing tests and CI lists, and leaves you a short list of what needs a person.
It is a Node.js CLI and skills package for Claude Code, Codex CLI, Gemini CLI, Cursor, Copilot, Windsurf, Pi and other tools that read AGENTS.md (14 install targets plus Pi; see the harness matrix for what each one enforces). It has no runtime dependencies and needs no package.json in the repository it sets up. Node.js 20 or later is required.
| ❌ Without Agent Flow | ✅ With Agent Flow | |
|---|---|---|
| Who checks the work | The agent that wrote it | A second agent reviews, a third runs the checks |
| The new test isn't in the CI list | It edits the workflow or loosens the check to get green | It adds one line; every other CI change is refused |
Push to main, force-push, --no-verify |
Nothing stops it | Refused by the guard (Claude Code, Pi) or caught by the pre-commit hook |
| A check fails | You find out when you're back | It goes back to the Implementer, up to a round cap |
| When you're back | A whole diff to read | A draft pull request and a short list of what needs you |
| Docs and context | AGENTS.md describes the old behaviour |
agent-flow doctor flags it |
| What happened | Ask the agent | A hash-chained audit log |
Install the harness files, then ask your coding agent to bootstrap the repository:
npx @drix10/agent-flow install --harness claudeUse codex, gemini, cursor, copilot, windsurf or agents instead of claude for those tools. Pi users can install the package with pi install npm:@drix10/agent-flow.
In repositories without package.json, install vendors the small runtime needed by local hooks into .agent-flow-runtime/. Commit that directory so the hooks work for other clones; Agent Flow adds no package files or dependencies to the project.
To have your coding agent do the whole setup for you, give it the prompt in SETUP.md.
Bootstrap scans the code read-only and proposes AGENTS.md, CONTEXT_MANIFEST.json and related context files. It asks about protected paths and risk boundaries before writing. Review every proposal before accepting it.
For Claude Code, run a task with observable acceptance criteria:
npx @drix10/agent-flow run "Add a --json flag so CI can parse the output"The command runs in a worktree and leaves checked work local by default. Add --dry-run to see setup gaps without launching roles. Add --pr to push the branch and open a pull request. Add --auto-merge only when you want eligible low-risk PRs queued for merge and the repository also sets pipeline.auto_merge_low_risk: true.
On other harnesses, ask the agent to use the invoking-agents skill for an issue or task. It applies the same state, risk, review and QA checks using that harness's launch procedure. Pi also provides /implement <issue>.
npx @drix10/agent-flow@latest update --yes
npx @drix10/agent-flow doctor
npx @drix10/agent-flow audit-risk --fail-on-newupdate refreshes unedited skills, hooks and the vendored runtime; it preserves files you customized. doctor checks context paths, commands, links and manifest freshness. audit-risk compares dependencies, secrets and other risk surfaces with .risk-baseline.json. See Adoption for update behavior by install type.
The CLI works in Python, Go, C++, Rust and other repositories without adding package.json or node_modules. Node.js is needed in the CI job to run it.
name: agent-context
on: [push, pull_request]
jobs:
context:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: 22 }
- run: npx @drix10/agent-flow doctor
# Create and review .risk-baseline.json before enabling this gate.
- run: npx @drix10/agent-flow audit-risk --fail-on-newOr use the composite GitHub Action, which also classifies pull request changes and can upload SARIF:
- uses: actions/checkout@v4
with: { fetch-depth: 0 }
- uses: Drix10/agent-flow@v1.2.3| Area | Behavior | Command or enforcement |
|---|---|---|
| Context drift | Checks referenced paths, commands, relative links, cited commits and manifest timestamps. | agent-flow doctor |
| Setup | Scans the repository and proposes context files with confidence markers. | bootstrap skill; agent-flow scan and init |
| Risk | Classifies the actual diff using protected paths and risk boundaries. | agent-flow classify |
| Pipeline | Runs implementation, mechanical classification, review, required gates and QA with a round cap. | agent-flow run on Claude Code; invoking-agents skill elsewhere |
| Risk changes | Compares new dependencies, credentials, payment/auth code and other risk signals with a reviewed baseline. | agent-flow audit-risk --fail-on-new |
| Audit trail | Records state changes, guard blocks, gate results and role runs in a hash-chained log. | agent-flow audit verify and summary |
| Deferred shortcuts | Reads back every lean: comment (a ceiling and when to upgrade) and flags the ones that name no trigger. |
agent-flow debt; Lean |
| Token cost | Short skill descriptions, an orchestrator skill that a CLI run never loads, terse role output, and a warning when a context file is heavy. | agent-flow doctor; what changed |
| Other tools | A read-only MCP server and a status badge, so a host without our hook can still ask. | agent-flow mcp, agent-flow statusline |
Useful commands:
npx @drix10/agent-flow status # protection, checks, context and pending work
npx @drix10/agent-flow doctor # read-only context report
npx @drix10/agent-flow scan # read-only repository reconnaissance
npx @drix10/agent-flow hook install # install the pre-commit gate
npx @drix10/agent-flow debt # the shortcuts agents left, and which have no trigger to revisit them
npx @drix10/agent-flow uninstall # preview taking back out what install wrote (--yes to do it)
npx @drix10/agent-flow state dismiss --issue 100000 --reason "did it by hand" # drop an issue from the listtask -> worktree -> implement -> classify -> review -> gates -> QA -> optional PR
^ |
+--- bounded fix rounds --+
Each role is a separate process and receives files from .agent-flow/artifacts/issue-N/, not another role's reasoning. Issue text is escaped and treated as untrusted input. The state machine enforces transitions and the review-round cap. Repeated failures stop for a human. When you request a PR, critical changes create a draft and wait for human review. Local is the default; publishing needs --pr.
pipeline.lean asks the Implementer and Reviewer for the smallest correct change and marks the shortcuts it takes (Lean).
Protected paths stop a run; review-only paths (review_paths, usually .github/workflows/) don't. An agent may add a line there, such as a new test in the CI list, and the pull request waits as a draft for you. See Adoption.
The loop is not equally enforced on every harness. The guard blocks tool calls directly on Claude Code and Pi. Other harnesses rely on their available sandbox and the pre-commit hook, which runs at commit time and can be bypassed by a human. Reviewer write restrictions vary by harness. See the harness matrix before choosing a role setup.
Agent Flow uses AGENTS.md for shared agent instructions. Claude Code can load it through CLAUDE.md containing @AGENTS.md. Module-level AGENTS.md files add rules for that part of the repository. CONTEXT_MANIFEST.json records referenced paths, protected paths, risk boundaries and pipeline settings. The gardener skill checks and repairs this context; it re-reads the code before refreshing timestamps.
When an agent repeats a mistake, prefer a mechanical fix: an invariant in code, a test, a guard rule or a protected path. Prompts and prose are the least reliable enforcement.
- Risk classification and secret detection are heuristic; review their findings.
- Shell-write analysis is best-effort. Use a sandbox or remove shell access from read-only roles when a hard boundary is required.
- Per-tool enforcement is available on Claude Code and Pi. Other harnesses have different enforcement limits.
- Hooks do not access the network.
doctormay make optional, cached read-only lookups for package updates and GitHub Code Owner settings; CI and offline modes skip them. - Each issue gets its own git worktree. Measured on Windows with a synthetic 60,000-file repository: creating it took 16 seconds (git's checkout), Agent Flow's own commands each finished in about half a second or less, and the before-and-after check around QA added about 3 seconds in total. Very large monorepos should use git sparse-checkout or a partial clone; Agent Flow does not automate that.
- Critical changes are reviewed by a stronger model only if you set
pipeline.models.high_reasoning(andpipeline.models.fastfor the rest). Unset, every role uses the default model, andstatusandrunsay so. - Audit logs are tamper-evident, not tamper-proof. Keep an audit hash outside the repository if you need an external trust anchor.
See Security and known failure modes for details. Agent Flow works one repository at a time and does not provide model-provider failover.
Will agents get stuck waiting for me? Only on protected paths and critical changes. Everything else, including CI list additions, carries on and reaches you as a draft.
What if my harness has no hooks? The skills and the pre-commit gate still apply; the guard's per-call blocking does not. The harness matrix says which is which.
How do I remove it? agent-flow uninstall previews, --yes applies. Files you edited are kept.
npm install
npm testThe test suite exercises real temporary Git repositories and supported installation targets. See Contributing.
MIT