Skip to content

Repository files navigation

License Claude Code Codex

devloop

A test-driven development loop for Claude Code and Codex: specify, plan, implement, with independent review agents that catch design bugs before coding and correctness bugs before commit.

/spec          Clarify WHAT to build (user stories, acceptance criteria)
    ↓            spec validation gate (8-point checklist)
/plan          Plan HOW to build it (chunks, dependencies, tracker)
    ↓
review-plan    Independent agent validates the plan (8-point review)
    ↓
/implement     Build it with TDD (failing test → code → pass), per chunk
    ↓
review-impl    Independent agent verifies implementation matches plan
   +
red-team       Independent agent hunts bugs + cleanups in the diff
   │
   └─ converge  If a review finding invalidates the *plan* (not just the
      (back-edge) code), /implement appends corrective chunks, re-gates
                 via the review-plan agent in-session, then resumes

Three commands, three gates, one per command: /spec (WHAT), /plan (HOW), /implement (BUILD). The review agents run in a context that did not write the plan or code, so their verdict is independent, not the author grading itself.

Why not just prompt the AI?

You can tell a coding agent to "build feature X" in one prompt, and it will often return working code. What a single prompt cannot give you is everything around the code:

  • One pass conflates what, how, and build. A misread requirement or a wrong abstraction gets decided and written as code in the same breath, where it is most expensive to unwind.
  • The author grades its own work. "I verified it works," from the context that just wrote the code, is a self-report, not a review, it is blind to its own assumptions.
  • Context is fragile. A long session degrades and a resumed one starts cold, so the plan and the reasoning behind it evaporate.
  • Ungated agents report optimistically. With no hard gate, "done" tends to mean "I stopped," not "a test proves it."
  • Prompts do not compose or persist. They live in your head and drift from run to run and project to project.

devloop closes these gaps: it separates WHAT / HOW / BUILD into three gated steps, runs review in a context that did not write the code, persists a tracker so work survives resets, makes every gate an evidence-backed hard-stop, and ships as a versioned artifact that behaves the same across projects and harnesses.

What the review layer adds

  • Design bugs are caught before coding starts. An 8-point plan review runs before any code is written. Fixing a wrong abstraction in a plan costs minutes; in code, hours.
  • Two complementary reviews at the end, not one. review-impl checks conformance, does the code match the plan, and does a test prove each acceptance criterion (CONFIRMED / PLAUSIBLE / REFUTED, each with a quoted line)? red-team checks correctness, is the diff wrong or wasteful, regardless of the plan? The two are scoped to not overlap: review-impl defers correctness, robustness, standards, and cleanup to red-team. A bug that faithfully implements a flawed plan is caught only by red-team; a correct-but-off-spec change only by review-impl. On a large, multi-file diff the red-team half fans into a focused bugs pass and a focused cleanup pass in parallel; on a small diff it stays a single both pass (see below).
  • The bug hunter is recall-biased, then verified. red-team surfaces candidate defects freely (conservative reviewers under-report), then runs a verify pass that keeps only CONFIRMED/PLAUSIBLE findings and drops the rest, trading a noisier find phase for higher recall without shipping the false positives.

How the loop works

The loop is not strictly linear. If an /implement review finding invalidates the plan itself (a chunk's whole approach is wrong, or the review surfaces work no chunk covers), /implement converges: it appends corrective chunks to the tracker and re-runs the plan-review gate by spawning the review-plan agent in-session (not by re-invoking /plan, which would regenerate the tracker), then resumes building. That keeps the plan and the code from drifting apart.

Skills vs agents. The user-facing, interactive steps are skills (/spec, /plan, /implement); the isolated steps that return a verdict are agents (review-plan, review-impl, red-team). The split is by the nature of the work, not by a harness quirk, so it ports across harnesses. It also means the loop is three user invocations, one per command; a skill cannot call another skill, and coupling them into one agent would tie the loop to a single harness.

The gates depend on spawning those agents. On a harness with no subagent capability at all, each gate degrades to an in-context self-check and loses the author-evaluator isolation; the skills say so at each gate rather than pretending the guarantee still holds.

What's Included

Piece Type Invocation Purpose
spec Skill /spec <feature> User stories, acceptance criteria, edge cases
plan Skill /plan <feature> Chunk decomposition, dependency graph, JSON tracker, plan-review gate
implement Skill /implement <feature> TDD chunk cycle against the tracker, 8-point quality gate
review-plan Agent auto (/plan; /implement convergence) 8-point plan review in fresh context
review-impl Agent auto (/implement Phase 3) Verifies implementation matches plan
red-team Agent auto (/implement Phase 3) / manual Adversarial diff review, bugs + cleanup

Reviewing arbitrary changes. review-impl is a conformance gate: it checks code against a plan, so it needs a tracker to review against. To review an ad-hoc diff with no plan (a hotfix, someone else's branch), invoke red-team directly, it is plan-agnostic and discovers the project's standards at runtime.

The red-team agent

red-team is the plugin's bug-and-cleanup reviewer. It reads the diff in a fresh context and runs a two-family review:

  • Correctness (5 angles): line-by-line diff scan, removed-behavior auditor, cross-file caller/callee tracer, language-pitfall specialist, wrapper/proxy correctness.
  • Cleanup (4 angles): reuse, simplification, efficiency, altitude.
  • Conventions (runs in every mode): reads the project's own rules file (CLAUDE.md/AGENTS.md) and .devloop/config.md at runtime and flags only rules it can quote. A quotable violation with a real fault (security, data loss, a broken documented invariant, a robustness gap) FAILs the gate like a correctness bug; a pure style violation stays a warning.

It then verifies each candidate (recall-biased: PLAUSIBLE by default, REFUTED only when the code proves it) and sweeps once more for gaps the first pass missed.

It takes a mode:

Mode What it does
bugs the 5 correctness angles + conventions, then verify + sweep
cleanup the 4 cleanup angles + conventions, the tidy pass; can apply fixes (report-only when run as a gate)
both (default) everything
Use the red-team agent in mode: both to review the changed files
Use the red-team agent in mode: cleanup to tidy the changed files

Size-adaptive at the /implement gate. Phase 3 sizes the diff the same way /plan sizes work: a single-file change (or a trivial one with no new logic) runs one red-team in mode: both; a multi-file or cross-cutting diff splits the red-team half into parallel mode: bugs and mode: cleanup runs so neither family crowds the other out. At the gate the cleanup run is invoked report-only, so the whole review stays read-only and safe to run alongside review-impl.

Install

Claude Code. Install from the marketplace:

/plugin marketplace add KashZod/devloop
/plugin install devloop@kashzod

Claude Code auto-discovers the three skills (skills/) and three agents (agents/); the manifest is .claude-plugin/plugin.json. To vendor the plugin instead, copy skills/ and agents/ into your project's .claude/ directory.

Codex. Install from the same marketplace:

codex plugin marketplace add KashZod/devloop
codex plugin add devloop@kashzod

Codex reads .codex-plugin/plugin.json and its own catalog (.agents/plugins/marketplace.json). The same three skills power both harnesses; Codex registers skills, not agents, so the three review agents under agents/ ride along as prompt files: the skills spawn them as subagents where the harness supports delegation, and otherwise degrade to the in-context self-check the gates already document.

Other Agent Plugins clients. The repo root also carries a plugin.json in the vendor-neutral Agent Plugins 1.0.0 format, validated against the published schema, so the three skills are structured to install on clients that adopt the standard (Kiro, for one, has announced support). Clients that support subagent delegation get the isolated-context review; the rest degrade to the self-check, the same as Codex. Claude Code and Codex read their own manifests and ignore this one.

Then give the loop project context in a .devloop/ directory at your project root:

.devloop/
  config.md    # engineering: build/test/lint, architecture, standards,
               #   blindspots, commit conventions, and the spec/tracker
               #   directory settings. Read by /plan, /implement,
               #   review-plan, review-impl, red-team, and /spec (which
               #   reads the spec-directory setting from here).
  domain.md    # domain context, architecture overview, domain-specific
               #   concerns. Read by /spec.
  trackers/    # impl-tracker-<feature>.json, written by /plan

Each skill and agent resolves its config as .devloop/<file> in your project, else generic mode (the loop still runs, with less project-specific insight). This is why a plugin install works: the skills live in a read-only cache, but they read .devloop/ from your project, not the cache. There is no copied-in fallback; .devloop/ is the only project-config source.

The fastest start is to copy the closest examples/<stack>/ directory to .devloop/ in your project: each holds a config.md and a domain.md for one stack (typescript-node, python, rust, android-kotlin). Each is a concrete example for that stack; copy the closest to .devloop/ and adapt it to your project.

Commit config.md and domain.md so the whole team shares one context. .devloop/trackers/ holds in-progress work; commit it for cross-machine resumability or gitignore it, your call.

Optional: proof that the check actually ran (rung)

devloop's reviews are evidence-backed (review-impl quotes the test line that proves each acceptance criterion; red-team verifies each finding before reporting), but the green step still trusts that the agent ran the tests it says passed. rung closes that last gap: it records whether a check drove the real surface (not just an isolated test or a reading of the code) and whether an independent context ran it, then gates on that record deterministically.

rung's author-vs-independent split is the same one devloop's review agents enforce, and its check-level record hardens the TDD green step into a re-checkable artifact instead of a self-report.

This is optional and unbundled by design. devloop ships only markdown and a bash validator, with no runtime dependencies; rung is a separate pip install rung-ai CLI (or GitHub Action). To use it, wrap the regression run in rung run --rung 1 and gate CI on rung gate (exit 0 is the only pass). Keep it out of the core loop unless you want CI-enforceable proof that the checks were real.

Related Work

devloop sits inside the conversation about an AI-native software development lifecycle. Anthropic's AI-Native SDLC playbook reframes the classic six stages (Plan, Design, Build, Test, Deploy, Maintain) as a loop in which each stage commits a version-controlled artifact the next stage reads, and governance is enforced as the agent works rather than discovered in a late review. devloop is a concrete implementation of that loop's inner stages: /spec commits a spec, /plan commits a reviewed JSON tracker, and /implement builds test-first behind review agents that grade the work in a context that did not write it. The committed-artifact chain and the fresh-context reviewer are devloop's spine, not extras, so the playbook describes a shape devloop already has.

What devloop deliberately does not own: deploy gates, production monitoring, incident-to-eval loops, and harness-level enforcement hooks. Those are real parts of an AI-native SDLC, better served by the harness (managed settings, hooks) or your CI. devloop owns the specify-plan-build-review core and stops there on purpose.

GitHub Spec Kit is a spec-driven development toolkit in the same space, with a comparable specify -> plan -> tasks -> implement flow across several coding agents. devloop's /spec -> /plan -> /implement shape covers similar ground; where it differs is the review layer: three independent agents (review-plan, review-impl, red-team) run in author-isolated context as hard gates between the phases, and the loop is TDD-first with a JSON tracker that survives context resets. For the broader spec-driven tooling ecosystem, start with Spec Kit; for the gated, review-heavy loop, devloop is narrower by design.

License

Apache-2.0. See LICENSE.

About

A test-driven development loop for agentic coding tools such as Claude Code and Codex: /spec, /plan, /implement, with independent review agents that gate the plan before you code and the diff before you commit.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages