Build tools agents can discover quickly, use correctly, and recover with.
A tested set of language-neutral contracts with a small Rust reference. Rust is preferred for efficient native tools, not required. Choose the language that meets the product's measured needs. One executable, ordinary shell commands, structured results. No service to run alongside it.
Giving this link to a coding agent? Tell it:
Build this CLI using https://github.com/paperfoot/agent-cli-framework. Read AGENTS.md first. Rust is preferred, not required; use example/ as the behavioral reference. Implement only relevant patterns, measure speed and resource use, and run tests and conformance checks before handing it back.
Build instructions · Example · Command design · Research
An agent should be able to discover a command, execute it, and use its result. It should not need to read an entire manual before doing useful work.
greeter --help # Small command index
greeter agent-info --command hello # Just this command's contract
greeter hello Ada --style pirate # ExecuteWhen piped, the last command returns:
{"version":"1","status":"success","data":{"name":"Ada","style":"pirate","message":"Ahoy, Ada! Welcome aboard!"}}The same command in a terminal prints a readable greeting. --json makes the
machine format explicit. agent-info without a filter still returns the full
manifest, and agent-info --command "config show" handles nested commands.
Already know the command and its contract? Use it directly.
| Friction | Framework response |
|---|---|
| Reading every tool definition | Short help; discovery scoped to one command or group |
| Guessing flags and defaults | Generate syntax from Clap; test semantic metadata |
| Reading thousands of irrelevant results | Search, limits, cursors, and field selection where the domain needs them |
| Repeated lookups before an action | Return stable identifiers and useful next-step data |
| Retrying an uncertain write | Distinguish a local lock from provider idempotency and reconciliation |
| Waiting through dozens of calls | Bounded batches and resumable jobs when justified |
| Relearning after an update | Stable envelopes, explicit capabilities, compatible additions |
| Excess CPU and memory | Lazy initialization, bounded work, and resource measurements |
| Optimizing the wrong thing | Measure completed tasks, calls, bytes, and latency together |
These choices apply across model families. Model and harness changes belong in evaluations, not hardcoded model-specific instructions.
Every CLI has:
agent-info/info— raw JSON describing real commands, arguments, defaults, examples, and effects;--commandnarrows discovery without changing its shape.- Structured output — compact JSON when piped; readable output in a terminal;
failures on stderr;
--quietsuppresses human informational output only. - Semantic exit codes —
0success,1runtime/transient failure,2setup,3input,4rate limit. Help and version exit0. config show/config path— lazy configuration loading and masked secrets.skill install/skill status— a short embedded signpost to the binary.
Add doctor for external dependencies, update for distributed tools, and a
concurrency guard for expensive or irreversible operations. Destructive commands
require --confirm. No interactive prompts or implicit stdin reads.
The example demonstrates the core, scoped discovery, structured diagnostic failures, and a kernel-backed duplicate guard. The optional command patterns describe domain features to implement and test when needed; they are not extra commands in the greeter.
For a Rust CLI, copy the reference:
git clone https://github.com/paperfoot/agent-cli-framework.git
cp -R agent-cli-framework/example my-cli
cd my-cliRename the package, replace the REPLACE placeholders, and add your domain
commands. Keep the modules small:
src/
main.rs parse, dispatch, exit
cli.rs command grammar and help
config.rs defaults → file → environment
error.rs codes and recovery instructions
output.rs JSON and human output
guard.rs exclusion for concurrent operations
commands/ one module per domain
tests/ observable behavior and contracts
Then validate the built tool:
cargo fmt --check
cargo clippy --all-targets --all-features --locked -- -D warnings
cargo test --locked
cargo build --release --locked
../agent-cli-framework/conformance/conformance.sh ./target/release/my-cliReplace my-cli with your actual binary name. The probe uses Bash and jq.
The repository CI also validates emitted JSON against the published schemas.
Other languages use their native build/test tools and the same conformance and
schema checks. Keep the behavior; adapt the implementation.
| Working on | Read |
|---|---|
| A new CLI | AGENTS.md, then example/ |
| Discovery, output, errors, compatibility | Runtime contracts and schemas/ |
| Search, writes, batches, long jobs | Command design |
| Config, secrets, dependencies, concurrency | Implementation notes |
| Releases and package managers | Update standard |
| Measuring improvements across agents | Evaluation |
| Language choice, CPU, memory, startup | Performance |
| Updating an existing framework CLI | Migration notes |
| Why these choices | Research and source links |
SWE-agent studied how interface design affects agent behavior. More recently, a controlled CLI/MCP comparison found widely varying costs across scaffoldings on one software task. Neither establishes a universal winning interface or a speedup for this framework.
The practical lesson is to keep interfaces discoverable, responses relevant, and results verifiable. This repository chooses a local CLI as its interface; it does not require claims that every alternative is slower.
Use the included measurement tool for output bytes and process latency. Use paired task evaluations for claims about agent performance.
search-cli · autoresearch · xmaster · email-cli
Built by Boris Djordjevic at 199 Biotechnologies and Paperfoot AI.