A reproducible LLM token-optimization stack for coding agents (Claude Code, Codex, opencode, aider, …).
Two engines, two sources of context bloat:
| Engine | Cuts | Typical savings |
|---|---|---|
| Headroom proxy | operational output the agent reads mid-session — tool results, file reads, logs, web fetches, history | ~88% (logs/JSON), 68% high-entropy |
| RAG funnel | which knowledge gets fetched from a large note/doc collection | 80–300x fewer tokens than reading everything |
Quality is preserved — compression protects the active turn, your prompts, system, and code; originals are recoverable (CCR). Includes a self-contained A/B test that proves answers survive.
git clone https://github.com/Selrach84/headroom-rag-stack
cd headroom-rag-stack
./install.sh # installs headroom, starts a persistent proxy, prints routing stepsThen route your agent through it (the installer prints this), e.g. Claude Code:
headroom init --global --port 8788 claude # writes ANTHROPIC_BASE_URL + MCP tool
# restart Claude Codepython3 scripts/ab_quality.pyExpected — relevance-pruning keeps every answer at ~half the tokens; naive truncation at the same budget loses ~half:
method recall avg tokens
FULL 8/8 100% 9,905
HEADROOM 8/8 100% ~5,000
TRUNCATE 4/8 50% ~5,000 (same budget)
python3 scripts/rag_headroom.py "your question" ./sample_docs --k 4 --reportSelf-contained: BM25 retrieval (built in) picks top-K docs, then Headroom query-relevance-prunes within each. No external RAG service required.
install.sh one-shot installer (macOS + Linux)
scripts/ab_quality.py self-contained quality A/B (deterministic proof)
scripts/rag_headroom.py portable RAG + Headroom funnel over any .md folder
proxy/com.headroom.proxy.plist.template macOS launchd (persistent proxy)
proxy/headroom.service.template Linux systemd user service
sample_docs/ tiny corpus so the funnel runs out of the box
ARCHITECTURE.md tables, flow, design
- Python 3.10+ (3.13/3.14 work — installer handles the PyO3 build flag)
- An agent that honors
ANTHROPIC_BASE_URL(orOPENAI_BASE_URL) - macOS or Linux, one free TCP port (default 8788)
| Trap | Fix |
|---|---|
| Build fails on Python 3.13+ | PYO3_USE_ABI3_FORWARD_COMPATIBILITY=1 pip install … (installer does this) |
| Proxy "saves 0%" | last 4 messages are protected + tool-result compression is opt-in → run with --intercept-tool-results; only aged outputs compress |
headroom mcp status shows wrong port |
it checks 8787 by default — cosmetic; your proxy is on --port |
| Savings vary | content-dependent: repetitive ~88–97%, high-entropy ~68%, 98% = best case |
| Quality worry | active turn/prompts/system/code never compressed; headroom_retrieve restores originals |
Built on Headroom (headroom-ai). This repo is the glue, persistence, routing, RAG funnel, and reproducible proof.
MIT