Skip to content

About

Reproducible LLM token-optimization stack: Headroom proxy + RAG funnel + quality A/B proof

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

headroom-rag-stack

A reproducible LLM token-optimization stack for coding agents (Claude Code, Codex, opencode, aider, …).

Two engines, two sources of context bloat:

Engine Cuts Typical savings
Headroom proxy operational output the agent reads mid-session — tool results, file reads, logs, web fetches, history ~88% (logs/JSON), 68% high-entropy
RAG funnel which knowledge gets fetched from a large note/doc collection 80–300x fewer tokens than reading everything

Quality is preserved — compression protects the active turn, your prompts, system, and code; originals are recoverable (CCR). Includes a self-contained A/B test that proves answers survive.


Quick start (5 min)

git clone https://github.com/Selrach84/headroom-rag-stack
cd headroom-rag-stack
./install.sh          # installs headroom, starts a persistent proxy, prints routing steps

Then route your agent through it (the installer prints this), e.g. Claude Code:

headroom init --global --port 8788 claude   # writes ANTHROPIC_BASE_URL + MCP tool
# restart Claude Code

Prove it works (no API key, no vault needed)

python3 scripts/ab_quality.py

Expected — relevance-pruning keeps every answer at ~half the tokens; naive truncation at the same budget loses ~half:

method          recall    avg tokens
FULL            8/8 100%       9,905
HEADROOM        8/8 100%       ~5,000
TRUNCATE        4/8  50%       ~5,000   (same budget)

Try the RAG funnel on any folder of .md files

python3 scripts/rag_headroom.py "your question" ./sample_docs --k 4 --report

Self-contained: BM25 retrieval (built in) picks top-K docs, then Headroom query-relevance-prunes within each. No external RAG service required.


What's in here

install.sh                     one-shot installer (macOS + Linux)
scripts/ab_quality.py          self-contained quality A/B (deterministic proof)
scripts/rag_headroom.py        portable RAG + Headroom funnel over any .md folder
proxy/com.headroom.proxy.plist.template   macOS launchd (persistent proxy)
proxy/headroom.service.template           Linux systemd user service
sample_docs/                   tiny corpus so the funnel runs out of the box
ARCHITECTURE.md                tables, flow, design

Requirements

  • Python 3.10+ (3.13/3.14 work — installer handles the PyO3 build flag)
  • An agent that honors ANTHROPIC_BASE_URL (or OPENAI_BASE_URL)
  • macOS or Linux, one free TCP port (default 8788)

Traps (already solved for you)

Trap Fix
Build fails on Python 3.13+ PYO3_USE_ABI3_FORWARD_COMPATIBILITY=1 pip install … (installer does this)
Proxy "saves 0%" last 4 messages are protected + tool-result compression is opt-in → run with --intercept-tool-results; only aged outputs compress
headroom mcp status shows wrong port it checks 8787 by default — cosmetic; your proxy is on --port
Savings vary content-dependent: repetitive ~88–97%, high-entropy ~68%, 98% = best case
Quality worry active turn/prompts/system/code never compressed; headroom_retrieve restores originals

Credits

Built on Headroom (headroom-ai). This repo is the glue, persistence, routing, RAG funnel, and reproducible proof.

License

MIT

About

Reproducible LLM token-optimization stack: Headroom proxy + RAG funnel + quality A/B proof

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages