An autonomous, white-hat bug-bounty hunting harness for LLM agents.
It is not a scanner and not a prompt. It is an operating system for an AI agent that hunts vulnerabilities the way a senior researcher does: build an independent model of the target, find where the code diverges from it, and drive one thread down five layers until a bug either dies with a falsifier or proves out with a PoC. The methodology is the product — the model underneath is an interchangeable part. It runs on Claude Code today and is designed to port to any agent runtime that can read files, run a shell, spawn sub-agents, and drive a browser.
AI does the typing. The architecture, the calls, and the mistakes are the operator's.
Point it at a target and it runs a self-driving hunt loop — select one hypothesis, drive it deep, prove or kill it, repeat — until it has a cold-verified High/Critical, banking Medium and Low findings along the way. It covers both crypto (EVM, Solana, TON, Move) and web2, and drives live dApp frontends through a real browser.
Three entry points:
| Command | Surface | What it is |
|---|---|---|
/deephunt <target> |
Smart contracts (EVM / Solana / TON) | Hypothesis-driven hunting for High/Critical — deep manual analysis, invariant testing, fork PoC. The main mode. |
/dapphunt <domain> |
Live web3 dApp frontends | Divergence-first frontend hunting — cross-clone differential, data-flow divergence, runtime observation. |
/hunt <target> |
web2 apps and APIs | Access-model-first hunting with an authorization-diff harness at its core. |
- Divergence-first. It builds a model of the system before reading the implementation, then hunts the gap between the two — the one place the crowd of duplicate-submitters isn't looking.
- Depth over breadth. It refuses to call a scope "clean" from a surface sweep. A finding only counts when one interaction chain has been pulled down five layers to its seam.
- Enforcement, not aspiration. The methodology is wired into runtime hooks — a ledger that creates itself before the first action, a loop that won't let the agent quit early, a prompt-injection guard over every piece of external content. The discipline runs whether or not the agent "remembers" to follow it.
- Compounding. Every hunt banks reusable invariants and patterns, so the next target starts with a library the crowd doesn't have.
- Cold verification. Nothing is reported until a fresh-context agent re-derives it and a separate specialist checks scope and payability. It kills roughly a quarter of candidates before they waste a submission.
This is a white-hat system. It thinks like an attacker in order to find and prove bugs — and then stops.
- THINK ≠ ACT. Impact is measured on a fork, never against live state.
- No DoS, no market manipulation, no live-state mutation, no phishing, no RAT/C2, no mass-targeting, no destruction of data — ever.
- Funds are returned or never moved; findings are disclosed to the project team.
- Only run it against targets you are authorized to test, within a program's scope.
You are responsible for using it legally and within scope. See SKILL.md for the full operating doctrine.
# 1. Point your agent runtime at this repo (for Claude Code, SKILL.md is the brain —
# copy or symlink it to CLAUDE.md, and copy .claude/settings.example.json to
# .claude/settings.json, fixing the hook paths to your checkout).
cp SKILL.md CLAUDE.md
cp .claude/settings.example.json .claude/settings.json
# 2. Configure optional API keys (each missing key just disables that module).
cp .env.example .env # then edit
# 3. Heavy tools run in Docker; build the image once.
docker build -t bbt docker/
# 4. In your agent, invoke a skill:
# /deephunt <repo-or-target>Everything an agent needs to operate the system is in SKILL.md. The detailed
methodology lives in methodology/; the engines in scripts/.
Verify the engine on a fresh clone (no setup, no API keys needed):
python run_selftests.py
# -> 30 passed, 4 skipped, 0 failedThe 4 skipped tests need a private benchmark/calibration corpus that isn't part of the public release; everything else runs standalone.
The outward-facing surface is English — this README, the operating guide (SKILL.md),
the methodology (methodology/), and the web2 payload/framework references and
report templates. Parts of the internals are still Russian and are being localized progressively,
in the open: the three skill playbooks under .claude/commands/ (SKILL.md §4
summarizes each in English), the enforcement hooks and gate engines, the web2 engine's inline
comments, and the replay/self-test fixtures. Much of this is format-strings the runtime matches on,
so translating it in isolation would break the system — it gets its English pass over time. The
tool runs regardless of comment language.
Apache-2.0. Use it, fork it, build on it. Patent grant included.
Part of Augustus Forge — autonomous decision systems, built in public. Failures included, wins earned.