Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -16,10 +16,10 @@ jobs:
name: safety
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@v7
with:
fetch-depth: 0
- uses: actions/setup-python@v5
- uses: actions/setup-python@v7
with:
python-version: '3.12'
- run: pip install pyyaml
Expand All @@ -34,7 +34,7 @@ jobs:
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
run: python3 .github/scripts/safety_scan.py --history
- name: Scan for committed secrets
uses: gitleaks/gitleaks-action@v2
uses: gitleaks/gitleaks-action@v3
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}

Expand All @@ -50,8 +50,8 @@ jobs:
run:
working-directory: pipelines/job-assessment
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
- uses: actions/checkout@v7
- uses: actions/setup-python@v7
with:
python-version: ${{ matrix.python }}
- run: pip install -r requirements.txt
Expand All @@ -64,7 +64,7 @@ jobs:
env:
PYTHONDONTWRITEBYTECODE: '1'
steps:
- uses: actions/setup-python@v5
- uses: actions/setup-python@v7
with:
python-version: '3.12'
- name: Clone the repo fresh, at the commit under test
Expand Down
2 changes: 1 addition & 1 deletion LICENSE
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
MIT License

Copyright (c) 2026 Benny Goddard
Copyright (c) 2026 Ben Goddard

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
Expand Down
10 changes: 6 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,7 @@
# Context Engineering Toolkit

[![ci](https://github.com/darthrootbeer/context-engineering-toolkit/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/darthrootbeer/context-engineering-toolkit/actions/workflows/ci.yml)

Skills, pipelines, and technique write-ups for building AI-agent systems that stay correct as they scale — built by using Claude Code every day, not by reading about it.

**The 60-second tour.**
Expand All @@ -12,7 +14,7 @@ Skills, pipelines, and technique write-ups for building AI-agent systems that st

**1. [`pipelines/docs-pipeline`](pipelines/docs-pipeline/)** is the strongest piece in this repo, and a generic export of the version I used. It is a chain of Claude Code skills that takes a documentation change from the first request through structure review, voice, grammar, links, visuals, expert review, and publishing, with a real gate between every stage. I used a version of it every working day on real documentation. If you only open one thing, open this, and start with its [README](pipelines/docs-pipeline/README.md) and [ARCHITECTURE](pipelines/docs-pipeline/ARCHITECTURE.md).

**2. [`pipelines/job-assessment`](pipelines/job-assessment/)** scores a job posting against your own written criteria and gives a plain verdict: Apply, Apply with reservations, or Skip. A model reads the posting and writes down findings, quoting the posting for every claim. Ordinary code then checks those quotes, does the scoring and picks the verdict, so the same findings always give the same answer. It was built by directing Claude Code, not hand-typed. What is shown: the whole chain runs offline on three invented postings with no model, no network and no API key, and the tests in that folder (run again in a fresh clone by CI) compare its scores, verdicts and saved output files to written-down expected values. A separate run with the `claude` CLI on Claude Sonnet is saved in [`tests/`](pipelines/job-assessment/tests/): the interview built a file that passed the checker, and the assessment gave the expected verdicts on 2 of 2 postings in its second attempt, after the first attempt got one verdict wrong ([both records are kept](pipelines/job-assessment/tests/e2e-output-run1.md)). What is not shown: that a verdict predicts getting hired, or that it beats any other method. No real postings or real person's data are in it. Its README starts with a five-minute quickstart from a clean clone, and every README and guide in the folder ends with prompts tested on Claude Sonnet and Claude Haiku.
**2. [`pipelines/job-assessment`](pipelines/job-assessment/)** scores a job posting against your own written criteria and gives a plain verdict: Apply, Apply with reservations, or Skip. A model reads the posting and quotes it for every finding, then ordinary code checks the quotes, does the scoring and picks the verdict, so the same findings always give the same answer. It runs offline from a clean clone with no model and no API key, over 500 tests pin its scores, verdicts and output files, and CI re-runs them in a fresh clone on every change; its README has the five-minute quickstart, the saved live runs (including one that got a verdict wrong), and what it does not show.

**3. [`patterns/block-and-tell-hooks.md`](patterns/block-and-tell-hooks.md)** explains how to make an AI agent follow a rule every time instead of hoping it remembers. It is the idea behind most of the guards in my own setup.

Expand All @@ -26,21 +28,21 @@ Every piece in here started as a real problem: a rule that kept getting skipped,

## What's inside

**`skills/`** — two self-contained Claude Code skills, each in its own folder with a `SKILL.md` that defines one repeatable, well-scoped task for an AI agent: `docs-readability-check` (a readability pass on documentation) and `plan-this` (capture an in-progress plan so it survives a context reset). `plan-this` also has a README.
**`skills/`** — two self-contained Claude Code skills, each in its own folder with a `SKILL.md` that defines one repeatable, well-scoped task for an AI agent: `docs-readability-check` (a readability pass on documentation) and `plan-this` (capture an in-progress plan so it survives a context reset). Each has a README.

**`pipelines/`** — multi-stage systems, not single tasks. The anchor piece here is a documentation-engineering pipeline: a chain of skills that takes a raw content change through structure review, voice/style checks, grammar, link and visual verification, and a subject-matter-expert review gate before anything publishes. See `pipelines/docs-pipeline/` for its own README and current state. The second pipeline is [`pipelines/job-assessment/`](pipelines/job-assessment/), which scores a job posting against written criteria.

**`patterns/`** — written technique docs for ideas that are more valuable described in prose than shipped as literal runnable code, either because the real implementation is too specific to one project to be useful as-is, or because the idea itself is the point. Covers: how to make an "always do X first" instruction actually reliable instead of hoped-for (block-and-tell hooks), how to keep an agent's standing instructions from becoming an unmaintainable single file as they grow (rules-index architecture), and how to give an agent memory that survives months of use without turning into an unreadable dump (typed, size-bounded memory).

**Prompt blocks.** The README and guides of [`pipelines/docs-pipeline/`](pipelines/docs-pipeline/) and [`pipelines/job-assessment/`](pipelines/job-assessment/) and the three docs in `patterns/` each end with a "Prompt for your AI model" block, tested on Claude Sonnet and Claude Haiku. This top-level README, `ROADMAP.md` and the two skills do not have one.
**Prompt blocks.** The READMEs and guides of the two pipelines, the `docs-readability-check` README and the three pattern docs end with a "Prompt for your AI model" block. Each of those prompts has a written "A good answer ..." sentence. Records of runs on Claude Sonnet and Claude Haiku are saved in `pipelines/job-assessment/tests/prompt-runs/`, `pipelines/docs-pipeline/tests/prompt-runs/` (which also covers the readability README) and `patterns/tests/prompt-runs/`. A Claude model graded each answer against its sentence, and I have not re-read every answer. The prompts in the docs-pipeline ARCHITECTURE guide and style guides were checked without keeping the answers. This README, `ROADMAP.md` and `plan-this` have no prompt block.

## How the pieces relate

A skill is one task. A pipeline is several skills chained with real gates between them (nothing moves to the next stage until the current one passes). A pattern is the idea behind a mechanism, written down so it can be rebuilt in a different codebase without copying code that won't fit.

## A note on how this was built

Everything here was built using Claude Code, directed and reviewed by a human, not hand-typed line by line. That's not a caveat — it's the actual differentiator this repo is trying to demonstrate: designing the system, prompting and iterating to get it right, and evaluating the output rigorously enough to trust it. The skills and patterns in here are, among other things, examples of exactly that evaluation discipline applied to itself.
Everything here was built using Claude Code, directed by a human and checked with tests, not hand-typed line by line. That's not a caveat — it's the actual differentiator this repo is trying to demonstrate: designing the system, prompting and iterating to get it right, and evaluating the output rigorously enough to trust it. The skills and patterns in here are, among other things, examples of exactly that evaluation discipline applied to itself.

## What's next

Expand Down
35 changes: 33 additions & 2 deletions patterns/block-and-tell-hooks.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,32 @@ FILE=$(echo "$INPUT" | jq -r '.tool_input.file_path // ""')
exit 0
```

Register both in `~/.claude/settings.json` under `hooks`, each with a `matcher` for the tool names it should watch. The marker is keyed by session id so one session's read does not clear another session's block.
Register both in `~/.claude/settings.json` under `hooks`. Each event holds a list of entries, and each entry has a `matcher` (a tool name, or several joined with `|`) and a `hooks` list of commands:

```json
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{ "type": "command", "command": "~/.claude/hooks/deploy-changelog-guard.sh" }
]
}
],
"PostToolUse": [
{
"matcher": "Read|Edit|Write",
"hooks": [
{ "type": "command", "command": "~/.claude/hooks/deploy-changelog-marker.sh" }
]
}
]
}
}
```

The event name (`PreToolUse`, `PostToolUse`) is the key, not a field inside the entry. The matcher is one string, not a list. List every tool that can touch the file, or a write through the tool you left out slips past. The marker is keyed by session id so one session's read does not clear another session's block.

## Three sharp edges

Expand Down Expand Up @@ -124,6 +149,8 @@ Give any AI model this file plus one of the prompts below. Paste the file text w
Here is a design pattern document: [PASTE FILE]

Explain it to me as if I know what a script is but have never used hooks. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why exit code 2 matters and why a guard needs a release path. Wait for my answer before each next question.

A good answer uses an analogy that is not the document's, covers why exit 2 blocks and exit 1 does not, and explains why a guard with no release path locks the agent out. It asks exactly three questions, one at a time, and waits for my answer before the next.
```

**2. Review it against your setup**
Expand All @@ -134,6 +161,8 @@ Here is a design pattern document: [PASTE FILE]
Below is a description of my own agent setup (tool, hook support, rules I currently rely on): [DESCRIBE YOUR SETUP]

List which parts of the pattern apply to my setup and which do not. Flag any claim that may be false for my tool, for example how it treats exit codes. Name one rule of mine that is a good candidate for a guard and one that is not, with a reason for each.

A good answer sorts the pattern into parts that apply and parts that do not for my stated tool, says plainly where the tool's exit-code behavior should be checked and does not assume it, and names one rule that suits a guard and one that does not, with a reason for each.
```

**3. Adapt and test it**
Expand All @@ -144,6 +173,8 @@ Here is a design pattern document: [PASTE FILE]
My rule is: [YOUR "ALWAYS DO X FIRST" RULE]. My tool is: [YOUR AGENT TOOL].

Write the guard and marker scripts for my rule in my tool's hook format. Then give me a three-step test plan: one test where the guard must block, one where it must allow after the marker is written, and one where I confirm the release path works. Tell me what output proves each test passed.

A good answer gives a guard that exits 2 with a message on stderr and a marker script keyed by session id, shows the exact settings entry for my tool, and watches every tool that can change the file (for example both Edit and Write). Its three tests cover block, allow after the marker, and the release path, each with the output that proves it passed.
```

**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically. I did not save those answers, so there is no record to read here, unlike the saved runs in `pipelines/job-assessment/tests/`.
**How these prompts were checked.** On 2026-10-01 each of the three prompts was run once on Claude Sonnet and once on Claude Haiku, with this file pasted in and sample details filled in. A Claude model (Sonnet 5.5) graded each answer against the "A good answer ..." sentence under the prompt. I have not re-read every answer. Sonnet met all three. Haiku met the first, and met the second and third only in part: it told me not to check the exit-code claim for my tool, and its guard matched only a relative path, so an absolute path would pass it. The settings example above was added after an earlier Haiku run invented the wrong registration shape, and the saved Haiku run now has it right. The answers are in [`tests/prompt-runs/`](tests/prompt-runs/).
8 changes: 7 additions & 1 deletion patterns/rules-index-architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -141,6 +141,8 @@ Give any AI model this file plus one of the prompts below. Paste the file text w
Here is a design pattern document: [PASTE FILE]

Explain it to me as if I have one long instructions file and have never split it. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why splitting files does not shrink what is loaded, and what the two opt-outs are. Wait for my answer before each next question.

A good answer uses an analogy that is not the document's, says that splitting a file does not shrink what is loaded at startup, and names both opt-outs. It asks exactly three questions, one at a time, and waits for my answer before the next.
```

**2. Review it against your setup**
Expand All @@ -151,6 +153,8 @@ Here is a design pattern document: [PASTE FILE]
Below is the folder listing and approximate line counts of my own instruction files, plus the agent tool I use: [PASTE LISTING AND TOOL NAME]

Tell me whether my setup needs this pattern yet. Flag any claim in the document that may not be true for my tool, especially about what gets loaded at startup. Suggest which of my files should be split out, which should stay, and which belong in a reference folder, and give a reason for each.

A good answer judges from my listing whether the pattern is needed yet, says which startup-loading claims should be checked for my tool and does not assume them, and sorts my files into split out, stay, and reference folder with a reason for each.
```

**3. Adapt and test it**
Expand All @@ -163,6 +167,8 @@ My instructions file is pasted below. My tool is: [YOUR AGENT TOOL].
[PASTE YOUR INSTRUCTIONS FILE]

Propose a split into domain files with an index table, and tell me which files should be excluded from startup loading. Then give me a way to test that the split worked: how to measure total loaded size before and after, and one question to ask the agent that only a specific rule file can answer.

A good answer proposes domain files and an index table built from my file's own contents, names which files to exclude from startup loading, and gives a before-and-after way to measure loaded size plus one question that only a specific rule file can answer.
```

**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically. I did not save those answers, so there is no record to read here, unlike the saved runs in `pipelines/job-assessment/tests/`.
**How these prompts were checked.** On 2026-10-01 each of the three prompts was run once on Claude Sonnet and once on Claude Haiku, with this file pasted in and sample details filled in. A Claude model (Sonnet 5.5) graded each answer against the "A good answer ..." sentence under the prompt. I have not re-read every answer. Sonnet met all three. Haiku met the first and only partly met the other two: it checked only one of the document's claims about startup loading, invented a saving figure, and told me to split a file of about 15 lines when this document says that is too small to bother. The answers are in [`tests/prompt-runs/`](tests/prompt-runs/).
Loading
Loading