diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index a1ec998..11071e8 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -16,10 +16,10 @@ jobs: name: safety runs-on: ubuntu-latest steps: - - uses: actions/checkout@v4 + - uses: actions/checkout@v7 with: fetch-depth: 0 - - uses: actions/setup-python@v5 + - uses: actions/setup-python@v7 with: python-version: '3.12' - run: pip install pyyaml @@ -34,7 +34,7 @@ jobs: if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch' run: python3 .github/scripts/safety_scan.py --history - name: Scan for committed secrets - uses: gitleaks/gitleaks-action@v2 + uses: gitleaks/gitleaks-action@v3 env: GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} @@ -50,8 +50,8 @@ jobs: run: working-directory: pipelines/job-assessment steps: - - uses: actions/checkout@v4 - - uses: actions/setup-python@v5 + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 with: python-version: ${{ matrix.python }} - run: pip install -r requirements.txt @@ -64,7 +64,7 @@ jobs: env: PYTHONDONTWRITEBYTECODE: '1' steps: - - uses: actions/setup-python@v5 + - uses: actions/setup-python@v7 with: python-version: '3.12' - name: Clone the repo fresh, at the commit under test diff --git a/LICENSE b/LICENSE index b44f74d..dbb1616 100644 --- a/LICENSE +++ b/LICENSE @@ -1,6 +1,6 @@ MIT License -Copyright (c) 2026 Benny Goddard +Copyright (c) 2026 Ben Goddard Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal diff --git a/README.md b/README.md index 413be06..8a5796c 100644 --- a/README.md +++ b/README.md @@ -1,5 +1,7 @@ # Context Engineering Toolkit +[![ci](https://github.com/darthrootbeer/context-engineering-toolkit/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/darthrootbeer/context-engineering-toolkit/actions/workflows/ci.yml) + Skills, pipelines, and technique write-ups for building AI-agent systems that stay correct as they scale — built by using Claude Code every day, not by reading about it. **The 60-second tour.** @@ -12,7 +14,7 @@ Skills, pipelines, and technique write-ups for building AI-agent systems that st **1. [`pipelines/docs-pipeline`](pipelines/docs-pipeline/)** is the strongest piece in this repo, and a generic export of the version I used. It is a chain of Claude Code skills that takes a documentation change from the first request through structure review, voice, grammar, links, visuals, expert review, and publishing, with a real gate between every stage. I used a version of it every working day on real documentation. If you only open one thing, open this, and start with its [README](pipelines/docs-pipeline/README.md) and [ARCHITECTURE](pipelines/docs-pipeline/ARCHITECTURE.md). -**2. [`pipelines/job-assessment`](pipelines/job-assessment/)** scores a job posting against your own written criteria and gives a plain verdict: Apply, Apply with reservations, or Skip. A model reads the posting and writes down findings, quoting the posting for every claim. Ordinary code then checks those quotes, does the scoring and picks the verdict, so the same findings always give the same answer. It was built by directing Claude Code, not hand-typed. What is shown: the whole chain runs offline on three invented postings with no model, no network and no API key, and the tests in that folder (run again in a fresh clone by CI) compare its scores, verdicts and saved output files to written-down expected values. A separate run with the `claude` CLI on Claude Sonnet is saved in [`tests/`](pipelines/job-assessment/tests/): the interview built a file that passed the checker, and the assessment gave the expected verdicts on 2 of 2 postings in its second attempt, after the first attempt got one verdict wrong ([both records are kept](pipelines/job-assessment/tests/e2e-output-run1.md)). What is not shown: that a verdict predicts getting hired, or that it beats any other method. No real postings or real person's data are in it. Its README starts with a five-minute quickstart from a clean clone, and every README and guide in the folder ends with prompts tested on Claude Sonnet and Claude Haiku. +**2. [`pipelines/job-assessment`](pipelines/job-assessment/)** scores a job posting against your own written criteria and gives a plain verdict: Apply, Apply with reservations, or Skip. A model reads the posting and quotes it for every finding, then ordinary code checks the quotes, does the scoring and picks the verdict, so the same findings always give the same answer. It runs offline from a clean clone with no model and no API key, over 500 tests pin its scores, verdicts and output files, and CI re-runs them in a fresh clone on every change; its README has the five-minute quickstart, the saved live runs (including one that got a verdict wrong), and what it does not show. **3. [`patterns/block-and-tell-hooks.md`](patterns/block-and-tell-hooks.md)** explains how to make an AI agent follow a rule every time instead of hoping it remembers. It is the idea behind most of the guards in my own setup. @@ -26,13 +28,13 @@ Every piece in here started as a real problem: a rule that kept getting skipped, ## What's inside -**`skills/`** — two self-contained Claude Code skills, each in its own folder with a `SKILL.md` that defines one repeatable, well-scoped task for an AI agent: `docs-readability-check` (a readability pass on documentation) and `plan-this` (capture an in-progress plan so it survives a context reset). `plan-this` also has a README. +**`skills/`** — two self-contained Claude Code skills, each in its own folder with a `SKILL.md` that defines one repeatable, well-scoped task for an AI agent: `docs-readability-check` (a readability pass on documentation) and `plan-this` (capture an in-progress plan so it survives a context reset). Each has a README. **`pipelines/`** — multi-stage systems, not single tasks. The anchor piece here is a documentation-engineering pipeline: a chain of skills that takes a raw content change through structure review, voice/style checks, grammar, link and visual verification, and a subject-matter-expert review gate before anything publishes. See `pipelines/docs-pipeline/` for its own README and current state. The second pipeline is [`pipelines/job-assessment/`](pipelines/job-assessment/), which scores a job posting against written criteria. **`patterns/`** — written technique docs for ideas that are more valuable described in prose than shipped as literal runnable code, either because the real implementation is too specific to one project to be useful as-is, or because the idea itself is the point. Covers: how to make an "always do X first" instruction actually reliable instead of hoped-for (block-and-tell hooks), how to keep an agent's standing instructions from becoming an unmaintainable single file as they grow (rules-index architecture), and how to give an agent memory that survives months of use without turning into an unreadable dump (typed, size-bounded memory). -**Prompt blocks.** The README and guides of [`pipelines/docs-pipeline/`](pipelines/docs-pipeline/) and [`pipelines/job-assessment/`](pipelines/job-assessment/) and the three docs in `patterns/` each end with a "Prompt for your AI model" block, tested on Claude Sonnet and Claude Haiku. This top-level README, `ROADMAP.md` and the two skills do not have one. +**Prompt blocks.** The READMEs and guides of the two pipelines, the `docs-readability-check` README and the three pattern docs end with a "Prompt for your AI model" block. Each of those prompts has a written "A good answer ..." sentence. Records of runs on Claude Sonnet and Claude Haiku are saved in `pipelines/job-assessment/tests/prompt-runs/`, `pipelines/docs-pipeline/tests/prompt-runs/` (which also covers the readability README) and `patterns/tests/prompt-runs/`. A Claude model graded each answer against its sentence, and I have not re-read every answer. The prompts in the docs-pipeline ARCHITECTURE guide and style guides were checked without keeping the answers. This README, `ROADMAP.md` and `plan-this` have no prompt block. ## How the pieces relate @@ -40,7 +42,7 @@ A skill is one task. A pipeline is several skills chained with real gates betwee ## A note on how this was built -Everything here was built using Claude Code, directed and reviewed by a human, not hand-typed line by line. That's not a caveat — it's the actual differentiator this repo is trying to demonstrate: designing the system, prompting and iterating to get it right, and evaluating the output rigorously enough to trust it. The skills and patterns in here are, among other things, examples of exactly that evaluation discipline applied to itself. +Everything here was built using Claude Code, directed by a human and checked with tests, not hand-typed line by line. That's not a caveat — it's the actual differentiator this repo is trying to demonstrate: designing the system, prompting and iterating to get it right, and evaluating the output rigorously enough to trust it. The skills and patterns in here are, among other things, examples of exactly that evaluation discipline applied to itself. ## What's next diff --git a/patterns/block-and-tell-hooks.md b/patterns/block-and-tell-hooks.md index ee83c75..cb76fdf 100644 --- a/patterns/block-and-tell-hooks.md +++ b/patterns/block-and-tell-hooks.md @@ -83,7 +83,32 @@ FILE=$(echo "$INPUT" | jq -r '.tool_input.file_path // ""') exit 0 ``` -Register both in `~/.claude/settings.json` under `hooks`, each with a `matcher` for the tool names it should watch. The marker is keyed by session id so one session's read does not clear another session's block. +Register both in `~/.claude/settings.json` under `hooks`. Each event holds a list of entries, and each entry has a `matcher` (a tool name, or several joined with `|`) and a `hooks` list of commands: + +```json +{ + "hooks": { + "PreToolUse": [ + { + "matcher": "Bash", + "hooks": [ + { "type": "command", "command": "~/.claude/hooks/deploy-changelog-guard.sh" } + ] + } + ], + "PostToolUse": [ + { + "matcher": "Read|Edit|Write", + "hooks": [ + { "type": "command", "command": "~/.claude/hooks/deploy-changelog-marker.sh" } + ] + } + ] + } +} +``` + +The event name (`PreToolUse`, `PostToolUse`) is the key, not a field inside the entry. The matcher is one string, not a list. List every tool that can touch the file, or a write through the tool you left out slips past. The marker is keyed by session id so one session's read does not clear another session's block. ## Three sharp edges @@ -124,6 +149,8 @@ Give any AI model this file plus one of the prompts below. Paste the file text w Here is a design pattern document: [PASTE FILE] Explain it to me as if I know what a script is but have never used hooks. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why exit code 2 matters and why a guard needs a release path. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, covers why exit 2 blocks and exit 1 does not, and explains why a guard with no release path locks the agent out. It asks exactly three questions, one at a time, and waits for my answer before the next. ``` **2. Review it against your setup** @@ -134,6 +161,8 @@ Here is a design pattern document: [PASTE FILE] Below is a description of my own agent setup (tool, hook support, rules I currently rely on): [DESCRIBE YOUR SETUP] List which parts of the pattern apply to my setup and which do not. Flag any claim that may be false for my tool, for example how it treats exit codes. Name one rule of mine that is a good candidate for a guard and one that is not, with a reason for each. + +A good answer sorts the pattern into parts that apply and parts that do not for my stated tool, says plainly where the tool's exit-code behavior should be checked and does not assume it, and names one rule that suits a guard and one that does not, with a reason for each. ``` **3. Adapt and test it** @@ -144,6 +173,8 @@ Here is a design pattern document: [PASTE FILE] My rule is: [YOUR "ALWAYS DO X FIRST" RULE]. My tool is: [YOUR AGENT TOOL]. Write the guard and marker scripts for my rule in my tool's hook format. Then give me a three-step test plan: one test where the guard must block, one where it must allow after the marker is written, and one where I confirm the release path works. Tell me what output proves each test passed. + +A good answer gives a guard that exits 2 with a message on stderr and a marker script keyed by session id, shows the exact settings entry for my tool, and watches every tool that can change the file (for example both Edit and Write). Its three tests cover block, allow after the marker, and the release path, each with the output that proves it passed. ``` -**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically. I did not save those answers, so there is no record to read here, unlike the saved runs in `pipelines/job-assessment/tests/`. +**How these prompts were checked.** On 2026-10-01 each of the three prompts was run once on Claude Sonnet and once on Claude Haiku, with this file pasted in and sample details filled in. A Claude model (Sonnet 5.5) graded each answer against the "A good answer ..." sentence under the prompt. I have not re-read every answer. Sonnet met all three. Haiku met the first, and met the second and third only in part: it told me not to check the exit-code claim for my tool, and its guard matched only a relative path, so an absolute path would pass it. The settings example above was added after an earlier Haiku run invented the wrong registration shape, and the saved Haiku run now has it right. The answers are in [`tests/prompt-runs/`](tests/prompt-runs/). diff --git a/patterns/rules-index-architecture.md b/patterns/rules-index-architecture.md index 4467c4e..85fe385 100644 --- a/patterns/rules-index-architecture.md +++ b/patterns/rules-index-architecture.md @@ -141,6 +141,8 @@ Give any AI model this file plus one of the prompts below. Paste the file text w Here is a design pattern document: [PASTE FILE] Explain it to me as if I have one long instructions file and have never split it. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why splitting files does not shrink what is loaded, and what the two opt-outs are. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, says that splitting a file does not shrink what is loaded at startup, and names both opt-outs. It asks exactly three questions, one at a time, and waits for my answer before the next. ``` **2. Review it against your setup** @@ -151,6 +153,8 @@ Here is a design pattern document: [PASTE FILE] Below is the folder listing and approximate line counts of my own instruction files, plus the agent tool I use: [PASTE LISTING AND TOOL NAME] Tell me whether my setup needs this pattern yet. Flag any claim in the document that may not be true for my tool, especially about what gets loaded at startup. Suggest which of my files should be split out, which should stay, and which belong in a reference folder, and give a reason for each. + +A good answer judges from my listing whether the pattern is needed yet, says which startup-loading claims should be checked for my tool and does not assume them, and sorts my files into split out, stay, and reference folder with a reason for each. ``` **3. Adapt and test it** @@ -163,6 +167,8 @@ My instructions file is pasted below. My tool is: [YOUR AGENT TOOL]. [PASTE YOUR INSTRUCTIONS FILE] Propose a split into domain files with an index table, and tell me which files should be excluded from startup loading. Then give me a way to test that the split worked: how to measure total loaded size before and after, and one question to ask the agent that only a specific rule file can answer. + +A good answer proposes domain files and an index table built from my file's own contents, names which files to exclude from startup loading, and gives a before-and-after way to measure loaded size plus one question that only a specific rule file can answer. ``` -**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically. I did not save those answers, so there is no record to read here, unlike the saved runs in `pipelines/job-assessment/tests/`. +**How these prompts were checked.** On 2026-10-01 each of the three prompts was run once on Claude Sonnet and once on Claude Haiku, with this file pasted in and sample details filled in. A Claude model (Sonnet 5.5) graded each answer against the "A good answer ..." sentence under the prompt. I have not re-read every answer. Sonnet met all three. Haiku met the first and only partly met the other two: it checked only one of the document's claims about startup loading, invented a saving figure, and told me to split a file of about 15 lines when this document says that is too small to bother. The answers are in [`tests/prompt-runs/`](tests/prompt-runs/). diff --git a/patterns/tests/prompt-runs/README.md b/patterns/tests/prompt-runs/README.md new file mode 100644 index 0000000..763df32 --- /dev/null +++ b/patterns/tests/prompt-runs/README.md @@ -0,0 +1,28 @@ +# Saved prompt runs for the pattern docs + +These are the answers behind the "How these prompts were checked" paragraph at the end of each doc in `patterns/`. + +**How they were produced.** On 2026-10-01 each of the nine "Prompt for your AI model" prompts was run through the Claude Code command line (`claude -p`, version 2.1.287) with no tools, no project instructions and automatic memory turned off, once on `sonnet` and once on `haiku`. The doc's text was pasted where the prompt says `[PASTE FILE]`, and the other bracketed inputs were replaced with made-up samples. Each saved file shows the prompt that was sent, then the answer. The script that does this is `../run_prompts.py`. It needs the `claude` tool, so the normal test suite does not run it. + +**How they were graded.** A Claude model (Sonnet 5.5, working under my direction) read each answer against that prompt's own "A good answer ..." sentence, which was written before the run. One run per prompt per model, so a Pass here means "this run met the sentence", not "this prompt always does". I have not re-read every answer myself. + +File names are `--.md`. + +| Doc and prompt | Sonnet | Haiku | +|---|---|---| +| block-and-tell-hooks, understand and teach | Pass | Pass | +| block-and-tell-hooks, review against your setup | Pass | Partial. It said the exit-code claims did not need checking for my tool, which the prompt asks it to question. | +| block-and-tell-hooks, adapt and test | Pass | Partial. The settings entry has the right shape and watches both Edit and Write. The guard only matches the relative path `migrations/*`, so an absolute file path would slip past it. | +| rules-index-architecture, understand and teach | Pass | Pass | +| rules-index-architecture, review against your setup | Pass | Partial. It flagged one claim to check and no others, and gave a made-up saving of "15 to 20 percent". | +| rules-index-architecture, adapt and test | Pass | Partial. It proposed splitting a file of about 15 lines without saying the doc's "when not to" section applies, and its before and after measurements did not cover the same files. | +| typed-memory-system, understand and teach | Pass | Partial. It never explained why compaction keeps topic files. | +| typed-memory-system, review against your setup | Pass | Partial. It listed gaps it could not see from a file listing (frontmatter, index grouping). | +| typed-memory-system, adapt and test | Pass | Partial. The size script uses a variable it never sets, and its test expects exit code 1 from a branch that exits 0. | + +Sonnet met the sentence on all nine. Haiku met it on two and only partly on seven. Those Haiku misses are the reason the docs say "check what it gives you" and do not say the prompts are reliable on a small model. + +## What the first attempt taught + +- **A first set of 18 runs was thrown away.** The runs picked up Claude Code's automatic memory instructions: one Sonnet answer tried to call a file tool and returned nothing usable, and a Haiku answer put a scratch path in its script. I turned automatic memory off (`CLAUDE_CODE_DISABLE_AUTO_MEMORY=1`) and ran all 18 again. Only the second set is saved. +- **The hook registration.** Before the settings example was added to `block-and-tell-hooks.md`, a reviewer's Haiku run of the "adapt and test" prompt invented the registration (`"matcher": ["Edit"]` with an `"event"` key, which is not Claude Code's shape) and guarded only `Edit`. After the example was added, both models in these saved runs wrote the correct shape and watched both `Edit` and `Write`. diff --git a/patterns/tests/prompt-runs/block-and-tell-hooks-adapt-haiku.md b/patterns/tests/prompt-runs/block-and-tell-hooks-adapt-haiku.md new file mode 100644 index 0000000..28ab1db --- /dev/null +++ b/patterns/tests/prompt-runs/block-and-tell-hooks-adapt-haiku.md @@ -0,0 +1,104 @@ +# block-and-tell-hooks prompt 3 (Adapt and test it) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +My rule is: Before editing any file under migrations/, read docs/MIGRATIONS.md in the same session.. My tool is: Claude Code. + +Write the guard and marker scripts for my rule in my tool's hook format. Then give me a three-step test plan: one test where the guard must block, one where it must allow after the marker is written, and one where I confirm the release path works. Tell me what output proves each test passed. + +A good answer gives a guard that exits 2 with a message on stderr and a marker script keyed by session id, shows the exact settings entry for my tool, and watches every tool that can change the file (for example both Edit and Write). Its three tests cover block, allow after the marker, and the release path, each with the output that proves it passed. +``` + +Answer: + +# Guard and marker scripts for migrations rule + +I'll give you the two scripts, the settings entry, and a test plan. + +## Guard script: `~/.claude/hooks/migrations-guard.sh` + +```bash +#!/bin/bash +INPUT=$(cat) +SESSION_ID=$(echo "$INPUT" | jq -r '.session_id // "nosession"') +FILE=$(echo "$INPUT" | jq -r '.tool_input.file_path // ""') + +# Fast path: only block migrations/ edits +case "$FILE" in + migrations/*) ;; + *) exit 0 ;; +esac + +MARKER="/tmp/migrations-guide-read-${SESSION_ID}" +if [[ ! -f "$MARKER" ]]; then + echo "Blocked: docs/MIGRATIONS.md must be read before editing migration files." >&2 + echo "Read docs/MIGRATIONS.md first to understand the migration guidelines." >&2 + exit 2 +fi +exit 0 +``` + +## Marker script: `~/.claude/hooks/migrations-marker.sh` + +```bash +#!/bin/bash +INPUT=$(cat) +SESSION_ID=$(echo "$INPUT" | jq -r '.session_id // "nosession"') +FILE=$(echo "$INPUT" | jq -r '.tool_input.file_path // ""') + +[[ "$FILE" == "docs/MIGRATIONS.md" ]] && touch "/tmp/migrations-guide-read-${SESSION_ID}" +exit 0 +``` + +## Settings entry for `~/.claude/settings.json` + +Add this to your `hooks` object: + +```json +{ + "hooks": { + "PreToolUse": [ + { + "matcher": "Edit|Write", + "hooks": [ + { "type": "command", "command": "~/.claude/hooks/migrations-guard.sh" } + ] + } + ], + "PostToolUse": [ + { + "matcher": "Read", + "hooks": [ + { "type": "command", "command": "~/.claude/hooks/migrations-marker.sh" } + ] + } + ] + } +} +``` + +## Three-step test plan + +**Test 1: Block without reading (must fail)** +- Try to edit `migrations/001_initial.sql` +- Expected: Edit is blocked +- Proof: stderr shows exactly "Blocked: docs/MIGRATIONS.md must be read before editing migration files." followed by the instruction to read it + +**Test 2: Allow after reading (must succeed)** +- Read `docs/MIGRATIONS.md` +- Then try to edit `migrations/001_initial.sql` again in the same session +- Expected: Edit succeeds without the block message +- Proof: no stderr message about being blocked; the file content changes are applied + +**Test 3: Escape hatch (the release path)** +- Start a fresh session and try editing `migrations/002_add_users.sql` +- Get blocked with the "Blocked" message (Test 1 state) +- Read `docs/MIGRATIONS.md` +- Edit `migrations/002_add_users.sql` again +- Expected: the second edit succeeds +- Proof: first attempt shows block message in stderr; second attempt shows no block message and the edit is applied; marker file `/tmp/migrations-guide-read-` exists after the read + +The guard watches both `Edit` and `Write` because migrations might be written programmatically. The marker watches only `Read` and checks for the exact filename, so grepping for "MIGRATIONS" in a commit message won't accidentally clear the block. diff --git a/patterns/tests/prompt-runs/block-and-tell-hooks-adapt-sonnet.md b/patterns/tests/prompt-runs/block-and-tell-hooks-adapt-sonnet.md new file mode 100644 index 0000000..3cdabf9 --- /dev/null +++ b/patterns/tests/prompt-runs/block-and-tell-hooks-adapt-sonnet.md @@ -0,0 +1,158 @@ +# block-and-tell-hooks prompt 3 (Adapt and test it) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +My rule is: Before editing any file under migrations/, read docs/MIGRATIONS.md in the same session.. My tool is: Claude Code. + +Write the guard and marker scripts for my rule in my tool's hook format. Then give me a three-step test plan: one test where the guard must block, one where it must allow after the marker is written, and one where I confirm the release path works. Tell me what output proves each test passed. + +A good answer gives a guard that exits 2 with a message on stderr and a marker script keyed by session id, shows the exact settings entry for my tool, and watches every tool that can change the file (for example both Edit and Write). Its three tests cover block, allow after the marker, and the release path, each with the output that proves it passed. +``` + +Answer: + +# Migrations-doc guard for Claude Code + +## Files + +Save both in `.claude/hooks/` and run `chmod +x` on them. + +**`.claude/hooks/migrations-guard.sh`** (PreToolUse) + +```bash +#!/bin/bash +# Blocks edits under migrations/ until docs/MIGRATIONS.md has been read this session. +INPUT=$(cat) +SESSION_ID=$(echo "$INPUT" | jq -r '.session_id // "nosession"') +TOOL=$(echo "$INPUT" | jq -r '.tool_name // ""') +MARKER="/tmp/migrations-doc-read-${SESSION_ID}" + +case "$TOOL" in + Edit|Write|MultiEdit) + TARGET=$(echo "$INPUT" | jq -r '.tool_input.file_path // ""') + case "$TARGET" in + */migrations/*|migrations/*) ;; + *) exit 0 ;; # fast path: not a migrations file + esac + ;; + Bash) + CMD=$(echo "$INPUT" | jq -r '.tool_input.command // ""') + case "$CMD" in + *migrations/*) ;; + *) exit 0 ;; # fast path: command never mentions migrations/ + esac + # Only block write-shaped commands, so cat/ls/grep on migrations/ still work. + echo "$CMD" | grep -Eq '(sed[[:space:]]+(-[A-Za-z]+[[:space:]]+)*-i|perl[[:space:]].*-i|tee[[:space:]]|(^|[;&|[:space:]])(mv|rm|cp|touch)[[:space:]]|[^0-9&>]>>?[[:space:]]*[^&[:space:]]*migrations/)' || exit 0 + ;; + *) exit 0 ;; +esac + +if [[ ! -f "$MARKER" ]]; then + echo "Blocked: docs/MIGRATIONS.md has not been read this session." >&2 + echo "Read docs/MIGRATIONS.md with the Read tool, then retry this change to migrations/." >&2 + echo "Why: it holds the rules for writing and ordering migrations." >&2 + exit 2 +fi +exit 0 +``` + +**`.claude/hooks/migrations-marker.sh`** (PostToolUse) + +```bash +#!/bin/bash +# Releases the migrations block once docs/MIGRATIONS.md has been read. +INPUT=$(cat) +SESSION_ID=$(echo "$INPUT" | jq -r '.session_id // "nosession"') +FILE=$(echo "$INPUT" | jq -r '.tool_input.file_path // ""') +[[ "$FILE" == *docs/MIGRATIONS.md ]] && touch "/tmp/migrations-doc-read-${SESSION_ID}" +exit 0 +``` + +## Settings entry + +Put this in `.claude/settings.json` (project) or `~/.claude/settings.json` (global): + +```json +{ + "hooks": { + "PreToolUse": [ + { + "matcher": "Edit|Write|MultiEdit|Bash", + "hooks": [ + { "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/migrations-guard.sh" } + ] + } + ], + "PostToolUse": [ + { + "matcher": "Read", + "hooks": [ + { "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/migrations-marker.sh" } + ] + } + ] + } +} +``` + +**Notes:** +- Only `Read` releases the block. A `cat docs/MIGRATIONS.md` through Bash won't, and the block message tells the agent to use Read. +- The Bash branch is a heuristic that recognizes write-shaped commands. If it blocks something harmless, treat that as a bug and tighten the regex. +- Prefer not to guard Bash at all? Remove `Bash` from the matcher and delete the `Bash)` branch. You'd then accept that `sed -i` on a migration can slip through. + +## Test plan + +Use a throwaway session id (`t1`) and clear any old marker first: `rm -f /tmp/migrations-doc-read-t1`. + +**1. The guard must block** + +```bash +echo '{"session_id":"t1","tool_name":"Edit","tool_input":{"file_path":"/repo/migrations/0042_add_users.sql"}}' \ + | .claude/hooks/migrations-guard.sh; echo "exit=$?" +``` + +Proof it passed: +- `exit=2`. +- stderr says `Blocked: docs/MIGRATIONS.md has not been read this session.` and tells you to read it. + +Also check that unrelated edits and read-only commands are allowed (should print `exit=0`, no stderr): + +```bash +echo '{"session_id":"t1","tool_name":"Edit","tool_input":{"file_path":"/repo/src/app.py"}}' | .claude/hooks/migrations-guard.sh; echo "exit=$?" +echo '{"session_id":"t1","tool_name":"Bash","tool_input":{"command":"ls migrations/"}}' | .claude/hooks/migrations-guard.sh; echo "exit=$?" +``` + +**2. The guard must allow after the marker is written** + +```bash +echo '{"session_id":"t1","tool_name":"Read","tool_input":{"file_path":"/repo/docs/MIGRATIONS.md"}}' \ + | .claude/hooks/migrations-marker.sh; echo "exit=$?" +ls /tmp/migrations-doc-read-t1 +echo '{"session_id":"t1","tool_name":"Write","tool_input":{"file_path":"/repo/migrations/0043_x.sql"}}' \ + | .claude/hooks/migrations-guard.sh; echo "exit=$?" +``` + +Proof it passed: +- The marker script prints `exit=0`. +- `ls` shows the marker file. +- The guard now prints `exit=0` with no stderr. + +**3. The release path works** + +This checks that the block can actually be cleared in a real session and doesn't leak across sessions. + +1. Start a fresh `claude` session in the repo and ask it to edit a file in `migrations/`. It should be blocked, and the block message should be visible in the transcript. +2. The agent should then read `docs/MIGRATIONS.md` on its own and retry the edit in the same turn, with no human input. +3. In the same session, confirm the doc read itself was never blocked. The guard ignores paths outside `migrations/`. +4. Open a second session, which has a new session id. It should be blocked again, which shows the marker is per-session. + +Proof it passed: +- The first edit attempt is refused with your stderr text. +- After the Read, the retried edit succeeds. +- `ls /tmp/migrations-doc-read-*` shows one marker per session that read the doc. +- The second session is blocked until it reads the doc. + +If the first edit succeeds with no block, run `/hooks` to confirm both hooks are registered. Then rerun test 1 by hand, since a guard that exits anything other than 2 won't block. diff --git a/patterns/tests/prompt-runs/block-and-tell-hooks-review-haiku.md b/patterns/tests/prompt-runs/block-and-tell-hooks-review-haiku.md new file mode 100644 index 0000000..6556932 --- /dev/null +++ b/patterns/tests/prompt-runs/block-and-tell-hooks-review-haiku.md @@ -0,0 +1,40 @@ +# block-and-tell-hooks prompt 2 (Review it against your setup) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Below is a description of my own agent setup (tool, hook support, rules I currently rely on): I use Claude Code. Hooks are enabled in ~/.claude/settings.json. Today my rules are prose in a CLAUDE.md file: always run the tests before a merge, never edit files under migrations/ without reading docs/MIGRATIONS.md, and keep commit messages short. + +List which parts of the pattern apply to my setup and which do not. Flag any claim that may be false for my tool, for example how it treats exit codes. Name one rule of mine that is a good candidate for a guard and one that is not, with a reason for each. + +A good answer sorts the pattern into parts that apply and parts that do not for my stated tool, says plainly where the tool's exit-code behavior should be checked and does not assume it, and names one rule that suits a guard and one that does not, with a reason for each. +``` + +Answer: + +## What applies to your Claude Code setup + +Nearly all of this pattern applies directly: + +- **Core mechanism:** PreToolUse and PostToolUse hooks exist in Claude Code and match the document's behavior. +- **Exit codes:** The document is written for Claude Code and explicitly describes the exit-code behavior you have: code 2 blocks, code 1 does not. This is already verified for your tool. +- **Paired guard/marker pattern:** The session-id keying, JSON input format, tool matching (Bash, Edit, Read, Write), and hook registration in `~/.claude/settings.json` are all Claude Code features shown in the example. +- **The two guard shapes:** Read-first and write-back both work in your setup. + +## Where not to assume + +The document's exit-code claims are stated as Claude Code specific ("In Claude Code, exit code 2 is the one that blocks"). You don't need to verify those—they're already tailored to your tool. The hook event names and JSON format in the examples are also Claude Code native, not a guess. + +## One rule that suits a guard, one that doesn't + +**Good candidate: "Never edit files under migrations/ without reading docs/MIGRATIONS.md"** + +Why: This is destructive (migrations affect the whole team), the condition is mechanically checkable (read happened or not), and it has a clear release path (read the file, marker clears the block). A missed read here is genuinely costly. This matches the document's advice: "destructive, irreversible, or touches state another session or process shares" and "mechanically checkable." + +**Not a good candidate: "Keep commit messages short"** + +Why: The pattern says "The condition needs judgment about content, not just 'did an event occur.'" Enforcing message length requires deciding what "short" means—that's a judgment call, not a checkable event. A guard could block all commits until you manually edit the message, but that's busywork, not reliable enforcement. + +(The test-before-merge rule is borderline; it *could* be guarded, but a failed merge is easier to undo than a broken migration, so it's lower urgency than migrations.) diff --git a/patterns/tests/prompt-runs/block-and-tell-hooks-review-sonnet.md b/patterns/tests/prompt-runs/block-and-tell-hooks-review-sonnet.md new file mode 100644 index 0000000..f927f0c --- /dev/null +++ b/patterns/tests/prompt-runs/block-and-tell-hooks-review-sonnet.md @@ -0,0 +1,56 @@ +# block-and-tell-hooks prompt 2 (Review it against your setup) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Below is a description of my own agent setup (tool, hook support, rules I currently rely on): I use Claude Code. Hooks are enabled in ~/.claude/settings.json. Today my rules are prose in a CLAUDE.md file: always run the tests before a merge, never edit files under migrations/ without reading docs/MIGRATIONS.md, and keep commit messages short. + +List which parts of the pattern apply to my setup and which do not. Flag any claim that may be false for my tool, for example how it treats exit codes. Name one rule of mine that is a good candidate for a guard and one that is not, with a reason for each. + +A good answer sorts the pattern into parts that apply and parts that do not for my stated tool, says plainly where the tool's exit-code behavior should be checked and does not assume it, and names one rule that suits a guard and one that does not, with a reason for each. +``` + +Answer: + +## Applies to your setup + +- **Paired guard and marker.** This fits Claude Code's `PreToolUse` and `PostToolUse` events. Your `~/.claude/settings.json` is the right place to register them. +- **Settings shape.** Event name as the key, a `matcher` string, and a `hooks` list of commands. This matches how Claude Code configures hooks. +- **Session-keyed marker.** `session_id` arrives in the JSON on stdin, so per-session state files work. +- **Block-and-tell.** Stderr from an exit-2 `PreToolUse` hook is surfaced to the agent, so the message-as-fix idea holds. +- **Fast path, release-path-first testing, and escape hatches.** These are tool-agnostic good practice. + +## Does not apply, or needs adjusting + +- **Write-back guards.** None of your three rules needs one. +- **The deploy example.** It is a template. Your matchers and paths will differ. +- **Cross-session shared state.** Your rules don't involve it, so that motivation is weaker for you. + +## Claims to verify against your installed version + +- **Exit codes.** Exit 2 blocks and exit 1 does not, per the Claude Code docs as I know them. Don't take my word or the document's for it. Write a throwaway `PreToolUse` hook that does `exit 1`, then another that does `exit 2`, and confirm that only the second stops the call. +- **"Hooks cannot rewrite a tool call's input."** This may be outdated. Newer Claude Code versions support structured JSON hook output, including a permission decision and possibly input modification. Check the current hooks docs. The block-and-print-the-corrected-value approach still works regardless. +- **PostToolUse timing.** The marker only runs after the tool succeeds, which is what you want. Exit 2 there cannot undo the call, so don't put enforcement in it. +- **Tool coverage.** The document says to list every tool that can touch the file. Beyond `Edit|Write`, check whether `MultiEdit` or `NotebookEdit` exist in your version. Also, `Bash` can modify files (`sed -i`, `>`, `tee`), so an `Edit|Write`-only guard has a hole. Decide whether to match `Bash` commands mentioning `migrations/` too, and watch for the false-block problem from sharp edge 3. + +## Your rules + +**Good candidate: "never edit files under migrations/ without reading docs/MIGRATIONS.md."** +- It is mechanically checkable: either a `Read` of that file happened this session or it didn't. +- A skip is costly, since a bad migration is hard to undo. +- Guard: `PreToolUse` on `Edit|Write`, which exits 2 if `file_path` contains `/migrations/` and the marker is absent. The stderr message should say "Read docs/MIGRATIONS.md, then retry." +- Marker: `PostToolUse` on `Read` that touches the session marker when the path ends in `docs/MIGRATIONS.md`. +- The release path is trivial, which is what you want. + +**Poor candidate: "keep commit messages short."** +- "Short" is a judgment call, not an event. The pattern's own "when not to" section excludes this. +- You'd have to parse `git commit -m` out of arbitrary Bash, including heredocs and `-F`, which invites false blocks. +- The cost of a skip is trivial and easy to fix. Leave it in `CLAUDE.md`. + +**In between: "always run the tests before a merge."** +- It is checkable, but "tests ran" is weaker than "tests passed." A marker keyed on a `Bash` call containing your test command would be satisfied by a failing run unless you also inspect the tool's result. +- It is also stale if code changes after the run. Consider adding a guard on `git merge`/`gh pr merge` only after you've confirmed the exit-code behavior and settled what counts as passing. + +I'd start with the migrations guard, since it exercises the whole pattern with the least risk. diff --git a/patterns/tests/prompt-runs/block-and-tell-hooks-teach-haiku.md b/patterns/tests/prompt-runs/block-and-tell-hooks-teach-haiku.md new file mode 100644 index 0000000..5ea58f3 --- /dev/null +++ b/patterns/tests/prompt-runs/block-and-tell-hooks-teach-haiku.md @@ -0,0 +1,32 @@ +# block-and-tell-hooks prompt 1 (Understand and teach it) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Explain it to me as if I know what a script is but have never used hooks. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why exit code 2 matters and why a guard needs a release path. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, covers why exit 2 blocks and exit 1 does not, and explains why a guard with no release path locks the agent out. It asks exactly three questions, one at a time, and waits for my answer before the next. +``` + +Answer: + +# Understanding Block-and-Tell Hooks + +Let me explain this with a security checkpoint analogy. Imagine a building with a secure room. Before you enter to make a change, a guard at the door checks: "Have you read the current status board?" If no, the guard says "not yet—read it first, here's why it matters" and **blocks your entry**. Once you've read the board, you get a stamped pass. On your next visit that session, the guard sees the pass and lets you through. + +The key insight: the guard can only **truly block** if they use a specific exit code (2 in Claude Code). Using the wrong code (like 1) is like a guard who steps aside anyway—looks official but doesn't actually stop you. And the guard's job is pointless if there's no way to get the stamp and enter. You'd be locked out forever. + +Here's what happens technically: + +1. **PreToolUse hook** (the guard): Checks if a marker file exists for this session. If not → exits with code 2 and explains what to do. If yes → exits 0, lets the tool run. +2. **PostToolUse hook** (the stamp machine): After you've read or edited the file, writes a marker file for this session. + +The two hooks are a **matched pair**—one enforces, one releases. + +--- + +**Question 1 of 3:** + +In the security checkpoint analogy, why would it be a problem if the guard used a different signal instead of exit code 2? What would change about whether the tool call actually gets blocked? diff --git a/patterns/tests/prompt-runs/block-and-tell-hooks-teach-sonnet.md b/patterns/tests/prompt-runs/block-and-tell-hooks-teach-sonnet.md new file mode 100644 index 0000000..19e3a3f --- /dev/null +++ b/patterns/tests/prompt-runs/block-and-tell-hooks-teach-sonnet.md @@ -0,0 +1,48 @@ +# block-and-tell-hooks prompt 1 (Understand and teach it) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Explain it to me as if I know what a script is but have never used hooks. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why exit code 2 matters and why a guard needs a release path. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, covers why exit 2 blocks and exit 1 does not, and explains why a guard with no release path locks the agent out. It asks exactly three questions, one at a time, and waits for my answer before the next. +``` + +Answer: + +## The idea in plain terms + +Imagine a concert venue with a wristband rule: **you can't enter the VIP area until you've picked up a wristband at the ticket booth.** + +You could put up a sign saying "Please get a wristband first." Most people would follow it, but some would forget or walk past it. That's the situation with an AI agent and a written rule. + +A **hook** is a small script the harness runs automatically at certain moments, outside the model's judgment. There are two here: + +- **The door guard (`PreToolUse`)** runs just before the agent does something, like editing a file or running a command. It checks for a wristband. If there isn't one, it stops the agent and says, "Go to the ticket booth and get one." +- **The booth (`PostToolUse`)** runs just after the agent does the thing you wanted, like reading a particular file. It slips a wristband on, which here is a small file saved on disk and labeled with the session ID. + +The guard without the booth is a door nobody can ever get through. The booth without the guard is a free wristband nobody checks. + +## Why exit code 2 matters + +When a hook script finishes, it reports a number called an exit code. Claude Code reads that number like a guard's signal: + +- **`0`** means all clear, so the action goes ahead. +- **`2`** means stop. The action does not happen, and the guard's written message (stderr) is shown to the agent. +- **Anything else, including `1`**, means the guard itself malfunctioned. Claude Code logs an error but lets the action through anyway. + +So a guard that says `exit 1` looks like it's blocking, but the door stays open. The script is "refusing" while the system is waving everyone in. + +## Why the message matters + +A guard that just says "No" leaves the agent stuck. A good guard says what's missing and how to fix it, such as "Read CHANGELOG.md, then retry." The agent can then fix it in the same turn without bothering you. + +## Why a guard needs a release path + +Suppose the wristband booth is inside the VIP area. Nobody can get a wristband without already being in, so nobody gets in. The same thing happens when a guard blocks something, but the action that would clear the block is also blocked, or the marker's trigger is too narrow to ever fire. The agent is locked out of that whole class of action with no fix. That's why you write and test the way out before the way in. + +## Question 1 of 3 + +A teammate writes a guard script that detects the missing marker, prints a clear "go read the changelog first" message to stderr, and ends with `exit 1`. They test it by watching the terminal and see the message appear. **Will the agent's deploy command actually be stopped? Why or why not?** diff --git a/patterns/tests/prompt-runs/rules-index-architecture-adapt-haiku.md b/patterns/tests/prompt-runs/rules-index-architecture-adapt-haiku.md new file mode 100644 index 0000000..41c959c --- /dev/null +++ b/patterns/tests/prompt-runs/rules-index-architecture-adapt-haiku.md @@ -0,0 +1,90 @@ +# rules-index-architecture prompt 3 (Adapt and test it) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +My instructions file is pasted below. My tool is: Claude Code. + +# Project rules + +## Git +Use short commit messages. Never push to main. + +## Testing +Run pytest before every merge. Add a test for every bug fix. + +## Deploys +Deploys go through make deploy. Check the changelog first. Never deploy on Fridays. + +## Style +Use type hints. Keep functions under 40 lines. + +## Old migration notes +The 2023 migration moved the billing tables. See docs/billing-2023.md for the details. + +Propose a split into domain files with an index table, and tell me which files should be excluded from startup loading. Then give me a way to test that the split worked: how to measure total loaded size before and after, and one question to ask the agent that only a specific rule file can answer. + +A good answer proposes domain files and an index table built from my file's own contents, names which files to exclude from startup loading, and gives a before-and-after way to measure loaded size plus one question that only a specific rule file can answer. +``` + +Answer: + +I'll propose a split for your small ruleset. Even though you only have five domains, splitting makes sense here because each is distinct—a reader won't have to scan past deploy rules to find the style guide. + +## Proposed structure + +``` +rules/ + README.md <- index + git-workflow.md + testing-requirements.md + deploy-process.md + code-style.md + reference/ + billing-migration-2023.md <- excluded from startup +``` + +**`rules/README.md`** (index table) + +```markdown +# Rules Index + +| File | What it governs | +|---|---| +| git-workflow.md | Commit messages, branch protection, push safety | +| testing-requirements.md | Test coverage for bugs, pre-merge testing | +| deploy-process.md | Deployment workflow, timing restrictions | +| code-style.md | Type hints, function length limits | +``` + +**Exclude from startup:** `rules/reference/**` in your `settings.json`: +```json +{ + "claudeMdExcludes": ["rules/reference/**"] +} +``` + +The billing migration note is reference material you'll only need when working with that specific code, so it goes in `reference/` and stays out of every startup. + +## How to test it worked + +**Before/after measurement:** +```bash +# Before (current state) +wc -c CLAUDE.md + +# After (new state) +cat CLAUDE.md rules/*.md | wc -c +``` + +Both should be the same total—splitting doesn't shrink what loads. Then verify the exclude works: +```bash +cat CLAUDE.md rules/*.md | wc -c # with reference/ +cat CLAUDE.md rules/[^r]*.md | wc -c # without reference/ folder +``` + +The gap is what you saved by excluding. + +**Testing question:** Ask "What must I include when I fix a bug?" The answer should come only from `testing-requirements.md` and mention adding a test. If the agent answers correctly, the rule file was loaded and understood in isolation. diff --git a/patterns/tests/prompt-runs/rules-index-architecture-adapt-sonnet.md b/patterns/tests/prompt-runs/rules-index-architecture-adapt-sonnet.md new file mode 100644 index 0000000..f1a3c10 --- /dev/null +++ b/patterns/tests/prompt-runs/rules-index-architecture-adapt-sonnet.md @@ -0,0 +1,145 @@ +# rules-index-architecture prompt 3 (Adapt and test it) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +My instructions file is pasted below. My tool is: Claude Code. + +# Project rules + +## Git +Use short commit messages. Never push to main. + +## Testing +Run pytest before every merge. Add a test for every bug fix. + +## Deploys +Deploys go through make deploy. Check the changelog first. Never deploy on Fridays. + +## Style +Use type hints. Keep functions under 40 lines. + +## Old migration notes +The 2023 migration moved the billing tables. See docs/billing-2023.md for the details. + +Propose a split into domain files with an index table, and tell me which files should be excluded from startup loading. Then give me a way to test that the split worked: how to measure total loaded size before and after, and one question to ask the agent that only a specific rule file can answer. + +A good answer proposes domain files and an index table built from my file's own contents, names which files to exclude from startup loading, and gives a before-and-after way to measure loaded size plus one question that only a specific rule file can answer. +``` + +Answer: + +## Proposed layout + +Your file is about 15 lines, so the pattern's "when not to" section applies: five one-line files plus an index is more indirection than you need today. If you want the split anyway (for example because you expect these sections to grow), here is how I'd do it. + +``` +CLAUDE.md <- keeps only the pointers, plus the one rule you want every session +.claude/rules/ + README.md <- the index + git-workflow.md + testing.md + deploys.md + code-style.md <- has paths: frontmatter + reference/ <- excluded from startup loading + billing-2023-migration.md +``` + +**`rules/README.md`** + +```markdown +# Rules Index + +| File | What it governs | +|---|---| +| git-workflow.md | Commit message length, never pushing to main | +| testing.md | pytest before merge, a test for every bug fix | +| deploys.md | make deploy, changelog check, no Friday deploys | +| code-style.md | Type hints, 40-line function limit (Python files only) | +| reference/billing-2023-migration.md | Pointer to the 2023 billing-table migration notes (docs/billing-2023.md) | +``` + +**`rules/git-workflow.md`** +```markdown +# Git Workflow +- Use short commit messages. +- Never push to main. +``` + +**`rules/testing.md`** +```markdown +# Testing +- Run pytest before every merge. +- Add a test for every bug fix. +``` + +**`rules/deploys.md`** +```markdown +# Deploys +- Deploy only through `make deploy`. +- Check the changelog first. +- Never deploy on Fridays. +``` + +**`rules/code-style.md`** +```markdown +--- +paths: + - "**/*.py" +--- +# Code Style +- Use type hints. +- Keep functions under 40 lines. +``` + +**`rules/reference/billing-2023-migration.md`** +```markdown +# 2023 billing migration +The 2023 migration moved the billing tables. Details: docs/billing-2023.md +``` + +**`CLAUDE.md`** +```markdown +# Project rules +Index of all rules: .claude/rules/README.md +``` + +## What to exclude from startup loading + +1. **`reference/`**: the migration notes are history, not instructions. Exclude them in `settings.json`: + ```json + { "claudeMdExcludes": ["**/.claude/rules/reference/**"] } + ``` + The index row stays visible, so the agent knows the notes exist and can read them on demand. +2. **`code-style.md`**: it is scoped with `paths:` to `**/*.py`, so it loads only when Python files are in play. + +I would leave `git-workflow.md`, `testing.md`, and `deploys.md` always loaded. They aren't tied to file patterns and are tiny. + +`docs/billing-2023.md` is outside `rules/`, so it is not loaded at startup either way. The stub exists only so the index can point to it. + +The Friday deploy ban and the no-push-to-main rule are good candidates for a hook, since a prose rule is only a request. If you add one, note "hook-enforced" in the rule file. + +## How to test the split + +**Measure loaded size.** The document's claim that every file under `rules/` loads at startup is worth checking on your install rather than trusting. + +- Before the split, save your original file and run `wc -c CLAUDE.md`. +- After the split, run: + ```bash + cat CLAUDE.md .claude/rules/*.md | wc -c + ``` + This counts only top-level rule files. Add `.claude/rules/reference/*.md` to the command to see what the exclude is saving. +- To see what Claude Code actually loaded, start a fresh session and run `/memory` (lists loaded instruction files) or `/context` (shows memory-file token usage). The reference file should be absent from the list. Compare the numbers with and without the `claudeMdExcludes` entry. + +Expect the total to stay about the same or grow slightly from the index and frontmatter. At this size the split gives you findability, not savings. + +**One question only a specific file can answer.** In a fresh session, ask: + +> What day are deploys forbidden, and what must I check before running `make deploy`? + +The answer (Fridays, check the changelog) lives only in `deploys.md`. If the agent gets it right, the rule file loaded and the pointer chain works. Then check the opt-outs: + +- In a fresh session where you haven't touched any `.py` file, ask "What's the maximum function length?" It should not know, because `code-style.md` is path-scoped. Open a Python file and ask again, and it should say 40 lines. +- Ask "What did the 2023 migration move?" The agent should have to open the reference stub and `docs/billing-2023.md` to answer, which confirms the exclude kept it out of startup context. diff --git a/patterns/tests/prompt-runs/rules-index-architecture-review-haiku.md b/patterns/tests/prompt-runs/rules-index-architecture-review-haiku.md new file mode 100644 index 0000000..b2636d4 --- /dev/null +++ b/patterns/tests/prompt-runs/rules-index-architecture-review-haiku.md @@ -0,0 +1,46 @@ +# rules-index-architecture prompt 2 (Review it against your setup) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Below is the folder listing and approximate line counts of my own instruction files, plus the agent tool I use: Tool: Claude Code. +CLAUDE.md 410 lines +rules/git.md 90 lines +rules/testing.md 60 lines +rules/deploy.md 120 lines +rules/old-notes.md 300 lines (rarely needed) + +Tell me whether my setup needs this pattern yet. Flag any claim in the document that may not be true for my tool, especially about what gets loaded at startup. Suggest which of my files should be split out, which should stay, and which belong in a reference folder, and give a reason for each. + +A good answer judges from my listing whether the pattern is needed yet, says which startup-loading claims should be checked for my tool and does not assume them, and sorts my files into split out, stay, and reference folder with a reason for each. +``` + +Answer: + +Your setup is close to the pattern but not quite there yet. You're at ~980 lines spread across five files—the threshold where an index becomes useful, and where that 300-line rarely-needed file is dragging down startup load. + +**Does your setup need this pattern?** Yes, but just barely. You've already felt the pain enough to split into domains (good instinct), but you're missing the index and the exclusion mechanism to complete it. The payoff is immediate: removing old-notes.md from startup load could save ~15–20% of your instruction context at session start. + +**Verify this claim about Claude Code:** The document states that Claude Code loads CLAUDE.md plus *every* markdown file under `rules/` recursively at startup, with no default lazy loading. Before you reorganize, test this: ask the agent a specific detail from `old-notes.md` (something it shouldn't know unless it was loaded). If it answers correctly, the claim holds for your version. + +**What to do with each file:** + +| File | Action | Why | +|---|---|---| +| CLAUDE.md | Stay, but trim to ~150 lines | Keep only rules that apply every session. Move domain-specific stuff to the rules files where it belongs. | +| rules/git.md | Stay as-is | 90 lines is lean and focused. | +| rules/testing.md | Stay as-is | 60 lines is already minimal. | +| rules/deploy.md | Stay as-is | 120 lines is on the edge but specific enough to stay. | +| rules/old-notes.md | Move to rules/reference/old-notes.md, exclude from startup | 300 rarely-used lines is exactly what `claudeMdExcludes` is for. You can still consult it during a session if needed. | +| (new) rules/README.md | Add | One table: file name, one-line description. Five rows, takes 30 seconds. | + +**Settings:** Add to your `settings.json`: +```json +{ + "claudeMdExcludes": ["rules/reference/**"] +} +``` + +Then test: before excluding, count characters in your rule files. After the exclude, ask the agent a question only something in old-notes.md can answer—it shouldn't know it anymore. diff --git a/patterns/tests/prompt-runs/rules-index-architecture-review-sonnet.md b/patterns/tests/prompt-runs/rules-index-architecture-review-sonnet.md new file mode 100644 index 0000000..a37c4e4 --- /dev/null +++ b/patterns/tests/prompt-runs/rules-index-architecture-review-sonnet.md @@ -0,0 +1,58 @@ +# rules-index-architecture prompt 2 (Review it against your setup) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Below is the folder listing and approximate line counts of my own instruction files, plus the agent tool I use: Tool: Claude Code. +CLAUDE.md 410 lines +rules/git.md 90 lines +rules/testing.md 60 lines +rules/deploy.md 120 lines +rules/old-notes.md 300 lines (rarely needed) + +Tell me whether my setup needs this pattern yet. Flag any claim in the document that may not be true for my tool, especially about what gets loaded at startup. Suggest which of my files should be split out, which should stay, and which belong in a reference folder, and give a reason for each. + +A good answer judges from my listing whether the pattern is needed yet, says which startup-loading claims should be checked for my tool and does not assume them, and sorts my files into split out, stay, and reference folder with a reason for each. +``` + +Answer: + +**Short answer:** yes, but mainly for `CLAUDE.md`. Your `rules/` folder is already split by domain. The 410-line `CLAUDE.md` and the 300-line `old-notes.md` are the real problems. + +## Does your setup need it? + +- **`CLAUDE.md` (410 lines):** This is past the point where the document says to split, since you're scrolling past unrelated rules to find one. It's likely a mix of domains. +- **`rules/` (four files):** These are already single-topic. An index for four files is more indirection than it's worth. Add `rules/README.md` once the `CLAUDE.md` split brings you to around 8 or more files. +- **Context cost:** Total is about 980 lines. If everything loads at startup, `old-notes.md` is 300 lines of mostly dead weight every session. + +## Claims to check for your setup + +I haven't inspected your install, so verify these rather than assuming them: + +1. **Where `rules/` lives.** The document's exclude example uses `.claude/rules/`. If your folder is a plain `rules/` at the project root, it may not auto-load at all. Check whether it's under `.claude/`. +2. **"Loads every `.md` under `rules/` recursively."** The document says this was confirmed in one installed version. Confirm it in yours by asking the agent a question only `old-notes.md` can answer, in a fresh session. +3. **Pointers in `CLAUDE.md`.** A line like `Full rules: rules/git.md` is plain text, not an import. The agent reads the file only if it decides to. If you want guaranteed loading, check whether your version supports `@path` imports. +4. **`paths:` frontmatter and `claudeMdExcludes`.** Check that your version supports both and what glob syntax it expects. Test each on one file and compare loaded size. +5. **The character counts (150k to 177k).** These are from the author's setup and don't transfer. Measure your own with `wc -c CLAUDE.md rules/*.md` before and after. + +## File sorting + +| File | Verdict | Reason | +|---|---|---| +| `CLAUDE.md` (410) | **Split out, keep a core** | Keep only what's needed every session, plus pointers. Move each domain section into its own `rules/*.md`. Merge into `git.md`, `testing.md`, or `deploy.md` where the topic matches, so you don't duplicate rules. | +| `rules/git.md` (90) | **Stay** | Already one topic and a reasonable size. | +| `rules/testing.md` (60) | **Stay**, maybe add `paths:` | Add `paths:` for your test globs if it only matters when editing tests. | +| `rules/deploy.md` (120) | **Stay, trim**, maybe add `paths:` | Keep the rules here. Move long command listings and incident history to `reference/`. Consider `paths:` if it only applies to deploy or CI files. | +| `rules/old-notes.md` (300) | **Reference folder** | Move it to `rules/reference/`, split by topic if you can, and add `**/rules/reference/**` to `claudeMdExcludes`. If it's truly stale, archive or delete it instead (look at it first). | + +## Suggested order + +1. Measure the current load. +2. Move `old-notes.md` to `reference/` and add the exclude. This is the cheapest win. +3. Split `CLAUDE.md` by domain. +4. Add the index once you have 8 or more files, and add its row in the same change as each new file. +5. Re-measure. + +If you paste `CLAUDE.md`, I can propose the actual split and index table. diff --git a/patterns/tests/prompt-runs/rules-index-architecture-teach-haiku.md b/patterns/tests/prompt-runs/rules-index-architecture-teach-haiku.md new file mode 100644 index 0000000..02125ff --- /dev/null +++ b/patterns/tests/prompt-runs/rules-index-architecture-teach-haiku.md @@ -0,0 +1,27 @@ +# rules-index-architecture prompt 1 (Understand and teach it) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Explain it to me as if I have one long instructions file and have never split it. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why splitting files does not shrink what is loaded, and what the two opt-outs are. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, says that splitting a file does not shrink what is loaded at startup, and names both opt-outs. It asks exactly three questions, one at a time, and waits for my answer before the next. +``` + +Answer: + +## The Pattern Explained + +Imagine your kitchen has one enormous junk drawer where you keep everything: scissors, recipe cards, batteries, instruction manuals, takeout menus, rubber bands, expired coupons. When you need scissors, you have to dig through the whole drawer. When you add a new item, you just shove it in wherever. And if you try to find the scissors again next week, you're digging again. + +The rules-index pattern is like giving each category its own drawer: one for cooking tools, one for instruction manuals, one for supplies. Now when you need scissors, you know exactly which drawer to open. A new rule about cooking goes in the cooking drawer, not crammed between the batteries and rubber bands. + +**But here's the key:** You still own the same amount of stuff. Moving scissors from the junk drawer to the cooking drawer didn't make the scissors disappear. In Claude Code, splitting one big `CLAUDE.md` into thirty smaller rule files doesn't shrink what gets loaded into memory at startup—it all still loads, you've just organized it so it's easier to find and edit. + +To actually shrink what loads, you need one of two escape hatches: either mark a rule file to load only for certain files in your project (`paths:` frontmatter), or tell Claude Code to skip a whole folder via `claudeMdExcludes` in settings. + +--- + +**Question 1:** If I take my 5,000-line `CLAUDE.md` file and split it into ten 500-line rule files in a `rules/` folder, what happens to the total amount of instructions Claude Code loads at startup? diff --git a/patterns/tests/prompt-runs/rules-index-architecture-teach-sonnet.md b/patterns/tests/prompt-runs/rules-index-architecture-teach-sonnet.md new file mode 100644 index 0000000..b5f34eb --- /dev/null +++ b/patterns/tests/prompt-runs/rules-index-architecture-teach-sonnet.md @@ -0,0 +1,32 @@ +# rules-index-architecture prompt 1 (Understand and teach it) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Explain it to me as if I have one long instructions file and have never split it. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand why splitting files does not shrink what is loaded, and what the two opt-outs are. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, says that splitting a file does not shrink what is loaded at startup, and names both opt-outs. It asks exactly three questions, one at a time, and waits for my answer before the next. +``` + +Answer: + +## The idea, using a different analogy + +Imagine you keep every recipe you own in one giant stained notebook. Appetizers, desserts, your grandmother's soup, and a note about not using the blue pan are all jumbled together. Finding anything means flipping through pages, and when you add a new recipe you squeeze it into whatever blank space is nearest. Eventually the two banana bread recipes are forty pages apart, and a correction you scribble for one smudges into the recipe next to it. + +The fix is to move to a recipe box with dividers. Each recipe gets its own card, cards are grouped by type, and a short table of contents at the front tells you which divider to look under. Now you can find things fast, add a card without disturbing its neighbors, and edit one without touching the rest. + +Here is the catch, and it's the part people get wrong. Imagine that every morning, before you start cooking, you are required to read **every card in the box aloud**. The box is better organized, but the reading takes just as long as it did with the notebook. If you add a drawer of long, detailed cards, the reading takes even longer. + +That's how Claude Code behaves. At startup it loads `CLAUDE.md` plus every markdown file under `rules/`, including subfolders. Splitting makes your rules easier to find and edit, but it doesn't make the startup load smaller. There are two ways to actually cut it: + +1. Give a rule file `paths:` frontmatter, so it only loads when you're working on matching files. +2. Add a `claudeMdExcludes` entry in `settings.json`, so a whole folder (for example, a `reference/` folder of long history and examples) is skipped at startup. + +## Check your understanding + +**Question 1 of 3:** You take your single 2,000-line instructions file and split it into 30 smaller files under `rules/`, changing nothing else. Roughly how does the amount of text loaded at startup change, and why? + +I'll wait for your answer before asking the next one. diff --git a/patterns/tests/prompt-runs/typed-memory-system-adapt-haiku.md b/patterns/tests/prompt-runs/typed-memory-system-adapt-haiku.md new file mode 100644 index 0000000..8b42239 --- /dev/null +++ b/patterns/tests/prompt-runs/typed-memory-system-adapt-haiku.md @@ -0,0 +1,178 @@ +# typed-memory-system prompt 3 (Adapt and test it) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +My tool is: Claude Code. Here are five things I want it to remember: 1. I am a backend engineer who knows Go but not React. 2. Do not mock the database in integration tests; we got burned when mocks passed and the real migration failed. 3. Yes, the single bundled PR was the right call for the refactor. 4. The billing rewrite must finish before 2026-12-01 because of an audit. 5. Pipeline bugs are tracked in the Ops board in the issue tracker. + +Sort each into a memory type, write the topic files and the index for them, and then write a short script that checks the index against a size limit in bytes. Finish with a test: how I check the next session actually loaded the index, and how I confirm the size check fails when I add too much. + +A good answer sorts each of my five items into a memory type, writes topic files and an index for them, gives a script that measures the index in bytes, and finishes with a test that the next session loaded the index and a test that the size check fails when the index grows too large. +``` + +Answer: + +# Sorting your five items into memory types + +1. **Backend engineer, knows Go not React** → `user` +2. **Don't mock the database in integration tests** → `feedback` (a correction/lesson learned) +3. **Single bundled PR was right for the refactor** → `feedback` (a confirmation) +4. **Billing rewrite deadline: 2026-12-01, audit-driven** → `project` +5. **Pipeline bugs tracked in Ops board** → `reference` + +--- + +# Topic files + +I'll create these four files for you: + +**feedback_no_db_mocks.md** +```markdown +--- +name: feedback-no-db-mocks +description: don't mock the database in integration tests +metadata: + type: feedback +--- +Do not mock the database in integration tests. + +**Why:** mocks passed but the real migration failed, catching the bug only after deploy. + +**How to apply:** write integration tests against a real (test) database. Let the migration run. Mocks hide real database issues. +``` + +**feedback_bundled_pr_refactor.md** +```markdown +--- +name: feedback-bundled-pr-refactor +description: single bundled PR was the right call for the refactor +metadata: + type: feedback +--- +The single bundled PR approach for the refactor was correct. + +**Why:** it kept related changes visible together and made the review simpler than breaking it into many small PRs. + +**How to apply:** use this pattern again for similar refactors where the changes are interdependent. +``` + +**user_backend_engineer.md** +```markdown +--- +name: user-backend-engineer +description: backend engineer, Go expertise, no React +metadata: + type: user +--- +Backend engineer. Expertise in Go. Does not know React; steer away from frontend tasks or pair on them. +``` + +**project_billing_rewrite.md** +```markdown +--- +name: project-billing-rewrite +description: deadline 2026-12-01, audit-driven +metadata: + type: project +--- +Billing rewrite must finish by **2026-12-01** (audit deadline). Blocks other work. No slip room. +``` + +**reference_pipeline_bugs.md** +```markdown +--- +name: reference-pipeline-bugs +description: pipeline bugs tracked in Ops board +metadata: + type: reference +--- +Pipeline bugs are tracked in the Ops board in the issue tracker. Not in a separate system. Check there first. +``` + +--- + +# Memory index + +Create a file called `MEMORY.md` in a `.claude-memory` folder (or wherever you keep your memory): + +```markdown +# Memory Index + +## User +- [Backend engineer](user_backend_engineer.md) Go expertise, no React + +## Feedback +- [No DB mocks](feedback_no_db_mocks.md) mocks passed, real migration failed +- [Bundled PR refactor](feedback_bundled_pr_refactor.md) single PR was correct approach + +## Project +- [Billing rewrite](project_billing_rewrite.md) deadline 2026-12-01, audit-driven + +## Reference +- [Pipeline bugs location](reference_pipeline_bugs.md) Ops board in issue tracker +``` + +--- + +# Size check script + +Create `memory-size-check.sh`: + +```bash +#!/bin/bash +set -e + +MEMORY_FILE=".claude-memory/MEMORY.md" +BYTE_LIMIT=25000 +WARNING_THRESHOLD=20000 # 80% of limit + +if [[ ! -f "$MEMORY_FILE" ]]; then + echo "Error: $MEMORY_FILE not found" + exit 1 +fi + +BYTES=$(wc -c < "$MEMORY_FILE") +LINES=$(wc -l < "$MEMORY_FILE") + +echo "Memory index: $BYTES bytes, $LINES lines" + +if [[ $BYTES -gt $LIMIT ]]; then + echo "⚠️ WARNING: $MEMORY_FILE exceeds $BYTE_LIMIT byte limit by $((BYTES - BYTE_LIMIT)) bytes" + exit 1 +elif [[ $BYTES -gt $WARNING_THRESHOLD ]]; then + echo "⚠️ CAUTION: $MEMORY_FILE is at $((BYTES * 100 / BYTE_LIMIT))% of limit" + exit 0 +else + echo "✓ Memory index is healthy" + exit 0 +fi +``` + +Make it executable: `chmod +x memory-size-check.sh` + +--- + +# Tests + +**Test 1: Verify the index loads in the next session** + +At the start of the next Claude Code session: +1. Ask me: "What do I know about your backend stack?" +2. I should reference Go, React avoidance, and the database mocking lesson without you restating it. +3. If I mention these facts without you re-explaining them, the index loaded. + +Alternatively, check the session logs: if the `.claude-memory/MEMORY.md` file was read, it appears in the context window summary. + +**Test 2: Confirm the size check fails when you exceed the limit** + +1. Add 5,000+ bytes of dummy content to `MEMORY.md` (push it past 20,000 bytes). +2. Run `./memory-size-check.sh` +3. It should print a WARNING and exit with code 1. + +Restore the file and run again—it should pass. + +--- + +**Current state:** Your index is ~700 bytes. You have plenty of room. Run the script weekly (add to cron: `0 22 * * 0 /path/to/memory-size-check.sh`) to catch drift early. diff --git a/patterns/tests/prompt-runs/typed-memory-system-adapt-sonnet.md b/patterns/tests/prompt-runs/typed-memory-system-adapt-sonnet.md new file mode 100644 index 0000000..39c9e02 --- /dev/null +++ b/patterns/tests/prompt-runs/typed-memory-system-adapt-sonnet.md @@ -0,0 +1,209 @@ +# typed-memory-system prompt 3 (Adapt and test it) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +My tool is: Claude Code. Here are five things I want it to remember: 1. I am a backend engineer who knows Go but not React. 2. Do not mock the database in integration tests; we got burned when mocks passed and the real migration failed. 3. Yes, the single bundled PR was the right call for the refactor. 4. The billing rewrite must finish before 2026-12-01 because of an audit. 5. Pipeline bugs are tracked in the Ops board in the issue tracker. + +Sort each into a memory type, write the topic files and the index for them, and then write a short script that checks the index against a size limit in bytes. Finish with a test: how I check the next session actually loaded the index, and how I confirm the size check fails when I add too much. + +A good answer sorts each of my five items into a memory type, writes topic files and an index for them, gives a script that measures the index in bytes, and finishes with a test that the next session loaded the index and a test that the size check fails when the index grows too large. +``` + +Answer: + +I sorted the five items, then wrote the files, the script and the tests. Nothing has been written to disk yet. + +## Sorting + +| # | Item | Type | Why | +|---|---|---|---| +| 1 | Backend engineer, knows Go, not React | `user` | Who you are and what you know | +| 2 | Don't mock the DB in integration tests | `feedback` | A correction, with the reason it hurt | +| 3 | Single bundled PR was right for the refactor | `feedback` | A confirmation of a non-obvious call | +| 4 | Billing rewrite due before 2026-12-01 (audit) | `project` | Ongoing state with an absolute date | +| 5 | Pipeline bugs live on the Ops board | `reference` | A pointer to where information lives | + +## Files + +Put these in your memory folder. For Claude Code that is usually `~/.claude/projects//memory/`, but check yours. + +`MEMORY.md` +```markdown +# Memory Index + +## User +- [Backend engineer, Go not React](user_background.md) strong Go, new to React: explain frontend concepts, skip backend basics + +## Feedback +- [No DB mocks in integration tests](feedback_no_db_mocks_integration.md) mocks passed, real migration failed +- [Bundled PR for refactors](feedback_bundled_refactor_pr.md) confirmed: one PR was right for the refactor + +## Project +- [Billing rewrite](project_billing_rewrite.md) must finish before 2026-12-01, audit + +## Reference +- [Pipeline bugs: Ops board](reference_pipeline_bugs.md) tracked on the Ops board in the issue tracker +``` + +`user_background.md` +```markdown +--- +name: user-background +description: user's role and expertise, for tuning explanations +metadata: + type: user +--- +Backend engineer. Fluent in Go; does not know React. +**How to apply:** don't explain Go or backend basics. When touching frontend code, explain React concepts (hooks, state, rendering) briefly and relate them to backend ideas where possible. +``` + +`feedback_no_db_mocks_integration.md` +```markdown +--- +name: feedback-no-db-mocks-integration +description: integration tests must hit a real database, not mocks +metadata: + type: feedback +--- +Do not mock the database in integration tests. +**Why:** mocked tests passed but the real migration failed. +**How to apply:** when writing or editing integration tests, use a real database and run real migrations. Mocks are fine only in unit tests. +``` + +`feedback_bundled_refactor_pr.md` +```markdown +--- +name: feedback-bundled-refactor-pr +description: confirmed that one bundled PR was the right shape for the refactor +metadata: + type: feedback +--- +A single bundled PR was the right call for the refactor. +**Why:** confirmed by the user after the fact. Splitting would have been the default instinct, so this is a validated non-obvious choice. +**How to apply:** for similar refactors where the pieces only make sense together, bundle them without asking again. This is specific to tightly coupled refactors, not all changes. +``` + +`project_billing_rewrite.md` +```markdown +--- +name: project-billing-rewrite +description: billing rewrite deadline driven by an audit +metadata: + type: project +--- +The billing rewrite must finish before 2026-12-01. +**Why:** an audit depends on it. +**How to apply:** prioritize billing-rewrite work, and flag anything that risks the date. Update this file as scope or status changes. +``` + +`reference_pipeline_bugs.md` +```markdown +--- +name: reference-pipeline-bugs +description: where pipeline bugs are tracked +metadata: + type: reference +--- +Pipeline bugs are tracked on the Ops board in the issue tracker. +**How to apply:** look there for existing bugs before investigating, and point to it when a pipeline bug needs filing. +``` + +## Size check script + +`memory-size-check.sh` +```bash +#!/usr/bin/env bash +# Usage: memory-size-check.sh [memory_dir] +# Limits are the values observed in Claude Code 2.1.263; re-check for your version. +set -u + +DIR="${1:-$HOME/.claude/projects/PROJECT/memory}" +INDEX="$DIR/MEMORY.md" +MAX_LINES="${MAX_LINES:-200}" +MAX_BYTES="${MAX_BYTES:-25000}" +WARN_LINES=$((MAX_LINES * 80 / 100)) +WARN_BYTES=$((MAX_BYTES * 80 / 100)) + +if [ ! -f "$INDEX" ]; then + echo "FAIL: index not found: $INDEX" + exit 2 +fi + +lines=$(wc -l < "$INDEX" | tr -d ' ') +bytes=$(wc -c < "$INDEX" | tr -d ' ') # bytes, not characters +status=0 + +echo "index: $lines lines (limit $MAX_LINES), $bytes bytes (limit $MAX_BYTES)" + +if [ "$lines" -ge "$MAX_LINES" ] || [ "$bytes" -ge "$MAX_BYTES" ]; then + echo "FAIL: index is at or over a limit; content past the cut is not loaded" + status=1 +elif [ "$lines" -ge "$WARN_LINES" ] || [ "$bytes" -ge "$WARN_BYTES" ]; then + echo "WARN: index is over 80% of a limit; compact now" + status=1 +fi + +# Every link in the index must resolve +while IFS= read -r f; do + [ -f "$DIR/$f" ] || { echo "FAIL: broken link: $f"; status=1; } +done < <(grep -o '](\([^)]*\.md\))' "$INDEX" | sed 's/^](//; s/)$//') + +# Every topic file should be linked from the index +for f in "$DIR"/*.md; do + b=$(basename "$f") + [ "$b" = "MEMORY.md" ] && continue + grep -q "]($b)" "$INDEX" || { echo "WARN: not in index: $b"; status=1; } +done + +[ "$status" -eq 0 ] && echo "OK" +exit "$status" +``` + +Run `chmod +x memory-size-check.sh` first. The weekly cron entry should go through a wrapper that logs every run, as the document describes. + +## Tests + +**1. Did the next session load the index?** Add a canary as the last line of `MEMORY.md`, because the end is what gets cut first: + +```markdown +- [Canary](reference_canary.md) the canary word is "heron-4471" +``` + +Create `reference_canary.md` with any content so the link resolves. Start a fresh session and ask: + +> Without opening any files, what is my background, what is the billing deadline, and what is the canary word? + +- **Pass:** you get Go but not React, 2026-12-01, and "heron-4471". Those come from the index gists alone, and the canary proves the tail was read. +- **Fail:** the canary is missing or a truncation warning appears. Check the size script's output. + +Delete the canary line and file afterwards. + +**2. Does the size check fail when the index grows too large?** Run this on a copy so your real memory stays untouched: + +```bash +T=$(mktemp -d) && cp ~/.claude/projects/PROJECT/memory/*.md "$T/" + +# Baseline: should print OK, exit 0 +./memory-size-check.sh "$T"; echo "exit=$?" + +# Line limit: pad past 160 lines (80% of 200) -> expect WARN, exit 1 +for i in $(seq 1 170); do echo "- [Pad $i](reference_pipeline_bugs.md) filler" >> "$T/MEMORY.md"; done +./memory-size-check.sh "$T"; echo "exit=$?" + +# Byte limit with few lines: multibyte padding -> expect WARN/FAIL, exit 1 +cp ~/.claude/projects/PROJECT/memory/MEMORY.md "$T/MEMORY.md" +for i in $(seq 1 40); do printf -- '- [Pad](reference_pipeline_bugs.md) %s\n' "$(printf 'é%.0s' $(seq 1 450))" >> "$T/MEMORY.md"; done +./memory-size-check.sh "$T"; echo "exit=$?" + +# Broken link: expect FAIL, exit 1 +echo "- [Ghost](missing.md) nope" >> "$T/MEMORY.md" +./memory-size-check.sh "$T"; echo "exit=$?" +``` + +- **Pass:** the baseline exits 0, and each of the other three exits 1 with the matching message. +- **Third case:** it stays well under 200 lines, so only the byte count can trip it. If it doesn't fire, the script is counting characters instead of bytes. + +One caveat: the 200-line and 25,000-byte limits come from your document's observations on Claude Code 2.1.263, and they weren't confirmed on later versions. The script reads them from `MAX_LINES` and `MAX_BYTES`, so you can change them without editing the script. diff --git a/patterns/tests/prompt-runs/typed-memory-system-review-haiku.md b/patterns/tests/prompt-runs/typed-memory-system-review-haiku.md new file mode 100644 index 0000000..430b05f --- /dev/null +++ b/patterns/tests/prompt-runs/typed-memory-system-review-haiku.md @@ -0,0 +1,64 @@ +# typed-memory-system prompt 2 (Review it against your setup) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Below is a listing of my memory or notes folder with file sizes, plus the tool I use: Tool: Claude Code. +MEMORY.md 182 lines, 21,400 bytes +user_role.md 12 lines +feedback_testing.md 20 lines +feedback_commits.md 15 lines +project_billing_rewrite.md 40 lines +reference_dashboards.md 8 lines + +Tell me which parts of the pattern my setup already has and which it lacks. Flag any number in the document (line limit, byte limit) that may be different for my tool and tell me how to find the real one. Name my biggest risk of hitting a size limit and one check I could automate. + +A good answer says which parts of the pattern my listing already has and which it lacks, flags the line and byte limits as values that may differ for my tool and says how to find the real ones, and names one biggest size risk and one check I could automate. +``` + +Answer: + +## What your setup already has + +✓ **Four memory types** — user, feedback (2 files), project, reference — matching the pattern exactly +✓ **Linked topic files** — individual files per topic instead of one growing dump +✓ **A bounded index** (MEMORY.md as the table of contents) + +## What you lack + +✗ **Frontmatter in topic files** — The pattern specifies `---name: type---` metadata in each file so the system can validate them programmatically +✗ **Organized index sections** — No visible grouping by type (## Feedback, ## Project, etc.) in MEMORY.md +✗ **Topic index tier** — No INDEX-*.md files yet for when the main index grows too large +✗ **Automated size checks** — No script running weekly to verify limits and warn before overflow + +## Limits to verify for your version + +The document cites **200 lines** and **25,000 bytes** for Claude Code 2.1.263. Those were not re-confirmed on version 2.1.287, so **test your own version**: + +- Add 50 lines of test content to MEMORY.md, save, start a new Claude Code session, and check if the index still loads in full +- Repeat until the load truncates with a warning +- Check `wc -l` and `wc -c` to measure both limits precisely + +## Your biggest risk: line limit imminent + +**MEMORY.md is at 182/200 lines (91% full).** You have ~18 lines of headroom. Once you add a handful of new topics, the index will truncate at session start and hide everything past the cutoff. This is the highest-priority risk because a full index at session start is how the agent reads your memory—when it gets cut off, the session has no access to what was cut. + +## One check to automate + +A weekly script (or one you run before adding entries) that: + +```bash +#!/bin/bash +lines=$(wc -l < MEMORY.md) +bytes=$(wc -c < MEMORY.md) +if [ $lines -gt 160 ] || [ $bytes -gt 20000 ]; then + echo "⚠️ MEMORY.md is at $lines lines, $bytes bytes (80% threshold: 160 lines, 20000 bytes)" + exit 1 +fi +``` + +This gives you a warning at 80% instead of hitting the wall at 100%. Run it before adding entries, or set it to a weekly cron. + +**Next step:** You're close enough to the limit that you should either test your exact limits now, or start compacting MEMORY.md into bundled format (combining short entries on one line) before you hit it. diff --git a/patterns/tests/prompt-runs/typed-memory-system-review-sonnet.md b/patterns/tests/prompt-runs/typed-memory-system-review-sonnet.md new file mode 100644 index 0000000..cfed161 --- /dev/null +++ b/patterns/tests/prompt-runs/typed-memory-system-review-sonnet.md @@ -0,0 +1,64 @@ +# typed-memory-system prompt 2 (Review it against your setup) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Below is a listing of my memory or notes folder with file sizes, plus the tool I use: Tool: Claude Code. +MEMORY.md 182 lines, 21,400 bytes +user_role.md 12 lines +feedback_testing.md 20 lines +feedback_commits.md 15 lines +project_billing_rewrite.md 40 lines +reference_dashboards.md 8 lines + +Tell me which parts of the pattern my setup already has and which it lacks. Flag any number in the document (line limit, byte limit) that may be different for my tool and tell me how to find the real one. Name my biggest risk of hitting a size limit and one check I could automate. + +A good answer says which parts of the pattern my listing already has and which it lacks, flags the line and byte limits as values that may differ for my tool and says how to find the real ones, and names one biggest size risk and one check I could automate. +``` + +Answer: + +## Already in place + +- **Types:** the filename prefixes (`user_`, `feedback_`, `project_`, `reference_`) match the four types. I can't confirm the `metadata: type:` frontmatter from a listing. +- **One small file per topic:** the largest is 40 lines, so topic files aren't a concern. +- **Separated reference and project state:** `reference_dashboards.md` and `project_billing_rewrite.md` are their own files. + +## Missing or unverifiable + +- **The index is far larger than its files.** `MEMORY.md` has 182 lines but only 5 topic files are listed. An index should be about one line per file, so either the listing is partial or `MEMORY.md` holds content directly. That is the "ever-growing dump" failure the pattern describes. At about 118 bytes per line, the entries look like content rather than pointers. +- **No second tier:** no `INDEX-*.md` files, so overflow has nowhere to go except deletion. +- **No automated check:** nothing in the listing suggests a weekly size check, a wrapper that logs every run, or an orphan-file check. +- **No compaction contract:** I see no written rules for what is safe to drop or merge. +- **Few feedback files:** I see only `feedback_testing` and `feedback_commits`. I can't tell whether confirmations ("that approach was right") are captured alongside corrections. + +## Numbers to verify + +The document's **200 lines** and **25,000 bytes** were observed in Claude Code 2.1.263 only. The author could not re-confirm them on 2.1.287, so treat them as unverified for your version. Two ways to find the real values: + +1. **Empirical (more reliable):** in a scratch project, make a `MEMORY.md` of 300 short lines, then another that is over 30,000 bytes in fewer than 200 lines. Start a session in each. The harness should show a truncation warning naming the file and the limit that was hit. +2. **Static:** run `claude --version`, then search the installed program for `25000` and `200` near memory-loading code. That check failed for the document's author on 2.1.287, so don't rely on it alone. + +Also, the topic-file limit is untested in the document. + +## Biggest risk + +**Line count.** You're at 182 of 200 lines (91%), with 18 lines of headroom. Bytes are at 21,400 of 25,000 (86%). Both are past the document's 80% warning thresholds (160 lines, 20,000 bytes), so you're already in the "compact on purpose" zone. Lines will break first, and any entries past the cutoff won't load in that session. + +## One check to automate + +A weekly size check, measuring bytes with `wc -c`, not characters: + +```sh +#!/bin/sh +f="$HOME/.claude/projects//memory/MEMORY.md" # adjust path +lines=$(wc -l < "$f"); bytes=$(wc -c < "$f") +echo "$(date -u +%FT%TZ) lines=$lines bytes=$bytes" >> "$HOME/memory-check.log" +[ "$lines" -ge 160 ] || [ "$bytes" -ge 20000 ] && echo "MEMORY.md near limit: $lines lines, $bytes bytes" +``` + +Run it from cron through a wrapper that logs on every run and complains if the script is missing, so a job that never ran can't look like a clean result. + +The next step is to compact `MEMORY.md` into pointer lines and bundled entries, moving any inline content into topic files. Then add an `INDEX-*.md` tier, and don't delete anything. diff --git a/patterns/tests/prompt-runs/typed-memory-system-teach-haiku.md b/patterns/tests/prompt-runs/typed-memory-system-teach-haiku.md new file mode 100644 index 0000000..652dd2b --- /dev/null +++ b/patterns/tests/prompt-runs/typed-memory-system-teach-haiku.md @@ -0,0 +1,29 @@ +# typed-memory-system prompt 1 (Understand and teach it) on haiku + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Explain it to me as if I have never given an AI tool long-term memory. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand the difference between a correction and a confirmation, why the index has a size limit, and why compaction never deletes topic files. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, explains the difference between a correction and a confirmation, why the index has a size limit, and why compaction never deletes topic files. It asks exactly three questions, one at a time, and waits for my answer before the next. +``` + +Answer: + +I'll explain this pattern and check your understanding, one question at a time. + +## The idea, with a different analogy + +Imagine you're a journalist who writes articles across many topics. You keep a **desk notebook** that you open at the start of every assignment—it's your quick reference. In that notebook is a short index: "Recipe development (3 notes)", "Interview techniques (5 notes)", "Publishing deadlines (2 notes)". + +The actual detailed notes live in separate folders. You could stuff everything into the desk notebook instead, but then it gets too thick to flip open. Instead, you keep the index short and scannable by pointing to files elsewhere. Each file is typed by *why* you keep it: a correction you learned ("never rush an interview"), a confirmation ("open questions work better"), a project deadline, or a reference ("newsroom is in building 3"). + +The memory system works the same way: an index you read every session (must stay small), topic files you open as needed (the real content), and types that tell you *why* something matters. + +--- + +**Question 1 of 3:** + +What is the difference between recording a *correction* ("don't do X") and a *confirmation* ("yes, that approach was right")? Why does the document say you need both? diff --git a/patterns/tests/prompt-runs/typed-memory-system-teach-sonnet.md b/patterns/tests/prompt-runs/typed-memory-system-teach-sonnet.md new file mode 100644 index 0000000..db2ad9b --- /dev/null +++ b/patterns/tests/prompt-runs/typed-memory-system-teach-sonnet.md @@ -0,0 +1,42 @@ +# typed-memory-system prompt 1 (Understand and teach it) on sonnet + +Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples): + +```text +Here is a design pattern document: [PASTE FILE] + +Explain it to me as if I have never given an AI tool long-term memory. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand the difference between a correction and a confirmation, why the index has a size limit, and why compaction never deletes topic files. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, explains the difference between a correction and a confirmation, why the index has a size limit, and why compaction never deletes topic files. It asks exactly three questions, one at a time, and waits for my answer before the next. +``` + +Answer: + +# Long-term memory for an AI tool, explained with a kitchen + +Imagine a talented chef who wakes up every morning with no memory of yesterday. They are skilled, but they don't know your restaurant, your regulars, or what went wrong last week. Every conversation with an AI coding tool starts like that morning. + +**Long-term memory is the binder the chef reads before service.** Each day they read it, then work as if they remembered. + +- **Each page is one topic.** One page might say "Table 4's customer is allergic to shellfish." Another might say "Supplier changes deliveries on 2026-11-01." These are the topic files, and each has a *type* (who the customer is, what to do or avoid, what's in progress, where things are kept). +- **The cover sheet is the index.** It's one line per page: a title, where to find it, and a few words of gist. The chef reads the whole cover sheet first and opens a page only when needed. + +## Corrections vs. confirmations + +Suppose the chef over-salts a dish and you say, "Too salty, use less." That's a **correction**: it tells the chef what to stop doing. Later the chef tries something unusual, like a lemon finish on the fish, and you say, "Yes, exactly that, keep doing it." That's a **confirmation**: it tells the chef what to keep doing. + +If the binder only held corrections, the chef would grow timid. They'd avoid everything that ever got a complaint and never learn which risky choices were good. Confirmations stop tomorrow's chef from second-guessing a good call. + +## Why the cover sheet has a size limit + +The chef only gets the cover sheet in one glance before service starts, with no flipping back and forth. In this tool, the harness reads the index up to a fixed cutoff (in the document's version, 200 lines or 25,000 bytes, whichever comes first) and ignores the rest. If the cover sheet is too long, the bottom entries simply don't exist for that day. So the cover sheet stays short and the detail lives on the pages behind it. + +## Why compaction never deletes topic files + +When the cover sheet gets crowded, you tidy *the sheet*, not the binder. You merge several lines into one denser line or group related pages under a heading. The pages themselves stay, because deleting one destroys knowledge (the shellfish allergy) that you can't recover. A short pointer costs little, but a lost lesson means someone gets hurt again. + +--- + +Let's check your understanding, one question at a time. + +**Question 1:** Say the AI makes a mistake and you tell it, "Don't do that." Another day, it picks an unusual approach and you say, "Yes, that was right." What would the memory contain if it recorded only the first kind of feedback, and what problem would that cause in later sessions? diff --git a/patterns/tests/run_prompts.py b/patterns/tests/run_prompts.py new file mode 100644 index 0000000..2a72970 --- /dev/null +++ b/patterns/tests/run_prompts.py @@ -0,0 +1,76 @@ +#!/usr/bin/env python3 +"""Run the three "Prompt for your AI model" prompts of each pattern doc on Sonnet and Haiku. + +Same method as pipelines/docs-pipeline/tests/prompt-runs: `claude -p` with no tools, no +project instructions and no automatic memory, the doc pasted where the prompt says [PASTE FILE], the other bracketed +inputs replaced with made-up samples. Writes one record per doc, prompt and model into +prompt-runs/. Needs the `claude` command line tool; the normal test suite does not run this. + +Usage: python3 run_prompts.py [doc-stem ...] +""" +import os, re, subprocess, sys, tempfile +from concurrent.futures import ThreadPoolExecutor +from pathlib import Path + +HERE = Path(__file__).resolve().parent +PATTERNS = HERE.parent +OUT = HERE / "prompt-runs" +MODELS = ["sonnet", "haiku"] + +SAMPLES = { + "block-and-tell-hooks": { + 2: {"[DESCRIBE YOUR SETUP]": "I use Claude Code. Hooks are enabled in ~/.claude/settings.json. Today my rules are prose in a CLAUDE.md file: always run the tests before a merge, never edit files under migrations/ without reading docs/MIGRATIONS.md, and keep commit messages short."}, + 3: {"[YOUR \"ALWAYS DO X FIRST\" RULE]": "Before editing any file under migrations/, read docs/MIGRATIONS.md in the same session.", "[YOUR AGENT TOOL]": "Claude Code"}, + }, + "rules-index-architecture": { + 2: {"[PASTE LISTING AND TOOL NAME]": "Tool: Claude Code.\nCLAUDE.md 410 lines\nrules/git.md 90 lines\nrules/testing.md 60 lines\nrules/deploy.md 120 lines\nrules/old-notes.md 300 lines (rarely needed)"}, + 3: {"[YOUR AGENT TOOL]": "Claude Code", "[PASTE YOUR INSTRUCTIONS FILE]": "# Project rules\n\n## Git\nUse short commit messages. Never push to main.\n\n## Testing\nRun pytest before every merge. Add a test for every bug fix.\n\n## Deploys\nDeploys go through make deploy. Check the changelog first. Never deploy on Fridays.\n\n## Style\nUse type hints. Keep functions under 40 lines.\n\n## Old migration notes\nThe 2023 migration moved the billing tables. See docs/billing-2023.md for the details."}, + }, + "typed-memory-system": { + 2: {"[PASTE LISTING AND TOOL NAME]": "Tool: Claude Code.\nMEMORY.md 182 lines, 21,400 bytes\nuser_role.md 12 lines\nfeedback_testing.md 20 lines\nfeedback_commits.md 15 lines\nproject_billing_rewrite.md 40 lines\nreference_dashboards.md 8 lines"}, + 3: {"[YOUR AGENT TOOL]": "Claude Code", "[LIST FIVE FACTS, PREFERENCES AND CORRECTIONS]": "1. I am a backend engineer who knows Go but not React. 2. Do not mock the database in integration tests; we got burned when mocks passed and the real migration failed. 3. Yes, the single bundled PR was the right call for the refactor. 4. The billing rewrite must finish before 2026-12-01 because of an audit. 5. Pipeline bugs are tracked in the Ops board in the issue tracker."}, + }, +} + + +def extract_prompts(text): + sec = text.split("## Prompt for your AI model", 1)[1] + blocks = re.findall(r"\*\*(\d)\. ([^\n]*)\*\*\n\n```\n(.*?)\n```", sec, re.S) + return {int(n): (title, body) for n, title, body in blocks} + + +def run(stem, n, title, body, doc_text, model): + shown = body + for k, v in SAMPLES[stem].get(n, {}).items(): + shown = shown.replace(k, v) + sent = shown.replace("[PASTE FILE]", doc_text) + with tempfile.TemporaryDirectory() as scratch: + p = subprocess.run( + ["claude", "-p", "--model", model, "--tools", "", "--setting-sources", "project"], + input=sent, capture_output=True, text=True, cwd=scratch, timeout=600, + env={**os.environ, "CLAUDE_CODE_DISABLE_AUTO_MEMORY": "1"}) + ans = p.stdout.strip() if p.returncode == 0 else f"RUN FAILED (exit {p.returncode}): {p.stderr.strip()}" + kind = ["", "teach", "review", "adapt"][n] + rec = (f"# {stem} prompt {n} ({title}) on {model}\n\n" + "Prompt (the document's full text was pasted where it says [PASTE FILE]; other inputs filled with made-up samples):\n\n" + f"```text\n{shown}\n```\n\nAnswer:\n\n{ans}\n") + (OUT / f"{stem}-{kind}-{model}.md").write_text(rec) + return stem, kind, model, p.returncode + + +def main(): + OUT.mkdir(exist_ok=True) + stems = sys.argv[1:] or list(SAMPLES) + jobs = [] + with ThreadPoolExecutor(max_workers=6) as ex: + for stem in stems: + doc = (PATTERNS / f"{stem}.md").read_text() + for n, (title, body) in extract_prompts(doc).items(): + for m in MODELS: + jobs.append(ex.submit(run, stem, n, title, body, doc, m)) + for j in jobs: + print(*j.result()) + + +if __name__ == "__main__": + main() diff --git a/patterns/typed-memory-system.md b/patterns/typed-memory-system.md index 9a14150..81c3331 100644 --- a/patterns/typed-memory-system.md +++ b/patterns/typed-memory-system.md @@ -89,7 +89,7 @@ The agent reads the one that matches the task. The cost is that these are not au ### 5. A weekly automatic check -A size limit nobody checks gets crossed. Run a small script on a schedule that checks the index and every topic file against both limits, flags any topic file the index does not link to, and sends a message only when something needs a human. A weekly cron entry is enough: +A size limit nobody checks gets crossed. Run a small script on a schedule that checks the index against both limits (and keeps topic files short, since their limit is untested), flags any topic file the index does not link to, and sends a message only when something needs a human. A weekly cron entry is enough: ``` 0 22 * * 0 /path/to/memory-size-check.sh # Sundays, 10pm @@ -149,6 +149,8 @@ Give any AI model this file plus one of the prompts below. Paste the file text w Here is a design pattern document: [PASTE FILE] Explain it to me as if I have never given an AI tool long-term memory. Use a different everyday analogy than the one in the document. Then ask me three questions, one at a time, that check I understand the difference between a correction and a confirmation, why the index has a size limit, and why compaction never deletes topic files. Wait for my answer before each next question. + +A good answer uses an analogy that is not the document's, explains the difference between a correction and a confirmation, why the index has a size limit, and why compaction never deletes topic files. It asks exactly three questions, one at a time, and waits for my answer before the next. ``` **2. Review it against your setup** @@ -159,6 +161,8 @@ Here is a design pattern document: [PASTE FILE] Below is a listing of my memory or notes folder with file sizes, plus the tool I use: [PASTE LISTING AND TOOL NAME] Tell me which parts of the pattern my setup already has and which it lacks. Flag any number in the document (line limit, byte limit) that may be different for my tool and tell me how to find the real one. Name my biggest risk of hitting a size limit and one check I could automate. + +A good answer says which parts of the pattern my listing already has and which it lacks, flags the line and byte limits as values that may differ for my tool and says how to find the real ones, and names one biggest size risk and one check I could automate. ``` **3. Adapt and test it** @@ -169,6 +173,8 @@ Here is a design pattern document: [PASTE FILE] My tool is: [YOUR AGENT TOOL]. Here are five things I want it to remember: [LIST FIVE FACTS, PREFERENCES AND CORRECTIONS] Sort each into a memory type, write the topic files and the index for them, and then write a short script that checks the index against a size limit in bytes. Finish with a test: how I check the next session actually loaded the index, and how I confirm the size check fails when I add too much. + +A good answer sorts each of my five items into a memory type, writes topic files and an index for them, gives a script that measures the index in bytes, and finishes with a test that the next session loaded the index and a test that the size check fails when the index grows too large. ``` -**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically. I did not save those answers, so there is no record to read here, unlike the saved runs in `pipelines/job-assessment/tests/`. +**How these prompts were checked.** On 2026-10-01 each of the three prompts was run once on Claude Sonnet and once on Claude Haiku, with this file pasted in and sample details filled in. A Claude model (Sonnet 5.5) graded each answer against the "A good answer ..." sentence under the prompt. I have not re-read every answer. Sonnet met all three. Haiku met none of them in full: its explanation left out why compaction keeps topic files, it listed gaps it could not see from a file listing, and the size-check script it wrote uses a variable it never sets. The answers are in [`tests/prompt-runs/`](tests/prompt-runs/). diff --git a/pipelines/job-assessment/README.md b/pipelines/job-assessment/README.md index 8227f28..e0183fc 100644 --- a/pipelines/job-assessment/README.md +++ b/pipelines/job-assessment/README.md @@ -22,12 +22,13 @@ cd context-engineering-toolkit/pipelines/job-assessment python3 -m venv .venv && . .venv/bin/activate && pip install -r requirements.txt python3 -m pytest -q python3 scripts/validate_profile.py --fixture fixtures/robin-sample/career-profile.yaml +OUT=$(mktemp -d) python3 scripts/assess_offline.py fixtures/postings/03-unlisted-pay-perks.md \ --findings fixtures/findings/03-unlisted-pay-perks.findings.json \ - --profile fixtures/robin-sample/career-profile.yaml --out /tmp/ja-out + --profile fixtures/robin-sample/career-profile.yaml --out "$OUT" ``` -The tests pass. The validator prints `clean`. The last command runs the whole chain on fictional posting 03 and prints its verdict, **Apply with reservations**, with four score bars and a table of what it checked. Open the HTML card it wrote under `/tmp/ja-out/email/` in a browser to see the summary card. +The tests pass. The validator prints `clean`. The last command runs the whole chain on fictional posting 03 and prints its verdict, **Apply with reservations**, with four score bars and a table of what it checked. Open the HTML card it wrote under `$OUT/email/` in a browser to see the summary card. The fresh `mktemp` folder is what lets you run the last command again: a second run into the same folder stops with `DUPLICATE`. [`SETUP.md`](SETUP.md) has each step with what to expect, a fix for each common failure, and how to wire the skills into Claude Code. @@ -55,7 +56,7 @@ Built by directing Claude Code. The author designed the rules, reviewed the outp It is a public rebuild of a private tool the author has had in daily use since late August 2026. The earliest public trace is this repo's first job-screening commit, dated 2026-08-25 (`git log --reverse --format='%ad %s' --date=short | grep job-fit-screen`). That earlier tool has since been removed; git history keeps it. Everything personal was left out: no real postings, no real person's data, no real employers. The rules were rewritten from a written list of the private tool's behaviors, and the scoring and verdict were rebuilt as scripts so they can be tested from a clean clone. -**Timeline, so the dates make sense.** The public version was assembled and reviewed on 2026-10-01, from a private system in daily use since late August. The pull requests that built it (listed in `git log`) were built by Claude Code agents under the author's direction. They merged between 16:38 and 22:02 that day, several a minute apart, because they were prepared in parallel and merged after CI passed on each. The author's check on that work is the tests and saved run records in this folder. No claim is made here about how much of each diff a person read line by line. +**Timeline, so the dates make sense.** The public version was assembled and reviewed on 2026-10-01, from a private system in daily use since late August. The pull requests that built it (listed in `git log`) were built by Claude Code agents under the author's direction. They merged over the course of that day, several a minute apart in places, because they were prepared in parallel and merged after CI passed on each. The check on the code is the tests, and the CI that re-runs them in a fresh clone on every change. The check on the model's answers is the saved run records in this folder. No production machine-learning or retrieval (RAG) work is claimed. This is prompts, a schema, and ordinary scripts with tests. @@ -86,6 +87,8 @@ SKIPPED [1] tests/live/test_live_intake.py:38: live tests are off: set JA_LIVE=1 $ python3 scripts/validate_profile.py --fixture fixtures/robin-sample/career-profile.yaml validate_profile: clean, 0 warning(s) in career-profile.yaml $ python3 scripts/validate_profile.py fixtures/broken-profile.yaml; echo "exit=$?" +validate_profile: 5 problem(s), 0 warning(s) in broken-profile.yaml +(the 5 problem lines, one per planted mistake, are omitted here) exit=1 $ python3 scripts/assess_offline.py fixtures/postings/03-unlisted-pay-perks.md ... | head -3 ASSESSMENT: fixtures/postings/03-unlisted-pay-perks.md diff --git a/pipelines/job-assessment/SETUP.md b/pipelines/job-assessment/SETUP.md index 644bde1..76ac121 100644 --- a/pipelines/job-assessment/SETUP.md +++ b/pipelines/job-assessment/SETUP.md @@ -51,9 +51,10 @@ The first prints `validate_profile: clean, 0 warning(s)` and exits 0. The second **Step 5. Run the whole chain on one posting.** ```text +OUT=$(mktemp -d) python3 scripts/assess_offline.py fixtures/postings/03-unlisted-pay-perks.md \ --findings fixtures/findings/03-unlisted-pay-perks.findings.json \ - --profile fixtures/robin-sample/career-profile.yaml --out /tmp/ja-out + --profile fixtures/robin-sample/career-profile.yaml --out "$OUT" ``` It prints ten progress lines, then a summary that starts: @@ -64,9 +65,9 @@ ASSESSMENT: fixtures/postings/03-unlisted-pay-perks.md 🚦 VERDICT: Apply with reservations ``` -followed by four score bars (Fit 7, Comp 5, Qualifications 6, Culture 9) and a table of checks. In `/tmp/ja-out` you will find the archived note with the verdict emoji at the front of its name, the working files, and an email card in `email/` that you can open in a browser. Nothing is sent anywhere. +followed by four score bars (Fit 7, Comp 5, Qualifications 6, Culture 9) and a table of checks. In `$OUT` you will find the archived note with the verdict emoji at the front of its name, the working files, and an email card in `email/` that you can open in a browser. Nothing is sent anywhere. -On Windows, use any folder you like in place of `/tmp/ja-out`. +On Windows, use any empty folder you like in place of `$OUT`. ## Use your own file diff --git a/pipelines/job-assessment/assessment/hooks/README.md b/pipelines/job-assessment/assessment/hooks/README.md index a77d346..69b05a8 100644 --- a/pipelines/job-assessment/assessment/hooks/README.md +++ b/pipelines/job-assessment/assessment/hooks/README.md @@ -42,10 +42,10 @@ Add it to your Claude Code settings (`.claude/settings.json` in a project, or th } ``` -The guard needs `python3` and nothing else. You can also run the same check by hand. This example checks the card the quickstart in the main README wrote: +The guard needs `python3` and nothing else. You can also run the same check by hand. This example checks the card the quickstart in the main README wrote (run it in the same shell, so `$OUT` is still set): ```text -python3 assessment/scripts/check_email.py /tmp/ja-out/email/*.html +python3 assessment/scripts/check_email.py "$OUT"/email/*.html ``` It prints `check_email: card passes` and exits 0, or lists each problem and exits 2. diff --git a/skills/plan-this/SKILL.md b/skills/plan-this/SKILL.md index 3635b7e..8986e1f 100644 --- a/skills/plan-this/SKILL.md +++ b/skills/plan-this/SKILL.md @@ -28,7 +28,7 @@ This skill needs one thing from the host project: where plans live. Set it once PLANS_DIR="${PLANS_DIR:-./plans}" ``` -Adjust the default to match your repo's conventions — a `plans/` folder at the repo root, a `_plans/` folder inside a docs or notes directory, whatever already exists. If the project has more than one place plans could live (e.g. a product-work vault and a config/infra vault), ask the user which one applies before writing — one question, wait for the answer — rather than guessing. +Adjust the default to match your repo's conventions — a `plans/` folder at the repo root, a `_plans/` folder inside a docs or notes directory, whatever already exists. If the project has more than one place plans could live (e.g. two separate notes folders, one for product work and one for config), ask the user which one applies before writing — one question, wait for the answer — rather than guessing. ---