Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 1 addition & 4 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -56,10 +56,7 @@ jobs:
python-version: ${{ matrix.python }}
- run: pip install -r requirements.txt
- name: Run the tests
# pytest exits 5 when it finds no tests. That is expected until the
# first work package adds some, so only that exit code is tolerated.
shell: bash
run: python3 -m pytest -q || [ $? -eq 5 ]
run: python3 -m pytest -q

clean-clone:
name: clean-clone
Expand Down
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,4 +1,6 @@
.DS_Store
venv/
.venv/
.pytest_cache/
__pycache__/
*.pyc
18 changes: 13 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,11 +2,17 @@

Skills, pipelines, and technique write-ups for building AI-agent systems that stay correct as they scale — built by using Claude Code every day, not by reading about it.

**The 60-second tour.**

- **What this is:** working examples of directing an AI coding agent and then checking what it built, with tests and written grades. Not a claim to machine-learning engineering.
- **See it work:** the job assessment system runs from a clean clone with no model and no API key. The quickstart is in [`pipelines/job-assessment/README.md`](pipelines/job-assessment/README.md).
- **The one file to read first:** [`pipelines/docs-pipeline/README.md`](pipelines/docs-pipeline/README.md), the documentation pipeline.

## Start here: the best things to look at

**1. [`pipelines/docs-pipeline`](pipelines/docs-pipeline/)** is the strongest piece in this repo. It is a chain of Claude Code skills that takes a documentation change from the first request through structure review, voice, grammar, links, visuals, expert review, and publishing, with a real gate between every stage. I used a version of it every working day on real documentation. If you only open one thing, open this, and start with its [README](pipelines/docs-pipeline/README.md) and [ARCHITECTURE](pipelines/docs-pipeline/ARCHITECTURE.md).
**1. [`pipelines/docs-pipeline`](pipelines/docs-pipeline/)** is the strongest piece in this repo, and a generic export of the version I used. It is a chain of Claude Code skills that takes a documentation change from the first request through structure review, voice, grammar, links, visuals, expert review, and publishing, with a real gate between every stage. I used a version of it every working day on real documentation. If you only open one thing, open this, and start with its [README](pipelines/docs-pipeline/README.md) and [ARCHITECTURE](pipelines/docs-pipeline/ARCHITECTURE.md).

**2. [`pipelines/job-assessment`](pipelines/job-assessment/)** scores a job posting against your own written criteria and gives a plain verdict: Apply, Apply with reservations, or Skip. A model reads the posting and writes down findings, quoting the posting for every claim. Ordinary code then checks those quotes, does the scoring and picks the verdict, so the same findings always give the same answer. It was built by directing Claude Code, not hand-typed. What is shown: the whole chain runs offline on three invented postings with no model, no network and no API key, and the tests in that folder (run again in a fresh clone by CI) compare its scores, verdicts and saved output files to written-down expected values. A separate run with the `claude` CLI on Claude Sonnet is saved in [`tests/`](pipelines/job-assessment/tests/): the interview built a file that passed the checker, and the assessment gave the expected verdicts on 2 of 2 postings in its second attempt, after the first attempt got one verdict wrong ([both records are kept](pipelines/job-assessment/tests/e2e-output-run1.md)). What is not shown: that a verdict predicts getting hired, or that it beats any other method. No real postings or real person's data are in it. Its README starts with a five-minute quickstart from a clean clone, and every doc in the folder ends with prompts tested on Claude Sonnet and Claude Haiku.
**2. [`pipelines/job-assessment`](pipelines/job-assessment/)** scores a job posting against your own written criteria and gives a plain verdict: Apply, Apply with reservations, or Skip. A model reads the posting and writes down findings, quoting the posting for every claim. Ordinary code then checks those quotes, does the scoring and picks the verdict, so the same findings always give the same answer. It was built by directing Claude Code, not hand-typed. What is shown: the whole chain runs offline on three invented postings with no model, no network and no API key, and the tests in that folder (run again in a fresh clone by CI) compare its scores, verdicts and saved output files to written-down expected values. A separate run with the `claude` CLI on Claude Sonnet is saved in [`tests/`](pipelines/job-assessment/tests/): the interview built a file that passed the checker, and the assessment gave the expected verdicts on 2 of 2 postings in its second attempt, after the first attempt got one verdict wrong ([both records are kept](pipelines/job-assessment/tests/e2e-output-run1.md)). What is not shown: that a verdict predicts getting hired, or that it beats any other method. No real postings or real person's data are in it. Its README starts with a five-minute quickstart from a clean clone, and every README and guide in the folder ends with prompts tested on Claude Sonnet and Claude Haiku.

**3. [`patterns/block-and-tell-hooks.md`](patterns/block-and-tell-hooks.md)** explains how to make an AI agent follow a rule every time instead of hoping it remembers. It is the idea behind most of the guards in my own setup.

Expand All @@ -20,12 +26,14 @@ Every piece in here started as a real problem: a rule that kept getting skipped,

## What's inside

**`skills/`** — self-contained Claude Code skills. Each one is a single markdown file (plus a README) that defines a repeatable, well-scoped task for an AI agent to carry out — a readability pass on documentation, a way to capture an in-progress plan so it survives a context reset, an audit that checks whether an agent's own guardrails are actually strong enough to trust.
**`skills/`** — two self-contained Claude Code skills, each in its own folder with a `SKILL.md` that defines one repeatable, well-scoped task for an AI agent: `docs-readability-check` (a readability pass on documentation) and `plan-this` (capture an in-progress plan so it survives a context reset). `plan-this` also has a README.

**`pipelines/`** — multi-stage systems, not single tasks. The anchor piece here is a documentation-engineering pipeline: a chain of skills that takes a raw content change through structure review, voice/style checks, grammar, link and visual verification, and a subject-matter-expert review gate before anything publishes. This is a first pass, still being developed further — see `pipelines/docs-pipeline/` for its own README and current state.
**`pipelines/`** — multi-stage systems, not single tasks. The anchor piece here is a documentation-engineering pipeline: a chain of skills that takes a raw content change through structure review, voice/style checks, grammar, link and visual verification, and a subject-matter-expert review gate before anything publishes. See `pipelines/docs-pipeline/` for its own README and current state. The second pipeline is [`pipelines/job-assessment/`](pipelines/job-assessment/), which scores a job posting against written criteria.

**`patterns/`** — written technique docs for ideas that are more valuable described in prose than shipped as literal runnable code, either because the real implementation is too specific to one project to be useful as-is, or because the idea itself is the point. Covers: how to make an "always do X first" instruction actually reliable instead of hoped-for (block-and-tell hooks), how to keep an agent's standing instructions from becoming an unmaintainable single file as they grow (rules-index architecture), and how to give an agent memory that survives months of use without turning into an unreadable dump (typed, size-bounded memory).

**Prompt blocks.** The README and guides of [`pipelines/docs-pipeline/`](pipelines/docs-pipeline/) and [`pipelines/job-assessment/`](pipelines/job-assessment/) and the three docs in `patterns/` each end with a "Prompt for your AI model" block, tested on Claude Sonnet and Claude Haiku. This top-level README, `ROADMAP.md` and the two skills do not have one.

## How the pieces relate

A skill is one task. A pipeline is several skills chained with real gates between them (nothing moves to the next stage until the current one passes). A pattern is the idea behind a mechanism, written down so it can be rebuilt in a different codebase without copying code that won't fit.
Expand All @@ -36,4 +44,4 @@ Everything here was built using Claude Code, directed and reviewed by a human, n

## What's next

See `ROADMAP.md` for what's planned for the next batch — genericized versions of a few more utility scripts, further development on the documentation pipeline, and a handful of additional skills that need a lighter cleanup pass before they're ready to publish.
See `ROADMAP.md` for what is done and what is planned next: more genericized utility scripts, a fourth pattern doc, and a handful of additional skills that need a lighter cleanup pass before they are ready to publish.
10 changes: 3 additions & 7 deletions ROADMAP.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# Roadmap

What's here now is a first batch. This is what's planned next, so it's clear what's deliberately not done yet versus what's missing by accident.
What's here now is a first batch. This is what is done, and what's planned next, so it's clear what's deliberately not done yet versus what's missing by accident.

## In progress
## Done

**Job assessment system.** Built in small pull requests under `pipelines/job-assessment/`. The scripts, schemas, fictional sample data, both skills and the offline end-to-end run are merged and tested; the earlier `tools/job-fit-screen` it replaces has been removed. `ARCHITECTURE.md`, `SETUP.md` and the prompt blocks, tested on Claude Sonnet and Claude Haiku, are merged too, so every item on its status list is done.

Expand All @@ -17,7 +17,7 @@ What's here now is a first batch. This is what's planned next, so it's clear wha

**A fourth pattern doc: always-read-first / write-back state files.** A convention for keeping a piece of live state (a machine's current status, a project's current phase) trustworthy across many separate agent sessions touching it — read the state file before acting, write back immediately after any change, and enforce both halves with hooks rather than hoping the instruction gets followed every time.

**Further development on the documentation pipeline.** The version in `pipelines/docs-pipeline/` is a first cleaned pass, not the finished shape. Ongoing work on it will land here as it develops.
**Further development on the documentation pipeline.** The version in `pipelines/docs-pipeline/` is a generic export of the version I used, not the finished shape. Ongoing work on it will land here as it develops.

**A handful of additional skills**, pending a lighter genericization pass:
- An AI-writing-tell detector — scans prose for patterns that read as machine-written and proposes rewrites, backed by a pattern registry with a cited source per rule, dated retirements for tells that stopped working, a documented false-positive case per rule, and a scheduled refresh step that diffs several outside source authorities to keep the pattern list current. Two small genericization items: a hardcoded personal file path and a house-style config tuned to one person's preferences.
Expand All @@ -34,7 +34,3 @@ What's here now is a first batch. This is what's planned next, so it's clear wha
- Anything narratively specific to one product's fictional world or brand voice — general technique only, not product content
- Anything still under an employer's confidentiality terms without explicit permission to publish
- Scratch/throwaway scripts that were never meant to be reused

## Not yet decided

Whether to link this repo from a resume, portfolio site, or LinkedIn — that's a deliberate later decision, not part of standing this repo up.
2 changes: 1 addition & 1 deletion patterns/block-and-tell-hooks.md
Original file line number Diff line number Diff line change
Expand Up @@ -146,4 +146,4 @@ My rule is: [YOUR "ALWAYS DO X FIRST" RULE]. My tool is: [YOUR AGENT TOOL].
Write the guard and marker scripts for my rule in my tool's hook format. Then give me a three-step test plan: one test where the guard must block, one where it must allow after the marker is written, and one where I confirm the release path works. Tell me what output proves each test passed.
```

**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically.
**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically. I did not save those answers, so there is no record to read here, unlike the saved runs in `pipelines/job-assessment/tests/`.
2 changes: 1 addition & 1 deletion patterns/rules-index-architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -165,4 +165,4 @@ My instructions file is pasted below. My tool is: [YOUR AGENT TOOL].
Propose a split into domain files with an index table, and tell me which files should be excluded from startup loading. Then give me a way to test that the split worked: how to measure total loaded size before and after, and one question to ask the agent that only a specific rule file can answer.
```

**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically.
**How these prompts were checked.** Each of the three prompts was run once with a small model (Claude Haiku) through the `claude` command line, with the full text of this file pasted in and sample details filled in. All three gave an on-topic answer that matched what this file says. In two runs a placeholder was left unfilled by my test setup, and the model noticed and said so or asked for the missing text instead of making something up. That is the behavior you want. One run per prompt is a light check, not a benchmark, so read the answers critically. I did not save those answers, so there is no record to read here, unlike the saved runs in `pipelines/job-assessment/tests/`.
Loading
Loading