Problem shape
The problem beneath them
Give coding agents a written specification with acceptance tests, a small set of rules that outrank convenience, skills for each step of the work, hooks that make the dangerous steps impossible, and reviewers with no context of their own, then keep the human for the decisions that shape the design.
Why I built it
I wanted five open-source libraries built properly (tests first, documented, simple), and I didn't have the months to type them. Coding agents can write that code, but out of the box they fail in two opposite ways. Either they ask for approval on every small choice, which makes me the bottleneck, or they stop asking and make design decisions I would have rejected. Long tasks also outlive a session's context, so a fresh session starts without the reasons behind the last one's choices.
What I needed was a system where the agent always knows what "done" means, decides small things alone and writes them down, cannot take the steps that are mine (changing the design, merging, publishing), and gets its work checked by something other than itself.
How it works in two phases
Building (internal tooling, not published). An internal project-design folder, which stays on my machine, holds the rules every repository shares: simplicity first (with concrete limits such as complexity ≤ 10 per function), goals, naming, architecture principles, the exact cross-package interfaces (ports) with contract tests, the shared record format, a testing strategy and a worker protocol. Each repository gets a SPEC.md with numbered milestones and acceptance cases, and a KICKOFF.md: the one message that starts its worker. Five Claude Code sessions ran in parallel, one per repository. The protocol told each worker to decide rather than ask, log every judgement call in DECISIONS.md, keep STATUS.md current as its memory between sessions, never push or publish, and use one machine-wide GPU lock for real-model tests. Before closing a milestone, a worker ran three read-only subagents: a spec reviewer, a test auditor and a simplicity reviewer. At the end, the flagged decisions from all five repositories were collected into one review document for me: 23 entries, of which about a dozen needed a human decision.
Maintaining (public, in every repository). After publishing, the internal specs were replaced by a public design history per repository: design/current.md (the design and its acceptance cases), design/changes/NNNN-*.md (one record per design change with context, options, decision and migration), and design/decisions.md for small choices. Claude Code follows one flow of skills: start-task (branch, classify), plan-change (write a change record with status proposed and wait for approval), implement-change (tests first), sync-docs, verify-before-done (evidence from scripts/check.sh, not "should work"), open-pr, address-review and finish-task, plus skills for debugging, examples, adapters, deprecation, real-model tests, triage, releases and learning from past reviews. Hooks enforce the lines that must not be crossed: a guard-main hook blocks file edits on main; a post-push hook reminds the agent when a pushed branch has no pull request; settings deny uv publish and twine. On GitHub, claude[bot] reviews every push to a pull request by starting three fresh reviewer agents (a PR reviewer, a test auditor and a simplicity reviewer) who get only the PR number and commit range, never the author's opinion. The bot triages their findings itself and posts REQUEST_CHANGES or a comment, never an approval: approving is a person's decision.
Agents may decide implementation details and must log them; they may not change the public design, merge, tag or publish. Those stay with the maintainer.
This is a way of working, not a product. It depends on a specification that is actually good; agents follow a bad spec faithfully.
Different problems, same shape
The surface changes. The decision structure persists.
How do I get from an empty repository to a tested package without supervising every step?
Each session starts by re-explaining the goal.
Can five agents build five packages at once without breaking each other?
Cross-package interfaces change under each other.
Who decided to change the public API?
Behaviour changed in a commit that looked like a refactor.
Is the reviewer checking the rules, or agreeing with the author?
The same session writes and reviews the code.
All four questions converge into Spec-Driven Agent Development.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Shared rules
- Question
- What outranks everything except correctness?
- Responsibility
- Write the rules once: simplicity limits, explicit failure (`None` is not `0`), determinism, records, and the ports between packages.
- Input
- The owner's intent.
- Output
- Short documents every session reads first (`AGENTS.md`, `CONTRIBUTING.md`, family specs).
- Stops when
- A newcomer, human or agent, could follow them.
-
02
Specification with acceptance cases
- Question
- What does "done" mean for this package?
- Responsibility
- Number the milestones and give each behaviour an acceptance case that must exist as an end-to-end test through the public API.
- Input
- Goals and non-goals.
- Output
- `SPEC.md` during building; `design/current.md` afterwards.
- Stops when
- Every guarantee has a test that would fail without the feature.
-
03
Skills per step
- Question
- How does the agent do each kind of task the same way every time?
- Responsibility
- One plain-Markdown skill per step, each saying what to read, what to do, and what evidence to show.
- Input
- The repository's commands and rules.
- Output
- A repeatable flow any session can follow; other tools can follow the same Markdown.
- Stops when
- The step ends in its stated output (a branch, a change record, a green check, a PR).
-
04
Hooks for the hard lines
- Question
- Which mistakes must be impossible rather than discouraged?
- Responsibility
- Block edits on `main`, deny publishing commands, format code after edits, and remind about missing pull requests.
- Input
- Tool calls from the agent.
- Output
- Refusal with a reason, or silent pass.
- Stops when
- The tool call is allowed or blocked.
-
05
Decide and log
- Question
- What happens when the spec is ambiguous?
- Responsibility
- Choose the option most consistent with the goals, record question, options, choice and reason, and continue; flag entries that need the owner or affect other repositories.
- Input
- An ambiguity.
- Output
- A numbered decision the code can cite (`D-00N`).
- Stops when
- The decision is logged; the owner reviews flagged ones later in one batch.
-
06
Fresh-context review
- Question
- Does the change follow the rules, judged by someone who didn't write it?
- Responsibility
- Run reviewer agents with only the PR and range; the bot triages, keeps real findings on changed lines, and never approves.
- Input
- A pull request.
- Output
- Inline findings with severity, a verdict, dropped findings with reasons.
- Stops when
- The review is posted; the author answers each thread.
-
07
Human decisions
- Question
- What does the maintainer still do?
- Responsibility
- Accept or reject change records, answer review threads, merge, and ask for releases.
- Input
- Proposed changes, reviews, flagged decisions.
- Output
- Approved design, merged code, releases.
- Stops when
- Nothing waits on the maintainer.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Change touches public API, guarantees, records, config, ports or dependencies | Write a change record and wait | Design belongs to the maintainer | The maintainer rejects or asks for changes |
| Small implementation choice | Decide and log in `decisions.md` | Asking would make the owner a bottleneck | The choice affects other repositories |
| Spec is ambiguous during a build | Pick, log, continue | Stopping costs more than a reversible choice | Two options change public API equally: still pick, flag it |
| Test fails | Fix the cause, never weaken the test | A weakened test hides the bug | The cause is a design gap: change record |
| Edit attempted on `main` | Block it | Every task happens on a branch | The owner overrides for one session |
| Reviewer finding on an unchanged line | Drop it with a reason | Reviews judge this change | The same issue recurs across PRs |
Operating paths