Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 06 / Decisions, Agents & Authority

The agents write the code. I own the decisions.

Spec-Driven Agent Development

Let Claude Code build and maintain a codebase from written specifications, skills, hooks and fresh-context reviewers, so it keeps working toward exactly what I asked for while I only approve design changes, answer reviews and merge.

Maturity
Implemented / Operational
Critical uncertainty
Agents may decide implementation details and must log them; they may not change the public design, merge, tag or publish. Those stay with the maintainer.
Human boundary
Accept change records, merge, release, override a hook for a session
Applications
Building a library from a spec · Maintaining several repositories · Parallel agent sessions · Automated pull-request review
Terminal output captured 2026-09-30 in the hone-select repository: the Claude Code reviewer agents, hooks and sixteen skills, and the task flow from start-task through plan-change, implement-change, sync-docs, verify-before-done, open-pr, review and finish-task.

Problem shape

The problem beneath them

Give coding agents a written specification with acceptance tests, a small set of rules that outrank convenience, skills for each step of the work, hooks that make the dangerous steps impossible, and reviewers with no context of their own, then keep the human for the decisions that shape the design.

Why I built it

I wanted five open-source libraries built properly (tests first, documented, simple), and I didn't have the months to type them. Coding agents can write that code, but out of the box they fail in two opposite ways. Either they ask for approval on every small choice, which makes me the bottleneck, or they stop asking and make design decisions I would have rejected. Long tasks also outlive a session's context, so a fresh session starts without the reasons behind the last one's choices.

What I needed was a system where the agent always knows what "done" means, decides small things alone and writes them down, cannot take the steps that are mine (changing the design, merging, publishing), and gets its work checked by something other than itself.

How it works in two phases

Building (internal tooling, not published). An internal project-design folder, which stays on my machine, holds the rules every repository shares: simplicity first (with concrete limits such as complexity ≤ 10 per function), goals, naming, architecture principles, the exact cross-package interfaces (ports) with contract tests, the shared record format, a testing strategy and a worker protocol. Each repository gets a SPEC.md with numbered milestones and acceptance cases, and a KICKOFF.md: the one message that starts its worker. Five Claude Code sessions ran in parallel, one per repository. The protocol told each worker to decide rather than ask, log every judgement call in DECISIONS.md, keep STATUS.md current as its memory between sessions, never push or publish, and use one machine-wide GPU lock for real-model tests. Before closing a milestone, a worker ran three read-only subagents: a spec reviewer, a test auditor and a simplicity reviewer. At the end, the flagged decisions from all five repositories were collected into one review document for me: 23 entries, of which about a dozen needed a human decision.

Maintaining (public, in every repository). After publishing, the internal specs were replaced by a public design history per repository: design/current.md (the design and its acceptance cases), design/changes/NNNN-*.md (one record per design change with context, options, decision and migration), and design/decisions.md for small choices. Claude Code follows one flow of skills: start-task (branch, classify), plan-change (write a change record with status proposed and wait for approval), implement-change (tests first), sync-docs, verify-before-done (evidence from scripts/check.sh, not "should work"), open-pr, address-review and finish-task, plus skills for debugging, examples, adapters, deprecation, real-model tests, triage, releases and learning from past reviews. Hooks enforce the lines that must not be crossed: a guard-main hook blocks file edits on main; a post-push hook reminds the agent when a pushed branch has no pull request; settings deny uv publish and twine. On GitHub, claude[bot] reviews every push to a pull request by starting three fresh reviewer agents (a PR reviewer, a test auditor and a simplicity reviewer) who get only the PR number and commit range, never the author's opinion. The bot triages their findings itself and posts REQUEST_CHANGES or a comment, never an approval: approving is a person's decision.

Agents may decide implementation details and must log them; they may not change the public design, merge, tag or publish. Those stay with the maintainer.

This is a way of working, not a product. It depends on a specification that is actually good; agents follow a bad spec faithfully.

Different problems, same shape

The surface changes. The decision structure persists.

01 / A new package

How do I get from an empty repository to a tested package without supervising every step?

Each session starts by re-explaining the goal.

02 / Parallel work

Can five agents build five packages at once without breaking each other?

Cross-package interfaces change under each other.

03 / Design drift

Who decided to change the public API?

Behaviour changed in a commit that looked like a refactor.

04 / Review

Is the reviewer checking the rules, or agreeing with the author?

The same session writes and reviews the code.

All four questions converge into Spec-Driven Agent Development.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Shared rules

    Question
    What outranks everything except correctness?
    Responsibility
    Write the rules once: simplicity limits, explicit failure (`None` is not `0`), determinism, records, and the ports between packages.
    Input
    The owner's intent.
    Output
    Short documents every session reads first (`AGENTS.md`, `CONTRIBUTING.md`, family specs).
    Stops when
    A newcomer, human or agent, could follow them.
  2. 02

    Specification with acceptance cases

    Question
    What does "done" mean for this package?
    Responsibility
    Number the milestones and give each behaviour an acceptance case that must exist as an end-to-end test through the public API.
    Input
    Goals and non-goals.
    Output
    `SPEC.md` during building; `design/current.md` afterwards.
    Stops when
    Every guarantee has a test that would fail without the feature.
  3. 03

    Skills per step

    Question
    How does the agent do each kind of task the same way every time?
    Responsibility
    One plain-Markdown skill per step, each saying what to read, what to do, and what evidence to show.
    Input
    The repository's commands and rules.
    Output
    A repeatable flow any session can follow; other tools can follow the same Markdown.
    Stops when
    The step ends in its stated output (a branch, a change record, a green check, a PR).
  4. 04

    Hooks for the hard lines

    Question
    Which mistakes must be impossible rather than discouraged?
    Responsibility
    Block edits on `main`, deny publishing commands, format code after edits, and remind about missing pull requests.
    Input
    Tool calls from the agent.
    Output
    Refusal with a reason, or silent pass.
    Stops when
    The tool call is allowed or blocked.
  5. 05

    Decide and log

    Question
    What happens when the spec is ambiguous?
    Responsibility
    Choose the option most consistent with the goals, record question, options, choice and reason, and continue; flag entries that need the owner or affect other repositories.
    Input
    An ambiguity.
    Output
    A numbered decision the code can cite (`D-00N`).
    Stops when
    The decision is logged; the owner reviews flagged ones later in one batch.
  6. 06

    Fresh-context review

    Question
    Does the change follow the rules, judged by someone who didn't write it?
    Responsibility
    Run reviewer agents with only the PR and range; the bot triages, keeps real findings on changed lines, and never approves.
    Input
    A pull request.
    Output
    Inline findings with severity, a verdict, dropped findings with reasons.
    Stops when
    The review is posted; the author answers each thread.
  7. 07

    Human decisions

    Question
    What does the maintainer still do?
    Responsibility
    Accept or reject change records, answer review threads, merge, and ask for releases.
    Input
    Proposed changes, reviews, flagged decisions.
    Output
    Approved design, merged code, releases.
    Stops when
    Nothing waits on the maintainer.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
Change touches public API, guarantees, records, config, ports or dependencies Write a change record and wait Design belongs to the maintainer The maintainer rejects or asks for changes
Small implementation choice Decide and log in `decisions.md` Asking would make the owner a bottleneck The choice affects other repositories
Spec is ambiguous during a build Pick, log, continue Stopping costs more than a reversible choice Two options change public API equally: still pick, flag it
Test fails Fix the cause, never weaken the test A weakened test hides the bug The cause is a design gap: change record
Edit attempted on `main` Block it Every task happens on a branch The owner overrides for one session
Reviewer finding on an unchanged line Drop it with a reason Reviews judge this change The same issue recurs across PRs

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Start task, implement tests first, sync docs, verify, open PR, bot review with no blocking findings, maintainer merges.

Final state
Merged change with passing checks.
Owner
The agent until the PR; the maintainer for the merge.
Evidence
Commits, `scripts/check.sh` output in the PR, review with its range marker.
Recovery
Revert or follow-up PR.
Ambiguous path

The task needs a design change; the agent writes a `proposed` change record and stops.

Final state
Proposed design change awaiting approval.
Owner
The maintainer.
Evidence
The record: context, problem, options including "leave it as it is", decision, consequences, migration.
Recovery
Accepted: implement. Rejected: close with the reason kept.
Failure path

Checks fail, or the review requests changes.

Final state
Open PR with requested changes, or a decision logged for later review.
Owner
The agent, then the maintainer for disputed findings.
Evidence
Failing command output, review threads and replies.
Recovery
`address-review` fixes and replies per thread; recurring findings become rules through `learn-from-reviews`.

Authority map

Capability does not grant authority.

RULE

May decide
Block edits on main and publishing commands; require green checks and resolved threads to merge
May not decide
Judge whether a design is good
Required evidence
Hook and branch-protection configuration

MODEL

May decide
Implement, test, document, decide small choices, review others' code
May not decide
Accept design changes, approve PRs, merge, tag, publish
Required evidence
Commits, logged decisions, review findings

SYSTEM

May decide
Run CI, run the review bot, enforce rulesets
May not decide
Approve a pull request
Required evidence
CI logs, review posts

HUMAN

May decide
Accept change records, merge, release, override a hook for a session
May not decide
Skip the recorded design history
Required evidence
Change record status, merge

EXCEPTION

May decide
Hold flagged decisions and proposed changes for the owner
May not decide
Be applied silently
Required evidence
Flagged entries, `proposed` status

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Agent follows a flawed spec Owner review of results Design history shows the spec Change record fixing the spec Maintainer
Context runs out mid-task Session ends `STATUS.md` / change records carry state New session resumes from them Agent
Reviewer false positives Rejected comments recur Findings dropped with reasons `learn-from-reviews` tunes the reviewer Maintainer
Tests weakened to pass Test auditor, rule 7 Review blocks Restore test, fix cause Agent
Parallel agents break an interface Contract tests Build fails in the offending repo Fix against the ports spec Agent
Shared GPU contention GPU lock One real-model run at a time Wait on the lock System

Invariants and guarantees

Properties the structure is designed to preserve.

  • No edit happens on main through the agent's edit tools.
  • A behaviour change without an accepted change record does not merge.
  • Every acceptance case has an end-to-end test through the public API.
  • The review bot never approves.
  • Agents never push to main, tag or publish unless the maintainer asks.
  • Judgement calls are written down where the next session and the owner can read them.

What changes between implementations

The constraints determine the final mechanism.

The rules change with the project: honeworks uses Python, uv, ruff and pyright and a GPU lock; a web app would have different gates. The size of the spec changes with the risk: a library with public guarantees needs acceptance cases per behaviour; an internal script may need a page. The review changes with the team: a solo maintainer relies on the bot and answers threads; a team adds human reviewers on top. The same structure is visible in this website's own repository, which has a working agreement and project skills for design, Django, data and writing.

Open-source alternatives I compared

This is a practice more than a tool, and it borrows openly. Two of the skills are adapted from obra/superpowers (MIT): debug-failure from its systematic-debugging skill and verify-before-done from its verification-before-completion skill, as the skill files state. The review workflow runs Claude in GitHub Actions. Plugins recommended in the settings (a pyright language server and a PR review toolkit) come from Anthropic's official plugin catalog. What I added is the combination: specs with acceptance cases, one flow of skills per repository, hooks for the hard lines, a public design history instead of chat memory, and a reviewer that is never the author.

Evidence chain

Follow the pattern into systems and software.

This is how the five honeworks repositories were built and are maintained; the skills, hooks, reviewer agents and review workflow are public in each repository. The build-phase kit is internal tooling and is not published. No productivity figures are claimed.

The model-routed code review design uses the same idea for pull requests in general: deterministic rules decide how much review a change gets, and a model can raise the bar but never lower it. It's a reference design; this pattern is how I actually work.

Every honeworks repository (hone-select, hone-flow, hone-models, hone-taste, hone-lens) contains its AGENTS.md, CLAUDE.md, .claude/skills/, .claude/agents/, .claude/hooks/, the claude[bot] review procedure in .github/claude-review.md, and its design/ history. The README of each says it was specified by a human and built by AI coding agents against written specifications and acceptance tests, with a human reviewing their decisions.

The screenshot is real terminal output from the hone-select repository on 2026-09-30: its reviewer agents, hooks, sixteen skills and the task flow from its CLAUDE.md. The build-phase kit (family specs, per-repo specs and kickoff messages, the worker protocol) is internal tooling: it isn't published, downloadable or open source, and I describe it here from my own files.

  • ImplementedSpec-Driven Agent DevelopmentThis pattern, as built in the open-source code below.
  • Reference designCode review matched to change riskRelated case-study system.
  • Alphahone-selectSource: https://github.com/honeworks/hone-select
  • Alphahone-flowSource: https://github.com/honeworks/hone-flow
  • Alphahone-modelsSource: https://github.com/honeworks/hone-models
  • Alphahone-tasteSource: https://github.com/honeworks/hone-taste
  • Alphahone-lensSource: https://github.com/honeworks/hone-lens

Known boundaries

Limitations and non-fit

  • The quality ceiling is the spec. Writing good acceptance cases is the hard part, and it's human work.
  • The owner still has to read change records and review findings; "minimal interaction" isn't "no interaction".
  • Reviewer agents miss things and flag false positives; the triage step and learn-from-reviews reduce that over time, they don't eliminate it.
  • Heavy for a small script or a one-off prototype.
  • Hooks limit the agent's edit tools; shell commands can still change files, which is why branch protection on GitHub is the final line.

Related patterns

Continue through the adjacent decision structures.

The agent doesn't need to be trusted with everything to be trusted with most things. Write down what done means, make the few irreversible steps impossible, let the agent decide the rest and log it, and have someone other than the author check the work. Then the owner's time goes to the decisions only the owner can make.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real