System patterns
11 guarantees behind reliable AI systems.
Choose what must stay true.
These are design guarantees: properties each pattern is built to keep. They are not a service warranty.
Each guarantee names something that must never fail. Next to it is the pattern that holds it: the decisions, authority, and failure modes behind it. The first six are built into my open-source honeworks packages and marked "Open source"; the others are reference designs, marked as such.
-
A crashed run continues where it stopped, with its evidence.
Item 37 of 50 failed and the script starts again from item 1.
Resumable AI Workflow RunsA run is a folder you can open. Open source
-
A failed score is missing, never zero, and every win can be explained.
A judge that crashes gives a good candidate a score of 0.
Gated Best-of-N SelectionNo single judge picks the winner. Open source
-
A prompt that doesn't fit fails loudly instead of being truncated.
A long JSON prompt was silently cut at the context limit.
Capability-Checked Model CallsRefuse the call that can't work. Open source
-
Every taste score says whose taste it learned and why it scored.
The output is correct and still nobody would want it.
Stand-in Human JudgmentAsk people rarely, and ask them well. Open source
-
A cause stays suspected until a replay test confirms it.
No single run looks broken, but the outputs keep getting worse.
Cross-Run Failure AnalysisSome bugs only exist across a thousand runs. Open source
-
Agents decide small things and log them; design changes wait for approval.
The agent keeps asking questions, or stops asking and quietly changes the design.
Spec-Driven Agent DevelopmentThe agents write the code. I own the decisions. Open source
-
Accepted work always ends in an owned state.
Failures have no final state.
Durable, Replayable ExecutionAccepted is a promise. Reference design
-
Merge decisions stay reversible and evidence-based.
Merge decisions cannot be reversed.
Entity Resolution with Reversible CanonicalizationTwo records. One company. Maybe. Reference design
-
Conclusions never outrun their evidence.
Citations do not entail conclusions.
Evidence-Bounded InferenceNo evidence. No conclusion. Reference design
-
A confident answer never skips required review.
Confident outputs bypass required review.
Risk- and Novelty-Aware Model CascadesUse the smallest model that knows when to stop. Reference design
-
Sensitive data never exits, even during fallback.
Public fallback activates during capacity pressure.
Private Inference with No-Egress FallbackPrivate means the fallback stays private too. Reference design
No pattern matches that.
The symptom may use different language. Try a broader term, or clear the filters.
Will be published soon
Three more patterns from the same packages, written up next.
-
Will be published soon
Missing Is Not Zero
Keep 'could not score', 'could not answer' and 'could not call' as distinct, typed outcomes through a whole pipeline, so failures never turn into numbers that decisions are made on.
-
Will be published soon
Offline-First Tests for AI Code
Test model-calling code with public fakes, recorded HTTP and contract checkers by default, and run real-model tests only behind a lock and a marker.
-
Will be published soon
Plan, Approve, Then Run
Declare a model or prompt comparison as a file, show its plan and cost, get a person's approval, and run it under recorded conditions so the result is worth trusting.
Not sure which fits?
Let's build something real.
Tell me what must stay true. I will suggest the right pattern.