Reliability engineering · Reference design
Incident triage with linked evidence
The deploy looked guilty. The replica had been lagging for nine minutes.
A read-only assistant collects evidence into one timeline, names what it couldn't check, and keeps at least three explanations alive until the engineer decides.
Company Parcelwise is a hypothetical company. Every figure is illustrative.
About Company Parcelwise
Company Parcelwise is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed.
- Industry
- Checkout and payments platform for online merchants
- Size
- About 200 engineers in eight teams, plus a six-person SRE team
- Systems
- About 40 services on Kubernetes, hosted metrics and logs, a paging tool, CI/CD history, runbooks in a wiki
- Volume
- About 25 pages a month become incidents
- Constraints
- Assistants get read-only access; no customer personal data in prompts; observability queries are metered
In brief
How the work ran before
Company Parcelwise runs checkout for online merchants: about 40 services on Kubernetes, 200 engineers in eight product teams, and a six-person SRE team. Around 25 pages a month turn into real incidents.
When the pager fires, the on-call engineer does the same thing every time. Open the metrics tool. Open the log search. Check the deployment history. Ask in the incident channel whether anyone deployed or flipped a feature flag. Find the runbook. Paste screenshots into the channel so the next person can see them. The first 15 to 25 minutes of every incident go to collecting context, and when the incident crosses a shift, half of what was checked lives only in the first engineer's head.
Where it broke
At 02:14 UTC a timeout alert fires on the checkout API. The on-call engineer sees that checkout-api was deployed at 02:09. Pool-wait time on the database connection pool started climbing at 02:12. The story writes itself: the deploy broke something. The engineer starts comparing the new revision with the old one.
Forty minutes later someone checks the database replicas. One replica has been lagging since 02:05, and the new revision had moved a read path onto replicas. The deploy mattered, but not the way everyone assumed. The replica-lag graph existed the whole time. Nobody looked, because the first plausible explanation filled the room.
Parcelwise's postmortem finds the same pattern as three earlier ones: the evidence was available early, and the team anchored on the first story.
What I would build
Not an AI that finds the root cause. An assistant that assembles the evidence and keeps the alternatives open.
- Collect. When an incident opens, the assistant queries a fixed set of read-only tools for the affected service and a time window around the alert: deployments, feature-flag changes, key metrics, error logs, dependency health. Each query has a scope, a timeout and a cost limit.
- Lay out a timeline. Every result becomes a timestamped entry with a link back to the source. Queries that failed or weren't allowed appear in the timeline too, as gaps.
- Keep at least three hypotheses. The model proposes explanations, each with the evidence for it, the evidence against it or missing, and the next read-only check that would tell them apart.
- Hand it to a person. The on-call engineer reads the packet, runs checks, and decides. The assistant cannot restart, roll back, scale, or change anything.
The value is not intelligence. It is that nothing checkable gets forgotten, and the obvious story has to compete with the others.
Following one incident through
The 02:14 incident, re-run with the assistant. Sample data.
| Time (UTC) | Source | Entry |
|---|---|---|
| 02:05 | Database metrics | Replica 2 lag rising (queried at 02:15) |
| 02:09 | Deployment history | checkout-api revision deployed |
| 02:12 | Service metrics | Connection-pool wait time rising |
| 02:14 | Alerting | Checkout timeout alert opened |
| 02:15 | Request metrics | Request rate normal for the hour |
| none | Feature flags | Query failed: token lacks permission (gap) |
Sample data.
At 02:16 the packet shows three hypotheses:
| Hypothesis | Supporting | Against or missing | Next check |
|---|---|---|---|
| The deploy changed pool behavior | Deploy precedes pool waits by 3 min | No config diff inspected yet | Diff pool settings between revisions |
| Traffic spike exhausted the pool | Timeouts overlap pool waits | Request rate is normal | None needed unless rate changes |
| Replica degradation | Replica 2 lag started 02:05, before the deploy | Unknown whether checkout reads from replicas | Check which queries the new revision sends to replicas |
Sample data.
The engineer sees replica lag that started before the deploy, on the first screen. The traffic hypothesis is already weakened by data, and the feature-flag gap is visible instead of silently missing. The decision about what to do stays with the engineer.
Replica 2 lag rising
database metrics Observed02:05 · starts before the deploy
checkout-api revision deployed
deployment history Observed02:09 · a lead, not a cause
Connection-pool wait rising
service metrics Observed02:12
Checkout timeout alert
alerting Triggered02:14
Feature-flag changes
token lacks permission Gapquery failed · shown, not hidden
Three hypotheses stay open. Deploy changed pool behaviour, a traffic spike, or replica degradation. Request rate is normal, so the traffic hypothesis is already weak.
When things go wrong
A tool is down or the token lacks permission. The query appears as a gap with the error. An empty graph and a failed query must never look the same.
The model invents a connection. Every hypothesis must cite timeline entries by ID. A claim without a cited entry is dropped. The engineer can click through to the raw data behind every entry.
The incident is somewhere the assistant can't see. A third-party payment provider's outage won't show up in internal metrics. The packet says which systems it covered, so an empty packet reads as "not found in these sources", never "nothing wrong".
Someone asks the assistant to fix it. It can't. It has no write credentials. The runbook link is there, and the person runs it.
What changes for the team
| Moment | Before | After |
|---|---|---|
| First five minutes | Open five tools, ask in chat | Read one packet with links |
| Shift handover | Retell from memory | Timeline and hypotheses carry over |
| Obvious-looking cause | Team anchors on it | Competes with at least two alternatives |
| Postmortem timeline | Reconstructed days later | Mostly assembled during the incident |
Illustrative figures from assumptions in the technical design. Not measured.
| Illustrative incident | Manual context gathering | Evidence packet |
|---|---|---|
| Time until key signals are on one screen | 15 to 25 min | 2 to 4 min, after the queries return |
| Tool queries per incident | ad hoc | about 12, fixed set, budgeted |
| Gaps (queries that failed) visible | no | yes |
I would not promise a shorter time to resolve. That depends on many things the packet doesn't touch. The honest claim is narrower: less time collecting, and fewer incidents where an available signal went unread.
What this does not solve
It doesn't know your system better than your engineers; it knows only what its queries return. It is weakest on novel failures, where nobody has built a query for the relevant signal yet. And it adds cost: every packet runs a dozen API queries against metered observability tools.
How I would prove it
Replay ten past incidents from their stored data, with the start time frozen at the moment of the page. For each, record whether the packet contained the signal the postmortem identified as key, and how early. Then run it silently alongside real incidents for a month and ask each on-call engineer one question afterwards: did the packet show you anything you hadn't checked yet?
The technical design
The same system for engineers: architecture, records, failure handling, evaluation, the options I rejected, and the stack. About 4 minutes.
Slow first minutes on call?
Bring one incident timeline.
One postmortem and the list of tools your engineers opened is enough to map which evidence could arrive sooner.