Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Reliability engineering · Reference design

Incident triage with linked evidence

The deploy looked guilty. The replica had been lagging for nine minutes.

A read-only assistant collects evidence into one timeline, names what it couldn't check, and keeps at least three explanations alive until the engineer decides.

Company Parcelwise is a hypothetical company. Every figure is illustrative.

An on-call notebook with a hand-drawn incident timeline and three circled hypotheses.

About Company Parcelwise

Company Parcelwise is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed.

Industry
Checkout and payments platform for online merchants
Size
About 200 engineers in eight teams, plus a six-person SRE team
Systems
About 40 services on Kubernetes, hosted metrics and logs, a paging tool, CI/CD history, runbooks in a wiki
Volume
About 25 pages a month become incidents
Constraints
Assistants get read-only access; no customer personal data in prompts; observability queries are metered

In brief

ProblemOn-call engineers spend the first 15 to 25 minutes collecting context from many tools, and teams anchor on the first plausible cause.
SystemA service registry of approved read-only queries, a budgeted query runner, an evidence timeline with gaps, and model-proposed hypotheses that must cite timeline entries.
People decideThe on-call engineer interprets evidence and decides every action; service owners approve registry and runbook changes.

How the work ran before

Company Parcelwise runs checkout for online merchants: about 40 services on Kubernetes, 200 engineers in eight product teams, and a six-person SRE team. Around 25 pages a month turn into real incidents.

When the pager fires, the on-call engineer does the same thing every time. Open the metrics tool. Open the log search. Check the deployment history. Ask in the incident channel whether anyone deployed or flipped a feature flag. Find the runbook. Paste screenshots into the channel so the next person can see them. The first 15 to 25 minutes of every incident go to collecting context, and when the incident crosses a shift, half of what was checked lives only in the first engineer's head.

Where it broke

At 02:14 UTC a timeout alert fires on the checkout API. The on-call engineer sees that checkout-api was deployed at 02:09. Pool-wait time on the database connection pool started climbing at 02:12. The story writes itself: the deploy broke something. The engineer starts comparing the new revision with the old one.

Forty minutes later someone checks the database replicas. One replica has been lagging since 02:05, and the new revision had moved a read path onto replicas. The deploy mattered, but not the way everyone assumed. The replica-lag graph existed the whole time. Nobody looked, because the first plausible explanation filled the room.

Parcelwise's postmortem finds the same pattern as three earlier ones: the evidence was available early, and the team anchored on the first story.

What I would build

Not an AI that finds the root cause. An assistant that assembles the evidence and keeps the alternatives open.

  1. Collect. When an incident opens, the assistant queries a fixed set of read-only tools for the affected service and a time window around the alert: deployments, feature-flag changes, key metrics, error logs, dependency health. Each query has a scope, a timeout and a cost limit.
  2. Lay out a timeline. Every result becomes a timestamped entry with a link back to the source. Queries that failed or weren't allowed appear in the timeline too, as gaps.
  3. Keep at least three hypotheses. The model proposes explanations, each with the evidence for it, the evidence against it or missing, and the next read-only check that would tell them apart.
  4. Hand it to a person. The on-call engineer reads the packet, runs checks, and decides. The assistant cannot restart, roll back, scale, or change anything.

The value is not intelligence. It is that nothing checkable gets forgotten, and the obvious story has to compete with the others.

Following one incident through

The 02:14 incident, re-run with the assistant. Sample data.

Time (UTC) Source Entry
02:05 Database metrics Replica 2 lag rising (queried at 02:15)
02:09 Deployment history checkout-api revision deployed
02:12 Service metrics Connection-pool wait time rising
02:14 Alerting Checkout timeout alert opened
02:15 Request metrics Request rate normal for the hour
none Feature flags Query failed: token lacks permission (gap)

Sample data.

At 02:16 the packet shows three hypotheses:

Hypothesis Supporting Against or missing Next check
The deploy changed pool behavior Deploy precedes pool waits by 3 min No config diff inspected yet Diff pool settings between revisions
Traffic spike exhausted the pool Timeouts overlap pool waits Request rate is normal None needed unless rate changes
Replica degradation Replica 2 lag started 02:05, before the deploy Unknown whether checkout reads from replicas Check which queries the new revision sends to replicas

Sample data.

The engineer sees replica lag that started before the deploy, on the first screen. The traffic hypothesis is already weakened by data, and the feature-flag gap is visible instead of silently missing. The decision about what to do stays with the engineer.

Company Parcelwise · incident channel · evidence packet

Evidence timeline · checkout-api 02:05 to 02:16 UTC read-only

Replica 2 lag rising

database metrics Observed

02:05 · starts before the deploy

checkout-api revision deployed

deployment history Observed

02:09 · a lead, not a cause

Connection-pool wait rising

service metrics Observed

02:12

Checkout timeout alert

alerting Triggered

02:14

Feature-flag changes

token lacks permission Gap

query failed · shown, not hidden

Three hypotheses stay open. Deploy changed pool behaviour, a traffic spike, or replica degradation. Request rate is normal, so the traffic hypothesis is already weak.

Worked example · sample data. The packet posted at 02:16 for Company Parcelwise, a hypothetical company.

When things go wrong

A tool is down or the token lacks permission. The query appears as a gap with the error. An empty graph and a failed query must never look the same.

The model invents a connection. Every hypothesis must cite timeline entries by ID. A claim without a cited entry is dropped. The engineer can click through to the raw data behind every entry.

The incident is somewhere the assistant can't see. A third-party payment provider's outage won't show up in internal metrics. The packet says which systems it covered, so an empty packet reads as "not found in these sources", never "nothing wrong".

Someone asks the assistant to fix it. It can't. It has no write credentials. The runbook link is there, and the person runs it.

What changes for the team

Moment Before After
First five minutes Open five tools, ask in chat Read one packet with links
Shift handover Retell from memory Timeline and hypotheses carry over
Obvious-looking cause Team anchors on it Competes with at least two alternatives
Postmortem timeline Reconstructed days later Mostly assembled during the incident

Illustrative figures from assumptions in the technical design. Not measured.

Illustrative incident Manual context gathering Evidence packet
Time until key signals are on one screen 15 to 25 min 2 to 4 min, after the queries return
Tool queries per incident ad hoc about 12, fixed set, budgeted
Gaps (queries that failed) visible no yes

I would not promise a shorter time to resolve. That depends on many things the packet doesn't touch. The honest claim is narrower: less time collecting, and fewer incidents where an available signal went unread.

What this does not solve

It doesn't know your system better than your engineers; it knows only what its queries return. It is weakest on novel failures, where nobody has built a query for the relevant signal yet. And it adds cost: every packet runs a dozen API queries against metered observability tools.

How I would prove it

Replay ten past incidents from their stored data, with the start time frozen at the moment of the page. For each, record whether the packet contained the signal the postmortem identified as key, and how early. Then run it silently alongside real incidents for a month and ask each on-call engineer one question afterwards: did the packet show you anything you hadn't checked yet?

The technical design

The same system for engineers: architecture, records, failure handling, evaluation, the options I rejected, and the stack. About 4 minutes.

  1. Scope and assumptions
  2. Architecture
  3. Tool boundary
  4. Failure handling
  5. Evaluation
  6. Improving over time
  7. Trade-offs I considered
  8. Stack
  9. Security and operations

Read the technical design

Slow first minutes on call?

Bring one incident timeline.

One postmortem and the list of tools your engineers opened is enough to map which evidence could arrive sooner.