Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Technical design

Incident triage with linked evidence

The deploy looked guilty. The replica had been lagging for nine minutes.

← Back to the story

Company Parcelwise is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed. The story explains the problem and follows one item through the system; this page is the engineering detail behind it.

Scope and assumptions

The assistant runs when an incident is opened from a page. It covers evidence collection and hypothesis presentation for services the SRE team has registered. It has no write access to any production system.

Illustrative figures assume: 12 read-only queries per packet (deploys, flags, 6 metrics, 2 log queries, 2 dependency health checks), each with a 20-second timeout, run in parallel; model summarization under 30 seconds; the 15-25 minute baseline comes from asking on-call engineers, not from measurement.

Architecture

paging tool ──webhook──► incident coordinator
                              │
               service registry (owners, dashboards, queries, dependencies)
                              │
                  ┌───────────┼─────────────────────────┐
                  ▼           ▼             ▼           ▼
            deploy API   flags API   metrics API   logs API     (read-only tokens, scoped)
                  └───────────┴──────┬──────┴───────────┘
                                     ▼
                           evidence timeline (entries + gaps)
                                     │
                                     ▼
                          hypothesis model (cites entry IDs)
                                     │
                                     ▼
                  packet in incident channel + link to full view
  • Service registry. A version-controlled file per service: owners, the metrics that matter, log queries, upstream and downstream dependencies, runbook links. The assistant can only query what the registry names. This is the most valuable part and the part that takes work to build.
  • Query runner. Executes registry queries with per-query timeouts and a per-incident budget (number of queries and, where the vendor bills by query, cost). Results are stored as entries; failures as gaps.
  • Timeline. Entries with id, timestamp, source, summary, link, and for gaps, error.
  • Hypothesis model. Receives the timeline, the service's dependency list and runbook titles. Returns structured hypotheses: statement, supporting entry IDs, contradicting entry IDs, missing evidence, next read-only check. Output is validated: every cited ID must exist.
  • Presentation. A short message in the incident channel (top three hypotheses, gaps, link) and a full page with the timeline.

Tool boundary

Every token is read-only and scoped to the observability and deployment APIs. The time window defaults to 30 minutes before the alert to now and can't exceed 6 hours. Log queries return aggregates and a capped number of sample lines with customer identifiers masked before they reach the model. Anything outside the registry is refused.

Failure handling

Failure Response
Query timeout or error Entry becomes a gap with the error text
Budget exhausted Remaining queries skipped and listed as gaps
Model output cites unknown entry IDs Hypothesis dropped; logged for evaluation
Model unavailable Packet posts the timeline without hypotheses
Registry missing the service Packet says so and lists what a registry entry needs
Engineer requests more Follow-up queries from the registry, same limits

Evaluation

Two sets. First, ten to fifteen past incidents replayed from retained data with the clock frozen at page time. For each: did the timeline contain the postmortem's key signal; which rank did the eventually-correct hypothesis get; how many cited claims were wrong. Second, live shadow mode for a month with a one-question survey.

Metrics I would report: key-signal coverage, correct hypothesis in the top three, unsupported-claim rate, gap rate per tool (a high rate means a permissions or reliability problem), and cost per packet.

Improving over time

When an engineer corrects a hypothesis ("replica lag is always noise on this service"), that becomes a proposed change to the service registry or runbook, reviewed by the service owner. The assistant does not learn silently. Knowledge changes go through the same review as code.

Trade-offs I considered

  • An autonomous agent that chooses its own queries. More flexible for novel incidents, much harder to bound. I would start with the registry and add model-chosen queries later, still restricted to registry tools.
  • Correlation dashboards only. Cheaper and useful, but they still require someone to decide where to look. The packet's job is to make the alternatives explicit.
  • Auto-remediation for known patterns. Possible later for a small number of well-understood cases, as a separate system with its own approvals. Mixing it into the evidence assistant would change what people trust it for.

Stack

A small Python service receiving the paging webhook, the observability and CI/CD APIs through read-only tokens, a hosted model with structured output, the chat integration for the packet, and Postgres for timelines so postmortems can export them. Services register themselves through a YAML file reviewed in the same repository as their code.

Security and operations

Prompts contain service names, metrics and masked log samples; no customer PII. The model provider must not retain data. The assistant's own health is monitored like any service: packet latency, gap rate and cost per incident, reviewed weekly by the SRE team.