Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 05 / Evaluation & Controlled Change

Some bugs only exist across a thousand runs.

Cross-Run Failure Analysis

Read the traces an AI workflow already records, find problems that only show across many runs, rank the recorded inputs that likely cause them, and confirm a cause with a replay test before anyone changes the workflow.

Maturity
Tested / Prototype
Critical uncertainty
Correlation is not cause. A proposed fix is tested by replaying recorded calls with only that input changed, and a person decides whether to apply it; the tool never changes the workflow.
Human boundary
Review the taxonomy, set finding status, apply fixes
Applications
LLM pipelines in production · Batch generation · Judge and scorer health · Prompt and model changes
A hone-lens HTML report generated locally on 2026-09-30 from the package's synthetic runs: seven findings ranked by severity, including a same-family judge, 38% of outputs sharing one pattern with a confirmed prompt-section cause, output repairs, zero scores from failed parses and truncated calls.

Problem shape

The problem beneath them

Ingest existing traces, run cheap statistics over every run, cluster outputs, use a language model only within a budget to describe and hypothesize, rank the recorded inputs that separate affected runs, confirm a suspected cause by replaying recorded calls with that input changed, and hand a short list of findings to a person.

Why I built it

My generation pipeline had every one of those problems, and I found each by hand, late. Song ideas slowly converged on the same imagery. JSON was repaired quietly. A judge from the same model family as the generator graded its own work. The context window cut answers short. A score of 0 turned out to mean "couldn't parse". None of it made any single run look broken. Reading thousands of traces wasn't going to happen, and a dashboard could tell me that a number moved but not why or what to do.

hone-lens reads the traces a workflow already records (OpenTelemetry GenAI traces as OTLP JSON, Phoenix and Langfuse exports, honeworks span stores and hone-flow run folders) and turns them into findings. Each finding says what goes wrong and how often, where, the recorded input that most likely causes it, a proposed fix, and the evidence behind every number.

Correlation is not cause. A proposed fix is tested by replaying recorded calls with only that input changed, and a person decides whether to apply it; the tool never changes the workflow.

It reads records and writes a static report; storing and browsing traces stay with your observability tools.

Different problems, same shape

The surface changes. The decision structure persists.

01 / Output quality

Why do 40% of the ideas share one pattern?

Each output looks fine alone; the sameness only shows in aggregate.

02 / Reliability

How many calls were cut off at the context limit last week?

Truncations are recorded but nobody counts them.

03 / Evaluation

Are these zeros bad outputs or failed judges?

A scorer records 0 when it can't parse its own answer.

04 / Change control

Did prompt version 3 make repairs more frequent?

A dashboard shows a change, not its cause.

All four questions converge into Cross-Run Failure Analysis.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Ingest

    Question
    What was recorded, and what's new since last time?
    Responsibility
    Read OTLP JSON, Phoenix Parquet, Langfuse exports, honeworks span stores and hone-flow runs incrementally.
    Input
    Trace sources.
    Output
    Normalized runs, calls and outputs in the workspace.
    Stops when
    All new records are read.
  2. 02

    Statistics

    Question
    What stands out without any model?
    Responsibility
    Run detectors for failure, retry and repair rates, truncation, time and cost hot spots, GPU thrash, score spikes and gate rejections, human overrides, setup smells and regressions after a change; users add their own with `@tl.detector`.
    Input
    All runs.
    Output
    Candidate findings with counts.
    Stops when
    Every detector has run.
  3. 03

    Output analysis

    Question
    Are outputs converging?
    Responsibility
    Embed and cluster outputs; measure homogeneity and closeness to prompt sections.
    Input
    Outputs and an embedder.
    Output
    Clusters with shares and examples.
    Stops when
    Clusters are computed reproducibly for the seed.
  4. 04

    Describe within a budget

    Question
    What do these clusters and samples mean?
    Responsibility
    Use an LLM only for describing clusters, reading samples and writing hypotheses, capped by a `Budget` in USD, tokens or calls; a person reviews the proposed failure taxonomy.
    Input
    Selected samples and clusters.
    Output
    Descriptions and hypotheses.
    Stops when
    Done or the budget is spent.
  5. 05

    Rank causes

    Question
    Which recorded input separates affected runs from the rest?
    Responsibility
    Rank prompt sections, models and parameters by how strongly they separate affected runs.
    Input
    Finding, recorded attributes.
    Output
    Suspected causes with effect sizes.
    Stops when
    Ranked, or "not explained yet" is stated.
  6. 06

    Replay test

    Question
    Does changing the suspected input actually fix it?
    Responsibility
    Re-run recorded calls with only that input changed, over several variants and samples, within a budget.
    Input
    Finding, cause, a replayer.
    Output
    Confirmed or refuted cause with the metric before and after.
    Stops when
    The test completes; a fix must not trade quality for the metric.
  7. 07

    Findings lifecycle

    Question
    What does a person do with this?
    Responsibility
    Keep stable finding IDs across re-analysis, statuses (new, dismissed and so on), and terminal, JSON and self-contained HTML reports.
    Input
    Findings and human status changes.
    Output
    A short ranked list and a report linking to the traces behind each number.
    Stops when
    The person has set a status.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
Statistic crosses its threshold Open a finding Cheap evidence first Many findings at once: rank by severity
Records lack rich attributes Say what can't be told Plain traces limit cause-finding The team adds prompt sections to records
A prompt section separates affected runs Mark it the suspected cause Correlation is a lead Replay test is requested
Replay removes the effect without quality loss Mark the cause confirmed The fix was tested A person applies it
Budget reached Stop model calls Analysis cost is capped More budget is approved
Person dismisses a finding Keep ID and status Re-analysis shouldn't resurface it as new The pattern gets worse

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Ingest, detect, cluster, rank the cause, replay, report.

Final state
Finding with confirmed cause and proposed fix.
Owner
The workflow owner.
Evidence
Report with metric before and after, traces and clusters.
Recovery
Apply the fix in the workflow; re-analyze later.
Ambiguous path

A statistic stands out but no recorded input explains it.

Final state
Finding with "not explained yet".
Owner
The workflow owner.
Evidence
Counts, affected runs, missing attributes.
Recovery
Record more detail (prompt sections, versions) and re-analyze.
Failure path

Replay refutes the suspected cause, or the budget ends early.

Final state
Finding with refuted or untested cause.
Owner
The workflow owner.
Evidence
Replay results or budget stop.
Recovery
Test the next cause or raise the budget.

Authority map

Capability does not grant authority.

RULE

May decide
Open findings from statistics, rank severity
May not decide
Declare a cause confirmed
Required evidence
Detector name, metric, baseline

MODEL

May decide
Describe clusters, propose hypotheses within budget
May not decide
Confirm causes or apply fixes
Required evidence
Budget use, sampled inputs

SYSTEM

May decide
Rank causes, run replay tests, keep IDs stable
May not decide
Change the workflow
Required evidence
Effect sizes, replay results

HUMAN

May decide
Review the taxonomy, set finding status, apply fixes
May not decide
Skip the replay when claiming a cause is confirmed
Required evidence
Status change with note

EXCEPTION

May decide
Report "not explained" or "refuted"
May not decide
Hide weak evidence behind a confident sentence
Required evidence
What's missing

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Spurious cause Replay refutes it Stays suspected Test the next cause Workflow owner
Fix games the metric Quality check in the test Not confirmed Different fix Workflow owner
Analysis cost runs away Budget cap Stops Raise budget explicitly Workflow owner
Traces lack attributes Detector reports limits Fewer findings Add attributes Workflow owner
Duplicate findings after re-analysis Stable IDs Same ID Status carries over System
Non-reproducible clusters Seeds Same data, same result Pin seeds System

Invariants and guarantees

Properties the structure is designed to preserve.

  • The tool never changes the workflow.
  • A cause is "confirmed" only after a replay test.
  • Model calls during analysis stay inside a stated budget.
  • Findings keep stable IDs across re-analysis.
  • Every number in a report links to the traces behind it.
  • The same data and seeds give the same clusters, findings and IDs.

What changes between implementations

The constraints determine the final mechanism.

The sources change: plain OpenTelemetry traces allow statistics and clustering; honeworks records add prompt sections, selection scores and step statuses, which unlock section-level causes and score spikes. The models change: an OpenAI-compatible embedder and LLM, or local ones through hone-models, or the deterministic fakes for tests. The replayer changes: hone-models can replay recorded calls; the flow extra forks hone-flow runs.

Open-source alternatives I compared

Tool What it does well Why it didn't fit this job
Phoenix, Langfuse Store and show traces A person still has to find the pattern
Clio, Kura, OpenClio, BERTopic Group and describe many outputs Don't link a cluster to the input that caused it
Docent Finds and counts behaviours in agent transcripts Aimed at agent transcripts, not general workflows
DSPy, GEPA Search for better prompts against a metric Don't explain what was wrong or test one hypothesis

hone-lens reads what Phoenix and Langfuse export, so it complements them rather than replacing them.

Evidence chain

Follow the pattern into systems and software.

Implemented in hone-lens (Apache-2.0, alpha 0.1.0) with acceptance tests in CI. The screenshot shows its report over the package's synthetic runs with planted issues; no production result is claimed.

The incident-triage design shares the discipline of keeping hypotheses open until evidence tells them apart. The support-drafts design names the metrics (stale-source rate, abstention rate) that a cross-run analysis would watch. Both are reference designs.

Implemented in hone-lens, Apache-2.0, version 0.1.0 (alpha). The core needs only numpy and pydantic. The repository holds the design, its acceptance cases and five recorded design changes.

The screenshot is the real self-contained HTML report produced by hone-lens' own quickstart, which I ran locally on 2026-09-30. Its data is synthetic: 2,000 generated runs with planted issues and deterministic fake models. The report is real output; the findings describe the planted issues, not a real workflow.

Known boundaries

Limitations and non-fit

  • Multi-step replay through a workflow runner, hand-off to a prompt optimizer, a watch mode with alerts, and analysis of image and audio outputs are not in this version.
  • It needs volume: with a few dozen runs, read them yourself.
  • Cause ranking depends on what was recorded; unrecorded inputs can't be blamed.
  • It's not a trace store, BI tool or monitoring system.

Related patterns

Continue through the adjacent decision structures.

Watching a workflow and fixing it are different jobs. Statistics find the pattern, clusters show its shape, a ranked cause gives a lead, and a replay test turns the lead into evidence. Only then does a person change anything.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real