Problem shape
The problem beneath them
Ingest existing traces, run cheap statistics over every run, cluster outputs, use a language model only within a budget to describe and hypothesize, rank the recorded inputs that separate affected runs, confirm a suspected cause by replaying recorded calls with that input changed, and hand a short list of findings to a person.
Why I built it
My generation pipeline had every one of those problems, and I found each by hand, late. Song ideas slowly converged on the same imagery. JSON was repaired quietly. A judge from the same model family as the generator graded its own work. The context window cut answers short. A score of 0 turned out to mean "couldn't parse". None of it made any single run look broken. Reading thousands of traces wasn't going to happen, and a dashboard could tell me that a number moved but not why or what to do.
hone-lens reads the traces a workflow already records (OpenTelemetry GenAI traces as OTLP JSON, Phoenix and Langfuse exports, honeworks span stores and hone-flow run folders) and turns them into findings. Each finding says what goes wrong and how often, where, the recorded input that most likely causes it, a proposed fix, and the evidence behind every number.
Correlation is not cause. A proposed fix is tested by replaying recorded calls with only that input changed, and a person decides whether to apply it; the tool never changes the workflow.
It reads records and writes a static report; storing and browsing traces stay with your observability tools.
Different problems, same shape
The surface changes. The decision structure persists.
Why do 40% of the ideas share one pattern?
Each output looks fine alone; the sameness only shows in aggregate.
How many calls were cut off at the context limit last week?
Truncations are recorded but nobody counts them.
Are these zeros bad outputs or failed judges?
A scorer records 0 when it can't parse its own answer.
Did prompt version 3 make repairs more frequent?
A dashboard shows a change, not its cause.
All four questions converge into Cross-Run Failure Analysis.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Ingest
- Question
- What was recorded, and what's new since last time?
- Responsibility
- Read OTLP JSON, Phoenix Parquet, Langfuse exports, honeworks span stores and hone-flow runs incrementally.
- Input
- Trace sources.
- Output
- Normalized runs, calls and outputs in the workspace.
- Stops when
- All new records are read.
-
02
Statistics
- Question
- What stands out without any model?
- Responsibility
- Run detectors for failure, retry and repair rates, truncation, time and cost hot spots, GPU thrash, score spikes and gate rejections, human overrides, setup smells and regressions after a change; users add their own with `@tl.detector`.
- Input
- All runs.
- Output
- Candidate findings with counts.
- Stops when
- Every detector has run.
-
03
Output analysis
- Question
- Are outputs converging?
- Responsibility
- Embed and cluster outputs; measure homogeneity and closeness to prompt sections.
- Input
- Outputs and an embedder.
- Output
- Clusters with shares and examples.
- Stops when
- Clusters are computed reproducibly for the seed.
-
04
Describe within a budget
- Question
- What do these clusters and samples mean?
- Responsibility
- Use an LLM only for describing clusters, reading samples and writing hypotheses, capped by a `Budget` in USD, tokens or calls; a person reviews the proposed failure taxonomy.
- Input
- Selected samples and clusters.
- Output
- Descriptions and hypotheses.
- Stops when
- Done or the budget is spent.
-
05
Rank causes
- Question
- Which recorded input separates affected runs from the rest?
- Responsibility
- Rank prompt sections, models and parameters by how strongly they separate affected runs.
- Input
- Finding, recorded attributes.
- Output
- Suspected causes with effect sizes.
- Stops when
- Ranked, or "not explained yet" is stated.
-
06
Replay test
- Question
- Does changing the suspected input actually fix it?
- Responsibility
- Re-run recorded calls with only that input changed, over several variants and samples, within a budget.
- Input
- Finding, cause, a replayer.
- Output
- Confirmed or refuted cause with the metric before and after.
- Stops when
- The test completes; a fix must not trade quality for the metric.
-
07
Findings lifecycle
- Question
- What does a person do with this?
- Responsibility
- Keep stable finding IDs across re-analysis, statuses (new, dismissed and so on), and terminal, JSON and self-contained HTML reports.
- Input
- Findings and human status changes.
- Output
- A short ranked list and a report linking to the traces behind each number.
- Stops when
- The person has set a status.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Statistic crosses its threshold | Open a finding | Cheap evidence first | Many findings at once: rank by severity |
| Records lack rich attributes | Say what can't be told | Plain traces limit cause-finding | The team adds prompt sections to records |
| A prompt section separates affected runs | Mark it the suspected cause | Correlation is a lead | Replay test is requested |
| Replay removes the effect without quality loss | Mark the cause confirmed | The fix was tested | A person applies it |
| Budget reached | Stop model calls | Analysis cost is capped | More budget is approved |
| Person dismisses a finding | Keep ID and status | Re-analysis shouldn't resurface it as new | The pattern gets worse |
Operating paths