Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Cross-run analysis

Find the problems no single run shows, and test the fix before anyone applies it.

Analyze thousands of AI workflow runs to find issues, their causes and replay-tested fixes.

Prevents: Outputs slowly converging, silent repairs, zeros that mean 'failed' and cost creep going unnoticed because no single trace looks broken.

Why it exists

The failure appears after the happy path.

Workflow problems that only exist across many runs.

Scope. Reads records and writes reports; never changes a workflow.

Five-minute orientation

See the smallest complete path.

pip install "git+https://github.com/honeworks/hone-lens"

Real output

What it looks like running.

A hone-lens HTML report: seven findings ranked by severity, each with category, affected share, status and cause.
Real report produced by the package's own quickstart on 2026-09-30. Its data is synthetic (2,000 generated runs with planted issues and fake models).

Architecture

The architecture is a set of promises.

A funnel: statistics over all runs, embeddings of all outputs, an LLM only for descriptions and hypotheses within a budget, then replay tests through a Replayer port.

Core concepts

The few concepts you need before reading the code.

Finding

What goes wrong, how often, where, the suspected cause, a proposed fix and the evidence.

Why. A short ranked list is readable; a dashboard isn't an answer.

Boundary. A person sets its status.

Replay test

Recorded calls re-run with only the suspected input changed.

Why. Correlation isn't cause.

Boundary. Needs a replayer for the models involved.

Operations

What happens after install.

`hone-lens ingest / analyze / findings / explain / test / report / set-status`.

Current boundary

What this project does not solve.

Alpha. No multi-step replay through a workflow runner, watch mode or image/audio analysis yet. Cause-finding depends on what was recorded.

Near-term roadmap. Follow the design history in design/changes/.

What it implements

  • ingest OTLP, Phoenix, Langfuse, honeworks records and hone-flow runs
  • statistics detectors
  • output clustering
  • budgeted LLM descriptions
  • cause ranking
  • replay tests
  • findings lifecycle
  • HTML reports

Engineering checklist

What the repository ships.

Quickstart
Present
Success tests
Present
Failure tests
Present
Architecture notes
Present
Security notes
Not published
Runbook
Not published
Limitations
Present
Changelog
Present

Recent changes

What changed, with the commits.

  • Initial release: hone-lens 0.1.0 Commit
  • Pull requests reviewed by claude[bot] Commit

Need the control, not just the component?

Fit it to the system that has to survive.

The repository exposes the mechanism. Production work is defining the permissions, data, failure costs, evidence, and owners around it.

Let's build something real