Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Best-of-N selection

Pick the best of several model outputs without letting one judge decide alone.

Generate, score and select the best of N candidates with pluggable code or prompt scorers, and record why the winner won.

Prevents: A failed judge scores a good candidate 0, a hard rule gets averaged away, the expensive judge reads everything, and nobody can explain the winner later.

Why it exists

The failure appears after the happy path.

Best-of-N is dependable but every project rewrites it by hand with the same bugs.

Scope. One reusable selection step and an experiments harness; not an orchestrator or model client.

Five-minute orientation

See the smallest complete path.

pip install "git+https://github.com/honeworks/hone-select"

Real output

What it looks like running.

The hone-select dashboard: a selection run's decision trace with gates, two scoring stages, escalation to a pairwise judge and the winner.
Real hone-select dashboard, run locally over records from the package's own offline examples. The judges in those examples are scripted fakes, so the scores show the mechanism, not model quality.
The hone-select dashboard: an experiment summary with the recommended setup, its 95% interval, the comparison with the baseline, and all four setups ranked.
Real dashboard view of the package's experiments example (a code scorer over two test cases). It shows the report format; the numbers are from a toy example.

Architecture

The architecture is a set of promises.

A small engine over pydantic: generator, gates, staged scorers, aggregation, policy and fallback, with ports for judges, text clients, embedders, record sinks and score caches. Adapters for OpenAI-compatible servers, LangChain and hone-models are optional extras.

  • A scorer that fails gives None with an error, never 0. Inspect test
  • Gates reject; they are never averaged into scores. Design note
  • A pairwise judgement counts only when both orders agree. Inspect test
  • Every run can be explained from the span store alone. Inspect test

Core concepts

The few concepts you need before reading the code.

Gate

A pass/fail check that rejects a candidate.

Why. Rules that can't be traded off shouldn't be scores.

Boundary. A gate never contributes to a total.

Cascade

Scoring stages from cheap to expensive, each keeping its top candidates.

Why. Expensive judges should see finalists only.

Boundary. A cut can drop a candidate a person would keep; stage sizes are configuration.

Experiment

A declared comparison of setups over fixed test cases with a plan a person approves.

Why. Model and prompt choices need evidence, not a demo.

Boundary. Results are for your cases; they're not a general benchmark.

Operations

What happens after install.

Records every step to a local SQLite span store; `hone-select explain` and `hone-select dashboard` read it.

Current boundary

What this project does not solve.

Alpha. LLM judges stay noisy; pairwise agreement and cross-family judges reduce, not remove, bias. Best-of-N multiplies generation cost.

Near-term roadmap. Follow the design history in design/changes/.

What it implements

  • gates
  • cheap-first scoring cascades
  • missing scores as None
  • pairwise tie escalation in both orders
  • budgets and score cache
  • span records with explain
  • experiments with plan and approval
  • local dashboard

Engineering checklist

What the repository ships.

Quickstart
Present
Success tests
Present
Failure tests
Present
Architecture notes
Present
Security notes
Not published
Runbook
Not published
Limitations
Present
Changelog
Present

Recent changes

What changed, with the commits.

  • Blind A/B: a person picks the better of two outputs (design change 0012) Commit
  • Run conditions and model-aware experiments (0010, 0011) Commit
  • Experiments: define, plan, approve, run, read (0009) Commit

Need the control, not just the component?

Fit it to the system that has to survive.

The repository exposes the mechanism. Production work is defining the permissions, data, failure costs, evidence, and owners around it.

Let's build something real