Problem shape
The problem beneath them
Generate several candidates, drop those that fail pass/fail gates, score the rest in stages from cheap to expensive, treat a failed score as missing rather than zero, escalate near-ties to a pairwise judge that must agree with itself in both orders, pick the winner by an explicit policy, and record every step.
Why I built it
Best-of-N is the most dependable way I know to improve a generative model's output without changing the model. My local song pipeline ran that loop in five places: ideas, lyrics, rendered songs, visual styles and keyframes. Each was written by hand, and each had its own version of the same bugs. A judge returned something unreadable and the parser gave it 0.0. "The hook must appear" was a score, so a draft without the hook could win on average. A slow model judged all ten candidates, including the ones with the wrong line count. Ties went to whichever candidate was generated first. The judge was the same model family as the generator, which prefers its own output. And only the winner was saved.
hone-select turns that loop into one tested engine. You write a generator and scorers as plain functions, prompt judges, or external commands. The engine does the rest in a fixed order and writes every decision as a span to a local store, so hone-select explain <run_id> can rebuild why a candidate won from the store alone.
A judge's score is evidence, not a verdict. Gates reject, scores rank, and a policy picks; a model is one possible source of scores and never decides alone.
The engine knows nothing about the domain: a candidate is data plus optional files. It doesn't call models itself; it asks the clients you inject. It is one step, not an orchestrator.
Different problems, same shape
The surface changes. The decision structure persists.
The judge timed out. Why did the best draft come last?
A scorer error is stored as a score of 0.
The chorus lost its hook. Why did it still win on average?
A pass/fail rule is averaged with quality scores.
Why did the expensive judge read ten drafts a regex could reject?
Every scorer runs on every candidate.
Two drafts scored almost the same. Who decided, and why?
Only the winner is kept; the comparison is lost.
All four questions converge into Gated Best-of-N Selection.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Generate
- Question
- Which candidates exist, and how were they varied?
- Responsibility
- Call the generator with an explicit seed and variation schedule; optionally remove exact or embedding duplicates.
- Input
- Task, `n`, variation parameters, seed.
- Output
- Candidates with IDs derived from their content (`hashlib`, never `hash()`).
- Stops when
- `n` candidates exist, or the budget stops generation.
-
02
Gate
- Question
- Which candidates break a rule that can't be traded off?
- Responsibility
- Run pass/fail gates; a failed gate rejects the candidate with a reason. Gates are never averaged into scores.
- Input
- Candidates and gate functions (code, command or prompt gate).
- Output
- Passing candidates and a rejection record per failure.
- Stops when
- Every candidate has passed or been rejected.
-
03
Cascade
- Question
- Which candidates deserve the expensive judges?
- Responsibility
- Score in stages; each stage keeps its top `keep_top` candidates for the next. Cached scores are keyed by scorer version.
- Input
- Passing candidates, staged scorers with costs, budgets for cost, time and money.
- Output
- Per-stage scores; a scorer error becomes `Score(None, error=...)`.
- Stops when
- The last stage finishes or a budget cap is reached.
-
04
Aggregate
- Question
- How do several scores become one total?
- Responsibility
- Combine with weights; a `missing` policy decides what a `None` does (renormalize by default).
- Input
- Stage scores, weights, missing policy, floor.
- Output
- Totals with their parts.
- Stops when
- Every finalist has a total or is marked unscorable.
-
05
Escalate
- Question
- Is the top of the ranking too close or too uncertain to trust?
- Responsibility
- Send near-ties within `tie_margin`, or low-confidence scores, to a pairwise judge, asked in both orders; a judgement counts only when both orders agree.
- Input
- The contested candidates and a pairwise judge.
- Output
- An agreed preference or "no agreement".
- Stops when
- The tie is settled or left explicitly unsettled.
-
06
Select and record
- Question
- Which candidate wins, and can we explain it later?
- Responsibility
- Apply the policy (`argmax`, `first_above`, `pairwise_tournament`); apply a fallback when everything was rejected; write spans for every step.
- Input
- Totals, escalation results, policy, fallback.
- Output
- Winner (or none) and a span record in `.hone/select/spans.db`.
- Stops when
- The result and its record are written.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Gate fails | Reject the candidate | A broken rule can't be compensated by quality | Every candidate is rejected: apply the fallback |
| Scorer raises or returns junk | Record `None` with the error | "Couldn't score" is not "bad" | Too many `None` values leave no usable total |
| Candidate below stage cut | Drop before expensive stages | Expensive judges see only finalists | The cut removes a candidate a person would keep |
| Top two within `tie_margin` | Ask a pairwise judge in both orders | Close totals are noise-level differences | The two orders disagree |
| Judge shares the generator's model family | Warn | Self-preference bias | The warning is ignored in production |
| Budget reached | Stop and select from what exists | Cost caps are hard limits | The stop happens before any finalist is scored |
Operating paths