Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 02 / Decisions, Agents & Authority

No single judge picks the winner.

Gated Best-of-N Selection

Generate several candidates, reject the ones that fail hard rules, score the rest cheapest-first, send near-ties to a pairwise judge asked in both orders, and record every decision so the winner can be explained later.

Maturity
Tested / Prototype
Critical uncertainty
A judge's score is evidence, not a verdict. Gates reject, scores rank, and a policy picks; a model is one possible source of scores and never decides alone.
Human boundary
Define gates, scorers, weights, policy; judge A/B blind
Applications
Picking one of several LLM drafts · Choosing generated images or audio · Comparing models, prompts or settings · Code patch selection
The hone-select dashboard, captured locally on 2026-09-30 over records from its offline examples: a selection run whose decision trace shows four candidates generated, three passing gates, two scoring stages, escalation to a pairwise judge for a near-tie, and the winner.

Problem shape

The problem beneath them

Generate several candidates, drop those that fail pass/fail gates, score the rest in stages from cheap to expensive, treat a failed score as missing rather than zero, escalate near-ties to a pairwise judge that must agree with itself in both orders, pick the winner by an explicit policy, and record every step.

Why I built it

Best-of-N is the most dependable way I know to improve a generative model's output without changing the model. My local song pipeline ran that loop in five places: ideas, lyrics, rendered songs, visual styles and keyframes. Each was written by hand, and each had its own version of the same bugs. A judge returned something unreadable and the parser gave it 0.0. "The hook must appear" was a score, so a draft without the hook could win on average. A slow model judged all ten candidates, including the ones with the wrong line count. Ties went to whichever candidate was generated first. The judge was the same model family as the generator, which prefers its own output. And only the winner was saved.

hone-select turns that loop into one tested engine. You write a generator and scorers as plain functions, prompt judges, or external commands. The engine does the rest in a fixed order and writes every decision as a span to a local store, so hone-select explain <run_id> can rebuild why a candidate won from the store alone.

A judge's score is evidence, not a verdict. Gates reject, scores rank, and a policy picks; a model is one possible source of scores and never decides alone.

The engine knows nothing about the domain: a candidate is data plus optional files. It doesn't call models itself; it asks the clients you inject. It is one step, not an orchestrator.

Different problems, same shape

The surface changes. The decision structure persists.

01 / Draft selection

The judge timed out. Why did the best draft come last?

A scorer error is stored as a score of 0.

02 / Hard requirements

The chorus lost its hook. Why did it still win on average?

A pass/fail rule is averaged with quality scores.

03 / Cost

Why did the expensive judge read ten drafts a regex could reject?

Every scorer runs on every candidate.

04 / Audit

Two drafts scored almost the same. Who decided, and why?

Only the winner is kept; the comparison is lost.

All four questions converge into Gated Best-of-N Selection.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Generate

    Question
    Which candidates exist, and how were they varied?
    Responsibility
    Call the generator with an explicit seed and variation schedule; optionally remove exact or embedding duplicates.
    Input
    Task, `n`, variation parameters, seed.
    Output
    Candidates with IDs derived from their content (`hashlib`, never `hash()`).
    Stops when
    `n` candidates exist, or the budget stops generation.
  2. 02

    Gate

    Question
    Which candidates break a rule that can't be traded off?
    Responsibility
    Run pass/fail gates; a failed gate rejects the candidate with a reason. Gates are never averaged into scores.
    Input
    Candidates and gate functions (code, command or prompt gate).
    Output
    Passing candidates and a rejection record per failure.
    Stops when
    Every candidate has passed or been rejected.
  3. 03

    Cascade

    Question
    Which candidates deserve the expensive judges?
    Responsibility
    Score in stages; each stage keeps its top `keep_top` candidates for the next. Cached scores are keyed by scorer version.
    Input
    Passing candidates, staged scorers with costs, budgets for cost, time and money.
    Output
    Per-stage scores; a scorer error becomes `Score(None, error=...)`.
    Stops when
    The last stage finishes or a budget cap is reached.
  4. 04

    Aggregate

    Question
    How do several scores become one total?
    Responsibility
    Combine with weights; a `missing` policy decides what a `None` does (renormalize by default).
    Input
    Stage scores, weights, missing policy, floor.
    Output
    Totals with their parts.
    Stops when
    Every finalist has a total or is marked unscorable.
  5. 05

    Escalate

    Question
    Is the top of the ranking too close or too uncertain to trust?
    Responsibility
    Send near-ties within `tie_margin`, or low-confidence scores, to a pairwise judge, asked in both orders; a judgement counts only when both orders agree.
    Input
    The contested candidates and a pairwise judge.
    Output
    An agreed preference or "no agreement".
    Stops when
    The tie is settled or left explicitly unsettled.
  6. 06

    Select and record

    Question
    Which candidate wins, and can we explain it later?
    Responsibility
    Apply the policy (`argmax`, `first_above`, `pairwise_tournament`); apply a fallback when everything was rejected; write spans for every step.
    Input
    Totals, escalation results, policy, fallback.
    Output
    Winner (or none) and a span record in `.hone/select/spans.db`.
    Stops when
    The result and its record are written.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
Gate fails Reject the candidate A broken rule can't be compensated by quality Every candidate is rejected: apply the fallback
Scorer raises or returns junk Record `None` with the error "Couldn't score" is not "bad" Too many `None` values leave no usable total
Candidate below stage cut Drop before expensive stages Expensive judges see only finalists The cut removes a candidate a person would keep
Top two within `tie_margin` Ask a pairwise judge in both orders Close totals are noise-level differences The two orders disagree
Judge shares the generator's model family Warn Self-preference bias The warning is ignored in production
Budget reached Stop and select from what exists Cost caps are hard limits The stop happens before any finalist is scored

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Generate, gate, cascade, aggregate, select a clear leader.

Final state
Winner with a recorded reason.
Owner
The calling code or workflow step.
Evidence
Spans for generation, gates, each stage's scores and the selection.
Recovery
`explain` the run; rerun with a seed or changed scorer version if needed.
Ambiguous path

Two finalists within the tie margin go to a pairwise judge asked in both orders.

Final state
Winner after escalation, or the policy's choice with "no agreement" recorded.
Owner
The calling code; a person when the result is used for something consequential.
Evidence
Both pairwise answers and whether they agreed.
Recovery
Human A/B judgement in the dashboard, blind to the setup.
Failure path

All candidates fail gates or cannot be scored.

Final state
Fallback winner flagged as fallback, or no winner.
Owner
The calling code.
Evidence
Rejection reasons and scorer errors per candidate.
Recovery
Fix the generator or gate, or choose a different fallback (`best_rejected`, `first_valid`, `none`).

Authority map

Capability does not grant authority.

RULE

May decide
Reject a candidate through a gate; cut candidates between stages
May not decide
Turn a failed check into a low score
Required evidence
Gate name, result and detail

MODEL

May decide
Provide a score, a gate verdict or a pairwise preference
May not decide
Pick the winner alone, or have a failure count as 0
Required evidence
Judge ID, answer, parse result, error

SYSTEM

May decide
Aggregate, escalate, apply the policy and fallback, record
May not decide
Hide a fallback or an escalation
Required evidence
Spans for every step, config hash

HUMAN

May decide
Define gates, scorers, weights, policy; judge A/B blind
May not decide
Be bypassed when the result is consequential
Required evidence
Configuration file and experiment approval

EXCEPTION

May decide
Return no winner or a flagged fallback
May not decide
Present a fallback as a normal win
Required evidence
Fallback flag, rejection reasons

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Judge output unparseable Parse error `Score(None, error)` Fix the prompt or client; rerun Pipeline author
Judge biased toward first position Pairwise orders disagree Judgement not counted Different judge; A/B check Pipeline author
Every candidate rejected Empty pass set Fallback flagged Adjust generator or gate Pipeline author
Budget exhausted early Budget span Select from scored candidates Raise cap or cheapen early stages Pipeline author
Scorer changed behaviour Cache keyed by scorer version Old scores not reused Re-score Pipeline author
Same-family judge Engine warning Visible in records Use another model family Pipeline author

Invariants and guarantees

Properties the structure is designed to preserve.

  • "Could not score" is None, never 0.
  • Gates reject; they are never averaged into scores.
  • A pairwise judgement counts only when both orders agree.
  • Candidate IDs and variations are deterministic for a given seed.
  • Every run can be explained from the span store alone.
  • A fallback winner is always flagged as a fallback.

What changes between implementations

The constraints determine the final mechanism.

The scorers change with the domain: word counts and rhyme checks for lyrics, compile-and-test for code, taste models for images. The judges change with the budget: a local model through hone-models, any OpenAI-compatible server, LangChain, or a human. The policy changes with latency needs: first_above stops generating once a candidate is good enough.

For comparisons of models, prompts or settings over fixed test cases, hone-select also runs experiments: a declared experiment.toml, a plan with an estimate that a person approves before anything runs, resumable execution, and results that say which setup wins against a baseline with intervals. Run conditions keep a busy machine out of the numbers.

Open-source alternatives I compared

Tool What it does well Why it didn't fit this job
DSPy BestOfN / Refine Good stopping rules Tied to DSPy modules and one reward function
Inspect AI Clean solver/scorer split Built for evaluation, not for picking a winner in production
DeepEval (G-Eval), Promptfoo Rubrics and LLM judges No selection engine
Instructor's Universal Self-Consistency One selection strategy One strategy, not an engine with gates, cascades and records

None combined gates, cascaded scorers, code-or-prompt scorers, pluggable judges, selection policies, tie escalation, budgets, caching and a full record of the decision in one small library.

Evidence chain

Follow the pattern into systems and software.

Implemented in hone-select (Apache-2.0, alpha 0.1.0) with acceptance tests in CI. The screenshot shows its dashboard over records from the package's own offline examples with scripted judges; no production result is claimed.

The small-model marketplace design uses the experiments feature to compare rules, a large model and a small model on a locked test set before choosing. The model-routed code review design is related through its refusal to let one model's opinion stand as a verdict. Both are reference designs.

Implemented in hone-select, Apache-2.0, version 0.1.0 (alpha). The core depends only on pydantic; adapters for OpenAI-compatible servers, LangChain and hone-models are optional extras. The design, the guarantees it tests as acceptance cases, and thirteen recorded design changes (among them the dashboard, experiments, run conditions and blind A/B judgement) are in the repository.

The screenshot is the real hone-select dashboard, which I ran locally on 2026-09-30 over records produced by the package's own offline examples. The judges in those examples are scripted fakes, so the scores illustrate the mechanism, not model quality. The run shown generated four drafts, rejected one at a gate, scored the rest in two stages, escalated a near-tie (0.95 against 0.86 with a 0.1 margin) to a pairwise judge, and recorded the winner.

Known boundaries

Limitations and non-fit

  • Not an orchestrator: running selection inside a larger workflow is the caller's job (hone-flow can do that).
  • Not a model client: it never calls a model itself.
  • Not an evaluation benchmark for model releases; experiments compare setups for your task.
  • LLM judges remain noisy and biased; pairwise agreement and cross-family judges reduce the damage, they don't remove it.
  • Best-of-N multiplies generation cost by N. It is worth it when a better pick matters more than the extra calls.
  • Alpha: the API and config format may change before 1.0.

Related patterns

Continue through the adjacent decision structures.

Selection goes wrong in the gaps between steps: a failure that turns into a number, a rule that turns into a weight, a tie that turns into list order. Naming each step and recording it closes those gaps, and it makes "why did this win?" a question with an answer.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real