Gate
A pass/fail check that rejects a candidate.
Why. Rules that can't be traded off shouldn't be scores.
Boundary. A gate never contributes to a total.
Best-of-N selection
Generate, score and select the best of N candidates with pluggable code or prompt scorers, and record why the winner won.
Prevents: A failed judge scores a good candidate 0, a hard rule gets averaged away, the expensive judge reads everything, and nobody can explain the winner later.
Why it exists
Best-of-N is dependable but every project rewrites it by hand with the same bugs.
Scope. One reusable selection step and an experiments harness; not an orchestrator or model client.
Five-minute orientation
pip install "git+https://github.com/honeworks/hone-select"
Real output
Architecture
A small engine over pydantic: generator, gates, staged scorers, aggregation, policy and fallback, with ports for judges, text clients, embedders, record sinks and score caches. Adapters for OpenAI-compatible servers, LangChain and hone-models are optional extras.
Core concepts
Operations
Records every step to a local SQLite span store; `hone-select explain` and `hone-select dashboard` read it.
Current boundary
Alpha. LLM judges stay noisy; pairwise agreement and cross-family judges reduce, not remove, bias. Best-of-N multiplies generation cost.
Near-term roadmap. Follow the design history in design/changes/.
What it implements
Engineering checklist
Need the control, not just the component?
The repository exposes the mechanism. Production work is defining the permissions, data, failure costs, evidence, and owners around it.