Slop score
Finds phrasing that language models overuse, with character spans.
Why. Generic phrasing is the most common taste failure in generated text.
Boundary. Lists are domain presets, not a universal truth.
Taste scoring
Stand-in scorers for human judgment: taste models, human-likeness checks, audience panels and personal taste.
Prevents: Outputs that pass every correctness check and are still generic, and human rating that doesn't scale.
Why it exists
Correctness metrics don't say whether people would like an output.
Scope. Scorers only; selection and model calls are left to the caller.
Five-minute orientation
pip install "git+https://github.com/honeworks/hone-taste"
Real output
Architecture
Four scorer families behind one Score type; heavy models are optional extras that load lazily inside an injected GPU lease.
Core concepts
Operations
`hone-taste score / models / agreement / profile`.
Current boundary
Alpha. Some wrapped models have non-commercial or unclear licences and refuse to load without accept_license=True. Stand-ins carry the biases of whoever rated their training data.
Near-term roadmap. Follow the design history in design/changes/.
What it implements
Engineering checklist
Need the control, not just the component?
The repository exposes the mechanism. Production work is defining the permissions, data, failure costs, evidence, and owners around it.