Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Taste scoring

Estimate what people would like, and say whose taste the estimate came from.

Stand-in scorers for human judgment: taste models, human-likeness checks, audience panels and personal taste.

Prevents: Outputs that pass every correctness check and are still generic, and human rating that doesn't scale.

Why it exists

The failure appears after the happy path.

Correctness metrics don't say whether people would like an output.

Scope. Scorers only; selection and model calls are left to the caller.

Five-minute orientation

See the smallest complete path.

pip install "git+https://github.com/honeworks/hone-taste"

Real output

What it looks like running.

Terminal: hone-taste scores a cliched lyric line 0.173 and a concrete line 1.000, then lists its wrapped models with licences and training data.
Real CLI output on 2026-09-30, using the two lines from the package's quickstart.

Architecture

The architecture is a set of promises.

Four scorer families behind one Score type; heavy models are optional extras that load lazily inside an injected GPU lease.

  • Could-not-score is None, never 0. Inspect test
  • Each wrapped model states whose ratings it learned and its licence. Design note
  • Detectors are signals, never gates. Design note

Core concepts

The few concepts you need before reading the code.

Slop score

Finds phrasing that language models overuse, with character spans.

Why. Generic phrasing is the most common taste failure in generated text.

Boundary. Lists are domain presets, not a universal truth.

Personal profile

A person's taste learned from 5 to 10 picks.

Why. Ask people rarely, about the closest calls.

Boundary. It re-weights other scorers; it isn't a trained reward model.

Operations

What happens after install.

`hone-taste score / models / agreement / profile`.

Current boundary

What this project does not solve.

Alpha. Some wrapped models have non-commercial or unclear licences and refuse to load without accept_license=True. Stand-ins carry the biases of whoever rated their training data.

Near-term roadmap. Follow the design history in design/changes/.

What it implements

  • one Score shape with reason and confidence
  • slop and pattern checks
  • taste-model wrappers
  • LLM audience panels
  • personal profiles from a few picks
  • combining and agreement checks
  • licence gating

Engineering checklist

What the repository ships.

Quickstart
Present
Success tests
Present
Failure tests
Present
Architecture notes
Present
Security notes
Not published
Runbook
Not published
Limitations
Present
Changelog
Present

Recent changes

What changed, with the commits.

  • Initial release: hone-taste 0.1.0 Commit
  • Claude Code workflow added: skills, hooks, settings Commit

Need the control, not just the component?

Fit it to the system that has to survive.

The repository exposes the mechanism. Production work is defining the permissions, data, failure costs, evidence, and owners around it.

Let's build something real