Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 04 / Evaluation & Controlled Change

Ask people rarely, and ask them well.

Stand-in Human Judgment

Estimate whether people would like an output and whether it feels human-made with four kinds of stand-in scorers, all returning one score shape with a reason, and use a few human picks to check and re-weight them.

Maturity
Tested / Prototype
Critical uncertainty
Stand-in scores estimate taste; they don't replace it. Detectors are signals, never gates, and optimizing hard against one taste model finds outputs that fool it.
Human boundary
Make picks, set weights, accept licences
Applications
Lyrics and copy drafts · Generated images · Generated music · Shortlists before a human decision
Terminal output captured 2026-09-30: hone-taste scores a cliched lyric line 0.173 with two slop hits and a concrete line 1.000, then lists its wrapped taste models with their licences and whose ratings each learned from.

Problem shape

The problem beneath them

Score outputs for taste with cheap stand-ins first (taste models trained on human ratings, human-likeness checks, simulated audience panels, and a profile learned from a few personal picks), return one score shape with a reason and an honest statement of whose taste it reflects, treat failures as missing, and spend real human attention only on the closest calls.

Why I built it

My song pipeline could check that lyrics were singable, that JSON parsed and that audio didn't clip. It couldn't tell me whether the chorus was any good, and every draft leaned on the same tired phrases. Rating everything myself didn't scale. The building blocks existed but were scattered: models trained on thousands of human ratings of songs, text and images; lists of overused LLM phrasing; zero-shot AI-text detectors; persona simulations. Each had its own code, input format and score range, and most came with no clear statement of whose taste they had learned or what their licence allowed.

hone-taste puts them behind one Score: a value from 0 to 1, a confidence, a reason, per-aspect details, and an error. That makes them combinable, lets them plug straight into best-of-N selection, and lets a few of your own picks check which of them agree with you.

Stand-in scores estimate taste; they don't replace it. Detectors are signals, never gates, and optimizing hard against one taste model finds outputs that fool it.

hone-taste scores; it doesn't select, call models on its own, or train taste models from scratch. It is not a tool for making AI text undetectable: human-likeness is used to improve quality, not to evade disclosure.

Different problems, same shape

The surface changes. The decision structure persists.

01 / Copy and lyrics

Every draft passes the checks. Why do they all sound the same?

Metrics measure correctness, not whether people would like it.

02 / Images

Which of these twelve renders would people pick?

A person rates every batch by hand.

03 / Music

The audio doesn't clip. Is the song any good?

Objective audio checks are the only automatic signal.

04 / Personal taste

Can the system learn what this one editor likes from a few picks?

Preference learning asks for a hundred comparisons before it helps.

All four questions converge into Stand-in Human Judgment.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Human-likeness checks

    Question
    Does this read like generic model output?
    Responsibility
    Find overused phrasing (`slop_score`, with character spans), your own banned patterns, and optionally a zero-shot detector signal (Binoculars).
    Input
    Text and a domain preset (general, lyrics, email).
    Output
    Score with the phrases found.
    Stops when
    The text is scored or an input error is returned.
  2. 02

    Taste models

    Question
    What would many people say about this, going by models trained on their ratings?
    Responsibility
    Wrap small models trained on human ratings (song aesthetics, audio aesthetics, text reward, image preference) and objective audio checks; load lazily, run inside an injected GPU lease, free on `close()`.
    Input
    Text, image or audio, plus the model's extra installed.
    Output
    Score with per-aspect details.
    Stops when
    Scored, or `None` with an error (wrong input kind, licence not accepted).
  3. 03

    Audience panel

    Question
    How would the target audience react?
    Responsibility
    Have an LLM play several personas; each rates with a reason. Panels count less than models trained on real ratings.
    Input
    Personas, a question, and any decision or text client.
    Output
    Combined score with each persona's reason.
    Stops when
    Every persona answered or the missing ones are recorded.
  4. 04

    Personal taste

    Question
    What does this one person like?
    Responsibility
    Learn a profile from 5 to 10 pairwise picks; re-weight the other scorers; ask next about the pairs the scorers are least sure of. Style scorers learn an author's voice from about ten texts.
    Input
    Picks or sample texts.
    Output
    A profile that scores like any other scorer.
    Stops when
    The profile exists; it improves with more picks.
  5. 05

    Combine and check agreement

    Question
    How much do these stand-ins agree with a real person?
    Responsibility
    Combine scorers with weights, dropping missing values; report how often each scorer prefers what the person picked.
    Input
    Scorers, weights, and a person's picks.
    Output
    A combined scorer and an agreement report.
    Stops when
    Scores and agreement are recorded as spans.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
Scorer fails or input is the wrong kind Return `None` with an error Missing is not bad Too many missing values to rank
Model licence unclear or non-commercial Refuse unless `accept_license=True` The user must read the terms Commercial use is intended
Detector says "likely AI" Record as a signal only Detectors are unreliable as gates Someone wants to gate on it
Persona panel disagrees with taste model Weight the panel lower One model pretending is weaker evidence A person's picks side with the panel
Scorers disagree on the top two Ask the person about that pair The closest calls teach the most The person has no time: keep both
One scorer dominates selection Combine families Single-model optimization is gamed Outputs start to look alike

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Score with cheap checks and a taste model, combine, pass to selection.

Final state
Score with reason and confidence.
Owner
The calling code.
Evidence
`hone.taste.score` spans with details.
Recovery
Adjust weights after an agreement check.
Ambiguous path

Scorers disagree; the person is asked about the most uncertain pair.

Final state
Human pick recorded and used to re-weight.
Owner
The person whose taste matters.
Evidence
The pick, the pair, the scorers' prior preferences.
Recovery
More picks over time; the agreement report shows whether they help.
Failure path

Model unavailable, wrong input, or licence not accepted.

Final state
Missing score with an error.
Owner
The calling code.
Evidence
Error on the score.
Recovery
Combine without it, or install the extra and accept the terms knowingly.

Authority map

Capability does not grant authority.

RULE

May decide
Refuse unlicensed models, drop missing scores in combinations
May not decide
Count a failure as zero
Required evidence
Licence flag, error

MODEL

May decide
Estimate taste or human-likeness with a reason
May not decide
Gate content or stand in for the final human choice
Required evidence
Model ID, whose ratings it learned

SYSTEM

May decide
Combine scorers, compute agreement, choose which pair to ask about
May not decide
Hide which scorer drove the result
Required evidence
Details per aspect and scorer

HUMAN

May decide
Make picks, set weights, accept licences
May not decide
Be replaced silently by a persona panel
Required evidence
Picks and their context

EXCEPTION

May decide
Return `Score(None, error)`
May not decide
Masquerade as a low score
Required evidence
Error text

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Goodhart: outputs fool one taste model Agreement with human picks drops Combine families Re-weight, add picks Pipeline author
Detector false positive Signal only No gate on it Ignore or re-check Pipeline author
Persona panel stereotypes Reasons read as generic Lower weight Better personas or drop panel Pipeline author
Licence misuse `accept_license` required Model won't load Read upstream terms User
Heavy model on small GPU Lease wait Lazy loading, `close()` Run later or on CPU Operator
Taste drift Agreement report over time Visible in spans New picks Person

Invariants and guarantees

Properties the structure is designed to preserve.

  • Every scorer returns the same Score shape.
  • "Could not score" is None, never 0; combining drops missing values.
  • Each wrapped model states whose ratings it learned and its licence.
  • Models with unclear or non-commercial terms need explicit acceptance.
  • Detectors are signals, never gates.
  • Every score is recorded as a span unless recording is turned off.

What changes between implementations

The constraints determine the final mechanism.

The media decide the families: text uses slop checks, reward models and panels; images use preference models and character consistency; songs use aesthetics models and audio checks. The budget decides the order: the core (slop, patterns, panels, profiles, combining) needs no GPU and no torch; the taste models are optional extras. The person decides the weights through a handful of picks.

Open-source alternatives I compared

There was no single alternative covering this job, only parts: individual taste models released with research papers (each with its own loader and score range), public lists of overused LLM phrases, zero-shot detectors such as Binoculars, and evaluation frameworks with LLM-judge rubrics. hone-taste wraps several of those parts rather than replacing them. What it adds is the shared score shape, the honest provenance and licence notes, missing-as-None, and the small loop that checks stand-ins against a real person's picks.

Evidence chain

Follow the pattern into systems and software.

Implemented in hone-taste (Apache-2.0, alpha 0.1.0) with acceptance tests in CI. Some wrapped models have non-commercial or unclear licences. The screenshot shows real CLI output; no production result is claimed.

The small-model marketplace design is about policy classification rather than taste, but it shares the stance: automated scores route work to people and are checked against human labels before anyone trusts them. It's a reference design.

Implemented in hone-taste, Apache-2.0, version 0.1.0 (alpha). It ships no model weights; each taste model downloads from its upstream source on first use. The README warns prominently that SongEval's licence is unclear, that the MuQ encoder weights are non-commercial, and that PickScore states no licence, and those models refuse to load without accept_license=True.

The screenshot is real CLI output from 2026-09-30: hone-taste score --domain lyrics gives a cliched line 0.173 (two overused phrases in 13 words) and a concrete line 1.000, and hone-taste models lists each wrapped model with its licence and whose ratings it learned from. The two lines are the package's own quickstart examples.

  • TestedStand-in Human JudgmentThis pattern, as built in the open-source code below.
  • Reference designA small model for listing reviewRelated case-study system.
  • Alphahone-tasteSource: https://github.com/honeworks/hone-taste

Known boundaries

Limitations and non-fit

  • Stand-ins are estimates of other people's taste; they carry those people's biases.
  • No video taste and no web page for collecting picks in this version; the CLI asks.
  • Taste models need a GPU to be fast and some have licence limits.
  • It's not an experiment or calibration platform; the agreement check is deliberately small.
  • For outputs where correctness is all that matters, you don't need this.

Related patterns

Continue through the adjacent decision structures.

Taste can't be automated away, but it can be spent carefully. Cheap stand-ins sort the obvious cases, their reasons make them checkable, and the person's attention goes to the few comparisons where it changes the outcome.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real