Problem shape
The problem beneath them
Score outputs for taste with cheap stand-ins first (taste models trained on human ratings, human-likeness checks, simulated audience panels, and a profile learned from a few personal picks), return one score shape with a reason and an honest statement of whose taste it reflects, treat failures as missing, and spend real human attention only on the closest calls.
Why I built it
My song pipeline could check that lyrics were singable, that JSON parsed and that audio didn't clip. It couldn't tell me whether the chorus was any good, and every draft leaned on the same tired phrases. Rating everything myself didn't scale. The building blocks existed but were scattered: models trained on thousands of human ratings of songs, text and images; lists of overused LLM phrasing; zero-shot AI-text detectors; persona simulations. Each had its own code, input format and score range, and most came with no clear statement of whose taste they had learned or what their licence allowed.
hone-taste puts them behind one Score: a value from 0 to 1, a confidence, a reason, per-aspect details, and an error. That makes them combinable, lets them plug straight into best-of-N selection, and lets a few of your own picks check which of them agree with you.
Stand-in scores estimate taste; they don't replace it. Detectors are signals, never gates, and optimizing hard against one taste model finds outputs that fool it.
hone-taste scores; it doesn't select, call models on its own, or train taste models from scratch. It is not a tool for making AI text undetectable: human-likeness is used to improve quality, not to evade disclosure.
Different problems, same shape
The surface changes. The decision structure persists.
Every draft passes the checks. Why do they all sound the same?
Metrics measure correctness, not whether people would like it.
Which of these twelve renders would people pick?
A person rates every batch by hand.
The audio doesn't clip. Is the song any good?
Objective audio checks are the only automatic signal.
Can the system learn what this one editor likes from a few picks?
Preference learning asks for a hundred comparisons before it helps.
All four questions converge into Stand-in Human Judgment.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Human-likeness checks
- Question
- Does this read like generic model output?
- Responsibility
- Find overused phrasing (`slop_score`, with character spans), your own banned patterns, and optionally a zero-shot detector signal (Binoculars).
- Input
- Text and a domain preset (general, lyrics, email).
- Output
- Score with the phrases found.
- Stops when
- The text is scored or an input error is returned.
-
02
Taste models
- Question
- What would many people say about this, going by models trained on their ratings?
- Responsibility
- Wrap small models trained on human ratings (song aesthetics, audio aesthetics, text reward, image preference) and objective audio checks; load lazily, run inside an injected GPU lease, free on `close()`.
- Input
- Text, image or audio, plus the model's extra installed.
- Output
- Score with per-aspect details.
- Stops when
- Scored, or `None` with an error (wrong input kind, licence not accepted).
-
03
Audience panel
- Question
- How would the target audience react?
- Responsibility
- Have an LLM play several personas; each rates with a reason. Panels count less than models trained on real ratings.
- Input
- Personas, a question, and any decision or text client.
- Output
- Combined score with each persona's reason.
- Stops when
- Every persona answered or the missing ones are recorded.
-
04
Personal taste
- Question
- What does this one person like?
- Responsibility
- Learn a profile from 5 to 10 pairwise picks; re-weight the other scorers; ask next about the pairs the scorers are least sure of. Style scorers learn an author's voice from about ten texts.
- Input
- Picks or sample texts.
- Output
- A profile that scores like any other scorer.
- Stops when
- The profile exists; it improves with more picks.
-
05
Combine and check agreement
- Question
- How much do these stand-ins agree with a real person?
- Responsibility
- Combine scorers with weights, dropping missing values; report how often each scorer prefers what the person picked.
- Input
- Scorers, weights, and a person's picks.
- Output
- A combined scorer and an agreement report.
- Stops when
- Scores and agreement are recorded as spans.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Scorer fails or input is the wrong kind | Return `None` with an error | Missing is not bad | Too many missing values to rank |
| Model licence unclear or non-commercial | Refuse unless `accept_license=True` | The user must read the terms | Commercial use is intended |
| Detector says "likely AI" | Record as a signal only | Detectors are unreliable as gates | Someone wants to gate on it |
| Persona panel disagrees with taste model | Weight the panel lower | One model pretending is weaker evidence | A person's picks side with the panel |
| Scorers disagree on the top two | Ask the person about that pair | The closest calls teach the most | The person has no time: keep both |
| One scorer dominates selection | Combine families | Single-model optimization is gamed | Outputs start to look alike |
Operating paths