Marketplace trust and safety · Reference design
A small model for listing review
The big model was right. At full volume, it was also unaffordable.
Settle whether a small model is good enough with a locked comparison against rules and a large model, then use it only to route listings, never to enforce.
Company Kestrel is a hypothetical company. Every figure is illustrative.
About Company Kestrel
Company Kestrel is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed.
- Industry
- Online marketplace for health, wellness and personal-care products
- Size
- About 70 people; twelve trust and safety reviewers and a policy lead
- Systems
- Listing service, a keyword rules engine, a review-queue tool
- Volume
- About 25,000 new or edited listings a week in six languages
- Constraints
- Statements of reasons and appeals for restrictions; no automatic removal for health claims; about a tenth of a large hosted model's inference budget
In brief
How the work ran before
Company Kestrel runs an online marketplace for health, wellness and personal-care products with a few thousand third-party sellers. It gets about 25,000 new or edited listings a week, in six languages. A trust and safety team of twelve reviewers and one policy lead keeps prohibited listings off the site. The most common problem is a health claim the seller can't back up: "cures joint pain", "reverses hair loss".
The first line of defence is a keyword engine. It flags about 15% of listings. Reviewers clear most of those flags as harmless, because "pain" and "recovery" appear in plenty of legitimate listings. Meanwhile, sellers who know the rules write around them: "supports your body's natural recovery from…" passes the keywords every time.
Where it broke
Kestrel's trust and safety team tried a large hosted model on a sample. It was good: it understood coded claims, explained its reasoning, and handled all six languages. Then they did the arithmetic. At full volume, with each listing's title, description and bullet points, the monthly cost was far outside the trust and safety budget, and the response time was too slow to review listings before they went live.
So the question became: is there a smaller model, cheap enough to run on every listing, that is good enough? Everyone had an opinion. Nobody had evidence, and a vendor demo on hand-picked examples isn't evidence.
What I would build
First, an experiment that answers the question honestly. Then, only if the answer is yes, a narrow classifier with clear limits.
The experiment compares four approaches on the same locked test set:
- the current keyword rules (the baseline the others must beat);
- the large prompted model (the quality ceiling we can afford to test but not to run);
- a small model with a prompt only;
- the same small model, fine-tuned on labelled listings.
If option 3 is close to option 2, fine-tuning isn't needed. If none of the small options beats the rules by a margin that matters, the answer is "improve the rules", and that is a perfectly good outcome.
The classifier, if it earns its place, does one thing: for each listing, propose a policy class (for example "unverified medical claim") with the sentence that triggered it, or "no candidate violation". It never removes a listing, never decides an appeal, and never penalizes a seller. Candidates go to reviewers. Clear negatives go live and a sample is audited.
Listings in a language or category the model wasn't tested on go to people, not to the model's best guess.
Following three listings through
Challenge cases for the test set. Constructed examples, not model outputs.
| Listing text (translated) | What it tests | Expected route |
|---|---|---|
| "Herbal capsules that cure arthritis in 30 days." | Explicit claim | Candidate: unverified medical claim, sentence quoted |
| "Supports your body's natural recovery from joint inflammation." | Coded claim, no trigger word | Candidate, if the model understands meaning; the keywords miss it |
| "A book about the history of arthritis treatment, from willow bark to aspirin." | Allowed, same vocabulary | No candidate; a model that flags this is copying the keywords |
Constructed challenge cases.
The third row matters as much as the first two. A model that flags every listing mentioning a disease gives reviewers the keyword engine's queue again, at a higher cost.
Herbal capsules: cure arthritis in 30 days
Constructed listing text
Expected route: candidate, sentence quoted
Joint support: natural recovery from joint inflammation
Constructed listing text
Expected route: candidate; keywords miss it
History of arthritis treatment: willow bark to aspirin
Constructed educational listing
Expected route: no candidate
When things go wrong
The model is unsure. It shouldn't pretend. The classifier's score is calibrated on held-out data, and listings between two thresholds go to review. A raw model score isn't a probability until it has been checked against real outcomes.
A new trick appears. Sellers adapt. A new phrasing that works around the model shows up in reviewer overrides and appeals first. Those feed the next labelled batch, and the locked test set gets a new challenge slice.
The language is unsupported. If the test set had no Polish supplements, Polish supplement listings route to reviewers. The model isn't allowed to extrapolate into places where its errors have never been measured.
The policy changes. A new version of the policy means new labels. The model is retested against the new version before it's allowed to route anything under it.
What changes for the team
| Step | Before | After (if the experiment says yes) |
|---|---|---|
| Flagging | Keyword hits, mostly harmless | Candidates with the triggering sentence quoted |
| Coded claims | Pass unnoticed | Caught when the model understands them |
| Clearly allowed listings | Also flagged by keywords | Go live; audited by sampling |
| Removal, appeals, penalties | Reviewers | Reviewers, unchanged |
Illustrative figures from assumptions in the technical design. Not measured.
| Illustrative week (25,000 listings) | Keyword rules | Small classifier with abstention |
|---|---|---|
| Listings sent to reviewers | about 3,750 | about 1,600 (candidates plus abstentions) |
| Of those, true policy problems | about 450 | about 500 |
| Coded claims missed | many, not counted | fewer; measured in the pilot |
These numbers are placeholders for what the experiment must measure. The design's promise is only that the decision will be made with a comparison, not an opinion.
What this does not solve
A small model won't match the large one on the rarest and most creative cases; those still need people. The classifier's quality is capped by label quality, and labelling needs reviewer time. It doesn't make legal decisions about what a claim means in each country. And a fine-tuned model is a thing to maintain: retraining when policy changes, monitoring when seller behaviour drifts.
How I would prove it
That's the experiment itself: 3,000 listings labelled by two reviewers each, with disagreements settled by the policy lead, split so that no seller or listing template appears in both training and test. Report results per policy class, per language, and on the challenge slices, with confidence intervals. The decision rule is written down before the results come in.
The technical design
The same system for engineers: architecture, records, failure handling, evaluation, the options I rejected, and the stack. About 4 minutes.
Is a smaller model good enough?
Bring your label set and a sample.
The policy taxonomy, a few hundred labelled items and your current process are enough to design the comparison before anyone picks a model.