Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Marketplace trust and safety · Reference design

A small model for listing review

The big model was right. At full volume, it was also unaffordable.

Settle whether a small model is good enough with a locked comparison against rules and a large model, then use it only to route listings, never to enforce.

Company Kestrel is a hypothetical company. Every figure is illustrative.

A grid of printed product listings with pencil marks separating flagged, allowed and unsure items.

About Company Kestrel

Company Kestrel is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed.

Industry
Online marketplace for health, wellness and personal-care products
Size
About 70 people; twelve trust and safety reviewers and a policy lead
Systems
Listing service, a keyword rules engine, a review-queue tool
Volume
About 25,000 new or edited listings a week in six languages
Constraints
Statements of reasons and appeals for restrictions; no automatic removal for health claims; about a tenth of a large hosted model's inference budget

In brief

ProblemKeyword rules flag harmless listings and miss coded health claims; a large hosted model handles both but costs too much at full volume.
SystemA four-way comparison on a leakage-resistant test set with a pre-written decision rule; if adopted, a calibrated classifier that routes candidates and abstentions to reviewers.
People decideReviewers decide every restriction and appeal; the policy lead owns the taxonomy, labels and thresholds.

How the work ran before

Company Kestrel runs an online marketplace for health, wellness and personal-care products with a few thousand third-party sellers. It gets about 25,000 new or edited listings a week, in six languages. A trust and safety team of twelve reviewers and one policy lead keeps prohibited listings off the site. The most common problem is a health claim the seller can't back up: "cures joint pain", "reverses hair loss".

The first line of defence is a keyword engine. It flags about 15% of listings. Reviewers clear most of those flags as harmless, because "pain" and "recovery" appear in plenty of legitimate listings. Meanwhile, sellers who know the rules write around them: "supports your body's natural recovery from…" passes the keywords every time.

Where it broke

Kestrel's trust and safety team tried a large hosted model on a sample. It was good: it understood coded claims, explained its reasoning, and handled all six languages. Then they did the arithmetic. At full volume, with each listing's title, description and bullet points, the monthly cost was far outside the trust and safety budget, and the response time was too slow to review listings before they went live.

So the question became: is there a smaller model, cheap enough to run on every listing, that is good enough? Everyone had an opinion. Nobody had evidence, and a vendor demo on hand-picked examples isn't evidence.

What I would build

First, an experiment that answers the question honestly. Then, only if the answer is yes, a narrow classifier with clear limits.

The experiment compares four approaches on the same locked test set:

  1. the current keyword rules (the baseline the others must beat);
  2. the large prompted model (the quality ceiling we can afford to test but not to run);
  3. a small model with a prompt only;
  4. the same small model, fine-tuned on labelled listings.

If option 3 is close to option 2, fine-tuning isn't needed. If none of the small options beats the rules by a margin that matters, the answer is "improve the rules", and that is a perfectly good outcome.

The classifier, if it earns its place, does one thing: for each listing, propose a policy class (for example "unverified medical claim") with the sentence that triggered it, or "no candidate violation". It never removes a listing, never decides an appeal, and never penalizes a seller. Candidates go to reviewers. Clear negatives go live and a sample is audited.

Listings in a language or category the model wasn't tested on go to people, not to the model's best guess.

Following three listings through

Challenge cases for the test set. Constructed examples, not model outputs.

Listing text (translated) What it tests Expected route
"Herbal capsules that cure arthritis in 30 days." Explicit claim Candidate: unverified medical claim, sentence quoted
"Supports your body's natural recovery from joint inflammation." Coded claim, no trigger word Candidate, if the model understands meaning; the keywords miss it
"A book about the history of arthritis treatment, from willow bark to aspirin." Allowed, same vocabulary No candidate; a model that flags this is copying the keywords

Constructed challenge cases.

The third row matters as much as the first two. A model that flags every listing mentioning a disease gives reviewers the keyword engine's queue again, at a higher cost.

Company Kestrel · challenge set · unverified medical claim three constructed cases · no model outputs shown

Explicit claim Product photo of an amber glass supplement bottle with a blank label.

Herbal capsules: cure arthritis in 30 days

Constructed listing text

Expected route: candidate, sentence quoted

Coded claim Product photo of a black neoprene knee support brace.

Joint support: natural recovery from joint inflammation

Constructed listing text

Expected route: candidate; keywords miss it

Allowed Product photo of a hardcover textbook with a plain blue cover.

History of arthritis treatment: willow bark to aspirin

Constructed educational listing

Expected route: no candidate

Constructed challenge cases for Company Kestrel, a hypothetical company. Photos are illustrations, not real listings.

When things go wrong

The model is unsure. It shouldn't pretend. The classifier's score is calibrated on held-out data, and listings between two thresholds go to review. A raw model score isn't a probability until it has been checked against real outcomes.

A new trick appears. Sellers adapt. A new phrasing that works around the model shows up in reviewer overrides and appeals first. Those feed the next labelled batch, and the locked test set gets a new challenge slice.

The language is unsupported. If the test set had no Polish supplements, Polish supplement listings route to reviewers. The model isn't allowed to extrapolate into places where its errors have never been measured.

The policy changes. A new version of the policy means new labels. The model is retested against the new version before it's allowed to route anything under it.

What changes for the team

Step Before After (if the experiment says yes)
Flagging Keyword hits, mostly harmless Candidates with the triggering sentence quoted
Coded claims Pass unnoticed Caught when the model understands them
Clearly allowed listings Also flagged by keywords Go live; audited by sampling
Removal, appeals, penalties Reviewers Reviewers, unchanged

Illustrative figures from assumptions in the technical design. Not measured.

Illustrative week (25,000 listings) Keyword rules Small classifier with abstention
Listings sent to reviewers about 3,750 about 1,600 (candidates plus abstentions)
Of those, true policy problems about 450 about 500
Coded claims missed many, not counted fewer; measured in the pilot

These numbers are placeholders for what the experiment must measure. The design's promise is only that the decision will be made with a comparison, not an opinion.

What this does not solve

A small model won't match the large one on the rarest and most creative cases; those still need people. The classifier's quality is capped by label quality, and labelling needs reviewer time. It doesn't make legal decisions about what a claim means in each country. And a fine-tuned model is a thing to maintain: retraining when policy changes, monitoring when seller behaviour drifts.

How I would prove it

That's the experiment itself: 3,000 listings labelled by two reviewers each, with disagreements settled by the policy lead, split so that no seller or listing template appears in both training and test. Report results per policy class, per language, and on the challenge slices, with confidence intervals. The decision rule is written down before the results come in.

The technical design

The same system for engineers: architecture, records, failure handling, evaluation, the options I rejected, and the stack. About 4 minutes.

  1. Scope and assumptions
  2. Evaluation: the experiment
  3. Decision rule (written before results)
  4. Architecture of the classifier service
  5. Failure handling
  6. Monitoring after launch
  7. Trade-offs I considered
  8. Stack
  9. Security and operations

Read the technical design

Is a smaller model good enough?

Bring your label set and a sample.

The policy taxonomy, a few hundred labelled items and your current process are enough to design the comparison before anyone picks a model.