Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Technical design

A small model for listing review

The big model was right. At full volume, it was also unaffordable.

← Back to the story

Company Kestrel is a hypothetical company, written for this reference design; the details below are its constraints. Every figure is illustrative, computed from stated assumptions. No client work or measured result is claimed. The story explains the problem and follows one item through the system; this page is the engineering detail behind it.

Scope and assumptions

Two deliverables: a comparison experiment, and, conditionally, a classifier service that routes listings to review. Enforcement, appeals and seller communication stay in the existing review tool.

Illustrative figures assume: 25,000 listings a week; keyword rules flag 15% (3,750) of which 12% are true problems (450); the classifier sends 4% as candidates (1,000) with 45% true problems (450), plus 2.4% abstentions (600) of which about 8% are problems (≈50). Totals: 1,600 to review, about 500 true problems found. These are assumptions for planning the experiment, not results.

Evaluation: the experiment

labelled pool (3,000 listings, double-labelled, adjudicated)
      │ group split by seller and template; time-based holdout for the last 4 weeks
      ▼
 train (fine-tuning only) · validation · calibration · locked test (+ challenge slices)
      │
      ▼
 four setups × same test set ──► per-class precision/recall, hard-negative errors,
                                  calibration, abstention, latency, cost per 1,000 listings
  • Labels. Each listing gets a policy class (from a versioned taxonomy with stable IDs), the sentence that triggered it, and the policy version. Two reviewers label independently; the policy lead settles disagreements. Agreement between reviewers is reported, because it's the ceiling for any model.
  • Splits. Group splits by seller and by listing template prevent the model from memorizing a seller's wording. A time-based holdout checks drift.
  • Challenge slices. Coded claims, allowed hard negatives (books, historical content, cosmetics that mention skin conditions), each language, and each rare class, reported separately.
  • Setups. Rules; large hosted model with the policy in the prompt; a small open-weight model (7B to 9B class) prompted; the same model fine-tuned with LoRA. Quantized variants are tested separately, because quantization can change behaviour on exactly the edge cases that matter.

I would run this with hone-select's experiments feature, which I built for comparisons like this: it declares setups and test cases in a file, shows the plan and cost before running, requires an approval step, and reports which setup wins against a baseline with intervals. Any harness that locks the test set and reports intervals will do.

Decision rule (written before results)

Adopt the small classifier only if, on the locked test set:

  • recall on "unverified medical claim" is at least as good as the large prompted model's minus 5 points;
  • hard-negative false-positive rate is below the keyword engine's;
  • calibration error is small enough that thresholds can be set per class;
  • cost per 1,000 listings is within budget at the p95 latency the review flow needs.

If it fails, report why and recommend the next best option (often: better rules for the explicit cases plus the large model on a small, pre-filtered slice).

Architecture of the classifier service

listing created/edited ──► language + category check ──► supported? ─no──► review queue
                                                              │yes
                                                              ▼
                                                  classifier (candidate class, sentence, score)
                                                              │
                                              calibrated score per class
                                   ┌──────────────┼───────────────────┐
                                   ▼              ▼                   ▼
                          candidate (≥ high)  abstain (between)   no candidate (< low)
                                   │              │                   │
                              review queue    review queue      publish; sampled audit

The service returns structured output: class ID, policy version, triggering sentence (verified to exist in the listing), and calibrated score. Thresholds are per class and set on the calibration split. The model version, taxonomy version and thresholds are recorded with every routing decision, so a statement of reasons can cite them.

Failure handling

Failure Response
Unsupported language or category Route to review, never to the model
Triggering sentence not found in listing Treat as abstention
Model service down Fall back to keyword rules; queue grows, nothing auto-publishes that rules would flag
Policy version changed Classifier disabled for the changed classes until retested
Reviewer overrides spike for one class Alert policy lead; lower that class's auto-publish threshold

Monitoring after launch

Weekly: reviewer override rate per class, appeal outcomes, abstention rate, and a random audit of 200 auto-published listings reviewed blind. A drop in audit quality or a rise in overrides triggers a retest on fresh labels. Retraining happens on a schedule, not on every complaint.

Trade-offs I considered

  • Large model on everything. Best quality, unaffordable at volume. Kept as the quality reference.
  • Large model only on keyword-flagged listings. Cheaper, but inherits the keywords' blind spot for coded claims.
  • Embedding similarity to known violations. Fast and cheap; good as a feature, weak on new phrasings and hard negatives.

Stack

A small open-weight model served with vLLM or llama.cpp on one mid-range GPU (or a cheap hosted endpoint), LoRA fine-tuning with a standard training library, Postgres for labels and routing decisions, the existing review tool's API for queues, and hone-select (or an equivalent harness) for the comparison. Latency and cost are measured on the actual serving setup, not quoted from benchmarks.

Security and operations

Listings are public content, but reviewer notes and seller data are not; the model sees only listing text. Every restriction a reviewer applies keeps the model's suggestion and the reviewer's reason, so statements of reasons and appeals can show both.