The default move in 2026 is to send every language task to the biggest general model you can afford. For open-ended writing and reasoning that's often right. For narrow, repetitive jobs (finding names in text, filling a JSON template, ranking passages, transcribing speech) there are small models trained for exactly that job, and some of them match or beat general models many times their size. Most teams never hear of these small task-specific models because they don't top chat leaderboards.
Four models worth knowing
GLiNER for entity extraction
GLiNER treats named-entity recognition as matching entity-type descriptions to spans of text, using a small bidirectional encoder instead of a generative model. You pass the text and the labels you care about ("invoice number", "medication", "supplier"), including labels it never saw in training. In the NAACL 2024 paper, the smallest version (about 50M parameters) outperformed ChatGPT on zero-shot NER benchmarks, and a 90M version was comparable to UniNER-13B, which is about 140 times larger (Zaratiana et al., 2024).
It runs on a CPU and returns spans with scores. That makes it a good first pass before a larger model, and an easy thing to evaluate.
NuExtract for structured extraction
NuExtract models are fine-tuned specifically to fill a JSON template from text: you pass the document and an empty template, and it returns the filled template. The first version was a fine-tuned Phi-3-mini; later versions added more template types (dates, enums, verbatim strings), and the current release is a 4B vision-language model that also reads document images (NuExtract model page, NuExtract3).
For extraction, a model trained to leave fields empty when the text doesn't support them is worth more than a bigger model that writes plausible values. You still need the checks: arithmetic, reference data, page evidence.
A reranker for retrieval
A cross-encoder reranker reads the query and one passage together and outputs a relevance score. It's slower than vector search per pair, so you use it on the top 20 or 50 results to reorder them. BAAI's bge-reranker-v2-m3 is a multilingual example with about 568M parameters under an Apache-2.0 licence.
When retrieval finds the right passage but ranks it fifth, a reranker usually fixes more than a bigger embedding model or a bigger generator would.
Distil-Whisper for speech
Distil-Whisper keeps Whisper's encoder and cuts the decoder to two layers. Its model card reports that distil-large-v3 is 6.3 times faster than Whisper large-v3 and within 1% word error rate on long-form English audio (model card). For English transcription at volume, that's the difference between one GPU and several.
Why small specialists can win
A general model spends its capacity on everything: code, poems, trivia, chat. A specialist spends all of its capacity on one input and output shape, trained on data built for that shape. It also constrains the output: an NER model can only return spans that exist in the text, which removes a whole class of hallucination by construction.
They also change the economics. They run on CPUs or small GPUs, answer in milliseconds, and can sit inside a private network without a GPU cluster. For a workflow that runs a million times a month, that matters more than a benchmark point.
Where they lose
- Anything open-ended: explanations, summaries with judgement, writing.
- Tasks that need world knowledge the small model never saw.
- Inputs that drift far from their training data: a new document layout, a new language, a new domain vocabulary.
- Anything where the task definition keeps changing. A prompt is faster to edit than a fine-tune.
How to decide
Put the specialist and your current general model side by side on the same test set, the way I describe in the marketplace review design study: rules first as a baseline, then the large model, then the small one. For extraction, score each field. For ranking, measure recall at k. For speech, word error rate on your own audio, including the numbers and negations that matter to you.
Often the best system uses both: the specialist handles the bulk and flags low-confidence cases, and the general model or a person handles those. That split is usually cheaper than either model alone, and it gives you a place to look when quality drops.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real