"Open models are a year behind the frontier" is true on broad benchmarks and often irrelevant to a specific job. A business process rarely needs a model that can do everything. It needs one task done well, thousands of times, at a cost and privacy level the business accepts. On a narrow task, a system built from open-source LLMs can get close to a frontier model's quality. The system is what gets there; no single open model does.
This piece is the method I use to find out how close, for one task, before anyone commits to an architecture. It's written as an experiment you can reproduce. I'm deliberately not giving you my numbers, because they would describe my tasks, my hardware and my test sets, and they'd mislead you on yours.
Step 1: define the task narrowly enough to measure
"Summarize customer emails" is not measurable. "From a customer email, extract the order number, the complaint category from a fixed list of nine, and whether a refund is requested, or say that a field isn't present" is.
Write down:
- the input, with its real variety (languages, lengths, formats);
- the output schema, including when a field must be empty;
- what "correct" means per field, and who decides when it's ambiguous;
- the cost of each kind of error. A wrong refund flag and a wrong category rarely cost the same.
Step 2: build a test set before building anything else
Collect 100 to 300 real inputs. Have a person label them, and have a second person label a sample to see how often people disagree. That disagreement rate is the ceiling for any model: if two careful people agree 92% of the time on a field, chasing 99% is measuring noise.
Split off a locked test set that nothing gets tuned on. Add a challenge slice: the weird formats, the negations, the inputs where the answer is "not present".
Step 3: measure the frontier once, as a ceiling
Run the best hosted model you can use on the test set, with a good prompt and structured output. This gives you a ceiling and a reference for cost and latency. If the frontier model is only marginally better than your rules-based baseline, stop: the task doesn't need a model. If it's far better, you now know the gap you're trying to close.
Use a hosted model here only if the test data is allowed to leave your environment. If it isn't, the ceiling is your best local model, and the rest of the method still applies.
Step 4: try the plain open-model baseline
Run two or three open models that fit your hardware, each at a sensible quantization (how to choose one), with the same prompt and schema. Record quality per field, tokens per second and memory. Most tasks show a gap here; the next steps are about closing it with structure instead of a bigger model.
Step 5: close the gap with structure
These techniques are ordered by how often they help in my experience, cheapest first. Measure each one on its own.
Constrain the output. Use JSON-schema constrained decoding where the server supports it, and validate everything. Small models lose more quality to format errors than big ones, and constraints remove most of those errors. Record whether each answer was constrained, parsed, retried or repaired, so the evaluation can see it.
Decompose the task. One call that extracts six fields and classifies the email asks a small model to do several things at once. Two or three focused calls often do better, and each can use the best tool: a span extractor for the order number, a classifier for the category, a model call only for the part that needs reading comprehension.
Use a specialist where one exists. Entity extractors, structured-extraction models, rerankers and speech models trained for one job can beat much larger general models at that job (I wrote about four of them).
Retrieve the context the model lacks. Many "the small model doesn't know" failures are really "the small model wasn't told". Add the relevant policy, product list or definitions to the prompt, and filter them for eligibility before ranking.
Generate several and select. For tasks with more than one acceptable output (drafts, rewrites, suggestions), generate N candidates, reject those that fail hard checks, score the rest, and escalate near-ties to a pairwise judge asked in both orders. Use a judge from a different model family than the generator; models prefer their own style. This is the loop hone-select implements, and the reason it treats a failed judgement as missing rather than as zero.
Check with code. Arithmetic, dates, IDs against a database, quotes against the source text: every check you can do in code turns a model error into a routed exception instead of a wrong answer.
Step 6: add a route to the frontier, or to a person
Even after all that, some inputs will be hard. Measure how well the system's own signals (validation failures, low scores, judge disagreement, the specialist's confidence) predict errors on the test set. If they do, route those cases to the frontier model or a person, and let the open-model system handle the rest.
This is where "close to frontier" becomes an economic statement. If 85% of cases run locally at frontier-level quality and 15% go to a hosted model or a reviewer, the blended quality can match the frontier model's while cost and data exposure drop. Whether your task has such a split is exactly what this experiment finds out.
Step 7: compare fairly
A comparison is only worth something if both sides ran under comparable conditions:
- Same test set, same scoring, locked before you start tuning.
- Several samples per case for anything with sampling, and intervals instead of single numbers.
- Record the model version, quantization, prompt version and sampling settings for every run.
- On a shared machine, note when something else was using the GPU, and exclude or mark those samples.
- Decide the adoption rule before you see the results: "adopt the open system if field accuracy is within two points of the ceiling on every critical field and within budget on cost".
I use hone-select's experiments for this: an experiment file declares the setups and cases, a plan with an estimate is approved before anything runs, run conditions keep a busy machine out of the numbers, and the results compare every setup against a baseline with intervals. A spreadsheet and discipline also work.
What you end up with
At the end you have one of three answers, each useful:
- The open system matches the ceiling on your task. Run it locally, keep the test set, and rerun it whenever a model or prompt changes.
- It matches on most cases and the rest can be detected. Build the routed system.
- It doesn't get close, even with structure. Use the hosted model, and you'll know exactly why, which makes the privacy and cost conversation concrete.
Each of those answers is specific to your task. The broad benchmark gap between open and frontier models never gave you any of them.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real