The Right-Sized Model Portfolio
Stop Paying for Quality You Cannot Measure: Route Each Job to the Cheapest Model That Clears the Bar
Most teams pick one model and send it everything. Production data from our benchmark says that is the expensive option: on several kinds of work, the cheapest model on the shortlist scores within the margin of error of the priciest one, and the only thing that separates them is the bill. Here are two of those cases, plus the measurement caveat that makes them safe to act on.
Set the bar by task, not by vendor
The portfolio approach starts by admitting that “good enough” differs from one kind of work to the next. High-volume identification, translation, sequencing and portfolio work can stay on the cheapest models that already score above eight, while label accuracy, structured summarization and section analysis that everything downstream reuses justify paying more [^1][^2].
That split matters because quality moves with task shape more than with model name. A model that is a bargain on one job can be the wrong choice on the next, and the examples below show both directions.
What a ranked score actually means
Every ranked score rests on 239,838 scored judgments collected between 2026-06-03 and 2026-08-21 on real production work, not a fixed question set [^3][^4]. A model is called ranked on a task only once the 95% confidence interval on its mean quality is ±0.2 points on a 10-point scale [^3][^4].
The surprising part is how much that rule leaves out. Even with that volume, 8,401 of our 11,284 measured model-task cells — 74% — are still too uncertain to rank, and none of them is published [^3][^4]. So when two models land within a few tenths of each other on the pages that are published, the intervals are what tell you whether the gap is real.
Example one: voice profiles, where the cheap model ties the expensive ones
Building a reusable voice profile is high-quality across many models, so cost and serving mode decide. Tencent Hy4 Preview scores 9.74 out of 10 (±0.27) and costs $0.03 per task run in sync mode, while GLM-5.3 scores 9.65 out of 10 (±0.25) and costs $0.02 per task run in sync mode [^1][^2].
MiniMax M3 scores 9.58 out of 10 (±0.04) and costs $0.0014 per task run in sync mode, with confidence intervals that overlap the two higher scores [^1][^2]. On this evidence there is no clear winner on quality, and MiniMax M3 is cheaper than either. Routing this work to the top nominal score means paying far more for a difference the data cannot confirm.
Example two: decision synthesis, where four models are indistinguishable
Synthesizing multiple perspectives into a decision scores lower than voice profiles across the board, and the leaders sit close together. The table shows the four models at the top.
| Model | Score out of 10 | Cost per task run, sync mode |
|---|---|---|
| Moonshot Kimi K3 | 8.82 (±0.15) | $0.21 [^5] |
| GLM-5.3 | 8.73 (±0.16) | $0.10 [^5] |
| Grok 4.5 | 8.72 (±0.09) | $0.05 [^5][^6] |
| Thinking Machines Inkling | 8.72 (±0.27) | $0.06 [^5][^6] |
Moonshot Kimi K3 has the highest nominal score, but its confidence interval overlaps the next group [^5]. On this evidence those four are indistinguishable on quality, while Grok 4.5 and Thinking Machines Inkling are cheaper than Moonshot Kimi K3 and GLM-5.3 [^5][^6]. Price is the decider, and the top of the leaderboard is the wrong place to look.
The same logic cuts the other way
Cheap models earn their place only where the evidence says they clear the bar. Tencent Hy3 writes a short promotional post at 8.68 out of 10 (±0.11) and $0.0007 per task run in sync mode [^7][^8][^9] but chooses authoritative sources at 6.7 out of 10 (±0.34) and $0.0017 per task run in sync mode [^10][^11].
Same model, same price class, and a quality step-down large enough to matter. A per-task route keeps Hy3 on the promotional copy where it is cheap and good, and sends source selection to a model proven there. The discipline is not “which model is best” but “which is the cheapest model ranked above the bar this job needs.”
Where to look next
The bar you need is yours to set. Check the numbers against it at https://llm-bench.kapualabs.com/, where you can set the quality threshold yourself, and compare batch against sync pricing there before you lock routing, because serving mode decides cost where quality is tied. The task pages in the sources below, on llm-bench.kapualabs.com and fronset.ai, are where each number above lives.