Scores You Can — and Cannot — Trust Yet
When a Benchmark Gap Is Real, When It Is Noise, and What “Not Ranked Yet” Means
Most benchmark tables hand you one number per model and leave you to guess how much of it is signal. Ours attach a confidence interval to every score, and the reading rule is short: a model earns a rank only when its interval is tight enough, and a gap between two models only counts when their intervals do not overlap. Here is how to apply it, with ties that look like gaps and rankings that are still moving.
The ±0.2 rule
A model is called ranked on a task only once the 95% confidence interval on its mean quality is ±0.2 points on a 10-point scale, whereas most benchmarks publish a single number with no interval at all [^1][^2]. Every ranked score rests on 239,838 scored judgments collected between 2026-06-03 and 2026-08-21 on real production work, not a fixed question set [^1][^2].
Even with that volume, 8,401 of our 11,284 measured model-task cells — 74% — are still too uncertain to rank, and none of them is published [^1][^2]. When a model is missing from a task page, that is usually what it means: the interval has not closed yet, not that the model failed.
When a gap is real
A real gap is one that survives the intervals. Claude Sonnet 5 writes newsletter copy at 9.32 out of 10, ±0.03 after 30 independent judgments [^3][^4], but discovers thematic topics at 1.02 out of 10, ±0.08 after 30 independent judgments [^5][^6]. Both cells rest on small counts, yet the intervals are narrow and nowhere near each other, so the difference is safe to build on.
The habit to form is to read the interval and the judgment count first and the point score last.
When a gap is noise
The more common case is a point score that suggests an order the data cannot support. For sorting and prioritizing engagement opportunities, three models land within a few hundredths of each other:
| Model | Score | 95% interval | Judgments |
|---|---|---|---|
| Claude Opus 4.8 | 8.71 | ±0.13 | 255 [^7][^8] |
| GLM-5.3 | 8.73 | ±0.21 | 30 [^8] |
| Grok 4.5 | 8.72 | ±0.16 | 167 [^7][^8] |
GLM-5.3 has the highest point score, the widest interval and the fewest judgments. With overlapping intervals around the same level, this is a tie, and choosing it because 8.73 beats 8.71 would be reading noise as a ranking.
Ties and real gaps can share a table. On triaging engagement opportunities, GPT-5.6 Luna scores 8.73 out of 10, ±0.16 after 194 independent judgments, while GPT-5.6 Terra scores 8.45 out of 10, ±0.17 after 164 [^7][^8]. Luna is higher than Terra, but Terra, GPT-5.4 Nano and GPT-5.6 Sol are tied with each other on the same task [^7][^8]. That table supports preferring Luna and refuses to support a choice among the other three.
The scoring process itself has measurable drift. Across 54,334 position-stamped judgments the first-shown candidate averages 7.87 against 7.73 for the last — a 0.14-point drift that is recorded rather than assumed away [^1][^2]. Keep that scale in mind whenever a lead is measured in hundredths.
When thin measurement can flip a ranking
Cells with few judgments move as more arrive, and not always upward. Since 2026-09-05, 829 model-task cells were compared like for like, with 4 added and 199 dropped [^2]. Several of those moves were demotions: GLM-5.3 Flash fell back from ranked to high confidence on translating text in batch [^9], MiniMax M3 fell back from ranked to high confidence on matching content to a topic taxonomy [^10] and from high to medium confidence on drafting public responses [^11], and Gemini 3.5 Flash fell back from high to medium confidence on assigning topics to sections [^12].
Confidence tiers track the interval, so a demotion tells you the measurement became less certain, not necessarily that the outputs got worse. The reverse holds as well. For identifying language, GLM-5.3 Flash gained 0.18 points since 2026-09-05, to 9.72 on a 10-point scale, while the 95% interval on its mean narrowed from ±0.5 to ±0.28 points [^13]. The interval closed far enough to move it from medium into high confidence, not because it scored better [^13]. At ±0.28 it is still outside the ±0.2 bar, so it remains a high-confidence cell rather than a ranked one.
Price can be thin in its own way: a strong quality score can sit beside a price that was never measured. On scoring how relevant generated content is, Qwen 3.7 Plus scores 9.12 out of 10 (±0.28) at $0.02 per task run in sync mode, but that cell is estimated with no qualifying usage history of its own [^14][^15][^16]. Read a cell like this as two claims of different strength: a quality figure with its own interval, and a cost figure that stays a projection until the price can be measured on the model’s own usage.
Where to look next
Check the live numbers before you commit, because these cells keep moving as intervals close and reopen. You can set the quality bar you actually need and compare batch against sync pricing yourself at https://llm-bench.kapualabs.com/, and the same tables are published per task at https://fronset.ai/benchmark/.