A Benchmark Score Is Four Numbers, Not One

A quality score on our benchmark only means something with three things attached: its confidence interval, the serving mode the cost was quoted in, and a flag saying whether that cost was measured or estimated. Read that way, the current data supports exactly one clean claim that paying more buys better output, and it declines to support several claims people would like a benchmark to make.

What a cell actually contains

The figures behind this rest on 57,538 scored judgments across 833 measured model-and-task combinations, covering 30 models on 52 tasks. They were published on a single date, but the measurement behind them is continuous.

Each score carries an interval, and it does real work. Where two models’ intervals overlap, the result is a tie even when the nominal scores differ. The width reflects the evidence behind the number: the most confidently measured translation result in its set is GPT-5.6 Terra in batch mode at 8.97 out of 10, ±0.14 after 50 independent judgments [^1][^2].

Each cost carries two more labels. One is the serving mode, sync or batch, and a batch price is not comparable to a sync price. The other is firmness: a cost is either measured on the model’s own usage, repriced at current rates from an older observation, or estimated with no qualifying usage history.

The one place paying more buys a measured lead

Creating title packages for published content is the one task where the data supports a quality lead for the more expensive option. All four cells below are measured and quoted in the same mode, which is what makes the comparison legitimate.

ModelQualitySync cost per runCost basis
Moonshot Kimi K39.12 out of 10 (±0.13)$0.07measured on this model’s own usage [^3][^4]
Thinking Machines Inkling8.75 out of 10 (±0.15)$0.03measured on this model’s own usage [^3][^4]
Grok 4.58.68 out of 10 (±0.1)$0.02measured on this model’s own usage, repriced at current rates [^3][^4]
Meta Muse Spark 1.18.67 out of 10 (±0.15)$0.02measured on this model’s own usage [^3][^4]

Among the three lower-priced options the confidence intervals overlap, so they are a tie on quality, with Inkling slightly higher on cost. Kimi K3 sits above that pack by more than the intervals allow, and that gap, not the raw score, is why the extra money can be said to buy something.

Notice how much had to line up: same task, same mode, non-overlapping intervals, firm costs on every row. Most cells do not clear all four at once, and then the right reading is different.

When a tie is the honest answer

Identifying language shows the opposite pattern: quality is very high, cost is very low, and the top results are a tie once uncertainty is taken into account. Claude Opus 4.8 scores 9.96 out of 10 (±0.15) at $0.0037 per task run in batch mode, and MiniMax M3 scores 9.95 out of 10 (±0.08) at $0.0005 per task run in sync mode, both estimated with no qualifying usage history [^5][^6]. GLM-5.3 Flash follows at 9.72 out of 10 (±0.28) at $0.0002 per task run in sync mode, and that one is measured on this model’s own usage [^6].

The interval says there is no quality gap to act on, so the choice turns on cost mode and deployment fit instead. Even that comparison needs care, because the cheapest figure is a sync price, the Opus figure is a batch price, and the firmness flags differ.

The firmness flag can also change how you read a low score. Gemini 3.5 Flash Lite summarizing structured content scores 5.73 out of 10 (±0.31) at $0.0066 per task run in batch mode, but the cell is estimated with no qualifying usage history of its own, so the lower score should be treated as provisional [^7][^8].

What the coverage cannot support

This is a deliberate disclosure about the data, not a ranking. The material does not supply a possible-total grid alongside the measured combinations, so measured-against-possible coverage cannot be calculated. It does not provide a model-by-task matrix showing which models are measured broadly and which narrowly; uneven call volumes are differences in observed attempts, not a ranking of breadth.

It also lacks a per-task coverage table that would identify the thinnest tasks. And although every cost cell is labeled measured, repriced or estimated, those labels are not aggregated into shares, so no freshness accounting can be stated.

So the coverage supports routing hypotheses conditioned on task, mode and usability gates, not a universal ranking, a claim about unmeasured model breadth, a thinnest-task list, or a freshness share for cost data. Answer availability, length effects, latency tails and parse failures are further gates outside the score, each checked task by task.

Where to look next

The live figures, with intervals, serving modes and firmness labels on every cell, are at https://llm-bench.kapualabs.com/, and the per-task pages on the fronset.ai benchmark site let you compare batch against synchronous pricing before selecting a route. Set the quality bar your workflow requires first, then look for cells that clear it with the interval, the mode and the flag all in your favour.