The number is not the decision

A benchmark score is trustworthy only for the exact task, operating mode, and population of calls it measures.

Fronset ranks a model-task result only when the 95% confidence interval on mean quality is within ±0.2 points on the 10-point scale. Of 11,284 measured cells, 8,401—74%—were still too uncertain to rank.[^1][^34] An unranked result is not a failure. It is unresolved evidence.

Absolute rubric scores and same-input head-to-head selections answer different questions: one measures performance against a standard; the other measures preference between outputs on shared inputs.[^1][^34] They should inform each other, not be treated as interchangeable leaderboards.

Tight intervals support narrow decisions

GPT-5.6 Sol scored 10.0 (±0.02) for language identification and 9.74 (±0.13) for predefined structured extraction.[^14][^36][^2][^33] Claude Sonnet 5 scored 9.44 (±0.15) for metadata rewriting and 8.93 (±0.13) for batch translation.[^25][^46][^22][^41]

These are strong routing signals for those named operations. They say little about adjacent work.

Claude Sonnet 5 makes the boundary obvious: 8.38 for labeling existing topic clusters, but 1.02 for discovering themes in a corpus.[^8][^16][^27][^52] It also scores 9.32 for newsletter copy.[^11][^53] Discovering a direction, labeling material after the direction exists, and writing within it are three separate acceptance problems.

GPT-5.4 Nano shows the same split: 9.46 for structured extraction, but 2.75 for short-post relevance and 4.53 for atomic factual-claim extraction.[^2][^4][^7][^31][^32][^33]

A precise score can still describe incomplete delivery

DeepSeek V4 Pro scored 9.54 for structured fields and answered on 80% of attempts. GPT-5.4 Nano scored 9.46 and answered on 76%; MiniMax M3 scored 9.36 and answered on 58%; Qwen 3.7 Plus scored 9.67 and answered on 40%.[^2][^33]

The quality score describes returned answers. The answer rate describes how much of the workload receives one. Unattended processing needs both, plus a plan for malformed output.

Failure composition adds another layer. MiniMax M3 had fewer judged failures than Gemini 3.5 Flash on structured summarization, where Gemini’s failures were mostly invention and MiniMax’s mostly brevity.[^37] The ordering reversed on content-relevance scoring.[^30] “Safer model” is therefore meaningful only after the task and unacceptable failure have been defined.

Small score differences are often ties

DeepSeek V4 Flash scored 5.91 (±0.41) for short-post relevance and GLM-5.3 Flash 6.53 (±0.40), at the same stated cost. The intervals overlap, so the evidence supports a tie—not a 0.62-point quality lead.[^7][^31]

Engagement-opportunity triage is similarly unresolved among GPT-5.6 Luna, Terra, GPT-5.4 Nano, and Sol.[^5][^49] DeepSeek V4 Pro and Flash are exactly tied at 9.54 (±0.14) for structured extraction.[^2][^33]

When quality is tied, decide on completion, latency, serving mode, failure handling, or cost. Do not manufacture a winner from the point estimates.

Cost is narrower than it appears

Only 50 of 68 tasks are batch-eligible.[^1][^34] A price belongs to a specific task and mode: GPT-5.6 Sol’s reference-preserving analysis costs $0.18 in batch versus $0.64 synchronously, while its theme-discovery prices are $0.15 and $0.51.[^8][^12][^13][^51][^52][^58] Those figures do not establish equal quality across modes or a model-wide price advantage.

Usable-output cost may also exceed the price card when output cannot be parsed or requires another billed call.[^1][^9][^34][^38]

Routing conclusion

Trust a score when its interval is sufficiently narrow, the task matches the production operation, and completion and output validity are acceptable. Treat overlapping intervals as ties, unranked cells as unresolved, and prices as task-and-mode-specific.

Inspect the live evidence at https://fronset.ai/benchmark/ before turning a number into a production route.