What the published measurement covers

The published benchmark is broad enough to support decisions about named model-task-mode combinations. It is not a complete model leaderboard.

The current published set contains 55,732 scored judgments across 812 model-task combinations, covering 30 models and 53 tasks. A simple 30-by-53 grid would contain 1,590 pairings, but the evidence does not establish that every pairing is feasible, eligible, or measured. The 812 published combinations should therefore not be read as a coverage percentage.

Task results also vary sharply within the same model. GPT-5.6 Terra scored 9.99 for language identification but 7.99 for research-community selection. Claude Sonnet 5 scored 9.32 for newsletter copy and 1.02 for thematic discovery.[^4][^6][^13][^15][^18][^28][^32][^34] These are not contradictions. They are evidence that capability belongs to the transformation, not the model name.

Missing cells do not mean weak models

The material reports that 8,401 of 11,284 measured cells—74%—were too uncertain to rank and were not published.[^1][^23] That measurement universe cannot be reconciled with a simple 30-by-53 grid from the supplied information, so it should not be used as the denominator for published coverage.

A result becomes rankable only when its 95% confidence interval is within ±0.2 points on the 10-point scale.[^1][^23] Absence from the ranked set may therefore mean insufficient precision, not poor performance.

The evidence also cannot identify which models are measured broadly, which tasks have the thinnest coverage, or how the 53-task published scope relates to the 68-task batch-eligibility scope. Eighteen of those 68 tasks are not batch-eligible.[^1][^23]

What task-level comparisons can tell us

Some tasks are operationally interchangeable at the top. For short promotional messages, 96.8% of leading pairs were within half a point, with a median gap of 0.1.[^21][^22] Section-to-cluster assignment had a 0.0 median gap.[^9][^12][^27]

Other tasks separate models more strongly. Content-domain suggestion had a median top-two gap of 1.9, and paragraph-level metadata improvement 0.8.[^16][^29][^35] In these cases, task-specific quality deserves more weight.

When quality is tied, operating constraints can decide. Language identification is a reported quality tie across several models, while quoted synchronous costs range materially.[^4][^18] That supports a cost choice for language identification in the stated mode—not a broader model ranking.

Delivery, parseability, and latency are separate evidence

Structured-field scores apply to returned answers. DeepSeek V4 Pro scored 9.54 and answered on 80% of attempts; GPT-5.4 Nano scored 9.46 and answered on 76%; MiniMax M3 scored 9.36 and answered on 58%.[^2][^25] DeepSeek V4 Flash answered on only 12% of atomic-claim extraction attempts.[^26]

Parseability can dominate cost. Choosing vetted sites spent $32.39, with $19.04—58.8%—going to generated, billed tokens that could not be parsed.[^5][^8][^19][^24]

Serving mode and latency are equally local. GPT-5.6 Terra and Luna were tied on image-prompt quality, but Terra’s result was synchronous at $0.01 and Luna’s batch result cost $0.0006.[^3][^20] GPT-5.4 Mini ranged from sub-second author matching to 18-second referenced analyst prose.[^7][^10][^11][^17][^30][^33][^36]

What the coverage cannot support

The published evidence cannot support a complete ranking of all models on all tasks, a list of the least-covered tasks, a model-wide latency claim, or an overall breakdown of price freshness.

Use each result only at the level measured. Where quality is tied, choose on the binding measured constraint. Where completion, latency, parseability, or like-for-like cost is missing, treat it as unresolved—not as evidence for a default winner.

Set the required quality bar and inspect the exact comparisons at https://fronset.ai/benchmark/.