Model Under the Microscope
The Top Two Models Usually Tie, So Route on Cost
If you pick a model for analysis work by its judged score, the benchmark has an awkward message for you: the winner and the runner-up are usually a rounding error apart. On catalyst and scenario analysis, a judged win tells you far less than the leaderboard position suggests. The practical move is to let cost and operational fit decide within that cluster, and to stop treating output length as a dial you can set once, because judges flip their length preference between neighboring tasks.
How close the top really is
On catalyst and scenario analysis, the top two answers scored exactly the same in 20.0% of 30 judged occasions, and finished within half a point in 73.3%, with a median gap of 0.2 on a 10-point scale [^1][^2]. That is not a leaderboard with a clear front-runner. It is a pack, and a judged win inside it separates very little.
Claude Opus 5 was judged the best answer in 10 of the 30 occasions it competed in, 33.3%, every one of them against other models answering the same live input [^2]. A one-in-three win rate is a real result, but it is the only direct quality win the benchmark material attributes to the model, and it sits inside those near-ties rather than above them.
What that means for routing
When the gap between first and second place is a fraction of a point, the quality column stops being the deciding one. Route catalyst and scenario work on cost or operational fit unless you have a quality bar high enough to justify chasing a small gap, and send it to Claude Opus 5 only when your workflow can tolerate near-ties and one of those two factors favors it.
Be careful not to stretch that result. The benchmark gives no cost per run or latency for Claude Opus 5, and no head-to-head for it outside catalyst and scenario work, so nothing here tells you how it does on summarization, translation or extraction. A win on one task is evidence for that task and nothing else.
Length is not a lever you can set once
The tempting shortcut is to correct for verbosity: ask every model for shorter answers, save output tokens, and assume the judges will not mind. The benchmark says they do mind, but not in a consistent direction. On catalyst and scenario work, longer answers were favored at 4,745 against 4,177 output tokens [^1][^2], so trimming that task would cost you quality while saving tokens.
Step sideways to a neighboring task and the preference reverses. Generating research queries favored shorter answers, while validating research queries favored longer ones, which means two halves of the same pipeline want opposite length settings. The table shows how large and how two-directional the effect is, with the favored answer’s length listed first.
| Task | Judges favored | Output tokens, favored vs. other |
|---|---|---|
| Generating research queries | shorter | 1,617 vs. 1,908 [^3] |
| Validating research queries | longer | 1,565 vs. 1,375 [^4] |
| Language detection | shorter | 58 vs. 268 [^5][^6] |
| Suggesting content domains | longer | 2,257 vs. 574 [^7][^8][^9] |
| Translation | shorter | 1,150 vs. 1,699 [^10][^11] |
| Summarization | longer | 3,333 vs. 2,203 [^12][^13][^14] |
The lesson is that a length assumption does not travel. If you tune a system prompt for brevity on translation and reuse it for summarization, you have made the summaries worse in the judges’ eyes while believing you optimized. Check the length preference for each task on its own before treating any judged win as a general lead for concise or verbose models.
Where to look next
Every number above comes from the live benchmark, which scores models with judges on the same live inputs. Before you commit a route, open the task pages at https://llm-bench.kapualabs.com/ and the matching pages under https://fronset.ai/benchmark/ to see the current scores for the quality bar you actually need, and compare batch against sync pricing there. If the top of a task is as tightly packed as catalyst and scenario analysis, let cost decide.