A high benchmark score can hide a weak production route. The score describes the answers a model returned; the answer rate tells you how often the workflow received one.

Structured extraction needs both quality and coverage

Gemini 3.5 Flash has the highest reported structured-fact score at 9.78 out of 10, but it answers on 77% of attempts. Qwen 3.7 Plus scores 9.67 and answers on only 40%.[^1][^12] DeepSeek V4 Pro is slightly lower at 9.54, but answers on 80% of attempts—the strongest reported balance among those alternatives.[^1][^12]

That changes the engineering question. Instead of asking only, “Which model produces the best answer?”, ask, “How many items reach a usable answer without a retry or fallback?” A 9.67-quality route with 40% coverage is difficult to run unattended.

Strong performance also does not transfer automatically to adjacent tasks. DeepSeek V4 Flash scores 7.43 for atomic factual-claim extraction but answers on 12% of attempts; for claim refinement, it scores 7.71 and answers on 36%.[^14][^16] DeepSeek V4 Pro was selected best in none of 91 matched claim-refinement comparisons.[^4][^13] Structured-field extraction and claim refinement need separate routes.

Topic analysis is not one operation

Claude Opus 5 has the strongest reported cluster-labeling result in its direct comparison: 8.91 out of 10 and the best answer in 10 of 25 matched occasions. It still returned an answer on only 76% of attempts.[^11]

Other stages point elsewhere. Gemini 3.8 Flash scores 7.81 for authoritative-source selection with an 87% answer rate, while DeepSeek V4 Flash scores 8.12 for cluster labeling with an 86% answer rate.[^9][^11] These results measure different decisions. A model that labels a cluster well has not thereby shown that it can judge source authority or content relevance.

Markdown cleanup makes the trade-off visible

Gemini 3.5 Flash scores 9.33 for Markdown newline repair when it answers, but answers on 65% of attempts.[^2][^10] MiniMax M3 scores 8.76, answers on 52%, and has a 40-second median response time on this task.[^2][^6][^7][^10]

For a pipeline that must repair every document, validation, retry, fallback, and review are part of the model choice—not implementation details added afterward.

Routing decision

Use DeepSeek V4 Pro for structured extraction when its reported quality-and-coverage balance fits the service requirement. Use Claude Opus 5 for cluster labeling when returned-answer quality is the priority, but provide a fallback. Keep source selection, relevance judgment, claim refinement, and document analysis on separate task-specific routes.

Explore the live task results at https://fronset.ai/benchmark/ and the methodology at https://fronset.ai/benchmark/methodology/.