Best Models for Structured Data & Fact Extraction
Structured Data & Fact Extraction: quality is high, but delivery conditions decide the route
Structured extraction looks nearly solved if you read only the top scores. It looks much less settled once answer rate, repair burden, and failure type enter the decision.
For predefined fields, Gemini 3.5 Flash leads the reported quality results at 9.78 out of 10, but answers on 77% of attempts. GPT-5.6 Sol follows at 9.74, while GPT-5.6 Luna and Qwen 3.7 Plus reach 9.67; Qwen answers on only 40% of attempts.[^1][^11]
DeepSeek V4 Flash and DeepSeek V4 Pro are tied at 9.54 (±0.14). In a separately reported completion comparison, Pro answers on 80% of attempts, versus 76% for GPT-5.4 Nano and 58% for MiniMax M3.[^1][^11]
That creates two different leaders:
- Maximum returned-answer quality: Gemini 3.5 Flash, with recovery handling for the missing 23%.
- Stronger reported quality-and-coverage balance: DeepSeek V4 Pro, at 9.54 quality and 80% answer rate.[^1][^11]
The supplied evidence contains no comparable pricing, so it cannot identify the cheapest qualifying route.
Clean output is a separate requirement
MiniMax M3 shows why “the model returned something” is not the same as “the task completed.” Its structured output required repair on 25.1% of attempts and was usable as returned on 74.6%.[^2][^5][^13][^14] On Markdown newline repair, it scored 8.76 when it answered, answered on 52% of attempts, and took 40 seconds at the median.[^5][^6][^7][^12]
For structured pipelines, validation and repair burden belong in the cost model. A cheaper call can be more expensive end to end if a large share of outputs requires another pass.
Fixed schemas do not predict fact extraction
Atomic factual decomposition is a different task. DeepSeek V4 Flash falls from 9.54 on structured fields to about 7.44 on atomic-claim extraction, with a separately reported answer rate of only 12%.[^1][^4][^9][^11] Gemini 3.1 Flash Lite scores 6.3 on the same kind of work.[^1][^4][^8][^9][^11]
Known-URL extraction reverses the comparison again. Gemini 3 Flash Preview had a 10.6% judged failure rate, versus 54.3% for Gemini 3.1 Flash Lite.[^3][^10] The first model tended to refuse—a visible retry condition—while the second more often produced other errors that could pass unnoticed. In high-stakes extraction, a detectable refusal may be safer than a plausible wrong answer.
Routing rule
Use Gemini 3.5 Flash when maximum returned-answer quality matters and recovery is available. Prefer DeepSeek V4 Pro when its slightly lower score and higher reported answer rate better fit the service requirement. Treat atomic-claim extraction and known-URL extraction as separate routes with their own failure tests.
Set the quality and completion bar at https://fronset.ai/benchmark/ before comparing batch and synchronous pricing.