Best LLMs for Evidence-Grounded Hypothesis Generation
Derive distinct, arguable hypotheses or claims from a source document while grounding and citing each one strictly in a supplied set of approved evidence.
Derive distinct, arguable hypotheses or claims from a source document while grounding and citing each one strictly in a supplied set of approved evidence.
Task-by-task breakdown
Bench Article Hypothesis Generation
Draw the distinct hypotheses a benchmark journal article supports, grounded in the approved claims it was written from.
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 96% | MEDIUM | best value |
| DeepSeek V4 Flash | 100% | HIGH | 4.6x |
| Thinking Machines Inkling Small | 96% | MEDIUM | 6.8x |
| MiniMax M3 | 98% | HIGH | 7.6x |
| Gemini 3.8 Flash | 95% | MEDIUM | 9x |
| Qwen 3.7 Plus | 96% | MEDIUM | 9.3x |
| GPT-5.6 Terra | 96% | HIGH | 10x |
| GPT-5.6 Sol | 97% | HIGH | 15x |
| Meta Muse Spark 1.3 | 98% | HIGH | 19x |
| Gemini 3.5 Flash | 94% | MEDIUM | 23x |
| Thinking Machines Inkling | 97% | MEDIUM | 30x |
| DeepSeek V4 Pro | 95% | MEDIUM | 31x |
| Grok 4.6 | 96% | MEDIUM | 42x |
| Tencent Hy4 Preview | 98% | HIGH | 69x |
| Moonshot Kimi K3 best | 100% | HIGH | 106x |
| Qwen 3.8 Max | 91% | MEDIUM | 146x |
Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.