Cost mode:

Derive distinct, arguable hypotheses or claims from a source document while grounding and citing each one strictly in a supplied set of approved evidence.

1 capability in this category.

Task-by-task breakdown

Bench Article Hypothesis Generation

Draw the distinct hypotheses a benchmark journal article supports, grounded in the approved claims it was written from.

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 96%MEDIUMbest value
DeepSeek V4 Flash100%HIGH4.6x
Thinking Machines Inkling Small96%MEDIUM6.8x
MiniMax M398%HIGH7.6x
Gemini 3.8 Flash95%MEDIUM9x
Qwen 3.7 Plus96%MEDIUM9.3x
GPT-5.6 Terra96%HIGH10x
GPT-5.6 Sol97%HIGH15x
Meta Muse Spark 1.398%HIGH19x
Gemini 3.5 Flash94%MEDIUM23x
Thinking Machines Inkling97%MEDIUM30x
DeepSeek V4 Pro95%MEDIUM31x
Grok 4.696%MEDIUM42x
Tencent Hy4 Preview98%HIGH69x
Moonshot Kimi K3 best100%HIGH106x
Qwen 3.8 Max91%MEDIUM146x

Task detail →

Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.