Best LLMs for Bench Article Hypothesis Generation
Draw the distinct hypotheses a benchmark journal article supports, grounded in the approved claims it was written from.
Run this task on Fronset — request an invitation
Models
Frontier on this task: Moonshot Kimi K3 at 9.01 / 10. Quality bar at 90%: 8.11.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate does not clear the bar.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GPT-5.6 Luna | 8.62 / 10 | 8.31 | $0.84 | best value |
| DeepSeek V4 Flash | 9.01 / 10 | 8.78 | $3.85 | 4.6x more expensive |
| Thinking Machines Inkling Small | 8.62 / 10 | 8.24 | $5.72 | 6.8x more expensive |
| MiniMax M3 | 8.84 / 10 | 8.55 | $6.40 | 7.6x more expensive |
| Gemini 3.8 Flash | 8.58 / 10 | 8.26 | $7.54 | 9x more expensive |
| Qwen 3.7 Plus | 8.62 / 10 | 8.29 | $7.83 | 9.3x more expensive |
| GPT-5.6 Terra | 8.66 / 10 | 8.39 | $8.49 | 10x more expensive |
| GPT-5.6 Sol | 8.75 / 10 | 8.48 | $12.74 | 15x more expensive |
| Meta Muse Spark 1.3 | 8.86 / 10 | 8.61 | $16.00 | 19x more expensive |
| Gemini 3.5 Flash | 8.45 / 10 | 8.05 | $19.76 | 23x more expensive |
| Thinking Machines Inkling | 8.71 / 10 | 8.37 | $25.01 | 30x more expensive |
| DeepSeek V4 Pro | 8.58 / 10 | 8.20 | $26.14 | 31x more expensive |
| Grok 4.6 | 8.66 / 10 | 8.26 | $34.98 | 42x more expensive |
| Tencent Hy4 Preview | 8.84 / 10 | 8.60 | $58.01 | 69x more expensive |
| Moonshot Kimi K3 | 9.01 / 10 | 8.81 | $89.12 | 106x more expensive |
| Qwen 3.8 Max | 8.18 / 10 | 7.69 | $122.59 | 146x more expensive |
| Gemini 3.5 Flash Lite | 7.76 / 10 | 7.28 | $0.90 | 1.1x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GPT-5.6 Luna ★ OpenAI | 8.62 / 10 CI [8.31, 8.92] | MEDIUM | $0.84 | best value | batch |
| DeepSeek V4 Flash DeepSeek | 9.01 / 10 CI [8.78, 9.24] | HIGH | $3.85 | 4.6x | batch |
| Thinking Machines Inkling Small OpenRouter | 8.62 / 10 CI [8.24, 9.00] | MEDIUM | $5.72 | 6.8x | batch |
| MiniMax M3 OpenRouter | 8.84 / 10 CI [8.55, 9.12] | HIGH | $6.40 | 7.6x | batch |
| Gemini 3.8 Flash Gemini | 8.58 / 10 CI [8.26, 8.91] | MEDIUM | $7.54 | 9x | batch |
| Qwen 3.7 Plus Alibaba Cloud (DashScope) | 8.62 / 10 CI [8.29, 8.95] | MEDIUM | $7.83 | 9.3x | batch |
| GPT-5.6 Terra OpenAI | 8.66 / 10 CI [8.39, 8.93] | HIGH | $8.49 | 10x | batch |
| GPT-5.6 Sol OpenAI | 8.75 / 10 CI [8.48, 9.01] | HIGH | $12.74 | 15x | batch |
| Meta Muse Spark 1.3 OpenRouter | 8.86 / 10 CI [8.61, 9.10] | HIGH | $16.00 | 19x | batch |
| Gemini 3.5 Flash Gemini | 8.45 / 10 CI [8.05, 8.84] | MEDIUM | $19.76 | 23x | batch |
| Thinking Machines Inkling OpenRouter | 8.71 / 10 CI [8.37, 9.05] | MEDIUM | $25.01 | 30x | batch |
| DeepSeek V4 Pro DeepSeek | 8.58 / 10 CI [8.20, 8.97] | MEDIUM | $26.14 | 31x | batch |
| Grok 4.6 xAI | 8.66 / 10 CI [8.26, 9.05] | MEDIUM | $34.98 | 42x | batch |
| Tencent Hy4 Preview OpenRouter | 8.84 / 10 CI [8.60, 9.09] | HIGH | $58.01 | 69x | batch |
| Moonshot Kimi K3 best Moonshot AI | 9.01 / 10 CI [8.81, 9.22] | HIGH | $89.12 | 106x | batch |
| Qwen 3.8 Max Alibaba Cloud (DashScope) | 8.18 / 10 CI [7.69, 8.66] | MEDIUM | $122.59 | 146x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 2947 input tokens → 4147 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Output schema
Every answer on this task is checked against this JSON Schema, whichever model wrote it. An answer that doesn't fit counts as a model failure, and the call is retried on another model.
{
"$defs": {
"JournalHypothesisV1": {
"additionalProperties": false,
"description": "One thesis the article supports, and the claims behind it.",
"properties": {
"claim_refs": {
"description": "1-based indices into the numbered evidence block, naming the claims this hypothesis rests on.",
"items": {
"type": "integer"
},
"title": "Claim Refs",
"type": "array"
},
"headline": {
"description": "The hypothesis in one line — a claim, not a topic.",
"title": "Headline",
"type": "string"
},
"rationale": {
"description": "Why the cited claims support it, in two or three sentences.",
"title": "Rationale",
"type": "string"
}
},
"required": [
"headline",
"rationale"
],
"title": "JournalHypothesisV1",
"type": "object"
}
},
"additionalProperties": false,
"description": "Every hypothesis drawn from one article.",
"properties": {
"hypotheses": {
"description": "The requested number of distinct, non-overlapping hypotheses.",
"items": {
"$ref": "#/$defs/JournalHypothesisV1"
},
"title": "Hypotheses",
"type": "array"
}
},
"required": [
"hypotheses"
],
"title": "JournalHypothesesV1",
"type": "object"
}Prompt templates
This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.
LLMB_BENCH_ARTICLE_HYPOTHESIS_GENERATION_SYSTEM +
LLMB_BENCH_ARTICLE_HYPOTHESIS_GENERATION_USER
(1293 calls in window)
System prompt
Read one benchmark journal article together with the approved evidence it was written from, and state the small number of distinct hypotheses the article supports. A hypothesis is a claim someone could agree or disagree with, not a topic or a summary. Ground every hypothesis in the numbered evidence and cite the entries it rests on by their numbers; never assert a figure the evidence does not carry, and never restate a figure without the denominator it is given with. Prefer hypotheses that differ in kind rather than in wording, so that each one could be argued on its own. Say nothing about publication, channels or promotion. Your response must conform exactly to this output schema: {schema_json_string}.
User prompt
Produce exactly {hypothesis_count} hypotheses.
Article title: {title}
Article excerpt: {excerpt}
Article body:
{article_body}
Approved evidence (cite these by number):
{evidence_block}
JSON_REPAIR_SYSTEM +
JSON_REPAIR_USER
(12 calls in window)
System prompt
You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.
User prompt
JSON Schema:
{schema_json}
Malformed output to repair:
{raw_text}
Return only the corrected JSON object.