Best LLMs for Bench Hypothesis Post Generation
Write the short social messages that argue one hypothesis, using only the approved evidence behind it.
Run this task on Fronset — request an invitation
Models
Frontier on this task: Claude Opus 5 at 9.19 / 10. Quality bar at 90%: 8.27.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate does not clear the bar.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GPT-5.4 Nano | 8.88 / 10 | 8.69 | $0.36 | best value |
| Gemini 3.1 Flash Lite | 8.72 / 10 | 8.47 | $0.37 | 1x more expensive |
| GPT-5.6 Luna | 8.96 / 10 | 8.79 | $0.49 | 1.4x more expensive |
| Gemini 3.5 Flash Lite | 8.79 / 10 | 8.60 | $0.49 | 1.4x more expensive |
| Claude Haiku 4.5 | 8.70 / 10 | 8.50 | $1.81 | 5.1x more expensive |
| NVIDIA Nemotron-3 Super 120B | 8.31 / 10 | 7.93 | $3.36 | 9.4x more expensive |
| Tencent Hy3 | 8.56 / 10 | 8.15 | $3.40 | 9.5x more expensive |
| DeepSeek V4 Flash | 8.97 / 10 | 8.80 | $3.68 | 10x more expensive |
| Gemini 3.8 Flash | 9.00 / 10 | 8.84 | $3.71 | 10x more expensive |
| GPT-5.6 Terra | 8.90 / 10 | 8.72 | $3.72 | 10x more expensive |
| Claude Sonnet 5 | 8.48 / 10 | 8.05 | $3.87 | 11x more expensive |
| Thinking Machines Inkling Small | 8.66 / 10 | 8.34 | $4.98 | 14x more expensive |
| MiniMax M3 | 8.90 / 10 | 8.67 | $5.46 | 15x more expensive |
| Qwen 3.8 Flash | 8.50 / 10 | 8.15 | $7.10 | 20x more expensive |
| Qwen 3.7 Plus | 8.82 / 10 | 8.66 | $7.35 | 21x more expensive |
| GPT-5.6 Sol | 8.95 / 10 | 8.80 | $7.40 | 21x more expensive |
| Claude Opus 5 | 9.19 / 10 | 9.06 | $8.49 | 24x more expensive |
| Gemini 3.5 Flash | 8.95 / 10 | 8.79 | $9.42 | 26x more expensive |
| Meta Muse Spark 1.3 | 8.84 / 10 | 8.65 | $12.73 | 36x more expensive |
| DeepSeek V4 Pro | 8.77 / 10 | 8.53 | $14.68 | 41x more expensive |
| Thinking Machines Inkling | 8.75 / 10 | 8.48 | $16.58 | 46x more expensive |
| Moonshot Kimi K3 | 9.00 / 10 | 8.74 | $45.19 | 126x more expensive |
| Tencent Hy4 Preview | 8.76 / 10 | 8.53 | $56.78 | 158x more expensive |
| Grok 4.6 | 8.79 / 10 | 8.59 | $62.00 | 173x more expensive |
| Qwen 3.8 Max | 8.88 / 10 | 8.70 | $122.84 | 343x more expensive |
| NVIDIA Nemotron-3 Nano 30B-A3B | 7.73 / 10 | 7.26 | $0.98 | 2.7x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GPT-5.4 Nano ★ OpenAI | 8.88 / 10 CI [8.69, 9.06] | RANKED | $0.36 | best value | batch |
| Gemini 3.1 Flash Lite Gemini | 8.72 / 10 CI [8.47, 8.97] | HIGH | $0.37 | 1x | batch |
| GPT-5.6 Luna OpenAI | 8.96 / 10 CI [8.79, 9.13] | RANKED | $0.49 | 1.4x | batch |
| Gemini 3.5 Flash Lite Gemini | 8.79 / 10 CI [8.60, 8.98] | RANKED | $0.49 | 1.4x | batch |
| Claude Haiku 4.5 Anthropic | 8.70 / 10 CI [8.50, 8.90] | HIGH | $1.81 | 5.1x | batch |
| NVIDIA Nemotron-3 Super 120B OpenRouter | 8.31 / 10 CI [7.93, 8.69] | MEDIUM | $3.36 | 9.4x | batch |
| Tencent Hy3 OpenRouter | 8.56 / 10 CI [8.15, 8.97] | MEDIUM | $3.40 | 9.5x | batch |
| DeepSeek V4 Flash DeepSeek | 8.97 / 10 CI [8.80, 9.14] | RANKED | $3.68 | 10x | batch |
| Gemini 3.8 Flash Gemini | 9.00 / 10 CI [8.84, 9.16] | RANKED | $3.71 | 10x | batch |
| GPT-5.6 Terra OpenAI | 8.90 / 10 CI [8.72, 9.08] | RANKED | $3.72 | 10x | batch |
| Claude Sonnet 5 Anthropic | 8.48 / 10 CI [8.05, 8.92] | MEDIUM | $3.87 | 11x | batch |
| Thinking Machines Inkling Small OpenRouter | 8.66 / 10 CI [8.34, 8.99] | MEDIUM | $4.98 | 14x | batch |
| MiniMax M3 OpenRouter | 8.90 / 10 CI [8.67, 9.14] | HIGH | $5.46 | 15x | batch |
| Qwen 3.8 Flash Alibaba Cloud (DashScope) | 8.50 / 10 CI [8.15, 8.86] | MEDIUM | $7.10 | 20x | batch |
| Qwen 3.7 Plus Alibaba Cloud (DashScope) | 8.82 / 10 CI [8.66, 8.98] | RANKED | $7.35 | 21x | batch |
| GPT-5.6 Sol OpenAI | 8.95 / 10 CI [8.80, 9.11] | RANKED | $7.40 | 21x | batch |
| Claude Opus 5 best Anthropic | 9.19 / 10 CI [9.06, 9.33] | RANKED | $8.49 | 24x | batch |
| Gemini 3.5 Flash Gemini | 8.95 / 10 CI [8.79, 9.12] | RANKED | $9.42 | 26x | batch |
| Meta Muse Spark 1.3 OpenRouter | 8.84 / 10 CI [8.65, 9.02] | RANKED | $12.73 | 36x | batch |
| DeepSeek V4 Pro DeepSeek | 8.77 / 10 CI [8.53, 9.01] | HIGH | $14.68 | 41x | batch |
| Thinking Machines Inkling OpenRouter | 8.75 / 10 CI [8.48, 9.01] | HIGH | $16.58 | 46x | batch |
| Moonshot Kimi K3 Moonshot AI | 9.00 / 10 CI [8.74, 9.26] | HIGH | $45.19 | 126x | batch |
| Tencent Hy4 Preview OpenRouter | 8.76 / 10 CI [8.53, 8.99] | HIGH | $56.78 | 158x | batch |
| Grok 4.6 xAI | 8.79 / 10 CI [8.59, 9.00] | HIGH | $62.00 | 173x | batch |
| Qwen 3.8 Max Alibaba Cloud (DashScope) | 8.88 / 10 CI [8.70, 9.05] | RANKED | $122.84 | 343x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1319 input tokens → 4279 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Output schema
Every answer on this task is checked against this JSON Schema, whichever model wrote it. An answer that doesn't fit counts as a model failure, and the call is retried on another model.
{
"$defs": {
"FeedPostV1": {
"additionalProperties": false,
"description": "One social message arguing a hypothesis.",
"properties": {
"angle": {
"description": "How this message approaches the hypothesis: CLAIM, CONTRAST, IMPLICATION, QUESTION or CAVEAT.",
"title": "Angle",
"type": "string"
},
"text": {
"description": "The message. No link and no placeholder — the article link belongs to the announcement post, not to these.",
"title": "Text",
"type": "string"
}
},
"required": [
"angle",
"text"
],
"title": "FeedPostV1",
"type": "object"
}
},
"additionalProperties": false,
"description": "Every message drafted for one hypothesis.",
"properties": {
"posts": {
"description": "The requested number of messages, each standing on its own.",
"items": {
"$ref": "#/$defs/FeedPostV1"
},
"title": "Posts",
"type": "array"
}
},
"required": [
"posts"
],
"title": "FeedPostsV1",
"type": "object"
}Prompt templates
This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.
LLMB_BENCH_HYPOTHESIS_POST_GENERATION_SYSTEM +
LLMB_BENCH_HYPOTHESIS_POST_GENERATION_USER
(2070 calls in window)
System prompt
Write short social messages that each argue one given hypothesis about an AI model benchmark. Each message must stand on its own, be readable by someone who has not seen the article, and stay within the stated character limit. Vary the approach across the set: state the claim, draw a contrast, name an implication, ask the question it raises, or give the caveat that bounds it. Use only figures present in the supplied evidence, and carry each figure's denominator with it. Do not include a link, a link placeholder, a hashtag pile, or an invitation to read the article — the article is promoted separately. Avoid clickbait, invented urgency, unsupported superlatives and engagement bait. Your response must conform exactly to this output schema: {schema_json_string}.
User prompt
Produce exactly {post_count} messages, each at most {max_characters} characters.
Hypothesis: {hypothesis}
Why it holds: {rationale}
Source article: {article_title}
Approved evidence:
{evidence_block}
JSON_REPAIR_SYSTEM +
JSON_REPAIR_USER
(37 calls in window)
System prompt
You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.
User prompt
JSON Schema:
{schema_json}
Malformed output to repair:
{raw_text}
Return only the corrected JSON object.