Cost mode:

Category: Social & Promotional Content · Typical I/O: 1319→4279 tokens

Run this task on Fronset — request an invitation

Models

Frontier on this task: Claude Opus 5 at 9.19 / 10. Quality bar at 90%: 8.27.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate does not clear the bar.

ModelQuality scoreCI lowCost / 1k runsvs best value
GPT-5.4 Nano8.88 / 108.69$0.36best value
Gemini 3.1 Flash Lite8.72 / 108.47$0.371x more expensive
GPT-5.6 Luna8.96 / 108.79$0.491.4x more expensive
Gemini 3.5 Flash Lite8.79 / 108.60$0.491.4x more expensive
Claude Haiku 4.58.70 / 108.50$1.815.1x more expensive
NVIDIA Nemotron-3 Super 120B8.31 / 107.93$3.369.4x more expensive
Tencent Hy38.56 / 108.15$3.409.5x more expensive
DeepSeek V4 Flash8.97 / 108.80$3.6810x more expensive
Gemini 3.8 Flash9.00 / 108.84$3.7110x more expensive
GPT-5.6 Terra8.90 / 108.72$3.7210x more expensive
Claude Sonnet 58.48 / 108.05$3.8711x more expensive
Thinking Machines Inkling Small8.66 / 108.34$4.9814x more expensive
MiniMax M38.90 / 108.67$5.4615x more expensive
Qwen 3.8 Flash8.50 / 108.15$7.1020x more expensive
Qwen 3.7 Plus8.82 / 108.66$7.3521x more expensive
GPT-5.6 Sol8.95 / 108.80$7.4021x more expensive
Claude Opus 59.19 / 109.06$8.4924x more expensive
Gemini 3.5 Flash8.95 / 108.79$9.4226x more expensive
Meta Muse Spark 1.38.84 / 108.65$12.7336x more expensive
DeepSeek V4 Pro8.77 / 108.53$14.6841x more expensive
Thinking Machines Inkling8.75 / 108.48$16.5846x more expensive
Moonshot Kimi K39.00 / 108.74$45.19126x more expensive
Tencent Hy4 Preview8.76 / 108.53$56.78158x more expensive
Grok 4.68.79 / 108.59$62.00173x more expensive
Qwen 3.8 Max8.88 / 108.70$122.84343x more expensive
NVIDIA Nemotron-3 Nano 30B-A3B7.73 / 107.26$0.982.7x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GPT-5.4 Nano OpenAI8.88 / 10 CI [8.69, 9.06]RANKED$0.36best valuebatch
Gemini 3.1 Flash Lite Gemini8.72 / 10 CI [8.47, 8.97]HIGH$0.371xbatch
GPT-5.6 Luna OpenAI8.96 / 10 CI [8.79, 9.13]RANKED$0.491.4xbatch
Gemini 3.5 Flash Lite Gemini8.79 / 10 CI [8.60, 8.98]RANKED$0.491.4xbatch
Claude Haiku 4.5 Anthropic8.70 / 10 CI [8.50, 8.90]HIGH$1.815.1xbatch
NVIDIA Nemotron-3 Super 120B OpenRouter8.31 / 10 CI [7.93, 8.69]MEDIUM$3.369.4xbatch
Tencent Hy3 OpenRouter8.56 / 10 CI [8.15, 8.97]MEDIUM$3.409.5xbatch
DeepSeek V4 Flash DeepSeek8.97 / 10 CI [8.80, 9.14]RANKED$3.6810xbatch
Gemini 3.8 Flash Gemini9.00 / 10 CI [8.84, 9.16]RANKED$3.7110xbatch
GPT-5.6 Terra OpenAI8.90 / 10 CI [8.72, 9.08]RANKED$3.7210xbatch
Claude Sonnet 5 Anthropic8.48 / 10 CI [8.05, 8.92]MEDIUM$3.8711xbatch
Thinking Machines Inkling Small OpenRouter8.66 / 10 CI [8.34, 8.99]MEDIUM$4.9814xbatch
MiniMax M3 OpenRouter8.90 / 10 CI [8.67, 9.14]HIGH$5.4615xbatch
Qwen 3.8 Flash Alibaba Cloud (DashScope)8.50 / 10 CI [8.15, 8.86]MEDIUM$7.1020xbatch
Qwen 3.7 Plus Alibaba Cloud (DashScope)8.82 / 10 CI [8.66, 8.98]RANKED$7.3521xbatch
GPT-5.6 Sol OpenAI8.95 / 10 CI [8.80, 9.11]RANKED$7.4021xbatch
Claude Opus 5 best Anthropic9.19 / 10 CI [9.06, 9.33]RANKED$8.4924xbatch
Gemini 3.5 Flash Gemini8.95 / 10 CI [8.79, 9.12]RANKED$9.4226xbatch
Meta Muse Spark 1.3 OpenRouter8.84 / 10 CI [8.65, 9.02]RANKED$12.7336xbatch
DeepSeek V4 Pro DeepSeek8.77 / 10 CI [8.53, 9.01]HIGH$14.6841xbatch
Thinking Machines Inkling OpenRouter8.75 / 10 CI [8.48, 9.01]HIGH$16.5846xbatch
Moonshot Kimi K3 Moonshot AI9.00 / 10 CI [8.74, 9.26]HIGH$45.19126xbatch
Tencent Hy4 Preview OpenRouter8.76 / 10 CI [8.53, 8.99]HIGH$56.78158xbatch
Grok 4.6 xAI8.79 / 10 CI [8.59, 9.00]HIGH$62.00173xbatch
Qwen 3.8 Max Alibaba Cloud (DashScope)8.88 / 10 CI [8.70, 9.05]RANKED$122.84343xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1319 input tokens → 4279 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Output schema

Every answer on this task is checked against this JSON Schema, whichever model wrote it. An answer that doesn't fit counts as a model failure, and the call is retried on another model.

{
  "$defs": {
    "FeedPostV1": {
      "additionalProperties": false,
      "description": "One social message arguing a hypothesis.",
      "properties": {
        "angle": {
          "description": "How this message approaches the hypothesis: CLAIM, CONTRAST, IMPLICATION, QUESTION or CAVEAT.",
          "title": "Angle",
          "type": "string"
        },
        "text": {
          "description": "The message. No link and no placeholder — the article link belongs to the announcement post, not to these.",
          "title": "Text",
          "type": "string"
        }
      },
      "required": [
        "angle",
        "text"
      ],
      "title": "FeedPostV1",
      "type": "object"
    }
  },
  "additionalProperties": false,
  "description": "Every message drafted for one hypothesis.",
  "properties": {
    "posts": {
      "description": "The requested number of messages, each standing on its own.",
      "items": {
        "$ref": "#/$defs/FeedPostV1"
      },
      "title": "Posts",
      "type": "array"
    }
  },
  "required": [
    "posts"
  ],
  "title": "FeedPostsV1",
  "type": "object"
}

Prompt templates

This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.

LLMB_BENCH_HYPOTHESIS_POST_GENERATION_SYSTEM + LLMB_BENCH_HYPOTHESIS_POST_GENERATION_USER (2070 calls in window)

System prompt

Write short social messages that each argue one given hypothesis about an AI model benchmark. Each message must stand on its own, be readable by someone who has not seen the article, and stay within the stated character limit. Vary the approach across the set: state the claim, draw a contrast, name an implication, ask the question it raises, or give the caveat that bounds it. Use only figures present in the supplied evidence, and carry each figure's denominator with it. Do not include a link, a link placeholder, a hashtag pile, or an invitation to read the article — the article is promoted separately. Avoid clickbait, invented urgency, unsupported superlatives and engagement bait. Your response must conform exactly to this output schema: {schema_json_string}.

User prompt

Produce exactly {post_count} messages, each at most {max_characters} characters.

Hypothesis: {hypothesis}

Why it holds: {rationale}

Source article: {article_title}

Approved evidence:
{evidence_block}
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (37 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.