Cost mode:

Category: Relevance, Classification & Matching · Typical I/O: 13947→1679 tokens

Run this task on Fronset — request an invitation

Models

Frontier on this task: Tencent Hy4 Preview at 8.83 / 10. Quality bar at 90%: 7.95.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate does not clear the bar.

ModelQuality scoreCI lowCost / 1k runsvs best value
GLM-5.3 Flash8.72 / 108.35$9.44best value
Gemini 3.5 Flash8.31 / 107.92$17.691.9x more expensive
DeepSeek V4 Pro8.08 / 107.82$22.772.4x more expensive
Meta Muse Spark 1.38.53 / 108.09$25.932.7x more expensive
Thinking Machines Inkling8.20 / 107.72$32.223.4x more expensive
Claude Opus 58.50 / 108.03$36.403.9x more expensive
Tencent Hy4 Preview8.83 / 108.49$38.124x more expensive
Grok 4.68.41 / 107.97$52.095.5x more expensive
GLM-5.38.47 / 107.99$104.5611x more expensive
Gemini 3.1 Flash Lite7.44 / 107.14$2.1577% cheaper
Gemini 3.5 Flash Lite6.93 / 106.53$2.2177% cheaper
Claude Haiku 4.57.51 / 107.14$7.9815% cheaper
DeepSeek V4 Flash7.59 / 107.28$8.2712% cheaper
GPT-5.4 Nano6.80 / 106.42$2.0379% cheaper
MiniMax M37.74 / 107.51$6.1934% cheaper
Claude Sonnet 57.80 / 107.41$20.752.2x more expensive
NVIDIA Nemotron-3 Ultra 550B7.13 / 106.63$13.281.4x more expensive
Qwen 3.7 Plus7.90 / 107.47$10.471.1x more expensive
Tencent Hy37.43 / 107.09$2.9169% cheaper
Thinking Machines Inkling Small7.91 / 107.52$9.951.1x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GLM-5.3 Flash ★ Z.AI8.72 / 10 CI [8.35, 9.09]MEDIUM$9.44best valuebatch
Gemini 3.5 Flash Gemini8.31 / 10 CI [7.92, 8.70]MEDIUM$17.691.9xbatch
DeepSeek V4 Pro DeepSeek8.08 / 10 CI [7.82, 8.35]HIGH$22.772.4xbatch
Meta Muse Spark 1.3 OpenRouter8.53 / 10 CI [8.09, 8.97]MEDIUM$25.932.7xbatch
Thinking Machines Inkling OpenRouter8.20 / 10 CI [7.72, 8.69]MEDIUM$32.223.4xbatch
Claude Opus 5 Anthropic8.50 / 10 CI [8.03, 8.97]MEDIUM$36.403.9xbatch
Tencent Hy4 Preview best OpenRouter8.83 / 10 CI [8.49, 9.17]MEDIUM$38.124xbatch
Grok 4.6 xAI8.41 / 10 CI [7.97, 8.85]MEDIUM$52.095.5xbatch
GLM-5.3 Z.AI8.47 / 10 CI [7.99, 8.96]MEDIUM$104.5611xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 13947 input tokens → 1679 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge candidate fit under the supplied dimensions, eligibility compliance, evidence for each match, ranking and threshold calibration, appropriate reuse of the existing pool, correct unmatched behavior, and restraint and usefulness when proposing a coverage-gap profile. Penalize superficial keyword matching, popularity or identity bias, invented attributes, and invented real candidates. Output cardinality, identifier validity, and schema conformance are deterministic.

Output schema

Every answer on this task is checked against this JSON Schema, whichever model wrote it. An answer that doesn't fit counts as a model failure, and the call is retried on another model.

{
  "$defs": {
    "NewAuthorData": {
      "description": "Data for creating a new fictional author.\nUsed when the LLM decides to create a new author rather than assign to existing.",
      "properties": {
        "biography": {
          "description": "Short bio under 200 chars (e.g., 'Clinical researcher specializing in women's health with 15 years of industry experience.')",
          "maxLength": 200,
          "title": "Biography",
          "type": "string"
        },
        "content_domains": {
          "anyOf": [
            {
              "items": {
                "type": "string"
              },
              "type": "array"
            },
            {
              "type": "null"
            }
          ],
          "description": "Names of content domains this author belongs to. Must be a subset of the template's content domains. Empty = generalist.",
          "title": "Content Domains"
        },
        "content_themes": {
          "description": "Keywords for future matching (5-10 keywords)",
          "items": {
            "type": "string"
          },
          "title": "Content Themes",
          "type": "array"
        },
        "expertise_areas": {
          "description": "List of expertise areas (3-5 areas)",
          "items": {
            "type": "string"
          },
          "title": "Expertise Areas",
          "type": "array"
        },
        "gender": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Portrait gender for image generation: 'male', 'female', or 'neutral'. Use the historical figure's gender (e.g. 'male' for Adam Smith, 'female' for Marie Curie). Omit if uncertain.",
          "title": "Gender"
        },
        "inspiration_note": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Internal note: who/what inspired this persona",
          "title": "Inspiration Note"
        },
        "interests": {
          "description": "List of interests that humanize the author",
          "items": {
            "type": "string"
          },
          "title": "Interests",
          "type": "array"
        },
        "name": {
          "description": "Author's display name (e.g., 'Dr. Elara Stanton')",
          "title": "Name",
          "type": "string"
        },
        "writing_style": {
          "description": "Writing style: 'academic', 'conversational', 'investigative', 'analytical', or 'accessible'",
          "title": "Writing Style",
          "type": "string"
        }
      },
      "required": [
        "name",
        "biography",
        "expertise_areas",
        "content_themes",
        "writing_style"
      ],
      "title": "NewAuthorData",
      "type": "object"
    }
  },
  "description": "Output model for LLM-based author matching.\nThe LLM either assigns content to an existing author or creates a new one.",
  "properties": {
    "action": {
      "description": "Either 'assign' to use existing author or 'create' to make a new one",
      "title": "Action",
      "type": "string"
    },
    "author_id": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "UUID of the existing author if action is 'assign'",
      "title": "Author Id"
    },
    "new_author_data": {
      "anyOf": [
        {
          "$ref": "#/$defs/NewAuthorData"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "New author details if action is 'create'"
    }
  },
  "required": [
    "action"
  ],
  "title": "AuthorMatchingResult",
  "type": "object"
}

Prompt templates

This task has 4 published templates; the default is shown first.

LLMB_PROFILE_POOL_MATCHING_SYSTEM + LLMB_PROFILE_POOL_MATCHING_USER (617 calls in window)

System prompt

Compare the requirements and relevant attributes of source_object with the candidates in profile_pool. Apply matching_profile exactly for candidate type, eligibility, comparison dimensions, weights, thresholds, ranking, result count, and gap policy. Preserve candidate identifiers and support every match with evidence from both the source object and candidate profile. Reuse qualifying candidates before proposing a new profile. If no candidate qualifies and profile proposals are disabled, return the configured unmatched or coverage-gap result. If proposals are enabled, describe the missing profile requirements without inventing a real person, organization, asset, or other existing entity, and without implying that any real entity holds the described views or produced the described work. Where matching_profile requires the proposed profile to be named after a real figure, follow that naming rule exactly — a named persona is a described role, not a claim about that figure. Do not match on name recognition, protected characteristics, popularity, or unsupported attributes unless the supplied policy explicitly makes an attribute relevant and lawful. Treat empty optional values as absent and return only the requested result. Your response must conform exactly to this output schema: {schema_json_string}.

User prompt

Inputs — source_object: {source_object}; profile_pool: {profile_pool}; matching_profile: {matching_profile}. Use only these inputs to complete the task defined by the system prompt.
AUTHOR_MATCHING_SYSTEM_ASSIGN_ONLY + AUTHOR_MATCHING_USER_ASSIGN_ONLY (318 calls in window)

System prompt

You are an editor assigning articles to AI author personas at a publication.

The author pool has reached its maximum size. You MUST select from the existing authors below. Creating new authors is NOT an option.

Given a content piece and a pool of available AI authors, select the BEST matching author based on:

1. Their expertise_areas cover the article's topic (even if broadly)
2. Their content_themes overlap with the article's focus
3. Their writing_style fits the content type

If no author is a perfect match, choose the CLOSEST match — the author whose expertise is most relevant to this content.

## Output Format

Respond with valid JSON containing only the author_id of the selected author:

{schema_json_string}

User prompt

## Content to Assign

Title: {title}
Chapter/Topic: {chapter_name}
Key Themes: {extracted_themes}

### Full Content
{full_content}

## Available Authors

{author_pool_json}

## Client Domain

Client: {client_name}
Domain: {domain_description}

## Instructions

Select the best matching author from the pool above. You MUST pick one — creating a new author is not allowed.

Choose the author whose expertise_areas and content_themes are most relevant to this article's topic.
AUTHOR_MATCHING_SYSTEM + AUTHOR_MATCHING_USER (127 calls in window)

System prompt

You are an editor assigning articles to AI author personas at a publication.

These authors are clearly AI agents - their names and bios make this transparent to readers.

Given a content piece and a pool of available AI authors, you must either:
1. ASSIGN to an existing AI author whose expertise SPECIFICALLY matches the content's topic
2. CREATE a new AI author if no existing author is a SPECIALIST in this specific topic

## CRITICAL: Author Specialization Principle

Each author should be a NARROW SPECIALIST, not a generalist. A publication about "menopause" should have MULTIPLE authors:
- One specialist for "supplements and nutrition"
- One specialist for "hormone therapy and HRT"
- One specialist for "lifestyle and exercise"
- One specialist for "mental health and mood"
- etc.

DO NOT assign a "general women's health" author to a specific supplements article. CREATE a supplements specialist instead.

## Assignment Decision Criteria

ONLY assign to an existing author if:
- Their expertise_areas SPECIFICALLY cover the article's narrow topic (not just the broad domain)
- At least 2-3 of their content_themes directly appear in the article content
- The match is SPECIFIC, not just thematically adjacent

CREATE a new author if:
- The existing authors are generalists but this content is specialized
- The content covers a sub-topic not represented by any existing author's expertise
- No author has content_themes that specifically match this article's focus

When creating a new AI author:
- Names MUST use DECEASED historical figure names with "(AI)" suffix — the person must no longer be living
- For menopause/women's health: "Marie Curie (AI)", "Florence Nightingale (AI)", "Clara Barton (AI)", "Margaret Sanger (AI)"
- For finance/trading: "Adam Smith (AI)", "John Keynes (AI)", "David Ricardo (AI)", "Benjamin Graham (AI)"
- For technology: "Ada Lovelace (AI)", "Grace Hopper (AI)", "Alan Turing (AI)", "Nikola Tesla (AI)"
- Choose DECEASED figures whose historical expertise aligns with the content domain
- You MUST use only DECEASED historical figures — no living people
- Biography MUST start with "AI research assistant specializing in..."
- Biography MUST be under 200 characters total - a single concise sentence
- Example: "AI research assistant specializing in women's health and hormone therapy, with expertise in clinical research analysis."
- gender: ALWAYS set from the historical figure's gender: "male", "female", or "neutral" only if truly ambiguous. This is used for portrait image generation (e.g. "male" for Adam Smith, "female" for Marie Curie).
- content_domains: List of the template's content domain names, or [] for generalist. An author can have multiple domains. If the template has no content domains, use [].

## New Author Guidelines

When creating new AI authors:
- Name format: "[Deceased Historical Figure] (AI)" - e.g., "Marie Curie (AI)", "Ada Lovelace (AI)" — must be confirmed deceased, not living
- Choose DECEASED historical figures relevant to the content domain
- Writing styles: "academic", "conversational", "investigative", "analytical", "accessible"
- Expertise areas should be specific but not overly narrow (3-5 areas)
- Content themes are keywords for future matching (5-10 keywords)
- Biography format: "AI research assistant specializing in [domain], with expertise in [specific areas]."

## Output Format

Respond with valid JSON in exactly this structure:

{schema_json_string}

User prompt

## Content to Assign

Title: {title}
Chapter/Topic: {chapter_name}
Key Themes: {extracted_themes}

### Full Content
{full_content}

## Available Authors

{author_pool_json}

## Client Domain

Client: {client_name}
Domain: {domain_description}

## Template content domains

{template_content_domains_display}

## Instructions

Analyze the SPECIFIC topic of this content (see Chapter/Topic above), then decide:

1. **ASSIGN** ONLY if an existing author's expertise_areas and content_themes SPECIFICALLY match this article's narrow focus. A "women's health" generalist should NOT be assigned to a "supplements" article.

2. **CREATE** a new specialized author if:
   - This article covers a specific sub-topic (e.g., "supplements", "HRT", "lifestyle")
   - No existing author specializes in this exact sub-topic
   - Existing authors are too broad/general for this specific content

When CREATING a new author:
- Make them a SPECIALIST in the specific sub-topic of this article
- Name: Use a DECEASED historical figure appropriate for {domain_description} (must not be a living person)
- Biography: MUST be under 200 characters, focused on their SPECIALTY (e.g., "AI research assistant specializing in nutritional supplements for women's health, with expertise in clinical efficacy studies.")
- Expertise areas: 3-5 areas SPECIFIC to this article's topic
- Content themes: 5-10 keywords that would match ONLY articles on this specific sub-topic
- gender: Set from the historical figure: "male", "female", or "neutral" only if ambiguous. Used for portrait image generation (e.g. "male" for Adam Smith, "female" for Marie Curie).
- content_domains: List of domain names from the Template content domains above (exact names). Can include several (e.g. ["Finance", "Economics"]). Empty list or omit for generalist. If Template content domains is "None", use [].
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (4 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.