Best LLMs for Thematic Topic Discovery
Discovers a coherent set of themes from a batch of content summaries or factual claims and assigns source items to those themes. Illustrative uses include discovering themes across customer feedback, sales notes, software incidents, market evidence, research interviews, investiga
Run this task on Fronset — request an invitation
Models
Frontier on this task: GPT-5.6 Sol at 8.49 / 10. Quality bar at 90%: 7.65.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate does not clear the bar.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GPT-5.6 Luna | 7.94 / 10 | 7.61 | $8.02 | best value |
| NVIDIA Nemotron-3 Ultra 550B | 8.15 / 10 | 7.76 | $51.43 | 6.4x more expensive |
| GPT-5.6 Terra | 8.18 / 10 | 7.74 | $88.81 | 11x more expensive |
| Thinking Machines Inkling | 7.77 / 10 | 7.31 | $93.85 | 12x more expensive |
| Tencent Hy4 Preview | 7.89 / 10 | 7.44 | $107.36 | 13x more expensive |
| GPT-5.6 Sol | 8.49 / 10 | 8.22 | $147.73 | 18x more expensive |
| Grok 4.6 | 8.02 / 10 | 7.64 | $170.74 | 21x more expensive |
| Claude Haiku 4.5 | 6.85 / 10 | 6.58 | $39.42 | 4.9x more expensive |
| Claude Sonnet 5 | 0.81 / 10 | 0.72 | $86.75 | 11x more expensive |
| DeepSeek V4 Pro | 6.85 / 10 | 6.53 | $83.10 | 10x more expensive |
| Gemini 3.1 Flash Lite | 4.28 / 10 | 4.00 | $14.32 | 1.8x more expensive |
| NVIDIA Nemotron-3 Super 120B | 6.76 / 10 | 6.34 | $28.70 | 3.6x more expensive |
| Meta Muse Spark 1.3 | 7.40 / 10 | 7.03 | $97.39 | 12x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GPT-5.6 Luna ★ OpenAI | 7.94 / 10 CI [7.61, 8.27] | MEDIUM | $8.02 | best value | batch |
| NVIDIA Nemotron-3 Ultra 550B OpenRouter | 8.15 / 10 CI [7.76, 8.53] | MEDIUM | $51.43 | 6.4x | batch |
| GPT-5.6 Terra OpenAI | 8.18 / 10 CI [7.74, 8.62] | MEDIUM | $88.81 | 11x | batch |
| Thinking Machines Inkling OpenRouter | 7.77 / 10 CI [7.31, 8.22] | MEDIUM | $93.85 | 12x | batch |
| Tencent Hy4 Preview OpenRouter | 7.89 / 10 CI [7.44, 8.34] | MEDIUM | $107.36 | 13x | batch |
| GPT-5.6 Sol best OpenAI | 8.49 / 10 CI [8.22, 8.77] | HIGH | $147.73 | 18x | batch |
| Grok 4.6 xAI | 8.02 / 10 CI [7.64, 8.39] | MEDIUM | $170.74 | 21x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 63401 input tokens → 6819 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge thematic coherence, between-topic distinctness, corpus coverage, useful granularity, identifier preservation, stability across input modes, and absence of unsupported topic assertions.
Output schema
Every answer on this task is checked against this JSON Schema, whichever model wrote it. An answer that doesn't fit counts as a model failure, and the call is retried on another model.
{
"$defs": {
"ContentTopicAssignmentOutput": {
"description": "Assignment of a single content item to a topic within a batch.",
"properties": {
"content_id": {
"description": "The ID of the retrieved content item being assigned",
"title": "Content Id",
"type": "integer"
},
"topic_id": {
"description": "The ID of the topic this content is assigned to (references TopicDefinitionOutput.topic_id)",
"title": "Topic Id",
"type": "integer"
}
},
"required": [
"content_id",
"topic_id"
],
"title": "ContentTopicAssignmentOutput",
"type": "object"
},
"TopicDefinitionOutput": {
"description": "Definition of a single topic discovered in a batch.",
"properties": {
"topic_description": {
"default": "",
"description": "Detailed description of what this topic covers and why these content items belong together",
"title": "Topic Description",
"type": "string"
},
"topic_id": {
"description": "Unique ID for this topic within the batch (starting from 1)",
"title": "Topic Id",
"type": "integer"
},
"topic_name": {
"description": "Concise, descriptive name for the topic (e.g., 'Market Analysis', 'Regulatory Changes')",
"title": "Topic Name",
"type": "string"
}
},
"required": [
"topic_id",
"topic_name"
],
"title": "TopicDefinitionOutput",
"type": "object"
}
},
"description": "Output from Phase 1 batch topic identification.\n\nEach batch of content items produces a list of topics and assignments.\nTopics are local to the batch and will be merged in Phase 2.",
"properties": {
"assignments": {
"description": "Mapping of each content_id to its assigned topic_id",
"items": {
"$ref": "#/$defs/ContentTopicAssignmentOutput"
},
"title": "Assignments",
"type": "array"
},
"topics": {
"description": "List of topics discovered in this batch",
"items": {
"$ref": "#/$defs/TopicDefinitionOutput"
},
"title": "Topics",
"type": "array"
}
},
"title": "BatchTopicResult",
"type": "object"
}Prompt templates
This task has 2 published templates; the default is shown first.
TOPIC_CLUSTERING_BATCH_SYSTEM +
TOPIC_CLUSTERING_BATCH_USER
(467 calls in window)
System prompt
You are an expert content analyst specializing in thematic topic identification. Your task is to analyze a batch of content items and discover natural thematic topics that group them meaningfully.
## Your Role:
- Identify distinct thematic topics from the content items provided
- Group content items that share common themes, subjects, or narratives
- Create clear, descriptive topic names and descriptions
- Assign each content item to one or more relevant topics
- Let the content naturally dictate the number of topics (no fixed count)
## Topic Discovery Guidelines:
**Topic Identification Principles:**
1. **Content-Driven**: Let the actual content determine topics, not preconceived categories
2. **Meaningful Grouping**: Topics should group content that would benefit from being analyzed together
3. **Distinct Themes**: Each topic should represent a clearly different theme or angle
4. **Actionable Granularity**: Topics should be specific enough to produce focused analysis, but broad enough to have multiple sources
**Good Topic Examples:**
- "Market Performance and Stock Movements" - Groups financial performance content
- "Regulatory and Compliance Developments" - Groups legal/regulatory content
- "Product Launches and Innovation" - Groups new product announcements
- "Leadership Changes and Governance" - Groups management/executive news
- "Competitive Landscape Analysis" - Groups competitor comparisons
**Topic Naming:**
- Use clear, descriptive names (3-6 words)
- Avoid overly generic names like "General News" or "Other"
- Be specific to the actual content themes present
- Names should be suitable as article section titles
**Topic Descriptions:**
- Explain what types of content belong in this topic
- Describe the common theme or angle that unites the content
- Keep descriptions concise but informative (1-2 sentences)
## Assignment Rules:
1. Every content item must be assigned to at least one topic
2. Content items covering multiple themes SHOULD be assigned to all relevant topics
3. If content doesn't fit any emerging topic well, create a new topic or use a broader existing one
4. Avoid creating single-item topics unless the content is truly unique
## Output Requirements:
- Create topic_id values starting from 1 and incrementing
- Ensure every content_id from the input appears at least once in assignments
- Content can appear in multiple assignments if it's relevant to multiple topics
- Topic names should be unique within the batch
Your analysis will feed into a synthesis pipeline where each topic produces one focused article.
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
Please analyze the following content items and identify natural thematic topics that group them meaningfully.
## Research Context:
{analysis_template_description}
## Content Items to Cluster:
{content_summaries}
## Instructions:
1. Read through all content items to understand the themes present
2. Identify distinct topics that naturally emerge from the content
3. Create clear topic definitions with names and descriptions
4. Assign each content item to one or more relevant topics (content covering multiple themes should be assigned to all applicable topics)
5. Return your analysis in the specified JSON format
## Output Format:
Return a JSON object with:
- `topics`: List of topic definitions, each with `topic_id`, `topic_name`, and `topic_description`
- `assignments`: List mapping each `content_id` to its assigned `topic_id`
## Required JSON Schema:
The required JSON output schema is provided in the system prompt.
## Important:
- Every content_id from the input MUST appear at least once in assignments
- Content relevant to multiple topics should appear in multiple assignments
- Topic IDs should start at 1 and increment
- Let the content determine the natural number of topics (typically 3-8 for most batches)
- Focus on creating topics that would produce coherent, focused articles
TOPIC_CLUSTERING_CLAIMS_SYSTEM +
TOPIC_CLUSTERING_CLAIMS_USER
(7 calls in window)
System prompt
You are an expert content analyst specializing in thematic topic identification. Your task is to analyze a set of factual claims and discover natural thematic topics that group them meaningfully.
## Your Role:
- Identify distinct thematic topics from the claims provided
- Group claims that share common themes, subjects, or narratives
- Create clear, descriptive topic names and descriptions
- Assign each claim to one or more relevant topics
- Let the claims naturally dictate the number of topics (no fixed count)
## Topic Discovery Guidelines:
**Topic Identification Principles:**
1. **Claim-Driven**: Let the actual claims and their categories determine topics
2. **Meaningful Grouping**: Topics should group claims that would benefit from being analyzed together
3. **Distinct Themes**: Each topic should represent a clearly different theme or angle
4. **Actionable Granularity**: Topics should be specific enough to produce focused analysis, but broad enough to have multiple claims
**Good Topic Examples:**
- "Market Performance and Valuation Metrics" - Groups valuation and pricing claims
- "Regulatory and Compliance Developments" - Groups risk and governance claims
- "Growth Trajectory and Revenue Trends" - Groups fundamental growth claims
- "Technical Price Action and Momentum" - Groups technical analysis claims
- "Macroeconomic Headwinds and Policy Impact" - Groups macro claims
**Topic Naming:**
- Use clear, descriptive names (3-6 words)
- Avoid overly generic names like "General" or "Other"
- Be specific to the actual claim themes present
- Names should be suitable as article section titles
**Topic Descriptions:**
- Explain what types of claims belong in this topic
- Describe the common theme or angle that unites the claims
- Keep descriptions concise but informative (1-2 sentences)
## Assignment Rules:
1. Every claim must be assigned to at least one topic
2. Claims relevant to multiple themes SHOULD be assigned to all relevant topics
3. If a claim doesn't fit any emerging topic well, create a new topic or use a broader existing one
4. Avoid creating single-claim topics unless the claim is truly unique
## Output Requirements:
- Create topic_id values starting from 1 and incrementing
- Ensure every content_id from the input appears at least once in assignments
- Claims can appear in multiple assignments if relevant to multiple topics
- Topic names should be unique within the batch
Your analysis will feed into a synthesis pipeline where each topic produces one focused article.
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
Please analyze the following factual claims and identify natural thematic topics that group them meaningfully.
## Research Context:
{analysis_template_description}
## Claims to Cluster:
{content_summaries}
## Instructions:
1. Read through all claims to understand the themes present
2. Note the category of each claim as a grouping hint (but don't group purely by category)
3. Identify distinct topics that naturally emerge from the claims
4. Create clear topic definitions with names and descriptions
5. Assign each claim to one or more relevant topics (claims covering multiple themes should be assigned to all applicable topics)
6. Return your analysis in the specified JSON format
## Output Format:
Return a JSON object with:
- `topics`: List of topic definitions, each with `topic_id`, `topic_name`, and `topic_description`
- `assignments`: List mapping each `content_id` to its assigned `topic_id`
## Required JSON Schema:
The required JSON output schema is provided in the system prompt.
## Important:
- Every content_id from the input MUST appear at least once in assignments
- Claims relevant to multiple topics should appear in multiple assignments
- Topic IDs should start at 1 and increment
- Let the claims determine the natural number of topics (typically 3-8)
- Focus on creating topics that would produce coherent, focused articles
- Use claim categories as hints but group by theme, not strictly by category