Quality and Cost Benchmarking
Theme discovery is a good example of why “cheapest model” is the wrong starting point. The task, serving mode, and quality bar must be fixed before the prices mean anything.
Choose the model for the actual discovery task
The reported theme-discovery results do not establish a quality winner among GPT-5.6 Sol, Luna, and Terra. Sol scores 8.52 out of 10 (±0.32), Luna 7.97 (±0.35), and Terra 8.21 (±0.47); the uncertainty intervals overlap.[^2][^18]
Their listed costs are not directly comparable. Sol’s result is synchronous at $0.51 per task, while Luna and Terra are batch results at $0.0085 and $0.16.[^2][^18] Luna is the lowest-listed-cost option when the work can wait and its measured quality clears the application bar. Sol is the directly measured route when synchronous discovery is required.
Do not replace discovery with a cheaper adjacent operation. Generating research queries, labeling clusters, assigning material to sections, and selecting sources all begin from a more defined problem. They are different stages with different acceptance criteria.[^1][^8][^16][^21]
Batch reduces cost; it does not prove equal quality
Lower batch prices appear across several models. Gemini 3.6 Flash is listed at $0.03 in batch versus $0.05 synchronously for theme discovery; Claude Haiku 4.5 at $0.04 versus $0.08; and Claude Sonnet 5 at $0.08 versus $0.18.[^2][^8][^13][^16][^18][^22][^23][^28][^29]
Those figures support batch for deferred work. They do not establish that batch and synchronous outputs have equal quality. The same applies to cheaper surrounding steps such as query generation, cluster labeling, and source selection.[^4][^10][^12][^13][^20][^21][^23][^26][^27][^29]
Low cost cannot rescue poor task fit
Claude Sonnet 5 scores 1.02 out of 10 for open-ended theme discovery despite strong results on more bounded work.[^2][^4][^5][^6][^7][^13][^14][^15][^18][^25] GPT-5.4 Nano similarly performs much better at cluster labeling than at content-set relevance scoring.[^3][^4][^9][^11][^13][^17][^19][^24]
The lesson is broader than either model: a low price on the wrong transformation is not a saving.
Routing rule
Use batch for delay-tolerant discovery only after the model clears the discovery-specific quality bar. Keep synchronous service for work that cannot wait. Never use query generation, labeling, source selection, or section assignment as proxy evidence for open-ended discovery.
Set the required bar and compare the exact model-task-mode combinations at https://fronset.ai/benchmark/. Methodology is available at https://fronset.ai/benchmark/methodology/.