Batch versus sync: cost is material, but routing remains task-specific
Batch can reduce the price of LLM work substantially. It does not make the wrong model right, preserve quality by definition, or satisfy a user who needs an immediate answer.
Of 68 tasks, 50 are batch-eligible and 18 are not. Batch handled 410,019 of 1,109,798 successful calls, or 37%.[^1][^49] Eligibility belongs to a particular model-task configuration; a sync-only route does not become cheaper just because the surrounding pipeline uses batch.
The discounts are real
Matched price pairs repeatedly favor batch. Gemini 3.5 Flash costs $0.07 rather than $0.12 for taxonomy matching and $0.05 rather than $0.09 for content-set relevance.[^5][^12][^17][^19][^40][^47][^51][^52]
GPT-5.6 Sol falls from $0.64 to $0.18 for reference-preserving analysis, while GPT-5.5 falls from $0.08 to $0.03 for image-prompt generation.[^3][^9][^25][^31][^45][^53] Claude Sonnet 5 also has lower listed batch prices for cluster labeling, factual-claim extraction, relevance scoring, profile matching, decision synthesis, and metadata rewriting.[^10][^13][^14][^18][^20][^21][^23][^29][^33][^39][^42][^44][^50][^52]
These comparisons establish a price direction. They do not establish that batch and sync outputs have equal quality.
Batch can change the economic winner
Image-prompt generation provides the clearest example. GPT-5.6 Terra scores 9.08 at $0.01 synchronously; GPT-5.6 Luna scores 8.90 at $0.0006 in batch. Their confidence intervals overlap, so the measured quality is a tie.[^3][^31]
For a deferrable workload, Luna’s much lower listed cost can decide the route. GPT-5.4 Nano is cheaper than Terra but scores only 7.38, so price alone does not make it interchangeable with the tied higher-quality options.[^3][^31]
Other batch results remain useful without proving cross-mode equivalence. Gemini 3.5 Flash scores 9.78 for batch structured extraction at $0.05, and Claude Sonnet 5 scores 9.32 for batch newsletter copy at $0.0041.[^2][^11][^36][^43] Without same-model synchronous quality results, neither comparison proves that moving the task to batch is quality-neutral.
The price card is not total cost
Gemini 3.1 Pro Preview required a separately billed repair call on 1.7% of 15,389 attempts. MiniMax M2.5 required one on 17.2% of 40,876 attempts, in addition to substantial local repair.[^1][^49] Shadow traffic to models whose answers will not be selected can add more cost.
A production comparison should therefore include repair calls, local processing, parse failures, redispatch, and comparison traffic—not just the displayed per-task price.
Routing conclusion
Use batch when all three conditions hold:
The work can wait.
The exact model-task configuration supports batch and clears the quality bar.
Total completed-task cost remains lower after repairs and comparison traffic.
Keep synchronous service for immediate-response paths. Do not compare prices across different tasks or assume that a batch discount changes the quality ranking.