Models That Get It Right First Time

For an unattended workflow, success is not “the API returned text.” Success is “the output can enter the next step without repair, retry, or human salvage.”

First-pass reliability can be measured directly

Meta Muse Spark 1.3 Contributor produced usable output on the first attempt in 99.9% of 30,332 calls. Tencent Hy3 reached 99.6% across 23,887 attempts. Other strong reported rates include GPT-5.6 Luna at 98.8%, Claude Opus 4.7 at 98.7%, Claude Sonnet 4.6 at 98.5%, Gemini 3.1 Flash Lite at 96.5%, and Qwen 3.5 Flash at 96.0%.[^1][^3][^4][^5][^6][^12][^16][^17][^18][^19]

These figures identify candidates for low-intervention paths. They do not prove that every output follows a particular schema or meets the quality bar for every task. First-pass usability should be an admission gate, followed by task-specific validation.

A returned response can still create work

MiniMax M3 failed outright on only 0.3% of 23,451 attempts, yet 25.1% of its outputs required repair. Only 74.6% were usable as returned.[^2][^3][^14][^16] Counting only hard failures would therefore make the route look much cleaner than it was.

The cost is not theoretical. In Markdown newline repair, $11.06 of $62.12 in billed generation—17.8%—could not be parsed.[^8][^11] That figure does not include the unknown cost and delay of a second repair call. Headline price is not completed-task cost when invalid output triggers more work.

Quality and completion can disagree

For public-response drafting, GPT-5.6 Sol scored 8.53 when it answered but answered on 67% of attempts. GPT-5.6 Terra scored 8.2 and answered on 83%.[^10][^15] Sol offers the stronger returned answer; Terra offers a better chance of receiving one. The preferred route depends on which failure the product can tolerate.

High availability also does not create a universal quality default. Terra was selected best in none of 47 claim-extraction comparisons, and Luna in none of 42 promotional social-post comparisons.[^7][^9][^13][^20] Reliability and task quality must both clear the bar.

Routing conclusion

For workflows that cannot absorb a retry, start with models whose usable-first-response rate meets the service requirement. On structured paths, validate parsing and count only output usable as returned. Model the whole completion path—including billed unusable tokens, repair calls, fallbacks, and added latency—rather than selecting on headline price or response rate alone.

Compare first-pass behavior with task-specific quality at https://fronset.ai/benchmark/.