A model can return clean output almost every time and still be the wrong model for the job. Meta Muse Spark 1.3 contributor produced a usable first response on 99.9% of 30,332 attempts, while Tencent Hy3 did so on 99.6% of 23,887 attempts.[^20][^3][^21] Those rates reveal how much repair and retry handling a workflow may need. They do not establish task quality.

Reliability is a gate, not a ranking

Several models report high usable-first-response rates:

ModelUsable first response without repair or retry
Meta Muse Spark 1.3 contributor99.9% of 30,332 attempts [^20]
Tencent Hy399.6% of 23,887 attempts [^3][^21]
GPT-5.6 Luna98.8% of 95,320 attempts [^1][^15]
claude-opus-4-798.7% of 5,148 attempts [^2][^20]
claude-sonnet-4-698.5% of 22,204 attempts [^2][^20]
Gemini 3.1 Flash Lite96.5% of 157,620 attempts [^5][^6][^23][^25]
qwen3.5-flash96.0% of 71,276 attempts [^2][^20]

For a workflow that cannot tolerate a retry, this is a useful first filter. The next question is still task-specific: does the model perform the required work well enough?

Models in the same family can require different routes

Public-response drafting shows the trade-off clearly. GPT-5.6 Sol scores 8.53 out of 10 when it answers, versus 8.2 for GPT-5.6 Terra. Terra, however, answers on 83% of attempts, compared with 67% for Sol.[^13][^18] Sol offers the higher conditional score; Terra offers greater coverage.

The split extends beyond that task. GPT-5.6 Luna was selected as the best answer in 19 of 43 content-summarization comparisons.[^11][^22] Terra was selected in 13 of 43 claim-referenced analyst-writing comparisons, 13 of 39 query-validation comparisons, and 9 of 29 image-prompt comparisons.[^4][^8][^9][^17][^19][^29]

But Terra was selected in none of 47 claim-extraction comparisons, while Luna was selected in none of 42 promotional social-post comparisons.[^7][^10][^16][^26] A family name does not create a universal default. Each variant still needs evidence for the exact operation.

Tail latency can overturn an otherwise attractive choice

For Markdown newline repair, median synchronous latency was 16 seconds for Luna, 19 seconds for Terra, and 39 seconds for Sol.[^2][^12][^27] Luna’s content-summarization result also shows why the median is not enough: its 8.6-second median rises to 37 seconds at the 90th percentile.[^11][^12][^22]

A user-facing path sized only to the median will miss the slower calls that shape the actual experience.

Version numbers do not transfer performance across tasks

Meta Muse Spark illustrates the same boundary. Version 1.2 recorded a 5.9-second median for translation and 17 seconds for voice-profile generation, while muse-spark-1.1 recorded 116 seconds for Markdown newline repair.[^2][^12][^27][^28][^30]

The available 1.3 evidence concerns reliability and quality: prompt adaptation rose to 8.93, and language identification to 9.86.[^14][^31] Those gains do not show that version 1.3 is faster on translation, voice generation, or Markdown repair.

Routing decision

Use first-pass usability to narrow the field when repair or retry is unacceptable. Then choose on the exact task, and size the interaction budget to tail latency rather than a model-family reputation.

That means using Luna for summarization only when its slower tail fits the product, choosing between Sol and Terra for public responses based on quality versus coverage, and budgeting Meta Muse Spark by the measured task and version.

Explore the live task results at https://fronset.ai/benchmark/ and the methodology at https://fronset.ai/benchmark/methodology/.