How Long Will This Take?

A model does not have one response time. An operation has a response time on a particular model—and the difference between operations can be larger than the difference between models.

Across 699,136 synchronous calls, the median response time was 5.0 seconds, but the 90th percentile was 51 seconds.[^1][^2][^3][^4][^5][^6][^7][^8][^9][^10][^11][^12][^13][^14][^15][^16][^17][^18][^19][^20][^21][^22][^23][^24][^25] A product designed around the median would still miss the slower tenth of requests. Tail latency, not just the typical call, should set the deadline.

The interactive lane is real—but narrow

Bounded decisions can be fast. GPT-5.4 Nano detected language in 860 ms, matched authors in 970 ms, scored topic-report relevance in 1.2 seconds, and named topic clusters in 1.5 seconds at the median.[^3][^29][^3][^5][^41][^3][^20][^39][^3][^13][^32]

The same model family took 18 seconds for reference-preserving analyst prose.[^3][^9][^44] Claude Sonnet 5 assigned clusters to sections in 1.7 seconds and generated discovery clusters in 2.6 seconds, but theme generation took 11 seconds and topic-to-client matching 14 seconds.[^1][^3][^33][^37][^56][^59]

Gemini 3.5 Flash shows an even sharper split: 4.4 seconds to assign material to known sections, 7.4 seconds to name clusters, 19 seconds to sequence topics, and 51 seconds to discover and cluster them.[^1][^3][^33][^3][^13][^32][^1][^3][^40][^1][^3][^37]

The pattern: known labels, short classifications, and simple matching can sit in an interactive path. Open-ended discovery and evidence-linked writing often cannot.

Open-ended work needs a patient or asynchronous path

Qwen 3.7 Plus took 100 seconds for topic discovery and 168 seconds for topic-to-client matching. Tencent Hy4 Preview took 208 seconds for thematic discovery and 314 seconds for taxonomy matching.[^1][^3][^37][^59][^36][^58]

Long-document work behaves similarly. Moonshot Kimi K3 recorded 90 seconds for structured extraction, 117 seconds for synthesis, 191 seconds for claim-referenced writing, and 210 seconds for topic-client matching.[^1][^3][^62][^31][^35][^9][^11][^44][^51][^59]

These are not universal model rankings—the operations differ—but they consistently show why a “fast model” label is unsafe.

Matched work can change the wait dramatically

Where the same operation is measured, model choice can still matter. For atomic factual-claim extraction, Claude Opus 5 recorded an 8-second median, versus 54 seconds for GLM-5.3 Flash and 66 seconds for Grok 4.6; GLM-5.3 Flash reached 230 seconds at the 90th percentile.[^1][^3][^28]

GPT-5.6 Luna completed structured fact extraction in 14 seconds, while GPT-5.6 Sol took 44 seconds.[^1][^3][^62] DeepSeek V4 Flash was also materially faster than V4 Pro on the cited promotional writing, prompt-adaptation, and title-generation workloads.[^3][^8][^16][^19][^26][^27][^43][^54]

These are speed results, not overall recommendations. Quality and cost still need to be checked on the same operation.

Batch changes the question

Half of completed batches returned within 5.5 minutes and 75.1% within 15 minutes, but 5.6% took longer than an hour.[^1][^25] At that point the key question is no longer “Will a user wait?” but “Can the workflow tolerate delayed completion, cancellation, and redispatch?”

Routing conclusion

Keep detection, classification, short matching, cluster naming, and section assignment interactive only when the measured model-operation pair fits the product’s tail-latency budget. Move discovery, substantial extraction, synthesis, client matching, and citation-backed writing to asynchronous or explicitly patient flows.

Inspect task-level median and tail latency alongside quality and price at https://fronset.ai/benchmark/.