Latency is not a model-wide property. It belongs to a model–operation pair. In one Claude Sonnet 5 topic workflow, median latency ranges from 1.7 seconds for section assignment to 19 seconds for vetted-site selection.[^1][^2][^3][^4][^6][^7][^8][^9][^14][^15][^19][^21][^22][^23] A single timeout for “topic analysis” is therefore a poor production design.

One workflow needs several latency budgets

Claude Sonnet 5 assigned clusters to sections in 1.7 seconds, generated discovery clusters in 2.6 seconds, named clusters in 3.7 seconds, and sequenced topics in 6.9 seconds. Theme development took 11 seconds and topic-to-client matching 14 seconds.[^1][^2][^3][^4][^6][^7][^8][^9][^14][^15][^19]

The adjacent research stages vary too: community selection took 3.0 seconds, query generation 4.5 seconds, and vetted-site selection 19 seconds.[^21][^22][^23] The faster stages can remain interactive, while slower selection and matching steps need their own budgets.

Fast labeling does not predict discovery time

Bounded operations can be very fast. Gemini 3.1 Flash Lite named clusters in 1.2 seconds and scored topic-report relevance in 949 milliseconds. GPT-5.4 Nano recorded 1.5 seconds and 1.2 seconds on the same kinds of work.[^1][^4][^5][^6][^12][^13]

But even within one model, the workload changes the result. Gemini 3.5 Flash assigned material to cluster sections in 4.4 seconds and named clusters in 7.4 seconds, while topic discovery and clustering took 51 seconds and sequencing took 19 seconds.[^1][^3][^4][^6][^7][^8][^9]

Labeling an existing structure and discovering a new one should not share a latency assumption.

Discovery, taxonomy matching, and client matching can take minutes

Qwen 3.8 Flash labeled clusters in 27 seconds, but thematic discovery took 156 seconds and taxonomy matching 204 seconds.[^11][^16][^17] Tencent Hy4 Preview took 74 seconds for section assignment, 208 seconds for thematic discovery, and 314 seconds for taxonomy matching.[^1][^3][^10][^16][^17]

Client matching can become the bottleneck as well. Moonshot Kimi K3 named clusters in 24 seconds but matched topics to clients in 210 seconds. Thinking Machines Inkling Small moved from 14 seconds for naming to 68 seconds for client matching.[^1][^3][^4][^6][^15][^19]

These stages belong on a longer-running path whenever their medians exceed the product’s interaction budget.

What latency alone cannot decide

These measurements report response time, not output quality or cost. A slower model may still be justified by a higher task-specific quality bar, and a fast model is not automatically the better production route.

Keep bounded labeling, relevance checks, naming, and section assignment interactive only where the measured operation fits the user experience. Move open-ended discovery, taxonomy matching, and slower client matching to an asynchronous or explicitly patient path. Then compare quality and pricing for the same operation before selecting the model.

Explore the live task results at https://fronset.ai/benchmark/ and the methodology at https://fronset.ai/benchmark/methodology/.