The Right-Sized Model Portfolio
The Right-Sized Model Portfolio
A production portfolio should be organized around transformations, not model reputations. The same model can be excellent after the direction is known and poor at deciding that direction in the first place.
Claude Sonnet 5 is the clearest example: 1.02 out of 10 for discovering themes, 8.38 for labeling existing clusters, and 9.32 for newsletter copy.[^12][^13][^49][^61][^62] One “general-purpose” route would hide three very different capabilities.
The portfolio below assumes an 8 out of 10 quality bar for automated output.
Separate discovery from execution
GPT-5.6 Sol scored 8.52 for theme discovery, while GPT-5.6 Terra scored 8.06 for assigning material to sections after the structure already existed.[^8][^12][^54][^61] These are complementary stages, not competing answers to the same problem.
Domain suggestion, cluster labeling, and taxonomy matching also need separate routes. Thinking Machines Inkling Small was selected best in 28 of 33 domain-suggestion comparisons. Grok 4.5 scored 8.62 for cluster labeling but only 4.66 for domain suggestion. Claude Haiku 4.5 scored 8.03 for cluster labeling but 5.4 for taxonomy matching.[^13][^16][^21][^48][^49][^56][^64]
The useful portfolio therefore hands work from a qualified discovery route to lower-cost organization and deliverable-specific publishing routes.
Use fast models only where the task is bounded
GPT-5.4 Nano detected language in 860 ms, scored topic-report relevance in 1.2 seconds, and named clusters in 1.5 seconds.[^28][^31][^39][^60][^63][^65] It also scored 9.92 for language identification and 9.46 for structured extraction.[^1][^18][^44][^52]
But Nano scored only 3.42 for content-set relevance.[^10][^23][^27][^46][^50][^59] Use it to identify, name, and extract. Do not let its speed become a license to make relevance-sensitive judgments it has not demonstrated.
Keep a specialist lane for analytical writing
Moonshot Kimi K3 scored 9.65 for report outlines and 9.10 for reference-preserving analytical writing.[^17][^20][^22][^34][^47][^57][^58][^66] GPT-5.6 Terra is the lower-cost outline alternative at 9.48 and remains a plausible fallback for analytical writing when its lower measured quality clears the application bar.[^17][^22][^34][^47][^55][^57][^58]
This specialist lane is worth preserving because reusable structure and reference-linked prose are expensive places to accept silent degradation.
Recommended portfolio
| Activity | Default route | Fallback | Why |
|---|---|---|---|
| Theme discovery | GPT-5.6 Sol | Human review | Sol clears the assumed bar on the directly measured discovery task; no equally supported fallback is established here. [^12][^61] |
| Language identification and structured extraction | GPT-5.4 Nano | DeepSeek V4 Pro | Nano combines strong quality with sub-second language detection; Pro has a higher reported extraction score and answer rate. [^1][^18][^31][^44][^52][^63] |
| Content-domain suggestion | Thinking Machines Inkling Small | Human review | Inkling Small has direct best-answer evidence; Grok’s strong cluster-labeling result does not transfer to domain suggestion. [^21][^56][^64] |
| Topic-cluster labeling | Grok 4.5 | Claude Haiku 4.5 | Both have direct labeling evidence, with Grok stronger on the reported point result. [^13][^49] |
| Assigning material to established sections | GPT-5.6 Terra | Human review | Terra’s result directly covers section assignment after a structure exists. [^8][^54] |
| Report outlines and reference-preserving analysis | Moonshot Kimi K3 | GPT-5.6 Terra | Kimi is the quality-led route; Terra is the lower-cost outline alternative. [^17][^20][^22][^34][^47][^55][^57][^58][^66] |
| Newsletter copy and visual-theme configuration | Claude Sonnet 5 | Human review | Sonnet clears the bar on both publishing deliverables, but the evidence does not establish a like-for-like model fallback. [^11][^14][^43][^62] |
This is not a universal minimum portfolio. It is the smallest set supported by these measurements without pretending that discovery, organization, relevance, synthesis, and publishing are interchangeable.
Portfolio controls matter as much as the table
Batch can lower cost, but half of batches returned within 5.5 minutes, 75.1% within 15 minutes, and 5.6% took longer than an hour.[^2][^42] Synchronous calls had a 5-second median and 51-second 90th percentile across the reported population.[^2][^3][^5][^6][^7][^9][^15][^18][^24][^25][^26][^28][^29][^30][^31][^32][^33][^35][^36][^37][^38][^39][^40][^41][^42]
High quality can also be conditional: DeepSeek V4 Pro, GPT-5.4 Nano, and MiniMax M3 answered on 80%, 76%, and 58% of structured-extraction attempts.[^1][^44] MiniMax M3 required repair on 25.1% of structured outputs.[^2][^4][^42][^51]
A portfolio is complete only when each lane has deadlines, validation, retries, fallbacks, and usable-output accounting.
Set the quality bar and compare the exact routes at https://fronset.ai/benchmark/.