Best Models for Long-form Content Generation
Production-Grade Newsletter Copy for $0.0007 a Run, and Why You Still Test Every Task
Two findings from our long-form benchmarks stand out. The cheapest run that clears a production quality bar is MiniMax M3 writing newsletter copy, at a small fraction of a cent, and it lands only a quarter of a point below the best score in the whole set. The same cheap model then posts the weakest score in the set on a neighboring task, so a model’s price and its headline score tell you little about the next job you hand it.
How the numbers are produced
These results rest on 57,538 scored judgments across 833 measured model-and-task combinations, covering 30 models on 52 tasks. Each cell is a quality score out of 10 with a confidence interval, paired with a per-run cost that is either measured on that model’s own usage or estimated where the cell has no qualifying usage history of its own. The figures were published on a single date, but the measurement behind them is continuous and ongoing.
For this piece we treat around 9.0 out of 10 as production-ready for drafts that ship without heavy review. Scores in the 8.6 to 8.9 band are good, but they leave a reviewer in the loop.
The top score and the cheap score are a quarter point apart
The highest score in the set is Moonshot Kimi K3 generating taxonomies for publication sections, at 9.33 out of 10 (±0.06) and $0.03 per task run in sync mode, measured on this model’s own usage [^1][^2]. It is both the highest score and the highest per-run cost shown.
The cheapest run that still clears the bar is MiniMax M3 writing newsletter copy, at 9.08 out of 10 (±0.03) and $0.0007 per task run in sync mode, measured on this model’s own usage and repriced at current rates [^3][^4].
Taken at face value, the gap between those two runs is 0.25 points in quality and $0.0293 per task run in listed cost [^1][^2][^3][^4]. Both prices are sync-mode per-run costs, but they are for different deliverables, a publication-section taxonomy against newsletter copy, so treat the arithmetic as illustrative rather than as a like-for-like saving on the same job.
The newsletter result is not a one-off for MiniMax M3. Translating text in batches comes in at 8.87 out of 10 (±0.24) and $0.0011 per task run in sync mode, measured on this model’s own usage [^5][^6], and repairing newlines in Markdown scores 8.76 out of 10 (±0.42) and $0.004 per task run in sync mode, measured on this model’s own usage and repriced at current rates [^7][^8]. Neither clears the production bar, but at those prices a review pass is cheap.
The same model is the weakest in the set on taxonomy matching
Matching topics to an existing taxonomy is MiniMax M3’s clear weak spot, at 6.56 out of 10 (±0.2) and $0.0064 per task run in sync mode, measured on this model’s own usage and repriced at current rates [^9]. Same model, similar price, output you would not put near production.
Synthesizing a decision from multiple perspectives tells a similar story at 7.85 out of 10 (±0.1) and $0.02 per task run in sync mode, and that cost is estimated because the cell has no qualifying usage history of its own [^10][^11]. It is also the highest cost shown for MiniMax M3 in this set, so on that task you pay more and get less.
The lesson generalizes beyond one model. Strength does not transfer across neighboring long-form tasks, and the benchmark does not support standardizing taxonomy matching on the newsletter winner.
Cheap alternatives sit near the top, too
Testing per task also cuts the other way: the expensive peak is rarely the only option. On the same publication-section taxonomy task where Kimi K3 leads, Thinking Machines Inkling scores 8.97 out of 10 (±0.12) at $0.01 per task run in sync mode [^1][^2], a low-cost alternative close to the top.
Title packages show the same spread within one task, with every cost below a per-run price in sync mode [^12][^13].
| Model | Title-package quality | Cost per run |
|---|---|---|
| Moonshot Kimi K3 | 9.12 out of 10 (±0.13) | $0.07 |
| Thinking Machines Inkling | 8.75 out of 10 (±0.15) | $0.03 |
| Meta Muse Spark 1.1 | 8.67 out of 10 (±0.15) | $0.02 |
Kimi K3 buys the only score above the bar among these three, but whether its top quality justifies the higher cost is a decision about your product, not about the model. If a reviewer already reads every title package, the mid-8s options are hard to argue against.
Where to look next
The practical rule is short: pick the task first, then look at the cell, and read the cost qualifier before you trust the price. The live numbers are at llm-bench, with per-task pages such as the newsletter copy and taxonomy tasks linked in the sources below. Set the quality bar your documents actually need there and compare batch against sync pricing yourself before you lock in a route.