Model intelligence briefing
How the leading models actually compare
A concise, sourced view of where each provider is stronger, where they trail, and what that means for the work you want built. Every figure links to its source.
Data as of 2026-08-13
Ranking order matters less than fit.
Seven labs place a model in the Epoch Capabilities Index top twenty, and the top ten sit inside a band of under six points. At that spacing the question worth asking is not which model is best but which is best at the job in front of you. That is the question this briefing is organised around.
Epoch Capabilities Index
Aggregate general-capabilities score. Where a model has several reasoning-effort variants, the best scoring one is used.
LMArena text rating
Live blind pairwise voting; higher is better. Where a model has several variants, the best scoring one is used.
Maximum context window
Tokens the model can attend to in a single request. Where a model has multiple listings, the largest is shown.
Context window ceiling
Maximum input tokens from the LiteLLM price table, in thousands. Where a model has multiple listings, the largest is shown.
Context stopped being a differentiator.
A million-token window is table stakes at the top of the catalog now, and the ceiling sits at two million. What varies is whether long context stays coherent across the window, and whether you pay a surcharge above 200K. Treat the headline number as a ceiling rather than a working range.
SWE-bench Verified
Verified pass rate on SWE-bench, expressed as a percentage. Where a model has several variants, the best scoring one is used.
Frontend code arena, Elo
LMArena webdev leaderboard. Where a model has several style-control variants, the highest rating is used.
Benchmarks and human preference disagree.
The models that top automated issue-resolution benchmarks are not the models people prefer when they cannot see which is which. Kimi K3 leads the blind frontend arena while trailing on SWE-bench, and the reverse holds elsewhere. If your work is repository-scale refactoring, weight the benchmarks. If it is interface work a human will look at, weight the arena.
Image generation and editing, Elo
LMArena image leaderboards. A model missing from the editing board is left blank for that series.
Not every lab competes here.
Anthropic, Moonshot and DeepSeek generate no images or video at all, which is a product decision rather than a capability gap. Where a matrix cell is hatched in this briefing it means the capability is not offered, not that it performs poorly.
Security ordering differs most sharply from raw capability.
Capability and security posture are close to uncorrelated across this field. A model can sit near the top of every intelligence board and near the bottom of every red-team board. If a model is going to touch customer data, run unsupervised, or execute tool calls against production systems, this is the axis that should decide the choice.
Output pricing per million tokens
USD per 1M output tokens. Lower is cheaper. Where a model has multiple listings, the cheapest is shown.
API output price per million tokens
USD per 1M output tokens from the LiteLLM price table, scoped to models in the current ECI top 20. Where a model has multiple listings, the cheapest is shown.
List price is not cost.
Output price per million tokens spans two orders of magnitude across models that score within twenty points of each other. But list price is a poor proxy for cost per completed task: a cheaper model that needs three attempts is not cheaper. Batch processing typically discounts fifty percent, and cache reads can run a tenth of base, so the effective figure for a production workload often looks nothing like the headline.
Model roster
Current frontier and frontier-adjacent models, with published context, pricing and licence posture.
| Label | Vendor | Context (K) | Input USD/1M | Output USD/1M |
|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 1000 | 5 | 25 |
| Claude Fable 5 | Anthropic | 1000 | 10 | 50 |
| Claude Sonnet 5 | Anthropic | 1000 | 3 | 15 |
| GPT-5.6 Sol | OpenAI | 1050 | 5 | 30 |
| GPT-5.6 Terra | OpenAI | 1050 | 2.5 | 15 |
| GPT-5.6 Luna | OpenAI | 1050 | 1 | 6 |
| Gemini 3.1 Pro | 1000 | 2 | 12 | |
| Gemini 3.6 Flash | 1000 | 1.5 | 7.5 | |
| Grok 4.5 | SpaceXAI | 500 | 2 | 6 |
| Grok 4.1 Fast | SpaceXAI | 2000 | 0.2 | 0.5 |
| Kimi K3 | Moonshot | 1000 | 3 | 15 |
| DeepSeek V4 Pro | DeepSeek | 1000 | 0.435 | 0.87 |
| DeepSeek V4 Flash | DeepSeek | 1000 | 0.14 | 0.28 |
| Qwen 3.7 Max | Alibaba | 1000 | 2.5 | 7.5 |
| Mistral Large 3 | Mistral | 256 | 0.5 | 1.5 |
| Muse Spark 1.1 | Meta | 1000 | 1.25 | 4.25 |
| Llama 4 Scout | Meta | 10000 | 0 | 0 |
What developers actually run
Top 20 models by tokens processed through OpenRouter on a single day. Where a model has multiple entries, the highest total is shown. Source: OpenRouter (openrouter.ai/rankings), as of 2026-08-12.