Model intelligence briefing

How the leading models actually compare

A concise, sourced view of where each provider is stronger, where they trail, and what that means for the work you want built. Every figure links to its source.

Data as of 2026-08-13

The shape of the field

Ranking order matters less than fit.

Seven labs place a model in the Epoch Capabilities Index top twenty, and the top ten sit inside a band of under six points. At that spacing the question worth asking is not which model is best but which is best at the job in front of you. That is the question this briefing is organised around.

Epoch Capabilities Index

Aggregate general-capabilities score. Where a model has several reasoning-effort variants, the best scoring one is used.

OpenAIAnthropicMoonshotAlibabaGoogle DeepMindGooglexAI

LMArena text rating

Live blind pairwise voting; higher is better. Where a model has several variants, the best scoring one is used.

AnthropicAlibabaGoogle DeepMindMeta AIOpenAIXiaomi CorpZ.aiMoonshotDeepSeekxAI
Source:LMArena leaderboard datasetCC-BY-4.0

Maximum context window

Tokens the model can attend to in a single request. Where a model has multiple listings, the largest is shown.

xAIMetaOpenAIXiaomi CorpDeepSeekGoogle DeepMindMoonshotMeta AIZ.aiMiniMaxGoogle
Source:OpenRouter model catalogPublic catalog

Context window ceiling

Maximum input tokens from the LiteLLM price table, in thousands. Where a model has multiple listings, the largest is shown.

xAIOpenAIDeepSeekGoogle DeepMind,GoogleGoogle DeepMindGoogleMeta AI
Context

Context stopped being a differentiator.

A million-token window is table stakes at the top of the catalog now, and the ceiling sits at two million. What varies is whether long context stays coherent across the window, and whether you pay a surcharge above 200K. Treat the headline number as a ceiling rather than a working range.