Models
Frontier & open-weight models compared by capability — columns are flagship models, rows are the headline benchmarks that measure each one. Every number is the published figure, restated with a link to the source it came from · as of 2026-06-14. Curated matrix is 70d old — a refresh is due; live leaderboard moves still flow into Model changes below.
We cite the benchmark authorities rather than re-rank them — Epoch AI, LMArena and Artificial Analysis. ChangeRadar's job is the changes: what gets released, repriced, deprecated, or quietly shifts behind a stable model id.
Vendor champions · best flagship per vendor · GPQA Diamond
Each vendor's most recent publicly-available flagship, ordered by score (highest first) — Google · OpenAI · Anthropic · xAI highlighted. The weekly market-watch surfaces new releases automatically; one tagged reported is the latest release shown with vendor-reported scores (linked to source) until we independently cite it. Score = GPQA Diamond; every number links to its source.
Compare two models
Agentic coding
| Benchmark | Claude Fable 5 | Gemini 3.1 Pro | Δ |
|---|---|---|---|
| SWE-bench Verified i% resolved (pass@1) | 95 | 80.6 | +14.4 |
| SWE-bench Pro i% resolved (pass@1) | 80.3 | 54.2 | +26.1 |
3 coverage gaps — only one model reports these
- Terminal-Bench — Gemini 3.1 Pro 68.5
- LiveCodeBench — Gemini 3.1 Pro 2887
- FrontierCode — Claude Fable 5 29.3
Tool use & agents
5 coverage gaps — only one model reports these
- TAU-bench — Gemini 3.1 Pro 99.3
- OSWorld — Claude Fable 5 85
- BrowseComp — Gemini 3.1 Pro 85.9
- GDPval-AA — Claude Fable 5 1932
- MCP Atlas — Gemini 3.1 Pro 69.2
Science & reasoning
| Benchmark | Claude Fable 5 | Gemini 3.1 Pro | Δ |
|---|---|---|---|
| GPQA Diamond i% accuracy | 92.6 | 94.3 | -1.7 |
2 coverage gaps — only one model reports these
- Humanity's Last Exam — Gemini 3.1 Pro 44.4
- ARC-AGI-2 — Gemini 3.1 Pro 77.1
General knowledge
1 coverage gap — only one model reports these
- MMMLU — Gemini 3.1 Pro 92.6
Multimodal
2 coverage gaps — only one model reports these
- MMMU — Gemini 3.1 Pro 80.5
- GDP.pdf — Claude Fable 5 29.8
Long context
1 coverage gap — only one model reports these
- MRCR — Gemini 3.1 Pro 84.9
Shared benchmarks first (with Δ when both report the same scale); one-sided coverage collapses below. Pick any two models — or a champion above — and the URL becomes shareable.
All models · benchmark matrix
reported columns (Claude Opus 5, GPT-5.6 Sol, Grok 4.6, Qwen3.8-Max, DeepSeek-V4-Pro-Max, Kimi K3) are auto-discovered by our weekly market-watch from each vendor's own reported numbers — not independently verified, and shown when a vendor ships a model newer than the hand-cited column beside it. Full claim sets are in Vendor-reported benchmarks below.
Agentic coding
Can the model fix real bugs, ship features, and operate a dev environment end-to-end as a coding agent — the single most-watched capability in 2026 vendor launches.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SWE-bench Verified i % resolved (pass@1) | 95 | 97 | 88.6 | — | — | — | 80.6 | — | — | 80.6 | 80.6 | 80.4 | — | 80.2 | 76.8 | — | — | 77.6 | — |
| SWE-bench Pro i % resolved (pass@1) | 80.3 | 79.2 | 69.2 | 77.8 | 58.6 | — | 54.2 | — | — | — | — | 60.6 | 67.7 | 58.6 | — | 62.1 | — | — | — |
| Terminal-Bench i % solved (pass@1) | — | 68.9 | 74.6 | 88 | 82.7 | 88.8 | 68.5 | — | 26 | 67.9 | — | 69.7 | 86.6 | 66.7 | 88.3 | 81 | — | — | — |
| LiveCodeBench i % pass@1 | — | — | — | — | — | — | 2887 | — | 88.4 | 93.5 | 93.5 | — | — | 89.6 | — | — | 43.4 | — | 29.7 |
Tool use & agents
Beyond writing code: can the model select and chain tools, follow policy, drive a computer/browser, and complete long-horizon multi-step tasks.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TAU-bench i % pass / pass^k | — | — | — | — | — | — | 99.3 | 98 | — | — | — | — | — | — | — | — | — | — | — |
| OSWorld i % success | 85 | 70.6 | 83.4 | — | 78.7 | — | — | — | — | — | — | — | 86.1 | 73.1 | — | — | — | — | — |
| BrowseComp i % accuracy | — | 90.8 | 84.3 | — | 90.1 | 90.4 | 85.9 | — | — | — | — | — | — | 83.2 | 91.2 | — | — | — | — |
Math
Competition and research-level mathematical reasoning, increasingly reported on uncontaminated/post-cutoff problem sets.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HMMT i % accuracy (pass@1) | — | — | — | — | — | — | — | — | — | — | — | 97.1 | — | 92.7 | — | — | — | — | — |
| MATH i % accuracy | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 61.2 | — | 89 |
Science & reasoning
Expert-level, Google-proof reasoning across the sciences and broad academia — the benchmarks vendors point to when claiming 'PhD-level' or 'frontier' reasoning.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPQA Diamond i % accuracy | 92.6 | 84.1 | 93.6 | — | 93.6 | — | 94.3 | 90.1 | — | 90.1 | 90.1 | 92.4 | 92.6 | 90.5 | 93.5 | 91.2 | 69.8 | — | 42.4 |
| Humanity's Last Exam i % accuracy | — | 64.7 | 57.9 | 64.5 | 57.2 | — | 44.4 | — | — | 37.7 | — | 41.4 | 43.6 | 54 | 43.5 | 40.5 | — | — | — |
| ARC-AGI-2 i % solved | — | 90.4 | — | — | 85 | — | 77.1 | — | — | — | — | — | — | — | — | — | — | — | — |
General knowledge
Broad multi-subject factual and reasoning coverage; the classic 'how much does it know' bucket, now reported via the harder Pro variant since base MMLU is saturated.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MMLU-Pro i % accuracy | — | — | — | — | — | — | — | — | — | 87.5 | 87.5 | — | — | — | — | — | 80.5 | — | 67.5 |
Multimodal
Vision + language understanding and visual reasoning — how well the model interprets images, diagrams, charts and figures.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MMMU i % accuracy | — | — | — | — | 83.2 | — | 80.5 | — | — | — | — | — | — | 79.4 | — | — | 73.4 | — | 64.9 |
Long context
Retrieval and reasoning quality as context length grows into the hundreds-of-thousands / millions of tokens — beyond simple needle-in-a-haystack.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MRCR i % accuracy | — | — | — | — | 74 | — | 84.9 | — | — | — | — | — | — | — | — | — | — | — | — |
Human preference
Aggregate real-user preference from blind head-to-head comparisons — the closest thing to a 'do people actually like the answers' metric, and the one number vendors love to top.
| Benchmark | Fable 5 | Opus 5 reported | Opus 4.8 | Mythos 5 | GPT-5.5 | GPT-5.6 Sol reported | Gemini 3.1 Pro | Grok 4.3 | Grok 4.6 reported | DeepSeek-V4-Pro open | DeepSeek-V4-Pro-Max reported | Qwen3.7-Max | Qwen3.8-Max reported | Kimi K2.6 open | Kimi K3 reported | GLM-5.2 open | Llama 4 Maverick open | Mistral Medium 3.5 open | Gemma 3 27B open |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LMArena Elo i Elo | — | — | — | 1458 | 1474 | — | — | — | — | — | — | — | — | 1466 | — | — | 1417 | — | 1338 |
Showing 13 flagship models across the headline benchmarks; 53 models and
84 benchmarks tracked in total from 60 primary & aggregator sources.
Claude Mythos 5 numbers are limited (access-restricted preview).
Numbers are published facts, restated with a per-cell source link; vendor benchmark charts are linked to their source, not rehosted.
Sources include: 9to5Google (Gemini 3 Flash launch coverage) · AIFire (citing OpenAI GPT-5.2 release) · Artificial Analysis · BenchLM.ai (citing LMArena) · BinaryVerse AI (xAI official figures) · BuildFastWithAI (citing OpenAI GPT-5.5 launch) · Caylent · Codersera (reporting Moonshot's figures) · DataCamp · DataCamp (reproducing Meta's Llama 4 launch chart) · DeepSeek-AI (arXiv 2512.02556) · DeepSeek-AI (Hugging Face model card) · Google (official Gemini 2.5 launch blog) · Google (official Gemini 3 Flash launch blog) · Google (official Gemini 3 launch blog) · Google DeepMind (Gemini 3.1 Pro model card) · +36 more.
Vendor-reported benchmarks
Numbers as claimed by the vendor on their own model/system card — not independently verified and often measured with a favourable harness. We track each vendor's claims over time and link to the source; cross-check against the cited matrix above.
| Frontier-Bench v0.1 · High effort | 43.3% |
| ARC-AGI-3 · High effort | 30.2% |
| SWE-bench Verified | 97.0% |
| SWE-bench Pro | 79.2% |
| GPQA Diamond | 84.1% |
| Terminal-Bench 2.0 | 68.9% |
| OSWorld 2.0 | 70.6% |
| Humanity's Last Exam · with tools | 64.7% |
| BrowseComp · agentic search | 90.8% |
| GDPval-AA v2 · knowledge work | 1861 Elo |
| ARC-AGI-1 · Max reasoning effort | 97.5% |
| ARC-AGI-2 · Semi-Private, Max reasoning effort | 90.4% |
| Terminal-Bench 2.1 · base Sol; Sol Ultra scored 91.9% | 88.8% |
| Agents' Last Exam · max reasoning setting | 53.6 score |
| ARC-AGI-3 · with retained reasoning and compaction enabled | 38.3% |
| ARC-AGI-3 · official harness (standard) | 13.3% |
| DeepSWE | 72.7% |
| BrowseComp · base Sol; Sol Ultra scored 92.2% | 90.4% |
| ExploitBench | 73.5% |
| Capture-the-Flag · OpenAI's curated internal task set; tool-enabled harness | 96.7% |
| HealthBench Professional | 60.5 length-adjusted score |
| Artificial Analysis Intelligence Index · nine-benchmark aggregate | 61 composite |
| CursorBench v3.2 · high thinking effort | 69.9% |
| DeepSWE v1.1 · high thinking effort | 65.9% |
| FrontierCode v1.1 Extended · extended split | 61.3% |
| APEX-Agents · high thinking effort | 57.5% |
| APEX-SWE · high thinking effort | 56.4% |
| Terminal-Bench v3.0 · Grok Build harness | 26% |
| GDPval-AA v2 | 1753 Elo |
| AA-Briefcase · long-horizon analyst work | 1577 Elo |
| Harvey Legal Agent Benchmark · Vals scoring | 15.8% |
| LiveCodeBench · xhigh thinking effort | 88.4% |
| Next.js Evals | 92% |
| GPQA Diamond · Graduate-level science | 94.3% |
| ARC-AGI-2 · Abstract reasoning | 77.1% |
| SWE-Bench Verified · Agentic code evaluation | 80.6% |
| Humanity's Last Exam · No tools | 44.4% |
| LiveCodeBench Pro · Competitive programming | 2887 Elo |
| Terminal-Bench 2.0 · Terminus-2 harness | 68.5% |
| SciCode · Scientific research coding | 59.0% |
| BrowseComp · Agentic task | 85.9% |
| MMMLU · Multilingual | 92.6% |
| MCP Atlas · Tool use coordination | 69.2% |
| APEX-Agents · Agentic multi-step tasks | 33.5% |
| GPQA Diamond · Vendor-run; independent replication 92.6% vs claimed 92.7% | 92.6% |
| SWE-bench Pro · Alibaba's own corrected task set with Claude Code harness | 67.7% |
| Terminal-Bench 2.1 · 5-hour timeout; official leaderboard caps at 83.8% | 86.6% |
| PaperBench · 28-point jump from Qwen3.7-Max; reproduces research paper results | 93.0% |
| IFBench · Instruction following benchmark | 82.8% |
| OSWorld-Verified · Desktop OS interaction; computer-use agentic workflows | 86.1% |
| Humanity's Last Exam · Weakest showing among four flagship comparisons | 43.6% |
| MathVision · Multimodal benchmark; vision-based math tasks | 95.2% |
| SWE-bench Verified · Pass@1 | 80.6% |
| GPQA Diamond · Pass@1 | 90.1% |
| MMLU-Pro · EM | 87.5% |
| LiveCodeBench | 93.5 pass@1 |
| MMLU Pro · 5-shot, CoT | 80.5 % |
| GPQA Diamond · 0-shot, CoT, multiple generations averaged | 69.8 % |
| LiveCodeBench · 10.01.2024-02.01.2025, multiple generations averaged | 43.4 % |
| MMMU · Image reasoning | 73.4 % |
| MathVista · Image reasoning | 73.7 % |
| ChartQA · Image understanding | 90.0 % |
| DocVQA · Image understanding, test set | 94.4 % |
| Multilingual MMLU · Multilingual | 84.6 % |
| MTOB Half Book · Long context, eng->kgv/kgv->eng | 54.0 / 46.4 % |
| MTOB Full Book · Long context, eng->kgv/kgv->eng | 50.8 / 46.7 % |
| GPQA Diamond | 93.5% |
| SWE-Bench Verified | 76.8% |
| Terminal-Bench 2.1 | 88.3 |
| FrontierSWE | 81.2 pass@1 |
| Program Bench · raw pass rate | 77.8% |
| SWE Marathon | 42.0 |
| DeepSWE · Kimi Code harness | 67.5 |
| BrowseComp · 300K context compaction | 91.2 |
| Humanity's Last Exam (HLE-Full) · text / text+tools | 43.5 / 56.0 |
| CritPt | 23.4 |
| AA-LCR · Long-context retrieval | 74.7 |
| SWE-bench Pro | 62.1% |
| Terminal-Bench 2.1 | 81.0% |
| GPQA Diamond | 91.2% |
| AIME 2026 | 99.2% |
| HLE · with tools | 54.7% |
Model changes
New releases, deprecations, and benchmark score moves we’ve recorded — newest first.
-
Meta reported benchmarks updated
Llama 4 Maverick: 10 benchmark claims (via web search)
-
OpenAI reported benchmarks updated
GPT-5.6 Sol: 9 benchmark claims (via web search)
-
Moonshot AI reported benchmarks updated
Kimi K3: 11 benchmark claims (via web search)
-
Zhipu AI (Z.ai) reported benchmarks updated
GLM-5.2: 5 benchmark claims (via web search)
-
Alibaba reported benchmarks updated
Qwen3.8-Max: 8 benchmark claims (via web search)
-
DeepSeek reported benchmarks updated
DeepSeek-V4-Pro-Max: 4 benchmark claims (via web search)
-
xAI reported benchmarks updated
Grok 4.6: 12 benchmark claims (via web search)
-
Google reported benchmarks updated
Gemini 3.1 Pro: 11 benchmark claims (via web search)
-
Anthropic reported benchmarks updated
Claude Opus 5: 12 benchmark claims (via web search)
-
LMArena (text) scores changed
Arena Elo (text, overall) updated.
-
NVIDIA Nemotron models changed
NVIDIA-Nemotron-Parse-2.0 license changed from 'other' to 'openmdw-1.1'
-
Google Gemma models changed
Removed six T5-Gemma model variants with gemma license dated 2025-06-19.
-
LMArena (text) scores changed
Arena Elo (text, overall) updated.
-
GPQA Diamond (Epoch) scores changed
GPQA Diamond accuracy (0–1) updated.
-
deepseek-coder-1.3b-base: Epoch Capabilities Index (ECI) ↑ 62.02 → 62.28
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
Cerebras-GPT-13B: Epoch Capabilities Index (ECI) ↑ 81.66 → 81.79
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
starcoder2-3b: Epoch Capabilities Index (ECI) ↑ 87.33 → 87.44
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
deepseek-coder-6.7b-base: Epoch Capabilities Index (ECI) ↑ 88.2 → 88.31
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
dolly-v2-12b: Epoch Capabilities Index (ECI) ↑ 88.25 → 88.36
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
Baichuan-7B: Epoch Capabilities Index (ECI) ↑ 89.05 → 89.15
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
phi-1_5: Epoch Capabilities Index (ECI) ↑ 90.19 → 90.29
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
xgen-7b-8k-base: Epoch Capabilities Index (ECI) ↑ 92.09 → 92.19
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
starcoder2-7b: Epoch Capabilities Index (ECI) ↑ 92.24 → 92.33
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
gemma-2b: Epoch Capabilities Index (ECI) ↑ 92.91 → 93
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.
-
mpt-7b: Epoch Capabilities Index (ECI) ↑ 93.31 → 93.4
Same model id, score moved on Epoch Capabilities Index — a silent re-evaluation or model swap.