Models

Frontier & open-weight models compared by capability — columns are flagship models, rows are the headline benchmarks that measure each one. Every number is the published figure, restated with a link to the source it came from · as of 2026-06-14. Curated matrix is 70d old — a refresh is due; live leaderboard moves still flow into Model changes below.

We cite the benchmark authorities rather than re-rank them — Epoch AI, LMArena and Artificial Analysis. ChangeRadar's job is the changes: what gets released, repriced, deprecated, or quietly shifts behind a stable model id.

Vendor champions · best flagship per vendor · GPQA Diamond

Google
Gemini 3.1 Pro
94.3GPQA Diamond
Moonshot AI reported
Kimi K3
93.5GPQA Diamond
Source ↗
Alibaba reported
Qwen3.8-Max
92.6GPQA Diamond
Source ↗
Zhipu AI (Z.ai) open
GLM-5.2
91.2GPQA Diamond
OpenAI reported
GPT-5.6 Sol
90.4BrowseComp
Source ↗
DeepSeek reported
DeepSeek-V4-Pro-Max
90.1GPQA Diamond
Source ↗
xAI reported
Grok 4.6
88.4LiveCodeBench
Source ↗
Anthropic reported
Claude Opus 5
84.1GPQA Diamond
Source ↗
Mistral AI open
Mistral Medium 3.5
77.6SWE-bench Verified
Meta open
Llama 4 Maverick
69.8GPQA Diamond

Each vendor's most recent publicly-available flagship, ordered by score (highest first) — Google · OpenAI · Anthropic · xAI highlighted. The weekly market-watch surfaces new releases automatically; one tagged reported is the latest release shown with vendor-reported scores (linked to source) until we independently cite it. Score = GPQA Diamond; every number links to its source.

Compare two models

2 wins · Claude Fable 5 1 wins · Gemini 3.1 Pro 0 ties

Agentic coding

BenchmarkClaude Fable 5Gemini 3.1 ProΔ
SWE-bench Verified i% resolved (pass@1) 95 80.6 +14.4
SWE-bench Pro i% resolved (pass@1) 80.3 54.2 +26.1
3 coverage gaps — only one model reports these
  • Terminal-Bench — Gemini 3.1 Pro 68.5
  • LiveCodeBench — Gemini 3.1 Pro 2887
  • FrontierCode — Claude Fable 5 29.3

Tool use & agents

5 coverage gaps — only one model reports these
  • TAU-bench — Gemini 3.1 Pro 99.3
  • OSWorld — Claude Fable 5 85
  • BrowseComp — Gemini 3.1 Pro 85.9
  • GDPval-AA — Claude Fable 5 1932
  • MCP Atlas — Gemini 3.1 Pro 69.2

Science & reasoning

BenchmarkClaude Fable 5Gemini 3.1 ProΔ
GPQA Diamond i% accuracy 92.6 94.3 -1.7
2 coverage gaps — only one model reports these
  • Humanity's Last Exam — Gemini 3.1 Pro 44.4
  • ARC-AGI-2 — Gemini 3.1 Pro 77.1

General knowledge

1 coverage gap — only one model reports these
  • MMMLU — Gemini 3.1 Pro 92.6

Multimodal

2 coverage gaps — only one model reports these
  • MMMU — Gemini 3.1 Pro 80.5
  • GDP.pdf — Claude Fable 5 29.8

Long context

1 coverage gap — only one model reports these
  • MRCR — Gemini 3.1 Pro 84.9

Shared benchmarks first (with Δ when both report the same scale); one-sided coverage collapses below. Pick any two models — or a champion above — and the URL becomes shareable.

All models · benchmark matrix

Vendor
Capability
19 of 19 columns shown

reported columns (Claude Opus 5, GPT-5.6 Sol, Grok 4.6, Qwen3.8-Max, DeepSeek-V4-Pro-Max, Kimi K3) are auto-discovered by our weekly market-watch from each vendor's own reported numbers — not independently verified, and shown when a vendor ships a model newer than the hand-cited column beside it. Full claim sets are in Vendor-reported benchmarks below.

Agentic coding

Can the model fix real bugs, ship features, and operate a dev environment end-to-end as a coding agent — the single most-watched capability in 2026 vendor launches.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
SWE-bench Verified i % resolved (pass@1) 95 97 88.6 80.6 80.6 80.6 80.4 80.2 76.8 77.6
SWE-bench Pro i % resolved (pass@1) 80.3 79.2 69.2 77.8 58.6 54.2 60.6 67.7 58.6 62.1
Terminal-Bench i % solved (pass@1) 68.9 74.6 88 82.7 88.8 68.5 26 67.9 69.7 86.6 66.7 88.3 81
LiveCodeBench i % pass@1 2887 88.4 93.5 93.5 89.6 43.4 29.7

Tool use & agents

Beyond writing code: can the model select and chain tools, follow policy, drive a computer/browser, and complete long-horizon multi-step tasks.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
TAU-bench i % pass / pass^k 99.3 98
OSWorld i % success 85 70.6 83.4 78.7 86.1 73.1
BrowseComp i % accuracy 90.8 84.3 90.1 90.4 85.9 83.2 91.2

Math

Competition and research-level mathematical reasoning, increasingly reported on uncontaminated/post-cutoff problem sets.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
HMMT i % accuracy (pass@1) 97.1 92.7
MATH i % accuracy 61.2 89

Science & reasoning

Expert-level, Google-proof reasoning across the sciences and broad academia — the benchmarks vendors point to when claiming 'PhD-level' or 'frontier' reasoning.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
GPQA Diamond i % accuracy 92.6 84.1 93.6 93.6 94.3 90.1 90.1 90.1 92.4 92.6 90.5 93.5 91.2 69.8 42.4
Humanity's Last Exam i % accuracy 64.7 57.9 64.5 57.2 44.4 37.7 41.4 43.6 54 43.5 40.5
ARC-AGI-2 i % solved 90.4 85 77.1

General knowledge

Broad multi-subject factual and reasoning coverage; the classic 'how much does it know' bucket, now reported via the harder Pro variant since base MMLU is saturated.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
MMLU-Pro i % accuracy 87.5 87.5 80.5 67.5

Multimodal

Vision + language understanding and visual reasoning — how well the model interprets images, diagrams, charts and figures.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
MMMU i % accuracy 83.2 80.5 79.4 73.4 64.9

Long context

Retrieval and reasoning quality as context length grows into the hundreds-of-thousands / millions of tokens — beyond simple needle-in-a-haystack.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
MRCR i % accuracy 74 84.9

Human preference

Aggregate real-user preference from blind head-to-head comparisons — the closest thing to a 'do people actually like the answers' metric, and the one number vendors love to top.

Benchmark Fable 5 Opus 5 reported Opus 4.8 Mythos 5 GPT-5.5 GPT-5.6 Sol reported Gemini 3.1 Pro Grok 4.3 Grok 4.6 reported DeepSeek-V4-Pro open DeepSeek-V4-Pro-Max reported Qwen3.7-Max Qwen3.8-Max reported Kimi K2.6 open Kimi K3 reported GLM-5.2 open Llama 4 Maverick open Mistral Medium 3.5 open Gemma 3 27B open
LMArena Elo i Elo 1458 1474 1466 1417 1338

Showing 13 flagship models across the headline benchmarks; 53 models and 84 benchmarks tracked in total from 60 primary & aggregator sources. Claude Mythos 5 numbers are limited (access-restricted preview). Numbers are published facts, restated with a per-cell source link; vendor benchmark charts are linked to their source, not rehosted.
Sources include: 9to5Google (Gemini 3 Flash launch coverage) · AIFire (citing OpenAI GPT-5.2 release) · Artificial Analysis · BenchLM.ai (citing LMArena) · BinaryVerse AI (xAI official figures) · BuildFastWithAI (citing OpenAI GPT-5.5 launch) · Caylent · Codersera (reporting Moonshot's figures) · DataCamp · DataCamp (reproducing Meta's Llama 4 launch chart) · DeepSeek-AI (arXiv 2512.02556) · DeepSeek-AI (Hugging Face model card) · Google (official Gemini 2.5 launch blog) · Google (official Gemini 3 Flash launch blog) · Google (official Gemini 3 launch blog) · Google DeepMind (Gemini 3.1 Pro model card) · +36 more.

Vendor-reported benchmarks

Numbers as claimed by the vendor on their own model/system card — not independently verified and often measured with a favourable harness. We track each vendor's claims over time and link to the source; cross-check against the cited matrix above.

Claude Opus 5 Anthropic
Frontier-Bench v0.1 · High effort 43.3%
ARC-AGI-3 · High effort 30.2%
SWE-bench Verified 97.0%
SWE-bench Pro 79.2%
GPQA Diamond 84.1%
Terminal-Bench 2.0 68.9%
OSWorld 2.0 70.6%
Humanity's Last Exam · with tools 64.7%
BrowseComp · agentic search 90.8%
GDPval-AA v2 · knowledge work 1861 Elo
ARC-AGI-1 · Max reasoning effort 97.5%
ARC-AGI-2 · Semi-Private, Max reasoning effort 90.4%
vendor card ↗
GPT-5.6 Sol OpenAI
Terminal-Bench 2.1 · base Sol; Sol Ultra scored 91.9% 88.8%
Agents' Last Exam · max reasoning setting 53.6 score
ARC-AGI-3 · with retained reasoning and compaction enabled 38.3%
ARC-AGI-3 · official harness (standard) 13.3%
DeepSWE 72.7%
BrowseComp · base Sol; Sol Ultra scored 92.2% 90.4%
ExploitBench 73.5%
Capture-the-Flag · OpenAI's curated internal task set; tool-enabled harness 96.7%
HealthBench Professional 60.5 length-adjusted score
vendor card ↗
Grok 4.6 xAI
Artificial Analysis Intelligence Index · nine-benchmark aggregate 61 composite
CursorBench v3.2 · high thinking effort 69.9%
DeepSWE v1.1 · high thinking effort 65.9%
FrontierCode v1.1 Extended · extended split 61.3%
APEX-Agents · high thinking effort 57.5%
APEX-SWE · high thinking effort 56.4%
Terminal-Bench v3.0 · Grok Build harness 26%
GDPval-AA v2 1753 Elo
AA-Briefcase · long-horizon analyst work 1577 Elo
Harvey Legal Agent Benchmark · Vals scoring 15.8%
LiveCodeBench · xhigh thinking effort 88.4%
Next.js Evals 92%
vendor card ↗
Gemini 3.1 Pro Google
GPQA Diamond · Graduate-level science 94.3%
ARC-AGI-2 · Abstract reasoning 77.1%
SWE-Bench Verified · Agentic code evaluation 80.6%
Humanity's Last Exam · No tools 44.4%
LiveCodeBench Pro · Competitive programming 2887 Elo
Terminal-Bench 2.0 · Terminus-2 harness 68.5%
SciCode · Scientific research coding 59.0%
BrowseComp · Agentic task 85.9%
MMMLU · Multilingual 92.6%
MCP Atlas · Tool use coordination 69.2%
APEX-Agents · Agentic multi-step tasks 33.5%
vendor card ↗
Qwen3.8-Max Alibaba
GPQA Diamond · Vendor-run; independent replication 92.6% vs claimed 92.7% 92.6%
SWE-bench Pro · Alibaba's own corrected task set with Claude Code harness 67.7%
Terminal-Bench 2.1 · 5-hour timeout; official leaderboard caps at 83.8% 86.6%
PaperBench · 28-point jump from Qwen3.7-Max; reproduces research paper results 93.0%
IFBench · Instruction following benchmark 82.8%
OSWorld-Verified · Desktop OS interaction; computer-use agentic workflows 86.1%
Humanity's Last Exam · Weakest showing among four flagship comparisons 43.6%
MathVision · Multimodal benchmark; vision-based math tasks 95.2%
vendor card ↗
DeepSeek-V4-Pro-Max DeepSeek
SWE-bench Verified · Pass@1 80.6%
GPQA Diamond · Pass@1 90.1%
MMLU-Pro · EM 87.5%
LiveCodeBench 93.5 pass@1
vendor card ↗
Llama 4 Maverick Meta
MMLU Pro · 5-shot, CoT 80.5 %
GPQA Diamond · 0-shot, CoT, multiple generations averaged 69.8 %
LiveCodeBench · 10.01.2024-02.01.2025, multiple generations averaged 43.4 %
MMMU · Image reasoning 73.4 %
MathVista · Image reasoning 73.7 %
ChartQA · Image understanding 90.0 %
DocVQA · Image understanding, test set 94.4 %
Multilingual MMLU · Multilingual 84.6 %
MTOB Half Book · Long context, eng->kgv/kgv->eng 54.0 / 46.4 %
MTOB Full Book · Long context, eng->kgv/kgv->eng 50.8 / 46.7 %
vendor card ↗
Mistral Medium 3.5 Mistral AI
SWE-Bench Verified 77.6%
τ³-Telecom 91.4%
vendor card ↗
Kimi K3 Moonshot AI
GPQA Diamond 93.5%
SWE-Bench Verified 76.8%
Terminal-Bench 2.1 88.3
FrontierSWE 81.2 pass@1
Program Bench · raw pass rate 77.8%
SWE Marathon 42.0
DeepSWE · Kimi Code harness 67.5
BrowseComp · 300K context compaction 91.2
Humanity's Last Exam (HLE-Full) · text / text+tools 43.5 / 56.0
CritPt 23.4
AA-LCR · Long-context retrieval 74.7
vendor card ↗
GLM-5.2 Zhipu AI (Z.ai)
SWE-bench Pro 62.1%
Terminal-Bench 2.1 81.0%
GPQA Diamond 91.2%
AIME 2026 99.2%
HLE · with tools 54.7%
vendor card ↗

Model changes

New releases, deprecations, and benchmark score moves we’ve recorded — newest first.