Benchmark

Anthropic — reported benchmarks

Anthropic · source page ↗ · last checked Aug 22, 2026, 12:02 PM

Reported benchmarks · Claude Opus 5

captured Aug 22, 2026, 12:02 PM
BenchmarkScore
Frontier-Bench v0.143.3% · High effort
ARC-AGI-330.2% · High effort
SWE-bench Verified97.0%
SWE-bench Pro79.2%
GPQA Diamond84.1%
Terminal-Bench 2.068.9%
OSWorld 2.070.6%
Humanity's Last Exam64.7% · with tools
BrowseComp90.8% · agentic search
GDPval-AA v21861 Elo · knowledge work
ARC-AGI-197.5% · Max reasoning effort
ARC-AGI-290.4% · Semi-Private, Max reasoning effort

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

In the news · Anthropic

Importance-filtered press coverage (Google News) mentioning Anthropic. Headlines link to the original; verify before acting.

Change history