Benchmark

OpenAI — reported benchmarks

OpenAI · source page ↗ · last checked Aug 22, 2026, 6:02 PM

Reported benchmarks · GPT-5.6 Sol

captured Aug 22, 2026, 6:02 PM
BenchmarkScore
Terminal-Bench 2.188.8% · base Sol; Sol Ultra scored 91.9%
Agents' Last Exam53.6 score · max reasoning setting
ARC-AGI-338.3% · with retained reasoning and compaction enabled
ARC-AGI-313.3% · official harness (standard)
DeepSWE72.7%
BrowseComp90.4% · base Sol; Sol Ultra scored 92.2%
ExploitBench73.5%
Capture-the-Flag96.7% · OpenAI's curated internal task set; tool-enabled harness
HealthBench Professional60.5 length-adjusted score

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

In the news · OpenAI

Importance-filtered press coverage (Google News) mentioning OpenAI. Headlines link to the original; verify before acting.

Change history