Benchmark

Qwen — reported benchmarks

Alibaba · source page ↗ · last checked Aug 22, 2026, 12:04 PM

Reported benchmarks · Qwen3.8-Max

captured Aug 22, 2026, 12:04 PM
BenchmarkScore
GPQA Diamond92.6% · Vendor-run; independent replication 92.6% vs claimed 92.7%
SWE-bench Pro67.7% · Alibaba's own corrected task set with Claude Code harness
Terminal-Bench 2.186.6% · 5-hour timeout; official leaderboard caps at 83.8%
PaperBench93.0% · 28-point jump from Qwen3.7-Max; reproduces research paper results
IFBench82.8% · Instruction following benchmark
OSWorld-Verified86.1% · Desktop OS interaction; computer-use agentic workflows
Humanity's Last Exam43.6% · Weakest showing among four flagship comparisons
MathVision95.2% · Multimodal benchmark; vision-based math tasks

Vendor-reported via automated web search — not independently verified. See the cited matrix on /models.

Change history