BenchAtlas

Benchmarks / Knowledge & Science

MMLU-Pro

Harder, reasoning-focused successor to MMLU.

maintainer unknown · 260 observations on record

Versions

1 version
VersionTasks
MMLU-Pro current
mmlu-pro

Current leaderboard

MMLU-Pro · Accuracy · top 25 of 259 systems
8789919294
Claude Fable 5
91.5
Gemini 3.1 Pro Preview
91.0
Gemini 3 Pro Preview
90.1
Gemini 3 Pro Preview
90.1
Claude Opus 4.7
89.9
Gemini 3 Pro Preview
89.8
Claude Opus 4.8
89.6
Gemini 3.5 Flash
89.5
Claude Opus 4.5
89.5
Gemini 3 Pro Preview
89.5
#SystemAccuracy %
1Claude Fable 5max effort
Anthropic
91.5%
2Gemini 3.1 Pro Previewhigh effort
Google DeepMind
91.0%
3Gemini 3 Pro Previewhigh effort
Google DeepMind
90.1%
4Gemini 3 Pro Preview
Google DeepMind
90.1%
5Claude Opus 4.7max effort
Anthropic
89.9%
6Gemini 3 Pro Previewhigh effort
Google DeepMind
89.8%
7Claude Opus 4.8max effort
Anthropic
89.6%
8Gemini 3.5 Flashhigh effort
Google DeepMind
89.5%
9Claude Opus 4.5thinking
Anthropic
89.5%
10Gemini 3 Pro Previewlow effort
Google DeepMind
89.5%
11Qwen3 7 Max
Alibaba (Qwen)
89.3%
12Grok 4.5high effort
xAI
89.2%
13Claude Opus 4.6max effort
Anthropic
89.1%
14GPT 5.6 Solmax effort
OpenAI
89.1%
15Gemini 3 Flash Previewthinking
Google DeepMind
89.0%
16Claude Opus 4.5no reasoning
Anthropic
88.9%
17Muse Spark 1.1xhigh effort
Meta AI
88.7%
18Gemini 3 Flash Previewhigh effort
Google DeepMind
88.6%
19Gemini 3 Flash Previewno reasoning
Google DeepMind
88.2%
20GPT 5.5 (2026-04-22)xhigh effort
OpenAI
88.1%
21Claude Opus 4.1 20250805thinking
Anthropic
87.9%
22Qwen3 6 Plus
Alibaba
87.7%
23Kimi K2 6
Moonshot AI
87.6%
24Kimi K2 6
Moonshot AI
87.6%
25Claude Sonnet 5max effort
Anthropic
87.5%

One row per evaluated system — reasoning-effort variants rank separately.

Source agreement

where evaluating organizations agree — and don't
overlap 27 systemsrank correlation 0.96mean |Δ| 2.17max |Δ| 14.67
most contested: Magistral Small 2509 (+14.67) · Magistral Medium 2509 (+12.84) · GPT OSS 20B (+3.16)

Possible reasons: run setting 'implementation' differs (artificial-analysis vs vals-ai).