BenchAtlas

Benchmarks / Knowledge & Science

GPQA Diamond

Graduate-level, Google-proof science questions (diamond subset).

maintainer unknown · 631 observations on record

Versions

1 version
VersionTasks
GPQA Diamond current
gpqa-diamond
198

Current leaderboard

GPQA Diamond · Accuracy · top 25 of 547 systems
9193959698
Gemini 3.1 Pro Preview
95.5
GPT 5.6 Sol
95.2
GPT 5.4 Pro (2026-03-05)
94.6
Gemini 3.1 Pro Preview
94.4
Gemini 3.1 Pro Preview
94.1
GPT 5.6 Sol
94.1
GPT 5.5 Pre Release
94.0
GPT 5.5 Pro Pre Release
93.9
GPT 5.5 (2026-04-22)
93.5
Grok 4.5
93.4
#SystemAccuracy %
1Gemini 3.1 Pro Previewhigh effort
Google DeepMind
95.5%
2GPT 5.6 Solmax effort
OpenAI
95.2%
3GPT 5.4 Pro (2026-03-05)xhigh effort
OpenAI
94.6%
4Gemini 3.1 Pro Previewhigh effort
Google DeepMind
94.4%
5Gemini 3.1 Pro Preview
Google DeepMind
94.1%
6GPT 5.6 Solmax effort
OpenAI
94.1%
7GPT 5.5 Pre Releasexhigh effort
OpenAI
94.0%
8GPT 5.5 Pro Pre Releasexhigh effort
OpenAI
93.9%
9GPT 5.5 (2026-04-22)xhigh effort
OpenAI
93.5%
10Grok 4.5high effort
xAI
93.4%
11GPT 5.6 Terramax effort
OpenAI
93.3%
12GPT 5.4 (2026-03-05)xhigh effort
OpenAI
93.3%
13GPT 5.5 (2026-04-22)high effort
OpenAI
93.2%
14Claude Fable 5max effort
Anthropic
93.2%
15GPT 5.5 (2026-04-22)xhigh effort
OpenAI
93.2%
16GPT 5.6 Solxhigh effort
OpenAI
93.1%
17Grok 4.5high effort
xAI
92.9%
18Minimax M3
MiniMax
92.9%
19Gemini 3.5 Flashhigh effort
Google DeepMind
92.8%
20GPT 5.6 Solhigh effort
OpenAI
92.8%
21Gemini 3.5 Flashhigh effort
Google DeepMind
92.7%
22Minimax M3
MiniMax
92.7%
23Gemini 3 Pro Preview
Google DeepMind
92.6%
24GPT 5.6 Solmedium effort
OpenAI
92.6%
25GPT 5.5 (2026-04-22)medium effort
OpenAI
92.6%

One row per evaluated system — reasoning-effort variants rank separately. Where a system has attr variants for this version (eval splits, style control), only the canonical variant is ranked: split=semi_private, style_control=true, or the unsplit run.

Source agreement

where evaluating organizations agree — and don't
overlap 61 systemsrank correlation 0.96mean |Δ| 2.56max |Δ| 26.18
most contested: O1 Preview (2024-09-12) (+26.18) · Magistral Small 2509 (+18.7) · Deepseek R1 Distill Llama 70B (-15.54)

Possible reasons: run setting 'implementation' differs (artificial-analysis vs —).

overlap 61 systemsrank correlation 0.95mean |Δ| 3.17max |Δ| 39.95
most contested: Mistral Medium 3.5 (+39.95) · Mimo V2 Flash (+25.26) · Magistral Medium 2509 (+11.53)

Possible reasons: run setting 'implementation' differs (artificial-analysis vs vals-ai).

overlap 68 systemsrank correlation 0.95mean |Δ| 2.51max |Δ| 10.73
most contested: Magistral Small 2509 (-10.73) · GPT OSS 20B (-8.14) · Claude Fable 5 (-7.32)

Possible reasons: run setting 'implementation' differs (— vs vals-ai).