Benchmarks / Knowledge & Science
GPQA Diamond
Graduate-level, Google-proof science questions (diamond subset).
maintainer unknown · 631 observations on record
Versions
1 version
| Version | Tasks |
|---|---|
| GPQA Diamond current gpqa-diamond | 198 |
Current leaderboard
GPQA Diamond · Accuracy · top 25 of 547 systems
| # | System | Accuracy % |
|---|---|---|
| 1 | Gemini 3.1 Pro Previewhigh effort Google DeepMind | 95.5% |
| 2 | GPT 5.6 Solmax effort OpenAI | 95.2% |
| 3 | GPT 5.4 Pro (2026-03-05)xhigh effort OpenAI | 94.6% |
| 4 | Gemini 3.1 Pro Previewhigh effort Google DeepMind | 94.4% |
| 5 | Gemini 3.1 Pro Preview Google DeepMind | 94.1% |
| 6 | GPT 5.6 Solmax effort OpenAI | 94.1% |
| 7 | GPT 5.5 Pre Releasexhigh effort OpenAI | 94.0% |
| 8 | GPT 5.5 Pro Pre Releasexhigh effort OpenAI | 93.9% |
| 9 | GPT 5.5 (2026-04-22)xhigh effort OpenAI | 93.5% |
| 10 | Grok 4.5high effort xAI | 93.4% |
| 11 | GPT 5.6 Terramax effort OpenAI | 93.3% |
| 12 | GPT 5.4 (2026-03-05)xhigh effort OpenAI | 93.3% |
| 13 | GPT 5.5 (2026-04-22)high effort OpenAI | 93.2% |
| 14 | Claude Fable 5max effort Anthropic | 93.2% |
| 15 | GPT 5.5 (2026-04-22)xhigh effort OpenAI | 93.2% |
| 16 | GPT 5.6 Solxhigh effort OpenAI | 93.1% |
| 17 | Grok 4.5high effort xAI | 92.9% |
| 18 | Minimax M3 MiniMax | 92.9% |
| 19 | Gemini 3.5 Flashhigh effort Google DeepMind | 92.8% |
| 20 | GPT 5.6 Solhigh effort OpenAI | 92.8% |
| 21 | Gemini 3.5 Flashhigh effort Google DeepMind | 92.7% |
| 22 | Minimax M3 MiniMax | 92.7% |
| 23 | Gemini 3 Pro Preview Google DeepMind | 92.6% |
| 24 | GPT 5.6 Solmedium effort OpenAI | 92.6% |
| 25 | GPT 5.5 (2026-04-22)medium effort OpenAI | 92.6% |
One row per evaluated system — reasoning-effort variants rank separately. Where a system has attr variants for this version (eval splits, style control), only the canonical variant is ranked: split=semi_private, style_control=true, or the unsplit run.
Source agreement
where evaluating organizations agree — and don't
GPQA DiamondArtificial Analysis vs Epoch AI Benchmarking Hub
overlap 61 systemsrank correlation 0.96mean |Δ| 2.56max |Δ| 26.18
most contested: O1 Preview (2024-09-12) (+26.18) · Magistral Small 2509 (+18.7) · Deepseek R1 Distill Llama 70B (-15.54)
Possible reasons: run setting 'implementation' differs (artificial-analysis vs —).
GPQA DiamondArtificial Analysis vs Vals AI
overlap 61 systemsrank correlation 0.95mean |Δ| 3.17max |Δ| 39.95
most contested: Mistral Medium 3.5 (+39.95) · Mimo V2 Flash (+25.26) · Magistral Medium 2509 (+11.53)
Possible reasons: run setting 'implementation' differs (artificial-analysis vs vals-ai).
GPQA DiamondEpoch AI Benchmarking Hub vs Vals AI
overlap 68 systemsrank correlation 0.95mean |Δ| 2.51max |Δ| 10.73
Possible reasons: run setting 'implementation' differs (— vs vals-ai).
