Benchmarks / Knowledge & Science
Humanity's Last Exam
Frontier-difficulty closed-ended academic questions (CAIS × Scale).
maintainer unknown · 517 observations on record
Versions
1 version
| Version | Tasks |
|---|---|
| Humanity's Last Exam current humanitys-last-exam | — |
Current leaderboard
Humanity's Last Exam · Accuracy · top 25 of 280 systems
| # | System | Accuracy % |
|---|---|---|
| 1 | GPT 5.6 Solmax effort OpenAI | 49.5% |
| 2 | Claude Opus 4.8max effort Anthropic | 48.7% |
| 3 | GPT 5.6 Solxhigh effort OpenAI | 47.3% |
| 4 | Gemini 3.1 Pro Preview Google DeepMind | 47.0% |
| 5 | Gemini 3.1 Pro Previewhigh effort Google DeepMind | 46.4% 44.5–48.4 |
| 6 | Muse Spark 1.1xhigh effort Meta AI | 46.2% |
| 7 | GPT 5.6 Solhigh effort OpenAI | 46.0% |
| 8 | GPT 5.5 (2026-04-22)xhigh effort OpenAI | 45.8% |
| 9 | GPT 5.5 (2026-04-22)high effort OpenAI | 45.0% |
| 10 | GPT 5.4 Pro (2026-03-05) OpenAI | 44.3% 42.4–46.3 |
| 11 | GPT 5.4 (2026-03-05)xhigh effort OpenAI | 43.7% |
| 12 | GPT 5.6 Terramax effort OpenAI | 42.9% |
| 13 | Gemini 3.5 Flashhigh effort Google DeepMind | 42.7% |
| 14 | Grok 4.5high effort xAI | 42.7% |
| 15 | GPT 5.3 Codexxhigh effort OpenAI | 42.5% |
| 16 | GPT 5.5 (2026-04-22)medium effort OpenAI | 42.4% |
| 17 | Claude Opus 4.7max effort Anthropic | 42.3% |
| 18 | GPT 5.6 Solmedium effort OpenAI | 42.2% |
| 19 | GPT 5.6 Terraxhigh effort OpenAI | 41.9% |
| 20 | Gemini 3.5 Flashmedium effort Google DeepMind | 41.3% |
| 21 | Claude Sonnet 5max effort Anthropic | 41.3% |
| 22 | GLM 5.2max effort Z.ai | 41.1% |
| 23 | Muse Spark Meta AI | 40.7% |
| 24 | Qwen3 7 Max Alibaba (Qwen) | 40.5% |
| 25 | Claude Opus 4.6max effort Anthropic | 39.9% |
One row per evaluated system — reasoning-effort variants rank separately. Where a system has attr variants for this version (eval splits, style control), only the canonical variant is ranked: split=semi_private, style_control=true, or the unsplit run.
Source agreement
where evaluating organizations agree — and don't
Humanity's Last ExamArtificial Analysis vs Scale Labs
overlap 18 systemsrank correlation 0.96mean |Δ| 3.49max |Δ| 9.9
most contested: GPT 5.2 (2025-12-11) (+9.9) · Gemini 3.1 Flash Lite Preview (+8.56) · GPT 5.4 (2026-03-05) (+7.46)
Possible reasons: run setting 'implementation' differs (artificial-analysis vs scale); run setting 'contamination' differs (— vs Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.).
