BenchAtlas

Benchmarks / Knowledge & Science

Humanity's Last Exam

Frontier-difficulty closed-ended academic questions (CAIS × Scale).

maintainer unknown · 517 observations on record

Versions

1 version
VersionTasks
Humanity's Last Exam current
humanitys-last-exam

Current leaderboard

Humanity's Last Exam · Accuracy · top 25 of 280 systems
4043464952
GPT 5.6 Sol
49.5
Claude Opus 4.8
48.7
GPT 5.6 Sol
47.3
Gemini 3.1 Pro Preview
47.0
Gemini 3.1 Pro Preview
46.4
Muse Spark 1.1
46.2
GPT 5.6 Sol
46.0
GPT 5.5 (2026-04-22)
45.8
GPT 5.5 (2026-04-22)
45.0
GPT 5.4 Pro (2026-03-05)
44.3
#SystemAccuracy %
1GPT 5.6 Solmax effort
OpenAI
49.5%
2Claude Opus 4.8max effort
Anthropic
48.7%
3GPT 5.6 Solxhigh effort
OpenAI
47.3%
4Gemini 3.1 Pro Preview
Google DeepMind
47.0%
5Gemini 3.1 Pro Previewhigh effort
Google DeepMind
46.4%
44.548.4
6Muse Spark 1.1xhigh effort
Meta AI
46.2%
7GPT 5.6 Solhigh effort
OpenAI
46.0%
8GPT 5.5 (2026-04-22)xhigh effort
OpenAI
45.8%
9GPT 5.5 (2026-04-22)high effort
OpenAI
45.0%
10GPT 5.4 Pro (2026-03-05)
OpenAI
44.3%
42.446.3
11GPT 5.4 (2026-03-05)xhigh effort
OpenAI
43.7%
12GPT 5.6 Terramax effort
OpenAI
42.9%
13Gemini 3.5 Flashhigh effort
Google DeepMind
42.7%
14Grok 4.5high effort
xAI
42.7%
15GPT 5.3 Codexxhigh effort
OpenAI
42.5%
16GPT 5.5 (2026-04-22)medium effort
OpenAI
42.4%
17Claude Opus 4.7max effort
Anthropic
42.3%
18GPT 5.6 Solmedium effort
OpenAI
42.2%
19GPT 5.6 Terraxhigh effort
OpenAI
41.9%
20Gemini 3.5 Flashmedium effort
Google DeepMind
41.3%
21Claude Sonnet 5max effort
Anthropic
41.3%
22GLM 5.2max effort
Z.ai
41.1%
23Muse Spark
Meta AI
40.7%
24Qwen3 7 Max
Alibaba (Qwen)
40.5%
25Claude Opus 4.6max effort
Anthropic
39.9%

One row per evaluated system — reasoning-effort variants rank separately. Where a system has attr variants for this version (eval splits, style control), only the canonical variant is ranked: split=semi_private, style_control=true, or the unsplit run.

Source agreement

where evaluating organizations agree — and don't
Humanity's Last ExamArtificial Analysis vs Scale Labs
overlap 18 systemsrank correlation 0.96mean |Δ| 3.49max |Δ| 9.9
most contested: GPT 5.2 (2025-12-11) (+9.9) · Gemini 3.1 Flash Lite Preview (+8.56) · GPT 5.4 (2026-03-05) (+7.46)

Possible reasons: run setting 'implementation' differs (artificial-analysis vs scale); run setting 'contamination' differs (— vs Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.).