BenchAtlas

BenchAtlas Index · snapshot 2026-09-11

The benchmark landscape, in one instrument panel

Ranked systems
450
clear the coverage gate
Live sources
10
of 10 registered — every number linked to its source

Top of the index

#SystemIndexΔ
1GPT 5.6 Solmax effort85.6·
2GPT 5.6 Solhigh effort84.4·
3GPT 5.6 Solxhigh effort84.3·
4Claude Fable 5max effort82.7·
5GPT 5.6 Solmax effort82.2·
6GPT 5.6 Solmedium effort81.6·
7GPT 5.5 (2026-04-22)xhigh effort81.5·
8Claude Opus 4.7max effort81.4·

Category leaders

as of 2026-09-11
Agentic Coding ×0.4
GPT 5.6 Solmax effort
99.5
1f
99.4
1f
Knowledge & Science ×0.15
GPT 5.6 Solmax effort
99.8
2f
Professional Domains ×0.1
O1 (2024-12-17)high effort
99.5
1f
Reasoning & Math ×0.1
99.5
1f
Long Context & Instruction ×0.05
98.8
1f
Multimodal ×0.05
Claude Fable 5max effort
99.4
1f

Best category percentile among all scored systems (×weight in the Index); “Nf” = benchmark families behind the score.

BenchAtlas Index, top 10

bands are leave-one-family-out ranges
6471778490
1. GPT 5.6 Sol (max effort)
85.6
2. GPT 5.6 Sol (high effort)
84.4
3. GPT 5.6 Sol (xhigh effort)
84.3
4. Claude Fable 5 (max effort)
82.7
5. GPT 5.6 Sol (max effort)
82.2
6. GPT 5.6 Sol (medium effort)
81.6
7. GPT 5.5 (2026-04-22) (xhigh effort)
81.5
8. Claude Opus 4.7 (max effort)
81.4
9. GPT 5.6 Sol (low effort)
80.6
10. Claude Opus 4.7 (max effort)
80.6

A band shows how far a system’s score moves when any single benchmark family is left out — wide bands mean the score leans on few families.

Biggest moves

vs previous snapshot
#193 → #192 1
#192 → #193 1
#256 → #255 1
#255 → #256 1
#389 → #388 1
#388 → #389 1

Recent releases

2026-07-09
proprietary
2026-07-09
proprietary
2026-07-09
license unknown
2026-07-09
license unknown
2026-07-08
proprietary
2026-06-30
proprietary
2026-06-13
open weights
2026-06-12
license unknown

Newly observed results

append-only fact table, newest first
SystemBenchmarkScore
glm-5.2_maxIntelligence Index v4.134.0 points
Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)Intelligence Index v4.134.7 points
gpt-5.6-luna_lowIntelligence Index v4.121.5 points
gpt-5.6-luna_mediumIntelligence Index v4.125.5 points
GLM-5.1 (Reasoning)Intelligence Index v4.126.4 points
claude-4-opus-thinkingAA-LCR69.3%
Claude Sonnet 5 (Adaptive Reasoning, High Effort)Intelligence Index v4.132.0 points
gpt-5.6-luna_highIntelligence Index v4.132.4 points

Source freshness

ARC Prize Leaderboard2h ago
Artificial Analysis3h ago
Epoch AI Benchmarking Hub8d ago
LiveBench<1h ago
LiveCodeBench Leaderboard6d ago
LMArena Leaderboard Dataset23h ago
Scale Labs21h ago
SWE-bench Leaderboard34d ago
Terminal-Bench Leaderboard10d ago
Vals AI22h ago

Health derived from ingestion runs: a streak of 3+ failures = failing, 1–2 = degraded. “Never run” sources are registered adapters awaiting their first scheduled ingestion.

The BenchAtlas Index is a percentile-based composite across independent benchmark families (methodology). Missing coverage is shown as absence, never as a low score, and every published number traces back to a raw source snapshot.