BenchAtlas Index · snapshot 2026-09-11
The benchmark landscape, in one instrument panel
Ranked systems
450
clear the coverage gate
Live sources
10
of 10 registered — every number linked to its source
Top of the index
| # | System | Index | Δ | |
|---|---|---|---|---|
| 1 | GPT 5.6 Solmax effort | OpenAI | 85.6 | · |
| 2 | GPT 5.6 Solhigh effort | OpenAI | 84.4 | · |
| 3 | GPT 5.6 Solxhigh effort | OpenAI | 84.3 | · |
| 4 | Claude Fable 5max effort | Anthropic | 82.7 | · |
| 5 | GPT 5.6 Solmax effort | OpenAI | 82.2 | · |
| 6 | GPT 5.6 Solmedium effort | OpenAI | 81.6 | · |
| 7 | GPT 5.5 (2026-04-22)xhigh effort | OpenAI | 81.5 | · |
| 8 | Claude Opus 4.7max effort | Anthropic | 81.4 | · |
Category leaders
as of 2026-09-11
Agentic Coding ×0.4
GPT 5.6 Solmax effort
99.5
1f
Coding ×0.15
99.4
1f
Knowledge & Science ×0.15
GPT 5.6 Solmax effort
99.8
2f
Professional Domains ×0.1
O1 (2024-12-17)high effort
99.5
1f
Reasoning & Math ×0.1
GPT 5.5 Pro Pre Releasehigh effort
99.5
1f
Long Context & Instruction ×0.05
Gemini 3.1 Pro Previewhigh effort
98.8
1f
Multimodal ×0.05
Claude Fable 5max effort
99.4
1f
Best category percentile among all scored systems (×weight in the Index); “Nf” = benchmark families behind the score.
BenchAtlas Index, top 10
bands are leave-one-family-out ranges
A band shows how far a system’s score moves when any single benchmark family is left out — wide bands mean the score leans on few families.
Biggest moves
vs previous snapshot
#193 → #192 ▲1
#192 → #193 ▼1
#256 → #255 ▲1
#255 → #256 ▼1
#389 → #388 ▲1
#388 → #389 ▼1
Recent releases
Muse Spark 1.1
Meta AI
2026-07-09
proprietaryGPT 5.6 Sol
OpenAI
2026-07-09
proprietaryGPT 5.6 Luna
OpenAI
2026-07-09
license unknownGPT 5.6 Terra
OpenAI
2026-07-09
license unknownGrok 4.5
xAI
2026-07-08
proprietaryClaude Sonnet 5
Anthropic
2026-06-30
proprietaryGLM 5.2
Z.ai
2026-06-13
open weightsKimi K2 7 Code
Moonshot
2026-06-12
license unknownNewly observed results
append-only fact table, newest first
| System | Benchmark | Score | |
|---|---|---|---|
| glm-5.2_max | Intelligence Index v4.1 | 34.0 points | Artificial Analysis |
| Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Intelligence Index v4.1 | 34.7 points | Artificial Analysis |
| gpt-5.6-luna_low | Intelligence Index v4.1 | 21.5 points | Artificial Analysis |
| gpt-5.6-luna_medium | Intelligence Index v4.1 | 25.5 points | Artificial Analysis |
| GLM-5.1 (Reasoning) | Intelligence Index v4.1 | 26.4 points | Artificial Analysis |
| claude-4-opus-thinking | AA-LCR | 69.3% | Artificial Analysis |
| Claude Sonnet 5 (Adaptive Reasoning, High Effort) | Intelligence Index v4.1 | 32.0 points | Artificial Analysis |
| gpt-5.6-luna_high | Intelligence Index v4.1 | 32.4 points | Artificial Analysis |
Source freshness
ARC Prize Leaderboard2h ago
Artificial Analysis3h ago
Epoch AI Benchmarking Hub8d ago
LiveBench<1h ago
LiveCodeBench Leaderboard6d ago
LMArena Leaderboard Dataset23h ago
Scale Labs21h ago
SWE-bench Leaderboard34d ago
Terminal-Bench Leaderboard10d ago
Vals AI22h ago
Health derived from ingestion runs: a streak of 3+ failures = failing, 1–2 = degraded. “Never run” sources are registered adapters awaiting their first scheduled ingestion.
The BenchAtlas Index is a percentile-based composite across independent benchmark families (methodology). Missing coverage is shown as absence, never as a low score, and every published number traces back to a raw source snapshot.
