BenchAtlas

BenchAtlas Index · 2026-09-11

Model rankings

Percentile-based composite across independent benchmark families. How it’s computed — coverage and uncertainty shown for every entry; models with too little evidence are never given a fake low score.

#SystemIndex
1 ·GPT 5.6 Solmax effort85.6compare
4 ·Claude Fable 5max effort82.7compare
7 ·GPT 5.5 (2026-04-22)xhigh effort81.5compare
8 ·Claude Opus 4.7max effort81.4compare
12 ·Claude Opus 4.8max effort80.3compare
14 ·GPT 5.6 Terraxhigh effort79.9compare
17 ·Gemini 3.1 Pro Preview79.0compare
19 ·Gemini 3.5 Flashmedium effort78.7compare
21 ·Claude Opus 4.6max effort78.5compare
23 ·Grok 4.5high effort78.0compare
25 ·GPT 5.5 Pre Releasexhigh effort76.7compare
26 ·GPT 5.2 (2025-12-11)xhigh effort76.6compare
27 ·GPT 5.3 Codexxhigh effort76.5compare
28 ·GPT 5.4 (2026-03-05)xhigh effort76.5compare
29 ·GLM 5.2max effort76.3compare
30 ·Claude Sonnet 5max effort76.0compare
33 ·Muse Spark 1.1xhigh effort75.6compare
35 ·GPT 5.6 Lunamax effort75.5compare
38 ·GPT 5.4 Pro (2026-03-05)75.0compare
44 ·Claude Opus 4.5thinking74.3compare
48 ·Gemini 3 Pro Previewhigh effort73.2compare
50 ·GPT 5.2 Codexxhigh effort73.2compare
51 ·Gemini 3 Flash Previewthinking73.1compare
52 ·GPT 5 Codexhigh effort73.0compare
53 ·Kimi K2 673.0compare
60 ·Deepseek V4 Promax effort72.1compare
66 ·Muse Spark71.2compare
70 ·Qwen3 7 Max70.5compare
79 ·Claude Sonnet 4.5 20250929thinking68.9compare
81 ·GPT 5 Pro (2025-10-06)68.6compare
82 ·Mimo V2 5 Pro68.6compare
89 ·Claude Opus 4.5 20251101high effort67.5compare
92 ·GPT 5 (2025-08-07)high effort67.3compare
97 ·Grok 467.1compare
99 ·Mimo V2 Pro66.4compare
100 ·Claude Sonnet 4.6max effort66.3compare
101 ·GLM 5thinking66.1compare
107 ·Claude Opus 4.1 20250805thinking65.7compare
115 ·Qwen3 7 Plus64.7compare
116 ·Minimax M364.5compare
120 ·Grok 4.20 0309 Reasoning64.2compare
121 ·Deepseek V4 Flashmax effort63.8compare
123 ·Qwen3 6 Max Preview63.6compare
127 ·Minimax M2 563.5compare
129 ·Gemini 2.5 Pro Exp 03.2563.4compare
131 ·GPT 5.1 Codexhigh effort63.1compare
134 ·GLM 5.162.6compare
151 ·Mimo V2 Flashthinking60.5compare
152 ·O1 (2024-12-17)high effort60.5compare
154 ·O3 (2025-04-16)high effort60.4compare
155 ·Deepseek R1 052860.2compare
158 ·Minimax M2 760.0compare
160 ·Kimi K2 Thinking59.9compare
161 ·Qwen3 5 Plusthinking59.9compare
167 ·Gemini 2.5 Pro Preview 06.0559.3compare
169 ·GPT 5.1 Codex Minihigh effort59.0compare
170 ·Grok 4.1 Fast Reasoning58.8compare
172 ·Deepseek V3 2thinking58.7compare
174 ·Qwen3 5 Flash58.4compare
175 ·Grok 4.3high effort58.1compare
177 ·Kimi K2 5thinking57.7compare
178 ·Qwen3 6 27Bthinking57.6compare
181 ·Grok 4 Fast Reasoning57.1compare
182 ·Claude Haiku 4.5 20251001thinking57.0compare
183 ·Qwen3 5 397B A17bthinking57.0compare
184 ·Qwen3 5 27Bthinking56.9compare
188 ·GLM 5v Turbothinking56.4compare
189 ·Grok 4 070956.3compare
190 ·Claude 4.5 Sonnetthinking56.3compare
191 ·GLM 4.7thinking56.3compare
192 1Claude Sonnet 4 20250514thinking56.1compare
196 ·Nemotron 3 Ultra 550B A55bthinking55.5compare
197 ·GPT 5.4 Mini (2026-03-17)high effort55.4compare
199 ·Mimo V2 Omni55.0compare
201 ·Step 3.5 Flash54.7compare
203 ·Minimax M2 154.5compare
205 ·Gemini 2.5 Pro (2025-06-17)54.4compare
207 ·GPT 5.1 (2025-11-13)high effort54.4compare
210 ·Mimo V2 553.5compare
214 ·GPT 5.4 Nano (2026-03-17)high effort53.2compare
218 ·Claude 4 Opusthinking52.8compare
220 ·O1 Preview (2024-09-12)52.4compare
226 ·Qwen3 5 122B A10bthinking51.6compare
229 ·Deepseek V3 2 Expthinking51.3compare
231 ·Kimi K2 7 Code50.9compare
239 ·Gemma 4 31Bthinking49.8compare
240 ·Grok 4 Fastthinking49.8compare
242 ·O4 Mini (2025-04-16)high effort49.8compare
243 ·Qwen3 6 Plus49.6compare
248 ·Mistral Medium 3.548.9compare
249 ·Moonshotai/kimi K2 Instruct48.9compare
250 ·GPT 5 Mini (2025-08-07)high effort48.6compare
257 ·Minimax M247.6compare
258 ·Deepseek V3 1thinking47.5compare
265 ·GLM 4.6thinking46.5compare
268 ·Claude Opus 4thinking45.8compare
270 ·Deepseek V3 1 Terminusthinking45.7compare
273 ·Qwen3 5 35B A3bthinking45.5compare
277 ·Mistral Large 251244.5compare
279 ·Gemma 4 26B A4bthinking44.2compare
280 ·Claude Opus 4 2025051443.7compare
281 ·Kimi K2thinking43.6compare
282 ·Claude 3.7 Sonnetthinking43.5compare
285 ·GLM 4.543.3compare
288 ·Gemini 3.1 Flash Lite Preview42.8compare
291 ·Claude 4 Sonnetthinking42.7compare
293 ·Qwen3 Max42.6compare
294 ·Grok Code Fast 142.5compare
295 ·Qwen3 Max Preview42.2compare
302 ·Qwen3 235B A22b 2507thinking41.5compare
303 ·Gemini 2.5 Flashthinking41.4compare
304 ·Nvidia Nemotron 3 Super 120B A12bthinking41.3compare
306 ·Claude Sonnet 4thinking40.7compare
307 ·Trinity Large Thinking40.1compare
309 ·GPT OSS 120Bhigh effort39.6compare
311 ·Grok 339.2compare
314 ·GLM 4.7 Flashthinking38.6compare
315 ·Mercury 238.5compare
320 ·Intellect 337.9compare
321 ·Langston/nim/nvidia/llama 3.3 Nemotron Super 49B V1 42e84561thinking37.8compare
322 ·Qwen3 235B A22b37.6compare
325 ·Qwq 32B37.3compare
326 ·GLM 4.5 Air37.2compare
328 ·GPT 4.1 (2025-04-14)high effort37.0compare
329 ·Llama4 Maverick Instruct Basic36.9compare
332 ·Claude 3.5 Sonnet 2024102236.3compare
333 ·GPT OSS 20Bhigh effort35.9compare
335 ·GPT 4o (2024-05-13)35.8compare
337 ·Magistral Medium 250935.5compare
338 ·Grok Build 0.135.3compare
342 ·Qwen3 Coder 480B A35b Instruct35.2compare
343 ·Command A Plus 05 202635.2compare
344 ·GLM 4 6vthinking35.0compare
345 ·O3 Mini (2025-01-31)high effort34.7compare
348 ·Claude 3.7 Sonnet 2025021934.5compare
351 ·Qwen3 235B A22b Instruct 250734.2compare
353 ·Gemini 2.0 Flash 00133.8compare
357 ·Qwq 32B Preview33.1compare
358 ·Grok 4 Fast Non Reasoning33.1compare
362 ·Claude 3.5 Sonnet 2024062032.2compare
363 ·Deepseek V3 032432.2compare
365 ·O1 Mini (2024-09-12)31.8compare
366 ·GPT 5 Nano (2025-08-07)high effort31.4compare
368 ·Grok 4.1 Fast Non Reasoning31.2compare
373 ·Exaone 4.0 32Bthinking30.5compare
374 ·Qwen3 VL 235B A22b Instruct30.5compare
375 ·GPT 4o (2024-08-06)high effort30.4compare
376 ·Ring Flash 2.030.3compare
378 ·Grok 2 121230.1compare
379 ·Qwen3 Next 80B A3b Instruct30.0compare
383 ·Gemini 1.5 Pro 00229.3compare
387 ·Qwen3 32Bthinking28.9compare
391 ·Qwen3 30B A3bthinking28.4compare
392 ·Ling Flash 2.028.3compare
394 ·GPT 4.1 Mini (2025-04-14)high effort27.4compare
398 ·Llama 3.1 Nemotron Ultra 253B V1thinking27.1compare
399 ·Qwen3 Coder 30B A3b Instruct26.9compare
400 ·GLM 4 5vthinking26.4compare
401 ·Deepseek V326.2compare
402 ·Gemini 2.0 Flash Lite Preview26.1compare
403 ·Deepseek R1 Distill Qwen 14B26.1compare
405 ·Magistral Small 250925.2compare
406 ·Mistral Large 325.0compare
407 ·Olmo 3.1 32B Think24.9compare
408 ·Gemini 2.0 Flash24.6compare
409 ·Olmo 3 32B Think24.6compare
410 ·GPT 4o Mini (2024-07-18)24.3compare
411 ·Gemini 1.5 Flash 00224.3compare
413 ·Meta Llama/llama 4 Scout 17B 16e Instruct23.9compare
414 ·Mistral Small 250323.8compare
417 ·Claude 223.3compare
419 ·Mistral Medium 323.1compare
421 ·Llama 3.2 90B Vision Instruct22.8compare
422 ·GPT 4o (2024-11-20)high effort22.7compare
424 ·Deepseek R1 Distill Llama 70B22.0compare
426 ·Qwen2 5 32B Instruct21.8compare
427 ·Qwen2 5 Coder 7B Instruct21.6compare
428 ·GPT 4 Turbo (2024-04-09)21.0compare
429 ·Jamba 1.5 Mini20.9compare
430 ·Llama 4 Maverick20.9compare
431 ·Qwen2 72B Instruct20.9compare
433 ·Jamba 1.5 Large20.2compare
437 ·Nova Pro19.3compare
439 ·Qwen2 5 72B Instruct19.2compare
440 ·Olmo 3.1 32B Instruct18.7compare
441 ·Qwen2 5 Coder 32B Instruct18.6compare
442 ·Claude 3 Opus18.6compare
443 ·Mistral Medium18.2compare
444 ·Granite 4.1 8B17.9compare
445 ·Phi 417.8compare
446 ·GPT 4.1 Nano (2025-04-14)high effort17.4compare
447 ·Nova Lite14.9compare
448 ·Claude 3 Haiku14.5compare
449 ·Mistral 7B Instruct14.5compare

450systems clear the coverage gate (≥4 benchmark families across ≥3 categories). Systems with less evidence appear on their model pages with full per-benchmark scores — absence of coverage is shown as absence, never as a low score. Reasoning-effort variants rank as separate systems; “best per model” (on by default) collapses them to each model’s best-ranked variant. Coverage dots scale against ~16 index families — hover for exact counts and confidence.