BenchAtlas

Rankings / OpenAI

GPT 5.6 Sol

proprietary

released 2026-07-09

BenchAtlas Index

as of 2026-09-11
85.6
max effort
rank #1
14 families · 5 categories · high
84.4
high effort
rank #2
9 families · 5 categories · high
84.3
xhigh effort
rank #3
13 families · 5 categories · high

Benchmark evidence

174 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort56.6 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort56.2 %independentLiveBench
τ²-Benchmax effort85.1 %independentArtificial Analysis
τ²-Benchxhigh effort84.8 %independentArtificial Analysis
τ²-Benchhigh effort83.3 %independentArtificial Analysis
τ²-Benchmedium effort81.0 %independentArtificial Analysis
τ²-Benchlow effort76.0 %independentArtificial Analysis
τ²-Bench
subset=banking
max effort44.3 %independentArtificial Analysis
τ²-Bench
subset=banking
xhigh effort38.1 %independentArtificial Analysis
τ²-Bench
subset=banking
high effort36.7 %independentArtificial Analysis
τ²-Bench
subset=banking
medium effort36.5 %independentArtificial Analysis
τ²-Bench
subset=banking
low effort29.1 %independentArtificial Analysis
τ²-Bench
subset=banking
no reasoning19.6 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
max effort83.9 %independentLiveBench
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
xhigh effort81.8 %independentLiveBench
SciCodehigh effort57.8 %independentArtificial Analysis
SciCodemedium effort57.4 %independentArtificial Analysis
SciCodexhigh effort57.1 %independentArtificial Analysis
SciCodemax effort57.1 %independentArtificial Analysis
SciCodelow effort56.4 %independentArtificial Analysis
SciCodeno reasoning47.1 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort80.3 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort79.8 %independentLiveBench
external indices
AA Coding Indexxhigh effort78.3 pointsindependentArtificial Analysis
AA Coding Indexmax effort77.4 pointsindependentArtificial Analysis
AA Coding Indexhigh effort77.2 pointsindependentArtificial Analysis
AA Coding Indexmedium effort76.3 pointsindependentArtificial Analysis
AA Coding Indexlow effort69.7 pointsindependentArtificial Analysis
AA Coding Indexno reasoning65.1 pointsindependentArtificial Analysis
Intelligence Index v4.1max effort47.1 pointsindependentArtificial Analysis
Intelligence Index v4.1xhigh effort44.1 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort42.5 pointsindependentArtificial Analysis
Intelligence Index v4.1medium effort39.5 pointsindependentArtificial Analysis
Intelligence Index v4.1low effort33.8 pointsindependentArtificial Analysis
Intelligence Index v4.1no reasoning28.3 pointsindependentArtificial Analysis
ECI162.0 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning162.0 pointsindependentEpoch AI Benchmarking Hub
ECIlow effort162.0 pointsindependentEpoch AI Benchmarking Hub
ECImax effort162.0 pointsindependentEpoch AI Benchmarking Hub
ECI162.0 pointsindependentEpoch AI Benchmarking Hub
ECImedium effort162.0 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort162.0 pointsindependentEpoch AI Benchmarking Hub
ECImax effort162.0 pointsindependentEpoch AI Benchmarking Hub
ECIxhigh effort162.0 pointsindependentEpoch AI Benchmarking Hub
Vals Indexmax effort63.7 pointsindependentVals AI
factuality
SimpleQA Verifiedmax effort71.6 %independentEpoch AI Benchmarking Hub
2026-07-09
SimpleQA Verifiedmax effort69.7 %independentEpoch AI Benchmarking Hub
2026-08-10
human preference
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1533.9
1518.21549.6
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1530.5
1518.31542.7
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
1528.6
1520.51536.7
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
xhigh effort1526.1
1504.21548.0
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
xhigh effort1522.5
1500.11544.8
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1518.5
1511.41525.5
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
xhigh effort1508.7
1490.51527.0
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1508.3
1486.01530.5
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1507.7
1490.61524.9
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1506.9
1500.91512.8
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1506.7
1498.61514.8
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
xhigh effort1504.8
1490.61519.0
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
xhigh effort1503.1
1485.31520.8
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
xhigh effort1498.0
1482.81513.2
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1496.9
1490.31503.5
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
xhigh effort1495.8
1469.71521.9
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1495.5
1488.61502.4
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1495.1
1481.21508.9
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1490.6
1481.21500.1
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1490.4
1478.01502.8
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1488.7
1481.81495.6
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1487.2
1477.11497.4
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
xhigh effort1487.0
1468.71505.4
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1485.5
1478.01493.0
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1485.3
1475.21495.4
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1484.8
1466.71502.9
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
xhigh effort1484.6
1473.21496.1
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1484.3
1475.91492.8
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1483.2
1478.01488.3
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
xhigh effort1482.3
1462.21502.5
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
xhigh effort1480.4
1457.91502.9
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1479.2
1464.31494.0
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1476.7
1467.11486.4
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
xhigh effort1475.4
1451.21499.6
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1473.9
1465.21482.6
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1472.7
1466.61478.8
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
xhigh effort1463.9
1449.31478.6
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1452.4
1428.41476.4
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
max effort95.2 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
max effort94.1 %independentArtificial Analysis
GPQA Diamondmax effort93.5 %independentEpoch AI Benchmarking Hub
2026-07-09
GPQA Diamond
implementation=artificial-analysis
xhigh effort93.1 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
high effort92.8 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
medium effort92.6 %independentArtificial Analysis
GPQA Diamondlow effort89.9 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
low effort89.8 %independentArtificial Analysis
GPQA Diamondno reasoning82.8 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
no reasoning79.0 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
max effort49.5 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
xhigh effort47.3 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
high effort46.0 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
medium effort42.2 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
low effort39.4 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
no reasoning16.7 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
max effort89.1 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort87.7 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort85.9 %independentLiveBench
long context instruction
AA-LCRmax effort84.0 %independentArtificial Analysis
AA-LCRxhigh effort82.3 %independentArtificial Analysis
AA-LCRhigh effort81.7 %independentArtificial Analysis
AA-LCRmedium effort80.3 %independentArtificial Analysis
AA-LCRlow effort78.0 %independentArtificial Analysis
AA-LCRno reasoning62.3 %independentArtificial Analysis
IFBenchmax effort72.7 %independentArtificial Analysis
IFBenchxhigh effort71.0 %independentArtificial Analysis
IFBenchmedium effort69.6 %independentArtificial Analysis
IFBenchhigh effort69.2 %independentArtificial Analysis
IFBenchlow effort66.5 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort71.8 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort67.4 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
max effort88.8 %independentVals AI
professional
CorpFinmax effort64.4 %independentVals AI
LegalBenchmax effort87.0 %independentVals AI
TaxEvalmax effort74.8 %independentVals AI
reasoning math
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort99.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
xhigh effort98.8 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
xhigh effort97.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort97.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort97.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort96.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort93.9 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort93.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort92.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort92.5 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
xhigh effort91.9 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
xhigh effort90.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort86.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort85.4 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort84.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort74.5 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort70.7 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort67.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort42.5 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort38.5 %independentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
max effort7.8 %independentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
xhigh effort7.0 %independentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
high effort2.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort1.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort1.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
medium effort1.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
xhigh effort1.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
xhigh effort1.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort0.8 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort0.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort0.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort0.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort0.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
xhigh effort0.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort0.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
low effort0.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
xhigh effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort0.1 usd_per_taskindependentARC Prize Leaderboard
EnigmaEvalhigh effort37.1 %
34.339.9
independentScale Labs
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort96.2 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort95.5 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort91.7 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort90.2 %independentLiveBench
OTIS Mock AIME 2024–2025max effort100.0 %independentEpoch AI Benchmarking Hub
2026-07-09
OTIS Mock AIME 2024–2025low effort95.6 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–2025no reasoning68.9 %independentEpoch AI Benchmarking Hub
2026-08-07

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench 2.189.5 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench 2.188.0 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench 2.187.3 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench 2.186.1 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench 2.176.8 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench 2.174.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench Hard65.9 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench Hard62.9 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench Hard62.1 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench Hard61.4 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 SolTerminal-Bench Hard60.6 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.