BenchAtlas

Rankings / OpenAI

GPT 5.5 (2026-04-22)

proprietary

released 2026-04-22

BenchAtlas Index

as of 2026-09-11
81.5
xhigh effort
rank #7
8 families · 4 categories · medium
79.8
xhigh effort
rank #15
12 families · 5 categories · high
79.5
medium effort
rank #16
7 families · 4 categories · medium

Benchmark evidence

174 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort54.0 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort46.8 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
41.1 %independentLiveBench
τ²-Benchxhigh effort93.9 %independentArtificial Analysis
τ²-Benchhigh effort93.0 %independentArtificial Analysis
τ²-Benchmedium effort91.8 %independentArtificial Analysis
τ²-Benchlow effort83.9 %independentArtificial Analysis
τ²-Benchno reasoning69.3 %independentArtificial Analysis
τ²-Bench
subset=banking
xhigh effort39.0 %independentArtificial Analysis
τ²-Bench
subset=banking
high effort36.7 %independentArtificial Analysis
τ²-Bench
subset=banking
medium effort29.9 %independentArtificial Analysis
τ²-Bench
subset=banking
low effort24.9 %independentArtificial Analysis
τ²-Bench
subset=banking
no reasoning14.8 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
xhigh effort82.2 %independentLiveBench
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
high effort80.0 %independentLiveBench
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
78.6 %independentLiveBench
SciCodehigh effort56.1 %independentArtificial Analysis
SciCodexhigh effort55.8 %independentArtificial Analysis
SciCodemedium effort54.5 %independentArtificial Analysis
SciCodelow effort51.6 %independentArtificial Analysis
SciCodeno reasoning47.3 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort81.6 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort80.4 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
77.0 %independentLiveBench
external indices
AA Coding Indexxhigh effort74.9 pointsindependentArtificial Analysis
AA Coding Indexhigh effort71.6 pointsindependentArtificial Analysis
AA Coding Indexmedium effort71.5 pointsindependentArtificial Analysis
AA Coding Indexlow effort60.9 pointsindependentArtificial Analysis
AA Coding Indexno reasoning56.5 pointsindependentArtificial Analysis
Intelligence Index v4.1xhigh effort38.6 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort37.3 pointsindependentArtificial Analysis
Intelligence Index v4.1medium effort34.2 pointsindependentArtificial Analysis
Intelligence Index v4.1low effort30.7 pointsindependentArtificial Analysis
Intelligence Index v4.1no reasoning23.2 pointsindependentArtificial Analysis
ECImedium effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECIlow effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECIxhigh effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning159.2 pointsindependentEpoch AI Benchmarking Hub
ECI159.2 pointsindependentEpoch AI Benchmarking Hub
Vals Indexxhigh effort57.4 pointsindependentVals AI
factuality
SimpleQA Verifiedxhigh effort63.0 %independentEpoch AI Benchmarking Hub
2026-08-27
human preference
Chinese (style control)older version
arena=text · category=chinese · style_control=true
high effort1522.7
1508.71536.8
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
high effort1518.1
1511.11525.1
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
high effort1517.5
1507.31527.7
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1515.5
1504.81526.1
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1511.4
1503.21519.7
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
high effort1510.0
1503.81516.2
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
1508.8
1503.01514.6
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1504.0
1498.81509.2
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
high effort1500.5
1495.11505.9
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
high effort1500.1
1487.01513.2
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
high effort1499.8
1493.21506.5
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1499.8
1488.91510.6
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
high effort1499.2
1491.11507.4
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1498.5
1492.91504.2
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1498.1
1493.61502.6
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1495.9
1489.11502.7
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1495.6
1481.11510.2
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
high effort1491.7
1485.81497.6
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1490.5
1479.61501.4
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
high effort1489.7
1476.71502.7
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
high effort1488.9
1481.21496.7
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
high effort1488.8
1477.61500.0
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
high effort1488.1
1479.81496.3
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
high effort1487.7
1477.91497.5
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
high effort1487.5
1470.01505.1
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1486.3
1481.31491.3
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1485.7
1476.71494.8
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1485.6
1479.01492.1
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
high effort1485.2
1479.01491.5
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1485.0
1479.81490.1
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1483.8
1459.51508.2
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
high effort1483.8
1478.01489.5
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1482.5
1475.71489.4
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1481.9
1476.91486.8
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
high effort1481.9
1477.31486.4
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
high effort1479.4
1457.71501.1
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1479.3
1471.11487.5
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
high effort1478.6
1467.21490.0
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1477.2
1467.71486.7
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1477.2
1473.41481.0
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
high effort1477.1
1470.51483.7
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
high effort1475.5
1469.81481.1
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1474.1
1468.71479.6
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1473.8
1453.71493.9
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
high effort1471.8
1464.61479.0
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1469.6
1451.41487.8
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1466.7
1462.11471.4
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1465.9
1459.91472.0
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1462.3
1447.41477.2
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
high effort1460.0
1434.61485.5
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
high effort1457.9
1439.61476.2
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
high effort1455.0
1429.81480.3
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1448.5
1441.71455.4
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
high effort1448.4
1440.11456.8
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
high effort1446.4
1438.71454.0
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1446.2
1439.91452.5
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1428.6
1409.01448.3
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=artificial-analysis
xhigh effort93.5 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
high effort93.2 %independentArtificial Analysis
GPQA Diamond
implementation=vals-ai
xhigh effort93.2 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
medium effort92.6 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
low effort91.0 %independentArtificial Analysis
GPQA Diamondlow effort90.7 %independentEpoch AI Benchmarking Hub
2026-05-05
GPQA Diamondno reasoning77.3 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
no reasoning76.8 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
xhigh effort45.8 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
high effort45.0 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
medium effort42.4 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
low effort32.7 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
no reasoning13.7 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
xhigh effort88.1 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort87.8 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort87.4 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
85.6 %independentLiveBench
long context instruction
AA-LCRxhigh effort84.3 %independentArtificial Analysis
AA-LCRhigh effort84.3 %independentArtificial Analysis
AA-LCRmedium effort83.0 %independentArtificial Analysis
AA-LCRlow effort81.0 %independentArtificial Analysis
AA-LCRno reasoning64.0 %independentArtificial Analysis
IFBenchxhigh effort75.8 %independentArtificial Analysis
IFBenchhigh effort71.6 %independentArtificial Analysis
IFBenchmedium effort71.0 %independentArtificial Analysis
IFBenchlow effort64.3 %independentArtificial Analysis
IFBenchno reasoning46.1 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort71.4 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort70.7 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
65.7 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
xhigh effort88.3 %independentVals AI
professional
CaseLawxhigh effort66.2 %independentVals AI
CorpFinxhigh effort68.4 %independentVals AI
LegalBenchxhigh effort86.5 %independentVals AI
TaxEvalxhigh effort75.0 %independentVals AI
reasoning math
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort97.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort97.4 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort96.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort95.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort94.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort92.2 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort90.6 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort85.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort83.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort82.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort82.2 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort76.2 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort70.7 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort70.4 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort34.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort33.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort1.9 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort1.9 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort1.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort1.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort1.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort0.9 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort0.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort0.6 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort0.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort0.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort0.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort0.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort0.2 usd_per_taskindependentARC Prize Leaderboard
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort95.9 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort95.2 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
69.8 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort89.7 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort89.7 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
87.3 %independentLiveBench
OTIS Mock AIME 2024–2025low effort84.4 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–2025no reasoning57.8 %independentEpoch AI Benchmarking Hub
2026-08-07

Agent + model results

systems, not bare-model scores
agent + model nexau + GPT 5.5 (2026-04-22)Terminal-Bench 2.084.7 %unverifiedTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench 2.184.3 %independentArtificial Analysis
agent + model Codex CLI + GPT 5.5 (2026-04-22)Terminal-Bench 2.183.4 %communityTerminal-Bench Leaderboard
agent + model capy-build + GPT 5.5 (2026-04-22)Terminal-Bench 2.083.2 %unverifiedTerminal-Bench Leaderboard
agent + model Codex CLI + GPT 5.5 (2026-04-22)Terminal-Bench 2.082.3 %communityTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench 2.180.5 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench 2.179.4 %independentArtificial Analysis
agent + model terminus-2 + GPT 5.5 (2026-04-22)Terminal-Bench 2.178.2 %communityTerminal-Bench Leaderboard
agent + model clnkr + GPT 5.5 (2026-04-22)Terminal-Bench 2.066.1 %unverifiedTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench 2.165.5 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench 2.161.0 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench Hard60.6 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench Hard59.9 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench Hard57.6 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench Hard52.3 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.5 (2026-04-22)Terminal-Bench Hard49.2 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants