BenchAtlas

Rankings / OpenAI

GPT 5.4 Mini (2026-03-17)

proprietary

released 2026-03-17

BenchAtlas Index

as of 2026-09-11
55.4
high effort
rank #197
5 families · 3 categories · medium
52.9
medium effort
rank #216
8 families · 5 categories · high
51.5
xhigh effort
rank #228
7 families · 4 categories · medium

Benchmark evidence

143 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort41.7 %independentLiveBench
τ²-Benchxhigh effort83.3 %independentArtificial Analysis
τ²-Benchmedium effort36.5 %independentArtificial Analysis
τ²-Bench
subset=banking
xhigh effort25.6 %independentArtificial Analysis
τ²-Benchno reasoning23.4 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
xhigh effort71.6 %independentLiveBench
SciCodexhigh effort52.1 %independentArtificial Analysis
SciCodemedium effort44.2 %independentArtificial Analysis
SciCodeno reasoning39.6 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort70.8 %independentLiveBench
external indices
AA Coding Indexxhigh effort56.1 pointsindependentArtificial Analysis
Intelligence Index v4.1xhigh effort24.6 pointsindependentArtificial Analysis
Intelligence Index v4.1medium effort19.7 pointsindependentArtificial Analysis
Intelligence Index v4.1no reasoning11.1 pointsindependentArtificial Analysis
ECImedium effort149.0 pointsindependentEpoch AI Benchmarking Hub
ECIxhigh effort149.0 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort149.0 pointsindependentEpoch AI Benchmarking Hub
ECI149.0 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning149.0 pointsindependentEpoch AI Benchmarking Hub
ECIlow effort149.0 pointsindependentEpoch AI Benchmarking Hub
Vals Indexxhigh effort39.6 pointsindependentVals AI
factuality
SimpleQA Verifiedhigh effort29.4 %independentEpoch AI Benchmarking Hub
2026-08-27
SimpleQA Verifiedhigh effort28.6 %independentEpoch AI Benchmarking Hub
2026-04-15
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
high effort1496.9
1490.61503.2
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
1495.9
1489.91501.9
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
high effort1487.2
1481.71492.8
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1486.0
1480.61491.3
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
high effort1480.9
1471.61490.3
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1480.6
1472.01489.3
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
high effort1480.5
1468.21492.9
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1478.5
1467.31489.6
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1474.8
1459.61490.0
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
high effort1472.1
1464.91479.2
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
high effort1471.7
1455.11488.3
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
high effort1471.6
1466.81476.4
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1469.9
1465.31474.5
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
high effort1469.6
1463.61475.6
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1468.3
1461.51475.1
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1466.8
1461.01472.6
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1466.2
1459.11473.2
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
high effort1466.0
1458.51473.4
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1460.9
1442.71479.1
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
high effort1460.9
1453.91467.9
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1459.6
1452.91466.2
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1456.2
1446.61465.8
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
high effort1455.2
1443.11467.4
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
high effort1454.4
1444.21464.7
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1454.2
1445.91462.5
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
high effort1453.7
1444.81462.5
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
high effort1453.1
1447.91458.3
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
high effort1453.0
1433.91472.2
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1452.5
1441.11463.9
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1451.7
1441.71461.7
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
high effort1451.2
1446.01456.4
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1450.9
1445.81455.9
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
high effort1450.6
1440.01461.2
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1450.0
1445.01455.1
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
high effort1449.1
1445.01453.1
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
high effort1448.8
1443.11454.5
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1448.2
1444.41452.1
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1448.0
1442.51453.4
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
high effort1445.2
1422.81467.6
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1442.3
1421.41463.2
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1440.7
1429.61451.7
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
high effort1439.4
1427.81451.0
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
high effort1439.3
1434.31444.3
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1438.9
1434.11443.7
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1436.2
1420.91451.5
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
high effort1435.8
1430.01441.7
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1434.5
1428.91440.0
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
high effort1430.4
1413.61447.2
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
high effort1425.3
1418.81431.8
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1425.0
1418.81431.2
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
high effort1411.3
1404.31418.3
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1410.8
1404.11417.5
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1408.9
1382.81435.1
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
high effort1404.3
1396.61411.9
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1404.1
1396.81411.4
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1397.8
1377.11418.4
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
high effort1395.6
1373.11418.1
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=artificial-analysis
xhigh effort87.5 %independentArtificial Analysis
GPQA Diamondxhigh effort86.9 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamondhigh effort83.6 %independentEpoch AI Benchmarking Hub
2026-04-15
GPQA Diamond
implementation=vals-ai
xhigh effort83.1 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
medium effort82.3 %independentArtificial Analysis
GPQA Diamondno reasoning64.1 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
no reasoning60.6 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
xhigh effort28.1 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
medium effort18.6 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
no reasoning5.9 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
xhigh effort84.6 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort71.0 %independentLiveBench
long context instruction
AA-LCRxhigh effort77.0 %independentArtificial Analysis
AA-LCRmedium effort67.0 %independentArtificial Analysis
AA-LCRno reasoning37.0 %independentArtificial Analysis
IFBenchxhigh effort73.3 %independentArtificial Analysis
IFBenchmedium effort64.8 %independentArtificial Analysis
IFBenchno reasoning38.8 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort59.8 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
xhigh effort79.2 %independentVals AI
professional
CaseLawxhigh effort51.7 %independentVals AI
CorpFinxhigh effort60.9 %independentVals AI
TaxEvalxhigh effort71.2 %independentVals AI
reasoning math
AIME
implementation=vals-ai
xhigh effort95.6 %independentVals AI
ARC-AGI-1older version
split=public_eval · model_type=CoT
xhigh effort75.1 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort66.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
xhigh effort63.7 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort58.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort55.4 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort40.8 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort31.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
xhigh effort18.9 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
xhigh effort17.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort13.2 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort13.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort7.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort5.4 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort4.4 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort1.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort0.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
xhigh effort0.8 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
xhigh effort0.8 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort0.6 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort0.6 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
xhigh effort0.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
xhigh effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort0.0 usd_per_taskindependentARC Prize Leaderboard
FrontierMathhigh effort28.3 %independentEpoch AI Benchmarking Hub
2026-04-15
FrontierMath Tier 4older versionhigh effort2.1 %independentEpoch AI Benchmarking Hub
2026-04-15
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort78.5 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort71.3 %independentLiveBench
OTIS Mock AIME 2024–2025xhigh effort88.9 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–2025high effort87.2 %independentEpoch AI Benchmarking Hub
2026-04-15
OTIS Mock AIME 2024–2025no reasoning26.7 %independentEpoch AI Benchmarking Hub
2026-08-07

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17)Terminal-Bench 2.159.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17)Terminal-Bench Hard52.3 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17)Terminal-Bench Hard34.1 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17)Terminal-Bench Hard18.2 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.