Rankings / OpenAI
GPT 5.4 Mini (2026-03-17)
proprietaryreleased 2026-03-17
BenchAtlas Index
as of 2026-09-11
high effort
rank #197
5 families · 3 categories · medium
medium effort
rank #216
8 families · 5 categories · high
xhigh effort
rank #228
7 families · 4 categories · medium
Benchmark evidence
143 results
agentic coding
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | xhigh effort | 41.7 % | independent | LiveBench |
| τ²-Bench | xhigh effort | 83.3 % | independent | Artificial Analysis |
| τ²-Bench | medium effort | 36.5 % | independent | Artificial Analysis |
| τ²-Bench subset=banking | xhigh effort | 25.6 % | independent | Artificial Analysis |
| τ²-Bench | no reasoning | 23.4 % | independent | Artificial Analysis |
coding
| Release 2026-06-25 tasks_counted=2 · livebench_version=2026-06-25 | xhigh effort | 71.6 % | independent | LiveBench |
| SciCode | xhigh effort | 52.1 % | independent | Artificial Analysis |
| SciCode | medium effort | 44.2 % | independent | Artificial Analysis |
| SciCode | no reasoning | 39.6 % | independent | Artificial Analysis |
data analysis
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | xhigh effort | 70.8 % | independent | LiveBench |
external indices
| AA Coding Index | xhigh effort | 56.1 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | xhigh effort | 24.6 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | medium effort | 19.7 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | no reasoning | 11.1 points | independent | Artificial Analysis |
| ECI | medium effort | 149.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | xhigh effort | 149.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | high effort | 149.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | 149.0 points | independent | Epoch AI Benchmarking Hub | |
| ECI | no reasoning | 149.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | low effort | 149.0 points | independent | Epoch AI Benchmarking Hub |
| Vals Index | xhigh effort | 39.6 points | independent | Vals AI |
factuality
| SimpleQA Verified | high effort | 29.4 % | independent | Epoch AI Benchmarking Hub 2026-08-27 |
| SimpleQA Verified | high effort | 28.6 % | independent | Epoch AI Benchmarking Hub 2026-04-15 |
human preference
| Coding (style control)older version arena=text · category=coding · style_control=true | high effort | 1496.9 1490.6–1503.2 | community | LMArena Leaderboard Dataset |
| Coding (style control)older version arena=text · category=coding · style_control=true | 1495.9 1489.9–1501.9 | community | LMArena Leaderboard Dataset | |
| Industry Software And It Services (style control)older version arena=text · category=industry_software_and_it_services · style_control=true | high effort | 1487.2 1481.7–1492.8 | community | LMArena Leaderboard Dataset |
| Industry Software And It Services (style control)older version arena=text · category=industry_software_and_it_services · style_control=true | 1486.0 1480.6–1491.3 | community | LMArena Leaderboard Dataset | |
| Expert (style control)older version arena=text · category=expert · style_control=true | high effort | 1480.9 1471.6–1490.3 | community | LMArena Leaderboard Dataset |
| Expert (style control)older version arena=text · category=expert · style_control=true | 1480.6 1472.0–1489.3 | community | LMArena Leaderboard Dataset | |
| Chinese (style control)older version arena=text · category=chinese · style_control=true | high effort | 1480.5 1468.2–1492.9 | community | LMArena Leaderboard Dataset |
| Chinese (style control)older version arena=text · category=chinese · style_control=true | 1478.5 1467.3–1489.6 | community | LMArena Leaderboard Dataset | |
| French (style control)older version arena=text · category=french · style_control=true | 1474.8 1459.6–1490.0 | community | LMArena Leaderboard Dataset | |
| Multi Turn (style control)older version arena=text · category=multi_turn · style_control=true | high effort | 1472.1 1464.9–1479.2 | community | LMArena Leaderboard Dataset |
| French (style control)older version arena=text · category=french · style_control=true | high effort | 1471.7 1455.1–1488.3 | community | LMArena Leaderboard Dataset |
| Hard Prompts (style control)older version arena=text · category=hard_prompts · style_control=true | high effort | 1471.6 1466.8–1476.4 | community | LMArena Leaderboard Dataset |
| Hard Prompts (style control)older version arena=text · category=hard_prompts · style_control=true | 1469.9 1465.3–1474.5 | community | LMArena Leaderboard Dataset | |
| Hard Prompts English (style control)older version arena=text · category=hard_prompts_english · style_control=true | high effort | 1469.6 1463.6–1475.6 | community | LMArena Leaderboard Dataset |
| Multi Turn (style control)older version arena=text · category=multi_turn · style_control=true | 1468.3 1461.5–1475.1 | community | LMArena Leaderboard Dataset | |
| Hard Prompts English (style control)older version arena=text · category=hard_prompts_english · style_control=true | 1466.8 1461.0–1472.6 | community | LMArena Leaderboard Dataset | |
| Industry Life And Physical And Social Science (style control)older version arena=text · category=industry_life_and_physical_and_social_science · style_control=true | 1466.2 1459.1–1473.2 | community | LMArena Leaderboard Dataset | |
| Industry Life And Physical And Social Science (style control)older version arena=text · category=industry_life_and_physical_and_social_science · style_control=true | high effort | 1466.0 1458.5–1473.4 | community | LMArena Leaderboard Dataset |
| Polish (style control)older version arena=text · category=polish · style_control=true | 1460.9 1442.7–1479.1 | community | LMArena Leaderboard Dataset | |
| Industry Business And Management And Financial Operations (style control)older version arena=text · category=industry_business_and_management_and_financial_operations · style_control=true | high effort | 1460.9 1453.9–1467.9 | community | LMArena Leaderboard Dataset |
| Industry Business And Management And Financial Operations (style control)older version arena=text · category=industry_business_and_management_and_financial_operations · style_control=true | 1459.6 1452.9–1466.2 | community | LMArena Leaderboard Dataset | |
| Industry Legal And Government (style control)older version arena=text · category=industry_legal_and_government · style_control=true | 1456.2 1446.6–1465.8 | community | LMArena Leaderboard Dataset | |
| Industry Mathematical (style control)older version arena=text · category=industry_mathematical · style_control=true | high effort | 1455.2 1443.1–1467.4 | community | LMArena Leaderboard Dataset |
| Industry Legal And Government (style control)older version arena=text · category=industry_legal_and_government · style_control=true | high effort | 1454.4 1444.2–1464.7 | community | LMArena Leaderboard Dataset |
| Russian (style control)older version arena=text · category=russian · style_control=true | 1454.2 1445.9–1462.5 | community | LMArena Leaderboard Dataset | |
| Russian (style control)older version arena=text · category=russian · style_control=true | high effort | 1453.7 1444.8–1462.5 | community | LMArena Leaderboard Dataset |
| English (style control)older version arena=text · category=english · style_control=true | high effort | 1453.1 1447.9–1458.3 | community | LMArena Leaderboard Dataset |
| Polish (style control)older version arena=text · category=polish · style_control=true | high effort | 1453.0 1433.9–1472.2 | community | LMArena Leaderboard Dataset |
| Industry Mathematical (style control)older version arena=text · category=industry_mathematical · style_control=true | 1452.5 1441.1–1463.9 | community | LMArena Leaderboard Dataset | |
| Industry Medicine And Healthcare (style control)older version arena=text · category=industry_medicine_and_healthcare · style_control=true | 1451.7 1441.7–1461.7 | community | LMArena Leaderboard Dataset | |
| Exclude Ties (style control)older version arena=text · category=exclude_ties · style_control=true | high effort | 1451.2 1446.0–1456.4 | community | LMArena Leaderboard Dataset |
| English (style control)older version arena=text · category=english · style_control=true | 1450.9 1445.8–1455.9 | community | LMArena Leaderboard Dataset | |
| Industry Medicine And Healthcare (style control)older version arena=text · category=industry_medicine_and_healthcare · style_control=true | high effort | 1450.6 1440.0–1461.2 | community | LMArena Leaderboard Dataset |
| Exclude Ties (style control)older version arena=text · category=exclude_ties · style_control=true | 1450.0 1445.0–1455.1 | community | LMArena Leaderboard Dataset | |
| Overall (style control) arena=text · category=overall · style_control=true | high effort | 1449.1 1445.0–1453.1 | community | LMArena Leaderboard Dataset |
| Longer Query (style control)older version arena=text · category=longer_query · style_control=true | high effort | 1448.8 1443.1–1454.5 | community | LMArena Leaderboard Dataset |
| Overall (style control) arena=text · category=overall · style_control=true | 1448.2 1444.4–1452.1 | community | LMArena Leaderboard Dataset | |
| Longer Query (style control)older version arena=text · category=longer_query · style_control=true | 1448.0 1442.5–1453.4 | community | LMArena Leaderboard Dataset | |
| German (style control)older version arena=text · category=german · style_control=true | high effort | 1445.2 1422.8–1467.6 | community | LMArena Leaderboard Dataset |
| German (style control)older version arena=text · category=german · style_control=true | 1442.3 1421.4–1463.2 | community | LMArena Leaderboard Dataset | |
| Math (style control)older version arena=text · category=math · style_control=true | 1440.7 1429.6–1451.7 | community | LMArena Leaderboard Dataset | |
| Math (style control)older version arena=text · category=math · style_control=true | high effort | 1439.4 1427.8–1451.0 | community | LMArena Leaderboard Dataset |
| Non English (style control)older version arena=text · category=non_english · style_control=true | high effort | 1439.3 1434.3–1444.3 | community | LMArena Leaderboard Dataset |
| Non English (style control)older version arena=text · category=non_english · style_control=true | 1438.9 1434.1–1443.7 | community | LMArena Leaderboard Dataset | |
| Spanish (style control)older version arena=text · category=spanish · style_control=true | 1436.2 1420.9–1451.5 | community | LMArena Leaderboard Dataset | |
| Instruction Following (style control)older version arena=text · category=instruction_following · style_control=true | high effort | 1435.8 1430.0–1441.7 | community | LMArena Leaderboard Dataset |
| Instruction Following (style control)older version arena=text · category=instruction_following · style_control=true | 1434.5 1428.9–1440.0 | community | LMArena Leaderboard Dataset | |
| Spanish (style control)older version arena=text · category=spanish · style_control=true | high effort | 1430.4 1413.6–1447.2 | community | LMArena Leaderboard Dataset |
| Industry Writing And Literature And Language (style control)older version arena=text · category=industry_writing_and_literature_and_language · style_control=true | high effort | 1425.3 1418.8–1431.8 | community | LMArena Leaderboard Dataset |
| Industry Writing And Literature And Language (style control)older version arena=text · category=industry_writing_and_literature_and_language · style_control=true | 1425.0 1418.8–1431.2 | community | LMArena Leaderboard Dataset | |
| Industry Entertainment And Sports And Media (style control)older version arena=text · category=industry_entertainment_and_sports_and_media · style_control=true | high effort | 1411.3 1404.3–1418.3 | community | LMArena Leaderboard Dataset |
| Industry Entertainment And Sports And Media (style control)older version arena=text · category=industry_entertainment_and_sports_and_media · style_control=true | 1410.8 1404.1–1417.5 | community | LMArena Leaderboard Dataset | |
| Japanese (style control)older version arena=text · category=japanese · style_control=true | 1408.9 1382.8–1435.1 | community | LMArena Leaderboard Dataset | |
| Creative Writing (style control)older version arena=text · category=creative_writing · style_control=true | high effort | 1404.3 1396.6–1411.9 | community | LMArena Leaderboard Dataset |
| Creative Writing (style control)older version arena=text · category=creative_writing · style_control=true | 1404.1 1396.8–1411.4 | community | LMArena Leaderboard Dataset | |
| Korean (style control)older version arena=text · category=korean · style_control=true | 1397.8 1377.1–1418.4 | community | LMArena Leaderboard Dataset | |
| Korean (style control)older version arena=text · category=korean · style_control=true | high effort | 1395.6 1373.1–1418.1 | community | LMArena Leaderboard Dataset |
knowledge science
| GPQA Diamond implementation=artificial-analysis | xhigh effort | 87.5 % | independent | Artificial Analysis |
| GPQA Diamond | xhigh effort | 86.9 % | independent | Epoch AI Benchmarking Hub 2026-08-07 |
| GPQA Diamond | high effort | 83.6 % | independent | Epoch AI Benchmarking Hub 2026-04-15 |
| GPQA Diamond implementation=vals-ai | xhigh effort | 83.1 % | independent | Vals AI |
| GPQA Diamond implementation=artificial-analysis | medium effort | 82.3 % | independent | Artificial Analysis |
| GPQA Diamond | no reasoning | 64.1 % | independent | Epoch AI Benchmarking Hub 2026-08-07 |
| GPQA Diamond implementation=artificial-analysis | no reasoning | 60.6 % | independent | Artificial Analysis |
| Humanity's Last Exam implementation=artificial-analysis | xhigh effort | 28.1 % | independent | Artificial Analysis |
| Humanity's Last Exam implementation=artificial-analysis | medium effort | 18.6 % | independent | Artificial Analysis |
| Humanity's Last Exam implementation=artificial-analysis | no reasoning | 5.9 % | independent | Artificial Analysis |
| MMLU-Pro implementation=vals-ai | xhigh effort | 84.6 % | independent | Vals AI |
language
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | xhigh effort | 71.0 % | independent | LiveBench |
long context instruction
| AA-LCR | xhigh effort | 77.0 % | independent | Artificial Analysis |
| AA-LCR | medium effort | 67.0 % | independent | Artificial Analysis |
| AA-LCR | no reasoning | 37.0 % | independent | Artificial Analysis |
| IFBench | xhigh effort | 73.3 % | independent | Artificial Analysis |
| IFBench | medium effort | 64.8 % | independent | Artificial Analysis |
| IFBench | no reasoning | 38.8 % | independent | Artificial Analysis |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | xhigh effort | 59.8 % | independent | LiveBench |
multimodal
| MMMU implementation=vals-ai | xhigh effort | 79.2 % | independent | Vals AI |
professional
| CaseLaw | xhigh effort | 51.7 % | independent | Vals AI |
| CorpFin | xhigh effort | 60.9 % | independent | Vals AI |
| TaxEval | xhigh effort | 71.2 % | independent | Vals AI |
reasoning math
| AIME implementation=vals-ai | xhigh effort | 95.6 % | independent | Vals AI |
| ARC-AGI-1older version split=public_eval · model_type=CoT | xhigh effort | 75.1 % | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=public_eval · model_type=CoT | high effort | 66.3 % | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | xhigh effort | 63.7 % | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | high effort | 58.0 % | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=public_eval · model_type=CoT | medium effort | 55.4 % | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | medium effort | 40.8 % | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=public_eval · model_type=CoT | low effort | 31.8 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | xhigh effort | 18.9 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | xhigh effort | 17.8 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | high effort | 13.2 % | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | low effort | 13.0 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | high effort | 7.0 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | medium effort | 5.4 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | medium effort | 4.4 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | low effort | 1.1 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | low effort | 0.8 % | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | xhigh effort | 0.8 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | xhigh effort | 0.8 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | high effort | 0.6 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | high effort | 0.6 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | xhigh effort | 0.5 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=public_eval · model_type=CoT | xhigh effort | 0.3 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | medium effort | 0.3 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | medium effort | 0.3 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | high effort | 0.3 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=public_eval · model_type=CoT | high effort | 0.2 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | medium effort | 0.2 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=public_eval · model_type=CoT | medium effort | 0.1 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=public_eval · model_type=CoT | low effort | 0.1 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-2 split=semi_private · model_type=CoT | low effort | 0.1 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=semi_private · model_type=CoT | low effort | 0.0 usd_per_task | independent | ARC Prize Leaderboard |
| ARC-AGI-1older version split=public_eval · model_type=CoT | low effort | 0.0 usd_per_task | independent | ARC Prize Leaderboard |
| FrontierMath | high effort | 28.3 % | independent | Epoch AI Benchmarking Hub 2026-04-15 |
| FrontierMath Tier 4older version | high effort | 2.1 % | independent | Epoch AI Benchmarking Hub 2026-04-15 |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | xhigh effort | 78.5 % | independent | LiveBench |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | xhigh effort | 71.3 % | independent | LiveBench |
| OTIS Mock AIME 2024–2025 | xhigh effort | 88.9 % | independent | Epoch AI Benchmarking Hub 2026-08-07 |
| OTIS Mock AIME 2024–2025 | high effort | 87.2 % | independent | Epoch AI Benchmarking Hub 2026-04-15 |
| OTIS Mock AIME 2024–2025 | no reasoning | 26.7 % | independent | Epoch AI Benchmarking Hub 2026-08-07 |
Agent + model results
systems, not bare-model scores
| agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17) | Terminal-Bench 2.1 | 59.2 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17) | Terminal-Bench Hard | 52.3 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17) | Terminal-Bench Hard | 34.1 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + GPT 5.4 Mini (2026-03-17) | Terminal-Bench Hard | 18.2 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
