BenchAtlas

Rankings / OpenAI

GPT 4.1 Mini (2025-04-14)

released 2025-04-14

BenchAtlas Index

as of 2026-09-11
27.4
high effort
rank #394
8 families · 4 categories · medium
27.4
base configuration
rank #395
15 families · 6 categories · high

Benchmark evidence

69 results
agentic coding
τ²-Bench52.9 %independentArtificial Analysis
τ²-Bench
subset=banking
5.4 %independentArtificial Analysis
coding
LiveCodeBench v6
implementation=artificial-analysis
48.3 %independentArtificial Analysis
SciCode40.4 %independentArtificial Analysis
external indices
AA Coding Index20.2 pointsindependentArtificial Analysis
Intelligence Index v4.110.2 pointsindependentArtificial Analysis
AA Math Index46.3 pointsindependentArtificial Analysis
ECI135.0 pointsindependentEpoch AI Benchmarking Hub
factuality
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
50.0 %
47.852.2
independentScale Labs
SimpleQA Verified12.7 %independentEpoch AI Benchmarking Hub
2026-08-31
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1433.0
1425.51440.6
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1418.2
1412.01424.3
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1412.8
1405.51420.1
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1408.3
1382.71433.8
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1401.7
1396.01407.4
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1397.2
1385.11409.3
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1396.2
1390.71401.7
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1395.3
1382.51408.1
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1393.1
1385.51400.8
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1392.0
1373.41410.5
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1388.8
1381.31396.3
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1384.4
1376.51392.3
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1383.9
1371.91395.9
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1383.0
1369.91396.0
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1382.3
1378.01386.7
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1379.0
1366.71391.3
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1378.6
1370.51386.7
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1372.4
1365.91378.9
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1368.6
1347.41389.8
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1365.8
1355.31376.4
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1363.3
1351.61374.9
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1360.6
1355.31365.9
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1359.7
1352.81366.7
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1358.2
1352.31364.1
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1353.1
1342.01364.2
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1348.5
1339.71357.3
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1347.6
1340.01355.2
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1346.6
1324.81368.4
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1329.8
1310.01349.5
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
high effort67.9 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
66.4 %independentArtificial Analysis
GPQA Diamond65.8 %independentEpoch AI Benchmarking Hub
2025-04-14
Humanity's Last Exam
implementation=artificial-analysis
5.0 %independentArtificial Analysis
MMLU-Pro
implementation=artificial-analysis
78.1 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
high effort77.2 %independentVals AI
long context instruction
AA-LCR44.0 %independentArtificial Analysis
IFBench38.3 %independentArtificial Analysis
multimodal
MMMU
implementation=vals-ai
high effort70.5 %independentVals AI
VISTA41.1 %
40.641.7
independentScale Labs
professional
CorpFinhigh effort57.9 %independentVals AI
LegalBenchhigh effort78.0 %independentVals AI
MedQAhigh effort84.6 %independentVals AI
TaxEvalhigh effort71.9 %independentVals AI
reasoning math
AIME
implementation=vals-ai
high effort49.4 %independentVals AI
AIME
year=2025 · implementation=artificial-analysis
46.3 %independentArtificial Analysis
AIME
implementation=artificial-analysis
43.0 %independentArtificial Analysis
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
7.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
3.5 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
0.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.0 %independentARC Prize Leaderboard
FrontierMath4.5 %independentEpoch AI Benchmarking Hub
2025-04-14
MATH-500
implementation=artificial-analysis
92.5 %independentArtificial Analysis
MATH-500
implementation=vals-ai
high effort88.0 %independentVals AI
MATH Level 587.3 %independentEpoch AI Benchmarking Hub
2025-04-14
OTIS Mock AIME 2024–202544.7 %independentEpoch AI Benchmarking Hub
2025-04-14

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + GPT 4.1 Mini (2025-04-14)Terminal-Bench 2.110.1 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 4.1 Mini (2025-04-14)Terminal-Bench Hard7.6 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants