BenchAtlas

Rankings / OpenAI

GPT 4.1 (2025-04-14)

released 2025-04-14

BenchAtlas Index

as of 2026-09-11
37.0
high effort
rank #328
10 families · 4 categories · medium
28.7
base configuration
rank #388
18 families · 6 categories · high

Benchmark evidence

72 results
agentic coding
τ²-Bench47.1 %independentArtificial Analysis
coding
LiveCodeBench v6
implementation=artificial-analysis
45.7 %independentArtificial Analysis
SciCode38.1 %independentArtificial Analysis
external indices
Intelligence Index v4.112.7 pointsindependentArtificial Analysis
AA Math Index34.7 pointsindependentArtificial Analysis
ECI136.8 pointsindependentEpoch AI Benchmarking Hub
factuality
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
51.1 %
50.152.2
independentScale Labs
SimpleQA Verified31.1 %independentEpoch AI Benchmarking Hub
2026-08-31
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1456.5
1449.91463.2
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1442.9
1437.51448.3
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1438.0
1431.61444.4
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1437.0
1426.41447.7
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1432.5
1420.71444.3
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1431.5
1426.61436.4
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1430.8
1424.01437.8
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1428.0
1416.41439.7
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1425.9
1419.51432.3
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1424.9
1418.01431.9
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1422.6
1417.81427.5
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1422.5
1411.81433.1
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1421.3
1398.91443.8
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1414.7
1407.61421.7
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1414.4
1410.71418.1
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1409.6
1391.21427.9
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1409.5
1392.81426.1
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1409.3
1399.91418.8
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1405.0
1393.21416.8
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1404.8
1398.61411.0
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1402.9
1397.11408.7
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1402.7
1397.61407.8
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1402.5
1394.81410.2
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1402.0
1397.31406.6
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1400.7
1393.91407.4
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1374.2
1363.91384.5
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1372.5
1361.91383.2
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1370.2
1351.41389.1
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1347.2
1328.61365.8
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond66.9 %independentEpoch AI Benchmarking Hub
2025-04-14
GPQA Diamond
implementation=artificial-analysis
66.6 %independentArtificial Analysis
GPQA Diamond
implementation=vals-ai
high effort65.4 %independentVals AI
Humanity's Last Exam
contamination=Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. · implementation=scale
5.4 %
4.56.3
independentScale Labs
Humanity's Last Exam
implementation=artificial-analysis
4.2 %independentArtificial Analysis
MMLU-Pro
implementation=artificial-analysis
80.6 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
high effort80.5 %independentVals AI
long context instruction
AA-LCR68.3 %independentArtificial Analysis
IFBench43.0 %independentArtificial Analysis
MultiChallenge39.4 %
38.540.4
independentScale Labs
multimodal
MMMU
implementation=vals-ai
high effort72.4 %independentVals AI
VISTA45.3 %
44.446.3
independentScale Labs
professional
CaseLawhigh effort69.9 %independentVals AI
CorpFinhigh effort63.1 %independentVals AI
LegalBenchhigh effort83.1 %independentVals AI
MedQAhigh effort91.2 %independentVals AI
TaxEvalhigh effort75.1 %independentVals AI
reasoning math
AIME
implementation=artificial-analysis
43.7 %independentArtificial Analysis
AIME
implementation=vals-ai
high effort39.6 %independentVals AI
AIME
year=2025 · implementation=artificial-analysis
34.7 %independentArtificial Analysis
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
11.8 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
5.5 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
0.4 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.0 %independentARC Prize Leaderboard
EnigmaEval
contamination=Potential contamination warning: This model was evaluated after the public release of EnigmaEval, allowing model builder access to the prompts and solutions.
2.2 %
1.33.0
independentScale Labs
FrontierMath5.5 %independentEpoch AI Benchmarking Hub
2025-04-14
FrontierMath Tier 4older version0.0 %independentEpoch AI Benchmarking Hub
2025-07-01
MATH-500
implementation=artificial-analysis
91.3 %independentArtificial Analysis
MATH-500
implementation=vals-ai
high effort87.2 %independentVals AI
MATH Level 583.0 %independentEpoch AI Benchmarking Hub
2025-04-14
OTIS Mock AIME 2024–202538.3 %independentEpoch AI Benchmarking Hub
2025-04-14

Agent + model results

systems, not bare-model scores
agent + model Epoch Inspect harness + GPT 4.1 (2025-04-14)SWE-bench Verified48.5 %independentEpoch AI Benchmarking Hub
agent + model GUIRepair + GPT 4.1 (2025-04-14)SWE-bench Multimodal31.1 %communitySWE-bench Leaderboard
agent + model camel + GPT 4.1 (2025-04-14)Terminal-Bench 1.035.0 %unverifiedTerminal-Bench Leaderboard
agent + model terminus-1 + GPT 4.1 (2025-04-14)Terminal-Bench 1.030.3 %communityTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + GPT 4.1 (2025-04-14)Terminal-Bench Hard13.6 %independentArtificial Analysis
agent + model Codex CLI + GPT 4.1 (2025-04-14)Terminal-Bench 1.08.3 %communityTerminal-Bench Leaderboard

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants