BenchAtlas

Rankings / Anthropic

Claude Opus 4.1 20250805

released 2025-08-05

BenchAtlas Index

as of 2026-09-11
65.7
thinking
rank #107
8 families · 5 categories · high
61.5
base configuration
rank #142
8 families · 4 categories · medium
58.7
thinking
rank #173
7 families · 5 categories · medium

Benchmark evidence

68 results
external indices
ECI144.1 pointsindependentEpoch AI Benchmarking Hub
ECI24K budget144.1 pointsindependentEpoch AI Benchmarking Hub
ECI32K budget144.1 pointsindependentEpoch AI Benchmarking Hub
ECI16K budget144.1 pointsindependentEpoch AI Benchmarking Hub
ECI27K budget144.1 pointsindependentEpoch AI Benchmarking Hub
factuality
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
thinking94.2 %
92.496.0
independentScale Labs
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
87.4 %
85.789.1
independentScale Labs
SimpleQA Verified27K budget34.8 %independentEpoch AI Benchmarking Hub
2025-12-09
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1505.5
1500.01510.9
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1486.6
1482.11491.1
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1481.9
1476.61487.2
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1477.5
1473.51481.4
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1472.5
1462.91482.1
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1472.2
1467.11477.3
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1469.0
1463.21474.9
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1468.4
1460.11476.7
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1467.1
1453.61480.7
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1465.4
1455.91474.9
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1463.7
1457.81469.6
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1461.5
1444.51478.5
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1458.2
1449.71466.8
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1456.0
1451.91460.2
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1454.7
1449.81459.6
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1451.3
1444.11458.6
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1450.2
1439.81460.6
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1448.2
1444.11452.4
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1447.7
1444.71450.7
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1446.5
1440.91452.1
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1443.9
1438.71449.1
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1442.5
1426.81458.1
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1442.1
1435.81448.4
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1440.0
1430.41449.7
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1433.9
1428.31439.4
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1433.6
1429.91437.3
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1432.0
1423.41440.7
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1411.6
1393.71429.6
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1402.1
1386.31417.8
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond16K budget77.3 %independentEpoch AI Benchmarking Hub
2025-08-05
GPQA Diamond27K budget76.8 %independentEpoch AI Benchmarking Hub
2025-08-05
GPQA Diamond
implementation=vals-ai
thinking76.3 %independentVals AI
GPQA Diamond73.2 %independentEpoch AI Benchmarking Hub
2025-08-05
GPQA Diamond
implementation=vals-ai
70.0 %independentVals AI
Humanity's Last Exam
contamination=Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. · implementation=scale
thinking11.5 %
10.312.8
independentScale Labs
Humanity's Last Exam
contamination=Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. · implementation=scale
7.9 %
6.99.0
independentScale Labs
MMLU-Pro
implementation=vals-ai
thinking87.9 %independentVals AI
MMLU-Pro
implementation=vals-ai
87.2 %independentVals AI
long context instruction
MultiChallengethinking57.2 %
56.358.1
independentScale Labs
multimodal
MMMU
implementation=vals-ai
thinking77.5 %independentVals AI
MMMU
implementation=vals-ai
73.7 %independentVals AI
VISTAthinking48.4 %
48.048.9
independentScale Labs
VISTA45.3 %
44.645.9
independentScale Labs
professional
LegalBench83.5 %independentVals AI
MedQAthinking93.6 %independentVals AI
MedQA92.5 %independentVals AI
TaxEvalthinking73.7 %independentVals AI
TaxEval71.5 %independentVals AI
reasoning math
AIME
implementation=vals-ai
thinking78.2 %independentVals AI
AIME
implementation=vals-ai
44.2 %independentVals AI
EnigmaEval
contamination=Potential contamination warning: This model was evaluated after the public release of EnigmaEval, allowing model builder access to the prompts and solutions.
thinking7.2 %
5.78.7
independentScale Labs
EnigmaEval
contamination=Potential contamination warning: This model was evaluated after the public release of EnigmaEval, allowing model builder access to the prompts and solutions.
4.8 %
3.66.0
independentScale Labs
FrontierMath27K budget7.2 %independentEpoch AI Benchmarking Hub
2025-08-05
FrontierMath5.9 %independentEpoch AI Benchmarking Hub
2025-08-05
FrontierMath Tier 4older version27K budget4.2 %independentEpoch AI Benchmarking Hub
2025-08-05
MATH-500
implementation=vals-ai
thinking95.4 %independentVals AI
MATH-500
implementation=vals-ai
93.0 %independentVals AI
OTIS Mock AIME 2024–202527K budget68.9 %independentEpoch AI Benchmarking Hub
2025-08-05
OTIS Mock AIME 2024–202516K budget64.4 %independentEpoch AI Benchmarking Hub
2025-08-05
OTIS Mock AIME 2024–202540.0 %independentEpoch AI Benchmarking Hub
2025-08-05

Agent + model results

systems, not bare-model scores
agent + model Epoch Inspect harness + Claude Opus 4.1 20250805SWE-bench Verified73.3 %independentEpoch AI Benchmarking Hub
agent + model terminus-2 + Claude Opus 4.1 20250805Terminal-Bench 2.038.0 %communityTerminal-Bench Leaderboard
agent + model OpenHands + Claude Opus 4.1 20250805Terminal-Bench 2.036.9 %communityTerminal-Bench Leaderboard
agent + model mini-SWE-agent + Claude Opus 4.1 20250805Terminal-Bench 2.035.1 %communityTerminal-Bench Leaderboard
agent + model Claude Code + Claude Opus 4.1 20250805Terminal-Bench 2.034.8 %communityTerminal-Bench Leaderboard

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants