BenchAtlas

Rankings / Anthropic

Claude Haiku 4.5 20251001

released 2025-10-01

BenchAtlas Index

as of 2026-09-11
57.0
thinking
rank #182
4 families · 3 categories · medium
53.9
32K budget
rank #208
6 families · 3 categories · medium
51.7
thinking
rank #225
10 families · 4 categories · medium

Benchmark evidence

99 results
external indices
ECIlow effort142.4 pointsindependentEpoch AI Benchmarking Hub
ECI1K budget142.4 pointsindependentEpoch AI Benchmarking Hub
ECI16K budget142.4 pointsindependentEpoch AI Benchmarking Hub
ECI8K budget142.4 pointsindependentEpoch AI Benchmarking Hub
ECI32K budget142.4 pointsindependentEpoch AI Benchmarking Hub
ECI142.4 pointsindependentEpoch AI Benchmarking Hub
Vals Indexthinking22.9 pointsindependentVals AI
factuality
SimpleQA Verified13.2 %independentEpoch AI Benchmarking Hub
2026-08-10
SimpleQA Verified32K budget12.6 %independentEpoch AI Benchmarking Hub
2026-08-27
SimpleQA Verified32K budget5.9 %independentEpoch AI Benchmarking Hub
2025-12-09
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1480.6
1476.31484.9
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1464.0
1460.21467.7
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1453.1
1448.91457.3
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1453.0
1446.51459.6
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1440.6
1437.41443.8
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1436.8
1432.71440.8
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1433.1
1425.81440.5
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1433.0
1424.91441.2
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1430.5
1427.11434.0
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1429.6
1424.51434.6
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1428.2
1416.01440.3
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1427.0
1415.71438.3
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1425.2
1420.31430.1
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1424.5
1417.01432.0
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1418.6
1413.91423.3
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1415.6
1411.61419.6
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1414.4
1407.51421.3
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1413.2
1410.61415.7
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1405.2
1390.81419.6
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1400.7
1397.21404.2
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1400.3
1394.61406.0
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1399.7
1395.31404.1
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1398.6
1391.01406.2
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1393.2
1390.01396.4
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1389.4
1384.81394.0
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1388.7
1383.61393.8
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1368.8
1357.41380.3
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1363.3
1348.41378.1
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1357.5
1339.21375.7
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
thinking72.2 %independentVals AI
GPQA Diamond32K budget71.2 %independentEpoch AI Benchmarking Hub
2025-10-22
GPQA Diamond60.5 %independentEpoch AI Benchmarking Hub
2025-10-16
MMLU-Pro
implementation=vals-ai
thinking78.7 %independentVals AI
long context instruction
MultiChallengethinking50.5 %
48.552.5
independentScale Labs
multimodal
MMMU
implementation=vals-ai
thinking46.1 %independentVals AI
professional
CaseLawthinking56.5 %independentVals AI
CorpFinthinking60.6 %independentVals AI
CorpFin60.3 %independentVals AI
LegalBenchthinking81.2 %independentVals AI
MedQAthinking79.6 %independentVals AI
TaxEvalthinking67.5 %independentVals AI
reasoning math
AIME
implementation=vals-ai
thinking82.7 %independentVals AI
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 32K budget62.9 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 16K budget51.4 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 32K budget47.7 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 8K budget45.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 16K budget37.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 1K budget27.1 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
26.6 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 8K budget25.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 1K budget16.8 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
14.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 32K budget5.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 16K budget4.2 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 32K budget4.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 16K budget2.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 8K budget2.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 8K budget1.7 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
1.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 1K budget1.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 32K budget0.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 32K budget0.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 32K budget0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 32K budget0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 16K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 16K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 16K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 16K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 8K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 8K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 8K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 8K budget0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 1K budget0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
thinking · 1K budget0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
thinking · 1K budget0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
thinking · 1K budget0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
thinking · 1K budget0.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.0 %independentARC Prize Leaderboard
FrontierMath32K budget5.9 %independentEpoch AI Benchmarking Hub
2025-10-22
FrontierMath4.1 %independentEpoch AI Benchmarking Hub
2025-10-16
FrontierMath Tier 4older version32K budget2.1 %independentEpoch AI Benchmarking Hub
2025-10-22
MATH Level 532K budget96.4 %independentEpoch AI Benchmarking Hub
2025-10-22
MATH Level 586.9 %independentEpoch AI Benchmarking Hub
2025-10-16
OTIS Mock AIME 2024–202532K budget66.7 %independentEpoch AI Benchmarking Hub
2025-10-22
OTIS Mock AIME 2024–202535.8 %independentEpoch AI Benchmarking Hub
2025-10-16

Agent + model results

systems, not bare-model scores
agent + model mini-SWE-agent + Claude Haiku 4.5 20251001SWE-bench bash-only66.6 %unverifiedSWE-bench Leaderboard
agent + model mini-SWE-agent + Claude Haiku 4.5 20251001SWE-bench Verified66.6 %unverifiedSWE-bench Leaderboard
agent + model mini-SWE-agent + Claude Haiku 4.5 20251001SWE-bench Multilingual64.7 %communitySWE-bench Leaderboard
agent + model mini-SWE-agent + Claude Haiku 4.5 20251001SWE-bench Multilingual0.4 usd_per_taskcommunitySWE-bench Leaderboard
agent + model mini-SWE-agent + Claude Haiku 4.5 20251001SWE-bench Verified0.3 usd_per_taskunverifiedSWE-bench Leaderboard
agent + model mini-SWE-agent + Claude Haiku 4.5 20251001SWE-bench bash-only0.3 usd_per_taskunverifiedSWE-bench Leaderboard
agent + model mini-SWE-agent + Claude Haiku 4.5 20251001Terminal-Bench 2.029.8 %communityTerminal-Bench Leaderboard
agent + model terminus-2 + Claude Haiku 4.5 20251001Terminal-Bench 2.028.3 %communityTerminal-Bench Leaderboard
agent + model Claude Code + Claude Haiku 4.5 20251001Terminal-Bench 2.027.5 %communityTerminal-Bench Leaderboard
agent + model OpenHands + Claude Haiku 4.5 20251001Terminal-Bench 2.013.9 %communityTerminal-Bench Leaderboard

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants