BenchAtlas

Rankings / Anthropic

Claude Opus 4.8

released 2026-05-28

BenchAtlas Index

as of 2026-09-11
80.3
max effort
rank #12
7 families · 4 categories · medium
79.9
max effort
rank #13
7 families · 4 categories · medium
72.9
low effort
rank #54
4 families · 3 categories · medium

Benchmark evidence

93 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort56.1 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort50.5 %independentLiveBench
τ²-Benchmax effort94.4 %independentArtificial Analysis
τ²-Bench
subset=banking
max effort34.2 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
max effort81.8 %independentLiveBench
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
xhigh effort79.3 %independentLiveBench
SciCodemax effort54.4 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort78.3 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort66.0 %independentLiveBench
external indices
AA Coding Indexmax effort74.3 pointsindependentArtificial Analysis
Intelligence Index v4.1max effort42.0 pointsindependentArtificial Analysis
ECImax effort158.2 pointsindependentEpoch AI Benchmarking Hub
ECI24K budget158.2 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning158.2 pointsindependentEpoch AI Benchmarking Hub
ECIlow effort158.2 pointsindependentEpoch AI Benchmarking Hub
ECI158.2 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort158.2 pointsindependentEpoch AI Benchmarking Hub
ECIxhigh effort158.2 pointsindependentEpoch AI Benchmarking Hub
ECImedium effort158.2 pointsindependentEpoch AI Benchmarking Hub
Vals Indexmax effort60.9 pointsindependentVals AI
factuality
SimpleQA Verifiedmax effort53.0 %independentEpoch AI Benchmarking Hub
2026-08-27
SimpleQA Verifiedmax effort39.5 %independentEpoch AI Benchmarking Hub
2026-05-29
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1527.3
1520.91533.8
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1523.4
1511.51535.3
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1515.1
1509.41520.8
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1514.5
1505.61523.4
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1504.4
1488.71520.1
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1504.4
1498.11510.6
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1503.2
1498.11508.2
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1499.5
1493.71505.2
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1499.2
1491.71506.6
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1497.4
1486.71508.1
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1495.3
1487.81502.7
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1494.6
1484.41504.7
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1488.3
1481.21495.4
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1487.1
1474.91499.4
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1486.8
1477.71495.9
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1482.8
1463.21502.4
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1482.4
1476.81488.0
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1479.6
1474.11485.1
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1476.9
1470.91482.8
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1475.6
1454.81496.5
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1474.0
1461.51486.5
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1472.8
1468.61477.1
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1469.0
1462.41475.5
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1468.0
1451.51484.5
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1465.3
1457.91472.7
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1462.9
1457.71468.1
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1455.5
1448.71462.3
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1448.5
1426.71470.4
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1448.0
1423.01473.0
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
max effort92.4 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
max effort92.0 %independentArtificial Analysis
GPQA Diamondmax effort91.0 %independentEpoch AI Benchmarking Hub
2026-06-07
GPQA Diamondlow effort88.4 %independentEpoch AI Benchmarking Hub
2026-08-06
GPQA Diamondno reasoning85.3 %independentEpoch AI Benchmarking Hub
2026-08-06
Humanity's Last Exam
implementation=artificial-analysis
max effort48.7 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
max effort89.6 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort81.4 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort79.7 %independentLiveBench
long context instruction
AA-LCRmax effort77.7 %independentArtificial Analysis
IFBenchmax effort62.2 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort72.4 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort72.0 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
max effort86.6 %independentVals AI
professional
CorpFinmax effort66.7 %independentVals AI
LegalBenchmax effort83.6 %independentVals AI
TaxEvalmax effort75.6 %independentVals AI
reasoning math
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort92.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort92.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort91.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort88.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort72.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort71.7 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort62.2 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort2.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort2.4 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort2.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort1.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
high effort1.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort1.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort0.9 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort0.7 usd_per_taskindependentARC Prize Leaderboard
EnigmaEvalxhigh effort23.5 %
21.125.9
independentScale Labs
FrontierMathmax effort47.2 %independentEpoch AI Benchmarking Hub
2026-06-08
FrontierMath Tier 4older versionmax effort31.3 %independentEpoch AI Benchmarking Hub
2026-06-08
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort95.3 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort94.3 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort89.7 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort89.2 %independentLiveBench
OTIS Mock AIME 2024–2025max effort98.3 %independentEpoch AI Benchmarking Hub
2026-06-07
OTIS Mock AIME 2024–2025low effort97.8 %independentEpoch AI Benchmarking Hub
2026-08-06
OTIS Mock AIME 2024–2025no reasoning84.4 %independentEpoch AI Benchmarking Hub
2026-08-06

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Claude Opus 4.8Terminal-Bench 2.184.6 %independentArtificial Analysis
agent + model Claude Code + Claude Opus 4.8Terminal-Bench 2.178.9 %communityTerminal-Bench Leaderboard
agent + model terminus-2 + Claude Opus 4.8Terminal-Bench 2.174.6 %communityTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + Claude Opus 4.8Terminal-Bench Hard58.3 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants