BenchAtlas

Rankings / Anthropic

Claude Sonnet 4.6

released 2026-02-17

BenchAtlas Index

as of 2026-09-11
66.3
max effort
rank #100
8 families · 4 categories · medium
65.8
max effort
rank #105
8 families · 4 categories · medium
61.6
32K budget
rank #140
4 families · 3 categories · medium

Benchmark evidence

105 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
medium effort42.6 %independentLiveBench
τ²-Benchhigh effort79.5 %independentArtificial Analysis
τ²-Benchlow effort79.0 %independentArtificial Analysis
τ²-Benchmax effort75.7 %independentArtificial Analysis
τ²-Bench
subset=banking
max effort34.4 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
medium effort79.3 %independentLiveBench
SciCodemax effort50.1 %independentArtificial Analysis
SciCodehigh effort46.9 %independentArtificial Analysis
SciCodelow effort44.1 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
medium effort78.0 %independentLiveBench
external indices
AA Coding Indexmax effort63.0 pointsindependentArtificial Analysis
Intelligence Index v4.1max effort30.5 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort24.7 pointsindependentArtificial Analysis
Intelligence Index v4.1low effort23.3 pointsindependentArtificial Analysis
ECImedium effort152.2 pointsindependentEpoch AI Benchmarking Hub
ECI152.2 pointsindependentEpoch AI Benchmarking Hub
ECI16K budget152.2 pointsindependentEpoch AI Benchmarking Hub
ECI32K budget152.2 pointsindependentEpoch AI Benchmarking Hub
ECImax effort152.2 pointsindependentEpoch AI Benchmarking Hub
ECIlow effort152.2 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort152.2 pointsindependentEpoch AI Benchmarking Hub
Vals Indexmax effort50.6 pointsindependentVals AI
factuality
SimpleQA Verifiedhigh effort35.5 %independentEpoch AI Benchmarking Hub
2026-08-10
SimpleQA Verifiedmax effort32.8 %independentEpoch AI Benchmarking Hub
2026-08-27
SimpleQA Verified32K budget29.0 %independentEpoch AI Benchmarking Hub
2026-02-21
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1528.3
1522.61534.1
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1514.9
1509.81520.0
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1508.2
1499.91516.4
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1506.9
1501.31512.6
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1506.1
1495.71516.5
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1503.8
1499.41508.2
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1500.7
1493.91507.4
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1493.4
1488.21498.5
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1491.2
1482.01500.5
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1485.7
1476.71494.7
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1483.9
1473.21494.6
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1483.5
1477.21489.9
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1483.1
1478.31487.9
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1482.3
1477.51487.0
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1480.7
1474.11487.3
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1479.6
1464.81494.4
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1476.1
1470.81481.5
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1472.4
1468.81476.0
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1471.0
1454.81487.2
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1467.9
1453.61482.3
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1465.7
1458.01473.4
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1463.0
1452.61473.4
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1457.4
1453.01461.9
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1456.2
1450.31462.0
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1450.4
1432.01468.7
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1449.6
1442.81456.4
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1445.7
1439.51452.0
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1438.5
1415.31461.8
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1434.9
1415.81453.9
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=artificial-analysis
max effort87.5 %independentArtificial Analysis
GPQA Diamond32K budget87.4 %independentEpoch AI Benchmarking Hub
2026-02-20
GPQA Diamond
implementation=vals-ai
max effort85.6 %independentVals AI
GPQA Diamondhigh effort83.3 %independentEpoch AI Benchmarking Hub
2026-07-13
GPQA Diamondmedium effort83.3 %independentEpoch AI Benchmarking Hub
2026-07-13
GPQA Diamond
implementation=artificial-analysis
high effort79.9 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
low effort79.7 %independentArtificial Analysis
GPQA Diamondmax effort78.8 %independentEpoch AI Benchmarking Hub
2026-08-06
Humanity's Last Exam
implementation=artificial-analysis
max effort33.6 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
high effort13.3 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
low effort11.2 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
max effort87.3 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
medium effort76.1 %independentLiveBench
long context instruction
AA-LCRmax effort80.0 %independentArtificial Analysis
AA-LCRlow effort69.3 %independentArtificial Analysis
AA-LCRhigh effort68.3 %independentArtificial Analysis
IFBenchmax effort56.6 %independentArtificial Analysis
IFBenchlow effort42.4 %independentArtificial Analysis
IFBenchhigh effort41.2 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
medium effort63.2 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
max effort83.6 %independentVals AI
professional
CaseLaw64.0 %independentVals AI
CorpFinmax effort65.3 %independentVals AI
LegalBenchmax effort82.1 %independentVals AI
MedQAthinking92.1 %independentVals AI
TaxEvalmax effort77.1 %independentVals AI
reasoning math
AIME
implementation=vals-ai
92.3 %independentVals AI
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort95.8 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort95.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort86.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort86.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort65.7 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort62.4 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort60.4 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort58.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort3.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort2.9 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort2.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort2.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort1.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort1.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort1.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort0.8 usd_per_taskindependentARC Prize Leaderboard
FrontierMath16K budget32.4 %independentEpoch AI Benchmarking Hub
2026-02-20
FrontierMath Tier 4older version16K budget8.3 %independentEpoch AI Benchmarking Hub
2026-02-21
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
medium effort87.0 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
medium effort84.8 %independentLiveBench
OTIS Mock AIME 2024–202532K budget85.8 %independentEpoch AI Benchmarking Hub
2026-02-20
OTIS Mock AIME 2024–2025medium effort82.2 %independentEpoch AI Benchmarking Hub
2026-07-13
OTIS Mock AIME 2024–2025high effort75.6 %independentEpoch AI Benchmarking Hub
2026-07-13
OTIS Mock AIME 2024–2025max effort71.1 %independentEpoch AI Benchmarking Hub
2026-08-06

Agent + model results

systems, not bare-model scores
agent + model Epoch Inspect harness + Claude Sonnet 4.6SWE-bench Verified75.2 %independentEpoch AI Benchmarking Hub
agent + model Artificial Analysis harness + Claude Sonnet 4.6Terminal-Bench 2.171.2 %independentArtificial Analysis
agent + model terminal-agent + Claude Sonnet 4.6Terminal-Bench 2.053.4 %unverifiedTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + Claude Sonnet 4.6Terminal-Bench Hard53.0 %independentArtificial Analysis
agent + model Artificial Analysis harness + Claude Sonnet 4.6Terminal-Bench Hard46.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + Claude Sonnet 4.6Terminal-Bench Hard42.4 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants