BenchAtlas

Rankings / DeepSeek

Deepseek V4 Pro

released 2026-04-24

BenchAtlas Index

as of 2026-09-11
72.1
max effort
rank #60
8 families · 4 categories · medium
69.0
high effort
rank #78
8 families · 4 categories · medium
67.4
high effort
rank #90
4 families · 3 categories · medium

Benchmark evidence

80 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
42.6 %independentLiveBench
τ²-Benchmax effort96.2 %independentArtificial Analysis
τ²-Benchhigh effort94.2 %independentArtificial Analysis
τ²-Benchno reasoning91.2 %independentArtificial Analysis
τ²-Bench
subset=banking
max effort30.1 %independentArtificial Analysis
τ²-Bench
subset=banking
high effort26.2 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
70.0 %independentLiveBench
SciCodemax effort50.8 %independentArtificial Analysis
SciCodehigh effort46.4 %independentArtificial Analysis
SciCodeno reasoning42.4 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
74.5 %independentLiveBench
external indices
AA Coding Indexmax effort59.4 pointsindependentArtificial Analysis
AA Coding Indexhigh effort58.7 pointsindependentArtificial Analysis
Intelligence Index v4.1max effort30.9 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort30.1 pointsindependentArtificial Analysis
Intelligence Index v4.1no reasoning20.8 pointsindependentArtificial Analysis
ECI149.2 pointsindependentEpoch AI Benchmarking Hub
ECImax effort149.2 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort149.2 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning149.2 pointsindependentEpoch AI Benchmarking Hub
Vals Indexmax effort42.9 pointsindependentVals AI
factuality
SimpleQA Verifiedmax effort57.0 %independentEpoch AI Benchmarking Hub
2026-06-16
SimpleQA Verifiedmax effort47.0 %independentEpoch AI Benchmarking Hub
2026-08-27
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1501.8
1495.81507.9
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1493.8
1481.91505.6
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1492.1
1486.71497.5
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1484.3
1478.31490.3
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1484.3
1469.21499.4
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1481.6
1472.81490.4
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1480.8
1473.51488.2
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1480.3
1475.61485.0
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1477.2
1466.71487.7
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1475.9
1468.71483.1
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1473.5
1463.61483.3
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1472.2
1466.71477.7
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1468.2
1448.21488.1
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1467.9
1449.51486.3
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1467.2
1462.01472.4
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1461.9
1437.61486.2
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1461.6
1456.31466.8
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1461.1
1454.21468.0
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1460.9
1449.41472.4
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1457.9
1449.31466.5
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1457.8
1453.81461.8
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1453.4
1437.71469.2
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1453.2
1447.51459.0
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1451.1
1444.61457.6
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1445.2
1440.31450.0
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1445.1
1433.61456.7
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1444.5
1437.01452.0
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1433.9
1427.11440.7
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1417.4
1397.91436.9
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamondhigh effort90.9 %independentEpoch AI Benchmarking Hub
2026-08-06
GPQA Diamond
implementation=artificial-analysis
high effort90.5 %independentArtificial Analysis
GPQA Diamondmax effort89.7 %independentEpoch AI Benchmarking Hub
2026-06-16
GPQA Diamond
implementation=vals-ai
max effort89.4 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
max effort88.8 %independentArtificial Analysis
GPQA Diamondno reasoning73.2 %independentEpoch AI Benchmarking Hub
2026-08-06
GPQA Diamond
implementation=artificial-analysis
no reasoning71.7 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
max effort37.5 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
high effort35.2 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
no reasoning8.2 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
max effort87.2 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
78.1 %independentLiveBench
long context instruction
AA-LCRmax effort74.7 %independentArtificial Analysis
AA-LCRhigh effort70.3 %independentArtificial Analysis
AA-LCRno reasoning53.0 %independentArtificial Analysis
IFBenchmax effort76.5 %independentArtificial Analysis
IFBenchhigh effort71.3 %independentArtificial Analysis
IFBenchno reasoning45.8 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
62.4 %independentLiveBench
professional
CaseLawmax effort59.4 %independentVals AI
CorpFinmax effort61.4 %independentVals AI
LegalBenchmax effort80.3 %independentVals AI
TaxEvalmax effort72.1 %independentVals AI
reasoning math
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
90.7 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
82.7 %independentLiveBench
OTIS Mock AIME 2024–2025max effort96.7 %independentEpoch AI Benchmarking Hub
2026-06-17
OTIS Mock AIME 2024–2025high effort95.6 %independentEpoch AI Benchmarking Hub
2026-08-06
OTIS Mock AIME 2024–2025no reasoning46.7 %independentEpoch AI Benchmarking Hub
2026-08-06

Agent + model results

systems, not bare-model scores
agent + model Epoch Inspect harness + Deepseek V4 ProSWE-bench Verified77.6 %independentEpoch AI Benchmarking Hub
agent + model Artificial Analysis harness + Deepseek V4 ProTerminal-Bench 2.164.8 %independentArtificial Analysis
agent + model Artificial Analysis harness + Deepseek V4 ProTerminal-Bench 2.164.0 %independentArtificial Analysis
agent + model Artificial Analysis harness + Deepseek V4 ProTerminal-Bench Hard46.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + Deepseek V4 ProTerminal-Bench Hard41.7 %independentArtificial Analysis
agent + model Artificial Analysis harness + Deepseek V4 ProTerminal-Bench Hard36.4 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.