BenchAtlas

Rankings / xiaomi

Mimo V2 5 Pro

open weights

released 2026-04-22

BenchAtlas Index

as of 2026-09-11
68.6
base configuration
rank #82
7 families · 4 categories · medium
60.4
base configuration
rank #153
6 families · 3 categories · medium
53.5
no reasoning
rank #211
7 families · 4 categories · medium

Benchmark evidence

52 results
agentic coding
τ²-Bench94.2 %independentArtificial Analysis
τ²-Benchno reasoning72.5 %independentArtificial Analysis
τ²-Bench
subset=banking
9.9 %independentArtificial Analysis
coding
SciCode50.6 %independentArtificial Analysis
SciCodeno reasoning39.1 %independentArtificial Analysis
external indices
AA Coding Index60.2 pointsindependentArtificial Analysis
Intelligence Index v4.126.4 pointsindependentArtificial Analysis
Intelligence Index v4.1no reasoning18.3 pointsindependentArtificial Analysis
ECI149.3 pointsindependentEpoch AI Benchmarking Hub
Vals Index41.0 pointsindependentVals AI
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1520.5
1514.31526.6
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1512.0
1500.71523.3
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1508.2
1502.81513.5
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1505.2
1496.41514.0
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1503.8
1497.81509.7
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1496.1
1491.51500.7
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1490.5
1475.61505.5
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1490.1
1483.11497.2
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1490.1
1478.71501.6
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1487.1
1476.91497.3
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1486.3
1480.91491.7
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1481.5
1476.41486.6
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1478.2
1471.11485.2
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1476.7
1464.91488.5
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1475.0
1465.51484.5
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1473.5
1468.51478.6
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1471.7
1465.01478.4
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1470.9
1465.31476.6
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1469.8
1451.71487.9
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1468.1
1464.21472.0
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1465.1
1444.71485.5
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1461.8
1453.71470.0
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1459.0
1443.41474.5
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1452.4
1446.31458.6
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1451.4
1446.71456.1
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1440.0
1433.51446.5
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1437.3
1417.71456.9
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1435.5
1428.51442.6
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1421.5
1397.41445.6
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=artificial-analysis
86.6 %independentArtificial Analysis
GPQA Diamond
implementation=vals-ai
82.6 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
no reasoning76.2 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
35.7 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
no reasoning14.8 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
84.6 %independentVals AI
long context instruction
AA-LCR79.7 %independentArtificial Analysis
AA-LCRno reasoning41.7 %independentArtificial Analysis
IFBench79.9 %independentArtificial Analysis
IFBenchno reasoning42.7 %independentArtificial Analysis
professional
CorpFin61.4 %independentVals AI
LegalBench77.1 %independentVals AI
TaxEval73.8 %independentVals AI

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Mimo V2 5 ProTerminal-Bench 2.165.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + Mimo V2 5 ProTerminal-Bench Hard43.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + Mimo V2 5 ProTerminal-Bench Hard35.6 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.