BenchAtlas

Rankings / xAI

Grok 4.5

proprietary

released 2026-07-08

BenchAtlas Index

as of 2026-09-11
78.0
high effort
rank #23
7 families · 5 categories · medium
74.5
high effort
rank #42
7 families · 4 categories · medium
60.9
base configuration
rank #147
6 families · 4 categories · medium

Benchmark evidence

54 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
56.5 %independentLiveBench
τ²-Bench
subset=banking
high effort42.1 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
68.6 %independentLiveBench
SciCodehigh effort55.0 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
73.0 %independentLiveBench
external indices
AA Coding Indexhigh effort72.4 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort39.1 pointsindependentArtificial Analysis
ECIlow effort153.9 pointsindependentEpoch AI Benchmarking Hub
ECImedium effort153.9 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort153.9 pointsindependentEpoch AI Benchmarking Hub
ECI153.9 pointsindependentEpoch AI Benchmarking Hub
Vals Indexhigh effort51.5 pointsindependentVals AI
factuality
SimpleQA Verifiedhigh effort53.5 %independentEpoch AI Benchmarking Hub
2026-07-08
SimpleQA Verifiedhigh effort48.3 %independentEpoch AI Benchmarking Hub
2026-08-27
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1522.8
1515.11530.6
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1511.8
1496.91526.8
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1509.0
1502.31515.6
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1497.5
1489.71505.2
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1496.7
1485.31508.2
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1494.9
1489.21500.6
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1493.2
1476.31510.0
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1490.7
1481.21500.2
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1487.1
1480.61493.7
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1479.2
1472.81485.6
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1478.8
1460.71496.8
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1477.2
1463.91490.4
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1476.3
1454.81497.8
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1476.3
1469.71482.9
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1475.7
1466.21485.1
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1473.7
1459.41488.0
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1471.7
1460.01483.3
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1471.1
1466.21476.0
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1466.3
1457.41475.2
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1466.2
1459.21473.3
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1461.3
1437.91484.7
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1460.3
1454.41466.1
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1459.8
1451.81467.8
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1451.3
1442.01460.5
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1446.4
1438.21454.7
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamondhigh effort93.4 %independentEpoch AI Benchmarking Hub
2026-07-08
GPQA Diamond
implementation=artificial-analysis
high effort93.1 %independentArtificial Analysis
GPQA Diamond
implementation=vals-ai
high effort92.9 %independentVals AI
Humanity's Last Exam
implementation=artificial-analysis
high effort42.7 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
high effort89.2 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
82.8 %independentLiveBench
long context instruction
AA-LCRhigh effort79.3 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
71.5 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
high effort61.8 %independentVals AI
professional
CorpFinhigh effort67.4 %independentVals AI
LegalBenchhigh effort86.0 %independentVals AI
TaxEvalhigh effort71.7 %independentVals AI
reasoning math
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
90.8 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
87.2 %independentLiveBench
OTIS Mock AIME 2024–2025high effort97.8 %independentEpoch AI Benchmarking Hub
2026-07-08

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Grok 4.5Terminal-Bench 2.181.7 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants