BenchAtlas

Rankings / Meta AI

Muse Spark

released 2026-04-08

BenchAtlas Index

as of 2026-09-11
71.2
base configuration
rank #66
11 families · 5 categories · high
65.4
base configuration
rank #110
8 families · 4 categories · medium

Benchmark evidence

49 results
agentic coding
τ²-Bench91.5 %independentArtificial Analysis
τ²-Bench
subset=banking
19.6 %independentArtificial Analysis
coding
SciCode51.5 %independentArtificial Analysis
external indices
AA Coding Index58.6 pointsindependentArtificial Analysis
Intelligence Index v4.131.3 pointsindependentArtificial Analysis
ECI152.1 pointsindependentEpoch AI Benchmarking Hub
factuality
SimpleQA Verified66.3 %independentEpoch AI Benchmarking Hub
2026-04-08
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1525.9
1515.81536.0
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1521.1
1499.71542.5
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1520.0
1511.41528.5
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1507.6
1498.21517.0
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1504.7
1497.71511.7
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1503.7
1484.11523.4
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1503.0
1483.21522.8
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1500.9
1493.31508.6
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1496.9
1484.01509.7
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1495.5
1487.71503.3
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1491.5
1478.01504.9
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1490.1
1478.31501.9
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1489.6
1472.91506.3
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1488.2
1482.31494.0
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1481.5
1454.21508.9
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1481.4
1465.61497.2
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1477.0
1469.41484.5
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1474.3
1465.51483.1
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1469.3
1448.51490.1
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1464.0
1450.01478.0
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1463.4
1454.21472.7
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1462.2
1449.91474.5
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1461.1
1441.21481.0
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1459.9
1448.91470.9
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond89.8 %independentEpoch AI Benchmarking Hub
2026-04-08
GPQA Diamond
implementation=vals-ai
89.6 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
88.4 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
40.7 %independentArtificial Analysis
Humanity's Last Exam
contamination=Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. · implementation=scale
40.6 %
38.642.5
independentScale Labs
MMLU-Pro
implementation=vals-ai
87.3 %independentVals AI
long context instruction
AA-LCR78.0 %independentArtificial Analysis
IFBench75.9 %independentArtificial Analysis
MultiChallenge75.5 %
71.579.6
independentScale Labs
multimodal
MMMU
implementation=vals-ai
87.4 %independentVals AI
professional
CaseLaw63.1 %independentVals AI
CorpFin65.1 %independentVals AI
LegalBench84.2 %independentVals AI
TaxEval77.7 %independentVals AI
reasoning math
AIME
implementation=vals-ai
96.9 %independentVals AI
FrontierMath39.0 %independentEpoch AI Benchmarking Hub
2026-04-08
FrontierMath Tier 4older version14.6 %independentEpoch AI Benchmarking Hub
2026-04-08
OTIS Mock AIME 2024–202588.9 %independentEpoch AI Benchmarking Hub
2026-04-08

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Muse SparkTerminal-Bench 2.162.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + Muse SparkTerminal-Bench Hard45.5 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants