BenchAtlas

Rankings / Meta AI

Muse Spark 1.1

proprietary

released 2026-07-09

BenchAtlas Index

as of 2026-09-11
75.6
xhigh effort
rank #33
7 families · 4 categories · medium
74.8
xhigh effort
rank #40
11 families · 5 categories · high
63.5
high effort
rank #125
6 families · 4 categories · medium

Benchmark evidence

58 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort58.5 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort51.6 %independentLiveBench
τ²-Bench
subset=banking
xhigh effort31.8 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
high effort80.4 %independentLiveBench
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
xhigh effort77.2 %independentLiveBench
SciCodexhigh effort58.8 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort73.2 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort72.5 %independentLiveBench
external indices
AA Coding Indexxhigh effort71.3 pointsindependentArtificial Analysis
Intelligence Index v4.1xhigh effort34.3 pointsindependentArtificial Analysis
ECImedium effort154.7 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort154.7 pointsindependentEpoch AI Benchmarking Hub
ECI154.7 pointsindependentEpoch AI Benchmarking Hub
Vals Indexxhigh effort54.8 pointsindependentVals AI
factuality
SimpleQA Verified57.8 %independentEpoch AI Benchmarking Hub
2026-08-31
human preference
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1534.2
1518.71549.8
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
1531.0
1523.11538.8
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1524.7
1517.81531.6
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1519.7
1498.21541.3
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1511.3
1505.51517.2
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1507.0
1500.41513.6
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1506.6
1496.81516.4
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1506.2
1489.41523.1
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1505.5
1497.51513.4
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1505.0
1493.31516.8
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1501.3
1487.51515.0
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1500.5
1491.41509.7
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1493.9
1483.71504.1
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1493.8
1481.71506.0
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1492.1
1487.01497.1
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1491.6
1484.91498.3
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1490.3
1471.91508.7
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1489.6
1474.91504.3
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1487.1
1464.01510.2
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1484.1
1478.11490.1
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1478.9
1472.11485.7
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1470.3
1463.11477.5
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1463.6
1455.41471.8
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1450.7
1442.21459.2
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1445.4
1436.01454.8
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
xhigh effort91.2 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
xhigh effort89.8 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
xhigh effort46.2 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
xhigh effort88.7 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort74.7 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort74.3 %independentLiveBench
long context instruction
AA-LCRxhigh effort77.7 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort70.1 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort69.6 %independentLiveBench
MultiChallenge75.3 %
74.775.9
independentScale Labs
multimodal
MMMU
implementation=vals-ai
xhigh effort86.6 %independentVals AI
professional
CorpFinxhigh effort71.3 %independentVals AI
LegalBenchxhigh effort85.0 %independentVals AI
TaxEvalxhigh effort79.7 %independentVals AI
reasoning math
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort87.3 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort87.1 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort88.4 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort87.7 %independentLiveBench

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Muse Spark 1.1Terminal-Bench 2.177.9 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants