Rankings / Mistral AI
Magistral Small 2509
released 2025-09-17
BenchAtlas Index
as of 2026-09-11
base configuration
rank #405
10 families · 5 categories · high
base configuration
rank #436
8 families · 4 categories · medium
Benchmark evidence
24 results
agentic coding
| τ²-Bench | 27.8 % | independent | Artificial Analysis | |
| τ²-Bench subset=banking | 4.5 % | independent | Artificial Analysis |
coding
| LiveCodeBench v6 implementation=artificial-analysis | 72.3 % | independent | Artificial Analysis | |
| SciCode | 35.2 % | independent | Artificial Analysis |
external indices
| AA Coding Index | 14.7 points | independent | Artificial Analysis | |
| Intelligence Index v4.1 | 8.6 points | independent | Artificial Analysis | |
| AA Math Index | 80.3 points | independent | Artificial Analysis | |
| ECI | 131.4 points | independent | Epoch AI Benchmarking Hub |
knowledge science
| GPQA Diamond implementation=artificial-analysis | 66.3 % | independent | Artificial Analysis | |
| GPQA Diamond implementation=vals-ai | 58.3 % | independent | Vals AI | |
| GPQA Diamond | 47.6 % | independent | Epoch AI Benchmarking Hub 2026-08-30 | |
| Humanity's Last Exam implementation=artificial-analysis | 6.4 % | independent | Artificial Analysis | |
| MMLU-Pro implementation=artificial-analysis | 76.8 % | independent | Artificial Analysis | |
| MMLU-Pro implementation=vals-ai | 62.1 % | independent | Vals AI |
long context instruction
| AA-LCR | 19.3 % | independent | Artificial Analysis | |
| IFBench | 44.4 % | independent | Artificial Analysis |
multimodal
| MMMU implementation=vals-ai | 65.2 % | independent | Vals AI |
professional
| CorpFin | 44.0 % | independent | Vals AI | |
| LegalBench | 40.0 % | independent | Vals AI | |
| MedQA | 82.4 % | independent | Vals AI | |
| TaxEval | 60.3 % | independent | Vals AI |
reasoning math
| AIME implementation=vals-ai | 80.7 % | independent | Vals AI | |
| AIME year=2025 · implementation=artificial-analysis | 80.3 % | independent | Artificial Analysis | |
| OTIS Mock AIME 2024–2025 | 28.1 % | independent | Epoch AI Benchmarking Hub 2026-08-30 |
Agent + model results
systems, not bare-model scores
| agent + model Artificial Analysis harness + Magistral Small 2509 | Terminal-Bench Hard | 4.5 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + Magistral Small 2509 | Terminal-Bench 2.1 | 4.5 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
