Rankings / ant-group
Ling Flash 2.0
open weightsreleased 2025-09-17
BenchAtlas Index
as of 2026-09-11
base configuration
rank #392
9 families · 5 categories · high
Benchmark evidence
27 results
agentic coding
| τ²-Bench | 20.8 % | independent | Artificial Analysis |
coding
| LiveCodeBench v6 implementation=artificial-analysis | 58.9 % | independent | Artificial Analysis | |
| SciCode | 28.9 % | independent | Artificial Analysis |
external indices
| Intelligence Index v4.1 | 7.8 points | independent | Artificial Analysis | |
| AA Math Index | 65.3 points | independent | Artificial Analysis |
human preference
| Coding (style control)older version arena=text · category=coding · style_control=true | 1410.7 1395.9–1425.5 | community | LMArena Leaderboard Dataset | |
| Industry Software And It Services (style control)older version arena=text · category=industry_software_and_it_services · style_control=true | 1404.0 1392.3–1415.7 | community | LMArena Leaderboard Dataset | |
| Hard Prompts English (style control)older version arena=text · category=hard_prompts_english · style_control=true | 1393.2 1378.6–1407.8 | community | LMArena Leaderboard Dataset | |
| English (style control)older version arena=text · category=english · style_control=true | 1373.2 1362.6–1383.8 | community | LMArena Leaderboard Dataset | |
| Hard Prompts (style control)older version arena=text · category=hard_prompts · style_control=true | 1365.2 1355.2–1375.2 | community | LMArena Leaderboard Dataset | |
| Industry Life And Physical And Social Science (style control)older version arena=text · category=industry_life_and_physical_and_social_science · style_control=true | 1361.0 1342.5–1379.5 | community | LMArena Leaderboard Dataset | |
| Industry Business And Management And Financial Operations (style control)older version arena=text · category=industry_business_and_management_and_financial_operations · style_control=true | 1347.1 1330.8–1363.3 | community | LMArena Leaderboard Dataset | |
| Overall (style control) arena=text · category=overall · style_control=true | 1344.2 1336.9–1351.5 | community | LMArena Leaderboard Dataset | |
| Longer Query (style control)older version arena=text · category=longer_query · style_control=true | 1324.5 1308.3–1340.8 | community | LMArena Leaderboard Dataset | |
| Instruction Following (style control)older version arena=text · category=instruction_following · style_control=true | 1315.5 1301.7–1329.3 | community | LMArena Leaderboard Dataset | |
| Multi Turn (style control)older version arena=text · category=multi_turn · style_control=true | 1313.8 1296.1–1331.4 | community | LMArena Leaderboard Dataset | |
| Non English (style control)older version arena=text · category=non_english · style_control=true | 1313.7 1304.2–1323.2 | community | LMArena Leaderboard Dataset | |
| Exclude Ties (style control)older version arena=text · category=exclude_ties · style_control=true | 1299.3 1288.8–1309.7 | community | LMArena Leaderboard Dataset | |
| Industry Writing And Literature And Language (style control)older version arena=text · category=industry_writing_and_literature_and_language · style_control=true | 1276.6 1260.9–1292.2 | community | LMArena Leaderboard Dataset | |
| Industry Entertainment And Sports And Media (style control)older version arena=text · category=industry_entertainment_and_sports_and_media · style_control=true | 1272.1 1254.9–1289.2 | community | LMArena Leaderboard Dataset | |
| Creative Writing (style control)older version arena=text · category=creative_writing · style_control=true | 1267.3 1246.7–1288.0 | community | LMArena Leaderboard Dataset |
knowledge science
| GPQA Diamond implementation=artificial-analysis | 65.7 % | independent | Artificial Analysis | |
| Humanity's Last Exam implementation=artificial-analysis | 6.2 % | independent | Artificial Analysis | |
| MMLU-Pro implementation=artificial-analysis | 77.7 % | independent | Artificial Analysis |
long context instruction
| AA-LCR | 17.7 % | independent | Artificial Analysis | |
| IFBench | 34.4 % | independent | Artificial Analysis |
reasoning math
| AIME year=2025 · implementation=artificial-analysis | 65.3 % | independent | Artificial Analysis |
Agent + model results
systems, not bare-model scores
| agent + model Artificial Analysis harness + Ling Flash 2.0 | Terminal-Bench Hard | 10.6 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
