BenchAtlas

Rankings / Alibaba

Qwen3 6 Plus

released 2026-03-31

BenchAtlas Index

as of 2026-09-11
49.6
base configuration
rank #243
9 families · 4 categories · medium
45.5
base configuration
rank #274
15 families · 5 categories · high

Benchmark evidence

64 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
41.4 %independentLiveBench
τ²-Bench97.7 %independentArtificial Analysis
τ²-Bench
subset=banking
20.8 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
78.2 %independentLiveBench
SciCode40.7 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
69.9 %independentLiveBench
external indices
AA Coding Index54.5 pointsindependentArtificial Analysis
Intelligence Index v4.127.0 pointsindependentArtificial Analysis
ECI147.7 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning147.7 pointsindependentEpoch AI Benchmarking Hub
Vals Index32.0 pointsindependentVals AI
factuality
SimpleQA Verified49.1 %independentEpoch AI Benchmarking Hub
2026-05-12
SimpleQA Verified44.1 %independentEpoch AI Benchmarking Hub
2026-08-27
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1495.0
1488.61501.4
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1482.8
1477.11488.5
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1477.3
1471.11483.6
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1477.2
1464.11490.4
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1476.7
1467.41486.1
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1467.7
1462.81472.7
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1463.5
1447.41479.7
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1460.5
1449.31471.7
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1457.0
1445.01469.0
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1456.4
1450.91461.8
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1455.2
1447.41463.0
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1454.9
1432.91477.0
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1454.8
1442.91466.8
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1454.5
1448.71460.4
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1447.7
1437.01458.4
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1446.6
1438.91454.2
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1446.2
1439.01453.4
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1443.7
1439.51448.0
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1442.2
1436.71447.7
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1438.9
1429.81448.0
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1436.5
1430.51442.6
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1433.2
1414.11452.3
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1432.9
1416.81449.0
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1427.4
1422.31432.5
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1423.7
1416.81430.6
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1407.9
1399.71416.0
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1407.4
1400.11414.6
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1389.6
1362.11417.1
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1383.9
1362.21405.6
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond88.4 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
88.2 %independentArtificial Analysis
GPQA Diamond
implementation=vals-ai
87.4 %independentVals AI
GPQA Diamond87.4 %independentEpoch AI Benchmarking Hub
2026-05-12
Humanity's Last Exam
implementation=artificial-analysis
27.8 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
87.7 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
75.0 %independentLiveBench
long context instruction
AA-LCR78.3 %independentArtificial Analysis
IFBench75.2 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
58.3 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
84.2 %independentVals AI
professional
CaseLaw51.4 %independentVals AI
CorpFin61.9 %independentVals AI
LegalBench84.2 %independentVals AI
TaxEval74.7 %independentVals AI
reasoning math
AIME
implementation=vals-ai
94.6 %independentVals AI
FrontierMath26.2 %independentEpoch AI Benchmarking Hub
2026-05-12
FrontierMath Tier 4older version8.3 %independentEpoch AI Benchmarking Hub
2026-05-12
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
83.7 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
75.8 %independentLiveBench
OTIS Mock AIME 2024–202593.3 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–202590.6 %independentEpoch AI Benchmarking Hub
2026-05-12

Agent + model results

systems, not bare-model scores
agent + model Epoch Inspect harness + Qwen3 6 PlusSWE-bench Verified57.9 %independentEpoch AI Benchmarking Hub
agent + model Artificial Analysis harness + Qwen3 6 PlusTerminal-Bench 2.161.4 %independentArtificial Analysis
agent + model Artificial Analysis harness + Qwen3 6 PlusTerminal-Bench Hard43.9 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.