BenchAtlas

Rankings / Alibaba (Qwen)

Qwen3 235B A22b Instruct 2507

released 2025-07-25

BenchAtlas Index

as of 2026-09-11
34.2
base configuration
rank #351
11 families · 5 categories · high

Benchmark evidence

51 results
agentic coding
τ²-Bench33.3 %independentArtificial Analysis
coding
LiveCodeBench v6
implementation=artificial-analysis
52.4 %independentArtificial Analysis
SciCode36.0 %independentArtificial Analysis
external indices
Intelligence Index v4.112.0 pointsindependentArtificial Analysis
AA Math Index71.7 pointsindependentArtificial Analysis
ECI138.9 pointsindependentEpoch AI Benchmarking Hub
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1472.3
1467.81476.9
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1468.6
1461.01476.1
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1463.8
1460.01467.6
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1455.9
1447.61464.1
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1454.5
1440.51468.5
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1450.7
1446.21455.2
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1449.1
1441.21457.0
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1447.9
1444.61451.2
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1445.9
1440.71451.1
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1437.2
1432.11442.3
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1434.1
1429.81438.4
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1433.6
1425.91441.2
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1431.8
1426.91436.7
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1430.7
1422.11439.4
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1429.9
1426.31433.4
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1426.3
1414.31438.4
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1424.5
1410.01439.0
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1423.1
1420.51425.7
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1417.7
1411.51423.8
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1417.5
1409.61425.4
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1415.5
1411.41419.7
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1415.1
1411.41418.7
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1411.2
1408.01414.4
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1403.7
1393.81413.5
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1396.3
1391.81400.8
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1393.2
1376.71409.7
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1383.4
1368.31398.6
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1379.3
1373.81384.9
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1376.5
1371.61381.5
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=artificial-analysis
75.3 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
11.1 %independentArtificial Analysis
MMLU-Pro
implementation=artificial-analysis
82.8 %independentArtificial Analysis
long context instruction
AA-LCR33.9 %independentArtificial Analysis
IFBench46.0 %independentArtificial Analysis
reasoning math
AIME
year=2025 · implementation=artificial-analysis
71.7 %independentArtificial Analysis
AIME
implementation=artificial-analysis
71.7 %independentArtificial Analysis
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
17.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
11.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
1.3 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=Base LLM
0.0 usd_per_taskindependentARC Prize Leaderboard
MATH-500
implementation=artificial-analysis
98.0 %independentArtificial Analysis

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Qwen3 235B A22b Instruct 2507Terminal-Bench Hard15.2 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.