BenchAtlas

Rankings / Alibaba

Qwen3 235B A22b 2507

released 2025-07-25

BenchAtlas Index

as of 2026-09-11
41.5
thinking
rank #302
11 families · 5 categories · high

Benchmark evidence

20 results
agentic coding
τ²-Benchthinking53.2 %independentArtificial Analysis
τ²-Bench
subset=banking
thinking7.8 %independentArtificial Analysis
coding
LiveCodeBench v6
implementation=artificial-analysis
thinking78.8 %independentArtificial Analysis
SciCodethinking41.4 %independentArtificial Analysis
external indices
AA Coding Indexthinking22.1 pointsindependentArtificial Analysis
Intelligence Index v4.1thinking12.7 pointsindependentArtificial Analysis
AA Math Indexthinking91.0 pointsindependentArtificial Analysis
ECIthinking143.9 pointsindependentEpoch AI Benchmarking Hub
factuality
SimpleQA Verifiedthinking50.1 %independentEpoch AI Benchmarking Hub
2025-12-11
SimpleQA Verifiedthinking40.4 %independentEpoch AI Benchmarking Hub
2026-08-27
knowledge science
GPQA Diamondthinking80.0 %independentEpoch AI Benchmarking Hub
2025-12-11
GPQA Diamond
implementation=artificial-analysis
thinking79.0 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
thinking15.9 %independentArtificial Analysis
MMLU-Pro
implementation=artificial-analysis
thinking84.3 %independentArtificial Analysis
long context instruction
AA-LCRthinking72.0 %independentArtificial Analysis
IFBenchthinking51.2 %independentArtificial Analysis
reasoning math
AIME
implementation=artificial-analysis
thinking94.0 %independentArtificial Analysis
AIME
year=2025 · implementation=artificial-analysis
thinking91.0 %independentArtificial Analysis
MATH-500
implementation=artificial-analysis
thinking98.4 %independentArtificial Analysis
OTIS Mock AIME 2024–2025thinking86.7 %independentEpoch AI Benchmarking Hub
2025-12-10

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Qwen3 235B A22b 2507Terminal-Bench Hard13.6 %independentArtificial Analysis
agent + model Artificial Analysis harness + Qwen3 235B A22b 2507Terminal-Bench 2.112.0 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants