Rankings / Alibaba
Qwen3 6 27B
released 2026-04-22
BenchAtlas Index
as of 2026-09-11
thinking
rank #178
7 families · 4 categories · medium
no reasoning
rank #234
8 families · 5 categories · high
base configuration
rank #339
8 families · 5 categories · high
Benchmark evidence
34 results
agentic coding
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 39.3 % | independent | LiveBench | |
| τ²-Bench | thinking | 94.2 % | independent | Artificial Analysis |
| τ²-Bench | no reasoning | 93.6 % | independent | Artificial Analysis |
| τ²-Bench subset=banking | thinking | 16.7 % | independent | Artificial Analysis |
| τ²-Bench subset=banking | no reasoning | 9.3 % | independent | Artificial Analysis |
coding
| Release 2026-06-25 tasks_counted=2 · livebench_version=2026-06-25 | 71.8 % | independent | LiveBench | |
| SciCode | thinking | 42.8 % | independent | Artificial Analysis |
| SciCode | no reasoning | 37.3 % | independent | Artificial Analysis |
data analysis
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 70.4 % | independent | LiveBench |
external indices
| AA Coding Index | thinking | 53.7 points | independent | Artificial Analysis |
| AA Coding Index | no reasoning | 46.6 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | thinking | 21.9 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | no reasoning | 19.8 points | independent | Artificial Analysis |
| ECI | 146.5 points | independent | Epoch AI Benchmarking Hub | |
| ECI | no reasoning | 146.5 points | independent | Epoch AI Benchmarking Hub |
knowledge science
| GPQA Diamond | 85.9 % | independent | Epoch AI Benchmarking Hub 2026-08-07 | |
| GPQA Diamond | no reasoning | 84.8 % | independent | Epoch AI Benchmarking Hub 2026-08-07 |
| GPQA Diamond implementation=artificial-analysis | thinking | 84.2 % | independent | Artificial Analysis |
| GPQA Diamond implementation=artificial-analysis | no reasoning | 82.9 % | independent | Artificial Analysis |
| Humanity's Last Exam implementation=artificial-analysis | thinking | 23.1 % | independent | Artificial Analysis |
| Humanity's Last Exam implementation=artificial-analysis | no reasoning | 15.1 % | independent | Artificial Analysis |
language
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 63.3 % | independent | LiveBench |
long context instruction
| AA-LCR | thinking | 77.3 % | independent | Artificial Analysis |
| AA-LCR | no reasoning | 66.7 % | independent | Artificial Analysis |
| IFBench | thinking | 67.5 % | independent | Artificial Analysis |
| IFBench | no reasoning | 45.7 % | independent | Artificial Analysis |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 53.2 % | independent | LiveBench |
professional
| CaseLaw | 53.2 % | independent | Vals AI | |
| CorpFin | 62.3 % | independent | Vals AI | |
| TaxEval | 71.3 % | independent | Vals AI |
reasoning math
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 79.9 % | independent | LiveBench | |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 70.3 % | independent | LiveBench | |
| OTIS Mock AIME 2024–2025 | 91.1 % | independent | Epoch AI Benchmarking Hub 2026-08-07 | |
| OTIS Mock AIME 2024–2025 | no reasoning | 66.7 % | independent | Epoch AI Benchmarking Hub 2026-08-07 |
Agent + model results
systems, not bare-model scores
| agent + model Artificial Analysis harness + Qwen3 6 27B | Terminal-Bench 2.1 | 60.7 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + Qwen3 6 27B | Terminal-Bench 2.1 | 51.3 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + Qwen3 6 27B | Terminal-Bench Hard | 34.9 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + Qwen3 6 27B | Terminal-Bench Hard | 21.2 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
