BenchAtlas

Rankings / Alibaba

Qwen3 6 27B

released 2026-04-22

BenchAtlas Index

as of 2026-09-11
57.6
thinking
rank #178
7 families · 4 categories · medium
50.5
no reasoning
rank #234
8 families · 5 categories · high
35.3
base configuration
rank #339
8 families · 5 categories · high

Benchmark evidence

34 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
39.3 %independentLiveBench
τ²-Benchthinking94.2 %independentArtificial Analysis
τ²-Benchno reasoning93.6 %independentArtificial Analysis
τ²-Bench
subset=banking
thinking16.7 %independentArtificial Analysis
τ²-Bench
subset=banking
no reasoning9.3 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
71.8 %independentLiveBench
SciCodethinking42.8 %independentArtificial Analysis
SciCodeno reasoning37.3 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
70.4 %independentLiveBench
external indices
AA Coding Indexthinking53.7 pointsindependentArtificial Analysis
AA Coding Indexno reasoning46.6 pointsindependentArtificial Analysis
Intelligence Index v4.1thinking21.9 pointsindependentArtificial Analysis
Intelligence Index v4.1no reasoning19.8 pointsindependentArtificial Analysis
ECI146.5 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning146.5 pointsindependentEpoch AI Benchmarking Hub
knowledge science
GPQA Diamond85.9 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamondno reasoning84.8 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
thinking84.2 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
no reasoning82.9 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
thinking23.1 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
no reasoning15.1 %independentArtificial Analysis
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
63.3 %independentLiveBench
long context instruction
AA-LCRthinking77.3 %independentArtificial Analysis
AA-LCRno reasoning66.7 %independentArtificial Analysis
IFBenchthinking67.5 %independentArtificial Analysis
IFBenchno reasoning45.7 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
53.2 %independentLiveBench
professional
CaseLaw53.2 %independentVals AI
CorpFin62.3 %independentVals AI
TaxEval71.3 %independentVals AI
reasoning math
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
79.9 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
70.3 %independentLiveBench
OTIS Mock AIME 2024–202591.1 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–2025no reasoning66.7 %independentEpoch AI Benchmarking Hub
2026-08-07

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Qwen3 6 27BTerminal-Bench 2.160.7 %independentArtificial Analysis
agent + model Artificial Analysis harness + Qwen3 6 27BTerminal-Bench 2.151.3 %independentArtificial Analysis
agent + model Artificial Analysis harness + Qwen3 6 27BTerminal-Bench Hard34.9 %independentArtificial Analysis
agent + model Artificial Analysis harness + Qwen3 6 27BTerminal-Bench Hard21.2 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants