BenchAtlas

Rankings / Alibaba

Qwen3 7

released 2026-05-19

BenchAtlas Index

as of 2026-09-11
Not enough benchmark coverage for a composite score (needs ≥4 families across ≥3 categories). The per-benchmark evidence below stands on its own — sparse coverage is an evidence gap, not a low score.

Benchmark evidence

7 results
external indices
ECImax effort153.5 pointsindependentEpoch AI Benchmarking Hub
factuality
SimpleQA Verifiedmax effort58.5 %independentEpoch AI Benchmarking Hub
2026-06-13
SimpleQA Verifiedmax effort55.8 %independentEpoch AI Benchmarking Hub
2026-08-27
knowledge science
GPQA Diamondmax effort91.6 %independentEpoch AI Benchmarking Hub
2026-06-12
GPQA Diamondmax effort90.9 %independentEpoch AI Benchmarking Hub
2026-08-07
reasoning math
OTIS Mock AIME 2024–2025max effort95.6 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–2025max effort95.0 %independentEpoch AI Benchmarking Hub
2026-06-13

Agent + model results

systems, not bare-model scores
agent + model Epoch Inspect harness + Qwen3 7SWE-bench Verified77.3 %independentEpoch AI Benchmarking Hub

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants