BenchAtlas

Rankings / OpenAI

GPT OSS 20B

open weights

released 2025-08-05

BenchAtlas Index

as of 2026-09-11
35.9
high effort
rank #333
10 families · 5 categories · high
31.3
base configuration
rank #367
8 families · 4 categories · medium
28.2
low effort
rank #393
10 families · 5 categories · high

Benchmark evidence

69 results
agentic coding
τ²-Benchhigh effort60.2 %independentArtificial Analysis
τ²-Benchlow effort50.3 %independentArtificial Analysis
τ²-Bench
subset=banking
high effort7.0 %independentArtificial Analysis
coding
LiveCodeBench v6
implementation=artificial-analysis
high effort77.7 %independentArtificial Analysis
LiveCodeBench v6
implementation=artificial-analysis
low effort65.2 %independentArtificial Analysis
SciCodehigh effort38.9 %independentArtificial Analysis
SciCodelow effort34.0 %independentArtificial Analysis
external indices
AA Coding Indexhigh effort20.7 pointsindependentArtificial Analysis
Intelligence Index v4.1low effort10.0 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort9.0 pointsindependentArtificial Analysis
AA Math Indexhigh effort89.3 pointsindependentArtificial Analysis
AA Math Indexlow effort62.3 pointsindependentArtificial Analysis
ECIlow effort137.8 pointsindependentEpoch AI Benchmarking Hub
ECI137.8 pointsindependentEpoch AI Benchmarking Hub
ECImedium effort137.8 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort137.8 pointsindependentEpoch AI Benchmarking Hub
factuality
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
86.5 %
84.488.5
independentScale Labs
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1369.2
1356.11382.2
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1358.1
1348.01368.1
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1351.9
1328.91374.9
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1350.0
1324.21375.7
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1336.6
1323.21349.9
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1335.3
1313.21357.5
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1331.0
1321.91340.1
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1329.0
1304.91353.2
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1326.7
1311.61341.9
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1322.7
1308.81336.7
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1322.2
1313.31331.1
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1320.4
1296.81344.0
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1317.1
1310.71323.5
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1313.6
1286.01341.3
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1308.2
1287.11329.3
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1301.8
1288.51315.2
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1298.6
1290.51306.6
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1289.3
1274.91303.8
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1281.2
1269.11293.4
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1274.8
1252.51297.0
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1274.0
1260.81287.2
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1261.6
1252.51270.7
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1254.7
1239.61269.8
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1239.6
1221.31257.9
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
68.9 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
high effort68.8 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
low effort61.1 %independentArtificial Analysis
GPQA Diamondmedium effort60.8 %independentEpoch AI Benchmarking Hub
2026-08-27
GPQA Diamondlow effort53.2 %independentEpoch AI Benchmarking Hub
2026-08-27
GPQA Diamondhigh effort46.0 %independentEpoch AI Benchmarking Hub
2026-08-06
Humanity's Last Exam
implementation=artificial-analysis
high effort11.0 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
low effort5.3 %independentArtificial Analysis
MMLU-Pro
implementation=artificial-analysis
high effort74.8 %independentArtificial Analysis
MMLU-Pro
implementation=artificial-analysis
low effort71.8 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
71.6 %independentVals AI
long context instruction
AA-LCRhigh effort34.7 %independentArtificial Analysis
AA-LCRlow effort30.7 %independentArtificial Analysis
IFBenchhigh effort65.1 %independentArtificial Analysis
IFBenchlow effort57.8 %independentArtificial Analysis
professional
CaseLaw43.8 %independentVals AI
CorpFin53.1 %independentVals AI
LegalBench70.8 %independentVals AI
MedQA82.9 %independentVals AI
TaxEval63.7 %independentVals AI
reasoning math
AIME
year=2025 · implementation=artificial-analysis
high effort89.3 %independentArtificial Analysis
AIME
implementation=vals-ai
86.0 %independentVals AI
AIME
year=2025 · implementation=artificial-analysis
low effort62.3 %independentArtificial Analysis
MATH-500
implementation=vals-ai
94.2 %independentVals AI
OTIS Mock AIME 2024–2025medium effort65.3 %independentEpoch AI Benchmarking Hub
2026-08-27
OTIS Mock AIME 2024–2025high effort53.9 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–2025high effort50.8 %independentEpoch AI Benchmarking Hub
2026-08-27
OTIS Mock AIME 2024–2025low effort40.3 %independentEpoch AI Benchmarking Hub
2026-08-27

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + GPT OSS 20BTerminal-Bench 2.113.9 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT OSS 20BTerminal-Bench Hard10.6 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT OSS 20BTerminal-Bench Hard4.5 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants