BenchAtlas

Rankings / OpenAI

GPT OSS 120B

released 2025-08-05

BenchAtlas Index

as of 2026-09-11
39.6
high effort
rank #309
10 families · 5 categories · high
38.0
base configuration
rank #317
4 families · 3 categories · medium
37.0
base configuration
rank #327
9 families · 4 categories · medium

Benchmark evidence

68 results
agentic coding
τ²-Benchhigh effort65.8 %independentArtificial Analysis
τ²-Benchlow effort45.0 %independentArtificial Analysis
τ²-Bench
subset=banking
high effort12.8 %independentArtificial Analysis
τ²-Bench
subset=banking
low effort2.9 %independentArtificial Analysis
coding
LiveCodeBench v6
implementation=artificial-analysis
high effort87.8 %independentArtificial Analysis
LiveCodeBench v6
implementation=artificial-analysis
low effort70.7 %independentArtificial Analysis
SciCodelow effort36.0 %independentArtificial Analysis
SciCodehigh effort34.0 %independentArtificial Analysis
external indices
AA Coding Indexhigh effort30.4 pointsindependentArtificial Analysis
AA Coding Indexlow effort21.2 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort12.3 pointsindependentArtificial Analysis
Intelligence Index v4.1low effort10.2 pointsindependentArtificial Analysis
AA Math Indexhigh effort93.4 pointsindependentArtificial Analysis
AA Math Indexlow effort66.7 pointsindependentArtificial Analysis
ECI140.1 pointsindependentEpoch AI Benchmarking Hub
ECImedium effort140.1 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort140.1 pointsindependentEpoch AI Benchmarking Hub
factuality
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
92.0 %
91.192.9
independentScale Labs
human preference
Coding (style control)older version
arena=text · category=coding · style_control=true
1390.8
1383.01398.5
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1386.3
1379.91392.8
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1381.9
1367.01396.8
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1379.6
1365.81393.4
communityLMArena Leaderboard Dataset
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1377.1
1364.31389.9
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1366.9
1359.21374.6
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1365.8
1350.71381.0
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1365.4
1344.11386.8
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1362.1
1356.41367.8
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1361.7
1355.81367.6
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1360.1
1351.21368.9
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1354.6
1338.31370.9
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1352.5
1348.11356.8
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1351.6
1343.21360.0
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1341.0
1327.11354.8
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1337.9
1332.51343.2
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1334.1
1321.61346.6
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1331.9
1317.21346.5
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1329.1
1306.11352.1
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1328.0
1302.71353.3
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1327.0
1318.21335.8
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1324.9
1317.11332.7
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1324.2
1317.11331.3
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1313.7
1307.61319.9
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1309.3
1301.61317.0
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1285.1
1276.81293.3
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1277.7
1267.81287.7
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1265.1
1240.61289.6
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
78.5 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
high effort78.2 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
low effort67.2 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
high effort19.6 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
low effort5.9 %independentArtificial Analysis
MMLU-Pro
implementation=artificial-analysis
high effort80.8 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
79.2 %independentVals AI
MMLU-Pro
implementation=artificial-analysis
low effort77.5 %independentArtificial Analysis
long context instruction
AA-LCRhigh effort52.0 %independentArtificial Analysis
AA-LCRlow effort46.0 %independentArtificial Analysis
IFBenchhigh effort69.0 %independentArtificial Analysis
IFBenchlow effort58.3 %independentArtificial Analysis
MultiChallenge45.3 %
43.247.5
independentScale Labs
professional
CaseLaw48.8 %independentVals AI
CorpFin58.2 %independentVals AI
LegalBench75.9 %independentVals AI
MedQA91.4 %independentVals AI
TaxEval71.6 %independentVals AI
reasoning math
AIME
year=2025 · implementation=artificial-analysis
high effort93.4 %independentArtificial Analysis
AIME
implementation=vals-ai
92.6 %independentVals AI
AIME
year=2025 · implementation=artificial-analysis
low effort66.7 %independentArtificial Analysis
MATH-500
implementation=vals-ai
94.8 %independentVals AI

Agent + model results

systems, not bare-model scores
agent + model mini-SWE-agent + GPT OSS 120BSWE-bench Verified26.0 %communitySWE-bench Leaderboard
agent + model mini-SWE-agent + GPT OSS 120BSWE-bench bash-only26.0 %communitySWE-bench Leaderboard
agent + model mini-SWE-agent + GPT OSS 120BSWE-bench bash-only0.1 usd_per_taskcommunitySWE-bench Leaderboard
agent + model mini-SWE-agent + GPT OSS 120BSWE-bench Verified0.1 usd_per_taskcommunitySWE-bench Leaderboard
agent + model Artificial Analysis harness + GPT OSS 120BTerminal-Bench 2.126.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT OSS 120BTerminal-Bench Hard23.5 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT OSS 120BTerminal-Bench 2.113.9 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT OSS 120BTerminal-Bench Hard5.3 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.

Related variants