BenchAtlas

Rankings / Google DeepMind

Gemini 3.1 Pro Preview

released 2026-02-19

BenchAtlas Index

as of 2026-09-11
79.0
base configuration
rank #17
12 families · 5 categories · high
68.5
high effort
rank #84
5 families · 4 categories · medium
57.8
high effort
rank #176
10 families · 5 categories · high

Benchmark evidence

81 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort44.1 %independentLiveBench
τ²-Bench95.6 %independentArtificial Analysis
τ²-Bench
subset=banking
21.4 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
high effort76.5 %independentLiveBench
SciCode58.7 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort78.5 %independentLiveBench
external indices
AA Coding Index68.7 pointsindependentArtificial Analysis
Intelligence Index v4.130.4 pointsindependentArtificial Analysis
ECIhigh effort154.9 pointsindependentEpoch AI Benchmarking Hub
ECI154.9 pointsindependentEpoch AI Benchmarking Hub
ECIlow effort154.9 pointsindependentEpoch AI Benchmarking Hub
ECImedium effort154.9 pointsindependentEpoch AI Benchmarking Hub
Vals Indexhigh effort41.9 pointsindependentVals AI
factuality
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
42.4 %
41.143.7
independentScale Labs
SimpleQA Verified77.3 %independentEpoch AI Benchmarking Hub
2026-02-19
SimpleQA Verifiedhigh effort73.5 %independentEpoch AI Benchmarking Hub
2026-08-10
human preference
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1531.1
1522.51539.6
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
1521.1
1516.11526.0
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1511.8
1507.31516.2
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1511.5
1505.71517.2
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1508.6
1501.71515.6
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1507.0
1503.11510.9
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1505.5
1500.61510.4
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1502.5
1494.61510.4
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1502.0
1489.71514.4
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1500.9
1496.51505.2
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1499.4
1492.71506.1
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1499.4
1494.81503.9
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1496.3
1488.81503.8
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1494.5
1488.81500.2
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1494.4
1474.91513.9
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1494.1
1480.71507.5
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1493.7
1478.01509.4
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1488.7
1480.01497.4
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1488.5
1484.31492.8
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1486.7
1483.51489.9
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1484.9
1475.91493.8
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1480.8
1476.21485.5
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1480.3
1468.71492.0
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1480.1
1474.91485.2
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1479.7
1475.81483.7
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1479.1
1473.21485.0
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1476.7
1471.31482.2
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1466.4
1461.01471.8
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1460.0
1444.21475.9
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=vals-ai
high effort95.5 %independentVals AI
GPQA Diamondhigh effort94.4 %independentEpoch AI Benchmarking Hub
2026-08-06
GPQA Diamond
implementation=artificial-analysis
94.1 %independentArtificial Analysis
GPQA Diamond94.1 %independentEpoch AI Benchmarking Hub
2026-02-20
Humanity's Last Exam
implementation=artificial-analysis
47.0 %independentArtificial Analysis
Humanity's Last Exam
contamination=Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. · implementation=scale
high effort46.4 %
44.548.4
independentScale Labs
MMLU-Pro
implementation=vals-ai
high effort91.0 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
high effort85.4 %independentLiveBench
long context instruction
AA-LCR82.0 %independentArtificial Analysis
IFBench77.1 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort79.1 %independentLiveBench
MultiChallenge71.4 %
69.673.1
independentScale Labs
multimodal
MMMU
implementation=vals-ai
high effort88.2 %independentVals AI
professional
CaseLawhigh effort64.8 %independentVals AI
CorpFinhigh effort64.5 %independentVals AI
LegalBenchhigh effort87.4 %independentVals AI
MedQAhigh effort96.4 %independentVals AI
TaxEvalhigh effort72.9 %independentVals AI
reasoning math
AIME
implementation=vals-ai
high effort98.1 %independentVals AI
ARC-AGI-1older version
split=semi_private · model_type=CoT
98.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
97.2 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
88.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
77.1 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
1.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
1.0 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
0.5 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
0.4 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
0.4 usd_per_taskindependentARC Prize Leaderboard
EnigmaEvalhigh effort36.8 %
34.139.5
independentScale Labs
EnigmaEval
contamination=Potential contamination warning: This model was evaluated after the public release of EnigmaEval, allowing model builder access to the prompts and solutions.
19.8 %
17.522.0
independentScale Labs
FrontierMath36.9 %independentEpoch AI Benchmarking Hub
2026-02-19
FrontierMath Tier 4older version16.7 %independentEpoch AI Benchmarking Hub
2026-02-19
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort91.0 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
high effort84.0 %independentLiveBench
OTIS Mock AIME 2024–202595.6 %independentEpoch AI Benchmarking Hub
2026-02-20
OTIS Mock AIME 2024–2025high effort95.6 %independentEpoch AI Benchmarking Hub
2026-08-06

Agent + model results

systems, not bare-model scores
agent + model judy + Gemini 3.1 Pro PreviewTerminal-Bench 2.080.2 %unverifiedTerminal-Bench Leaderboard
agent + model terminus-3-3 + Gemini 3.1 Pro PreviewTerminal-Bench 2.074.8 %unverifiedTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + Gemini 3.1 Pro PreviewTerminal-Bench 2.173.8 %independentArtificial Analysis
agent + model gemini-cli + Gemini 3.1 Pro PreviewTerminal-Bench 2.170.7 %communityTerminal-Bench Leaderboard
agent + model terminus-2 + Gemini 3.1 Pro PreviewTerminal-Bench 2.170.3 %communityTerminal-Bench Leaderboard
agent + model gemini-cli + Gemini 3.1 Pro PreviewTerminal-Bench 2.061.4 %unverifiedTerminal-Bench Leaderboard
agent + model gemini-cli + Gemini 3.1 Pro PreviewTerminal-Bench 2.059.4 %unverifiedTerminal-Bench Leaderboard
agent + model Artificial Analysis harness + Gemini 3.1 Pro PreviewTerminal-Bench Hard53.8 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.