BenchAtlas

Rankings / Google

Gemini 3.1 Flash Lite Preview

proprietary

released 2026-03-03

BenchAtlas Index

as of 2026-09-11
42.8
base configuration
rank #288
11 families · 6 categories · high
41.7
high effort
rank #300
8 families · 4 categories · medium

Benchmark evidence

53 results
agentic coding
τ²-Bench31.3 %independentArtificial Analysis
τ²-Bench
subset=banking
9.7 %independentArtificial Analysis
coding
SciCode43.4 %independentArtificial Analysis
external indices
AA Coding Index34.7 pointsindependentArtificial Analysis
Intelligence Index v4.116.0 pointsindependentArtificial Analysis
Vals Indexhigh effort15.5 pointsindependentVals AI
factuality
MASK
contamination=Potential contamination warning: This model was evaluated after the public release of MASK, allowing model builder access to the prompts and solutions.
48.4 %
47.549.3
independentScale Labs
human preference
Chinese (style control)older version
arena=text · category=chinese · style_control=true
1477.7
1467.21488.2
communityLMArena Leaderboard Dataset
Coding (style control)older version
arena=text · category=coding · style_control=true
1456.9
1451.01462.8
communityLMArena Leaderboard Dataset
Industry Medicine And Healthcare (style control)older version
arena=text · category=industry_medicine_and_healthcare · style_control=true
1455.3
1445.81464.8
communityLMArena Leaderboard Dataset
Polish (style control)older version
arena=text · category=polish · style_control=true
1454.9
1438.21471.7
communityLMArena Leaderboard Dataset
Industry Software And It Services (style control)older version
arena=text · category=industry_software_and_it_services · style_control=true
1454.2
1449.01459.4
communityLMArena Leaderboard Dataset
Industry Life And Physical And Social Science (style control)older version
arena=text · category=industry_life_and_physical_and_social_science · style_control=true
1450.8
1444.01457.5
communityLMArena Leaderboard Dataset
French (style control)older version
arena=text · category=french · style_control=true
1448.4
1433.21463.6
communityLMArena Leaderboard Dataset
Hard Prompts (style control)older version
arena=text · category=hard_prompts · style_control=true
1444.6
1440.11449.2
communityLMArena Leaderboard Dataset
Expert (style control)older version
arena=text · category=expert · style_control=true
1442.5
1433.91451.1
communityLMArena Leaderboard Dataset
Hard Prompts English (style control)older version
arena=text · category=hard_prompts_english · style_control=true
1441.7
1436.01447.3
communityLMArena Leaderboard Dataset
Industry Legal And Government (style control)older version
arena=text · category=industry_legal_and_government · style_control=true
1440.2
1431.01449.3
communityLMArena Leaderboard Dataset
Spanish (style control)older version
arena=text · category=spanish · style_control=true
1439.0
1424.71453.2
communityLMArena Leaderboard Dataset
Math (style control)older version
arena=text · category=math · style_control=true
1438.0
1427.61448.4
communityLMArena Leaderboard Dataset
Multi Turn (style control)older version
arena=text · category=multi_turn · style_control=true
1435.7
1429.11442.3
communityLMArena Leaderboard Dataset
German (style control)older version
arena=text · category=german · style_control=true
1434.8
1415.31454.3
communityLMArena Leaderboard Dataset
Russian (style control)older version
arena=text · category=russian · style_control=true
1433.7
1425.91441.5
communityLMArena Leaderboard Dataset
English (style control)older version
arena=text · category=english · style_control=true
1433.4
1428.51438.4
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1432.5
1428.81436.2
communityLMArena Leaderboard Dataset
Industry Mathematical (style control)older version
arena=text · category=industry_mathematical · style_control=true
1431.8
1420.91442.6
communityLMArena Leaderboard Dataset
Longer Query (style control)older version
arena=text · category=longer_query · style_control=true
1428.2
1422.81433.5
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
1427.1
1422.11432.2
communityLMArena Leaderboard Dataset
Non English (style control)older version
arena=text · category=non_english · style_control=true
1424.4
1419.81429.0
communityLMArena Leaderboard Dataset
Industry Business And Management And Financial Operations (style control)older version
arena=text · category=industry_business_and_management_and_financial_operations · style_control=true
1420.9
1414.51427.3
communityLMArena Leaderboard Dataset
Industry Writing And Literature And Language (style control)older version
arena=text · category=industry_writing_and_literature_and_language · style_control=true
1420.5
1414.51426.5
communityLMArena Leaderboard Dataset
Japanese (style control)older version
arena=text · category=japanese · style_control=true
1419.6
1393.91445.3
communityLMArena Leaderboard Dataset
Creative Writing (style control)older version
arena=text · category=creative_writing · style_control=true
1413.1
1406.01420.2
communityLMArena Leaderboard Dataset
Instruction Following (style control)older version
arena=text · category=instruction_following · style_control=true
1406.7
1401.31412.2
communityLMArena Leaderboard Dataset
Industry Entertainment And Sports And Media (style control)older version
arena=text · category=industry_entertainment_and_sports_and_media · style_control=true
1403.4
1397.01409.8
communityLMArena Leaderboard Dataset
Korean (style control)older version
arena=text · category=korean · style_control=true
1401.0
1381.31420.6
communityLMArena Leaderboard Dataset
knowledge science
GPQA Diamond
implementation=artificial-analysis
82.2 %independentArtificial Analysis
GPQA Diamond
implementation=vals-ai
high effort81.1 %independentVals AI
Humanity's Last Exam
implementation=artificial-analysis
17.2 %independentArtificial Analysis
Humanity's Last Exam
contamination=Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. · implementation=scale
8.6 %
7.59.7
independentScale Labs
MMLU-Pro
implementation=vals-ai
86.2 %independentVals AI
MMLU-Pro
implementation=vals-ai
high effort86.2 %independentVals AI
long context instruction
AA-LCR74.3 %independentArtificial Analysis
IFBench77.2 %independentArtificial Analysis
MultiChallenge60.6 %
60.261.0
independentScale Labs
multimodal
MMMU
implementation=vals-ai
high effort82.5 %independentVals AI
VISTA46.9 %
44.249.7
independentScale Labs
professional
CaseLawhigh effort55.0 %independentVals AI
CorpFinhigh effort59.4 %independentVals AI
LegalBenchhigh effort83.8 %independentVals AI
TaxEvalhigh effort71.8 %independentVals AI
reasoning math
AIME
implementation=vals-ai
high effort83.3 %independentVals AI
EnigmaEval
contamination=Potential contamination warning: This model was evaluated after the public release of EnigmaEval, allowing model builder access to the prompts and solutions.
3.0 %
2.14.0
independentScale Labs

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + Gemini 3.1 Flash Lite PreviewTerminal-Bench 2.131.1 %independentArtificial Analysis
agent + model Artificial Analysis harness + Gemini 3.1 Flash Lite PreviewTerminal-Bench Hard24.2 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.