BenchAtlas

Rankings / Google DeepMind

Gemma 4 31B IT

released 2026-04-02

BenchAtlas Index

as of 2026-09-11
Not enough benchmark coverage for a composite score (needs ≥4 families across ≥3 categories). The per-benchmark evidence below stands on its own — sparse coverage is an evidence gap, not a low score.

Benchmark evidence

7 results
external indices
ECIminimal effort142.7 pointsindependentEpoch AI Benchmarking Hub
ECI142.7 pointsindependentEpoch AI Benchmarking Hub
factuality
SimpleQA Verified10.4 %independentEpoch AI Benchmarking Hub
2026-08-27
SimpleQA Verified9.6 %independentEpoch AI Benchmarking Hub
2026-04-04
knowledge science
GPQA Diamondminimal effort75.8 %independentEpoch AI Benchmarking Hub
2026-08-06
professional
CaseLawhigh effort52.6 %independentVals AI
reasoning math
OTIS Mock AIME 2024–2025minimal effort73.3 %independentEpoch AI Benchmarking Hub
2026-08-06