BenchAtlas

Benchmarks / Reasoning & Math

OTIS Mock AIME

Mock AIME 2024–2025 problems evaluated by Epoch AI.

maintained by Epoch AI · 250 observations on record

Versions

1 version
VersionTasks
OTIS Mock AIME 2024–2025 current
otis-mock-aime

Current leaderboard

OTIS Mock AIME 2024–2025 · Accuracy · top 25 of 236 systems
95969899100
GPT 5.6 Sol
100.0
GPT 5.5 Pro Pre Release
100.0
Claude Fable 5
100.0
GPT 5.5 Pre Release
100.0
GPT 5.6 Terra
99.7
Claude Fable 5
99.7
GPT 5.6 Luna
98.3
Claude Opus 4.8
98.3
Claude Opus 4.7
97.8
Claude Fable 5
97.8
#SystemAccuracy %
1GPT 5.6 Solmax effort
OpenAI
100.0%
2GPT 5.5 Pro Pre Releasexhigh effort
OpenAI
100.0%
3Claude Fable 5high effort
Anthropic
100.0%
4GPT 5.5 Pre Releasexhigh effort
OpenAI
100.0%
5GPT 5.6 Terramax effort
OpenAI
99.7%
6Claude Fable 5max effort
Anthropic
99.7%
7GPT 5.6 Lunamax effort
OpenAI
98.3%
8Claude Opus 4.8max effort
Anthropic
98.3%
9Claude Opus 4.7xhigh effort
Anthropic
97.8%
10Claude Fable 5low effort
Anthropic
97.8%
11GPT 5.4 (2026-03-05)high effort
OpenAI
97.8%
12Claude Opus 4.8low effort
Anthropic
97.8%
13Grok 4.5high effort
xAI
97.8%
14Deepseek V4 Promax effort
DeepSeek
96.7%
15Kimi K2 7 Code
Moonshot
96.4%
16Kimi K2 6
Moonshot AI
96.1%
17GPT 5.2 (2025-12-11)high effort
OpenAI
96.1%
18GPT 5.2 (2025-12-11)xhigh effort
OpenAI
96.1%
19Gemini 3.1 Pro Preview
Google DeepMind
95.6%
20Gemini 3.1 Pro Previewhigh effort
Google DeepMind
95.6%
21GPT 5.4 (2026-03-05)medium effort
OpenAI
95.6%
22GPT 5.6 Sollow effort
OpenAI
95.6%
23Qwen3 7max effort
Alibaba
95.6%
24Deepseek V4 Prohigh effort
DeepSeek
95.6%
25Gemini 3 Flash Previewhigh effort
Google DeepMind
95.6%

One row per evaluated system — reasoning-effort variants rank separately. Where a system has attr variants for this version (eval splits, style control), only the canonical variant is ranked: split=semi_private, style_control=true, or the unsplit run.