Rankings / Anthropic
Claude Opus 4.7
released 2026-04-16
BenchAtlas Index
as of 2026-09-11
max effort
rank #8
8 families · 4 categories · medium
max effort
rank #10
8 families · 4 categories · medium
base configuration
rank #34
4 families · 3 categories · medium
Benchmark evidence
77 results
agentic coding
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | xhigh effort | 50.7 % | independent | LiveBench |
| τ²-Bench | max effort | 88.6 % | independent | Artificial Analysis |
| τ²-Bench | high effort | 74.0 % | independent | Artificial Analysis |
| τ²-Bench subset=banking | max effort | 34.6 % | independent | Artificial Analysis |
coding
| Release 2026-06-25 tasks_counted=2 · livebench_version=2026-06-25 | xhigh effort | 82.1 % | independent | LiveBench |
| SciCode | max effort | 54.5 % | independent | Artificial Analysis |
| SciCode | high effort | 50.1 % | independent | Artificial Analysis |
data analysis
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | xhigh effort | 78.3 % | independent | LiveBench |
external indices
| AA Coding Index | max effort | 73.6 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | max effort | 40.7 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | high effort | 30.9 points | independent | Artificial Analysis |
| ECI | medium effort | 156.1 points | independent | Epoch AI Benchmarking Hub |
| ECI | xhigh effort | 156.1 points | independent | Epoch AI Benchmarking Hub |
| ECI | max effort | 156.1 points | independent | Epoch AI Benchmarking Hub |
| ECI | 156.1 points | independent | Epoch AI Benchmarking Hub | |
| ECI | high effort | 156.1 points | independent | Epoch AI Benchmarking Hub |
| ECI | low effort | 156.1 points | independent | Epoch AI Benchmarking Hub |
| Vals Index | high effort | 56.1 points | independent | Vals AI |
factuality
| SimpleQA Verified | xhigh effort | 51.7 % | independent | Epoch AI Benchmarking Hub 2026-08-27 |
| SimpleQA Verified | xhigh effort | 50.6 % | independent | Epoch AI Benchmarking Hub 2026-04-16 |
human preference
| Coding (style control)older version arena=text · category=coding · style_control=true | 1547.2 1541.2–1553.2 | community | LMArena Leaderboard Dataset | |
| Expert (style control)older version arena=text · category=expert · style_control=true | 1533.7 1525.2–1542.1 | community | LMArena Leaderboard Dataset | |
| Chinese (style control)older version arena=text · category=chinese · style_control=true | 1531.6 1520.7–1542.6 | community | LMArena Leaderboard Dataset | |
| Industry Software And It Services (style control)older version arena=text · category=industry_software_and_it_services · style_control=true | 1530.8 1525.5–1536.2 | community | LMArena Leaderboard Dataset | |
| Hard Prompts English (style control)older version arena=text · category=hard_prompts_english · style_control=true | 1520.8 1515.0–1526.5 | community | LMArena Leaderboard Dataset | |
| Hard Prompts (style control)older version arena=text · category=hard_prompts · style_control=true | 1518.8 1514.2–1523.4 | community | LMArena Leaderboard Dataset | |
| Multi Turn (style control)older version arena=text · category=multi_turn · style_control=true | 1514.8 1507.8–1521.9 | community | LMArena Leaderboard Dataset | |
| Industry Medicine And Healthcare (style control)older version arena=text · category=industry_medicine_and_healthcare · style_control=true | 1513.0 1503.2–1522.8 | community | LMArena Leaderboard Dataset | |
| French (style control)older version arena=text · category=french · style_control=true | 1512.0 1497.2–1526.8 | community | LMArena Leaderboard Dataset | |
| Industry Life And Physical And Social Science (style control)older version arena=text · category=industry_life_and_physical_and_social_science · style_control=true | 1511.4 1504.4–1518.4 | community | LMArena Leaderboard Dataset | |
| Exclude Ties (style control)older version arena=text · category=exclude_ties · style_control=true | 1508.9 1503.8–1514.0 | community | LMArena Leaderboard Dataset | |
| Longer Query (style control)older version arena=text · category=longer_query · style_control=true | 1507.2 1501.9–1512.6 | community | LMArena Leaderboard Dataset | |
| Industry Mathematical (style control)older version arena=text · category=industry_mathematical · style_control=true | 1500.5 1489.6–1511.5 | community | LMArena Leaderboard Dataset | |
| Polish (style control)older version arena=text · category=polish · style_control=true | 1500.2 1481.9–1518.6 | community | LMArena Leaderboard Dataset | |
| English (style control)older version arena=text · category=english · style_control=true | 1499.6 1494.6–1504.6 | community | LMArena Leaderboard Dataset | |
| Russian (style control)older version arena=text · category=russian · style_control=true | 1497.3 1488.8–1505.8 | community | LMArena Leaderboard Dataset | |
| German (style control)older version arena=text · category=german · style_control=true | 1495.1 1475.1–1515.1 | community | LMArena Leaderboard Dataset | |
| Industry Business And Management And Financial Operations (style control)older version arena=text · category=industry_business_and_management_and_financial_operations · style_control=true | 1495.0 1488.4–1501.7 | community | LMArena Leaderboard Dataset | |
| Overall (style control) arena=text · category=overall · style_control=true | 1494.5 1490.6–1498.4 | community | LMArena Leaderboard Dataset | |
| Industry Legal And Government (style control)older version arena=text · category=industry_legal_and_government · style_control=true | 1494.2 1485.0–1503.5 | community | LMArena Leaderboard Dataset | |
| Instruction Following (style control)older version arena=text · category=instruction_following · style_control=true | 1492.9 1487.3–1498.4 | community | LMArena Leaderboard Dataset | |
| Math (style control)older version arena=text · category=math · style_control=true | 1491.6 1480.4–1502.7 | community | LMArena Leaderboard Dataset | |
| Non English (style control)older version arena=text · category=non_english · style_control=true | 1483.9 1479.1–1488.6 | community | LMArena Leaderboard Dataset | |
| Industry Writing And Literature And Language (style control)older version arena=text · category=industry_writing_and_literature_and_language · style_control=true | 1483.5 1477.2–1489.7 | community | LMArena Leaderboard Dataset | |
| Creative Writing (style control)older version arena=text · category=creative_writing · style_control=true | 1481.3 1474.1–1488.4 | community | LMArena Leaderboard Dataset | |
| Japanese (style control)older version arena=text · category=japanese · style_control=true | 1478.2 1454.4–1501.9 | community | LMArena Leaderboard Dataset | |
| Spanish (style control)older version arena=text · category=spanish · style_control=true | 1473.6 1459.1–1488.0 | community | LMArena Leaderboard Dataset | |
| Industry Entertainment And Sports And Media (style control)older version arena=text · category=industry_entertainment_and_sports_and_media · style_control=true | 1473.0 1466.5–1479.5 | community | LMArena Leaderboard Dataset | |
| Korean (style control)older version arena=text · category=korean · style_control=true | 1467.5 1448.2–1486.8 | community | LMArena Leaderboard Dataset |
knowledge science
| GPQA Diamond implementation=artificial-analysis | max effort | 91.4 % | independent | Artificial Analysis |
| GPQA Diamond implementation=vals-ai | max effort | 90.2 % | independent | Vals AI |
| GPQA Diamond | xhigh effort | 90.2 % | independent | Epoch AI Benchmarking Hub 2026-04-17 |
| GPQA Diamond implementation=artificial-analysis | high effort | 88.5 % | independent | Artificial Analysis |
| GPQA Diamond | max effort | 86.4 % | independent | Epoch AI Benchmarking Hub 2026-08-06 |
| Humanity's Last Exam implementation=artificial-analysis | max effort | 42.3 % | independent | Artificial Analysis |
| Humanity's Last Exam contamination=Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. · implementation=scale | 36.2 % 34.3–38.1 | independent | Scale Labs | |
| Humanity's Last Exam implementation=artificial-analysis | high effort | 33.3 % | independent | Artificial Analysis |
| MMLU-Pro implementation=vals-ai | max effort | 89.9 % | independent | Vals AI |
language
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | xhigh effort | 77.9 % | independent | LiveBench |
long context instruction
| AA-LCR | max effort | 78.7 % | independent | Artificial Analysis |
| AA-LCR | high effort | 75.7 % | independent | Artificial Analysis |
| IFBench | max effort | 58.6 % | independent | Artificial Analysis |
| IFBench | high effort | 43.6 % | independent | Artificial Analysis |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | xhigh effort | 66.7 % | independent | LiveBench |
multimodal
| MMMU implementation=vals-ai | max effort | 85.5 % | independent | Vals AI |
professional
| CaseLaw | 68.4 % | independent | Vals AI | |
| CorpFin | max effort | 66.1 % | independent | Vals AI |
| LegalBench | max effort | 85.3 % | independent | Vals AI |
| TaxEval | max effort | 75.3 % | independent | Vals AI |
reasoning math
| AIME implementation=vals-ai | 96.3 % | independent | Vals AI | |
| ARC-AGI-3older version split=semi_private · model_type=CoT | high effort | 0.2 % | independent | ARC Prize Leaderboard |
| FrontierMath | xhigh effort | 43.8 % | independent | Epoch AI Benchmarking Hub 2026-04-17 |
| FrontierMath Tier 4older version | xhigh effort | 22.9 % | independent | Epoch AI Benchmarking Hub 2026-04-17 |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | xhigh effort | 92.8 % | independent | LiveBench |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | xhigh effort | 87.2 % | independent | LiveBench |
| OTIS Mock AIME 2024–2025 | xhigh effort | 97.8 % | independent | Epoch AI Benchmarking Hub 2026-04-17 |
| OTIS Mock AIME 2024–2025 | max effort | 86.7 % | independent | Epoch AI Benchmarking Hub 2026-08-06 |
Agent + model results
systems, not bare-model scores
| agent + model Epoch Inspect harness + Claude Opus 4.7 | SWE-bench Verified | 83.5 % | independent | Epoch AI Benchmarking Hub |
| agent + model Artificial Analysis harness + Claude Opus 4.7 | Terminal-Bench 2.1 | 83.2 % | independent | Artificial Analysis |
| agent + model claude-wozcode-plugin + Claude Opus 4.7 | Terminal-Bench 2.0 | 80.2 % | unverified | Terminal-Bench Leaderboard |
| agent + model Claude Code + Claude Opus 4.7 | Terminal-Bench 2.1 | 69.7 % | community | Terminal-Bench Leaderboard |
| agent + model terminus-2 + Claude Opus 4.7 | Terminal-Bench 2.1 | 66.1 % | community | Terminal-Bench Leaderboard |
| agent + model Artificial Analysis harness + Claude Opus 4.7 | Terminal-Bench Hard | 54.5 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + Claude Opus 4.7 | Terminal-Bench Hard | 51.5 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
