BenchAtlas

Rankings / Anthropic

Claude 2.1

released 2023-11-21

BenchAtlas Index

as of 2026-09-11
Not enough benchmark coverage for a composite score (needs ≥4 families across ≥3 categories). The per-benchmark evidence below stands on its own — sparse coverage is an evidence gap, not a low score.

Benchmark evidence

3 results
external indices
ECI119.2 pointsindependentEpoch AI Benchmarking Hub
knowledge science
GPQA Diamond33.0 %independentEpoch AI Benchmarking Hub
2025-01-27
reasoning math
OTIS Mock AIME 2024–20251.9 %independentEpoch AI Benchmarking Hub
2025-03-07

Related variants