BenchAtlas

Rankings / DeepSeek

Deepseek Chat

released 2025-12-01

BenchAtlas Index

as of 2026-09-11
Not enough benchmark coverage for a composite score (needs ≥4 families across ≥3 categories). The per-benchmark evidence below stands on its own — sparse coverage is an evidence gap, not a low score.

Benchmark evidence

3 results
external indices
ECI146.2 pointsindependentEpoch AI Benchmarking Hub
knowledge science
GPQA Diamond71.2 %independentEpoch AI Benchmarking Hub
2026-07-16
reasoning math
OTIS Mock AIME 2024–202548.9 %independentEpoch AI Benchmarking Hub
2026-07-16

Agent + model results

systems, not bare-model scores
agent + model Moatless Tools + Deepseek ChatSWE-bench Lite30.7 %communitySWE-bench Leaderboard

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.