Rankings / Moonshot
Kimi K2 7 Code
released 2026-06-12
BenchAtlas Index
as of 2026-09-11
base configuration
rank #231
13 families · 5 categories · high
Benchmark evidence
23 results
agentic coding
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 45.7 % | independent | LiveBench | |
| τ²-Bench | 90.1 % | independent | Artificial Analysis | |
| τ²-Bench subset=banking | 20.2 % | independent | Artificial Analysis |
coding
| Release 2026-06-25 tasks_counted=2 · livebench_version=2026-06-25 | 74.0 % | independent | LiveBench | |
| SciCode | 47.8 % | independent | Artificial Analysis |
data analysis
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 62.7 % | independent | LiveBench |
external indices
| AA Coding Index | 60.8 points | independent | Artificial Analysis | |
| Intelligence Index v4.1 | 26.3 points | independent | Artificial Analysis | |
| ECI | 150.3 points | independent | Epoch AI Benchmarking Hub |
factuality
| SimpleQA Verified | 39.2 % | independent | Epoch AI Benchmarking Hub 2026-06-12 | |
| SimpleQA Verified | 36.5 % | independent | Epoch AI Benchmarking Hub 2026-08-27 |
knowledge science
| GPQA Diamond implementation=artificial-analysis | 89.6 % | independent | Artificial Analysis | |
| GPQA Diamond | 89.5 % | independent | Epoch AI Benchmarking Hub 2026-06-12 | |
| GPQA Diamond | 87.9 % | independent | Epoch AI Benchmarking Hub 2026-08-07 | |
| Humanity's Last Exam implementation=artificial-analysis | 35.0 % | independent | Artificial Analysis |
language
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 77.9 % | independent | LiveBench |
long context instruction
| AA-LCR | 79.3 % | independent | Artificial Analysis | |
| IFBench | 63.1 % | independent | Artificial Analysis | |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 56.3 % | independent | LiveBench |
reasoning math
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 79.6 % | independent | LiveBench | |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 82.8 % | independent | LiveBench | |
| OTIS Mock AIME 2024–2025 | 96.4 % | independent | Epoch AI Benchmarking Hub 2026-06-13 | |
| OTIS Mock AIME 2024–2025 | 95.6 % | independent | Epoch AI Benchmarking Hub 2026-08-07 |
Agent + model results
systems, not bare-model scores
| agent + model Artificial Analysis harness + Kimi K2 7 Code | Terminal-Bench 2.1 | 67.4 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + Kimi K2 7 Code | Terminal-Bench Hard | 44.7 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
