BenchAtlas

Rankings / Meta

Codellama 70B Instruct

open weights

release date unknown

BenchAtlas Index

as of 2026-09-11
Not enough benchmark coverage for a composite score (needs ≥4 families across ≥3 categories). The per-benchmark evidence below stands on its own — sparse coverage is an evidence gap, not a low score.

Benchmark evidence

3 results
human preference
English (style control)older version
arena=text · category=english · style_control=true
1152.3
1130.31174.3
communityLMArena Leaderboard Dataset
Overall (style control)
arena=text · category=overall · style_control=true
1118.6
1100.51136.7
communityLMArena Leaderboard Dataset
Exclude Ties (style control)older version
arena=text · category=exclude_ties · style_control=true
957.8
930.0985.6
communityLMArena Leaderboard Dataset