BenchAtlas

Sources / scale-labs

Scale Labs

healthydefault: independent

labs.scale.com/leaderboard · trust rank 70 · adapter scale-labs

Observations
262
append-only score facts
Raw snapshots
313
immutable fetched payloads
Last success
2026-09-11 10:30Z
end of last succeeded/partial run
Freshness
21.8 h
hours since last success

License & attribution

scale.com terms grant an internal-purposes license and robots.txt allows the leaderboard, but republication is a legal gray zone — a wave-2 source, attributed and linked, to be reassessed before public launch.

What we ingest

Embedded payloads from labs.scale.com leaderboard pages: score plus 95% CI half-width per model configuration across their SEAL-lineage benchmarks (HLE, MultiChallenge, Fortress, …). Scale runs the evals, so rows land as independent.

Ingestion runs

last 10 of the run history
StartedDurationStatus
2026-09-11 10:30Z5.5 spartial
2026-09-10 10:30Z5.1 spartial
2026-09-09 10:30Z5.0 spartial
2026-09-08 10:30Z4.5 spartial
2026-09-07 10:30Z4.6 spartial
2026-09-06 10:30Z5.3 spartial
2026-09-05 10:30Z5.5 spartial
2026-09-04 10:30Z4.6 spartial
2026-09-03 10:30Z4.5 ssucceeded
2026-09-02 10:30Z4.7 ssucceeded

partial = run finished but some rows are blocked on identity review — nothing is published until every referenced model/agent/benchmark is resolved.

Raw snapshots

latest 5
FetchedURLSHA-256
2026-09-11 10:30Zhttps://labs.scale.com/leaderboard/visual_language_understanding8286387e6b3d
2026-09-11 10:30Zhttps://labs.scale.com/leaderboard/maska3a261f32c7b
2026-09-11 10:30Zhttps://labs.scale.com/leaderboard/enigma_eval578d120bd254
2026-09-11 10:30Zhttps://labs.scale.com/leaderboard/multichallengeeb1aa57cadd4
2026-09-11 10:30Zhttps://labs.scale.com/leaderboard/humanitys_last_exama09255106d50

Snapshots are immutable raw payloads — every published number traces back to one. Re-fetching an identical payload bumps the re-fetch time instead of storing a duplicate. Scores from this source appear on benchmark profiles and model pages with this source linked.