BenchAtlas

Rankings / OpenAI

GPT 5.6 Terra

released 2026-07-09

BenchAtlas Index

as of 2026-09-11
79.9
xhigh effort
rank #14
7 families · 4 categories · medium
78.8
high effort
rank #18
8 families · 5 categories · high
78.6
max effort
rank #20
14 families · 5 categories · high

Benchmark evidence

136 results
agentic coding
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort55.0 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort53.3 %independentLiveBench
τ²-Benchmax effort86.3 %independentArtificial Analysis
τ²-Benchxhigh effort80.4 %independentArtificial Analysis
τ²-Benchhigh effort78.4 %independentArtificial Analysis
τ²-Benchmedium effort72.8 %independentArtificial Analysis
τ²-Benchlow effort60.5 %independentArtificial Analysis
τ²-Bench
subset=banking
max effort40.2 %independentArtificial Analysis
τ²-Bench
subset=banking
xhigh effort29.7 %independentArtificial Analysis
τ²-Bench
subset=banking
high effort28.7 %independentArtificial Analysis
τ²-Bench
subset=banking
medium effort25.6 %independentArtificial Analysis
τ²-Bench
subset=banking
low effort18.8 %independentArtificial Analysis
τ²-Bench
subset=banking
no reasoning15.7 %independentArtificial Analysis
coding
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
max effort78.3 %independentLiveBench
Release 2026-06-25
tasks_counted=2 · livebench_version=2026-06-25
xhigh effort75.4 %independentLiveBench
SciCodemax effort55.0 %independentArtificial Analysis
SciCodehigh effort52.4 %independentArtificial Analysis
SciCodexhigh effort52.3 %independentArtificial Analysis
SciCodemedium effort50.5 %independentArtificial Analysis
SciCodelow effort49.9 %independentArtificial Analysis
SciCodeno reasoning44.6 %independentArtificial Analysis
data analysis
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort79.3 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort77.4 %independentLiveBench
external indices
AA Coding Indexmax effort76.7 pointsindependentArtificial Analysis
AA Coding Indexxhigh effort70.6 pointsindependentArtificial Analysis
AA Coding Indexhigh effort67.1 pointsindependentArtificial Analysis
AA Coding Indexmedium effort64.7 pointsindependentArtificial Analysis
AA Coding Indexlow effort58.1 pointsindependentArtificial Analysis
AA Coding Indexno reasoning52.3 pointsindependentArtificial Analysis
Intelligence Index v4.1max effort42.3 pointsindependentArtificial Analysis
Intelligence Index v4.1xhigh effort38.2 pointsindependentArtificial Analysis
Intelligence Index v4.1high effort34.5 pointsindependentArtificial Analysis
Intelligence Index v4.1medium effort30.4 pointsindependentArtificial Analysis
Intelligence Index v4.1low effort27.9 pointsindependentArtificial Analysis
Intelligence Index v4.1no reasoning22.3 pointsindependentArtificial Analysis
ECIlow effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECIhigh effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECI159.2 pointsindependentEpoch AI Benchmarking Hub
ECImedium effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECIno reasoning159.2 pointsindependentEpoch AI Benchmarking Hub
ECImax effort159.2 pointsindependentEpoch AI Benchmarking Hub
ECIxhigh effort159.2 pointsindependentEpoch AI Benchmarking Hub
Vals Indexmax effort59.6 pointsindependentVals AI
Vals Indexxhigh effort56.5 pointsindependentVals AI
factuality
SimpleQA Verifiedmax effort43.2 %independentEpoch AI Benchmarking Hub
2026-08-10
SimpleQA Verifiedmax effort43.1 %independentEpoch AI Benchmarking Hub
2026-07-09
knowledge science
GPQA Diamondmax effort93.3 %independentEpoch AI Benchmarking Hub
2026-07-09
GPQA Diamond
implementation=artificial-analysis
max effort92.5 %independentArtificial Analysis
GPQA Diamond
implementation=vals-ai
xhigh effort90.9 %independentVals AI
GPQA Diamond
implementation=artificial-analysis
xhigh effort90.8 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
high effort89.6 %independentArtificial Analysis
GPQA Diamondlow effort87.4 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
medium effort87.2 %independentArtificial Analysis
GPQA Diamond
implementation=artificial-analysis
low effort84.3 %independentArtificial Analysis
GPQA Diamondno reasoning77.3 %independentEpoch AI Benchmarking Hub
2026-08-07
GPQA Diamond
implementation=artificial-analysis
no reasoning74.6 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
max effort42.9 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
xhigh effort41.9 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
high effort38.5 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
medium effort33.3 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
low effort29.2 %independentArtificial Analysis
Humanity's Last Exam
implementation=artificial-analysis
no reasoning11.4 %independentArtificial Analysis
MMLU-Pro
implementation=vals-ai
xhigh effort86.7 %independentVals AI
language
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
max effort82.9 %independentLiveBench
Release 2026-06-25
tasks_counted=3 · livebench_version=2026-06-25
xhigh effort79.9 %independentLiveBench
long context instruction
AA-LCRmax effort83.0 %independentArtificial Analysis
AA-LCRxhigh effort79.0 %independentArtificial Analysis
AA-LCRhigh effort77.7 %independentArtificial Analysis
AA-LCRmedium effort74.0 %independentArtificial Analysis
AA-LCRlow effort71.3 %independentArtificial Analysis
AA-LCRno reasoning58.7 %independentArtificial Analysis
IFBenchmax effort71.2 %independentArtificial Analysis
IFBenchxhigh effort66.3 %independentArtificial Analysis
IFBenchhigh effort64.4 %independentArtificial Analysis
IFBenchmedium effort62.2 %independentArtificial Analysis
IFBenchlow effort59.7 %independentArtificial Analysis
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort64.6 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort59.4 %independentLiveBench
multimodal
MMMU
implementation=vals-ai
xhigh effort86.5 %independentVals AI
MMMU
implementation=vals-ai
max effort86.5 %independentVals AI
professional
CorpFinxhigh effort65.3 %independentVals AI
LegalBenchxhigh effort85.1 %independentVals AI
LegalBenchmax effort85.1 %independentVals AI
TaxEvalxhigh effort76.2 %independentVals AI
reasoning math
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort98.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort96.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort95.8 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
xhigh effort95.6 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
xhigh effort94.0 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort92.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort91.3 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort84.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort83.9 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
xhigh effort82.6 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort77.0 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
xhigh effort74.2 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort73.9 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort70.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort67.1 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort60.2 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort38.9 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort37.5 %independentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort18.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort14.9 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
max effort1.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
max effort1.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
max effort0.8 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
xhigh effort0.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
xhigh effort0.7 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
xhigh effort0.7 %independentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
high effort0.6 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
high effort0.6 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
max effort0.6 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
high effort0.5 %independentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
max effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
medium effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
medium effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
xhigh effort0.3 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
high effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
xhigh effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=public_eval · model_type=CoT
low effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-2
split=semi_private · model_type=CoT
low effort0.2 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
high effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
medium effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
medium effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=semi_private · model_type=CoT
low effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-1older version
split=public_eval · model_type=CoT
low effort0.1 usd_per_taskindependentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
medium effort0.1 %independentARC Prize Leaderboard
ARC-AGI-3older version
split=semi_private · model_type=CoT
low effort0.0 %independentARC Prize Leaderboard
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort94.9 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort89.5 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
max effort90.6 %independentLiveBench
Release 2026-06-25
tasks_counted=4 · livebench_version=2026-06-25
xhigh effort84.9 %independentLiveBench
OTIS Mock AIME 2024–2025max effort99.7 %independentEpoch AI Benchmarking Hub
2026-07-09
OTIS Mock AIME 2024–2025low effort88.9 %independentEpoch AI Benchmarking Hub
2026-08-07
OTIS Mock AIME 2024–2025no reasoning53.3 %independentEpoch AI Benchmarking Hub
2026-08-07

Agent + model results

systems, not bare-model scores
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench 2.188.0 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench 2.180.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench 2.175.7 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench 2.172.3 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench Hard62.9 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench 2.162.5 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench Hard57.6 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench Hard57.6 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench 2.156.2 %independentArtificial Analysis
agent + model Artificial Analysis harness + GPT 5.6 TerraTerminal-Bench Hard43.9 %independentArtificial Analysis

These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.