Rankings / Z.ai
GLM 5.2
open weightsreleased 2026-06-13
BenchAtlas Index
as of 2026-09-11
max effort
rank #29
9 families · 5 categories · high
max effort
rank #55
7 families · 3 categories · medium
base configuration
rank #91
13 families · 6 categories · high
Benchmark evidence
111 results
agentic coding
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 51.8 % | independent | LiveBench | |
| τ²-Bench | max effort | 99.1 % | independent | Artificial Analysis |
| τ²-Bench subset=banking | max effort | 34.6 % | independent | Artificial Analysis |
| τ²-Bench subset=banking | no reasoning | 16.7 % | independent | Artificial Analysis |
coding
| Release 2026-06-25 tasks_counted=2 · livebench_version=2026-06-25 | 79.7 % | independent | LiveBench | |
| SciCode | max effort | 51.2 % | independent | Artificial Analysis |
| SciCode | no reasoning | 36.1 % | independent | Artificial Analysis |
data analysis
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 73.7 % | independent | LiveBench |
external indices
| AA Coding Index | max effort | 68.8 points | independent | Artificial Analysis |
| AA Coding Index | no reasoning | 46.5 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | max effort | 38.6 points | independent | Artificial Analysis |
| Intelligence Index v4.1 | no reasoning | 22.4 points | independent | Artificial Analysis |
| ECI | max effort | 152.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | minimal effort | 152.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | medium effort | 152.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | low effort | 152.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | no reasoning | 152.0 points | independent | Epoch AI Benchmarking Hub |
| ECI | 152.0 points | independent | Epoch AI Benchmarking Hub | |
| ECI | high effort | 152.0 points | independent | Epoch AI Benchmarking Hub |
| Vals Index | max effort | 65.0 points | independent | Vals AI |
| Vals Index | 53.1 points | independent | Vals AI |
factuality
| SimpleQA Verified | max effort | 38.1 % | independent | Epoch AI Benchmarking Hub 2026-06-19 |
| SimpleQA Verified | max effort | 34.2 % | independent | Epoch AI Benchmarking Hub 2026-08-27 |
human preference
| Chinese (style control)older version arena=text · category=chinese · style_control=true | 1517.0 1503.5–1530.4 | community | LMArena Leaderboard Dataset | |
| Coding (style control)older version arena=text · category=coding · style_control=true | max effort | 1510.9 1500.8–1520.9 | community | LMArena Leaderboard Dataset |
| Coding (style control)older version arena=text · category=coding · style_control=true | 1508.9 1501.8–1516.0 | community | LMArena Leaderboard Dataset | |
| Chinese (style control)older version arena=text · category=chinese · style_control=true | max effort | 1503.9 1481.5–1526.4 | community | LMArena Leaderboard Dataset |
| Industry Software And It Services (style control)older version arena=text · category=industry_software_and_it_services · style_control=true | 1500.8 1494.6–1507.0 | community | LMArena Leaderboard Dataset | |
| Industry Software And It Services (style control)older version arena=text · category=industry_software_and_it_services · style_control=true | max effort | 1499.4 1490.8–1508.1 | community | LMArena Leaderboard Dataset |
| Industry Mathematical (style control)older version arena=text · category=industry_mathematical · style_control=true | max effort | 1496.7 1475.3–1518.1 | community | LMArena Leaderboard Dataset |
| Hard Prompts English (style control)older version arena=text · category=hard_prompts_english · style_control=true | 1493.7 1486.8–1500.7 | community | LMArena Leaderboard Dataset | |
| Hard Prompts (style control)older version arena=text · category=hard_prompts · style_control=true | 1492.0 1486.5–1497.4 | community | LMArena Leaderboard Dataset | |
| Hard Prompts English (style control)older version arena=text · category=hard_prompts_english · style_control=true | max effort | 1491.5 1482.1–1500.8 | community | LMArena Leaderboard Dataset |
| Expert (style control)older version arena=text · category=expert · style_control=true | 1490.6 1480.3–1501.0 | community | LMArena Leaderboard Dataset | |
| Industry Life And Physical And Social Science (style control)older version arena=text · category=industry_life_and_physical_and_social_science · style_control=true | 1490.4 1481.9–1498.8 | community | LMArena Leaderboard Dataset | |
| French (style control)older version arena=text · category=french · style_control=true | 1489.3 1470.8–1507.7 | community | LMArena Leaderboard Dataset | |
| Industry Life And Physical And Social Science (style control)older version arena=text · category=industry_life_and_physical_and_social_science · style_control=true | max effort | 1488.3 1475.9–1500.6 | community | LMArena Leaderboard Dataset |
| Hard Prompts (style control)older version arena=text · category=hard_prompts · style_control=true | max effort | 1487.1 1479.8–1494.4 | community | LMArena Leaderboard Dataset |
| French (style control)older version arena=text · category=french · style_control=true | max effort | 1486.7 1460.2–1513.2 | community | LMArena Leaderboard Dataset |
| Industry Mathematical (style control)older version arena=text · category=industry_mathematical · style_control=true | 1486.7 1472.6–1500.8 | community | LMArena Leaderboard Dataset | |
| Industry Medicine And Healthcare (style control)older version arena=text · category=industry_medicine_and_healthcare · style_control=true | 1485.0 1472.3–1497.6 | community | LMArena Leaderboard Dataset | |
| Expert (style control)older version arena=text · category=expert · style_control=true | max effort | 1483.1 1467.3–1498.9 | community | LMArena Leaderboard Dataset |
| Longer Query (style control)older version arena=text · category=longer_query · style_control=true | 1481.8 1475.6–1487.9 | community | LMArena Leaderboard Dataset | |
| Industry Legal And Government (style control)older version arena=text · category=industry_legal_and_government · style_control=true | 1480.3 1468.4–1492.2 | community | LMArena Leaderboard Dataset | |
| Industry Medicine And Healthcare (style control)older version arena=text · category=industry_medicine_and_healthcare · style_control=true | max effort | 1480.0 1461.4–1498.7 | community | LMArena Leaderboard Dataset |
| Exclude Ties (style control)older version arena=text · category=exclude_ties · style_control=true | 1479.9 1473.8–1486.1 | community | LMArena Leaderboard Dataset | |
| Spanish (style control)older version arena=text · category=spanish · style_control=true | 1479.8 1460.2–1499.4 | community | LMArena Leaderboard Dataset | |
| English (style control)older version arena=text · category=english · style_control=true | 1479.5 1473.5–1485.5 | community | LMArena Leaderboard Dataset | |
| Math (style control)older version arena=text · category=math · style_control=true | max effort | 1477.3 1455.4–1499.2 | community | LMArena Leaderboard Dataset |
| Math (style control)older version arena=text · category=math · style_control=true | 1476.3 1461.5–1491.2 | community | LMArena Leaderboard Dataset | |
| Longer Query (style control)older version arena=text · category=longer_query · style_control=true | max effort | 1474.7 1466.2–1483.2 | community | LMArena Leaderboard Dataset |
| English (style control)older version arena=text · category=english · style_control=true | max effort | 1474.6 1466.7–1482.6 | community | LMArena Leaderboard Dataset |
| Industry Legal And Government (style control)older version arena=text · category=industry_legal_and_government · style_control=true | max effort | 1474.4 1456.8–1492.0 | community | LMArena Leaderboard Dataset |
| Overall (style control) arena=text · category=overall · style_control=true | 1471.7 1467.0–1476.4 | community | LMArena Leaderboard Dataset | |
| Exclude Ties (style control)older version arena=text · category=exclude_ties · style_control=true | max effort | 1470.7 1462.4–1479.0 | community | LMArena Leaderboard Dataset |
| German (style control)older version arena=text · category=german · style_control=true | 1470.5 1446.2–1494.9 | community | LMArena Leaderboard Dataset | |
| Multi Turn (style control)older version arena=text · category=multi_turn · style_control=true | 1469.0 1460.4–1477.6 | community | LMArena Leaderboard Dataset | |
| Russian (style control)older version arena=text · category=russian · style_control=true | 1466.0 1455.7–1476.3 | community | LMArena Leaderboard Dataset | |
| Multi Turn (style control)older version arena=text · category=multi_turn · style_control=true | max effort | 1465.6 1452.8–1478.5 | community | LMArena Leaderboard Dataset |
| Instruction Following (style control)older version arena=text · category=instruction_following · style_control=true | 1465.4 1458.9–1471.9 | community | LMArena Leaderboard Dataset | |
| Russian (style control)older version arena=text · category=russian · style_control=true | max effort | 1465.3 1449.4–1481.1 | community | LMArena Leaderboard Dataset |
| Overall (style control) arena=text · category=overall · style_control=true | max effort | 1464.3 1458.0–1470.6 | community | LMArena Leaderboard Dataset |
| Instruction Following (style control)older version arena=text · category=instruction_following · style_control=true | max effort | 1460.8 1451.7–1469.9 | community | LMArena Leaderboard Dataset |
| Industry Business And Management And Financial Operations (style control)older version arena=text · category=industry_business_and_management_and_financial_operations · style_control=true | 1459.0 1451.0–1467.1 | community | LMArena Leaderboard Dataset | |
| Non English (style control)older version arena=text · category=non_english · style_control=true | 1457.2 1451.6–1462.7 | community | LMArena Leaderboard Dataset | |
| Industry Writing And Literature And Language (style control)older version arena=text · category=industry_writing_and_literature_and_language · style_control=true | 1456.4 1449.1–1463.6 | community | LMArena Leaderboard Dataset | |
| Industry Business And Management And Financial Operations (style control)older version arena=text · category=industry_business_and_management_and_financial_operations · style_control=true | max effort | 1455.5 1444.0–1466.9 | community | LMArena Leaderboard Dataset |
| Polish (style control)older version arena=text · category=polish · style_control=true | 1454.7 1430.1–1479.2 | community | LMArena Leaderboard Dataset | |
| Japanese (style control)older version arena=text · category=japanese · style_control=true | 1453.0 1424.7–1481.3 | community | LMArena Leaderboard Dataset | |
| Non English (style control)older version arena=text · category=non_english · style_control=true | max effort | 1451.6 1443.9–1459.3 | community | LMArena Leaderboard Dataset |
| Creative Writing (style control)older version arena=text · category=creative_writing · style_control=true | 1451.4 1443.0–1459.8 | community | LMArena Leaderboard Dataset | |
| Industry Writing And Literature And Language (style control)older version arena=text · category=industry_writing_and_literature_and_language · style_control=true | max effort | 1451.1 1440.6–1461.5 | community | LMArena Leaderboard Dataset |
| Creative Writing (style control)older version arena=text · category=creative_writing · style_control=true | max effort | 1442.9 1430.5–1455.4 | community | LMArena Leaderboard Dataset |
| Industry Entertainment And Sports And Media (style control)older version arena=text · category=industry_entertainment_and_sports_and_media · style_control=true | 1441.2 1433.6–1448.7 | community | LMArena Leaderboard Dataset | |
| Korean (style control)older version arena=text · category=korean · style_control=true | 1436.7 1413.3–1460.1 | community | LMArena Leaderboard Dataset | |
| Industry Entertainment And Sports And Media (style control)older version arena=text · category=industry_entertainment_and_sports_and_media · style_control=true | max effort | 1434.1 1423.1–1445.0 | community | LMArena Leaderboard Dataset |
knowledge science
| GPQA Diamond | max effort | 91.9 % | independent | Epoch AI Benchmarking Hub 2026-06-24 |
| GPQA Diamond implementation=artificial-analysis | max effort | 89.5 % | independent | Artificial Analysis |
| GPQA Diamond | low effort | 87.9 % | independent | Epoch AI Benchmarking Hub 2026-08-10 |
| GPQA Diamond implementation=vals-ai | max effort | 85.6 % | independent | Vals AI |
| GPQA Diamond implementation=vals-ai | 85.6 % | independent | Vals AI | |
| GPQA Diamond | no reasoning | 71.2 % | independent | Epoch AI Benchmarking Hub 2026-08-10 |
| GPQA Diamond implementation=artificial-analysis | no reasoning | 68.6 % | independent | Artificial Analysis |
| Humanity's Last Exam implementation=artificial-analysis | max effort | 41.1 % | independent | Artificial Analysis |
| Humanity's Last Exam implementation=artificial-analysis | no reasoning | 9.8 % | independent | Artificial Analysis |
| MMLU-Pro implementation=vals-ai | 86.7 % | independent | Vals AI | |
| MMLU-Pro implementation=vals-ai | max effort | 86.7 % | independent | Vals AI |
language
| Release 2026-06-25 tasks_counted=3 · livebench_version=2026-06-25 | 76.2 % | independent | LiveBench |
long context instruction
| AA-LCR | max effort | 78.3 % | independent | Artificial Analysis |
| AA-LCR | no reasoning | 42.3 % | independent | Artificial Analysis |
| IFBench | max effort | 73.3 % | independent | Artificial Analysis |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 62.3 % | independent | LiveBench |
professional
| CorpFin | 66.1 % | independent | Vals AI | |
| CorpFin | max effort | 66.1 % | independent | Vals AI |
| LegalBench | max effort | 84.1 % | independent | Vals AI |
| LegalBench | 84.1 % | independent | Vals AI | |
| TaxEval | 73.3 % | independent | Vals AI | |
| TaxEval | max effort | 73.3 % | independent | Vals AI |
reasoning math
| ARC-AGI-1older version split=public_eval · model_type=CoT | 80.4 % | independent | ARC Prize Leaderboard | |
| ARC-AGI-1older version split=semi_private · model_type=CoT | 77.0 % | independent | ARC Prize Leaderboard | |
| ARC-AGI-2 split=semi_private · model_type=CoT | 22.8 % | independent | ARC Prize Leaderboard | |
| ARC-AGI-2 split=public_eval · model_type=CoT | 20.8 % | independent | ARC Prize Leaderboard | |
| ARC-AGI-2 split=semi_private · model_type=CoT | 0.3 usd_per_task | independent | ARC Prize Leaderboard | |
| ARC-AGI-2 split=public_eval · model_type=CoT | 0.2 usd_per_task | independent | ARC Prize Leaderboard | |
| ARC-AGI-1older version split=semi_private · model_type=CoT | 0.2 usd_per_task | independent | ARC Prize Leaderboard | |
| ARC-AGI-1older version split=public_eval · model_type=CoT | 0.2 usd_per_task | independent | ARC Prize Leaderboard | |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 89.8 % | independent | LiveBench | |
| Release 2026-06-25 tasks_counted=4 · livebench_version=2026-06-25 | 78.6 % | independent | LiveBench | |
| OTIS Mock AIME 2024–2025 | max effort | 86.4 % | independent | Epoch AI Benchmarking Hub 2026-06-25 |
| OTIS Mock AIME 2024–2025 | low effort | 75.6 % | independent | Epoch AI Benchmarking Hub 2026-08-10 |
| OTIS Mock AIME 2024–2025 | no reasoning | 28.9 % | independent | Epoch AI Benchmarking Hub 2026-08-10 |
Agent + model results
systems, not bare-model scores
| agent + model Epoch Inspect harness + GLM 5.2 | SWE-bench Verified | 78.7 % | independent | Epoch AI Benchmarking Hub |
| agent + model Artificial Analysis harness + GLM 5.2 | Terminal-Bench 2.1 | 77.9 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + GLM 5.2 | Terminal-Bench 2.1 | 51.7 % | independent | Artificial Analysis |
| agent + model Artificial Analysis harness + GLM 5.2 | Terminal-Bench Hard | 50.8 % | independent | Artificial Analysis |
These scores measure the whole agent system (scaffold, tools, budgets) — they are never merged into the bare model’s numbers.
