GLM-5.2: The Open Model Built for Long-Horizon Coding
Z.ai shipped a 744B MoE with 40B active, MIT-licensed weights and the first open-weight Terminal-Bench 2.1 score above 80. A technical read on what that means.
Most model launches are pitched at single-turn quality: give it a prompt, look at the answer. GLM-5.2 is pitched at something else entirely — tasks that run for hundreds of tool calls without a human in the loop. Z.ai released it in June 2026 and the benchmark selection makes the target obvious.
The number that got attention was Terminal-Bench 2.1. Z.ai reports 81.0 under the Terminus-2 harness, with a best-reported figure of 82.7. At its June 2026 release that was the first open-weight model past 80 on that evaluation, up from 62.0 for the previous generation; later models have since scored higher. A 19-point jump on a long-horizon agentic benchmark is a different kind of improvement from a couple of points on a knowledge quiz.
Architecture: large total, small active
GLM-5.2 is a mixture-of-experts model with roughly 744B total parameters and about 40B active per token. Z.ai's own release post cites 753B total; the discrepancy across sources is small enough not to matter, but it is a reminder to check which figure a comparison table is using.
What does matter is the ratio. Around 40B active parameters is a comparatively lean forward pass for a model in this capability class — DeepSeek V4 Pro activates 49B out of 1.6T, and Kimi K3 activates roughly 104B out of 2.8T. Fewer active parameters generally means cheaper serving and lower latency per token, which is exactly what you want when a single task might involve two hundred model calls.
Context is 1M tokens. Maximum output is 128K on the coding configuration and up to 163,840 tokens for reasoning work. There are two reasoning effort levels, high and xhigh, so you can trade depth for latency rather than being stuck with one setting.
The benchmark profile is unusually coherent
Z.ai's published numbers cluster around agentic software engineering rather than spreading thinly across everything:
- SWE-bench Pro — 62.1
- FrontierSWE — 74.4
- Terminal-Bench 2.1 — 81.0 (Terminus-2), 82.7 best reported
- DeepSWE — 46.2
- NL2Repo — 48.9
- ProgramBench — 63.7
- PostTrainBench — 34.3
- SWE-Marathon — 13.0
- MCP-Atlas (public set) — 76.8, Tool-Decathlon — 48.2
On the reasoning side it reports GPQA-Diamond 91.2, AIME 2026 99.2, HMMT February 2026 92.5, IMOAnswerBench 91.0, and Humanity's Last Exam at 40.5 without tools and 54.7 with them.
Two caveats. These are vendor-run numbers on vendor scaffolds, and vendor scaffolds routinely score several points above a standardised harness. And SWE-Marathon at 13.0 is worth staring at — it is a reminder that "long-horizon" has a ceiling, and that the very longest tasks remain unsolved by everything, not just this model.
On the Artificial Analysis Intelligence Index, GLM-5.2 sits around 51, which puts it near the top of the open-weight field and within range of the closed frontier without matching it.
MIT licensing is the quiet advantage
The weights are on Hugging Face under the MIT licence, with no regional restriction and no revenue threshold. That is a materially different proposition from Kimi K3, whose custom licence requires a separate agreement for Model-as-a-Service businesses above $20 million in annual revenue, and from MiniMax M3, which ships under a custom community licence.
If you are building a product that serves inference to customers, the licence is not a footnote — it decides whether you can ship at all. MIT means you can fine-tune, redistribute, and resell without asking anyone.
Serving is also more tractable than the headline parameter count suggests, because 40B active is what determines your throughput. It is still a serious deployment, but it is closer to feasible on a single well-specified node than a 2.8T model is.
Where it fits, and where it does not
The profile suggests a specific job: agent loops that run long, call many tools, and need to hold a plan together across a session. Repository-scale refactors, CI-driven fix loops, and terminal-heavy automation are the natural fit.
What the profile does not claim is a lead on one-shot generation quality or on the interactive feel of a chat session. Kimi K3 tops Arena's blind Frontend Code voting; that is a different property and GLM-5.2 does not compete for it. If your workload is a developer iterating in a chat window on UI code, the benchmark that predicts your experience is the human preference one, not Terminal-Bench.
Price sits between the cheap open models and the frontier. Z.ai's official API rates are around $1.40 per million input tokens and $4.40 per million output, with third-party providers listing lower. That is several times the cost of DeepSeek V4 Pro and a fraction of Kimi K3 or the closed frontier.
How to test the claim yourself
The benchmark that matters to you is the one built from your own repository. To check whether the long-horizon claim holds:
- Pick five tasks from your git history that took more than twenty tool calls to complete.
- Run each through the same agent harness with GLM-5.2 at high effort and again at xhigh.
- Record turns to completion, not just pass or fail. Long-horizon strength shows up as fewer wasted turns, not as a better first draft.
- Watch what happens after a failed test. Does it re-read the error and change approach, or restate the same fix?
If the answer is that it grinds through without losing the thread, the benchmark selection was honest. If it drifts after turn thirty, you have learned something no leaderboard would have told you.
Common questions
Is GLM-5.2 genuinely open source?
Yes, in the sense that matters commercially. The weights are published on Hugging Face under the MIT licence with no revenue thresholds or regional limits, so you can fine-tune, redistribute and resell without a separate agreement.
What does 744B total but 40B active actually mean for cost?
Total parameters set your memory requirement; active parameters set your compute per token. About 40B active is lean for this capability class, which is why serving costs and per-token latency are lower than the headline size implies.
Should I use high or xhigh reasoning effort?
Default to high and reserve xhigh for tasks where you have measured a real difference. Higher effort spends more output tokens and adds latency, which compounds badly across a long agent loop unless the task genuinely needs the depth.