Z.ai did not train a new model. GLM-5.3 keeps the exact base model of GLM-5.2, changes nothing about pretraining, and multiplies Terminal Bench 3.0 by six: 4.6 to 28.3. DeepSWE goes 46.2 to 66.9. Agents' Last Exam goes 23.8 to 28.5. Everything came from post-training in one month.
The mechanism is the story. As agent capability improves, the hard part of scaling moves from the model to the environment. A useful training environment has to be executable, verifiable, and close to real professional work, and you need thousands, not a handful of hand-built tasks. Some of the new environments represent several days of an experienced engineer's time, like being handed a full ML infrastructure stack with compute clusters, storage, docs, and experiment results, and having to deliver a measurable end-to-end speedup.
The load-bearing piece is the verifier. For each environment, a verifier is synthesized without ever seeing the reference solution. Before its reward is trusted, it must pass three checks: oracle, it accepts a known-good solution; no-op, it rejects doing nothing; unsolved-state, it rejects an outcome where the real task never got done. Then solver trajectories probe for reward shortcuts and close them. What survives emits a binary reward, zero or one, reliable enough to train on directly. Environment creation becomes a verification problem.
In this video:
why a frozen base model makes this release a controlled experiment
the environment factory: research agent, judge agent, reference-isolated verifiers
the three sanity checks and the reward hacks each one kills
what solver probes catch that the checks cannot
slime's single dataflow, and why environments plug in as data instead of code
the 1e-7 training-rollout logprob alignment and the 2.3x throughput gain
the cyber staircase: CyberGym 84.5, ExploitBench 54.4, ExploitGym 105 tasks in two hours
2,436 real vulnerabilities found, 26.6 years average time to discovery, 2,383 still under embargo
what to discount: private benchmark, modified anti-cheat checks, single-run evals, no ablation behind the word emergent
The cyber result deserves the skepticism. The capability grew fastest exactly where the gap to closed frontier models is widest, and the strongest real-world numbers are almost entirely embargoed. The video walks through which claims you can verify today and which you have to take on faith until the weights land.
For engineers and ML practitioners tracking how frontier agents actually improve.
Source: GLM-5.3: Frontier Coding with Emergent Cyber Capabilities, Z.ai, 2026 (z.ai/blog/glm-5.3).
More paper breakdowns in the Latent Papers playlist.