GLM-5.3: Post-Training Scaling Boosts Coding and Cyber
Z.ai announces GLM-5.3, claiming all gains come from post-training scaling on the same base model as GLM-5.2. The release post details environment-synthesis pipelines, reward verification, and the slime RL stack. Public coding/agent benchmarks improve markedly: Terminal-Bench 3.0 rises from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9. The standout claim is an "emergent cyber capability": ExploitBench jumps from 24.4 to 54.4, and real-world testing across 269 projects found 2,436 vulnerabilities, the oldest from 1981, now tracked in a public disclosure ledger. On the systems side, slime adds local storage caching, 1e-7 training-rollout logprob alignment, and workload-aware scheduling, yielding 2.3x throughput for long-horizon coding RL. The API removes thinking disabled and introduces reasoning_effort low/high/max. Weights arrive in two weeks after safety hardening. All scores are self-reported, and closed models still lead on several cyber and coding suites.