Glean 拾遗
Daily /2026-08-15 / GLM-5.3: Post-Training Scaling Boosts Coding and Cyber

GLM-5.3: Post-Training Scaling Boosts Coding and Cyber

Source z.ai Glean’d 2026-08-15 06:00 Read 22 min
AI summary

Z.ai announces GLM-5.3, claiming all gains come from post-training scaling on the same base model as GLM-5.2. The release post details environment-synthesis pipelines, reward verification, and the slime RL stack. Public coding/agent benchmarks improve markedly: Terminal-Bench 3.0 rises from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9. The standout claim is an "emergent cyber capability": ExploitBench jumps from 24.4 to 54.4, and real-world testing across 269 projects found 2,436 vulnerabilities, the oldest from 1981, now tracked in a public disclosure ledger. On the systems side, slime adds local storage caching, 1e-7 training-rollout logprob alignment, and workload-aware scheduling, yielding 2.3x throughput for long-horizon coding RL. The API removes thinking disabled and introduces reasoning_effort low/high/max. Weights arrive in two weeks after safety hardening. All scores are self-reported, and closed models still lead on several cyber and coding suites.

Original · 22 min
z.ai ↗
§ 1

Image 1

2026-08-14 · Research

* Image 2Z.ai Coding Plan* Image 3Code with ZCode* Image 4HuggingFace (Coming Soon)

Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them.

Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

  • Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam.
  • Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.
  • Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.

Image 1

2026-08-14 · 研究

* Image 2Z.ai Coding Plan* Image 3Code with ZCode* Image 4HuggingFace(即将上线)

GLM-5.3 的全部工作就是扩展后训练。 在 GLM-5.2 上我们搭好了整套技术栈:IndexShare 负责高效长上下文处理,SAO 负责长程任务的强化学习,slime 负责大规模异步训练——它们都跑在我们持续积累的长程任务环境上。过去一个月,我们在这套技术栈上继续扩展:更多的环境、更多样的任务、更多的训练算力。

今天发布的 GLM-5.3 与 GLM-5.2 使用同一个基座模型——所有提升都来自后训练。相比 GLM-5.2,它在复杂编码和长程任务上进步明显:

  • 更强的编码能力: GLM-5.3 是目前最强的开源权重编码模型,在我们内部的 Z.ai Code Bench 上比 GLM-5.2 提升 50%;在 Terminal Bench 3.0、Agents' Last Exam 等公开基准上也达到开源 SOTA。
  • 涌现的网络攻防能力: 随着后训练规模扩大,网络能力的成长速度超出预期。GLM-5.3 在 CyberGym 漏洞发现上达到 SOTA,而且越往利用链上游走提升越大,在漏洞利用基准上比 GLM-5.2 翻了一倍多。
  • 开源: 我们将于发布两周后、安全评估与加固完成后公开权重。
§ 2

Image 5: img_v3_0214i_9fec3a82-38c0-4674-8f7d-41f356c4fa6g

Performance across comparison models

Benchmark GLM-5.3 GLM-5.2 Kimi K3 DeepSeek-V4 Pro-0813 Qwen3.8-Max Opus 4.8 Fable 5 (w/ fallback) GPT-5.6 Sol
Coding
Terminal Bench 2.1 88.2 81.0 88.3 87.9 86.6 85.0 88.0 88.8
Terminal Bench 3.0 28.3 4.6 17.4 - - 21.1 33.7 34.6
DeepSWE v1.1 66.9 46.2 67.5 62.7 56.6 58.0 69.7 72.7
NL2Repo 58.0 48.9 58.0 61.1 55.9 69.7 - -
ProgramBench Almost Solved 19.0 9.5 17.5 - 10.5 15.5 33.0 23.0
FrontierSWE 78.1 67.5 - - - 66.5 88.2 -
SWE-Marathon v1.1 42.5 19.4 48.1 - - 48.8 33.1 42.5
PostTrainBench 39.8 31.7 32.0 - - 32.9 41.8 36.2
Cyber
CyberGym 84.5 77.2 80.0 83.3 78.5 78.1 83.8 83.6
ExploitGym 2h / 6h 105 / 130 29 / 39 36 / 70 - 14 / 26 80 / 120 181 / 247 216 / 293
ExploitBench 54.4 24.4 32.2 - 28.8 40.0 78.0 76.5
Agentic
Toolathlon Verified 73.0 59.9 76.5 74.1 72.5 76.2 74.7 74.9
AutomationBench v1.0.6 48.2 26.2 46.7 43.2 39.8 41.0 46.2 45.8
Agents' Last Exam ALE-CLI 28.5 23.8 27.6 25.7 27.0 25.7 23.8 28.6
HLE w/ Tools 62.5 54.7 59.8 60.0 56.2 57.9 63.9 64.5
GDPval-AA v2 1769 1508 1682 1590 1739 1588 1743 1730

Image 5: img_v3_0214i_9fec3a82-38c0-4674-8f7d-41f356c4fa6g

各模型性能对比

基准 GLM-5.3 GLM-5.2 Kimi K3 DeepSeek-V4 Pro-0813 Qwen3.8-Max Opus 4.8 Fable 5(带 fallback) GPT-5.6 Sol
编程
Terminal Bench 2.1 88.2 81.0 88.3 87.9 86.6 85.0 88.0 88.8
Terminal Bench 3.0 28.3 4.6 17.4 - - 21.1 33.7 34.6
DeepSWE v1.1 66.9 46.2 67.5 62.7 56.6 58.0 69.7 72.7
NL2Repo 58.0 48.9 58.0 61.1 55.9 69.7 - -
ProgramBench Almost Solved 19.0 9.5 17.5 - 10.5 15.5 33.0 23.0
FrontierSWE 78.1 67.5 - - - 66.5 88.2 -
SWE-Marathon v1.1 42.5 19.4 48.1 - - 48.8 33.1 42.5
PostTrainBench 39.8 31.7 32.0 - - 32.9 41.8 36.2
网络攻防
CyberGym 84.5 77.2 80.0 83.3 78.5 78.1 83.8 83.6
ExploitGym 2h / 6h 105 / 130 29 / 39 36 / 70 - 14 / 26 80 / 120 181 / 247 216 / 293
ExploitBench 54.4 24.4 32.2 - 28.8 40.0 78.0 76.5
智能体
Toolathlon Verified 73.0 59.9 76.5 74.1 72.5 76.2 74.7 74.9
AutomationBench v1.0.6 48.2 26.2 46.7 43.2 39.8 41.0 46.2 45.8
Agents' Last Exam ALE-CLI 28.5 23.8 27.6 25.7 27.0 25.7 23.8 28.6
HLE w/ Tools 62.5 54.7 59.8 60.0 56.2 57.9 63.9 64.5
GDPval-AA v2 1769 1508 1682 1590 1739 1588 1743 1730
§ 3

For GLM-5.3, we pushed environment scaling toward tasks that look less like coding exercises and more like real units of expert work. The environments now cover a much broader range of production workflows, with tasks designed around how engineering and research work is actually carried out in practice. Some represent several days of work for an experienced engineer. In an ML infrastructure task, for example, the model may be given the same working environment as an engineer, with access to compute clusters, storage systems, internal documentation, codebases, and experiment results. It must diagnose bottlenecks across the training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup while preserving correctness. Training on environments at this level pushes the model toward taking ownership of substantial work end to end, rather than relying on users to decompose the problem and supervise each step.

在 GLM-5.3 上,我们把环境扩展推向了那些不像编程练习、而更像真实专家工作单元的任务。环境覆盖的生产工作流范围广得多,任务设计也围绕工程和研究在实际中如何展开。有些任务相当于一位资深工程师数天的工作量。例如在一个 ML 基础设施任务中,模型会拿到与工程师相同的工作环境,包括计算集群、存储系统、内部文档、代码库和实验记录。它必须诊断训练栈中的瓶颈、实现优化、运行实验,并在保证正确性的前提下交付可衡量的端到端加速。在这种级别的环境上训练,会推动模型对完整工作承担起端到端的责任,而不是依赖用户拆解问题、逐步监督。

§ 4

As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment. A useful task environment has to be executable, verifiable, and close to real professional work — and we need many of them, not a handful of hand-built ones. To scale this process, we built pipelines that synthesize environments end to end, and for a subset of tasks, the RL reward signal as well. Research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state; a judge agent then attempts each task to verify that it is actually solvable. Verifiers are synthesized without access to the reference solution, while solver trajectories are used to discover and close reward shortcuts. A verifier that passes oracle, no-op, and unsolved-state checks produces a binary reward reliable enough to train on directly.

智能体能力提升后,后训练扩展的难点大多从模型转移到了环境。一个有用的任务环境必须可执行、可验证、贴近真实专业工作——而且数量要足够多,不能只靠手工打造几个。为了规模化这个过程,我们构建了端到端合成环境的流水线,对部分任务还一并合成 RL 奖励信号。研究智能体从真实工作中收集任务模式,把它们变成可运行的长程环境,带有多步依赖和隐藏状态;随后一个裁判智能体会试做每个任务,确认它确实可解。验证器在合成时不接触参考解法,求解轨迹则用来发现并堵住奖励捷径。一个能通过 oracle、no-op 和未解状态检查的验证器,会产生足够可靠的二元奖励,可以直接用于训练。

§ 5

It carries over the RL strategies introduced in GLM-5.2, including SAO with compaction, which helps these gains hold on long-horizon tasks rather than only on short ones. The effect shows up across both coding and general agent tasks. GLM-5.3 improves from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 23.8 to 28.5 on Agents' Last Exam. These pipelines still require a meaningful amount of human-in-the-loop work; making environment generation and verification more autonomous is one of the next steps.

这套方法沿用了 GLM-5.2 引入的 RL 策略,包括带压缩的 SAO;它让这些收益在长程任务上也能保持,而不只是体现在短任务上。效果同时体现在编码任务和通用智能体任务上。GLM-5.3 在 Terminal-Bench 3.0 上从 4.6 提升到 28.3,在 DeepSWE v1.1 上从 46.2 提升到 66.9,在 Agents' Last Exam 上从 23.8 提升到 28.5。这些流水线仍然需要相当多的人工参与;让环境生成与验证更加自动化是下一步工作之一。

§ 6

Beyond public benchmarks, we introduce Z.ai Code Bench, an in-house benchmark designed to evaluate coding agents under realistic user scenarios. It covers diverse task categories and places agents in complex local development environments. At different effort levels, we evaluate agents along two dimensions: end-to-end task completion rate and fine-grained checklist accuracy. As a private benchmark, Z.ai Code Bench also reduces the risk of contamination from public test sets and gives us a more faithful measure of real-world user experience.

Image 6: img_v3_0214i_ee4cacaa-cd74-4e24-a430-2c3e09e30c9g

As shown in the figure, GLM-5.3 improves both performance and token efficiency. It delivers markedly stronger agentic coding results than GLM-5.2 at every effort level while consuming fewer output tokens. At Max effort, GLM-5.3 reaches 34.5% at roughly 75K output tokens per task, compared with 23.4% at 96K for GLM-5.2. The same shift holds against closed models. At High effort, GLM-5.3 reaches 31.4% at around 50K output tokens, surpassing Claude Opus 4.8 at 29.5% with 120K. GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort.

在公开基准之外,我们推出了内部基准 Z.ai Code Bench,用于在贴近真实用户场景下评估编码智能体。它覆盖多样的任务类别,让智能体身处复杂的本地开发环境。我们按不同投入程度,从两个维度评估智能体:端到端任务完成率和细粒度清单准确率。作为私有基准,Z.ai Code Bench 还降低了公开测试集污染的风险,能更真实地反映实际用户体验。

Image 6: img_v3_0214i_ee4cacaa-cd74-4e24-a430-2c3e09e30c9g

如图所示,GLM-5.3 在性能和 token 效率上都有提升。它在每个投入档位下都比 GLM-5.2 交出明显更强的智能体编码结果,同时消耗更少的输出 token。在 Max 档,GLM-5.3 每个任务约 75K 输出 token 即达到 34.5%,而 GLM-5.2 用 96K 才到 23.4%。同样的转变也适用于闭源模型。在 High 档,GLM-5.3 用约 50K 输出 token 达到 31.4%,超过 Claude Opus 4.8 的 29.5%(120K)。GLM-5.3 仍落后于 Claude Fable 5,后者在 Max 档达到 39.5%。

§ 7

As part of post-training, we introduced vulnerability discovery data and environments into the training mix. We expected this to make the model better at finding and reasoning about vulnerabilities. What surprised us was how quickly the capability continued to develop as training scaled. GLM-5.3 did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains.

Image 7: img_v3_0214i_d001f597-8f56-48f0-bf15-f22ce6be653g

在后训练中,我们把漏洞发现数据和环境加入了训练组合。我们原本预期,模型在发现和推理漏洞方面会变得更好。让我们意外的是,随着训练规模扩大,这项能力成长得如此之快。GLM-5.3 不只是更擅长识别孤立的缺陷,它开始跨多个利用阶段进行推理,为完整的利用链制定连贯计划。

Image 7: img_v3_0214i_d001f597-8f56-48f0-bf15-f22ce6be653g

§ 8

We evaluate GLM-5.3 across three benchmarks covering different stages of vulnerability analysis and exploitation. On CyberGym, which starts from white-box source code and tests whether the model can identify and validate vulnerabilities by triggering faults, GLM-5.3 scores 84.5%, up from GLM-5.2's 77.2% — the best result on the benchmark, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). On ExploitBench, which requires deeper reasoning about real vulnerabilities and their exploitation, GLM-5.3 reaches 54.4%, more than doubling GLM-5.2's 24.4%, while Mythos 5 and GPT-5.6 Sol score 78.0% and 76.5%, respectively. On ExploitGym, which measures how many exploitation tasks a model can complete under time-normalized budgets, GLM-5.3 completes 105 tasks within two hours and 130 within six hours, compared with 29 and 39 for GLM-5.2; budgets are normalized across models using per-model throughput figures, detailed in the footnotes. Mythos 5 remains well ahead at 181 and 247 tasks. The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2 — and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.

我们用三个覆盖漏洞分析与利用不同阶段的基准来评估 GLM-5.3。CyberGym 从白盒源代码出发,测试模型能否通过触发故障来发现并验证漏洞;GLM-5.3 得到 84.5%,高于 GLM-5.2 的 77.2%——这是该基准的最好成绩,领先 Mythos 5(83.8%)和 GPT-5.6 Sol(83.6%)。ExploitBench 要求对真实漏洞及其利用做更深层的推理,GLM-5.3 达到 54.4%,是 GLM-5.2(24.4%)的两倍多,而 Mythos 5 和 GPT-5.6 Sol 分别为 78.0% 和 76.5%。ExploitGym 衡量模型在按时间归一化的预算内能完成多少利用任务;GLM-5.3 在两小时内完成 105 个任务,六小时内完成 130 个,而 GLM-5.2 分别是 29 和 39;预算按各模型自身吞吐量归一化,细节见脚注。Mythos 5 仍遥遥领先,为 181 和 247 个。三个基准呈现出一致的规律:基准越靠近利用链上游,相对 GLM-5.2 的提升越大——与闭源前沿的差距也越大。能力增长最快的,恰恰是我们落后最多的地方。

§ 9

We then tested whether these capabilities transfer beyond controlled benchmarks. Since GLM-5.2, we have been working with several security teams in China to run our models against real-world codebases. After expert review, screening, and deduplication, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues. The findings span system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols. Many had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years.

接下来我们检验了这些能力能否迁移到受控基准之外。自 GLM-5.2 以来,我们一直与中国多个安全团队合作,让模型在真实代码库上运行。经过专家复核、筛选和去重,模型在 269 个项目中识别出 2,436 个漏洞,其中 1,097 个为中高危问题。发现范围涵盖系统内核、操作系统、浏览器引擎、开源基础设施、Web 应用和网络协议。许多漏洞多年来甚至数十年都未被注意,最久远的可追溯到约 40 年前。

§ 10

This work has since grown into an ongoing disclosure effort. We built the Z.ai Security Disclosure Ledger to maintain a public record of the findings as they move through the disclosure process. The ledger is continuously updated as new vulnerabilities are reviewed and disclosed, distinguishing issues that have already been made public from those that are still under disclosure. For disclosed issues, it records information including the affected project, severity, CVE where available, and how long the vulnerability had remained in the codebase.

2,436

FINDINGS TRACKED

53

PUBLICLY DISCLOSED

2,383

UNDER EMBARGO

1,097

CRITICAL & HIGH

269

OSS PROJECTS

45

YEARS OF IMPACT

Findings span 45 years of impact - the oldest flaw was introduced in 1981, and on average a vulnerability lived 26.6 years before discovery.

SEVERITY DISTRIBUTION

Critical 107 High 990 Medium 1,286 Low 53

WHEN THE FLAWS WERE INTRODUCED

1981 2026

ViewZ.ai Security Disclosure Ledger ↗

这项工作此后发展成一项持续进行的披露行动。我们建立了 Z.ai Security Disclosure Ledger,在发现进入披露流程时公开记录各项结果。随着新漏洞被评审和披露,台账会持续更新,并区分已经公开的问题和仍在保密期的问题。对于已披露的问题,它会记录受影响项目、严重程度、可用的 CVE 编号,以及漏洞在代码库中已存在多久。

2,436

已追踪发现

53

已公开披露

2,383

处于保密期

1,097

严重/高危

269

开源项目

45

影响年数

发现横跨 45 年的影响——最古老的缺陷引入于 1981 年,漏洞在被发现前平均潜伏 26.6 年。

严重度分布

严重 107 高危 990 中危 1,286 低危 53

缺陷引入时间

1981 2026

查看 Z.ai Security Disclosure Ledger ↗

§ 11

All of this runs on slime, our open-source post-training framework for RL scaling, with Megatron on the training side and SGLang on the rollout side. Its design keeps training, rollout, and the data buffer on a single dataflow, so math, code, sandboxes, verifiers, and long-horizon agentic environments plug in as data generation rather than as changes to the training loop. That is what let us keep adding environments through GLM-5.2 and GLM-5.3 without rebuilding the training stack each time.

这一切都跑在 slime 上——我们开源的 RL 扩展后训练框架,训练侧用 Megatron,推理侧用 SGLang。它的设计把训练、推理和数据缓冲放在同一条数据流里,因此数学、代码、沙箱、验证器和长程智能体环境都是作为数据生成接入,而不是改动训练循环本身。这正是我们能在 GLM-5.2 和 GLM-5.3 中不断新增环境、却无需每次重建训练栈的原因。

§ 12

Through GLM-5.3 we kept building it out on two fronts. On the algorithmic side we added capabilities aimed at RL research: top-p mask, top-k and full-vocabulary OPD, and configurations that improve training–rollout consistency, including R3-style setups and full numerical alignment between the training and rollout paths, which give us finer control over sampling, training, and teacher signals, and make it fast to run controlled comparisons. In our training–rollout consistency evaluation, the average difference in log probabilities (logprob) was controlled at the 1e-7 level, representing a reduction of more than 99.99% compared with previous setups.

在 GLM-5.3 的开发中,我们从两个方向继续完善它。算法侧新增了面向 RL 研究的能力:top-p mask、top-k 与全词表 OPD,以及改进训练-推理一致性的配置,包括 R3 风格设置和训练路径与推理路径的完全数值对齐。这让我们能更精细地控制采样、训练和教师信号,也能快速做受控对比。在我们的训练-推理一致性评估中,log 概率(logprob)的平均差异被控制在 1e-7 量级,相比此前设置减少了 99.99% 以上。

§ 13

We also worked on resource efficiency and system throughput for large-scale RL. Local storage now serves as an additional caching layer, holding model states and data hierarchically that would otherwise sit in host memory. This matters most for multi-teacher OPD: with dynamic teacher switching and prefetching on the training side, several teachers can be used without standing up a dedicated long-running inference service for each, at limited added overhead and substantially lower resource consumption. For agentic and asynchronous workloads, we improved joint scheduling and load balancing between the router and slime, so that rollout requests with widely varying lengths and completion times make better use of inference resources. We added workload-aware heuristics that derive throughput-oriented configurations — prefill/decode resource ratio, concurrency settings, and other throughput-critical parameters — from the characteristics of each rollout environment. As a result, for long-horizon coding RL tasks, these system-level optimizations improved end-to-end RL training throughput by more than 2.3×, allowing us to scale training over longer trajectories and more complex environments with substantially higher efficiency.

Taken together, these give us more experimental flexibility, lower resource cost, and higher throughput — which is what makes it practical to keep scaling RL.

我们还着力于大规模 RL 的资源效率和系统吞吐。本地存储现在作为额外缓存层,把本应驻留在主机内存中的模型状态和数据分层保存。这对多教师 OPD 尤其重要:借助训练侧动态切换教师和预取,无需为每个教师单独常驻一个推理服务,就能同时使用多个教师,新增开销有限,资源消耗大幅降低。针对智能体和异步工作负载,我们改进了 router 与 slime 之间的联合调度和负载均衡,让长度和完成时间差异很大的推理请求更充分地利用推理资源。我们还加入了工作负载感知启发式,根据每个推理环境的特点自动推导面向吞吐的配置——prefill/decode 资源比例、并发设置以及其他吞吐关键参数。结果是,在长程编码 RL 任务上,这些系统级优化将端到端 RL 训练吞吐提升了 2.3 倍以上,让我们能以高得多的效率扩展更长轨迹和更复杂环境的训练。

综合起来,我们获得了更强的实验灵活性、更低的资源成本和更高的吞吐——这正是持续扩展 RL 变得切实可行的原因。

§ 14

GLM-5.3 supports three thinking effort levels: low, high, and max. Disabling thinking is no longer supported by GLM-5.3.

Thinking Parameters

Parameter Values Default Description
thinking.type enabled enabled Enables thinking. disabled is no longer supported.
reasoning_effort low, high, max max low: light; high: enhanced; max: deep.

max is recommended for coding tasks.

{
  "model": "glm-5.3",
  "thinking": { "type": "enabled" },
  "reasoning_effort": "max"
}

Migration required: If your application currently uses thinking.type: "disabled", change it to enabled and set reasoning_effort to low before updating the model ID to glm-5.3. Otherwise, the request will fail.

GLM-5.3 支持三档思考强度:lowhighmax。GLM-5.3 不再支持关闭思考。

思考参数

参数 取值 默认值 说明
thinking.type enabled enabled 启用思考。不再支持 disabled
reasoning_effort lowhighmax max low:轻度;high:加强;max:深度。

编码任务建议使用 max

{
  "model": "glm-5.3",
  "thinking": { "type": "enabled" },
  "reasoning_effort": "max"
}

需要迁移:如果你的应用当前使用 thinking.type: "disabled",请先改为 enabled 并将 reasoning_effort 设为 low,然后再把模型 ID 更新为 glm-5.3。否则请求将失败。

§ 15

Try GLM-5.3 in your favorite coding agents—ZCode, Claude Code, OpenCode, and more. https://docs.z.ai/devpack/overview

For GLM Coding Plan subscribers: We’ve rolled out GLM-5.3 to all GLM Coding Plan users. The new GLM Coding Plan now uses a points-based quota system. Point usage is calculated separately for input, cached input, and output tokens. Model calls made outside peak hours consume 50% of the standard points. Peak hours are 14:00–18:00 (UTC+8), Monday through Friday; all other hours, including weekends, receive the 50% off-peak rate. Start building now: https://z.ai/subscribe

Get more from GLM-5.3 with ZCode

  • 98%+ cache hit rate — repeated context billed at the lower cached rate, ~30% more effective tokens;
  • 1.5x limited-time quota boost — stack it with the cache savings for up to 180% your standard quota through August 31.
  • Long-horizon mastery — Goal mode plans, codes, tests, and verifies until the target is met;
  • Remote Control — monitor and steer long-running tasks from your phone via WeChat or Feishu.

Try ZCode: https://zcode.z.ai

Serve GLM-5.3 Locally

The model weights of GLM-5.3 will be publicly available soon in two weeks.

在您常用的编码智能体——ZCode、Claude Code、OpenCode 等——中试试 GLM-5.3https://docs.z.ai/devpack/overview

GLM Coding Plan 订阅用户: 我们已向所有 GLM Coding Plan 用户开放 GLM-5.3。新版 GLM Coding Plan 采用基于点数的额度系统,输入、缓存输入和输出 token 分别计点。非高峰时段调用仅消耗标准点数的 50%。高峰时段为北京时间周一至周五 14:00–18:00;其余时间(含周末)均按 50% 的低峰费率计点。立即开始:https://z.ai/subscribe

用 ZCode 充分释放 GLM-5.3

  • 缓存命中率 98%+——重复上下文按更低的缓存费率计费,有效 token 多约 30%;
  • 限时 1.5 倍额度加成——叠加缓存节省,到 8 月 31 日前最高可达标准额度的 180%。
  • 长程任务掌控——Goal 模式会持续规划、编码、测试和验证,直到目标达成;
  • 远程控制——通过微信或飞书在手机上监控和引导长时任务。

试试 ZCode:https://zcode.z.ai

本地部署 GLM-5.3

GLM-5.3 的模型权重将于两周后公开可用。

§ 16

Footnotes

  • HLE w/ tools: We use sampling parameters of temperature=1.0 and top_p=0.95 for evaluation, with a maximum generation length of 163,840 tokens. The evaluation is conducted with a maximum context length of 300,000 tokens, using a context management strategy. We use GPT-5.6-luna (medium) as the judge model.
  • NL2Repo: We evaluated NL2Repo with temperature=1.0, top_p=1.0, and max_new_tokens=64k under 1M context. To prevent hacking, we use rule-based and a LLM-based judgement to prevent malicious behaviors (e.g., unauthorized pip or curl operations).
  • DeepSWE: We run DeepSWE using the mini-swe-agent harness with temperature=0.95, top_p=1.0, timeout=6h and 400K context.
  • Terminal-Bench 2.1: We evaluate in Claude Code 2.1.207 with temperature=1.0, top_p=1, max_new_tokens=65536 with 6h timeout.
  • Terminal-Bench 3.0: We evaluate Terminal-Bench-3 tasks with the Claude Code 2.1.207 harness (reasoning effort=max, 400K context, and 128K maximum output), reporting avg@3 over three rollouts per task. Each rollout runs in an isolated container built from the task's official image, and is capped at 600 agent turns with a 10-hour timeout. Tool Search is disabled, and the artifacts each agent produces are scored by the task's official separate verifier.
  • Agent’s Last Exam (CLI): We evaluate ALE using the official evaluation protocol with the Claude Code harness (reasoning effort=max, 1M context, and 64K maximum output). Each of the 105 tasks runs in an isolated Docker container using the resources declared in its Task Card. The default timeout is 4 hours, with task-specific limits taking precedence (up to 8 hours). Tool Search is disabled, and results are scored by the official ALE evaluators.
  • Toolathlon Verified: We obtain all results via the official evaluation service and report pass@1 averaged over 3 independent runs.
  • AutomationBench: We evaluate on AutomationBench v1.0.6, incorporating the fix for the null-type handling issue introduced in PR #13.
  • GDPval-AA v2: Models are evaluated by Artificial Analysis.
  • CyberGym: We evaluate GLM-5.3 in Claude Code 2.1.207 (max reasoning effort, no web tools with temperature=1.0, top_p=1.0, max_new_tokens=128000). All evaluations are under unlimited timeout per task and results are single-run Pass@1 over 1,507 tasks. To simulate real-world usage scenarios, we place the agent inside the task container. We also remove all Git-related information and apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
  • ExploitGym: We evaluate GLM-5.3, Kimi-K3 and Qwen3.8 Max in Claude Code 2.1.207 (max reasoning effort, no web tools with temperature=1.0, top_p=1.0, max_new_tokens=128000). The reported results are single-run Pass@1 on 869 tasks under two timeout budgets: 2 hours and 6 hours, which are calculated as the API inference time rescaled by per-model tokens per second rate (per-model TPS sourced from Artificial Analysis; that is, we rescale GLM-5.3's results by 115 TPS, Kimi K3's results by 40 TPS and Qwen3.8 Max's results by 47 TPS), plus the non-API overhead. We also apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
  • ExploitBench: We evaluate GLM-5.3 in Claude Code 2.1.207 (max reasoning effort, no web tools with temperature=1.0, top_p=1.0, max_new_tokens=128000). Following the official evaluation settings, we limit the maximum number of interaction rounds between the agent and the environment to 300, and compute the average coverage score over all 41 tasks across 3 revisions. The coverage result of a task is determined by taking the union of capabilities achieved across all revisions, and the average score is obtained by averaging the results. We also apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
  • FrontierSWE: The evaluation was conducted by Proximal with 1M context length, max effort level, and 128K maximum output tokens. Dominance score reported as of 2026/08/14.
  • PostTrainBench: We evaluate GLM-5.3 using Claude Code 2.1.207 with max effort level, temperature = 1.0, top_p = 1.0, max_new_tokens = 128000, and a 1M-token context window. We report the weighted average over 3 runs. Runs that fail to produce a score fall back to the official zero-shot base-model baseline score. For checks intended to prevent the use of third-party APIs, we removed the original pattern-matching-based checks, as they produced false positives when a local vLLM endpoint was accessed through the OpenAI SDK. Instead, we use an LLM agent to inspect solutions for external API usage.
  • SWE-Marathon: We evaluate GLM-5.3 using Claude Code 2.1.207 with maximum effort level, temperature = 1.0, top_p = 0.95, max_new_tokens = 128000, and a 1M-token context window. For strip-clone, the original anti-cheat checks used overly broad import detection that could reject valid implementations. We removed the affected checks and performed llm-based inspection instead to avoid false positives. For parameter-golf and trimul-cuda, changes to the NVIDIA wheels caused the Docker image builds to fail, so we added --extra-index-url https://pypi.org/simple to restore successful builds.

Image 8

脚注

  • HLE w/ tools:我们使用 temperature=1.0top_p=0.95 的采样参数进行评估,最大生成长度为 163,840 tokens。评估在最大上下文长度 300,000 tokens 下进行,并采用上下文管理策略。我们使用 GPT-5.6-luna(medium)作为裁判模型。
  • NL2Repo:我们在 1M 上下文下,以 temperature=1.0、top_p=1.0、max_new_tokens=64k 评估 NL2Repo。为防止作弊,我们使用基于规则和基于 LLM 的判断来阻止恶意行为(例如未经授权的 pip 或 curl 操作)。
  • DeepSWE:我们使用 mini-swe-agent harness 运行 DeepSWE,设置为 temperature=0.95top_p=1.0timeout=6h 和 400K 上下文。
  • Terminal-Bench 2.1:我们在 Claude Code 2.1.207 中进行评估,temperature=1.0、top_p=1、max_new_tokens=65536,超时 6 小时。
  • Terminal-Bench 3.0:我们使用 Claude Code 2.1.207 harness(reasoning effort=max、400K 上下文、128K 最大输出)评估 Terminal-Bench-3 任务,报告每个任务三次 rollout 的 avg@3。每次 rollout 在基于任务官方镜像构建的隔离容器中运行,上限为 600 个智能体回合,超时 10 小时。Tool Search 被禁用,每个智能体产生的工件由任务的官方独立验证器评分。
  • Agent’s Last Exam (CLI):我们使用官方评估协议和 Claude Code harness(reasoning effort=max、1M 上下文、64K 最大输出)评估 ALE。105 个任务中的每一个都在隔离的 Docker 容器中运行,使用其 Task Card 中声明的资源。默认超时为 4 小时,任务特定限制优先(最长 8 小时)。Tool Search 被禁用,结果由官方 ALE 评估器评分。
  • Toolathlon Verified:我们通过官方评估服务获得所有结果,报告 3 次独立运行平均的 pass@1。
  • AutomationBench:我们评估的是 AutomationBench v1.0.6,其中包含 PR #13 中对 null 类型处理问题的修复。
  • GDPval-AA v2:模型由 Artificial Analysis 评估。
  • CyberGym:我们在 Claude Code 2.1.207(max reasoning effort、无 web 工具,temperature=1.0、top_p=1.0、max_new_tokens=128000)中评估 GLM-5.3。所有评估均不限制每个任务的超时时间,结果是在 1,507 个任务上的单次运行 Pass@1。为了模拟真实使用场景,我们把智能体放进任务容器内。我们还移除了所有 Git 相关信息,并应用域名白名单(仅允许 pypi.orgdeb.debian.org 等基本工具安装所需域名),防止智能体作弊。
  • ExploitGym:我们在 Claude Code 2.1.207(max reasoning effort、无 web 工具,temperature=1.0、top_p=1.0、max_new_tokens=128000)中评估 GLM-5.3、Kimi-K3 和 Qwen3.8 Max。报告结果是 869 个任务在 2 小时和 6 小时两个超时预算下的单次运行 Pass@1;预算按 API 推理时间乘以各模型每秒 token 速率重新换算(各模型 TPS 来自 Artificial Analysis;即 GLM-5.3 按 115 TPS、Kimi K3 按 40 TPS、Qwen3.8 Max 按 47 TPS 重新换算),再加上非 API 开销。我们还应用域名白名单(仅允许 pypi.orgdeb.debian.org 等基本工具安装所需域名),防止智能体作弊。
  • ExploitBench:我们在 Claude Code 2.1.207(max reasoning effort、无 web 工具,temperature=1.0、top_p=1.0、max_new_tokens=128000)中评估 GLM-5.3。按照官方评估设置,我们限制智能体与环境的交互轮数最多为 300,并计算全部 41 个任务在 3 个版本上的平均覆盖率得分。任务的覆盖率结果取所有版本实现能力上限的并集,平均得分由各结果取平均得到。我们还应用域名白名单(仅允许 pypi.orgdeb.debian.org 等基本工具安装所需域名),防止智能体作弊。
  • FrontierSWE:评估由 Proximal 完成,上下文长度 1M、最高 effort 档、最大输出 128K tokens。支配得分的报告日期为 2026/08/14。
  • PostTrainBench:我们使用 Claude Code 2.1.207 评估 GLM-5.3,设置为 max effort、temperature = 1.0、top_p = 1.0、max_new_tokens = 128000,以及 1M-token 上下文窗口。我们报告 3 次运行的加权平均。未能产生分数的运行会回退到官方零样本基座模型基线。对于旨在防止使用第三方 API 的检查,我们移除了原有的模式匹配检查,因为当本地 vLLM 端点通过 OpenAI SDK 访问时它们会产生误报;取而代之的是用 LLM 智能体检查解决方案是否使用了外部 API。
  • SWE-Marathon:我们使用 Claude Code 2.1.207 评估 GLM-5.3,设置为最高 effort 档、temperature = 1.0、top_p = 0.95、max_new_tokens = 128000,以及 1M-token 上下文窗口。对于 strip-clone,原有的反作弊检查使用了过于宽泛的 import 检测,可能拒绝有效实现。我们移除了受影响的检查,改用基于 LLM 的检查来避免误报。对于 parameter-golftrimul-cuda,NVIDIA wheels 的变更导致 Docker 镜像构建失败,因此我们添加了 --extra-index-url https://pypi.org/simple 来恢复成功构建。

Image 8

Open source ↗