GLM-5.3
modelYour notes
Z.ai's agentic coding and cybersecurity successor to GLM-5.2. It uses the same base model — a 744B MoE (40B active) with a 1M-token context — and attributes every gain to one month of scaled post-training across more numerous, diverse, and longer-horizon task environments. The training stack combines IndexShare for long-context processing, SAO with compaction for long-horizon RL, and the open-source slime framework; new scheduling, caching, and workload-aware configurations improved end-to-end RL throughput by more than 2.3×.
Z.ai reports large gains over GLM-5.2 on complex work: Terminal-Bench 3.0 rises from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5. Cybersecurity environments produced especially steep improvements: 84.5% on CyberGym, 54.4% on ExploitBench (versus 24.4%), and 105/130 completed ExploitGym tasks under normalized two-/six-hour budgets (versus 29/39). The API requires thinking and offers low, high, and max reasoning effort. Open weights followed on 2026-08-28 (FP8 GLM-5.3 and BF16 GLM-5.3-BF16 on HuggingFace, ~753B by safetensors count) under the custom GLM-5.3 License — MIT-style permissions with one added condition: a licensee whose affiliates run a "Model as a Service" business with more than US$10B in aggregate revenue over any 12 months must pass a Z.AI security review before commercial use.
In a launch-day Interconnects post, Nathan Lambert reads GLM-5.3 — a text-only ~750B model, roughly a third the size of Kimi K3, that surpasses K3 on many benchmarks and matches or beats Claude Fable 5 and GPT-5.6-Sol on some agentic-coding evals — as evidence that Chinese labs keep stride with the frontier through means other than distillation: release cycles measured in days rather than months, heavier optimization toward public benchmarks under fundraising pressure, narrower specialization on high-value use cases like coding, an emerging domestic RL-environment data industry, and strong organizational talent and compute efficiency.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| Terminal-Bench 2.1 | 88.2 | — |
| Terminal-Bench 3.0 | 28.3 | — |
| DeepSWE v1.1 | 66.9 | — |
| FrontierSWE | 78.1 | — |
| CyberGym | 84.5 | — |
| ExploitBench | 54.4 | — |
| Agents' Last Exam | 28.5 | — |