Most powerful Z.ai model family. GLM-5 is 744B total params (40B active) via MoE with 256 experts. Hybrid Attention and Multi-Token Prediction. First frontier-scale model trained entirely on 100,000 Huawei Ascend 910B chips. GLM-5-Turbo is the fast variant optimized for the OpenClaw agent ecosystem.

Outputs 3

GLM-5: From Vibe Coding to Agentic Engineering

model

744B total params (40B active) via MoE with 256 experts (8 routed per token; ~754B by safetensors count). Hybrid Attention and Multi-Token Prediction. First frontier-scale model trained entirely on 100,000 Huawei Ascend 910B chips (zero American hardware).

Architecture MOE
Parameters 744B
Active params 40B
Training tokens 28.5T
AA Intelligence 28 was 32 on v4.2

GLM-5 Technical Report

paper

GLM-5-Turbo

model

Specialized "fast" version optimized for the OpenClaw agent ecosystem, focusing on continuous task execution and tool reliability.

Architecture MOE
AA Intelligence 27 was 31 on v4.2
moeagenticcodingfrontierefficiency