Runs on Data centre · Hugging Face ↗
Coding index 55.8 (Artificial Analysis)
The previous Z.ai flagship, ~754B MoE — useful as the generation-gap reference against GLM-5.2 (which scores ~13 points higher on coding).
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈42 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q2_K | 252.4 GB |
| Q3_K_M | 338.4 GB |
| IQ4_XS | 361.3 GB |
| Q4_K_S | 433.9 GB |
| MXFP4 | 450.9 GB |
| Q4_K_M | 464.5 GB |
| Q4 | 466.0 GB |
| Q5_K_M | 558.5 GB |
| Q8_0 | 801.3 GB |
| Q8 | 810.6 GB |
| Q6_K | 1305.7 GB |
| BF16 | 1508.0 GB |
~465 GB at Q4 — data-centre. If you can host a GLM at all, host 5.2 instead; 5.1 is here for comparison, not as the pick.