Runs on Multi-GPU · Hugging Face ↗
AA coding index 68.8 (Artificial Analysis, frozen since 2026-09-11)
Z.ai's flagship — a very large Mixture-of-Experts (hundreds of billions of total params). One of the strongest open models by coding index, but firmly data-centre scale to run locally, even at Q4 (hundreds of GB).
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈6 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ1_S | 216.7 GB |
| IQ1_M | 228.5 GB |
| IQ2_XXS | 238.5 GB |
| IQ2_M | 238.6 GB |
| Q2_K_XL | 253.9 GB |
| IQ3_XXS | 281.7 GB |
| IQ3_S | 308.6 GB |
| Q3_K_M | 342.7 GB |
| Q3_K_XL | 343.0 GB |
| IQ4_XS | 365.3 GB |
| IQ4_NL | 372.7 GB |
| Q4_K_S | 436.4 GB |
| Q4_K_M | 465.8 GB |
| Q4_K_XL | 467.3 GB |
| Q5_K_S | 527.3 GB |
| Q5_K_M | 560.8 GB |
| Q5_K_XL | 562.5 GB |
| Q6_K | 625.9 GB |
| Q6_K_XL | 684.4 GB |
| Q8_0 | 801.4 GB |
| Q8_K_XL | 819.7 GB |
| BF16 | 1508.0 GB |
Multi-GPU or a serious server only — a Q4 build is still hundreds of GB of weights plus a large KV cache. Serve with vLLM or SGLang tensor-parallel across cards. For desktops it is far cheaper to rent it per token (see the leaderboard).