Runs on Multi-GPU · Hugging Face ↗
AA coding index 60.8 (Artificial Analysis, frozen since 2026-09-11)
Moonshot's ~1-trillion-parameter MoE, tuned specifically for coding — and unlike Kimi K3 the weights ARE openly released. One of the strongest open coders.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈4 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ1_M | 303.9 GB |
| IQ2_XXS | 317.8 GB |
| IQ2_M | 318.0 GB |
| Q2_K_XL | 339.5 GB |
| IQ3_S | 418.8 GB |
| Q3_K_M | 463.6 GB |
| Q3_K_XL | 463.9 GB |
| IQ4_XS | 495.1 GB |
| Q4_K_XL | 583.7 GB |
| Q8_K_XL | 594.5 GB |
~495 GB at Q4 — cluster-scale, tensor / pipeline parallel across many GPUs. Open but not local for most; the hosted per-token price is the realistic route.