Runs on Data centre · Hugging Face ↗
Coding index 60.2 (Artificial Analysis)
Xiaomi's ~1-trillion-parameter MoE flagship — a strong open coder at the very top of the size range.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈14 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q2_K | 338.4 GB |
| Q3_K_M | 459.7 GB |
| IQ4_XS | 490.7 GB |
| Q4_K_S | 588.5 GB |
| MXFP4 | 610.5 GB |
| Q4_K_M | 629.6 GB |
| Q4 | 631.0 GB |
| Q5_K_M | 758.0 GB |
| Q8_0 | 1087.6 GB |
| Q8 | 1101.5 GB |
| Q6_K | 1775.0 GB |
| BF16 | 2046.7 GB |
~630 GB at Q4 — multi-node only. Serve with vLLM / SGLang tensor-parallel. Another "open but data-centre" model; rent it per token unless you have a fleet.