Runs on Data centre · Hugging Face ↗
Coding index 48.2 (Artificial Analysis)
A 397B-total MoE (~17B active) — strong quality, data-centre footprint.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| IQ4_XS | 189.7 GB |
| MXFP4 | 237.4 GB |
| Q4 | 245.3 GB |
| Q3_K_M | 354.8 GB |
| Q8_0 | 421.5 GB |
| Q8 | 427.7 GB |
| Q4_K_S | 456.0 GB |
| Q4_K_M | 488.2 GB |
| Q5_K_M | 587.3 GB |
| BF16 | 793.0 GB |
| Q6_K | 1015.5 GB |
~260 GB at Q4, multi-GPU tensor-parallel. The ~17B active params make decoding faster than the total size implies, but all experts must be resident.