Runs on Data centre · Hugging Face ↗
Coding index 49.3 (Artificial Analysis)
NVIDIA's flagship 550B-total MoE (~55B active), with a 1M-token context. Also offered free on some hosts, but self-hosting is cluster-scale.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| Q2_K | 202.2 GB |
| Q3_K_M | 274.0 GB |
| IQ4_XS | 286.1 GB |
| Q4_K_S | 327.7 GB |
| MXFP4 | 352.3 GB |
| Q4_K_M | 359.2 GB |
| Q4 | 360.2 GB |
| Q5_K_M | 426.5 GB |
| Q8_0 | 584.3 GB |
| Q8 | 594.5 GB |
| Q6_K | 983.9 GB |
| BF16 | 1099.0 GB |
~360 GB at Q4 — multi-node. The 1M context makes the KV cache enormous, so long-context runs need even more headroom. Rent it unless you have the fleet.