Runs on Data centre · Hugging Face ↗
Coding index 37.7 (Artificial Analysis)
NVIDIA's 120B-total MoE (~12B active) — the best measured model in the 80 GB tier, tuned to run well on NVIDIA's own stack.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q2_K | 54.7 GB |
| Q3_K_M | 61.7 GB |
| IQ4_XS | 64.5 GB |
| Q4_K_S | 79.0 GB |
| MXFP4 | 82.1 GB |
| Q4_K_M | 82.5 GB |
| Q4 | 83.8 GB |
| Q5_K_M | 107.3 GB |
| Q8_0 | 128.5 GB |
| Q8 | 132.5 GB |
| Q6_K | 232.6 GB |
| BF16 | 241.5 GB |
~79 GB at Q4 — a single 80 GB card or a 2-GPU split. vLLM / TensorRT-LLM give the best throughput on NVIDIA hardware.