Runs on Data centre · Hugging Face ↗
Coding index 45.7 (Artificial Analysis)
A 122B-total MoE (~10B active) — the strongest model that fits a single 80 GB card, and far above the other 80 GB-class open models on coding.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| Q2_K | 41.8 GB |
| Q3_K_M | 56.4 GB |
| IQ4_XS | 60.2 GB |
| Q4_K_S | 71.7 GB |
| MXFP4 | 74.7 GB |
| Q4_K_M | 76.5 GB |
| Q4 | 77.0 GB |
| Q5_K_M | 91.5 GB |
| Q8_0 | 129.9 GB |
| Q8 | 170.8 GB |
| Q6_K | 213.4 GB |
| BF16 | 244.3 GB |
~78 GB at Q4 on one 80 GB card (H100 / A100), or a 2-GPU split with headroom. The ~10B active params keep decoding brisk despite the size.