Runs on ≤80 GB · Hugging Face ↗
Coding index 17.4 (Artificial Analysis)
An 80B-total MoE (~3B active) with a hybrid-attention backbone built for very long context at low active-compute cost.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q3_K_M | 38.3 GB |
| IQ4_XS | 42.6 GB |
| Q4_0 | 45.3 GB |
| Q4_K_S | 45.5 GB |
| Q4 | 46.1 GB |
| Q4_K_M | 48.5 GB |
| Q4_1 | 50.1 GB |
| Q5_K_M | 56.8 GB |
| Q8_0 | 84.8 GB |
| Q2_K | 88.5 GB |
| Q8 | 93.1 GB |
| Q6_K | 134.1 GB |
| BF16 | 159.5 GB |
~50 GB at Q4 — fits an 80 GB card with lots of room for its long context. The hybrid attention keeps the KV cache smaller than a dense 80B would.