Runs on ≤32 GB · Hugging Face ↗
Coding index 10.4 (Artificial Analysis)
IBM's 30B-total MoE, Apache-2.0, enterprise-tuned for tool use and RAG rather than chat flair.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈9 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q2_K | 11.0 GB |
| Q3_K_M | 14.0 GB |
| IQ4_XS | 15.5 GB |
| Q4_0 | 16.4 GB |
| Q4_K_S | 16.5 GB |
| Q4_K_M | 17.5 GB |
| Q4 | 17.7 GB |
| Q4_1 | 18.1 GB |
| Q5_K_M | 20.5 GB |
| Q8_0 | 30.7 GB |
| Q8 | 33.5 GB |
| Q6_K | 48.5 GB |
| BF16 | 57.7 GB |
~17 GB at Q4 — fits a 24 GB card easily. A permissively-licensed option for commercial / on-prem deployment where model licence matters.