Runs on ≤32 GB · Hugging Face ↗
Coding index 39.3 (Artificial Analysis)
A Mixture-of-Experts Gemma — 26B total, ~4B active, so it is fast yet all experts must sit in VRAM. Gated on Hugging Face.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| F16 | 0.9 GB |
| Q2_K | 10.5 GB |
| Q3_K_M | 12.7 GB |
| IQ4_XS | 13.6 GB |
| Q4_K_S | 16.5 GB |
| MXFP4 | 16.6 GB |
| Q4_K_M | 16.9 GB |
| Q4 | 17.0 GB |
| Q5_K_M | 21.2 GB |
| Q8_0 | 27.3 GB |
| Q8 | 27.6 GB |
| Q6_K | 46.5 GB |
| BF16 | 51.4 GB |
Budget for the full 26B (~15 GB at Q4) even though only 4B compute. Strong quality-per-VRAM; MLX build is excellent on Macs.