Runs on ≤32 GB · Hugging Face ↗
Coding index 53.7 (Artificial Analysis)
The strongest measured model that fits a single 24-32 GB card — a dense 27B and the full-precision parent of the ternary Bonsai build.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| Q2_K | 12.0 GB |
| Q3_K_M | 13.8 GB |
| IQ4_XS | 15.7 GB |
| Q4_0 | 16.1 GB |
| Q4_K_S | 16.1 GB |
| Q4_K_M | 17.1 GB |
| Q4_1 | 17.5 GB |
| Q4 | 17.9 GB |
| Q5_K_M | 19.8 GB |
| Q8_0 | 29.0 GB |
| Q8 | 35.8 GB |
| Q6_K | 48.9 GB |
| BF16 | 54.7 GB |
~18 GB at Q4 on a 24 GB (RTX 4090) or 32 GB card. Native long context; use YaRN to push further. The sensible local flagship for most people.