Runs on ≤32 GB · Hugging Face ↗
Coding index 43.4 (Artificial Analysis)
The largest dense Gemma 4 — a strong general model just inside the 32 GB tier. Gated on Hugging Face.
BigCodeBench-Hard pass@1 43% (52/121), via rented RTX PRO 6000 Blackwell 96 GB (vLLM, BF16) — comparable full-148, conc64. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| F16 | 1.0 GB |
| Q2_K | 11.8 GB |
| Q3_K_M | 14.7 GB |
| IQ4_XS | 16.4 GB |
| Q4_0 | 17.3 GB |
| Q4_K_S | 17.4 GB |
| Q4_K_M | 18.3 GB |
| Q4 | 18.8 GB |
| Q4_1 | 19.1 GB |
| Q5_K_M | 21.7 GB |
| Q8_0 | 33.2 GB |
| Q8 | 35.0 GB |
| Q6_K | 52.7 GB |
| BF16 | 62.4 GB |
~20 GB at Q4 — a 32 GB card, or a 24 GB one at a tighter quant / short context. Excellent on Apple Silicon via MLX.