Model Spend Arena2ND ED.
31B · dense · 2026-08-14

Gemma 4 31B

Runs on ≤32 GB · Hugging Face ↗

llama.cppMLXvLLM

Coding index 43.4 (Artificial Analysis)

The largest dense Gemma 4 — a strong general model just inside the 32 GB tier. Gated on Hugging Face.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 43% (52/121), via rented RTX PRO 6000 Blackwell 96 GB (vLLM, BF16) — comparable full-148, conc64. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
F161.0 GB
Q2_K11.8 GB
Q3_K_M14.7 GB
IQ4_XS16.4 GB
Q4_017.3 GB
Q4_K_S17.4 GB
Q4_K_M18.3 GB
Q418.8 GB
Q4_119.1 GB
Q5_K_M21.7 GB
Q8_033.2 GB
Q835.0 GB
Q6_K52.7 GB
BF1662.4 GB

Config tips

~20 GB at Q4 — a 32 GB card, or a 24 GB one at a tighter quant / short context. Excellent on Apple Silicon via MLX.