Model Spend Arena2ND ED.
12B · dense · 2026-08-14

Gemma 4 12B

Runs on ≤8 GB · Hugging Face ↗

llama.cppMLXOllama

Coding index 31.0 (Artificial Analysis)

Google's dense 12B — the strongest measured model that still fits a small card. Weights are gated on Hugging Face (accept the licence to download).

First-party test · not the AA coding index

HumanEval pass@1 63% (19/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 36% (44/121), via local GPU (Ollama, num_ctx 16384, budget 12K) — comparable full-148, 0/148 truncated. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

First-party measurement

~31 tok/s on RTX 4060 Ti, measured by us.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
F160.9 GB
Q2_K4.7 GB
Q3_K_M5.7 GB
IQ4_XS6.4 GB
Q4_06.7 GB
Q4_K_S6.8 GB
Q4_K_M7.1 GB
Q47.4 GB
Q4_17.4 GB
Q5_K_M8.4 GB
Q8_013.1 GB
Q813.6 GB
Q6_K20.5 GB
BF1624.7 GB

Config tips

Q4 lands around 8 GB, so a busy KV cache can push it to a 12-16 GB card. Follow Gemma's exact chat template. Runs very well on Apple Silicon via MLX.