Model Spend Arena2ND ED.
26B · dense · 2026-08-14

Gemma 4 26B A4B

Runs on ≤32 GB · Hugging Face ↗

llama.cppMLXvLLM

Coding index 39.3 (Artificial Analysis)

A Mixture-of-Experts Gemma — 26B total, ~4B active, so it is fast yet all experts must sit in VRAM. Gated on Hugging Face.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
F160.9 GB
Q2_K10.5 GB
Q3_K_M12.7 GB
IQ4_XS13.6 GB
Q4_K_S16.5 GB
MXFP416.6 GB
Q4_K_M16.9 GB
Q417.0 GB
Q5_K_M21.2 GB
Q8_027.3 GB
Q827.6 GB
Q6_K46.5 GB
BF1651.4 GB

Config tips

Budget for the full 26B (~15 GB at Q4) even though only 4B compute. Strong quality-per-VRAM; MLX build is excellent on Macs.