Model Spend Arena2ND ED.
29B · dense · 131,072 ctx · 2026-08-14

Granite 4.1 30B

Runs on ≤32 GB · Hugging Face ↗

llama.cppvLLMOllama

Coding index 10.4 (Artificial Analysis)

IBM's 30B-total MoE, Apache-2.0, enterprise-tuned for tool use and RAG rather than chat flair.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈9 GB at 32K, fp16).

QuantizationSize
Q2_K11.0 GB
Q3_K_M14.0 GB
IQ4_XS15.5 GB
Q4_016.4 GB
Q4_K_S16.5 GB
Q4_K_M17.5 GB
Q417.7 GB
Q4_118.1 GB
Q5_K_M20.5 GB
Q8_030.7 GB
Q833.5 GB
Q6_K48.5 GB
BF1657.7 GB

Config tips

~17 GB at Q4 — fits a 24 GB card easily. A permissively-licensed option for commercial / on-prem deployment where model licence matters.