Model Spend Arena2ND ED.
8B · dense · 131,072 ctx · 2026-08-14

Granite 4.1 8B

Runs on ≤8 GB · Hugging Face ↗

llama.cppvLLMOllama

Coding index 9.5 (Artificial Analysis)

IBM's enterprise-focused 8B, Apache-2.0 licensed, tuned for tool use and RAG rather than chat flair.

First-party test · not the AA coding index

HumanEval pass@1 67% (20/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 17% (21/121), via local GPU (Ollama, Q4_K_M). The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

First-party measurement

~46 tok/s on RTX 4060 Ti, measured by us.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈5 GB at 32K, fp16).

QuantizationSize
Q2_K3.4 GB
Q3_K_M4.3 GB
Q4_05.1 GB
Q4_K_S5.1 GB
Q4_K_M5.3 GB
Q4_15.6 GB
Q5_K_M6.3 GB
Q6_K7.2 GB
Q8_09.3 GB
BF1617.6 GB

Config tips

~5-6 GB at Q4 with room for long context. A safe, permissively-licensed default for on-prem / commercial use where model licence matters.