Model Spend Arena2ND ED.
9B · dense · 131,072 ctx · 2026-08-14

Nemotron Nano 9B V2

Runs on ≤16 GB · Hugging Face ↗

llama.cppvLLMOllama

Not independently scored by Artificial Analysis — its base model nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base is the closest proxy

NVIDIA's open 9B on the Nemotron-H hybrid architecture, with a reasoning mode. Free on OpenRouter's :free routes, and small enough for any card.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 22% (27/121), via OpenRouter :free route (hosted) — comparable full-148. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. A modest 9B on hard problems — fine as a fast free draft, not a top coder.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈8 GB at 32K, fp16).

QuantizationSize
Q2_K5.0 GB
Q4_05.3 GB
Q4_15.8 GB
Q4_K_S6.2 GB
Q4_K_M6.5 GB
Q5_K_M7.1 GB
Q6_K9.1 GB
Q8_09.5 GB
F1617.8 GB

Config tips

~6 GB at Q4 — fits an 8 GB card. Free on OpenRouter :free, which is how we measured it.