Model Spend Arena2ND ED.
32B · MoE · 262,144 ctx · 2026-08-14

Nemotron Cascade 2 30B A3B

Runs on ≤32 GB · Hugging Face ↗

vLLMllama.cpp

Coding index 25.3 (Artificial Analysis)

A newer NVIDIA 30B-total MoE (~3B active) using cascade routing — fast decoding for its capability, tuned for the NVIDIA stack.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16).

QuantizationSize
IQ4_XS18.2 GB
Q4_018.3 GB
Q3_K_M19.1 GB
Q4_120.1 GB
Q4_K_S22.4 GB
Q4_K_M24.7 GB
Q424.9 GB
Q5_K_M26.2 GB
Q8_033.6 GB
Q2_K36.5 GB
BF1663.2 GB
Q6_K67.1 GB

Config tips

~25 GB at Q4 — a 32 GB card, or a 24 GB one with CPU offload of idle experts. vLLM / TensorRT-LLM give the best throughput on NVIDIA hardware.