Model Spend Arena2ND ED.
120B · MoE · 262,144 ctx · 2026-08-14

Nemotron 3 Super 120B

Runs on Data centre · Hugging Face ↗

vLLMllama.cpp

Coding index 37.7 (Artificial Analysis)

NVIDIA's 120B-total MoE (~12B active) — the best measured model in the 80 GB tier, tuned to run well on NVIDIA's own stack.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16).

QuantizationSize
Q2_K54.7 GB
Q3_K_M61.7 GB
IQ4_XS64.5 GB
Q4_K_S79.0 GB
MXFP482.1 GB
Q4_K_M82.5 GB
Q483.8 GB
Q5_K_M107.3 GB
Q8_0128.5 GB
Q8132.5 GB
Q6_K232.6 GB
BF16241.5 GB

Config tips

~79 GB at Q4 — a single 80 GB card or a 2-GPU split. vLLM / TensorRT-LLM give the best throughput on NVIDIA hardware.