Model Spend Arena2ND ED.
550B · MoE · 262,144 ctx · 2026-08-14

Nemotron 3 Ultra 550B

Runs on Data centre · Hugging Face ↗

vLLMSGLang

Coding index 49.3 (Artificial Analysis)

NVIDIA's flagship 550B-total MoE (~55B active), with a 1M-token context. Also offered free on some hosts, but self-hosting is cluster-scale.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q2_K202.2 GB
Q3_K_M274.0 GB
IQ4_XS286.1 GB
Q4_K_S327.7 GB
MXFP4352.3 GB
Q4_K_M359.2 GB
Q4360.2 GB
Q5_K_M426.5 GB
Q8_0584.3 GB
Q8594.5 GB
Q6_K983.9 GB
BF161099.0 GB

Config tips

~360 GB at Q4 — multi-node. The 1M context makes the KV cache enormous, so long-context runs need even more headroom. Rent it unless you have the fleet.