Model Spend Arena2ND ED.
80B · MoE · 262,144 ctx · 2026-08-14

Qwen3 Next 80B A3B

Runs on ≤80 GB · Hugging Face ↗

vLLMllama.cppSGLang

Coding index 17.4 (Artificial Analysis)

An 80B-total MoE (~3B active) with a hybrid-attention backbone built for very long context at low active-compute cost.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16).

QuantizationSize
Q3_K_M38.3 GB
IQ4_XS42.6 GB
Q4_045.3 GB
Q4_K_S45.5 GB
Q446.1 GB
Q4_K_M48.5 GB
Q4_150.1 GB
Q5_K_M56.8 GB
Q8_084.8 GB
Q2_K88.5 GB
Q893.1 GB
Q6_K134.1 GB
BF16159.5 GB

Config tips

~50 GB at Q4 — fits an 80 GB card with lots of room for its long context. The hybrid attention keeps the KV cache smaller than a dense 80B would.