Model Spend Arena2ND ED.
125B · dense · 2026-08-14

Qwen3.5 122B A10B

Runs on Data centre · Hugging Face ↗

vLLMllama.cppSGLang

Coding index 45.7 (Artificial Analysis)

A 122B-total MoE (~10B active) — the strongest model that fits a single 80 GB card, and far above the other 80 GB-class open models on coding.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q2_K41.8 GB
Q3_K_M56.4 GB
IQ4_XS60.2 GB
Q4_K_S71.7 GB
MXFP474.7 GB
Q4_K_M76.5 GB
Q477.0 GB
Q5_K_M91.5 GB
Q8_0129.9 GB
Q8170.8 GB
Q6_K213.4 GB
BF16244.3 GB

Config tips

~78 GB at Q4 on one 80 GB card (H100 / A100), or a 2-GPU split with headroom. The ~10B active params keep decoding brisk despite the size.