Model Spend Arena2ND ED.
397B · dense · 2026-08-14

Qwen3.5 397B A17B

Runs on Data centre · Hugging Face ↗

vLLMSGLang

Coding index 48.2 (Artificial Analysis)

A 397B-total MoE (~17B active) — strong quality, data-centre footprint.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
IQ4_XS189.7 GB
MXFP4237.4 GB
Q4245.3 GB
Q3_K_M354.8 GB
Q8_0421.5 GB
Q8427.7 GB
Q4_K_S456.0 GB
Q4_K_M488.2 GB
Q5_K_M587.3 GB
BF16793.0 GB
Q6_K1015.5 GB

Config tips

~260 GB at Q4, multi-GPU tensor-parallel. The ~17B active params make decoding faster than the total size implies, but all experts must be resident.