Model Spend Arena2ND ED.
235B · MoE · 262,144 ctx · 2026-08-14

Qwen3 235B A22B 2507

Runs on Data centre · Hugging Face ↗

vLLMSGLang

Coding index 22.1 (Artificial Analysis)

A 235B-total MoE (~22B active) — a widely-served open flagship, Apache-2.0.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈6 GB at 32K, fp16).

QuantizationSize
Q3_K_M112.4 GB
IQ4_XS125.5 GB
Q4_0133.1 GB
Q4_K_S133.7 GB
Q4134.3 GB
Q4_K_M142.2 GB
Q4_1147.2 GB
Q5_K_M166.8 GB
Q8_0249.9 GB
Q2_K260.3 GB
Q8274.3 GB
Q6_K395.0 GB
BF16470.3 GB

Config tips

~142 GB at Q4 — multi-GPU tensor-parallel. All experts must be resident even though only ~22B compute. A common self-host target for those with two 80 GB cards.