Model Spend Arena2ND ED.
33B · dense · 40,960 ctx · 2026-08-14

Qwen3 32B

Runs on ≤32 GB · Hugging Face ↗

llama.cppvLLMOllamaLM Studio

Coding index 15.3 (Artificial Analysis)

The dense 32B of the original Qwen3 line — a strong single-card generalist with a reasoning mode.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈9 GB at 32K, fp16).

QuantizationSize
Q3_K_M16.0 GB
IQ4_XS17.7 GB
Q4_018.7 GB
Q4_K_S18.8 GB
Q4_K_M19.8 GB
Q420.0 GB
Q4_120.6 GB
Q5_K_M23.2 GB
Q8_034.8 GB
Q2_K37.7 GB
Q839.5 GB
Q6_K55.8 GB
BF1665.5 GB

Config tips

~20 GB at Q4, fits a 24-32 GB card. Turn thinking off for fast agent loops; use YaRN for context beyond the native window.