Runs on ≤32 GB · Hugging Face ↗
Coding index 15.3 (Artificial Analysis)
The dense 32B of the original Qwen3 line — a strong single-card generalist with a reasoning mode.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈9 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q3_K_M | 16.0 GB |
| IQ4_XS | 17.7 GB |
| Q4_0 | 18.7 GB |
| Q4_K_S | 18.8 GB |
| Q4_K_M | 19.8 GB |
| Q4 | 20.0 GB |
| Q4_1 | 20.6 GB |
| Q5_K_M | 23.2 GB |
| Q8_0 | 34.8 GB |
| Q2_K | 37.7 GB |
| Q8 | 39.5 GB |
| Q6_K | 55.8 GB |
| BF16 | 65.5 GB |
~20 GB at Q4, fits a 24-32 GB card. Turn thinking off for fast agent loops; use YaRN for context beyond the native window.