Runs on ≤80 GB · Hugging Face ↗
Coding index 30.4 (Artificial Analysis)
The larger open-weight OpenAI model, also MXFP4-native. ~63 GB at its shipped precision — a single 80 GB card (H100/A100) or a 2×48 GB / 2×40 GB split.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q4_0 | 62.6 GB |
| Q3_K_M | 62.6 GB |
| Q4_1 | 62.7 GB |
| Q4_K_S | 62.8 GB |
| Q4_K_M | 62.8 GB |
| Q5_K_M | 62.9 GB |
| Q4 | 63.0 GB |
| Q8_0 | 63.4 GB |
| Q8 | 64.5 GB |
| F16 | 65.4 GB |
| Q2_K | 125.4 GB |
| Q6_K | 126.6 GB |
Not a consumer-GPU model — plan for 80 GB or multi-GPU. MXFP4-native, so don't re-quantize. vLLM gives the best throughput; llama.cpp works with CPU/GPU offload if you are short on VRAM (slower).