Runs on Multi-GPU · Hugging Face ↗
AA coding index 68.8 (Artificial Analysis, frozen since 2026-09-11)
A very large MoE and one of the highest-scoring open coding models. The weights are open, but running it is a multi-node affair.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈0 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| Q2_K | 0.1 GB |
| Q3_K_S | 0.1 GB |
| Q3_K_M | 0.1 GB |
| IQ4_XS | 0.1 GB |
| Q4_K_S | 0.1 GB |
| Q3_K_L | 0.1 GB |
| Q4_K_M | 0.1 GB |
| Q5_K_S | 0.1 GB |
| Q5_K_M | 0.1 GB |
| Q6_K | 0.1 GB |
| Q8_0 | 0.1 GB |
| F16 | 0.2 GB |
Hundreds of GB even at Q4 — tensor / pipeline parallel across many GPUs with vLLM or SGLang. For anything short of a cluster, rent it per token (leaderboard).