Runs on ≤96 GB · Hugging Face ↗
AA coding index 46.9 (Artificial Analysis, frozen since 2026-09-11)
Mistral's 128B mid-flagship — the best coder in the batch that gets close to fitting a single big card rather than needing a cluster.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈12 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ2_XXS | 34.9 GB |
| IQ2_M | 44.1 GB |
| Q2_K | 46.6 GB |
| Q2_K_L | 47.0 GB |
| Q2_K_XL | 48.1 GB |
| IQ3_XXS | 49.3 GB |
| Q3_K_S | 54.4 GB |
| Q3_K_M | 60.6 GB |
| Q3_K_XL | 62.5 GB |
| IQ4_XS | 67.1 GB |
| IQ4_NL | 70.9 GB |
| Q4_0 | 71.0 GB |
| Q4_K_S | 71.2 GB |
| Q4_K_M | 74.9 GB |
| Q4_K_XL | 75.7 GB |
| Q4_1 | 78.5 GB |
| Q5_K_S | 86.2 GB |
| Q5_K_M | 88.3 GB |
| Q5_K_XL | 88.4 GB |
| Q6_K | 102.6 GB |
| Q6_K_XL | 109.2 GB |
| Q8_0 | 132.9 GB |
| Q8_K_XL | 144.7 GB |
| BF16 | 250.1 GB |
~75 GB at Q4 — a single 80 GB card at a tight quant / short context, or a 2-GPU split. The practical ceiling for one-node self-hosting among these.