Runs on ≤80 GB · Hugging Face ↗
Coding index 11.9 (Artificial Analysis)
Meta's dense 70B — the long-standing local workhorse. Lower coding score than the newer MoEs, but battle-tested and widely supported. Gated on Hugging Face.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| Q3_K_M | 34.3 GB |
| IQ4_XS | 37.9 GB |
| Q4_0 | 40.1 GB |
| Q4_K_S | 40.3 GB |
| Q4_K_M | 42.5 GB |
| Q4 | 42.7 GB |
| Q4_1 | 44.3 GB |
| Q5_K_M | 49.9 GB |
| Q8_0 | 75.0 GB |
| Q2_K | 80.0 GB |
| Q8 | 81.2 GB |
| BF16 | 141.1 GB |
| F16 | 141.1 GB |
| Q6_K | 176.9 GB |
~43 GB at Q4 — fits an 80 GB card, or 2×24 GB with tensor-parallel. Universally supported by every runtime and fine-tune toolchain.