Runs on ≤96 GB · Hugging Face ↗
AA coding index 11.9 (Artificial Analysis, frozen since 2026-09-11)
Meta's dense 70B — the long-standing local workhorse. Lower coding score than the newer MoEs, but battle-tested and widely supported. Gated on Hugging Face.
BigCodeBench-Hard pass@1 30% (44/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 29.7% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈11 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ1_S | 15.9 GB |
| IQ1_M | 17.1 GB |
| IQ2_XXS | 19.4 GB |
| IQ2_M | 24.3 GB |
| Q2_K | 26.4 GB |
| Q2_K_L | 26.6 GB |
| Q2_K_XL | 27.0 GB |
| IQ3_XXS | 27.7 GB |
| Q3_K_S | 30.9 GB |
| Q3_K_M | 34.3 GB |
| Q3_K_XL | 34.8 GB |
| IQ4_XS | 37.9 GB |
| IQ4_NL | 40.1 GB |
| Q4_0 | 40.1 GB |
| Q4_K_S | 40.3 GB |
| Q4_K_M | 42.5 GB |
| Q4_K_XL | 42.7 GB |
| Q4_1 | 44.3 GB |
| Q5_K_S | 48.7 GB |
| Q5_K_XL | 49.9 GB |
| Q5_K_M | 49.9 GB |
| Q6_K | 57.9 GB |
| Q6_K_XL | 61.2 GB |
| Q8_0 | 75.0 GB |
| Q8_K_XL | 81.2 GB |
| BF16 | 141.1 GB |
| F16 | 141.1 GB |
~43 GB at Q4 — fits an 80 GB card, or 2×24 GB with tensor-parallel. Universally supported by every runtime and fine-tune toolchain.