Model Spend Arena2ND ED.
70B · dense · 2026-10-03

Llama 3.3 70B

Runs on ≤96 GB · Hugging Face ↗

llama.cppvLLMOllama

AA coding index 11.9 (Artificial Analysis, frozen since 2026-09-11)

Meta's dense 70B — the long-standing local workhorse. Lower coding score than the newer MoEs, but battle-tested and widely supported. Gated on Hugging Face.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 30% (44/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard29.7%n=148 · 3 run(s)official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈11 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ1_S15.9 GB
IQ1_M17.1 GB
IQ2_XXS19.4 GB
IQ2_M24.3 GB
Q2_K26.4 GB
Q2_K_L26.6 GB
Q2_K_XL27.0 GB
IQ3_XXS27.7 GB
Q3_K_S30.9 GB
Q3_K_M34.3 GB
Q3_K_XL34.8 GB
IQ4_XS37.9 GB
IQ4_NL40.1 GB
Q4_040.1 GB
Q4_K_S40.3 GB
Q4_K_M42.5 GB
Q4_K_XL42.7 GB
Q4_144.3 GB
Q5_K_S48.7 GB
Q5_K_XL49.9 GB
Q5_K_M49.9 GB
Q6_K57.9 GB
Q6_K_XL61.2 GB
Q8_075.0 GB
Q8_K_XL81.2 GB
BF16141.1 GB
F16141.1 GB

Config tips

~43 GB at Q4 — fits an 80 GB card, or 2×24 GB with tensor-parallel. Universally supported by every runtime and fine-tune toolchain.