Model Spend Arena2ND ED.
70B · dense · 2026-08-14

Llama 3.3 70B

Runs on ≤80 GB · Hugging Face ↗

llama.cppvLLMOllama

Coding index 11.9 (Artificial Analysis)

Meta's dense 70B — the long-standing local workhorse. Lower coding score than the newer MoEs, but battle-tested and widely supported. Gated on Hugging Face.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q3_K_M34.3 GB
IQ4_XS37.9 GB
Q4_040.1 GB
Q4_K_S40.3 GB
Q4_K_M42.5 GB
Q442.7 GB
Q4_144.3 GB
Q5_K_M49.9 GB
Q8_075.0 GB
Q2_K80.0 GB
Q881.2 GB
BF16141.1 GB
F16141.1 GB
Q6_K176.9 GB

Config tips

~43 GB at Q4 — fits an 80 GB card, or 2×24 GB with tensor-parallel. Universally supported by every runtime and fine-tune toolchain.