Model Spend Arena2ND ED.
8B · dense · 2026-10-03

Llama 3.1 8B Instruct

Runs on — · Hugging Face ↗

llama.cppvLLMOllama

AA coding index 5.4 (Artificial Analysis, frozen since 2026-09-11)

Meta's 8B from July 2024, under the Llama 3.1 licence. It is here as a BASELINE: the model everyone still compares against, measured on the same protocol as the rest.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 17% (25/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.7 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 2 of 148 answers hit the token limit (effective ceiling 99%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard16.7%n=148 · 3 run(s)official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.

Size

No GGUF build yet — ~5 GB estimated at 4-bit from the parameter count. Run from safetensors (transformers / vLLM / SGLang).

Config tips

16.9% over three runs (16.9 / 16.9 / 16.2). Useful as a floor: a 2026 model of the same size that cannot beat this is not worth the download.