Runs on — · Hugging Face ↗
AA coding index 5.4 (Artificial Analysis, frozen since 2026-09-11)
Meta's 8B from July 2024, under the Llama 3.1 licence. It is here as a BASELINE: the model everyone still compares against, measured on the same protocol as the rest.
BigCodeBench-Hard pass@1 17% (25/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.7 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 2 of 148 answers hit the token limit (effective ceiling 99%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 16.7% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
No GGUF build yet — ~5 GB estimated at 4-bit from the parameter count. Run from safetensors (transformers / vLLM / SGLang).
16.9% over three runs (16.9 / 16.9 / 16.2). Useful as a floor: a 2026 model of the same size that cannot beat this is not worth the download.