Model Spend Arena2ND ED.
8B · dense · 40,960 ctx · 2026-10-03

Qwen3 8B

Runs on ≤8 GB · Hugging Face ↗

llama.cpp / llama-server (GGUF)Ollama (GGUF)LM Studiotransformers

AA coding index 9.0 (Artificial Analysis, frozen since 2026-09-11)

Alibaba's dense 8B from April 2025, Apache-2.0: 36 layers, 32 attention heads over 8 key-value heads, 40K native context. Old by this page's standards, and on it for a reason: it is the reference model the function-calling leaderboard publishes, which we use to check our own harness.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 25% (37/148), via official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard25.0%n=148 · 3 run(s)official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0
RULER (long context)88.4 ± 5.6N=5/task · at 32K1 quantisations measured
Long context · RULER · first-party
quantGiB4K8K32K
BF1615.390.8 ± 5.5 · N=589.4 ± 5.1 · N=588.4 ± 5.6 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops. Everything here clears the paper’s effective-length bar of 85.6 at every length shown.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈5 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
IQ1_S2.3 GB
IQ1_M2.4 GB
IQ2_XXS2.6 GB
IQ2_M3.1 GB
Q2_K3.3 GB
IQ3_XXS3.4 GB
Q2_K_L3.4 GB
Q2_K_XL3.5 GB
Q3_K_S3.8 GB
Q3_K_M4.1 GB
Q3_K_XL4.3 GB
IQ4_XS4.6 GB
IQ4_NL4.8 GB
Q4_K_S4.8 GB
Q4_K_M5.0 GB
Q4_K_XL5.1 GB
Q4_15.2 GB
Q5_K_S5.7 GB
Q5_K_M5.9 GB
Q5_K_XL5.9 GB
Q6_K6.7 GB
Q6_K_XL7.5 GB
Q8_08.7 GB
Q8_K_XL10.8 GB
BF1616.4 GB

Config tips

Loads everywhere. On this page it is mostly a yardstick: newer 8-9B models should beat it, and when one does not, that is worth noticing.