Model Spend Arena2ND ED.
12B · dense · 2026-10-03

Gemma 4 12B

Runs on — · Hugging Face ↗

llama.cppMLXOllama

AA coding index 31.0 (Artificial Analysis, frozen since 2026-09-11)

Google's dense 12B — the strongest measured model that still fits a small card. Weights are gated on Hugging Face (accept the licence to download).

First-party test · not the AA coding index

HumanEval pass@1 63% (19/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.

First-party test · not the AA coding index

BigCodeBench-Hard: no score, on purpose. We measured it three times and withdrew every run. At Q4_K_M it falls into a repetition loop on BigCodeBench/37 — it repeats 'ResponseBody' until it runs out of budget — and 92 of 148 answers hit the token ceiling. The resulting 4% measures the jam, not the model. Reproducible across all three lanes, so it is not a bad run.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BFCL v4 non-live87.5% ± 2.1n=1150 · 1 run(s)BFCL v4 · Q6_K GGUF via llama.cpp · one request at a time · a shared cluster GPU · prompting mode
RULER (long context)70.9 ± 8.1N=5/task · at 32K12 quantisations measured
Tool use · BFCL v4 non-live · first-party

87.5% ± 2.1 over 1150 problems

Q6_K · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Tool-use quantisation curve · first-party
quantGiBBFCL non-liveruns
BF1622.287.1%1
Q8_011.886.9%1
Q6_K9.187.5%1
Q5_K_M7.886.7%1
UD-Q4_K_XL6.986.9%1
Q4_K_M6.685.8%1
Q4_06.383.1%1
UD-Q3_K_XL5.684.5%1
UD-IQ3_XXS4.380.4%1

One request at a time, same server build and same KV cache throughout. Every point on the same class of card. Points measured once carry the sampling margin only; where the curve jumps we measure again and give the spread. A row is only here if all seven categories completed.

Long context · RULER · first-party
quantGiB4K8K32K
BF1622.286.5 ± 5.5 · N=582.1 ± 6.2 · N=570.9 ± 8.1 · N=5
Q8_011.887.7 ± 5.8 · N=586.2 ± 7.1 · N=568.9 ± 7.7 · N=5
Q6_K9.189.7 ± 5.5 · N=584.5 ± 7.1 · N=564.7 ± 7.4 · N=5
Q5_K_M7.882.1 ± 6.4 · N=582.2 ± 7.2 · N=569.6 ± 8.1 · N=5
UD-Q4_K_XL6.992.9 ± 5.6 · N=591.5 ± 5.1 · N=570.1 ± 9.4 · N=5
Q4_K_M6.688.8 ± 6.4 · N=589.5 ± 6.1 · N=560.6 ± 8.0 · N=5
Q4_06.377.4 ± 4.1 · N=579.0 ± 7.1 · N=563.6 ± 8.8 · N=5
UD-Q3_K_XL5.684.0 ± 6.9 · N=578.4 ± 7.5 · N=559.6 ± 9.4 · N=5
Q3_K_M5.380.6 ± 8.2 · N=576.5 ± 8.5 · N=540.0 ± 7.6 · N=5
UD-Q2_K_XL4.361.6 ± 6.9 · N=541.7 ± 8.4 · N=56.5 ± 5.7 · N=5
UD-IQ3_XXS4.369.3 ± 9.1 · N=549.9 ± 9.3 · N=514.7 ± 7.2 · N=5
UD-IQ2_M3.952.6 ± 10.4 · N=538.8 ± 7.9 · N=510.1 ± 6.3 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.

One or more rows ran on an engine that does not report whether the prompt was truncated, so for those we cannot prove the haystack went in whole.

First-party measurement

~31 tok/s on RTX 4060 Ti, measured by us.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
IQ2_M4.2 GB
IQ3_XXS4.6 GB
Q2_K_XL4.7 GB
Q3_K_S5.1 GB
Q3_K_M5.7 GB
Q3_K_XL6.0 GB
IQ4_XS6.4 GB
IQ4_NL6.7 GB
Q4_06.7 GB
Q4_K_S6.8 GB
Q4_K_M7.1 GB
Q4_K_XL7.4 GB
Q4_17.4 GB
Q5_K_S8.2 GB
Q5_K_M8.4 GB
Q5_K_XL8.6 GB
Q6_K9.8 GB
Q6_K_XL10.7 GB
Q8_012.7 GB
Q8_K_XL13.6 GB
BF1623.8 GB

Config tips

Not measurable with our current setup, and the reason is the model, not our rig. Google has an open, acknowledged bug: on long prompts the model falls into a deterministic "thought\nthought\n..." loop that runs for around a thousand iterations (huggingface.co/google/gemma-4-12B-it, discussion 41). Two things there matter if you are thinking of running it. It reproduces at F16 with no quantisation at all, shown across 48 seeds, so no GGUF build and no quant level avoids it. And it is probabilistic, not constant: roughly 3 runs in 5 fail, 43.8% with the embedded chat template and 35.4% with the best community one, which lowers the rate without removing it. Google confirmed it on 2026-07-09 and said a fix was coming; as of 2026-09-12 there is none. Against a 1,280-token budget a loop that frequent is fatal, so we publish no score. We also see a second, separate problem that is not the loop: at Q8_0, with no loop at all, it still exhausts the budget on some problems writing perfectly ordinary code. It is simply verbose. And llama.cpp crashes reproducibly on BigCodeBench/308 and /492, on upstream builds too, so that one is the GGUF rather than our patch.