Runs on ≤8 GB · Hugging Face ↗
AA coding index 2.7 (Artificial Analysis, frozen since 2026-09-11)
A vision-language model that declares Gemma3ForConditionalGeneration, so the GGUF repo ships an mmproj projector next to the weights - do not point llama-server at it by mistake. 34 layers with a 1024-token sliding window, which is short: the KV cache stays small as context grows, but most layers only ever see the last 1024 tokens. The weights are under the Gemma licence, not Apache, and the official repo is gated: you have to accept the terms before you can download the tokenizer.
BigCodeBench-Hard pass@1 11% (16/148), via official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 10.8% | n=148 · 3 run(s) | official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 |
| BFCL v4 non-live | 55.5% ± 3.0 | n=1390 · 1 run(s) | BFCL v4 · Q4_K_M GGUF via llama.cpp · one request at a time · a shared cluster GPU |
| RULER (long context) | 59.6 ± 8.3 | N=5/task · at 32K | 7 quantisations measured |
| quant | GiB | code pass@1 |
|---|---|---|
| BF16 | 7.2 | 10.8% |
| Q8_0 | 3.8 | 10.8% |
| Q6_K | 3.0 | 11.5% |
| Q5_K_M | 2.6 | 10.1% |
| Q4_K_M | 2.3 | 9.5% |
| Q4_0 | 2.2 | 8.1% |
| Q3_K_M | 2.0 | 4.7% |
| Q2_K | 1.6 | 2.0% |
Each quant links to the exact file it was measured from, in unsloth/gemma-3-4b-it-GGUF. Every point measured on the same class of card, one pass, one request at a time. One problem is 0.68 points, so anything under ~3.4 points apart is a tie — that is the spread we measure between GPU classes, and our rows are not all on the same card. Each card column is the context left for the KV cache after the weights, KV-Q4 / KV-Q8. A dash means the file does not fit that card, or fits with no room left to work. Context is usually the real constraint, not quality. Computed for one slot (--parallel 1), which is what you get serving yourself. llama-server opens four by default, and on a hybrid that costs real context — the table below each curve shows how much. The card figures assume a dedicated card. Measured on an idle one: the driver keeps 623 MiB of 143,771, so 99.6% is usable. Compute buffers are measured too — a fixed ~130 MiB plus 1 MiB per 1K of context, the same slope on every architecture we checked. If your card also runs your desktop you get noticeably less, and we have not measured that case yet.
55.5% ± 3.0 over 1390 problems
Q4_K_M · a shared cluster GPU · KV f16 · 1 run. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.
| quant | GiB | 4K | 8K | 32K |
|---|---|---|---|---|
| BF16 | 7.2 | 87.8 ± 7.0 · N=5 | 70.0 ± 8.5 · N=5 | 59.6 ± 8.3 · N=5 |
| Q8_0 | 3.8 | 87.8 ± 7.0 · N=5 | 70.8 ± 8.4 · N=5 | 61.7 ± 8.2 · N=5 |
| Q6_K | 3.0 | 87.2 ± 7.3 · N=5 | 71.1 ± 8.8 · N=5 | 59.4 ± 7.6 · N=5 |
| Q5_K_M | 2.6 | 87.5 ± 7.2 · N=5 | 72.8 ± 8.5 · N=5 | 62.7 ± 7.8 · N=5 |
| Q4_K_M | 2.3 | 82.4 ± 7.6 · N=5 | 68.5 ± 8.5 · N=5 | 60.9 ± 7.5 · N=5 |
| Q4_0 | 2.2 | 89.3 ± 6.2 · N=5 | 69.2 ± 9.3 · N=5 | 61.7 ± 7.4 · N=5 |
| Q3_K_M | 2.0 | 83.1 ± 7.2 · N=5 | 67.8 ± 8.9 · N=5 | 59.5 ± 8.0 · N=5 |
All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.
This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. On this model Q3_K_M scores 59.5 at 32K against 59.4 for Q6_K, a difference of 0.2 points — while on BigCodeBench-Hard the same two files score 4.7% and 11.5%. A page that read the long-context number as a general score would be telling the truth and misleading you.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ1_S | 1.2 GB |
| IQ1_M | 1.2 GB |
| IQ2_XXS | 1.3 GB |
| IQ2_M | 1.6 GB |
| IQ3_XXS | 1.7 GB |
| Q2_K | 1.7 GB |
| Q2_K_L | 1.7 GB |
| Q2_K_XL | 1.8 GB |
| Q3_K_S | 1.9 GB |
| Q3_K_M | 2.1 GB |
| Q3_K_XL | 2.2 GB |
| IQ4_XS | 2.3 GB |
| IQ4_NL | 2.4 GB |
| Q4_0 | 2.4 GB |
| Q4_K_S | 2.4 GB |
| Q4_K_M | 2.5 GB |
| Q4_K_XL | 2.5 GB |
| Q4_1 | 2.6 GB |
| Q5_K_S | 2.8 GB |
| Q5_K_M | 2.8 GB |
| Q5_K_XL | 2.8 GB |
| Q6_K | 3.2 GB |
| Q6_K_XL | 3.6 GB |
| Q8_0 | 4.1 GB |
| Q8_K_XL | 5.2 GB |
| BF16 | 7.8 GB |
Pull the GGUF from unsloth/gemma-3-4b-it-GGUF and serve it with llama-server: llama-server -m gemma-3-4b-it-Q4_K_M.gguf --n-gpu-layers 99 --flash-attn on --jinja. The chat template writes the BOS token as text and the server adds another, so strip it from the prompt if you drive the completions endpoint yourself.
We measured its whole quantisation curve on BigCodeBench-Hard, three runs per point, and it is well behaved: the score falls with the file size and never jumps back. Q6_K is the top at 11.5%, BF16 and Q8_0 tie at 10.8%, and it keeps dropping to 2.0% at Q2_K. One problem is 0.68 points, so the whole range from Q4_K_M up sits within three problems of itself and the useful choice is the smallest that fits. Two things worth knowing before you pick it for code. First, size is not what decides here: granite-4.0-micro scores 16.2% and MiniCPM5-2B 19.6%, both smaller than this. Second, the packing format costs more than a bit of width - Q4_K_M scores 9.5% and Q4_0, the same width and almost the same file size, 8.1%.