Model Spend Arena2ND ED.
5B · dense · 2026-10-03

Qwen3.5 4B

Runs on ≤8 GB · Hugging Face ↗

llama.cppMLXOllamaLM Studio

AA coding index 22.6 (Artificial Analysis, frozen since 2026-09-11)

A dense 4B with a thinking mode — a surprisingly capable reasoner at a size that fits almost anything, phone included.

First-party test · not the AA coding index

HumanEval pass@1 33% (10/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column. Reasoning model truncated at 2.5K tokens on 22/30 problems — a floor, not its ceiling.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 22% (33/148), via official protocol · llama.cpp Q8_0 · one request at a time · a shared cluster GPU · mean of 6 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. repeats itself until it runs out; 4 of 148 answers hit the token limit (effective ceiling 97%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard22.3%n=148 · 6 run(s)official protocol · llama.cpp Q8_0 · one request at a time · a shared cluster GPU · mean of 6 runs, range 0.0
BFCL v4 non-live82.3% ± 2.5n=1150 · 1 run(s)BFCL v4 · Q4_K_M GGUF via llama.cpp · one request at a time · a shared cluster GPU
RULER (long context)92.3 ± 5.2N=5/task · at 32K8 quantisations measured
Quantisation curve · first-party
quantGiBcode
pass@1
tools
BFCL
8 GB12 GB16 GB24 GB32 GB48 GB96 GB141 GB
BF167.821.6%——262K / 178K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
Q8_04.222.3%80.2%262K / 166K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
Q6_K3.323.0%80.8%262K / 222K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
Q5_K_M2.919.6%—262K / 246K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
Q4_K_M2.620.3%82.3%262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
Q4_02.419.6%—262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
Q3_K_M2.118.2%—262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
UD-Q2_K_XL1.88.1%—262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K

Two different questions. code is BigCodeBench-Hard: write a working program. tools is BFCL v4 non-live: call the right function with the right arguments. A model can be good at one and ordinary at the other, and quantisation does not hit them the same way — read each column on its own. Each quant links to the exact file it was measured from, in unsloth/Qwen3.5-4B-GGUF. Every point measured on the same class of card, one pass, one request at a time. One problem is 0.68 points, so anything under ~3.4 points apart is a tie — that is the spread we measure between GPU classes, and our rows are not all on the same card. Each card column is the context left for the KV cache after the weights, KV-Q4 / KV-Q8. A dash means the file does not fit that card, or fits with no room left to work. Context is usually the real constraint, not quality. Computed for one slot (--parallel 1), which is what you get serving yourself. llama-server opens four by default, and on a hybrid that costs real context — the table below each curve shows how much. The card figures assume a dedicated card. Measured on an idle one: the driver keeps 623 MiB of 143,771, so 99.6% is usable. Compute buffers are measured too — a fixed ~130 MiB plus 1 MiB per 1K of context, the same slope on every architecture we checked. If your card also runs your desktop you get noticeably less, and we have not measured that case yet.

Tool use · BFCL v4 non-live · first-party

82.3% ± 2.5 over 1150 problems

Q4_K_M · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Tool-use quantisation curve · first-party
quantGiBBFCL non-liveruns
Q8_04.280.2%1
Q6_K3.380.8%1
Q4_K_M2.682.3%1

One request at a time, same server build and same KV cache throughout. Every point on the same class of card. Points measured once carry the sampling margin only; where the curve jumps we measure again and give the spread. A row is only here if all seven categories completed.

Long context · RULER · first-party
quantGiB4K8K32K
BF167.886.2 ± 5.2 · N=590.3 ± 4.9 · N=592.3 ± 5.2 · N=5
Q8_04.286.2 ± 5.2 · N=589.9 ± 5.2 · N=592.3 ± 5.2 · N=5
Q6_K3.386.2 ± 5.2 · N=591.8 ± 4.2 · N=592.3 ± 5.2 · N=5
Q5_K_M2.986.2 ± 5.2 · N=588.2 ± 5.7 · N=592.3 ± 4.8 · N=5
Q4_K_M2.684.1 ± 6.1 · N=590.8 ± 5.5 · N=590.8 ± 4.3 · N=5
Q4_02.487.7 ± 5.2 · N=587.2 ± 5.1 · N=588.9 ± 4.1 · N=5
Q3_K_M2.183.1 ± 4.3 · N=591.3 ± 5.4 · N=592.3 ± 3.9 · N=5
Q2_K_XL1.892.0 ± 5.3 · N=590.2 ± 5.8 · N=582.0 ± 5.1 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. On this model Q3_K_M scores 92.3 at 32K against 92.3 for Q6_K, a difference of 0.0 points — while on BigCodeBench-Hard the same two files score 18.2% and 23.0%. A page that read the long-context number as a general score would be telling the truth and misleading you.

First-party measurement

~71 tok/s on RTX 4060 Ti, measured by us.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
IQ2_XXS1.5 GB
IQ2_M1.8 GB
Q2_K_XL1.9 GB
IQ3_XXS1.9 GB
Q3_K_S2.1 GB
Q3_K_M2.3 GB
Q3_K_XL2.4 GB
IQ4_XS2.5 GB
IQ4_NL2.6 GB
Q4_02.6 GB
Q4_K_S2.6 GB
Q4_K_M2.7 GB
Q4_12.8 GB
Q4_K_XL2.9 GB
Q5_K_S3.0 GB
Q5_K_M3.1 GB
Q5_K_XL3.3 GB
Q6_K3.5 GB
Q6_K_XL4.1 GB
Q8_04.5 GB
Q8_K_XL6.0 GB
BF168.4 GB

Config tips

Under 5 GB at Q4. Great as a fast draft / autocomplete model; toggle thinking off for latency. Runs well on Apple Silicon via MLX.