Model Spend Arena2ND ED.
27B · dense · 262,144 ctx · 2026-10-10

Qwen3.8-27B

Runs on ≤16 GB · Hugging Face ↗

llama.cpp / llama-server (MTP spec-decode)Ollama (GGUF)vLLM (BF16 / FP8 / NVFP4)

AA coding index 68.1 (Artificial Analysis, frozen since 2026-09-11)

A dense 27B (Apache-2.0) and a native vision-language model, released 2026-08-14. Dense means every parameter is read per token, so it is slower per GB than a mixture-of-experts of similar size. On a 16 GB card the quant to use is Q3_K_XL (~12.2 GiB), which on our tests ties the 27 GiB Q8_0. It is a reasoning model, but thinking traces are long, so give it room in the token budget.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 32% (47/148), via official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard31.8%n=148 · 3 run(s)official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0
BFCL v4 non-live88.1% ± 2.1n=1150 · 3 run(s)BFCL v4 · UD-Q3_K_XL GGUF via llama.cpp · one request at a time · a shared cluster GPU
RULER (long context)95.4 ± 4.3N=5/task · at 128K3 quantisations measured
Quantisation curve · first-party
quantGiBcode
pass@1
tools
BFCL
12 GB16 GB24 GB32 GB48 GB96 GB141 GB
BF1650.931.8%87.8%—————262K / 262K262K / 262K
Q8_027.131.8%87.8%———151K / 80K262K / 262K262K / 262K262K / 262K
UD-Q6_K20.529.3%87.4%——93K / 49K262K / 262K262K / 262K262K / 262K262K / 262K
UD-Q5_K_M18.431.8%———215K / 114K262K / 262K262K / 262K262K / 262K262K / 262K
UD-Q4_K_XL16.429.9%———262K / 175K262K / 262K262K / 262K262K / 262K262K / 262K
UD-Q4_K_M15.329.3%87.9%——262K / 209K262K / 262K262K / 262K262K / 262K262K / 262K
Q4_015.032.4%
pre-flight check failed and was overridden
———262K / 218K262K / 262K262K / 262K262K / 262K262K / 262K
UD-Q3_K_XL12.233.1%88.1%—133K / 70K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
UD-IQ3_S11.233.8%——192K / 101K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
UD-IQ3_XXS10.228.1%83.7%29K / 15K250K / 132K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
UD-Q2_K_XL9.228.1%76.8%87K / 46K262K / 163K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K
UD-IQ2_S7.823.0%77.9%168K / 89K262K / 206K262K / 262K262K / 262K262K / 262K262K / 262K262K / 262K

Two different questions. code is BigCodeBench-Hard: write a working program. tools is BFCL v4 non-live: call the right function with the right arguments. A model can be good at one and ordinary at the other, and quantisation does not hit them the same way — read each column on its own. Each quant links to the exact file it was measured from, in unsloth/Qwen3.8-27B-GGUF. Every point measured on the same class of card, one pass, one request at a time. One problem is 0.68 points, so anything under ~3.4 points apart is a tie — that is the spread we measure between GPU classes, and our rows are not all on the same card. Each card column is the context left for the KV cache after the weights, KV-Q4 / KV-Q8. A dash means the file does not fit that card, or fits with no room left to work. Context is usually the real constraint, not quality. Computed for one slot (--parallel 1), which is what you get serving yourself. llama-server opens four by default, and on a hybrid that costs real context — the table below each curve shows how much. The card figures assume a dedicated card. Measured on an idle one: the driver keeps 623 MiB of 143,771, so 99.6% is usable. Compute buffers are measured too — a fixed ~130 MiB plus 1 MiB per 1K of context, the same slope on every architecture we checked. If your card also runs your desktop you get noticeably less, and we have not measured that case yet.

Tool use · BFCL v4 non-live · first-party

88.1% ± 2.1 over 1150 problems

UD-Q3_K_XL · a shared cluster GPU · KV f16 · 3 runs. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Agentic coding · SWE-bench Verified · first-party

6 resolved · 7 real patches from 14 instances tried

An early run, and deliberately without a percentage: it covered 14 instances of a 30-instance pool picked because their reference patch passes in our container, so neither 30 nor the 500 of SWE-bench Verified is a denominator you can divide by. Over the real 500 it is 1.2%. Step budget 400, mini-swe-agent, temperature 0. A full 500-instance run with the official protocol is what replaces this.

Tool-use quantisation curve · first-party
quantGiBBFCL non-liveruns
BF1650.987.8%1
Q8_027.187.8%1
Q6_K20.587.4%1
UD-Q4_K_M15.387.9%1
UD-Q3_K_XL12.288.1%3
UD-IQ3_XXS10.283.7%1
UD-Q2_K_XL9.276.8%2
UD-IQ2_S7.877.9%3

One request at a time, same server build and same KV cache throughout. Every point on the same class of card. Points measured once carry the sampling margin only; where the curve jumps we measure again and give the spread. A row is only here if all seven categories completed.

Long context · RULER · first-party
quantGiB4K32K128K
UD-Q3_K_XL12.298.5 ± 2.8 · N=597.2 ± 0.8 · N=10095.4 ± 4.3 · N=5
IQ3_XXS10.296.5 ± 0.9 · N=10096.7 ± 0.8 · N=100—
IQ2_S7.897.1 ± 0.8 · N=10096.4 ± 0.9 · N=100—

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops. Everything here clears the paper’s effective-length bar of 85.6 at every length shown.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.

First-party speed · our home cards
R9700 32GB (Vulkan)41 tok/s
RTX 4060 Ti 16GB42 tok/s

UD-Q3_K_XL GGUF with MTP speculative decoding (--spec-draft-n-max 2), one request at a time, 16384 context, q4_0 KV cache. Without MTP: 33.0. The old 46.6 figure was aggregate throughput, not decode speed. One request at a time — the speed one person sees. Data-centre cards are measured but not published.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
IQ1_S6.2 GB
IQ1_M6.7 GB
IQ2_XXS7.3 GB
IQ2_S8.4 GB
Q2_K_XL9.8 GB
IQ3_XXS10.9 GB
IQ3_S12.0 GB
Q3_K_XL13.1 GB
IQ4_XS14.3 GB
Q4_K_S15.4 GB
Q4_016.1 GB
Q4_K_M16.5 GB
Q4_117.5 GB
Q4_K_XL17.6 GB
Q5_K_S18.7 GB
Q5_K_M19.8 GB
Q5_K_XL20.9 GB
Q6_K22.0 GB
Q6_K_M23.1 GB
Q6_K_L24.2 GB
Q6_K_XL25.3 GB
Q8_K_L28.0 GB
Q8_029.0 GB
Q8_K_XL31.5 GB
BF1654.7 GB

How to run it

Serving recipe

Pull the GGUF (unsloth/Qwen3.8-27B-GGUF) and serve with llama-server, NOT Ollama, to get speculative decoding: llama-server -m Qwen3.8-27B-UD-Q3_K_XL.gguf --n-gpu-layers 99 --spec-type draft-mtp --spec-draft-n-max 2 --flash-attn on --jinja. The MTP draft head ships inside the GGUF and roughly doubles single-user speed.

Config tips

Quant vs quality, measured with the official protocol (148 problems, one attempt each, three or four runs per point). The names matter: most of these files are unsloth's UD- builds, which are dynamic — they do not apply the same width to every layer — so UD-Q6_K and a plain Q6_K are not the same thing. Each row in the curve links to the exact file. From BF16 down to UD-IQ3_S at 11.2 GiB, every point sits between 29.3% and 33.8%: BF16, Q8_0 and UD-Q5_K_M score 31.8%, and the two highest are UD-IQ3_S at 33.8% and UD-Q3_K_XL at 33.1%. One problem is 0.68 points, so that whole range is about seven problems, and each point carries three to four points of margin either way: none of those gaps is solid. Below it the score drops: UD-IQ3_XXS and UD-Q2_K_XL reach 28.2%, UD-IQ2_S 23.0%. The cliff is below three bits, not at four. Pick the smallest quant that fits and spend the VRAM you save on context: on a 24 GB card UD-Q6_K leaves about 93K tokens with a q4 KV cache, UD-Q3_K_XL the full 262K. Our earlier BF16/FP8/NVFP4 numbers were measured with a different protocol over 121 problems and have been withdrawn.