Model Spend Arena2ND ED.
21B · MoE · 131,072 ctx · 2026-10-03

gpt-oss-20b

Runs on ≤16 GB · Hugging Face ↗

llama.cppvLLMOllamaLM Studio

AA coding index 20.7 (Artificial Analysis, frozen since 2026-09-11)

OpenAI's open-weight 20B, shipped natively in MXFP4 (~4-bit) — so its on-disk size already IS the quantized size. Uses the "harmony" chat format with separate reasoning channels.

First-party test · not the AA coding index

HumanEval pass@1 97% (29/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 27% (40/148), via official protocol · llama.cpp MXFP4 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 1 of 148 answers hit the token limit (effective ceiling 99%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard27.0%n=148 · 3 run(s)official protocol · llama.cpp MXFP4 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0
BFCL v4 non-live (partial)74.0% ± 3.9n=750 · 1 run(s)BFCL v4 · MXFP4 GGUF via llama.cpp · one request at a time · a shared cluster GPU
Tool use · BFCL v4 non-live · partial · first-party

74.0% ± 3.9 over 750 problems, 2 of BFCL’s 4 scoring groups

MXFP4 · a shared cluster GPU · 1 run. Not comparable with the other BFCL numbers on this site, which average all four groups. We publish no score for parallel and parallel_multiple: harmony, the format these models answer in, puts one tool call per message, so a correct answer to a parallel-call problem cannot be expressed at all. Both categories score exactly 0.00% with every problem attempted — that is the format hitting a wall, not the model failing. BFCL's own non-live score is the unweighted mean of four groups (simple, multiple, parallel, parallel multiple); we compute it over the two that can be measured here, so the arithmetic is BFCL's but the groups are half of them.

First-party speed · our home cards
R9700 32GB (Vulkan)159 tok/s
RTX 4060 Ti 16GB89 tok/s

MXFP4 GGUF, one request at a time, 16384 context, KV as in the model card. Re-measured 2026-08-29 with our speed kit; the run files (flags, GPU layers, device, llama.cpp build) are in results/speed/kit/r9700/. One request at a time — the speed one person sees. Data-centre cards are measured but not published.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
MXFP412.1 GB

Config tips

Do NOT re-quantize — it is already 4-bit; a Q8 re-quant only wastes VRAM. Fits ~13 GB. Needs a runtime with MXFP4 support (recent llama.cpp / vLLM / Ollama). Set the reasoning effort via the harmony system fields.