Model Spend Arena2ND ED.
117B · MoE · 131,072 ctx · 2026-10-03

gpt-oss-120b

Runs on ≤96 GB · Hugging Face ↗

vLLMllama.cppOllama

AA coding index 30.4 (Artificial Analysis, frozen since 2026-09-11)

The larger open-weight OpenAI model, also MXFP4-native. ~63 GB at its shipped precision — a single 80 GB card (H100/A100) or a 2×48 GB / 2×40 GB split.

First-party test · not the AA coding index

BigCodeBench-Hard: score withdrawn. WITHDRAWN: of 4 runs, the only one that passed the reliability filter scored 22.30% while the 3 discarded ones all agree on 16.89%. The filter kept the outlier.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BFCL v4 non-live (partial)74.0% ± 3.9n=750 · 1 run(s)BFCL v4 · MXFP4 GGUF via llama.cpp · one request at a time · a shared cluster GPU
Tool use · BFCL v4 non-live · partial · first-party

74.0% ± 3.9 over 750 problems, 2 of BFCL’s 4 scoring groups

MXFP4 · a shared cluster GPU · 1 run. Not comparable with the other BFCL numbers on this site, which average all four groups. We publish no score for parallel and parallel_multiple: harmony, the format these models answer in, puts one tool call per message, so a correct answer to a parallel-call problem cannot be expressed at all. Both categories score exactly 0.00% with every problem attempted — that is the format hitting a wall, not the model failing. BFCL's own non-live score is the unweighted mean of four groups (simple, multiple, parallel, parallel multiple); we compute it over the two that can be measured here, so the arithmetic is BFCL's but the groups are half of them.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
Q3_K_S62.6 GB
Q2_K62.6 GB
Q4_062.6 GB
Q3_K_M62.6 GB
Q4_162.7 GB
Q4_K_S62.8 GB
Q4_K_M62.8 GB
Q2_K_L62.9 GB
Q5_K_S62.9 GB
Q5_K_M62.9 GB
Q4_K_XL63.0 GB
Q6_K63.3 GB
Q6_K_XL63.3 GB
Q8_063.4 GB
Q8_K_XL64.5 GB
F1665.4 GB

Config tips

Not a consumer-GPU model — plan for 80 GB or multi-GPU. MXFP4-native, so don't re-quantize. vLLM gives the best throughput; llama.cpp works with CPU/GPU offload if you are short on VRAM (slower).