Runs on ≤96 GB · Hugging Face ↗
AA coding index 30.4 (Artificial Analysis, frozen since 2026-09-11)
The larger open-weight OpenAI model, also MXFP4-native. ~63 GB at its shipped precision — a single 80 GB card (H100/A100) or a 2×48 GB / 2×40 GB split.
BigCodeBench-Hard: score withdrawn. WITHDRAWN: of 4 runs, the only one that passed the reliability filter scored 22.30% while the 3 discarded ones all agree on 16.89%. The filter kept the outlier.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BFCL v4 non-live (partial) | 74.0% ± 3.9 | n=750 · 1 run(s) | BFCL v4 · MXFP4 GGUF via llama.cpp · one request at a time · a shared cluster GPU |
74.0% ± 3.9 over 750 problems, 2 of BFCL’s 4 scoring groups
MXFP4 · a shared cluster GPU · 1 run. Not comparable with the other BFCL numbers on this site, which average all four groups. We publish no score for parallel and parallel_multiple: harmony, the format these models answer in, puts one tool call per message, so a correct answer to a parallel-call problem cannot be expressed at all. Both categories score exactly 0.00% with every problem attempted — that is the format hitting a wall, not the model failing. BFCL's own non-live score is the unweighted mean of four groups (simple, multiple, parallel, parallel multiple); we compute it over the two that can be measured here, so the arithmetic is BFCL's but the groups are half of them.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.
| Quantization | Size |
|---|---|
| Q3_K_S | 62.6 GB |
| Q2_K | 62.6 GB |
| Q4_0 | 62.6 GB |
| Q3_K_M | 62.6 GB |
| Q4_1 | 62.7 GB |
| Q4_K_S | 62.8 GB |
| Q4_K_M | 62.8 GB |
| Q2_K_L | 62.9 GB |
| Q5_K_S | 62.9 GB |
| Q5_K_M | 62.9 GB |
| Q4_K_XL | 63.0 GB |
| Q6_K | 63.3 GB |
| Q6_K_XL | 63.3 GB |
| Q8_0 | 63.4 GB |
| Q8_K_XL | 64.5 GB |
| F16 | 65.4 GB |
Not a consumer-GPU model — plan for 80 GB or multi-GPU. MXFP4-native, so don't re-quantize. vLLM gives the best throughput; llama.cpp works with CPU/GPU offload if you are short on VRAM (slower).