Runs on — · Hugging Face ↗
AA coding index 39.3 (Artificial Analysis, frozen since 2026-09-11)
A Mixture-of-Experts Gemma — 26B total, ~4B active, so it is fast yet all experts must sit in VRAM. Gated on Hugging Face.
BigCodeBench-Hard: score withdrawn. WITHDRAWN by the publication gate: presupuesto:truncado-inservible, protocolo:ejecucion-incompleta, protocolo:suelo-de-infra
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BFCL v4 non-live | 89.8% ± 2.0 | n=1150 · 1 run(s) | BFCL v4 · Q8_0 GGUF via llama.cpp · one request at a time · a shared cluster GPU · prompting mode |
89.8% ± 2.0 over 1150 problems
1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.
| quant | GiB | BFCL non-live | runs |
|---|---|---|---|
| BF16 | 47.0 | 88.8% | 1 |
| Q8_0 | 25.0 | 89.8% | 1 |
| UD-Q5_K_M | 19.7 | 89.8% | 1 |
| UD-Q4_K_M | 15.8 | 89.3% | 1 |
| UD-Q4_K_XL | 15.8 | 89.6% | 1 |
| UD-Q3_K_XL | 12.0 | 89.4% | 3 |
| UD-IQ2_M | 9.3 | 85.4% | 3 |
One request at a time, same server build and same KV cache throughout. Every point on the same class of card. Points measured once carry the sampling margin only; where the curve jumps we measure again and give the spread. A row is only here if all seven categories completed.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| IQ2_XXS | 9.9 GB |
| IQ2_M | 10.0 GB |
| Q2_K_XL | 10.5 GB |
| IQ3_S | 11.3 GB |
| IQ3_XXS | 11.4 GB |
| Q3_K_M | 12.7 GB |
| Q3_K_XL | 12.9 GB |
| IQ4_XS | 13.6 GB |
| IQ4_NL | 13.6 GB |
| Q4_K_S | 16.5 GB |
| MXFP4 | 16.6 GB |
| Q4_K_M | 16.9 GB |
| Q4_K_XL | 17.0 GB |
| Q5_K_S | 18.9 GB |
| Q5_K_M | 21.2 GB |
| Q5_K_XL | 21.2 GB |
| Q6_K | 23.2 GB |
| Q6_K_XL | 23.3 GB |
| Q8_0 | 26.9 GB |
| Q8_K_XL | 27.6 GB |
| BF16 | 50.5 GB |
Budget for the full 26B (~15 GB at Q4) even though only 4B compute. Strong quality-per-VRAM; MLX build is excellent on Macs.