Runs on ≤24 GB · Hugging Face ↗
AA coding index 12.5 (Artificial Analysis, frozen since 2026-09-11)
Mistral's dense 24B, Apache-2.0, strong at instruction-following and function calling for its size.
HumanEval pass@1 83% (25/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.
BigCodeBench-Hard pass@1 23% (34/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 1 of 148 answers hit the token limit (effective ceiling 99%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 23.0% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
| BFCL v4 non-live | 61.6% ± 2.6 | n=1150 · 1 run(s) | BFCL v4 · Q4_K_M GGUF via llama.cpp · one request at a time · a shared cluster GPU · prompting mode |
61.6% ± 2.6 over 1150 problems
Q4_K_M · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting. In Java and JavaScript this is “does not call”, not “calls badly”. We ran the corrector's own parser over every saved answer to count how many produce a tool call at all: 93% in Python but 19% in Java and 14% in JavaScript, where it answers with an explanation and a code block instead. The scores follow — 83.75 / 3.00 / 4.00. The parallel categories are the opposite case: 86% and 90% of answers do call, and score 68% and 57%, so there it calls and gets it wrong. We drive it through the prompt, which is what /v1/completions gives us; the vendor's function-calling route hands the tools over natively and Gorilla scores the model that way, so it may behave differently — we have not measured that.
~11 tok/s on RTX 4060 Ti, measured by us.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈5 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ1_S | 5.6 GB |
| IQ1_M | 6.0 GB |
| IQ2_XXS | 6.8 GB |
| IQ2_M | 8.2 GB |
| Q2_K | 8.9 GB |
| Q2_K_L | 9.0 GB |
| Q2_K_XL | 9.3 GB |
| IQ3_XXS | 9.4 GB |
| Q3_K_S | 10.4 GB |
| Q3_K_M | 11.5 GB |
| Q3_K_XL | 11.9 GB |
| IQ4_XS | 12.8 GB |
| IQ4_NL | 13.5 GB |
| Q4_0 | 13.5 GB |
| Q4_K_S | 13.5 GB |
| Q4_K_M | 14.3 GB |
| Q4_K_XL | 14.5 GB |
| Q4_1 | 14.9 GB |
| Q5_K_S | 16.3 GB |
| Q5_K_M | 16.8 GB |
| Q5_K_XL | 16.8 GB |
| Q6_K | 19.3 GB |
| Q6_K_XL | 20.8 GB |
| Q8_0 | 25.1 GB |
| Q8_K_XL | 29.0 GB |
| BF16 | 47.2 GB |
~14 GB at Q4 — a 16 GB card is the comfortable home. Good tool-calling support; respect the v3 tokenizer / template. Driven through the prompt it rarely emits a tool call in Java or JavaScript, so for those give it the native function-calling route.