Runs on ≤16 GB · Hugging Face ↗
Coding index 12.5 (Artificial Analysis)
Mistral's dense 24B, Apache-2.0, strong at instruction-following and function calling for its size.
HumanEval pass@1 83% (25/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.
BigCodeBench-Hard pass@1 25% (30/121), via local GPU (Ollama, Q4_K_M). The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
~11 tok/s on RTX 4060 Ti, measured by us.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| Q3_K_M | 11.5 GB |
| IQ4_XS | 12.8 GB |
| Q4_0 | 13.5 GB |
| Q4_K_S | 13.5 GB |
| Q4_K_M | 14.3 GB |
| Q4 | 14.5 GB |
| Q4_1 | 14.9 GB |
| Q5_K_M | 16.8 GB |
| Q8_0 | 25.1 GB |
| Q2_K | 27.2 GB |
| Q8 | 29.0 GB |
| Q6_K | 40.1 GB |
| BF16 | 47.2 GB |
~14 GB at Q4 — a 16 GB card is the comfortable home. Good tool-calling support; respect the v3 tokenizer / template.