Model Spend Arena2ND ED.
24B · dense · 2026-08-14

Mistral Small 3.2 24B

Runs on ≤16 GB · Hugging Face ↗

llama.cppvLLMOllama

Coding index 12.5 (Artificial Analysis)

Mistral's dense 24B, Apache-2.0, strong at instruction-following and function calling for its size.

First-party test · not the AA coding index

HumanEval pass@1 83% (25/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 25% (30/121), via local GPU (Ollama, Q4_K_M). The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

First-party measurement

~11 tok/s on RTX 4060 Ti, measured by us.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q3_K_M11.5 GB
IQ4_XS12.8 GB
Q4_013.5 GB
Q4_K_S13.5 GB
Q4_K_M14.3 GB
Q414.5 GB
Q4_114.9 GB
Q5_K_M16.8 GB
Q8_025.1 GB
Q2_K27.2 GB
Q829.0 GB
Q6_K40.1 GB
BF1647.2 GB

Config tips

~14 GB at Q4 — a 16 GB card is the comfortable home. Good tool-calling support; respect the v3 tokenizer / template.