Model Spend Arena2ND ED.
14B · dense · 2026-08-14

Ministral 3 14B

Runs on ≤16 GB · Hugging Face ↗

llama.cppvLLMOllama

Coding index 14.4 (Artificial Analysis)

Mistral's dense 14B in the edge-oriented Ministral line — a compact generalist with solid instruction-following.

First-party test · not the AA coding index

HumanEval pass@1 63% (19/30), via local GPU (Ollama, Q4_K_M). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 14% (17/121), via local GPU (Ollama, Q4_K_M). The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

First-party measurement

~30 tok/s on RTX 4060 Ti, measured by us.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q3_K_M6.7 GB
IQ4_XS7.4 GB
Q4_07.8 GB
Q4_K_S7.8 GB
Q4_K_M8.2 GB
Q48.4 GB
Q4_18.6 GB
Q5_K_M9.6 GB
Q8_014.4 GB
Q2_K16.2 GB
Q817.1 GB
Q6_K23.2 GB
BF1627.0 GB

Config tips

~8 GB at Q4, comfortable on a 16 GB card with long context. Respect the Mistral v3 tokenizer / chat template.