Model Spend Arena2ND ED.
128B · dense · 2026-10-03

Mistral Medium 3.5

Runs on ≤96 GB · Hugging Face ↗

vLLMllama.cppSGLang

AA coding index 46.9 (Artificial Analysis, frozen since 2026-09-11)

Mistral's 128B mid-flagship — the best coder in the batch that gets close to fitting a single big card rather than needing a cluster.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈12 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ2_XXS34.9 GB
IQ2_M44.1 GB
Q2_K46.6 GB
Q2_K_L47.0 GB
Q2_K_XL48.1 GB
IQ3_XXS49.3 GB
Q3_K_S54.4 GB
Q3_K_M60.6 GB
Q3_K_XL62.5 GB
IQ4_XS67.1 GB
IQ4_NL70.9 GB
Q4_071.0 GB
Q4_K_S71.2 GB
Q4_K_M74.9 GB
Q4_K_XL75.7 GB
Q4_178.5 GB
Q5_K_S86.2 GB
Q5_K_M88.3 GB
Q5_K_XL88.4 GB
Q6_K102.6 GB
Q6_K_XL109.2 GB
Q8_0132.9 GB
Q8_K_XL144.7 GB
BF16250.1 GB

Config tips

~75 GB at Q4 — a single 80 GB card at a tight quant / short context, or a 2-GPU split. The practical ceiling for one-node self-hosting among these.