Model Spend Arena2ND ED.
128B · dense · 2026-08-14

Mistral Medium 3.5

Runs on Data centre · Hugging Face ↗

vLLMllama.cppSGLang

Coding index 46.9 (Artificial Analysis)

Mistral's 128B mid-flagship — the best coder in the batch that gets close to fitting a single big card rather than needing a cluster.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q3_K_M60.6 GB
IQ4_XS67.1 GB
Q4_071.0 GB
Q4_K_S71.2 GB
Q4_K_M74.9 GB
Q475.7 GB
Q4_178.5 GB
Q5_K_M88.3 GB
Q8_0132.9 GB
Q2_K141.7 GB
Q8144.7 GB
Q6_K211.7 GB
BF16250.1 GB

Config tips

~75 GB at Q4 — a single 80 GB card at a tight quant / short context, or a 2-GPU split. The practical ceiling for one-node self-hosting among these.