Model Spend Arena2ND ED.
1023B · MoE · 1,048,576 ctx · 2026-08-14

MiMo-V2.5-Pro

Runs on Data centre · Hugging Face ↗

vLLMSGLang

Coding index 60.2 (Artificial Analysis)

Xiaomi's ~1-trillion-parameter MoE flagship — a strong open coder at the very top of the size range.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈14 GB at 32K, fp16).

QuantizationSize
Q2_K338.4 GB
Q3_K_M459.7 GB
IQ4_XS490.7 GB
Q4_K_S588.5 GB
MXFP4610.5 GB
Q4_K_M629.6 GB
Q4631.0 GB
Q5_K_M758.0 GB
Q8_01087.6 GB
Q81101.5 GB
Q6_K1775.0 GB
BF162046.7 GB

Config tips

~630 GB at Q4 — multi-node only. Serve with vLLM / SGLang tensor-parallel. Another "open but data-centre" model; rent it per token unless you have a fleet.