Model Spend Arena2ND ED.
1023B · MoE · 1,048,576 ctx · 2026-10-03

MiMo-V2.5-Pro

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLang

AA coding index 60.2 (Artificial Analysis, frozen since 2026-09-11)

Xiaomi's ~1-trillion-parameter MoE flagship — a strong open coder at the very top of the size range.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ1_M304.1 GB
IQ2_XXS317.2 GB
IQ2_M317.3 GB
Q2_K_XL338.4 GB
IQ3_S377.9 GB
IQ3_XXS412.7 GB
Q3_K_M459.7 GB
Q3_K_XL459.9 GB
IQ4_XS490.7 GB
IQ4_NL501.3 GB
Q4_K_S588.5 GB
MXFP4610.5 GB
Q4_K_M629.6 GB
Q4_K_XL631.0 GB
Q5_K_S713.2 GB
Q5_K_M758.0 GB
Q5_K_XL759.3 GB
Q6_K846.6 GB
Q6_K_XL928.5 GB
Q8_01087.6 GB
Q8_K_XL1101.5 GB
BF162046.7 GB

Config tips

~630 GB at Q4 — multi-node only. Serve with vLLM / SGLang tensor-parallel. Another "open but data-centre" model; rent it per token unless you have a fleet.