Model Spend Arena2ND ED.
1024B · MoE · 1,048,576 ctx · 2026-10-10

MiMo-V2.6-Pro

Runs on Multi-GPU · Hugging Face ↗

llama.cpp

Not independently scored by Artificial Analysis

Xiaomi's flagship (MIT), the model it sells inside its token plan: a 1.02T mixture of experts with 384 experts and 8 active per token. We measured AesSedai's 2-bit-per-weight GGUF (BPW2.0, with an importance matrix), downloaded on 2026-09-25; that repository, AesSedai/MiMo-V2.6-Pro-RL-GGUF, is no longer public.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 22% (33/148), via official protocol · llama.cpp BPW2_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. different runtime; 3 of 148 answers hit the token limit (effective ceiling 98%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard22.3%n=148 · 3 run(s)official protocol · llama.cpp BPW2_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
Q8_04.9 GB
BF169.3 GB
Q2_K418.8 GB
MXFP4554.2 GB

Config tips

22.3% over three runs, the same number each time. Not a home model: 220 GiB even at 2 bits, and we ran it on a 141 GB data-centre card with the experts spilled to system RAM, at about 8 tokens per second. Read it as what the open flagship scores when squeezed to 2 bits, not as what Xiaomi's API serves.