Runs on Multi-GPU · Hugging Face ↗
Not independently scored by Artificial Analysis
Xiaomi's flagship (MIT), the model it sells inside its token plan: a 1.02T mixture of experts with 384 experts and 8 active per token. We measured AesSedai's 2-bit-per-weight GGUF (BPW2.0, with an importance matrix), downloaded on 2026-09-25; that repository, AesSedai/MiMo-V2.6-Pro-RL-GGUF, is no longer public.
BigCodeBench-Hard pass@1 22% (33/148), via official protocol · llama.cpp BPW2_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. different runtime; 3 of 148 answers hit the token limit (effective ceiling 98%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 22.3% | n=148 · 3 run(s) | official protocol · llama.cpp BPW2_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| Q8_0 | 4.9 GB |
| BF16 | 9.3 GB |
| Q2_K | 418.8 GB |
| MXFP4 | 554.2 GB |
22.3% over three runs, the same number each time. Not a home model: 220 GiB even at 2 bits, and we ran it on a 141 GB data-centre card with the experts spilled to system RAM, at about 8 tokens per second. Read it as what the open flagship scores when squeezed to 2 bits, not as what Xiaomi's API serves.