Runs on ≤12 GB · Hugging Face ↗
Not independently scored by Artificial Analysis
AMD's fully-open MoE reasoner — 16B total, ~3B active across 64 experts on a DeepSeek-V3-style architecture, with a visible think phase. A rare frontier-style open release straight from AMD; trending on HF but unrated by Artificial Analysis.
BigCodeBench-Hard pass@1 6% (9/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. repeats itself until it runs out; 1 of 148 answers hit the token limit (effective ceiling 99%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 6.1% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ2_M | 6.4 GB |
| Q2_K | 6.5 GB |
| IQ3_XXS | 7.1 GB |
| IQ3_XS | 7.2 GB |
| Q3_K_S | 7.6 GB |
| IQ3_M | 7.6 GB |
| Q3_K_M | 8.2 GB |
| Q3_K_L | 8.6 GB |
| IQ4_XS | 8.7 GB |
| Q4_0 | 9.0 GB |
| IQ4_NL | 9.0 GB |
| Q4_K_S | 9.6 GB |
| Q4_1 | 10.0 GB |
| Q4_K_M | 10.5 GB |
| Q5_K_S | 11.3 GB |
| Q5_K_M | 12.0 GB |
| Q6_K | 14.2 GB |
| Q8_0 | 16.9 GB |
| BF16 | 31.7 GB |
Load a community Q4_K_M GGUF in llama.cpp; that is the route behind our score. Before llama.cpp supported it, vLLM 0.26 failed on the custom `InstellaMoE` architecture (2026-07-31) and transformers in bf16 ran at about 17 tok/s (2026-08-06).
About 10 GB at Q4_K_M. A third-party GGUF now loads in stock llama.cpp, and that is how we measured it: 6.1% on BigCodeBench-Hard over three runs. The small active set keeps it fast once loaded.