Model Spend Arena2ND ED.
16B · MoE · 32,768 ctx · 2026-10-10

Instella MoE 16B A3B

Runs on ≤12 GB · Hugging Face ↗

transformers

Not independently scored by Artificial Analysis

AMD's fully-open MoE reasoner — 16B total, ~3B active across 64 experts on a DeepSeek-V3-style architecture, with a visible think phase. A rare frontier-style open release straight from AMD; trending on HF but unrated by Artificial Analysis.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 6% (9/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. repeats itself until it runs out; 1 of 148 answers hit the token limit (effective ceiling 99%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard6.1%n=148 · 3 run(s)official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ2_M6.4 GB
Q2_K6.5 GB
IQ3_XXS7.1 GB
IQ3_XS7.2 GB
Q3_K_S7.6 GB
IQ3_M7.6 GB
Q3_K_M8.2 GB
Q3_K_L8.6 GB
IQ4_XS8.7 GB
Q4_09.0 GB
IQ4_NL9.0 GB
Q4_K_S9.6 GB
Q4_110.0 GB
Q4_K_M10.5 GB
Q5_K_S11.3 GB
Q5_K_M12.0 GB
Q6_K14.2 GB
Q8_016.9 GB
BF1631.7 GB

How to run it

Serving recipe

Load a community Q4_K_M GGUF in llama.cpp; that is the route behind our score. Before llama.cpp supported it, vLLM 0.26 failed on the custom `InstellaMoE` architecture (2026-07-31) and transformers in bf16 ran at about 17 tok/s (2026-08-06).

Config tips

About 10 GB at Q4_K_M. A third-party GGUF now loads in stock llama.cpp, and that is how we measured it: 6.1% on BigCodeBench-Hard over three runs. The small active set keeps it fast once loaded.