Model Spend Arena2ND ED.
12B · MoE · 131,072 ctx · 2026-10-09

Mellum2.1 12B Thinking

Runs on ≤12 GB · Hugging Face ↗

llama.cpp / llama-server (GGUF)Ollama (GGUF)vLLMtransformers

Not independently scored by Artificial Analysis — its base model JetBrains/Mellum2-12B-A2.5B-Base is the closest proxy

Released by JetBrains, Apache-2.0. The next version of Mellum 2 Thinking: same MellumForCausalLM architecture (12B MoE, 64 experts with 8 active, about 2.5B active parameters, 28 layers, sliding window on 3 of every 4 layers, 128K context). The new part is post-training, mostly reinforcement learning in real repositories with a shell and tools.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
MXFP47.0 GB
Q4_K_M8.1 GB
Q6_K10.9 GB
Q8_012.9 GB
BF1624.3 GB

Config tips

JetBrains ships its own GGUF and stock llama.cpp loads the architecture, so no special build. About 8 GB at Q4_K_M: fits an 8-16 GB card. Read it next to the Mellum 2 12B card: same architecture, and the difference between the two is what the agentic RL added. It is a thinking model, so give it a token limit with room for the reasoning.