Model Spend Arena2ND ED.
30B · MoE · 500,000 ctx · 2026-10-03

North Mini Code 1.0

Runs on ≤24 GB · Hugging Face ↗

vLLMSGLangllama.cpp

AA coding index 36.5 (Artificial Analysis, frozen since 2026-09-11)

Cohere's first fully-open model — 30B total / ~3B active MoE, Apache-2.0, tuned for agentic SWE. AA coding index 33.4, said to top Devstral 2 (123B) and Nemotron 3 Super (120B). Free on OpenRouter and in opencode Zen's free set.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 24% (35/148), via official protocol · llama.cpp UD-Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard23.6%n=148 · 3 run(s)official protocol · llama.cpp UD-Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range
BFCL v4 non-live79.9% ± 2.4n=1390 · 1 run(s)BFCL v4 · UD-Q4_K_M GGUF via llama.cpp · one request at a time · a shared cluster GPU
Tool use · BFCL v4 non-live · first-party

79.9% ± 2.4 over 1390 problems

UD-Q4_K_M · a shared cluster GPU · KV f16 · 1 run. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Size

No GGUF build yet — ~18 GB estimated at 4-bit from the parameter count. Run from safetensors (transformers / vLLM / SGLang).

Config tips

~18 GB at Q4 — a 24 GB+ card or a single H100 FP8. But it's free on OpenRouter :free and opencode Zen, which is how we measured it.