Runs on ≤24 GB · Hugging Face ↗
Not independently scored by Artificial Analysis — its base model Qwen/Qwen3.6-35B-A3B is the closest proxy
A coder build of Qwen3.6 27B-A3B that Qwen never published under that name — the only weights on the Hub are third-party, so treat the lineage as claimed rather than confirmed. Worth a card anyway because we measured it on SWE-bench and the numbers should not sit in a table with no page behind them. Architecturally it is a hybrid: 184 experts with 10 active, only 2 KV heads, and SSM layers mixed in with attention, which is why its KV cache stays tiny at long context.
BigCodeBench-Hard pass@1 22% (32/148), via official protocol · llama.cpp Q6_K · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 1 of 148 answers hit the token limit (effective ceiling 99%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 21.6% | n=148 · 3 run(s) | official protocol · llama.cpp Q6_K · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 |
| BFCL v4 non-live | 85.9% ± 2.3 | n=1150 · 1 run(s) | BFCL v4 · Q6_K GGUF via llama.cpp · one request at a time · a shared cluster GPU |
| SWE-bench Verified | 26.6% ± 3.9 | n=500 · 1 run(s) | SWE-bench Verified · mini-swe-agent 2.4.5 · step_limit 250 · swebench 4.1.0 · a shared cluster GPU |
85.9% ± 2.3 over 1150 problems
Q6_K · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.
26.6% ± 3.9 of all 500 instances
SWE-bench Verified · mini-swe-agent 2.4.5 · step_limit 250 · swebench 4.1.0 · a shared cluster GPU. 1 run, patches graded by the official swebench harness in Docker, nothing filtered out of the denominator. server errors · some instances ran out of steps · context window exceeded · unreadable output format · agent config modified · network access inside the container · single run
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.
| Quantization | Size |
|---|---|
| Q2_K | 9.9 GB |
| Q3_K_S | 11.5 GB |
| Q3_K_M | 12.7 GB |
| Q3_K_L | 13.7 GB |
| IQ4_XS | 14.4 GB |
| Q4_K_S | 15.1 GB |
| Q4_K_M | 16.1 GB |
| Q5_K_S | 18.2 GB |
| Q5_K_M | 18.7 GB |
| Q6_K | 21.6 GB |
| Q8_0 | 27.9 GB |
15.0 GB at Q4_K_M and 20.1 GB at Q6_K, real GGUF bytes; the Q4 build needs about 19 GB of VRAM in total, so it fits a 24 GB card with room to spare. Plan for the full 27B in memory, not the 3B active.