Runs on ≤24 GB · Hugging Face ↗
Not independently scored by Artificial Analysis — its base model Qwen3.6-35B-A3B is the closest proxy
Kwaipilot's agentic coder — a 35B-total Qwen3.5-MoE (the same architecture family as Ornith 35B), Apache-2.0, tagged code + agent. Trending on Hugging Face yet unscored by AA, so exactly the kind of capable open coder the index misses.
BigCodeBench-Hard pass@1 25% (37/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 25.0% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
| BFCL v4 non-live | 77.5% ± 2.6 | n=1390 · 1 run(s) | BFCL v4 · Q4_K_M GGUF via llama.cpp · one request at a time · a shared cluster GPU |
77.5% ± 2.6 over 1390 problems
Q4_K_M · a shared cluster GPU · KV f16 · 1 run. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.
| Quantization | Size |
|---|---|
| IQ2_XXS | 9.8 GB |
| IQ2_XS | 10.8 GB |
| IQ2_S | 11.0 GB |
| IQ2_M | 12.1 GB |
| Q2_K | 12.6 GB |
| Q2_K_L | 13.1 GB |
| IQ3_XXS | 14.9 GB |
| Q3_K_S | 15.5 GB |
| IQ3_XS | 16.2 GB |
| Q3_K_M | 16.2 GB |
| Q3_K_L | 16.9 GB |
| IQ3_M | 16.9 GB |
| Q3_K_XL | 17.3 GB |
| IQ4_XS | 18.8 GB |
| IQ4_NL | 19.9 GB |
| Q4_0 | 19.9 GB |
| Q4_K_S | 20.6 GB |
| Q4_K_M | 21.4 GB |
| Q4_K_L | 21.8 GB |
| Q4_1 | 22.0 GB |
| Q5_K_S | 24.2 GB |
| Q5_K_M | 25.0 GB |
| Q5_K_L | 25.3 GB |
| Q6_K | 30.1 GB |
| Q6_K_L | 30.3 GB |
| Q8_0 | 36.9 GB |
| BF16 | 69.4 GB |
~20 GB at Q4 — spills past a 16 GB card, comfortable on 24 GB+ or the 32 GB tier. MoE, so only a few billion params compute per token; built for multi-turn agentic coding, so a single-shot BCB understates it. (Unmeasured — genuinely hard to benchmark. It's a vision-language model, so vLLM 0.26 won't load it (visual-tower weights missing) AND its text-only GGUF loads in Ollama but generates empty output. Needs a runtime that handles the full VL model.)