Runs on ≤24 GB · Hugging Face ↗
Not independently scored by Artificial Analysis — its base model Qwen/Qwen3.8-27B is the closest proxy
The most liked model of its week on Hugging Face (819 likes in days), a 26.9B on Qwen3.5. LICENCE cc-by-nc-4.0: non-commercial.
BigCodeBench-Hard pass@1 28% (41/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 1 of 148 answers hit the token limit (effective ceiling 99%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 27.7% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.
| Quantization | Size |
|---|---|
| IQ2_XXS | 8.9 GB |
| IQ2_XS | 9.1 GB |
| IQ2_S | 9.7 GB |
| IQ2_M | 10.5 GB |
| Q2_K | 10.8 GB |
| IQ3_XXS | 12.3 GB |
| Q3_K_S | 12.7 GB |
| IQ3_XS | 12.8 GB |
| Q3_K_M | 13.4 GB |
| Q3_K_L | 14.1 GB |
| IQ3_M | 14.9 GB |
| IQ4_XS | 15.5 GB |
| Q4_0 | 16.3 GB |
| Q4_K_S | 16.4 GB |
| IQ4_NL | 17.4 GB |
| Q4_K_M | 17.4 GB |
| Q4_1 | 17.8 GB |
| Q4_K_L | 18.8 GB |
| Q5_K_S | 19.6 GB |
| Q5_K_M | 20.9 GB |
| Q6_K_S | 22.9 GB |
| Q6_K | 23.9 GB |
| Q6_K_L | 25.0 GB |
| Q8_0 | 29.1 GB |
| BF16 | 54.7 GB |
27.7% over three runs, which ties the top of the small tiers. But the licence forbids commercial use, so read it before putting it in a product. Its GGUF is third-party (bartowski).