Runs on ≤32 GB · Hugging Face ↗
Not independently scored by Artificial Analysis
The 35B-total MoE middle sibling of DeepReinforce's self-scaffolding agentic coder (MIT), a Qwen3.5-MoE. A genuine step up from the 9B on our first-party BCB-Hard, and among the stronger open coders the AA index doesn't rank. Agentic by design, so a single-shot benchmark is a floor for its agent-loop behaviour.
BigCodeBench-Hard pass@1 35% (42/121), via rented RTX PRO 6000 Blackwell 96 GB (vLLM, BF16) — comparable full-148, conc64. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. The 35B MoE — a clear step up from the 9B. Agentic by design, so a single-shot benchmark is a floor for its agent-loop behaviour.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| Q2_K | 12.3 GB |
| Q3_K_M | 16.7 GB |
| IQ4_XS | 17.8 GB |
| Q4_K_S | 20.9 GB |
| MXFP4 | 21.7 GB |
| Q4_K_M | 22.1 GB |
| Q4 | 22.3 GB |
| Q5_K_M | 26.5 GB |
| Q8_0 | 36.9 GB |
| Q8 | 38.2 GB |
| Q6_K | 61.2 GB |
| BF16 | 69.4 GB |
~20 GB at Q4 — needs a 24 GB+ card or the 32 GB tier; the MoE keeps decode fast. We benchmarked it on a rented RTX PRO 6000 Blackwell 96 GB (vLLM, BF16) since it doesn't fit a 16 GB card.