Runs on ≤8 GB · Hugging Face ↗
Not independently scored by Artificial Analysis
Released 2026-09-22 by Microsoft. A repository-level coding agent derived from Qwen3.5-4B, post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. Same architecture as Qwen3.5-4B (hybrid Gated DeltaNet and gated attention, 32 layers). LICENCE unclear: the repo metadata says MIT, the model card text says Apache 2.0.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.
| Quantization | Size |
|---|---|
| IQ2_M | 1.7 GB |
| Q2_K | 1.8 GB |
| IQ3_XXS | 2.0 GB |
| Q3_K_S | 2.0 GB |
| IQ3_XS | 2.1 GB |
| Q3_K_M | 2.1 GB |
| Q3_K_L | 2.2 GB |
| IQ3_M | 2.3 GB |
| IQ4_XS | 2.5 GB |
| Q4_K_S | 2.6 GB |
| Q4_0 | 2.7 GB |
| IQ4_NL | 2.8 GB |
| Q4_K_M | 2.8 GB |
| Q4_1 | 2.9 GB |
| Q4_K_L | 3.0 GB |
| Q5_K_S | 3.2 GB |
| Q5_K_M | 3.4 GB |
| Q6_K_S | 3.6 GB |
| Q6_K | 3.8 GB |
| Q6_K_L | 3.9 GB |
| Q8_0 | 4.6 GB |
| BF16 | 8.7 GB |
Same architecture as Qwen3.5-4B, so every current GGUF runner loads it; Microsoft publishes no GGUF, the one here is bartowski's. The config allows 256K but Microsoft validated it at about 131K with up to 8,192 output tokens per turn. Built for its own five-tool Leaf harness, so a single-shot benchmark reads as a floor. Worth reading next to the Qwen3.5-4B card.