Runs on Multi-GPU · Hugging Face ↗
AA coding index 52.9 (Artificial Analysis, frozen since 2026-09-11)
Thinking Machines Lab's ~952B MoE — a notable open release from the lab, landing mid-pack on open coding.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈9 GB at 32K, fp16). Estimated from the original model config, not from the GGUF we serve. Less reliable than the rest of this page.
| Quantization | Size |
|---|---|
| IQ1_S | 270.2 GB |
| IQ1_M | 285.0 GB |
| Q2_K_XL | 317.3 GB |
| Q3_K_XL | 432.8 GB |
| Q4_K_XL | 587.0 GB |
| Q8_0 | 856.8 GB |
| BF16 | 1894.3 GB |
~590 GB at Q4 — cluster only. Serve tensor-parallel with vLLM / SGLang. Included as a landmark open model to track rather than a practical desktop option.