Model Spend Arena2ND ED.
952B · dense · 2026-10-03

Inkling

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLang

AA coding index 52.9 (Artificial Analysis, frozen since 2026-09-11)

Thinking Machines Lab's ~952B MoE — a notable open release from the lab, landing mid-pack on open coding.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈9 GB at 32K, fp16). Estimated from the original model config, not from the GGUF we serve. Less reliable than the rest of this page.

QuantizationSize
IQ1_S270.2 GB
IQ1_M285.0 GB
Q2_K_XL317.3 GB
Q3_K_XL432.8 GB
Q4_K_XL587.0 GB
Q8_0856.8 GB
BF161894.3 GB

Config tips

~590 GB at Q4 — cluster only. Serve tensor-parallel with vLLM / SGLang. Included as a landmark open model to track rather than a practical desktop option.