Model Spend Arena2ND ED.
952B · dense · 2026-08-14

Inkling

Runs on Data centre · Hugging Face ↗

vLLMSGLang

Coding index 52.9 (Artificial Analysis)

Thinking Machines Lab's ~952B MoE — a notable open release from the lab, landing mid-pack on open coding.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q2_K87.9 GB
Q3_K_M119.4 GB
IQ4_XS127.4 GB
Q4_K_S152.3 GB
MXFP4158.0 GB
Q4_K_M162.5 GB
Q4163.3 GB
Q5_K_M196.1 GB
Q8_0280.3 GB
Q8288.4 GB
Q6_K458.6 GB
BF16527.4 GB

Config tips

~590 GB at Q4 — cluster only. Serve tensor-parallel with vLLM / SGLang. Included as a landmark open model to track rather than a practical desktop option.