Model Spend Arena2ND ED.
1027B · dense · 2026-08-14

Kimi K2.6

Runs on Data centre · Hugging Face ↗

vLLMSGLang

Coding index 61.8 (Artificial Analysis)

Moonshot's ~1T general MoE — openly released (unlike K3), and a notch above its own K2.7-Code sibling on the coding index.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q2_K340.5 GB
Q4583.7 GB
Q8594.5 GB
BF162053.2 GB

Config tips

~580 GB at Q4 — cluster-scale. Serve tensor / pipeline parallel with vLLM or SGLang. Open but not local for most; rent it per token unless you run a fleet.