Model Spend Arena2ND ED.
1027B · dense · 2026-10-03

Kimi K2.7 Code

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLang

AA coding index 60.8 (Artificial Analysis, frozen since 2026-09-11)

Moonshot's ~1-trillion-parameter MoE, tuned specifically for coding — and unlike Kimi K3 the weights ARE openly released. One of the strongest open coders.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈4 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ1_M303.9 GB
IQ2_XXS317.8 GB
IQ2_M318.0 GB
Q2_K_XL339.5 GB
IQ3_S418.8 GB
Q3_K_M463.6 GB
Q3_K_XL463.9 GB
IQ4_XS495.1 GB
Q4_K_XL583.7 GB
Q8_K_XL594.5 GB

Config tips

~495 GB at Q4 — cluster-scale, tensor / pipeline parallel across many GPUs. Open but not local for most; the hosted per-token price is the realistic route.