Runs on Data centre · Hugging Face ↗
Coding index 61.8 (Artificial Analysis)
Moonshot's ~1T general MoE — openly released (unlike K3), and a notch above its own K2.7-Code sibling on the coding index.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| Q2_K | 340.5 GB |
| Q4 | 583.7 GB |
| Q8 | 594.5 GB |
| BF16 | 2053.2 GB |
~580 GB at Q4 — cluster-scale. Serve tensor / pipeline parallel with vLLM or SGLang. Open but not local for most; rent it per token unless you run a fleet.