Runs on Multi-GPU · Hugging Face ↗
AA coding index 58.8 (Artificial Analysis, frozen since 2026-09-11)
Tencent Hunyuan's ~299B MoE, open-weight — the same model whose hosted free window opened and closed on the leaderboard, here as a self-host option.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈11 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ1_S | 64.1 GB |
| IQ1_M | 71.1 GB |
| IQ2_XXS | 82.1 GB |
| IQ2_XS | 91.1 GB |
| IQ2_S | 92.8 GB |
| IQ2_M | 102.1 GB |
| Q2_K | 106.7 GB |
| Q2_K_L | 107.2 GB |
| IQ3_XXS | 126.1 GB |
| Q3_K_S | 131.3 GB |
| IQ3_XS | 137.4 GB |
| Q3_K_M | 137.5 GB |
| Q3_K_L | 143.1 GB |
| Q3_K_XL | 143.5 GB |
| IQ3_M | 143.7 GB |
| IQ4_XS | 160.8 GB |
| IQ4_NL | 169.8 GB |
| Q4_0 | 170.3 GB |
| Q4_K_S | 175.5 GB |
| Q4_K_M | 182.2 GB |
| Q4_K_L | 182.5 GB |
| Q4_1 | 187.8 GB |
| Q5_K_S | 206.0 GB |
| Q5_K_M | 212.8 GB |
| Q6_K | 257.2 GB |
| Q8_0 | 317.7 GB |
~182 GB at Q4 — multi-GPU. A GGUF exists, so llama.cpp with heavy offload is possible but slow; vLLM / SGLang tensor-parallel is the real route.