Model Spend Arena2ND ED.
299B · MoE · 262,144 ctx · 2026-10-03

Hy3

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLang

AA coding index 58.8 (Artificial Analysis, frozen since 2026-09-11)

Tencent Hunyuan's ~299B MoE, open-weight — the same model whose hosted free window opened and closed on the leaderboard, here as a self-host option.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈11 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ1_S64.1 GB
IQ1_M71.1 GB
IQ2_XXS82.1 GB
IQ2_XS91.1 GB
IQ2_S92.8 GB
IQ2_M102.1 GB
Q2_K106.7 GB
Q2_K_L107.2 GB
IQ3_XXS126.1 GB
Q3_K_S131.3 GB
IQ3_XS137.4 GB
Q3_K_M137.5 GB
Q3_K_L143.1 GB
Q3_K_XL143.5 GB
IQ3_M143.7 GB
IQ4_XS160.8 GB
IQ4_NL169.8 GB
Q4_0170.3 GB
Q4_K_S175.5 GB
Q4_K_M182.2 GB
Q4_K_L182.5 GB
Q4_1187.8 GB
Q5_K_S206.0 GB
Q5_K_M212.8 GB
Q6_K257.2 GB
Q8_0317.7 GB

Config tips

~182 GB at Q4 — multi-GPU. A GGUF exists, so llama.cpp with heavy offload is possible but slow; vLLM / SGLang tensor-parallel is the real route.