Model Spend Arena2ND ED.
753B · MoE · 1,048,576 ctx · 2026-10-03

GLM-5.2

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLangllama.cpp

AA coding index 68.8 (Artificial Analysis, frozen since 2026-09-11)

Z.ai's flagship — a very large Mixture-of-Experts (hundreds of billions of total params). One of the strongest open models by coding index, but firmly data-centre scale to run locally, even at Q4 (hundreds of GB).

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈6 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ1_S216.7 GB
IQ1_M228.5 GB
IQ2_XXS238.5 GB
IQ2_M238.6 GB
Q2_K_XL253.9 GB
IQ3_XXS281.7 GB
IQ3_S308.6 GB
Q3_K_M342.7 GB
Q3_K_XL343.0 GB
IQ4_XS365.3 GB
IQ4_NL372.7 GB
Q4_K_S436.4 GB
Q4_K_M465.8 GB
Q4_K_XL467.3 GB
Q5_K_S527.3 GB
Q5_K_M560.8 GB
Q5_K_XL562.5 GB
Q6_K625.9 GB
Q6_K_XL684.4 GB
Q8_0801.4 GB
Q8_K_XL819.7 GB
BF161508.0 GB

Config tips

Multi-GPU or a serious server only — a Q4 build is still hundreds of GB of weights plus a large KV cache. Serve with vLLM or SGLang tensor-parallel across cards. For desktops it is far cheaper to rent it per token (see the leaderboard).