Model Spend Arena2ND ED.
754B · MoE · 202,752 ctx · 2026-10-03

GLM-5.1

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLangllama.cpp

AA coding index 55.8 (Artificial Analysis, frozen since 2026-09-11)

The previous Z.ai flagship, ~754B MoE — useful as the generation-gap reference against GLM-5.2 (which scores ~13 points higher on coding).

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈6 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ1_M205.5 GB
IQ2_XXS220.6 GB
IQ2_M236.2 GB
Q2_K_XL252.4 GB
IQ3_XXS268.3 GB
IQ3_S279.6 GB
Q3_K_S313.5 GB
Q3_K_M338.4 GB
Q3_K_XL340.1 GB
IQ4_XS361.3 GB
IQ4_NL368.9 GB
Q4_K_S433.9 GB
MXFP4450.9 GB
Q4_K_M464.5 GB
Q4_K_XL466.0 GB
Q5_K_S525.9 GB
Q5_K_M558.5 GB
Q5_K_XL560.1 GB
Q6_K621.3 GB
Q6_K_XL684.4 GB
Q8_0801.3 GB
Q8_K_XL810.6 GB
BF161508.0 GB

Config tips

~465 GB at Q4 — data-centre. If you can host a GLM at all, host 5.2 instead; 5.1 is here for comparison, not as the pick.