Runs on — · Hugging Face ↗
AA coding index 71.5 (Artificial Analysis, frozen since 2026-09-11)
Z.ai's 321B Mixture-of-Experts (288 experts, 8 active), MIT-licensed, released 2026-08-25, with a 1M context. Attention is hybrid: none of its 45 layers use classic attention, so the KV cache grows far slower with context than the size suggests.
BigCodeBench-Hard pass@1 20% (29/148), via official protocol · llama.cpp UD-IQ3_XXS · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. repeats itself until it runs out; 3 of 148 answers hit the token limit (effective ceiling 98%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 19.6% | n=148 · 3 run(s) | official protocol · llama.cpp UD-IQ3_XXS · one request at a time · a shared cluster GPU · mean of 3 runs, rang |
| BFCL v4 non-live | 86.9% ± 2.2 | n=1390 · 1 run(s) | BFCL v4 · UD-Q2_K_XL GGUF via llama.cpp · one request at a time · a shared cluster GPU |
| quant | GiB | code pass@1 |
|---|---|---|
| UD-IQ3_XXS | 112.1 | 19.6% |
| UD-Q2_K_XL | 101.3 | 17.6% |
| UD-IQ2_XXS | 94.9 | 13.5% |
| UD-IQ1_M | 90.9 | 20.9% |
| UD-IQ1_S | 86.7 | 10.1% |
Each quant links to the exact file it was measured from, in unsloth/GLM-5.3-Flash-GGUF. Every point measured on the same class of card, one pass, one request at a time. One problem is 0.68 points, so anything under ~3.4 points apart is a tie — that is the spread we measure between GPU classes, and our rows are not all on the same card. Each card column is the context left for the KV cache after the weights, KV-Q4 / KV-Q8. A dash means the file does not fit that card, or fits with no room left to work. Context is usually the real constraint, not quality. Computed for one slot (--parallel 1), which is what you get serving yourself. llama-server opens four by default, and on a hybrid that costs real context — the table below each curve shows how much. The card figures assume a dedicated card. Measured on an idle one: the driver keeps 623 MiB of 143,771, so 99.6% is usable. Compute buffers are measured too — a fixed ~130 MiB plus 1 MiB per 1K of context, the same slope on every architecture we checked. If your card also runs your desktop you get noticeably less, and we have not measured that case yet.
86.9% ± 2.2 over 1390 problems
UD-Q2_K_XL · a shared cluster GPU · KV f16 · 1 run. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.
10 resolved · 10 real patches from 10 instances tried
An early run, and deliberately without a percentage: it covered 10 instances of a 30-instance pool picked because their reference patch passes in our container, so neither 30 nor the 500 of SWE-bench Verified is a denominator you can divide by. Over the real 500 it is 2.0%. Step budget 80, mini-swe-agent, temperature 0. A full 500-instance run with the official protocol is what replaces this.
Real GGUF file sizes = weight VRAM.
| Quantization | Size |
|---|---|
| IQ1_S | 93.1 GB |
| IQ1_M | 97.6 GB |
| IQ2_XXS | 101.8 GB |
| Q2_K_XL | 108.7 GB |
| IQ3_XXS | 120.4 GB |
| Q3_K_XL | 147.5 GB |
| IQ4_XS | 156.8 GB |
| Q4_K_XL | 199.7 GB |
| Q5_K_XL | 240.3 GB |
| Q6_K_XL | 291.8 GB |
| Q8_0 | 341.0 GB |
| BF16 | 641.6 GB |
We have measured its whole quantisation curve now, five points on the official protocol with three runs each, and it is the least well-behaved curve on this site. It does not fall as the files get smaller: UD-IQ1_M, at 90.9 GiB, scores 20.9% — the best of the five — while UD-IQ2_XXS, four gigabytes larger, scores 13.5%, and UD-IQ1_S, four gigabytes smaller, scores 10.1%. Between neighbours that is a swing of ten points with no pattern to it, so on this model you cannot pick a quant by size and expect the score to follow. Take UD-IQ1_M if it fits, and treat anything else as untested until we measure it. Every run here used unsloth's build of llama.cpp: the GGUF declares the architecture as `glm5next` and the reference build does not know it. Renting per token is still the cheap way in at $0.075/M in and $0.250/M out.