Runs on ≤32 GB · Hugging Face ↗
AA coding index 23.6 (Artificial Analysis, frozen since 2026-09-11)
LG AI Research's dense 33B, strong bilingual (Korean / English). Note its licence is research / non-commercial — read it before deploying.
BigCodeBench-Hard pass@1 21% (31/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 20.9% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
| BFCL v4 non-live | 76.1% ± 2.8 | n=1150 · 1 run(s) | BFCL v4 · Q4KM GGUF via llama.cpp · one request at a time · a shared cluster GPU |
| RULER (long context) | 92.0 ± 5.9 | N=5/task · at 32K | 1 quantisations measured |
76.1% ± 2.8 over 1150 problems
Q4KM · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.
Read this first: These three points were measured with a NEWER llama.cpp build (2026-09-17) than every other RULER figure on this site. With the reference build the model produced looping garbage at 8K and 32K (2.51 and 0.0), so those numbers were withdrawn. Same data, same seed, same server settings and temperature 0: only the binary changed. The 4K score also moved, and downwards, from 86.15 to 78.46, because with the new build the model starts reasoning before answering and runs out of tokens on two tasks (qa_1 and vt drop to 0). So read this row against itself, not against the rest of the table.
| quant | GiB | 4K | 8K | 32K |
|---|---|---|---|---|
| Q4_K_M | 18.7 | 78.5 ± 2.8 · N=5 | 95.1 ± 4.5 · N=5 | 92.0 ± 5.9 · N=5 |
All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.
This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈5 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| Q2_K | 12.4 GB |
| Q3_K_S | 14.5 GB |
| Q3_K_M | 16.1 GB |
| Q3_K_L | 17.4 GB |
| IQ4_XS | 18.0 GB |
| Q4_K_S | 19.0 GB |
| Q4_K_M | 20.0 GB |
| Q5_K_S | 22.8 GB |
| Q5_K_M | 23.5 GB |
| Q6_K | 27.1 GB |
| Q8_0 | 35.1 GB |
~21 GB at Q4, a 32 GB card. Good for Korean workloads; check the licence for any commercial use.