Runs on ≤12 GB · Hugging Face ↗
Not independently scored by Artificial Analysis — its base model Qwen/Qwen3.8-27B is the closest proxy
Released 2026-09-16 by PrismML, Apache-2.0: the second version of Ternary Bonsai 27B, again built from Qwen3.8-27B with ternary weights. Three builds: PTQ1_0 (6.0 GB), PQ2_0 (7.2 GB) and F16 (53.8 GB).
BigCodeBench-Hard pass@1 27% (40/148), via official protocol · llama.cpp PQ2_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. repeats itself until it runs out; 2 of 148 answers hit the token limit (effective ceiling 99%)
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 27.0% | n=148 · 3 run(s) | official protocol · llama.cpp PQ2_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 |
| RULER (long context) | 71.8 ± 4.9 | N=5/task · at 32K | 2 quantisations measured |
| quant | GiB | code pass@1 | tools BFCL | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB | 96 GB | 141 GB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| PQ2_0 | 6.7 | 27.0% | 86.0% | 11K / 6K | 233K / 123K | 262K / 240K | 262K / 262K | 262K / 262K | 262K / 262K | 262K / 262K | 262K / 262K |
| PTQ1_0 | 5.5 | 25.7% | 86.1% | 81K / 43K | 262K / 160K | 262K / 262K | 262K / 262K | 262K / 262K | 262K / 262K | 262K / 262K | 262K / 262K |
Two different questions. code is BigCodeBench-Hard: write a working program. tools is BFCL v4 non-live: call the right function with the right arguments. A model can be good at one and ordinary at the other, and quantisation does not hit them the same way — read each column on its own. We have not been able to tie these files back to a published repo, so take the quant names as the labels we ran them under. Every point measured on the same class of card, one pass, one request at a time. One problem is 0.68 points, so anything under ~3.4 points apart is a tie — that is the spread we measure between GPU classes, and our rows are not all on the same card. Each card column is the context left for the KV cache after the weights, KV-Q4 / KV-Q8. A dash means the file does not fit that card, or fits with no room left to work. Context is usually the real constraint, not quality. Computed for one slot (--parallel 1), which is what you get serving yourself. llama-server opens four by default, and on a hybrid that costs real context — the table below each curve shows how much. The card figures assume a dedicated card. Measured on an idle one: the driver keeps 623 MiB of 143,771, so 99.6% is usable. Compute buffers are measured too — a fixed ~130 MiB plus 1 MiB per 1K of context, the same slope on every architecture we checked. If your card also runs your desktop you get noticeably less, and we have not measured that case yet.
| quant | GiB | BFCL non-live | runs |
|---|---|---|---|
| PQ2_0 | 6.7 | 86.0% | 1 |
| PTQ1_0 | 5.5 | 86.1% | 1 |
One request at a time, same server build and same KV cache throughout. Every point on the same class of card. Points measured once carry the sampling margin only; where the curve jumps we measure again and give the spread. A row is only here if all seven categories completed.
| quant | GiB | 4K | 8K | 32K |
|---|---|---|---|---|
| PQ2_0 | 6.7 | 57.3 ± 7.6 · N=5 | 71.2 ± 5.4 · N=5 | 71.8 ± 4.9 · N=5 |
| PTQ1_0 | 5.5 | 57.3 ± 7.6 · N=5 | 71.2 ± 5.4 · N=5 | 72.3 ± 4.8 · N=5 |
All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.
This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. On this model PTQ1_0 scores 72.3 at 32K against 71.8 for PQ2_0, a difference of 0.5 points — while on BigCodeBench-Hard the same two files score 25.7% and 27.0%. A page that read the long-context number as a general score would be telling the truth and misleading you.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Estimated from the original model config, not from the GGUF we serve. Less reliable than the rest of this page.
| Quantization | Size |
|---|---|
| F16 | 53.8 GB |
Stock llama.cpp will not run it: it rejects PQ2_0 and PTQ1_0 as unknown types. The weights are stored in a rotated basis, so the runtime has to apply a Hadamard rotation to the activations, and only PrismML's own llama.cpp fork does that, on CUDA, Metal or CPU. It does not run on AMD cards through Vulkan, and vLLM does not apply the rotation either. We measured it with that fork.