Model Spend Arena2ND ED.
8B · dense · 65,536 ctx · 2026-10-03

Bonsai 8B

Runs on ≤8 GB · Hugging Face ↗

llama.cpp (patched)

Not independently scored by Artificial Analysis — its base model prism-ml/Bonsai-8B-unpacked is the closest proxy

PrismML's 1-bit 8B — ~1.15 GB of weights, small enough to run on a phone. With the patched runtime it is genuinely coherent (not the garbage stock Ollama produces) and surprisingly capable on easy problems for its size.

First-party test · not the AA coding index

HumanEval pass@1 63% (19/30), via local GPU (patched llama.cpp, Q1_0). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 7% (11/148), via official protocol · llama.cpp Q1_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard7.4%n=148 · 3 run(s)official protocol · llama.cpp Q1_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0
BFCL v4 non-live42.2% ± 3.1n=1150 · 1 run(s)BFCL v4 · Q1_0 ternary GGUF via llama.cpp · one request at a time · a shared cluster GPU
RULER (long context)68.0 ± 8.8N=5/task · at 32K1 quantisations measured
Tool use · BFCL v4 non-live · first-party

42.2% ± 3.1 over 1150 problems

Q1_0 · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Long context · RULER · first-party
quantGiB4K8K32K
Q1_01.186.1 ± 6.6 · N=577.4 ± 8.8 · N=568.0 ± 8.8 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.

First-party measurement

~156 tok/s on RTX 4060 Ti, measured by us.

Size

1-bit: 1.1 GB on disk.

How to run it

Serving recipe

Needs a 1-bit-capable (PrismML patched) llama.cpp build — not stock Ollama, where it loads but the output is incoherent garbage.
1. Build PrismML's fork (has q1_0 kernels): `cmake -B build -DGGML_CUDA=ON && cmake --build build -j`
2. `build/bin/llama-server -ngl 99 -m Bonsai-8B-Q1_0.gguf -c 8192`
With the right runtime it IS coherent — we measured HumanEval 63% and ~160 tok/s in 1.15 GB. The 1-bit quant only really shows on hard problems (BCB-Hard 7%). A superb size/speed trade for drafts and edge devices.

Config tips

Tiny footprint and blazing fast (~160 tok/s). Holds up on easy coding (HumanEval 63%) but the 1-bit quant shows on hard problems (BCB-Hard 7%) — a great draft/edge model, not a heavy agentic coder. Needs the 1-bit-capable PrismML build.