Model Spend Arena2ND ED.
9B · dense · 262,144 ctx · 2026-10-03

Ornith-1.5 9B

Runs on ≤8 GB · Hugging Face ↗

llama.cppvLLMOllama

Not independently scored by Artificial Analysis

The revision of Ornith-1.0, now published by ornith-ai. Same agentic idea as its predecessor: it learns its own scaffold instead of relying on the one you hand it. On our BCB-Hard it scores 23.6%, BELOW the 25.7% of the 1.0, so the newer version is not better at single-shot code. Any gain may be on the agentic side, which this benchmark does not measure.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 24% (35/148), via official protocol · llama.cpp Q6 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. does not fit the token budget; 5 of 148 answers hit the token limit (effective ceiling 97%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard23.6%n=148 · 3 run(s)official protocol · llama.cpp Q6 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp
BFCL v4 non-live79.1% ± 2.5n=1150 · 1 run(s)BFCL v4 · Q6 GGUF via llama.cpp · one request at a time · a shared cluster GPU
RULER (long context)73.4 ± 6.3N=5/task · at 32K1 quantisations measured
Tool use · BFCL v4 non-live · first-party

79.1% ± 2.5 over 1150 problems

Q6 · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Long context · RULER · first-party
quantGiB4K8K32K
Q66.978.1 ± 6.9 · N=583.6 ± 6.8 · N=573.4 ± 6.3 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.

One or more rows ran on an engine that does not report whether the prompt was truncated, so for those we cannot prove the haystack went in whole.

First-party measurement

~73 tok/s on R9700 32GB (Vulkan), measured by us. Q6_K GGUF, one request at a time, 16384 context, KV as in the model card. Re-measured 2026-08-29 with our speed kit; the run files (flags, GPU layers, device, llama.cpp build) are in results/speed/kit/r9700/.

Size

Q6_K GGUF (~6.9 GB, fits 12 GB with usable context): 6.9 GB on disk.

How to run it

Serving recipe

Grab the GGUF from unsloth/Ornith-1.5-9B-GGUF and serve it with llama-server, the server that ships with llama.cpp (build it from github.com/ggml-org/llama.cpp, or use the ghcr.io/ggml-org/llama.cpp:server-cuda container). Nothing custom is needed — a stock build runs it:
llama-server -m Ornith-1.5-9B-Q6_K.gguf --ctx-size 32768 -ngl 99 -fa on -ctk q4_0 -ctv q4_0 --jinja
The q4_0 KV cache is what buys you the 32K context on a 12 GB card; drop it and the same card only holds about half that.

Config tips

~6.9 GB at Q6_K. An 8 GB card does NOT fit it with usable context; from 12 GB up you get 110K. Same story as the rest of the 9B class: the limit is not quality, it is the context left over after the weights.