Model Spend Arena2ND ED.
9B · dense · 2026-10-03

Ornith-1.0 9B

Runs on ≤8 GB · Hugging Face ↗

llama.cppvLLMOllama

Not independently scored by Artificial Analysis

DeepReinforce's self-scaffolding agentic coder (MIT, Jun 2026), built on Gemma 4 / Qwen 3.5. Unusual idea: it learns its own agent scaffold — memory layout, retry logic, tool orchestration — during RL, rather than having engineers hard-code it. Coding-specialised, smallest of a family up to 397B.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 26% (38/148), via official protocol · llama.cpp Q6 · one request at a time · a shared cluster GPU · single run. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. does not fit the token budget; conditions not recorded; only 1 run; 10 of 148 answers hit the token limit (effective ceiling 93%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard25.7%n=148 · 1 run(s)official protocol · llama.cpp Q6 · one request at a time · a shared cluster GPU · single run
BFCL v4 non-live80.6% ± 2.5n=1150 · 1 run(s)BFCL v4 · Q6 GGUF via llama.cpp · one request at a time · a shared cluster GPU
RULER (long context)73.8 ± 4.3N=5/task · at 32K1 quantisations measured
Tool use · BFCL v4 non-live · first-party

80.6% ± 2.5 over 1150 problems

Q6 · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Long context · RULER · first-party
quantGiB4K8K32K
Q67.070.2 ± 3.9 · N=581.0 ± 4.2 · N=573.8 ± 4.3 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.

First-party measurement

~72 tok/s on R9700 32GB (Vulkan), measured by us. Q6_K GGUF, one request at a time, 16384 context, KV as in the model card. Re-measured 2026-08-29 with our speed kit; the run files (flags, GPU layers, device, llama.cpp build) are in results/speed/kit/r9700/.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
IQ2_M3.9 GB
IQ3_XXS4.2 GB
Q2_K_XL4.3 GB
Q3_K_S4.3 GB
Q3_K_M4.7 GB
Q3_K_XL5.1 GB
IQ4_XS5.3 GB
Q4_05.4 GB
Q4_K_S5.4 GB
IQ4_NL5.5 GB
Q4_K_M5.7 GB
Q4_15.9 GB
Q4_K_XL6.0 GB
Q5_K_S6.4 GB
Q5_K_M6.5 GB
Q5_K_XL6.7 GB
Q6_K7.5 GB
Q6_K_XL8.8 GB
Q8_09.5 GB
Q8_K_XL13.0 GB
BF1617.9 GB

Config tips

~6 GB at Q4, fits an 8 GB card. Built for AGENTIC coding (multi-turn, tools), so a single-shot benchmark like our BCB-Hard likely understates it — read its number as a floor for how it behaves in an agent loop.