Model Spend Arena2ND ED.
35B · dense · 2026-08-14

Ornith-1.0 35B

Runs on ≤32 GB · Hugging Face ↗

llama.cppvLLMOllama

Not independently scored by Artificial Analysis

The 35B-total MoE middle sibling of DeepReinforce's self-scaffolding agentic coder (MIT), a Qwen3.5-MoE. A genuine step up from the 9B on our first-party BCB-Hard, and among the stronger open coders the AA index doesn't rank. Agentic by design, so a single-shot benchmark is a floor for its agent-loop behaviour.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 35% (42/121), via rented RTX PRO 6000 Blackwell 96 GB (vLLM, BF16) — comparable full-148, conc64. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. The 35B MoE — a clear step up from the 9B. Agentic by design, so a single-shot benchmark is a floor for its agent-loop behaviour.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
Q2_K12.3 GB
Q3_K_M16.7 GB
IQ4_XS17.8 GB
Q4_K_S20.9 GB
MXFP421.7 GB
Q4_K_M22.1 GB
Q422.3 GB
Q5_K_M26.5 GB
Q8_036.9 GB
Q838.2 GB
Q6_K61.2 GB
BF1669.4 GB

Config tips

~20 GB at Q4 — needs a 24 GB+ card or the 32 GB tier; the MoE keeps decode fast. We benchmarked it on a rented RTX PRO 6000 Blackwell 96 GB (vLLM, BF16) since it doesn't fit a 16 GB card.