Model Spend Arena2ND ED.
33B · MoE · 1,048,576 ctx · 2026-10-03

CobrIX-1.5-Coder-Flash-33B-A13B

Runs on ≤24 GB · Hugging Face ↗

llama.cpp

Not independently scored by Artificial Analysis — its base model empero-ai/Qwythos-9B-v2 is the closest proxy

A 33B MoE with 13B active (MIT), sold as a coding model, built on Qwen3.5. It is a preview.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 3% (4/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard2.7%n=148 · 3 run(s)official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈1 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
Q4_K_M20.3 GB
Q5_K_M23.6 GB

Config tips

READ THE SCORE BEFORE DOWNLOADING: 2.7% on BigCodeBench-Hard over three runs, the lowest in this catalogue by a wide margin, from a model whose name says Coder. Either the preview is broken for this task or its chat template does not survive the GGUF conversion. We publish it because the measurement is the point.