Model Spend Arena2ND ED.
124B · MoE · 262,144 ctx · 2026-08-14

Ling-3.0-flash

Runs on Data centre · Hugging Face ↗

vLLMSGLang

Coding index 50.6 (Artificial Analysis)

inclusionAI's open MoE (Ling 3.0; 124B total, 5.1B active, hybrid reasoning on by default). AA now scores it — coding index 50.6 — and our first-party BCB-Hard (36%) confirms it punches above its active size. Free in opencode Zen's rotating set; on OpenRouter the :free route rotated over to Ling 3.0 Tiny, so flash is paid-only there now ($0.08/$0.22 per M).

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 36% (44/121), via OpenRouter :free route (hosted) — comparable full-148. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. First-party BCB-Hard via the free hosted route — punches above its size; AA now scores it 50.6.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈23 GB at 32K, fp16).

QuantizationSize
Q4_K_S74.2 GB
Q8_0133.1 GB
Q4_K_M154.6 GB
Q4156.3 GB
Q5_K_M177.8 GB
Q6_K209.8 GB
IQ4_XS209.8 GB
BF16248.9 GB

Config tips

~70 GB at Q4 — a single 80 GB card or a 2-GPU split; community GGUFs exist but are days old and unvetted. We tested it hosted, back when it rode the OpenRouter <code>:free</code> route.