Model Spend Arena2ND ED.
291B · MoE · 1,048,576 ctx · 2026-10-03

DeepSeek V4 Flash

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLangllama.cpp

AA coding index 69.1 (Artificial Analysis, frozen since 2026-09-11)

The lighter DeepSeek V4 — a 158B MoE, far more self-hostable than the 862B V4 Pro while still a capable open coder. MIT-licensed, 1M context.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 33% (49/148), via official protocol · llama.cpp UD-IQ3_S · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard33.1%n=148 · 3 run(s)official protocol · llama.cpp UD-IQ3_S · one request at a time · a shared cluster GPU · mean of 3 runs, range
Quantisation curve · first-party
quantGiBcode
pass@1
96 GB141 GB
UD-IQ3_S109.333.1%—1038K / 549K
UD-IQ3_XXS95.932.4%—1048K / 857K
UD-Q2_K_XL90.230.4%13K / 6K1048K / 988K
UD-IQ2_M84.734.5%251K / 133K1048K / 1048K
UD-IQ2_XXS84.627.0%255K / 135K1048K / 1048K
UD-IQ1_M80.927.7%416K / 220K1048K / 1048K
UD-IQ1_S76.929.0%589K / 312K1048K / 1048K

We have not been able to tie these files back to a published repo, so take the quant names as the labels we ran them under. Every point measured on the same class of card, one pass, one request at a time. One problem is 0.68 points, so anything under ~3.4 points apart is a tie — that is the spread we measure between GPU classes, and our rows are not all on the same card. Each card column is the context left for the KV cache after the weights, KV-Q4 / KV-Q8. A dash means the file does not fit that card, or fits with no room left to work. Context is usually the real constraint, not quality. Computed for one slot (--parallel 1), which is what you get serving yourself. llama-server opens four by default, and on a hybrid that costs real context — the table below each curve shows how much. The card figures assume a dedicated card. Measured on an idle one: the driver keeps 623 MiB of 143,771, so 99.6% is usable. Compute buffers are measured too — a fixed ~130 MiB plus 1 MiB per 1K of context, the same slope on every architecture we checked. If your card also runs your desktop you get noticeably less, and we have not measured that case yet.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈0 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
Q8_010.9 GB
BF1611.3 GB
IQ1_S82.5 GB
IQ1_M86.9 GB
IQ2_XXS90.9 GB
IQ2_M90.9 GB
Q2_K_XL96.8 GB
IQ3_XXS104.2 GB
IQ3_S116.1 GB
Q3_K_M128.1 GB
Q3_K_XL128.2 GB
IQ4_NL136.7 GB
IQ4_XS136.7 GB
Q4_K_XL155.1 GB
Q8_K_XL161.9 GB

Config tips

~138 GB at Q4 — a 2×80 GB node, or heavy CPU offload with llama.cpp on less. The most accessible of the strong DeepSeek line for local use.