Model Spend Arena2ND ED.
291B · MoE · 1,048,576 ctx · 2026-08-14

DeepSeek V4 Flash

Runs on Data centre · Hugging Face ↗

vLLMSGLangllama.cpp

Coding index 69.1 (Artificial Analysis)

The lighter DeepSeek V4 — a 158B MoE, far more self-hostable than the 862B V4 Pro while still a capable open coder. MIT-licensed, 1M context.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16).

QuantizationSize
Q8_010.9 GB
BF1611.3 GB
Q2_K96.8 GB
Q3_K_M128.1 GB
IQ4_XS136.7 GB
Q4155.1 GB
Q8161.9 GB

Config tips

~138 GB at Q4 — a 2×80 GB node, or heavy CPU offload with llama.cpp on less. The most accessible of the strong DeepSeek line for local use.