Model Spend Arena2ND ED.
862B · MoE · 1,048,576 ctx · 2026-10-03

DeepSeek V4 Pro

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLang

AA coding index 68.8 (Artificial Analysis, frozen since 2026-09-11)

A very large MoE and one of the highest-scoring open coding models. The weights are open, but running it is a multi-node affair.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈0 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
Q2_K0.1 GB
Q3_K_S0.1 GB
Q3_K_M0.1 GB
IQ4_XS0.1 GB
Q4_K_S0.1 GB
Q3_K_L0.1 GB
Q4_K_M0.1 GB
Q5_K_S0.1 GB
Q5_K_M0.1 GB
Q6_K0.1 GB
Q8_00.1 GB
F160.2 GB

Config tips

Hundreds of GB even at Q4 — tensor / pipeline parallel across many GPUs with vLLM or SGLang. For anything short of a cluster, rent it per token (leaderboard).