Model Spend Arena2ND ED.
397B · dense · 2026-10-03

Nex-N2-Pro

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLang

AA coding index 59.1 (Artificial Analysis, frozen since 2026-09-11)

Nex AGI's 397B MoE, open-weight, a strong non-frontier open coder.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
IQ1_S81.8 GB
IQ1_M91.4 GB
IQ2_XXS106.3 GB
IQ2_XS118.5 GB
IQ2_S120.5 GB
IQ2_M133.1 GB
Q2_K139.5 GB
Q2_K_L140.5 GB
IQ3_XXS166.1 GB
Q3_K_S172.9 GB
IQ3_XS181.4 GB
Q3_K_M181.5 GB
Q3_K_L189.5 GB
IQ3_M189.6 GB
Q3_K_XL190.4 GB
IQ4_XS212.2 GB
IQ4_NL224.5 GB
Q4_0225.4 GB
Q4_K_S232.8 GB
Q4_K_M241.8 GB
Q4_K_L242.6 GB
Q4_1249.1 GB
Q5_K_S274.0 GB
Q5_K_M283.3 GB
Q6_K342.4 GB
Q8_0421.5 GB

Config tips

~242 GB at Q4 — multi-node. Serve with vLLM / SGLang. Origin of the lab is not well documented, but the weights and the AA score are real.