Runs on Multi-GPU · Hugging Face ↗
AA coding index 59.1 (Artificial Analysis, frozen since 2026-09-11)
Nex AGI's 397B MoE, open-weight, a strong non-frontier open coder.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.
| Quantization | Size |
|---|---|
| IQ1_S | 81.8 GB |
| IQ1_M | 91.4 GB |
| IQ2_XXS | 106.3 GB |
| IQ2_XS | 118.5 GB |
| IQ2_S | 120.5 GB |
| IQ2_M | 133.1 GB |
| Q2_K | 139.5 GB |
| Q2_K_L | 140.5 GB |
| IQ3_XXS | 166.1 GB |
| Q3_K_S | 172.9 GB |
| IQ3_XS | 181.4 GB |
| Q3_K_M | 181.5 GB |
| Q3_K_L | 189.5 GB |
| IQ3_M | 189.6 GB |
| Q3_K_XL | 190.4 GB |
| IQ4_XS | 212.2 GB |
| IQ4_NL | 224.5 GB |
| Q4_0 | 225.4 GB |
| Q4_K_S | 232.8 GB |
| Q4_K_M | 241.8 GB |
| Q4_K_L | 242.6 GB |
| Q4_1 | 249.1 GB |
| Q5_K_S | 274.0 GB |
| Q5_K_M | 283.3 GB |
| Q6_K | 342.4 GB |
| Q8_0 | 421.5 GB |
~242 GB at Q4 — multi-node. Serve with vLLM / SGLang. Origin of the lab is not well documented, but the weights and the AA score are real.