Runs on Data centre · Hugging Face ↗
Coding index 63.5 (Artificial Analysis)
A 314B-total Mixture-of-Experts (384 routed experts) that scores a genuine 62.0 on AA's coding index — strong quality, but the "run it locally" framing around it is optimistic. No GGUF build exists yet, so the size is estimated from params. Weights are available under a non-commercial / research licence — "weights-available", not permissively open.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈21 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q2_K | 121.5 GB |
| Q3_K_M | 156.4 GB |
| Q4_K_M | 196.0 GB |
| Q5_K_M | 228.5 GB |
| Q6_K | 263.1 GB |
| Q8_0 | 338.2 GB |
Data-centre only (~190 GB even at Q4). No GGUF today, so it is transformers / vLLM / SGLang from safetensors, tensor-parallel across many GPUs. Check the licence before any commercial use. Listed here as the honest ceiling — excellent scores, not a desktop model.