Runs on Multi-GPU · Hugging Face ↗
AA coding index 63.5 (Artificial Analysis, frozen since 2026-09-11)
A 314B-total Mixture-of-Experts (384 routed experts) that scores a genuine 62.0 on AA's coding index — strong quality, but the "run it locally" framing around it is optimistic. No GGUF build exists yet, so the size is estimated from params. Weights are available under a non-commercial / research licence — "weights-available", not permissively open.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈5 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| Q2_K | 121.5 GB |
| Q3_K_M | 156.4 GB |
| Q4_K_M | 196.0 GB |
| Q5_K_M | 228.5 GB |
Data-centre only (~190 GB even at Q4). No GGUF today, so it is transformers / vLLM / SGLang from safetensors, tensor-parallel across many GPUs. Check the licence before any commercial use. Listed here as the honest ceiling — excellent scores, not a desktop model.