Runs on Data centre · Hugging Face ↗
Coding index 44.7 (Artificial Analysis)
Upstage's 250B-total MoE (320 experts, 1M context) — the Korean lab's first frontier-scale open release, trending on HF but not yet on the AA index.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈6 GB at 32K, fp16).
| Quantization | Size |
|---|---|
| Q2_K | 95.4 GB |
| IQ4_XS | 136.2 GB |
| Q6_K | 205.6 GB |
~140 GB at Q4 — a 2×80 GB node, or heavy CPU offload with a community GGUF on llama.cpp. All 320 experts must be resident despite the small per-token compute.