Runs on — · Hugging Face ↗
AA coding index 45.0 (Artificial Analysis, frozen since 2026-09-11)
Upstage's 250B-total MoE (320 experts, 1M context) — the Korean lab's first frontier-scale open release, trending on HF but not yet on the AA index.
No GGUF build yet — ~150 GB estimated at 4-bit from the parameter count. Run from safetensors (transformers / vLLM / SGLang).
~140 GB at Q4 — a 2×80 GB node, or heavy CPU offload with a community GGUF on llama.cpp. All 320 experts must be resident despite the small per-token compute.