Model Spend Arena2ND ED.
315B · MoE · 262,144 ctx · 2026-10-03

Motif-3 Beta

Runs on Multi-GPU · Hugging Face ↗

vLLMSGLang

AA coding index 63.5 (Artificial Analysis, frozen since 2026-09-11)

A 314B-total Mixture-of-Experts (384 routed experts) that scores a genuine 62.0 on AA's coding index — strong quality, but the "run it locally" framing around it is optimistic. No GGUF build exists yet, so the size is estimated from params. Weights are available under a non-commercial / research licence — "weights-available", not permissively open.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈5 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
Q2_K121.5 GB
Q3_K_M156.4 GB
Q4_K_M196.0 GB
Q5_K_M228.5 GB

Config tips

Data-centre only (~190 GB even at Q4). No GGUF today, so it is transformers / vLLM / SGLang from safetensors, tensor-parallel across many GPUs. Check the licence before any commercial use. Listed here as the honest ceiling — excellent scores, not a desktop model.