Model Spend Arena2ND ED.
118B · MoE · 1,048,576 ctx · 2026-10-03

Laguna S 2.1

Runs on ≤96 GB · Hugging Face ↗

vLLMSGLangllama.cpp

Not independently scored by Artificial Analysis

The 118B middle Laguna from Poolside (~8B active MoE), open-weight and built for agentic long-horizon coding — the bigger sibling of Laguna XS.2, reported to beat rivals 10x its size. Free on OpenRouter (laguna-s-2.1:free) and opencode Zen.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ1_S33.8 GB
IQ1_M35.6 GB
IQ2_XXS37.2 GB
IQ2_M37.3 GB
Q2_K_XL39.7 GB
IQ3_XXS44.3 GB
IQ3_S48.4 GB
Q3_K_M54.0 GB
Q3_K_XL54.1 GB
IQ4_XS57.6 GB
IQ4_NL58.7 GB
Q4_K_S68.6 GB
MXFP471.1 GB
Q4_K_M73.1 GB
Q4_K_XL73.4 GB
Q5_K_S82.7 GB
Q5_K_M87.9 GB
Q5_K_XL88.1 GB
Q6_K97.9 GB
Q6_K_XL107.1 GB
Q8_0125.0 GB
Q8_K_XL128.1 GB
BF16235.2 GB

Config tips

~70 GB at Q4 — a single 80 GB card or a DGX Spark. But it's free on OpenRouter :free and opencode Zen, so try it hosted first.