Runs on ≤96 GB · Hugging Face ↗
Not independently scored by Artificial Analysis
The 118B middle Laguna from Poolside (~8B active MoE), open-weight and built for agentic long-horizon coding — the bigger sibling of Laguna XS.2, reported to beat rivals 10x its size. Free on OpenRouter (laguna-s-2.1:free) and opencode Zen.
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| IQ1_S | 33.8 GB |
| IQ1_M | 35.6 GB |
| IQ2_XXS | 37.2 GB |
| IQ2_M | 37.3 GB |
| Q2_K_XL | 39.7 GB |
| IQ3_XXS | 44.3 GB |
| IQ3_S | 48.4 GB |
| Q3_K_M | 54.0 GB |
| Q3_K_XL | 54.1 GB |
| IQ4_XS | 57.6 GB |
| IQ4_NL | 58.7 GB |
| Q4_K_S | 68.6 GB |
| MXFP4 | 71.1 GB |
| Q4_K_M | 73.1 GB |
| Q4_K_XL | 73.4 GB |
| Q5_K_S | 82.7 GB |
| Q5_K_M | 87.9 GB |
| Q5_K_XL | 88.1 GB |
| Q6_K | 97.9 GB |
| Q6_K_XL | 107.1 GB |
| Q8_0 | 125.0 GB |
| Q8_K_XL | 128.1 GB |
| BF16 | 235.2 GB |
~70 GB at Q4 — a single 80 GB card or a DGX Spark. But it's free on OpenRouter :free and opencode Zen, so try it hosted first.