Model Spend Arena2ND ED.
33B · MoE · 262,144 ctx · 2026-10-03

Laguna XS.2

Runs on ≤24 GB · Hugging Face ↗

vLLMSGLang

Not independently scored by Artificial Analysis

Poolside's first open-weight coder — 33B total / ~3B active MoE, Apache-2.0, built for agentic long-horizon coding. Reportedly 68.2% SWE-bench Verified and beating dense models 10x its size; the local-runnable member of the Laguna family (S 2.1 / M.1 are data-centre). Custom `LagunaForCausalLM` arch.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 29% (43/148), via official protocol · vLLM · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. served artifact not verified; different runtime; 1 of 148 answers hit the token limit (effective ceiling 99%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard29.0%n=148 · 3 run(s)official protocol · vLLM · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈2 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
IQ2_XXS9.4 GB
IQ2_XS10.4 GB
IQ2_S10.6 GB
IQ2_M11.6 GB
Q2_K12.2 GB
Q2_K_L12.4 GB
IQ3_XXS14.3 GB
Q3_K_S14.9 GB
IQ3_XS15.6 GB
Q3_K_M15.6 GB
Q3_K_L16.1 GB
IQ3_M16.3 GB
Q3_K_XL16.3 GB
IQ4_XS18.2 GB
Q4_019.2 GB
IQ4_NL19.2 GB
Q4_K_S19.8 GB
Q4_K_M20.5 GB
Q4_K_L20.7 GB
Q4_121.2 GB
Q5_K_S23.2 GB
Q5_K_M24.0 GB
Q5_K_L24.1 GB
Q6_K29.0 GB
Q6_K_L29.1 GB
Q8_035.6 GB
BF1666.9 GB

Config tips

~20 GB at Q4 — a 24 GB+ card or the 32 GB tier. Also on OpenRouter (laguna-xs-2.1:free) and opencode Zen's free set, so you can try the family for $0 before committing local VRAM.