Model Spend Arena2ND ED.
8B · dense · 131,072 ctx · 2026-10-10

xLAM 2 8B

Runs on ≤8 GB · Hugging Face ↗

llama.cpp

Not independently scored by Artificial Analysis

The 8B of the xLAM 2 family, Salesforce's function-calling models on Llama 3.1, with its own GGUF. Licence cc-by-nc-4.0: not for commercial use.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 6% (9/148), via official protocol · llama.cpp F16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard6.1%n=148 · 3 run(s)official protocol · llama.cpp F16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 p

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈4 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
Q2_K3.2 GB
Q3_K_S3.7 GB
Q3_K_L4.3 GB
Q4_04.7 GB
Q4_K_S4.7 GB
Q4_K_M4.9 GB
Q5_05.6 GB
Q5_K_S5.6 GB
Q5_K_M5.7 GB
Q6_K6.6 GB
Q8_08.5 GB
F1616.1 GB

Config tips

Built for tool calls, not for writing code, and BigCodeBench shows it: 6.1% over three runs, at F16. Look at it for routing tools, next to its 3B sibling.