Runs on ≤8 GB · Hugging Face ↗
Not independently scored by Artificial Analysis
The 8B of the xLAM 2 family, Salesforce's function-calling models on Llama 3.1, with its own GGUF. Licence cc-by-nc-4.0: not for commercial use.
BigCodeBench-Hard pass@1 6% (9/148), via official protocol · llama.cpp F16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 6.1% | n=148 · 3 run(s) | official protocol · llama.cpp F16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 p |
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈4 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.
| Quantization | Size |
|---|---|
| Q2_K | 3.2 GB |
| Q3_K_S | 3.7 GB |
| Q3_K_L | 4.3 GB |
| Q4_0 | 4.7 GB |
| Q4_K_S | 4.7 GB |
| Q4_K_M | 4.9 GB |
| Q5_0 | 5.6 GB |
| Q5_K_S | 5.6 GB |
| Q5_K_M | 5.7 GB |
| Q6_K | 6.6 GB |
| Q8_0 | 8.5 GB |
| F16 | 16.1 GB |
Built for tool calls, not for writing code, and BigCodeBench shows it: 6.1% over three runs, at F16. Look at it for routing tools, next to its 3B sibling.