Model Spend Arena2ND ED.
4B · dense · 131,072 ctx · 2026-10-03

Phi-3 mini 128k

Runs on ≤8 GB · Hugging Face ↗

llama.cpp

Not independently scored by Artificial Analysis

A 3.8B from Microsoft, released in 2024 and one of the first small models people actually ran at home. We measure it as a reference point, not as a recommendation: it is the yardstick for how far small models have come at calling tools. It predates the tool-calling era, so it works in prompting mode — the functions go into the system prompt as text rather than through an API field.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 5% (7/148), via official protocol · llama.cpp Q8_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 1 of 148 answers hit the token limit (effective ceiling 99%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard4.7%n=148 · 3 run(s)official protocol · llama.cpp Q8_0 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0
BFCL v4 non-live32.6% ± 2.7n=1150 · 1 run(s)BFCL v4 · Q8_0 GGUF via llama.cpp · one request at a time · a shared cluster GPU · prompting mode
Tool use · BFCL v4 non-live · first-party

32.6% ± 2.7 over 1150 problems

Q8_0 · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈13 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
Q2_K1.4 GB
IQ3_XS1.6 GB
IQ3_S1.7 GB
Q3_K_S1.7 GB
IQ3_M1.8 GB
Q3_K_M1.9 GB
Q3_K_L2.0 GB
IQ4_XS2.1 GB
Q4_K_S2.2 GB
Q4_K_M2.3 GB
Q5_K_S2.6 GB
Q5_K_M2.7 GB
Q6_K3.1 GB
Q8_04.1 GB
F167.6 GB

Config tips

3.8 GB at Q8_0 and about 17 GB of VRAM in total with a 32K context, so it fits any card we test on. If you want a small model for tool use today, the ones measured above it here are two to three times better at it.