Runs on ≤8 GB · Hugging Face ↗
Not independently scored by Artificial Analysis
Released 2026-09-16 by Cactus Compute, Apache-2.0. Not a chat model: a 121M-parameter model built only for on-device tool calls, structured extraction and embeddings, on its own NeedleForToolCalling architecture (20 layers, 8K context). Its weights ship compressed to 2 bits in a single 35 MB file, and a grammar built from your function schemas constrains every token it emits.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BFCL v4 non-live | 37.8% ± 3.0 | n=1390 · 3 run(s) | BFCL v4 · cactus-needle 3.0.2, the vendor's own engine, on CPU · one request at a time · measured on a home ma |
37.8% ± 3.0 over 1390 problems
CQ2 · CPU (AMD Ryzen AI 9 HX 370) · 3 runs. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.
needle3.cact, CQ2-bit, the single file the engine loads: 35 MB on disk.
It does not run on llama.cpp or any GGUF runner: it needs the vendor's own engine, cactus-needle, which we used (version 3.0.2, on CPU). We measured it on BFCL only, the function-calling benchmark it was built for; it cannot write the programs BigCodeBench asks for. Compare it with the other small models on this page by that score, not by the coding columns.