Model Spend Arena2ND ED.
121M · dense · 8,192 ctx · 2026-10-03

Needle 3

Runs on ≤8 GB · Hugging Face ↗

cactus-needle, the vendor's own engine (Python package, iOS, Android, WebAssembly)

Not independently scored by Artificial Analysis

Released 2026-09-16 by Cactus Compute, Apache-2.0. Not a chat model: a 121M-parameter model built only for on-device tool calls, structured extraction and embeddings, on its own NeedleForToolCalling architecture (20 layers, 8K context). Its weights ship compressed to 2 bits in a single 35 MB file, and a grammar built from your function schemas constrains every token it emits.

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BFCL v4 non-live37.8% ± 3.0n=1390 · 3 run(s)BFCL v4 · cactus-needle 3.0.2, the vendor's own engine, on CPU · one request at a time · measured on a home ma
Tool use · BFCL v4 non-live · first-party

37.8% ± 3.0 over 1390 problems

CQ2 · CPU (AMD Ryzen AI 9 HX 370) · 3 runs. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Size

needle3.cact, CQ2-bit, the single file the engine loads: 35 MB on disk.

Config tips

It does not run on llama.cpp or any GGUF runner: it needs the vendor's own engine, cactus-needle, which we used (version 3.0.2, on CPU). We measured it on BFCL only, the function-calling benchmark it was built for; it cannot write the programs BigCodeBench asks for. Compare it with the other small models on this page by that score, not by the coding columns.