Model Spend Arena2ND ED.
4B · dense · 262,144 ctx · 2026-10-03

Nanbeige 4.2-3B

Runs on ≤8 GB · Hugging Face ↗

transformers

Not independently scored by Artificial Analysis — its base model Nanbeige/Nanbeige4.2-3B-Base is the closest proxy

A dense 3B from Nanbeige LLM Lab (BOSS Zhipin), pre-trained on ~23T tokens. Tiny — the "runs on anything" end of the table, laptop iGPU or CPU.

First-party test · not the AA coding index

HumanEval pass@1 81% (13/16), via local GPU (transformers). Measured by us on this hardware, as a rough sanity check — HumanEval is a different, easier, partly-contaminated benchmark than Artificial Analysis’ composite, so it is not comparable to the coding-index column. Would not load in llama.cpp/Ollama at all — its custom 'nanbeige' architecture is unknown to them. We ran it via transformers with trust_remote_code in a GPU container, which needed an old transformers pin (for a rope_scaling change), sentencepiece, and a prebuilt flash_attn wheel. Worth it: it codes coherently and well for a 3B — the opposite of the Bonsai quants.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 16% (24/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. does not fit the token budget; 7 of 148 answers hit the token limit (effective ceiling 95%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard16.2%n=148 · 3 run(s)official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.
BFCL v4 non-live75.2% ± 2.8n=1150 · 1 run(s)BFCL v4 · Q4KM GGUF via llama.cpp · one request at a time · a shared cluster GPU
RULER (long context)75.4 ± 6.5N=5/task · at 32K1 quantisations measured
Tool use · BFCL v4 non-live · first-party

75.2% ± 2.8 over 1150 problems

Q4KM · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Long context · RULER · first-party
quantGiB4K8K32K
Q4_K_M2.471.7 ± 6.4 · N=567.5 ± 5.2 · N=575.4 ± 6.5 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. Quantising hits the two differently, so a model that still retrieves at length may already have lost the code.

First-party measurement

~9 tok/s on RTX 4060 Ti, measured by us. RTX 4060 Ti via transformers+flash (bf16); ~1.0 tok/s CPU. Drops with length — the custom modeling's KV cache is ineffective (O(n^2)).

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
Q2_K1.8 GB
IQ2_M1.9 GB
IQ3_XXS1.9 GB
IQ3_XS2.0 GB
Q3_K_S2.1 GB
IQ3_M2.2 GB
Q3_K_M2.2 GB
Q3_K_L2.3 GB
Q2_K_L2.3 GB
IQ4_XS2.4 GB
Q4_02.5 GB
IQ4_NL2.5 GB
Q4_K_S2.6 GB
Q4_K_M2.7 GB
Q4_12.7 GB
Q3_K_XL2.8 GB
Q5_K_S3.0 GB
Q4_K_L3.1 GB
Q5_K_M3.1 GB
Q5_K_L3.4 GB
Q6_K3.6 GB
Q6_K_L3.8 GB
Q8_04.4 GB
BF168.3 GB

How to run it

Serving recipe

Won't load in llama.cpp / Ollama / LM Studio — the custom `nanbeige` architecture has no GGUF converter, so ignore the "runs on anything" first impression. Run it from safetensors with Hugging Face transformers and `trust_remote_code=True`.
- GPU: `pip install transformers==4.44.2` + the matching prebuilt flash-attn wheel; load with `attn_implementation="flash_attention_2"`, bf16.
- CPU: don't install flash-attn (it's CUDA-only). transformers still does a *static* import check, so drop a tiny empty `flash_attn` stub package on the path to satisfy it — `is_flash_attn_2_available()` is False on a CUDA-less torch, so the real import never fires. Load with `attn_implementation="eager"`, float32.
We measured ~9 tok/s on an RTX 4060 Ti and ~1 tok/s on a 24-core CPU. It's slower than a GGUF model its size because the custom modeling's KV cache is ineffective (throughput drops as the answer grows).

Config tips

Q4 fits well under 6 GB with room for long context. Good as a fast draft / autocomplete model, not for hard agentic coding. Runs on Apple Silicon via MLX.