Model Spend Arena2ND ED.
30B · MoE · 262,144 ctx · 2026-10-03

Nemotron 3.5 Lightning

Runs on — · Hugging Face ↗

llama.cpp / Ollama (GGUF)vLLM (BF16 / NVFP4)

AA coding index 26.8 (Artificial Analysis, frozen since 2026-09-11)

A 30B mixture-of-experts with only 3B active parameters, hybrid Mamba2-Transformer, built by NVIDIA for always-on coding agents (43 programming languages). The small active-parameter count is the point, it keeps generation cheap and lets the model run with expert-offload on a 16 GB card, or fast on one bigger GPU. It is a reasoning model, so it thinks a lot before it answers.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 29% (43/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 1 of 148 answers hit the token limit (effective ceiling 99%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard29.0%n=148 · 3 run(s)official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.
BFCL v4 non-live81.1% ± 2.5n=1150 · 1 run(s)BFCL v4 · Q4_K_M GGUF via llama.cpp · one request at a time · a shared cluster GPU
Tool use · BFCL v4 non-live · first-party

81.1% ± 2.5 over 1150 problems

Q4_K_M · a shared cluster GPU · KV f16 · 1 run · spread between runs assumed at 0.25 points, not measured. This is the non-agentic half of BFCL: does the model call the right function with the right arguments. It says nothing about how it behaves over a long agentic conversation, which is a separate measurement. The score is BFCL’s own: the unweighted mean of four groups (simple, multiple, parallel, parallel multiple), where simple is itself the mean of Python, Java and JavaScript. Every group counts the same whatever its size, and irrelevance detection — which we also measure, another 240 problems — is not part of it. The margin follows that same weighting.

Sizes on disk

Real GGUF file sizes = weight VRAM.

QuantizationSize
IQ1_M19.4 GB
IQ2_XXS19.4 GB
IQ2_M19.4 GB
IQ3_XXS19.8 GB
IQ3_S21.2 GB
IQ4_NL21.2 GB
Q3_K_XL21.2 GB
MXFP423.2 GB
Q4_K_S24.5 GB
Q4_K_M25.3 GB
Q4_K_XL25.5 GB
Q5_K_S26.2 GB
Q5_K_M30.2 GB
Q5_K_XL30.4 GB
Q8_035.0 GB
Q6_K_XL35.0 GB
Q8_K_XL38.6 GB
BF1665.9 GB

How to run it

Serving recipe

Pull the GGUF (unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF) in llama.cpp or Ollama. On a 16 GB card use a Q4 quant with expert-offload (-ot to keep the experts on CPU), or skip local hardware and hit the free opencode Zen endpoint at $0.

Config tips

Weights ship as a base model plus GGUF k-quants; the aligned version is the one served free on opencode Zen, which is what we measured. Give it a generous output budget, it reasons a lot before the code.