Runs on ≤24 GB · Hugging Face ↗
Not independently scored by Artificial Analysis
Agnes 3.0 Flash (apache-2.0), 33.1B on its own AgnesForConditionalGeneration architecture. Artificial Analysis lists the family but publishes no coding index for it, so our run is the only coding number it has.
BigCodeBench-Hard pass@1 26% (38/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability.
Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.
| test | score | over | conditions |
|---|---|---|---|
| BigCodeBench-Hard | 25.7% | n=148 · 3 run(s) | official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0. |
Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.
| Quantization | Size |
|---|---|
| Q4_K_M | 19.8 GB |
| Q5_K_M | 23.0 GB |
| Q6_K | 26.4 GB |
| Q8_0 | 34.2 GB |
Its GGUF comes from a third party (0xKitkat), not from Agnes: check the quantisation yourself before trusting a size. Own architecture, so a llama.cpp that loads it is not a given.