Model Spend Arena2ND ED.
27B · dense · 262,144 ctx · 2026-10-03

Hemmingway-1

Runs on ≤24 GB · Hugging Face ↗

llama.cpp

Not independently scored by Artificial Analysis — its base model Qwen/Qwen3.8-27B is the closest proxy

The most liked model of its week on Hugging Face (819 likes in days), a 26.9B on Qwen3.5. LICENCE cc-by-nc-4.0: non-commercial.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 28% (41/148), via official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. 1 of 148 answers hit the token limit (effective ceiling 99%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard27.7%n=148 · 3 run(s)official protocol · llama.cpp Q4_K_M · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈3 GB at 32K, fp16). Read from the GGUF header and checked against what llama.cpp actually allocates.

QuantizationSize
IQ2_XXS8.9 GB
IQ2_XS9.1 GB
IQ2_S9.7 GB
IQ2_M10.5 GB
Q2_K10.8 GB
IQ3_XXS12.3 GB
Q3_K_S12.7 GB
IQ3_XS12.8 GB
Q3_K_M13.4 GB
Q3_K_L14.1 GB
IQ3_M14.9 GB
IQ4_XS15.5 GB
Q4_016.3 GB
Q4_K_S16.4 GB
IQ4_NL17.4 GB
Q4_K_M17.4 GB
Q4_117.8 GB
Q4_K_L18.8 GB
Q5_K_S19.6 GB
Q5_K_M20.9 GB
Q6_K_S22.9 GB
Q6_K23.9 GB
Q6_K_L25.0 GB
Q8_029.1 GB
BF1654.7 GB

Config tips

27.7% over three runs, which ties the top of the small tiers. But the licence forbids commercial use, so read it before putting it in a product. Its GGUF is third-party (bartowski).