Model Spend Arena2ND ED.
9B · dense · 524,288 ctx · 2026-10-03

K2 Horizon 7B

Runs on ≤24 GB · Hugging Face ↗

llama.cpp built from the IFM branch (model/K2Horizon), GGUFtransformers

AA coding index 38.6 (Artificial Analysis, frozen since 2026-09-11)

Released 2026-09-01 by the Institute of Foundation Models (MBZUAI), Apache-2.0. A dense model despite the name: 9.0B parameters, 36 layers, 32 attention heads over 8 key-value heads, and a declared 512K context. It is one size of a family (0.9B, 3.7B, 7B, 32B, a 36B-A4B MoE and a 375B-A23B), all on the new K2HorizonForCausalLM architecture.

First-party test · not the AA coding index

BigCodeBench-Hard pass@1 28% (42/148), via official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0 pp. The brutal counterpart to HumanEval — where HumanEval saturates near the top, BCB-Hard spreads the field, so this is the number that actually separates coding ability. repeats itself until it runs out; 4 of 148 answers hit the token limit (effective ceiling 97%)

What we measured

Our own runs on this model. Each number carries how many problems it is over and how many times we ran it — these are not comparable with each other, and none of them is on the same scale as the third-party indices on the leaderboard.

testscoreoverconditions
BigCodeBench-Hard28.4%n=148 · 3 run(s)official protocol · llama.cpp BF16 · one request at a time · a shared cluster GPU · mean of 3 runs, range 0.0
RULER (long context)90.8 ± 5.5N=5/task · at 32K3 quantisations measured
Quantisation curve · first-party
quantGiBcode
pass@1
8 GB12 GB16 GB24 GB32 GB48 GB96 GB141 GB
BF1616.828.4%———137K / 72K333K / 176K524K / 385K524K / 524K524K / 524K
Q8_08.927.0%—46K / 24K144K / 76K341K / 180K524K / 285K524K / 493K524K / 524K524K / 524K
Q4_K_M5.223.0%44K / 23K142K / 75K240K / 127K437K / 231K524K / 335K524K / 524K524K / 524K524K / 524K

We have not been able to tie these files back to a published repo, so take the quant names as the labels we ran them under. Every point measured on the same class of card, one pass, one request at a time. One problem is 0.68 points, so anything under ~3.4 points apart is a tie — that is the spread we measure between GPU classes, and our rows are not all on the same card. Each card column is the context left for the KV cache after the weights, KV-Q4 / KV-Q8. A dash means the file does not fit that card, or fits with no room left to work. Context is usually the real constraint, not quality. Computed for one slot (--parallel 1), which is what you get serving yourself. llama-server opens four by default, and on a hybrid that costs real context — the table below each curve shows how much. The card figures assume a dedicated card. Measured on an idle one: the driver keeps 623 MiB of 143,771, so 99.6% is usable. Compute buffers are measured too — a fixed ~130 MiB plus 1 MiB per 1K of context, the same slope on every architecture we checked. If your card also runs your desktop you get noticeably less, and we have not measured that case yet.

Long context · RULER · first-party
quantGiB4K8K32K
BF1616.892.8 ± 5.0 · N=594.3 ± 5.0 · N=590.8 ± 5.5 · N=5
Q8_08.993.3 ± 4.7 · N=594.3 ± 5.0 · N=592.2 ± 4.8 · N=5
Q4_K_M5.291.9 ± 5.4 · N=589.2 ± 6.5 · N=588.8 ± 5.4 · N=5

All 13 RULER tasks, one request at a time, KV cache f16, bench commit c3f5e3b4, seed 42. The score is the mean of the 13; the margin is two standard deviations of the sampling, so it grows as the sample shrinks and as the score drops. Everything here clears the paper’s effective-length bar of 85.6 at every length shown.

This is not a coding score and must not be read as one. RULER asks whether the model finds and follows a fact buried in a long text. On this model Q4_K_M scores 88.8 at 32K against 90.8 for BF16, a difference of 2.0 points — while on BigCodeBench-Hard the same two files score 23.0% and 28.4%. A page that read the long-context number as a general score would be telling the truth and misleading you.

Sizes on disk

Real GGUF file sizes = weight VRAM. Add the KV cache for your context (≈5 GB at 32K, fp16). Read from the GGUF header, but this architecture has never been checked against llama.cpp — treat it as an estimate.

QuantizationSize
Q4_K_M5.6 GB
Q5_06.3 GB
Q5_K_M6.5 GB
Q6_K7.4 GB
Q8_09.6 GB
BF1618.0 GB

Config tips

Upstream llama.cpp did not load it when we measured: we used a build of IFM's own branch (model/K2Horizon, commit 42adf01). IFM only publishes a BF16 GGUF, so the Q8_0 and Q4_K_M points in our curve are quantised by us from that file with llama-quantize, using the same build.