Model Spend Arena2ND ED.
inference-host · 2026-08-14

Nscale

Official site ↗ · 🇬🇧 United Kingdom

inference-host

Serves: open-weight serverless — gpt-oss, Qwen2.5-Coder / Qwen3, Llama 4, Devstral, DeepSeek R1 distills (no full DeepSeek V4, no GLM, no Kimi)
Billing unit: tokens (per 1M, pay-as-you-go)
Converts to tokens: full
Verifiable pre-pay: partial
Pricing / discount: ~$5 free signup credit; no minimums

UK-domiciled European "sovereign" AI cloud (London HQ; data centres in Norway / UK / Iceland). Two products: a GPU-rental hyperscaler (per-hour, its main business) and a smaller per-token serverless API. The serverless roster is mid-tier for coding — Qwen2.5-Coder, Devstral and gpt-oss — but the current open flagships (DeepSeek V4, GLM-5.2, Kimi) are absent, and DeepSeek is offered only as R1 distills. Quantization not disclosed; the full catalogue sits behind a signup gate.

What it charges

List price per model (input / output per 1M tokens, USD), for the models it actually serves — the detail the five-model comparison table cannot show. Prices as published by the host; re-verify before relying on them.

ModelPrice in / outQuant
gpt-oss-120b$0.10 / $0.40
gpt-oss-20b$0.05 / $0.20
Qwen2.5-Coder 32B$0.06 / $0.20
Qwen3 235B A22B (2507)$0.20 / $0.60
Devstral Small 2505$0.10 / $0.30
Llama 4 Scout 17B$0.09 / $0.29
DeepSeek R1 Distill Qwen 14B$0.20