Model Spend Arena2ND ED.
inference-host · 2026-08-14

Scaleway Generative APIs

Official site ↗ · 🇫🇷 France

inference-host

Serves: open-weight, EU-sovereign — GLM-5.2, Qwen3 Coder, Mistral (Devstral 2, Large 3, Small 3.2), Llama, gpt-oss, MiniMax-M2.5 (DeepSeek R1 distills only, no DeepSeek V4, no Kimi)
Billing unit: tokens (per 1M, priced in EUR)
Converts to tokens: full
Verifiable pre-pay: yes
Pricing / discount: first 1M tokens free; Batch API -50%

France-based (Scaleway, Iliad group), served from Paris only; GDPR / EU residency and states it does not read or reuse prompt/output content. The most relevant of the HF-partner hosts for coding: it serves GLM-5.2 and Qwen3 Coder and — unusually — DISCLOSES quantization openly via model-id suffixes (:fp8, :bf16, :fp4, :int4, :awq). Prices are EUR (normalise to USD with FX). DeepSeek is only R1 distills, not V4. Single region (Paris), so no failover.

What it charges

List price per model (input / output per 1M tokens, EUR), for the models it actually serves — the detail the five-model comparison table cannot show. Prices as published by the host; re-verify before relying on them.

ModelPrice in / outQuant
GLM-5.2€1.80 / €5.50fp8
Qwen3 Coder 30B A3B€0.20 / €0.80fp8
Qwen3.6 35B A3B€0.25 / €1.50
Qwen3 235B A22B€0.75 / €2.25
Mistral Small 3.2 24B€0.15 / €0.35
Mistral Medium 3.5€1.50 / €7.50
Llama 3.3 70B€0.90
gpt-oss-120bfp4