Model Spend Arena2ND ED.
inference-host · 2026-08-14

OVHcloud AI Endpoints

Official site ↗ · 🇫🇷 France

inference-host

Serves: open-weight, EU-sovereign — Llama, Mistral (incl. Codestral / Small 3.2), Qwen3 Coder, gpt-oss, DeepSeek R1 distill (no DeepSeek V4, no GLM, no Kimi)
Billing unit: tokens (per 1M, priced in EUR)
Converts to tokens: full
Verifiable pre-pay: yes
Pricing / discount: sign-up credit for new Public Cloud projects; some models free; discounted batch tier

France-based (OVHcloud, Roubaix), GDPR / EU data residency, served from Gravelines; states customer data is never used to train models. Open per-token catalogue in EUR (from ~€0.04/M). Relevant coding options are Qwen3 Coder 30B and gpt-oss-120b — the current open coding SOTA (DeepSeek V4, GLM-5.2, Kimi) is not onboarded. Prices are EUR, so the USD-normalised cost floats with FX. Quantization not disclosed.

What it charges

List price per model (input / output per 1M tokens, EUR), for the models it actually serves — the detail the five-model comparison table cannot show. Prices as published by the host; re-verify before relying on them.

ModelPrice in / outQuant
gpt-oss-120b€0.08 / €0.40
gpt-oss-20b€0.04 / €0.15
Qwen3 Coder 30B A3B€0.07 / €0.26
Qwen3 32B€0.09 / €0.25
Mistral Small 3.2 24B€0.09 / €0.28
Llama 3.3 70B€0.67
Qwen3.5 397B A17B€0.60 / €3.60