Model Spend Arena2ND ED.
inference-host · 2026-08-14

RunPod

Official site ↗ · origin not confirmed

inference-host

Serves: Public Endpoints — a LIMITED set of pre-deployed OpenAI-compatible models, image/video/audio-first (Flux, Qwen Image, WAN 2.5, Kling, SORA 2) with only a few text/code models (Qwen3 32B, IBM Granite, Moonshot Kimi). Any other open weight you self-deploy on Serverless (per-second GPU) or Pods (per-hour GPU).
Billing unit: tokens
Converts to tokens: full
Verifiable pre-pay: partial
Pricing / discount: none

Primarily a GPU-rental / serverless platform, but it DOES expose an OpenAI-compatible inference API via Public Endpoints (api.runpod.ai/v2/...), so a coding agent can point at it. Caveats for coding: the managed text catalogue is thin — Qwen3 32B is the main code-capable one, at a steep $10 per 1M tokens — and the platform is image/video-first. It is NOT an OpenRouter provider (no routed named-model prices); for arbitrary weights you self-deploy and pay GPU-time (Serverless per-second, Pods per-hour), not per-named-token. Sits between raw GPU rental (Vast/Lambda) and a full per-token host (DeepInfra/Together).