Model Spend Arena2ND ED.
inference-host · 2026-08-14

Featherless AI

Official site ↗ · 🇸🇬 Singapore

inference-host

Serves: open-weight models only (40k+ pulled from Hugging Face) — DeepSeek V4 Pro/Flash, GLM-5.2, Qwen3, Kimi K2.x, Llama
Billing unit: flat monthly subscription by concurrency (not tokens)
Converts to tokens: none
Verifiable pre-pay: partial
Pricing / discount: flat rate — "unlimited tokens" within a concurrency budget, not a per-token discount

Serverless host of open weights (Singapore-founded by the RWKV team, AMD/Airbus-backed). Unusual model: you buy concurrency, not tokens — Premium/Chat $25/mo (4 units, 32K context), Agent $100/$200 (8 units, 256K). Overflow returns HTTP 429, no overage. It does serve DeepSeek V4 and GLM-5.2, but there is no per-token price to compare and no quantization disclosed, so cost per unit of real work is not derivable up front. A Feather Per-Request credit plan exists but its rate card is unpublished.