Model Spend Arena2ND ED.
router · 2026-08-14

Hugging Face Inference Providers

Official site ↗ · 🇺🇸 United States

router

Serves: open-weight models across partner hosts (Cerebras, Groq, Together, Fireworks, Novita, DeepInfra, Scaleway, Z.ai …) through one HF token
Billing unit: tokens (pass-through provider rate)
Converts to tokens: full
Verifiable pre-pay: yes
Pricing / discount: no markup — you pay the partner's published per-token rate; monthly credit included (Free $0.10, PRO $2, Team / Enterprise $2/seat)

A router, not a host: one HF token routes to leading inference providers, OpenAI-compatible (router.huggingface.co/v1), with no HF markup — you pay the partner's rate. The provider is chosen by policy suffix (:fastest default, :cheapest, :preferred, or an explicit :provider). A small monthly credit ($0.10 free / $2 on PRO at $9/mo) then pay-as-you-go. It serves whatever its partners serve (open weights only), so price and quantization come from the chosen partner, not HF. Closed models (GPT-5.x, Claude) are not routed.