Model Spend Arena2ND ED.
inference-host · 2026-08-14

NVIDIA Build

Official site ↗ · 🇺🇸 United States

inference-host

Serves: 100+ open-weight models via an OpenAI-compatible API on NVIDIA's own GPUs — Nemotron, Llama, Qwen, gpt-oss, DeepSeek and GLM (Zhipu) families
Billing unit: free developer credits (production via NIM / per-token)
Converts to tokens: full
Verifiable pre-pay: partial
Pricing / discount: generous free tier — 1,000 inference credits on signup (up to 5,000), no card, never expires, ~40 requests/min; several models are marked free and do not consume credits

NVIDIA's hosted API catalogue (build.nvidia.com), OpenAI-compatible — a genuinely generous free developer tier over 100+ open models on NVIDIA's own hardware, no credit card. It is not a per-token storefront: the hosted catalogue is free for prototyping (rate-limited ~40 RPM), and production means self-hosting the NIM microservice under NVIDIA AI Enterprise (~$4,500/GPU/year, with a 90-day free eval) or a quoted per-token range of ~$0.10-$10/M. Serves recent open models — strong on its own Nemotron line — but not always the very latest frontier (newest DeepSeek/GLM). US company.