Official site · 🇺🇸 United States
| Tier | $/month | Included usage |
|---|---|---|
| Free | $0.00 | — |
| Scale | $200.00 | — |
console only quota converts to units
As advertised. Free: 200 requests/month, then usage-based per-token with practically no rate limits. Scale: flat-rate reserved capacity from ~$200/mo (~2x modelled token capacity, marketed up to ~6.6-7.0x cheaper than pay-per-token); Dedicated = reserved B200 by the GPU-hour.
What that actually means. Now a pure usage-based inference host: pay per token, with a 200-request/month free allowance and an optional Scale flat-rate for steady volume. The old credit tiers (Free 250K / Starter $20 / Pro $60 / Scale $400) are GONE.
Where to check what is left. account dashboard
Models covered. GLM-5.2 744B, Qwen 27B dense, MiniMax 230B MoE, DeepSeek (1M context), Kimi K3, plus its own Fast Apply / WarpGrep / Compact / Reflex / Model Router endpoints
checked 2026-08-13 · high (read from the rendered pricing page via browser, 2026-08-13) · source
7 endpoints, cheapest first. This is the view the model-by-model tables cannot give you: what one provider actually offers, and at what precision.
| Model | Quant | $/M in | $/M out | Context | Uptime 30m |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | bf16 | $0.14 | $0.28 | 1,048,576 | 98.5% |
| Gemma 4 31B | fp4 | $0.14 | $0.40 | 175,000 | 96.5% |
| MiniMax-M3 | fp4 | $0.30 | $1.20 | 256,000 | 96.7% |
| Qwen3.6 27B | fp4 | $0.29 | $2.40 | 131,072 | 100.0% |
| GLM-5.2 | fp4 | $1.10 | $4.10 | 1,048,576 | 96.1% |
| Kimi K3 | fp4 | $2.80 | $14.00 | 1,048,576 | 99.9% |
| Kimi K3 | fp4 | $6.00 | $22.50 | 1,048,576 | — |