Official site · 🇺🇸 United States
| Tier | $/month | Included usage |
|---|---|---|
| Free | $0.00 | $0.00 |
| Scale | $200.00 | $400.00 |
console only quota converts to units
As advertised. Free: 200 requests/month, then usage-based per-token with practically no rate limits. Scale: $200/month with 40M credits ($400) included. Dedicated = reserved B200 by the GPU-hour.
What that actually means. Now a pure usage-based inference host: pay per token, with a 200-request/month free allowance and an optional Scale flat-rate for steady volume. The old credit tiers (Free 250K / Starter $20 / Pro $60 / Scale $400) are GONE.
Where to check what is left. account dashboard
Models covered. Kimi K3 2.8T (morph-kimik3), Kimi K3 2.8T, latency-tuned (morph-kimik3-fast), GLM-5.3 744B (morph-glm53-744b), GLM-5.3-Flash (morph-glm53flash), DeepSeek V4 Flash 0731 (morph-dsv4flash), DeepSeek V4.1 Flash (morph-dsv41flash), plus its own Fast Apply / WarpGrep / Compact / Reflex / Model Router endpoints
checked 2026-10-02 · high (read from the rendered pricing page via a real browser over CDP, 2026-08-28) · source
6 endpoints, cheapest first. This is the view the model-by-model tables cannot give you: what one provider actually offers, and at what precision.
| Model | Quant | $/M in | $/M out | Context | Uptime 30m |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | fp8 | $0.02 | $0.38 | 1,048,576 | 99.7% |
| DeepSeek V4 Flash | bf16 | $0.14 | $0.40 | 1,048,576 | 100.0% |
| GLM-5.3-Flash | fp8 | $0.14 | $0.48 | 1,048,576 | 100.0% |
| GLM-5.3 | fp8 | $0.13 | $2.99 | 1,048,576 | 99.2% |
| GLM-5.2 | fp8 | $0.17 | $3.07 | 1,048,576 | 97.5% |
| Kimi K3 | fp8 | $2.50 | $11.36 | 1,048,576 | 93.4% |