Model Spend Arena2ND ED.
serves endpoints · 2026-10-03

Morph

Official site · 🇺🇸 United States

Plan

Usage-based (Free / Pay-per-token / Scale flat-rate)

Tier$/monthIncluded usage
Free$0.00$0.00
Scale$200.00$400.00

console only quota converts to units

As advertised. Free: 200 requests/month, then usage-based per-token with practically no rate limits. Scale: $200/month with 40M credits ($400) included. Dedicated = reserved B200 by the GPU-hour.

What that actually means. Now a pure usage-based inference host: pay per token, with a 200-request/month free allowance and an optional Scale flat-rate for steady volume. The old credit tiers (Free 250K / Starter $20 / Pro $60 / Scale $400) are GONE.

Where to check what is left. account dashboard

Models covered. Kimi K3 2.8T (morph-kimik3), Kimi K3 2.8T, latency-tuned (morph-kimik3-fast), GLM-5.3 744B (morph-glm53-744b), GLM-5.3-Flash (morph-glm53flash), DeepSeek V4 Flash 0731 (morph-dsv4flash), DeepSeek V4.1 Flash (morph-dsv41flash), plus its own Fast Apply / WarpGrep / Compact / Reflex / Model Router endpoints

Pure usage-based inference host, re-read 2026-09-11 through a headless browser over CDP (the site sits behind a Vercel bot challenge: WebFetch, curl and r.jina.ai all get 429). Free = 200 requests/month, then pay-per-token ("practically no rate limits"); PAYG top-up starts at $10 with +$5 free on the first top-up, and carries a 5.5% credit purchase fee ($0.80 minimum) that dedicated capacity does not. Hosted-model per-M tokens (input / cache-read / output), all 1M context: Kimi K3 $2.50/$0.29/$14.00; Kimi K3 fast $6.00/$0.60/$22.50; GLM-5.3 744B $1.19/$0.20/$3.74; GLM-5.3-Flash $0.20/$0.04/$0.70; DeepSeek V4 Flash 0731 $0.14/$0.036/$0.40; DeepSeek V4.1 Flash $0.15/$0.01/$0.60 (text and images, reasoning on by default). Specialized: Fast Apply $0.80/$1.20 per M in/out (morph-v3-fast, 262K ctx) or $0.90/$1.90 (morph-v3-large); WarpGrep, now morph-warp-grep-v2, $0.80/100K, 100K ctx (1M on Pro); Compact $0.20/$0.50 per M in/out; Model Router $0.005/request; Reflex $0.001/event realtime or $0.0005/event batch (both halve past 1M events/mo). Dedicated reserves B200 OR B300 by the GPU-hour, from $9.06 per reserved B200-hour behind a 99.9% monthly availability SLA, with a 90-day initial commitment; $8.64 / $8.43 / $8.16 are the 2x / 4x / 8x card rates. The dedicated catalog is seven models and wider than the chat one: it still carries Qwen 3.6 27B, Gemma 4 31B and MiniMax M3 428B, which left chat but not the product. Careful with one figure the vendor contradicts itself on: the dedicated calculator quotes DeepSeek V4 Flash 0731 at $0.14/$0.0028/$0.28, against the $0.14/$0.04/$0.40 of its own price table (2026-09-18). The table is the more credible of the two. 2026-10-02: producto nuevo, System One (morph-systemone-v1) a $0,042 de entrada y $0,01 cacheada por millon, 64K de contexto: 'decisiones tipadas, un estado entra y sale una distribucion calibrada por pregunta'.

checked 2026-10-02 · high (read from the rendered pricing page via a real browser over CDP, 2026-08-28) · source

Models it serves

6 endpoints, cheapest first. This is the view the model-by-model tables cannot give you: what one provider actually offers, and at what precision.

ModelQuant$/M in$/M outContextUptime 30m
DeepSeek V4.1 Flashfp8$0.02$0.381,048,57699.7%
DeepSeek V4 Flashbf16$0.14$0.401,048,576100.0%
GLM-5.3-Flashfp8$0.14$0.481,048,576100.0%
GLM-5.3fp8$0.13$2.991,048,57699.2%
GLM-5.2fp8$0.17$3.071,048,57697.5%
Kimi K3fp8$2.50$11.361,048,57693.4%