Model Spend Arena2ND ED.
serves endpoints · 2026-08-14

Morph

Official site · 🇺🇸 United States

Plan

Usage-based (Free / Pay-per-token / Scale flat-rate)

Tier$/monthIncluded usage
Free$0.00
Scale$200.00

console only quota converts to units

As advertised. Free: 200 requests/month, then usage-based per-token with practically no rate limits. Scale: flat-rate reserved capacity from ~$200/mo (~2x modelled token capacity, marketed up to ~6.6-7.0x cheaper than pay-per-token); Dedicated = reserved B200 by the GPU-hour.

What that actually means. Now a pure usage-based inference host: pay per token, with a 200-request/month free allowance and an optional Scale flat-rate for steady volume. The old credit tiers (Free 250K / Starter $20 / Pro $60 / Scale $400) are GONE.

Where to check what is left. account dashboard

Models covered. GLM-5.2 744B, Qwen 27B dense, MiniMax 230B MoE, DeepSeek (1M context), Kimi K3, plus its own Fast Apply / WarpGrep / Compact / Reflex / Model Router endpoints

RESTRUCTURED to usage-based, CONFIRMED 2026-08-13 off the rendered pricing page (fetched through a real browser — the SPA still 429s plain HTTP). Free = 200 requests/month; then pay-per-token, marketed with "practically no rate limits". Hosted-model per-M (input / cache-read / output): Kimi K3 $2.80 / $0.29 / $14.00; GLM-5.2 $6.00 / $0.60 / $22.50; a ~180 tok/s model $0.50 / $0.30 / $3.50; DeepSeek $0.30 / — / $1.20; a low-cost model $0.14 / $0.07 / $0.28. Specialized: Fast Apply $0.80/M in, $1.20/M out (10,500+ tok/s); WarpGrep $0.80/100K; Model Router $0.005/request; Reflex $0.001/event ($0.0005 over 1M/mo). A flat-rate "Scale" option (from ~$200/mo, ~40M tokens) and GPU-hour "Dedicated" B200 capacity sit alongside pay-per-token, marketed up to ~7x cheaper for steady load.

checked 2026-08-13 · high (read from the rendered pricing page via browser, 2026-08-13) · source

Models it serves

7 endpoints, cheapest first. This is the view the model-by-model tables cannot give you: what one provider actually offers, and at what precision.

ModelQuant$/M in$/M outContextUptime 30m
DeepSeek V4 Flashbf16$0.14$0.281,048,57698.5%
Gemma 4 31Bfp4$0.14$0.40175,00096.5%
MiniMax-M3fp4$0.30$1.20256,00096.7%
Qwen3.6 27Bfp4$0.29$2.40131,072100.0%
GLM-5.2fp4$1.10$4.101,048,57696.1%
Kimi K3fp4$2.80$14.001,048,57699.9%
Kimi K3fp4$6.00$22.501,048,576
◆ Some links on this page are referral links: if you sign up through them this site may earn a commission, at no extra cost to you. Rankings, prices and measurements are taken from the sources listed under Method, and are not influenced by whether a provider has a referral programme.