Model Spend Arena2ND ED.
serves endpoints · 2026-10-03

BaseTen

Official site · 🇺🇸 United States

Models it serves

19 endpoints, cheapest first. This is the view the model-by-model tables cannot give you: what one provider actually offers, and at what precision.

ModelQuant$/M in$/M outContextUptime 30m
DeepSeek V4 Flashfp8$0.13$0.261,048,57699.9%
DeepSeek V4 Flashfp8$0.13$0.261,048,576100.0%
gpt-oss-120bfp4$0.10$0.50128,072100.0%
gpt-oss-120bfp4$0.10$0.50128,072100.0%
GLM-5.3-Flashfp8$0.15$0.501,048,576100.0%
GLM-5.3-Flashfp8$0.15$0.501,048,576100.0%
DeepSeek V4.1 Flashfp8$0.30$1.201,048,57699.8%
DeepSeek V4.1 Flashfp8$0.30$1.201,048,57699.9%
Nemotron 3 Ultra 550Bfp4$0.60$2.40202,800100.0%
Nemotron 3 Ultra 550Bfp4$0.60$2.40202,800100.0%
GLM-5.2fp8$1.40$4.401,048,576100.0%
GLM-5.2fp8$1.40$4.401,048,57699.4%
GLM-5.3fp4$1.40$4.401,048,57699.9%
GLM-5.3fp4$1.40$4.401,048,57699.6%
GLM-5.2fp8$2.10$6.601,048,576100.0%
GLM-5.2fp8$2.10$6.601,048,576100.0%
GLM-5.3fp8$2.10$6.601,048,576100.0%
GLM-5.3fp8$2.10$6.601,048,57699.9%
Kimi K3fp8$3.00$15.001,048,57698.5%