Model Spend Arena2ND ED.
serves endpoints · 2026-10-03

OpenInference

Official site · origin not confirmed

Models it serves

3 endpoints, cheapest first. This is the view the model-by-model tables cannot give you: what one provider actually offers, and at what precision.

ModelQuant$/M in$/M outContextUptime 30m
DeepSeek V4.1 Flashfp4$0.01$0.681,048,57694.4%
GLM-5.3-Flashfp4$0.03$0.931,048,57699.3%
DeepSeek V4 Flashfp8$0.00$1.041,048,57699.3%