Model Spend Arena2ND ED.
serves endpoints · 2026-08-14

OpenInference

origin not confirmed

Models it serves

2 endpoints, cheapest first. This is the view the model-by-model tables cannot give you: what one provider actually offers, and at what precision.

ModelQuant$/M in$/M outContextUptime 30m
DeepSeek V4 Flashfp4$0.08$0.18262,14496.2%
Gemma 4 31Bbf16$0.08$0.35262,14494.7%
◆ Some links on this page are referral links: if you sign up through them this site may earn a commission, at no extra cost to you. Rankings, prices and measurements are taken from the sources listed under Method, and are not influenced by whether a provider has a referral programme.