The column nobody else shows The same model is served by many providers at nearly the same price but at different numeric precision. Cheapest is often the most aggressively quantized — and sometimes higher precision costs the same or less.
bf16fp8fp4unknown
Buying direct from Anthropic costs $10.00 in and $50.00 out per million, cached input $1.000.
14 providers · bf16, fp4, fp8, mxfp4, unknown Buying direct from Moonshot costs $3.00 in and $15.00 out per million, cached input $0.300. That is 7% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Anthropic costs $5.00 in and $25.00 out per million, cached input $0.500.
Buying direct from Anthropic costs $5.00 in and $25.00 out per million, cached input $0.500.
Buying direct from xAI costs $2.00 in and $6.00 out per million, cached input $0.300.
4 providers · bf16, unknown Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Buying direct from Anthropic costs $2.00 in and $10.00 out per million, cached input $0.200.
32 providers · fp4, fp8, unknown Buying direct from Z.ai / Zhipu costs $1.40 in and $4.40 out per million, cached input $0.260. That is 186% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Alibaba / Qwen costs $2.50 in and $7.50 out per million, cached input $0.500. That is 69% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Anthropic costs $3.00 in and $15.00 out per million, cached input $0.300.
21 providers · bf16, fp4, fp8, int4, unknown Buying direct from Moonshot costs $0.95 in and $4.00 out per million, cached input $0.160. That is 69% dearer than the cheapest routed endpoint on a 3:1 blend.
15 providers · fp4, fp8, int4, unknown Buying direct from Moonshot costs $0.95 in and $4.00 out per million, cached input $0.190. That is 27% dearer than the cheapest routed endpoint on a 3:1 blend.
7 providers · bf16, fp8, unknown Buying direct from Xiaomi costs $0.43 in and $0.87 out per million, cached input $0.004.
2 providers · fp8, unknown Not listed above: Kwaipilot directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
7 providers · fp4, fp8, unknown Buying direct from DeepSeek costs $0.43 in and $0.87 out per million, cached input $0.004.
2 providers · fp8, unknown 6 providers · bf16, fp8, unknown Buying direct from Tencent costs $0.15 in and $0.59 out per million, cached input $0.037. That is 16% dearer than the cheapest routed endpoint on a 3:1 blend. Priced in RMB: 1 / 4 / 0.25 per million (input / output / cached), converted at about 6.8 CNY to the dollar.
12 providers · fp4, fp8, unknown Buying direct from MiniMax costs $0.30 in and $1.20 out per million, cached input $0.060. That is 27% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Xiaomi costs $0.14 in and $0.28 out per million, cached input $0.003.
28 providers · bf16, fp4, fp8, unknown Buying direct from DeepSeek costs $0.14 in and $0.28 out per million, cached input $0.003. That is 75% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Alibaba / Qwen costs $0.50 in and $3.00 out per million, cached input $0.050. That is 101% dearer than the cheapest routed endpoint on a 3:1 blend.
18 providers · fp4, fp8, unknown Buying direct from Z.ai / Zhipu costs $1.40 in and $4.40 out per million, cached input $0.260. That is 45% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Alibaba / Qwen costs $0.50 in and $3.00 out per million, cached input $0.050. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.
9 providers · fp4, fp8, unknown Buying direct from Alibaba / Qwen costs $0.60 in and $3.60 out per million. That is 86% dearer than the cheapest routed endpoint on a 3:1 blend.
11 providers · fp8, unknown Buying direct from MiniMax costs $0.30 in and $1.20 out per million, cached input $0.060. That is 25% dearer than the cheapest routed endpoint on a 3:1 blend.
3 providers · fp8, unknown Not listed above: Thinking Machines directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
11 providers · fp8, unknown Buying direct from Alibaba / Qwen costs $0.60 in and $3.60 out per million. That is 91% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Mistral costs $1.50 in and $7.50 out per million.
10 providers · fp4, fp8, int4, unknown Buying direct from Moonshot costs $0.60 in and $3.00 out per million, cached input $0.100. That is 52% dearer than the cheapest routed endpoint on a 3:1 blend.
5 providers · bf16, fp4, fp8 Buying direct from Z.ai / Zhipu costs $0.60 in and $2.20 out per million, cached input $0.110. That is 32% dearer than the cheapest routed endpoint on a 3:1 blend.
5 providers · bf16, fp4, fp8, unknown Buying direct from Alibaba / Qwen costs $0.40 in and $3.20 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.
Not listed above: Meituan directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
9 providers · fp16, fp4, fp8, unknown Buying direct from Z.ai / Zhipu costs $0.60 in and $2.20 out per million, cached input $0.110. That is 36% dearer than the cheapest routed endpoint on a 3:1 blend.
14 providers · fp4, fp8, unknown Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
5 providers · fp4, fp8, unknown Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
19 providers · bf16, fp16, fp4, fp8, unknown Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Not listed above: InclusionAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Buying direct from xAI costs $1.25 in and $2.50 out per million, cached input $0.200.
9 providers · fp8, unknown Buying direct from Alibaba / Qwen costs $0.25 in and $1.49 out per million. That is 79% dearer than the cheapest routed endpoint on a 3:1 blend.
Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
3 providers · fp8, unknown 8 providers · bf16, fp8, unknown 8 providers · fp8, unknown Buying direct from Alibaba / Qwen costs $0.25 in and $2.00 out per million. That is 94% dearer than the cheapest routed endpoint on a 3:1 blend.
4 providers · bf16, fp8, unknown 20 providers · bf16, fp16, fp4, fp8, unknown Not listed above: OpenAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
5 providers · bf16, fp8, unknown Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
2 providers · fp8, unknown Buying direct from Mistral costs $0.15 in and $0.60 out per million.
2 providers · fp4, unknown Not listed above: Arcee Ai directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
2 providers · bf16, unknown Not listed above: InclusionAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
3 providers · fp4, fp8, unknown Buying direct from DeepSeek costs $0.14 in and $0.28 out per million, cached input $0.003. That is 61% cheaper than the cheapest routed endpoint on a 3:1 blend.
Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
4 providers · bf16, fp4, fp8 Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
12 providers · bf16, fp4, fp8, unknown Not listed above: OpenAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
5 providers · fp8, unknown Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
4 providers · fp8, unknown Buying direct from Alibaba / Qwen costs $0.70 in and $2.80 out per million. That is 842% dearer than the cheapest routed endpoint on a 3:1 blend.
Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
3 providers · fp8, int4, unknown Buying direct from Alibaba / Qwen costs $0.35 in and $1.40 out per million. That is 371% dearer than the cheapest routed endpoint on a 3:1 blend.
Not listed above: IBM directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Buying direct from Alibaba / Qwen costs $0.18 in and $0.70 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.
4 providers · bf16, fp8, unknown Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Buying direct from xAI costs $1.00 in and $2.00 out per million, cached input $0.200.
4 providers · fp4, fp8, unknown Buying direct from Nvidia costs $0.50 in and $2.50 out per million, cached input $0.150. That is 8% dearer than the cheapest routed endpoint on a 3:1 blend.
Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
3 providers · bf16, fp4, unknown Buying direct from Nvidia costs $0.20 in and $0.80 out per million. That is 114% dearer than the cheapest routed endpoint on a 3:1 blend.
Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Buying direct from Cohere costs $0.00 in and $0.00 out per million.
Not listed above: Mistral directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Buying direct from Alibaba / Qwen costs $0.70 in and $2.80 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Mistral costs $0.50 in and $1.50 out per million. That is 75% cheaper than the cheapest routed endpoint on a 3:1 blend.
Buying direct from Nvidia costs $0.00 in and $0.00 out per million.
Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Buying direct from Nvidia costs $0.00 in and $0.00 out per million.
Buying direct from Mistral costs $0.10 in and $0.30 out per million. That is 41% dearer than the cheapest routed endpoint on a 3:1 blend.
Not listed above: Mistral directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
2 providers · fp8, unknown 5 providers · bf16, fp8, unknown Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.
Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.