Model Spend Arena2ND ED.
111 MODELS · 25 LABS · 44 HOSTS · 24 PLANS · 18 GATEWAYSUPDATED 2026-08-14
Coding-model intelligence · Updated 2026-08-14

The honest map of what LLM coding actually costs.

Metered per token or bought as a subscription — price, provider precision and published quality, aggregated, not re-benchmarked. Everything below is sourced and dated, including the column no other comparison shows: what numeric precision you are actually buying.

Best value · idx per $
1Ling-3.0-flash1606 idx/$
2Solar Pro 41004 idx/$
3gpt-oss-120b468 idx/$
Best value · top 20% (idx ≥ 69)
2GPT-5.6 Luna$0.22/M
Cheapest / session
Quality ceiling · coding idx
Cheapest open-weight · avg
Best plan deal · tok/$ @ $0.82/M

Editor’s picks

Hand-written · 8 on record

Written from the data alone; referral relationships carry no weight in what is picked or how it is described. Each pick keeps the date it was made and is not quietly rewritten later, so you can judge how well it aged.

Weekly notes

betathe run, in plain words

A short write-up each week on what changed and why — the same data as the tables, told as a story. Newest first.

2026-08-14

First weekly note: a free model that reasons, and prices that moved

I am starting a short note each week to explain what changed on the site, in plain words. Same data as the tables, just told as a story. This is the first one.

A new free model I could actually test. NVIDIA released Nemotron 3.5 Lightning this week, a 30B mixture-of-experts model with only 3B active parameters. The small active count is the interesting part: it means the model is cheap to run and it fits on modest hardware. It is also already free on opencode Zen, so I did not need to rent a GPU. I pointed our BigCodeBench-Hard harness at the free endpoint and measured it myself: 34.7% on BCB-Hard (42 of 121). It is a reasoning model, so it thinks before it answers and it is not fast, but for a free local-agent model that is a solid number. It is now both on the leaderboard (Artificial Analysis gives it a coding index of 26.8) and in the Run local catalogue with our own measurement.

Qwen3.8-27B landed, and I ran it. The Max is still datacenter-only (2.4T, 95B active), but the small dense 27B (Apache-2.0) that fits at home shipped this afternoon, so I tested it properly. On my 16 GB card it fits at Q3 (about 14 GB), and with MTP speculative decoding through llama-server it runs at 36 tok/s, twice the 18 it manages without spec-decode, because the model ships its own draft head and about 87% of the drafts get accepted. On our BCB-Hard it scores 27% at Q3 non-thinking; the full model at good precision (BF16/FP8) is around 37 to 43%, so the 3-bit quant you need to fit 16 GB costs about 10 points. It is not on the leaderboard yet, no Artificial Analysis index and not on OpenRouter, so for now it lives in Run local with our own numbers. It made me want a bigger GPU, but on the 16 GB card I actually have it does not beat what I already run, so I have not switched, the honest take is in the Run local pick. And it is dense, so if you cannot buy a card the one to wait for is the small A3B MoE Qwen will ship next, which runs fast on the hardware you already own.

The ceiling kept climbing. Grok 4.6 joined the board at a 76.8 coding index, near the top of everything tracked. Solar Pro 4 (52.7), Muse Glimmer 30B (49) and Nemotron 3.5 Lightning also went on. Qwen3.8-Max was already there at 71.8. DeepSeek also shipped a new revision of V4 Pro (0813) that jumped its coding index from 59.4 to 68.8, a big step for an already cheap open model.

GLM-5.3 dropped today too. Zhipu launched it on August 14, "built to code", and it looks strong: they claim +50% coding over GLM-5.2 and first place among open-source models on Terminal-Bench 3.0 (DeepSWE 1.1 66.9, CyberGym 84.5). The Z.ai coding plan already serves it, and GLM-5.2 and 5.1 requests now route to 5.3. But it is not on OpenRouter or the Artificial Analysis index yet, and the open weights are about two weeks out, so it does not go on the leaderboard yet. That is the rule here: a model only joins the board once it has a public coding index and a place to buy it. I will add it the moment it lands on both.

Prices moved in a few places. OpenAI turned on unlimited text with GPT-5.6 Luna for the Free and Go tiers, it went live this week. Note it is text only and the paid tiers still keep caps on Luna, the earlier claim that Luna was unlimited on Plus and Pro was wrong. Synthetic's $7 first-month promo ended, it is back to $30 a month. Mistral folded its old $5.99 Education tier into a student discount on the $14.99 Pro plan. Alibaba now shows a $2 new-user discount across its Individual token tiers.

DeepSeek is about to get more expensive. V4 Pro 0813 shipped at the old flat rate, but from August 16 (16:00 UTC) DeepSeek switches to peak/off-peak billing and the jumps are steep: cache-hit input rises up to 12x at peak, cache-miss up to 3x, output up to 4.5x, with peak on Beijing working hours (01:00-04:00 and 06:00-10:00 UTC). The cache-hit line is the one that stings, because a coding agent replays a big cached prefix every turn, so the effective bill tracks that 12x, not the headline flat price. If you lean on DeepSeek, front-load the heavy runs before the 16th, or push them to off-peak after.

Two corrections worth calling out. The "Cerebras free tier sunsets on Aug 17" note that went around is a false alarm: that date only kills preview models on the paid Developer tier, the free trial is not going away. And Morph quietly rebuilt its whole pricing to pure pay-per-token with 200 free requests a month, the old Free/Starter/Pro/Scale credit tiers are gone. Their page bot-walls plain requests, so I had to open it in a real browser to read it, now confirmed.

New this week. The free set on opencode Zen rotated: it added Hy3 and Nemotron 3.5 Lightning and dropped Ling-3.0-flash, LongCat-2.0 and North Mini Code. I added RunPod to the gateways, it has a real OpenAI-compatible inference API through its Public Endpoints, though the text catalogue is thin and it is mostly image and video. And I added this notes tab.

As always, free sets rotate and close without notice, so treat any free allowance as this week's, not a standing plan.

And one last thing: next week we bring a surprise. Stay tuned.

Leaderboard

Ranked by coding index. Value = coding index per $ (blended, 3:1 in:out).
#ModelCoding indexIntelligenceTBench-HLong-ctxtok/s$/M blendedValue idx/$Context
1GPT-5.6 Sol · OpenAI78.359.00.6140.76357.0$11.25071,050,000
2Claude Opus 5 · Anthropic78.063.10.75745.1$10.00081,000,000
3Grok 4.6 · xAI76.860.90.75052.1$3.00026500,000
4GPT-5.6 Terra · OpenAI76.756.60.5760.79789.8$2.250341,050,000
5Claude Fable 5 · Anthropic76.562.10.6290.76756.2$20.00041,000,000
6Kimi K3 · Moonshot76.259.70.82736.5$6.000131,048,576
7GPT-5.5 · OpenAI74.956.30.6060.7900.0$11.25071,050,000
8Claude Opus 4.8 · Anthropic74.357.30.5830.7300.0$10.00071,000,000
9Claude Opus 4.7 · Anthropic73.655.00.5150.7530.0$10.00071,000,000
10Grok 4.5 · xAI72.455.80.74050.9$3.00024500,000
11Muse Spark 1.2 · Meta72.256.80.8330.0$2.000361,048,576
12Qwen3.8 Max · Alibaba / Qwen71.858.10.74343.3$3.000241,000,000
13Claude Sonnet 5 · Anthropic71.555.30.77072.4$4.000181,000,000
14GPT-5.6 Luna · OpenAI71.452.30.783139.0$0.2253171,050,000
15Muse Spark 1.1 · Meta71.353.20.8130.0$2.000361,048,576
16GPT-5.4 · OpenAI71.153.10.5760.7770.0$5.625131,050,000
17Gemini 3.5 Flash · Google70.152.00.4090.8100.0$3.375211,048,576
18Gemini 3.6 Flash · Google69.251.60.790199.7$1.500461,048,576
19DeepSeek V4 Flash · DeepSeek69.151.80.743101.1$0.1753951,048,576
20GLM-5.2 · Z.ai / Zhipu68.852.60.5080.767104.3$1.828381,048,576
21Gemini 3.1 Pro Preview · Google68.847.70.5380.790128.4$4.500151,048,576
22DeepSeek V4 Pro · DeepSeek68.853.20.75366.6$0.5441271,048,576
23Qwen3.7 Max · Alibaba / Qwen66.046.70.5080.7470.0$2.213301,000,000
24Claude Sonnet 4.6 · Anthropic63.048.40.5300.7400.0$6.000101,000,000
25Kimi K2.6 · Moonshot61.845.10.4390.7670.0$1.01061262,144
26Kimi K2.7 Code · Moonshot60.843.00.4470.75036.8$1.40743262,144
27Xiaomi MiMo-V2.5-Pro · Xiaomi60.242.90.4320.77747.7$0.5441111,050,000
28KAT-Coder-Pro V2 · Kwaipilot59.534.50.4920.700109.7$0.525113262,144
29Nex-N2-Pro · Nex AGI59.141.70.763142.2$0.438135262,144
30Tencent Hy3 · Tencent58.842.20.74766.5$0.231255262,144
31MiniMax-M3 · MiniMax58.645.40.4240.80370.0$0.5251121,048,576
32Xiaomi MiMo-V2.5 · Xiaomi56.838.00.4170.68390.0$0.1753251,050,000
33GPT-5.4 Nano · OpenAI56.139.70.4240.7200.0$0.463121400,000
34GPT-5.4 Mini · OpenAI56.140.90.5230.7300.0$1.68833400,000
35Qwen3.7 Plus · Alibaba / Qwen55.939.40.4700.69056.3$0.5601001,000,000
36GLM-5.1 · Z.ai / Zhipu55.841.00.4320.6800.0$2.15026204,800
37Qwen3.6 Plus · Alibaba / Qwen54.540.50.4390.7230.0$0.731751,000,000
38Qwen3.6 27B · Alibaba / Qwen53.737.70.3480.73353.3$1.35040262,144
39Solar Pro 4 · Upstage52.741.60.70742.0$0.0521004524,288
40MiniMax M2.7 · MiniMax52.638.90.3940.7530.0$0.525100204,800
41Inkling · Thinking Machines52.142.30.73381.8$1.725301,048,576
42Grok Build 0.1 · xAI51.540.70.7000.0$1.25041256,000
43Ling-3.0-flash · InclusionAI50.637.80.670380.4$0.0321606262,144
44GPT-5.1 · OpenAI49.437.50.4550.7670.0$3.43814400,000
45Gemini 3.5 Flash-Lite · Google49.337.40.747305.6$0.850581,048,576
46Nemotron 3 Ultra 550B · Nvidia49.338.30.3640.710132.0$1.35037512,288
47Muse Glimmer · Meta49.035.10.80090.2$0.63777131,072
48Qwen3.5 397B A17B · Alibaba / Qwen48.234.30.4090.72780.3$0.87755262,144
49Mistral Medium 3.5 · Mistral46.930.40.3330.653122.2$3.00016262,144
50Kimi K2.5 · Moonshot46.836.00.3480.7300.0$1.14041262,144
51GLM 4.6 · Z.ai / Zhipu45.829.30.2500.5530.0$0.96348204,800
52Qwen3.5-122B-A10B · Alibaba / Qwen45.732.80.3110.703132.8$0.81756262,144
53LongCat 2.0 · Meituan45.334.00.62735.4$0.525861,048,756
54GLM 4.7 · Z.ai / Zhipu45.334.50.3180.6800.0$0.73861204,800
55DeepSeek V3.2 · DeepSeek44.232.80.3560.7070.0$0.302146163,840
56DeepSeek V3.1 Terminus · DeepSeek43.531.10.3030.6800.0$0.44099163,840
57Gemma 4 31B · Google43.429.70.3640.68335.7$0.160271262,144
58Ring-2.6-1T · InclusionAI42.831.70.2880.673121.3$0.212201262,144
59Grok 4.3 · xAI42.237.90.3790.6600.0$1.562271,000,000
60Qwen3.6 35B A3B · Alibaba / Qwen41.932.10.3480.667138.3$0.362116262,144
61o1 · OpenAI39.723.90.1290.6330.0$26.2502200,000
62Step 3.7 Flash · StepFun39.630.90.3560.697357.3$0.43891262,144
63Gemma 4 26B A4B · Google39.326.10.1360.6170.0$0.190207262,144
64GPT-5 · OpenAI37.835.30.3260.7630.0$3.43811400,000
65Qwen3.5-35B-A3B · Alibaba / Qwen37.024.30.1060.603157.1$0.61960262,144
66North Mini Code · Cohere36.520.20.3110.36026.5$0.000256,000
67Qwen3 Coder Next · Alibaba / Qwen36.221.30.1820.423134.3$0.290125262,144
68Gemini 3.1 Flash Lite · Google34.725.60.2420.7130.0$0.562621,048,576
69Gemini 2.5 Pro · Google33.325.90.2650.6600.0$3.438101,048,576
70Devstral 2 · Mistral31.319.20.1890.31344.4
71Mercury 2 · Inception31.121.90.2650.407761.5$0.37583128,000
72gpt-oss-120b · OpenAI30.424.10.2350.510163.3$0.065468131,072
73Qwen3.5-9B · Alibaba / Qwen28.721.80.2420.65384.6$0.112255262,144
74Command A · Cohere27.822.80.2500.487192.5$4.3756256,000
75Nemotron 3.5 Lightning · Nvidia26.823.60.553287.6$0.1381951,000,000
76Mistral Small 4 · Mistral26.619.70.1740.473144.2$0.262101262,144
77Ling 3.0 Tiny · InclusionAI26.524.50.587228.2
78Mistral Small 3.1 · Mistral26.314.90.0760.2200.0$0.40265128,000
79Trinity Large Thinking · Arcee Ai25.818.70.2270.383212.3$0.37868262,144
80DeepSeek R1 · DeepSeek24.618.60.0610.5600.0$1.1502164,000
81GPT-4o · OpenAI24.28.40.0$4.3756128,000
82DeepSeek V3 · DeepSeek23.014.20.0680.3170.0$0.45051163,840
83Qwen3 235B A22B 2507 · Alibaba / Qwen22.119.90.1360.7070.0$0.79628131,072
84GPT-4 Turbo · OpenAI21.57.70.0$15.0001128,000
85DeepSeek V3 0324 · DeepSeek21.215.20.1520.4130.0$0.48344163,840
86gpt-oss-20b · OpenAI20.715.20.1060.333125.6$0.055376131,072
87Mistral Medium 3.1 · Mistral20.514.70.1060.2130.0$0.80026131,072
88GPT-4.1 Mini · OpenAI20.214.80.0760.4530.0$0.700291,047,576
89Mistral Large 3 · Mistral20.115.90.1590.34756.5$3.0007128,000
90Llama 4 Maverick · Meta16.314.50.0680.500112.9$0.350471,048,576
91o3 Mini · OpenAI16.315.70.0610.4200.0$1.9258200,000
92Solar Pro 3 · Upstage16.214.50.0760.310152.8$0.26262131,072
93GPT-5 Mini · OpenAI15.625.80.3330.7100.0$0.68823400,000
94Qwen3 32B · Alibaba / Qwen15.311.40.0300.0000.0$0.130118131,072
95Nemotron 3 Nano 30B · Nvidia14.414.50.1360.373253.6$0.087165262,144
96Qwen3 14B · Alibaba / Qwen13.810.40.0380.0000.0$0.15092131,072
97Nemotron 3 Nano Omni 30B · Nvidia13.815.00.0830.407328.6$0.000256,000
98GPT-4 · OpenAI13.16.80.0$37.50008,191
99Mistral Small 3.2 · Mistral12.510.70.0680.1870.0$0.13394256,000
100Qwen3 30B A3B 2507 · Alibaba / Qwen12.114.60.0530.5970.0$0.21556131,072
101GPT-4o-mini · OpenAI11.46.70.0$0.26243128,000
102GPT-4.1 Nano · OpenAI11.19.60.0380.1930.0$0.175631,047,576
103GPT-3.5 Turbo · OpenAI10.73.20.0$0.7501416,385
104Gemma 3 27B · Google10.17.40.0380.0630.0$0.17259262,144
105Granite 4.1 8B · IBM9.56.40.0000.113128.6$0.062152131,072
106Qwen3 8B · Alibaba / Qwen9.08.30.0230.0000.0$0.20245131,072
107Llama 4 Scout · Meta8.210.30.0150.303131.1$0.150551,310,720
108Gemma 3 12B · Google5.85.50.0080.0800.0$0.07577131,072
109Gemma 3n E4B · Google3.21.00.0230.0000.0$0.0754332,768
110Gemma 3 4B · Google2.71.00.0080.0670.0$0.06243131,072
Nemotron 3 Super 120B · Nvidia$0.1641,000,000

Quality and speed are third-party measurements from Artificial Analysis — the only source publishing measured speed. A dash means the model has not been independently measured, not that it scored zero. Intelligence is AA’s general index, shown next to the coding one because they diverge: a model can code well and reason poorly, or the reverse, and only the coding column decides the ranking here. Long-ctx is AA’s long-context reasoning score — the column that matters on a large codebase, since Context only says how many tokens fit, not whether the model still reasons at that length.

Published quality

Mind the provenance

The flattering coding scores are vendor-reported. Where an independent lab measured the same benchmark the numbers agree closely — so the gap is about which benchmarks get published, not fabricated values.

ModelBenchmarkScoreProvenanceSource
MiniMax-M3SWE-bench Verified80.5%vendorminimax.io
MiniMax-M3SWE-bench Pro59.0%vendorminimax.io
MiniMax-M3Terminal-Bench 2.166.0%vendorminimax.io
MiniMax-M3Terminal-Bench 2.165.2%independentArtificial Analysis
MiniMax-M3TerminalBench-Hard42.4%independentArtificial Analysis
MiniMax-M3Aider polyglotnot listedindependentaider.chat
Xiaomi MiMo-V2.5-ProSWE-bench Verified78.9%vendormimo.xiaomi.com
Mistral Devstral Small 2SWE-bench Verified68.0%vendormistral.ai

Models the coding index can’t rank

Measured here, not by AA

The leaderboard ranks on Artificial Analysis’ coding index — a proprietary composite AA publishes only for models it chooses to test. Some models here have no such score and cannot get one from us: the index is not a benchmark we can run. So they were measured on tests we can run, and placed next to models that appear on both scales.

BigCodeBench-Hard saturation - the SAME 148 problems, every model comparable

ModelScoreAA coding index
Granite 4.1 8B17%9.5
Mistral Small 3.225%12.5
Ministral 3 14B14%14.4
Nemotron Nano 9B V222%not measured
gpt-oss-20b26%20.7
North Mini Code35%36.5
Nemotron 3 Super 120B38%37.7
Gemma 4 26B A4B36%39.3
Nemotron 3 Ultra 550B44%49.3
DeepSeek V4 Flash33%56.2
GLM-5.239%68.8

BigCodeBench-Hard (run here)

ModelScoreAA coding index
MiniMax40.0%58.6
Agnes AI35.0%not measured

SWE-bench Verified (run here, 25 instances, 80-step budget)

ModelScoreAA coding index
Tencent68% (17/25)58.8
Agnes AI52% (13/25)not measured

The saturation table is the honest shape of our own BigCodeBench-Hard, run on ONE fixed set of 148 problems so every model is directly comparable (an earlier pass compared different problem subsets and was invalid). It climbs with the coding index up to about index 49 (~44%), then FLATTENS: DeepSeek V4 Flash (index 56) and GLM-5.2 (index 69) land at 33-39%, no better than a 550B Nemotron at 49. It does NOT rise to 100% - the ceiling is ~40-44%, because 36% of these problems (44 of 121 that all models attempt) are passed by NOBODY: 38 are genuine wrong answers, only 6 are offline-sandbox errors. So the ceiling is real (hard problems + strict tests that reject valid alternative solutions), not a grading artefact. The lesson for this site: our BCB-Hard discriminates coders up to ~index 45; above that it saturates and the AA coding index is the better ranker - which is exactly how the local picks are ordered. Grades are offline (--network none), which caps everyone equally, so read absolute scores as a floor. Below, older single-benchmark brackets (SWE-bench Verified, 25 instances) - read as a bracket, not a coordinate: across those gold-gated instances Tencent Hy3 (index 58.8) resolved 68% and Agnes 52%, but most of that gap is our test - Agnes failed to converge in the 80-step budget on 9 of 25 vs 1 for Hy3, and 2 of 4 re-run at 200 steps then resolved. These stay off the leaderboard and never become an index.

Providers & quantization

The column nobody else shows

The same model is served by many providers at nearly the same price but at different numeric precision. Cheapest is often the most aggressively quantized — and sometimes higher precision costs the same or less.

bf16fp8fp4unknown

GPT-5.6 Sol

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$2.50$15.001,050,00099.6%
Azureunknown$5.00$30.001,050,00099.9%
OpenAIunknown$5.00$30.001,050,00099.6%
Amazon Bedrockunknown$5.50$33.001,050,000
Azureunknown$5.50$33.001,050,000100.0%
Azureunknown$5.50$33.001,050,000
OpenAIunknown$10.00$60.001,050,00099.6%

Claude Opus 5

10 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Amazon Bedrockunknown$5.00$25.001,000,0003.7%
Amazon Bedrockunknown$5.00$25.001,000,00099.4%
Anthropicunknown$5.00$25.001,000,00098.5%
Azureunknown$5.00$25.001,000,000100.0%
Azureunknown$5.00$25.001,000,000
Claude Platform on AWSunknown$5.00$25.001,000,000
Googleunknown$5.00$25.001,000,000100.0%
Amazon Bedrockunknown$5.50$27.501,000,000100.0%
Googleunknown$5.50$27.501,000,000100.0%
Googleunknown$5.50$27.501,000,000

GPT-5.6 Terra

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.50$3.001,050,00099.9%
OpenAIunknown$1.00$6.001,050,00099.9%
Azureunknown$2.00$12.001,050,00099.9%
OpenAIunknown$2.00$12.001,050,00099.9%
Amazon Bedrockunknown$2.20$13.201,050,000
Azureunknown$2.20$13.201,050,000100.0%
Azureunknown$2.20$13.201,050,000

Claude Fable 5

6 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$10.00$50.00
Amazon Bedrockunknown$10.00$50.001,000,000
Amazon Bedrockunknown$10.00$50.001,000,00099.6%
Anthropicunknown$10.00$50.001,000,000100.0%
Azureunknown$10.00$50.001,000,000
Googleunknown$10.00$50.001,000,000100.0%
Googleunknown$11.00$55.001,000,000

Buying direct from Anthropic costs $10.00 in and $50.00 out per million, cached input $1.000.

Kimi K3

14 providers · bf16, fp4, fp8, mxfp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Moonshot directfirst-party$3.00$15.00
Morphfp4$2.80$14.001,048,57699.9%
DeepInfrabf16$2.85$14.251,048,576
DigitalOceanunknown$2.85$14.251,048,57699.8%
BaseTenfp8$3.00$15.001,048,57699.0%
Chutesmxfp4$3.00$15.001,048,57698.1%
Fireworksunknown$3.00$15.001,048,57699.6%
Modalmxfp4$3.00$15.001,048,57698.9%
Moonshotmxfp4$3.00$15.001,048,576100.0%
Phalaunknown$3.00$15.001,048,57699.4%
Sail Researchfp4$3.00$15.00974,84299.6%
Togetherunknown$3.00$15.001,000,00099.1%
Waferunknown$3.00$15.001,048,57698.4%
Fireworksunknown$4.50$22.501,048,57699.6%
Morphfp4$6.00$22.501,048,576

Buying direct from Moonshot costs $3.00 in and $15.00 out per million, cached input $0.300. That is 7% dearer than the cheapest routed endpoint on a 3:1 blend.

GPT-5.5

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$2.50$15.001,050,00099.9%
Azureunknown$5.00$30.001,050,00099.9%
OpenAIunknown$5.00$30.001,050,00099.9%
Amazon Bedrockunknown$5.50$33.001,050,000
Azureunknown$5.50$33.001,050,000
Azureunknown$5.50$33.001,050,000100.0%
OpenAIunknown$12.50$75.001,050,00099.9%

Claude Opus 4.8

10 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$5.00$25.00
Amazon Bedrockunknown$5.00$25.001,000,000100.0%
Anthropicunknown$5.00$25.001,000,00099.7%
Anthropicunknown$5.00$25.001,000,00099.9%
Azureunknown$5.00$25.001,000,000
Azureunknown$5.00$25.001,000,000
Googleunknown$5.00$25.001,000,000100.0%
Amazon Bedrockunknown$5.50$27.501,000,000
Amazon Bedrockunknown$5.50$27.501,000,000
Googleunknown$5.50$27.501,000,000
Googleunknown$5.50$27.501,000,000

Buying direct from Anthropic costs $5.00 in and $25.00 out per million, cached input $0.500.

Claude Opus 4.7

8 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$5.00$25.00
Amazon Bedrockunknown$5.00$25.001,000,000100.0%
Anthropicunknown$5.00$25.001,000,00097.2%
Anthropicunknown$5.00$25.001,000,00095.6%
Azureunknown$5.00$25.001,000,000
Googleunknown$5.00$25.001,000,000100.0%
Amazon Bedrockunknown$5.50$27.501,000,000
Googleunknown$5.50$27.501,000,000
Googleunknown$5.50$27.501,000,000

Buying direct from Anthropic costs $5.00 in and $25.00 out per million, cached input $0.500.

Grok 4.5

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAI directfirst-party$2.00$6.00
xAIunknown$2.00$6.00500,000100.0%
xAIunknown$2.00$6.00500,00099.8%
xAIunknown$4.00$12.00500,000100.0%
xAIunknown$4.00$12.00500,00099.8%

Buying direct from xAI costs $2.00 in and $6.00 out per million, cached input $0.300.

Grok 4.6

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAIunknown$2.00$6.00500,00099.7%
xAIunknown$2.00$6.00500,00099.8%
xAIunknown$4.00$12.00500,00099.7%
xAIunknown$4.00$12.00500,00099.8%

Muse Glimmer

4 providers · bf16, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.30$1.20131,072100.0%
Fireworksunknown$0.35$1.50131,072100.0%
Phalaunknown$0.35$1.50131,07297.2%
Togetherunknown$0.35$1.50131,07299.7%

Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Solar Pro 4

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Upstageunknown$0.03$0.12524,288100.0%

Nemotron 3.5 Lightning

3 providers · bf16, fp4
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.08$0.20262,14499.6%
CoreWeavebf16$0.10$0.25262,144100.0%
Venicefp4$0.10$0.251,000,000100.0%

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Claude Sonnet 5

9 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$2.00$10.00
Amazon Bedrockunknown$2.00$10.001,000,000100.0%
Amazon Bedrockunknown$2.00$10.001,000,000100.0%
Anthropicunknown$2.00$10.001,000,000100.0%
Azureunknown$2.00$10.001,000,000100.0%
Azureunknown$2.00$10.001,000,000
Googleunknown$2.00$10.001,000,000100.0%
Amazon Bedrockunknown$2.20$11.001,000,000
Googleunknown$2.20$11.001,000,000100.0%
Googleunknown$2.20$11.001,000,000

Buying direct from Anthropic costs $2.00 in and $10.00 out per million, cached input $0.200.

GPT-5.6 Luna

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.05$0.301,050,00099.9%
OpenAIunknown$0.10$0.601,050,00099.9%
Azureunknown$0.20$1.201,050,000100.0%
OpenAIunknown$0.20$1.201,050,00099.9%
Amazon Bedrockunknown$0.22$1.321,050,000100.0%
Azureunknown$0.22$1.321,050,000100.0%
Azureunknown$0.22$1.321,050,000100.0%

Muse Spark 1.1

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Metaunknown$1.25$4.251,048,576100.0%

GPT-5.4

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.25$7.501,050,00076.0%
Azureunknown$2.50$15.001,050,000100.0%
OpenAIunknown$2.50$15.001,050,00076.0%
Amazon Bedrockunknown$2.75$16.501,050,000
Azureunknown$2.75$16.501,050,000100.0%
Azureunknown$2.75$16.501,050,000100.0%
OpenAIunknown$5.00$30.001,050,00076.0%

Gemini 3.5 Flash

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.75$4.501,048,57699.9%
Google AI Studiounknown$0.75$4.501,048,57699.7%
Googleunknown$1.50$9.001,048,57699.9%
Google AI Studiounknown$1.50$9.001,048,57699.7%
Googleunknown$1.65$9.901,048,576
Googleunknown$2.70$16.201,048,57699.9%
Google AI Studiounknown$2.70$16.201,048,57699.7%

Gemini 3.6 Flash

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.38$1.881,048,57699.8%
Google AI Studiounknown$0.38$1.881,048,57699.7%
Googleunknown$0.75$3.751,048,57699.8%
Google AI Studiounknown$0.75$3.751,048,57699.7%
Googleunknown$0.83$4.121,048,576
Googleunknown$1.35$6.751,048,57699.8%
Google AI Studiounknown$1.35$6.751,048,57699.7%

GLM-5.2

32 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$1.40$4.40
Baidufp8$0.49$1.541,048,576100.0%
DigitalOceanunknown$0.63$1.98262,144100.0%
Novitafp8$0.74$2.331,048,57699.3%
GMICloudfp8$0.74$2.331,048,57698.9%
StreamLakefp8$0.75$2.371,024,00099.7%
DeepInfrafp4$0.75$2.401,048,57699.6%
Sail Researchfp8$0.50$3.151,048,57699.6%
CoreWeavefp4$0.76$2.42262,14499.9%
AkashMLfp8$0.77$2.4296,89096.7%
Decartfp4$0.77$2.561,048,57699.9%
Inceptronfp4$0.75$2.901,048,57698.8%
Alibaba / Qwenfp8$0.97$3.041,048,576100.0%
Phalafp8$1.13$3.001,048,57699.7%
SiliconFlowfp8$1.19$3.741,048,576100.0%
Morphfp4$1.10$4.101,048,57696.1%
Ambientfp8$1.05$4.40202,75298.5%
AtlasCloudfp8$1.26$3.961,048,576100.0%
Waferfp4$1.26$3.961,048,57698.5%
BaseTenfp8$1.40$4.401,048,576100.0%
Cloudflareunknown$1.40$4.40262,144
Crusoefp8$1.40$4.401,048,57697.6%
Fireworksunknown$1.40$4.401,048,57698.7%
Friendliunknown$1.40$4.401,048,57698.6%
Parasailfp4$1.40$4.40262,14499.9%
Togetherunknown$1.40$4.40512,00099.1%
Venicefp8$1.40$4.401,000,00098.9%
Z.ai / Zhipufp8$1.40$4.401,048,57699.9%
BaseTenfp8$2.10$6.601,048,576100.0%
Cloudflareunknown$2.10$6.60262,144100.0%
Fireworksunknown$2.10$6.601,048,57698.5%
Waferfp4$2.10$6.601,048,57699.6%
Alibaba / Qwenfp8$2.31$7.261,048,576100.0%

Buying direct from Z.ai / Zhipu costs $1.40 in and $4.40 out per million, cached input $0.260. That is 186% dearer than the cheapest routed endpoint on a 3:1 blend.

Gemini 3.1 Pro Preview

6 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$1.00$6.001,048,57697.2%
Google AI Studiounknown$1.00$6.001,048,57699.8%
Googleunknown$2.00$12.001,048,57697.2%
Google AI Studiounknown$2.00$12.001,048,57699.8%
Googleunknown$3.60$21.601,048,57697.2%
Google AI Studiounknown$3.60$21.601,048,57699.8%

Qwen3.7 Max

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$2.50$7.50
Alibaba / Qwenunknown$1.48$4.421,000,000100.0%

Buying direct from Alibaba / Qwen costs $2.50 in and $7.50 out per million, cached input $0.500. That is 69% dearer than the cheapest routed endpoint on a 3:1 blend.

Claude Sonnet 4.6

9 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$3.00$15.00
Amazon Bedrockunknown$3.00$15.001,000,00099.9%
Anthropicunknown$3.00$15.001,000,000100.0%
Anthropicunknown$3.00$15.001,000,000
Azureunknown$3.00$15.001,000,000
Googleunknown$3.00$15.001,000,00099.9%
Amazon Bedrockunknown$3.30$16.501,000,000
Amazon Bedrockunknown$3.30$16.501,000,000
Googleunknown$3.30$16.501,000,000
Googleunknown$3.30$16.501,000,000

Buying direct from Anthropic costs $3.00 in and $15.00 out per million, cached input $0.300.

Kimi K2.6

21 providers · bf16, fp4, fp8, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Moonshot directfirst-party$0.95$4.00
Baidufp4$0.56$2.36262,144100.0%
Decartfp4$0.57$2.39262,14499.6%
StreamLakefp8$0.60$2.52256,00098.5%
Chutesint4$0.58$3.40262,14499.6%
Inceptronint4$0.60$3.41262,144100.0%
CoreWeavefp4$0.65$3.41262,14499.9%
DigitalOceanunknown$0.76$3.20262,14499.6%
Crusoebf16$0.70$3.50262,14488.2%
SiliconFlowfp8$0.77$3.40262,14499.7%
DeepInfrafp4$0.75$3.50262,144100.0%
Parasailint4$0.75$3.50262,14499.6%
Veniceint4$0.75$3.50256,00081.4%
Novitaunknown$0.80$3.40262,14499.9%
AtlasCloudint4$0.95$4.00262,14491.1%
BaseTenfp4$0.95$4.00262,000
Cloudflareunknown$0.95$4.00262,144100.0%
Fireworksunknown$0.95$4.00262,144
Moonshotint4$0.95$4.00262,144100.0%
Sail Researchint4$1.00$4.00262,144100.0%
Phalaunknown$1.09$4.60262,14496.3%
Togetherunknown$1.20$4.50262,14495.0%

Buying direct from Moonshot costs $0.95 in and $4.00 out per million, cached input $0.160. That is 69% dearer than the cheapest routed endpoint on a 3:1 blend.

Kimi K2.7 Code

15 providers · fp4, fp8, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Moonshot directfirst-party$0.95$4.00
Inceptronint4$0.67$3.40262,14499.9%
DeepInfrafp4$0.68$3.40262,144100.0%
Ambientunknown$0.69$3.49262,14499.7%
CoreWeaveint4$0.71$3.50262,144100.0%
Veniceint4$0.75$3.50256,000100.0%
Parasailint4$0.76$3.50262,14499.8%
ModelRunfp4$0.85$3.75262,14499.7%
SiliconFlowfp8$0.86$3.80262,144100.0%
Novitaint4$0.91$3.84262,144100.0%
Alibaba / Qwenfp8$0.95$4.00262,144100.0%
AtlasCloudint4$0.95$4.00262,144
Cloudflareunknown$0.95$4.00262,144
Moonshotint4$0.95$4.00262,144100.0%
Togetherunknown$0.95$4.00262,14498.7%
Moonshotint4$1.90$8.00262,144100.0%

Buying direct from Moonshot costs $0.95 in and $4.00 out per million, cached input $0.190. That is 27% dearer than the cheapest routed endpoint on a 3:1 blend.

Xiaomi MiMo-V2.5-Pro

7 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Xiaomi directfirst-party$0.43$0.87
AtlasCloudfp8$0.43$0.871,024,000100.0%
GMICloudbf16$0.43$0.871,050,00098.2%
Xiaomifp8$0.43$0.871,048,576100.0%
Novitaunknown$0.48$0.961,048,576100.0%
StreamLakeunknown$0.52$1.041,000,000
DigitalOceanunknown$0.40$1.50262,144100.0%
DeepInfrafp8$1.00$3.001,048,576100.0%

Buying direct from Xiaomi costs $0.43 in and $0.87 out per million, cached input $0.004.

KAT-Coder-Pro V2

2 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
AtlasCloudfp8$0.30$1.20262,144100.0%
StreamLakeunknown$0.30$1.20256,000

Not listed above: Kwaipilot directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

DeepSeek V4 Pro

7 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepSeek directfirst-party$0.43$0.87
DeepSeekunknown$0.43$0.871,048,576100.0%
GMICloudfp8$1.22$2.441,048,57599.9%
BaseTenfp4$1.32$3.961,048,576100.0%
Cloudflareunknown$1.32$3.961,048,57699.8%
Fireworksunknown$1.32$3.961,048,576100.0%
Novitafp8$1.32$3.961,048,57699.7%
SiliconFlowfp8$1.32$3.961,048,576100.0%

Buying direct from DeepSeek costs $0.43 in and $0.87 out per million, cached input $0.004.

Nex-N2-Pro

2 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Nex AGIfp8$0.25$1.00262,144
SiliconFlowunknown$0.50$2.50262,144

Tencent Hy3

6 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Tencent directfirst-party$0.15$0.59
Baidufp8$0.13$0.51262,14499.9%
GMICloudbf16$0.13$0.53262,14499.2%
Tencentfp8$0.13$0.53262,14499.9%
DeepInfrafp8$0.14$0.58262,14496.5%
Novitaunknown$0.14$0.58262,144100.0%
AtlasCloudfp8$0.20$0.80262,14499.8%

Buying direct from Tencent costs $0.15 in and $0.59 out per million, cached input $0.037. That is 16% dearer than the cheapest routed endpoint on a 3:1 blend. Priced in RMB: 1 / 4 / 0.25 per million (input / output / cached), converted at about 6.8 CNY to the dollar.

MiniMax-M3

12 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
MiniMax directfirst-party$0.30$1.20
CoreWeavefp4$0.23$0.96262,14499.9%
GMICloudfp8$0.24$0.961,048,57697.8%
DeepInfrafp8$0.28$1.10524,28898.8%
AtlasCloudfp8$0.30$1.20524,30099.7%
MiniMaxfp8$0.30$1.20524,28897.6%
Morphfp4$0.30$1.20256,00096.7%
Novitafp8$0.30$1.201,000,00099.7%
Parasailfp8$0.30$1.201,048,57699.0%
StreamLakefp8$0.30$1.201,000,00099.5%
Togetherunknown$0.30$1.20524,28899.5%
Venicefp8$0.30$1.20524,28899.8%
ModelRunfp4$0.75$3.001,048,57699.9%

Buying direct from MiniMax costs $0.30 in and $1.20 out per million, cached input $0.060. That is 27% dearer than the cheapest routed endpoint on a 3:1 blend.

Xiaomi MiMo-V2.5

5 providers · bf16, fp8
Served byQuant$/M in$/M outContextUptime 30m
Xiaomi directfirst-party$0.14$0.28
GMICloudfp8$0.14$0.281,050,00099.5%
Parasailfp8$0.14$0.281,048,57699.4%
Xiaomifp8$0.14$0.281,048,57699.8%
Novitafp8$0.17$0.341,048,57699.1%
DeepInfrabf16$0.40$2.00262,14497.2%

Buying direct from Xiaomi costs $0.14 in and $0.28 out per million, cached input $0.003.

DeepSeek V4 Flash

28 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepSeek directfirst-party$0.14$0.28
Baidufp8$0.08$0.161,048,57690.4%
Decartfp4$0.08$0.16262,14499.9%
GMICloudfp8$0.08$0.171,048,57599.7%
DeepInfrafp4$0.08$0.181,048,57699.1%
OpenInferencefp4$0.08$0.18262,14496.2%
Sail Researchfp4$0.09$0.18262,14499.4%
StreamLakefp8$0.09$0.181,024,00099.8%
DigitalOceanunknown$0.08$0.251,048,57699.1%
BaseTenfp8$0.13$0.261,048,57699.7%
CoreWeavefp8$0.13$0.28262,144100.0%
Inceptronfp4$0.13$0.281,048,57699.8%
Morphbf16$0.14$0.281,048,57698.5%
AkashMLfp8$0.14$0.28131,07299.5%
Ambientfp4$0.14$0.281,048,57697.9%
AtlasCloudfp4$0.14$0.281,048,57699.5%
Cloudflareunknown$0.14$0.281,048,576100.0%
DeepSeekfp8$0.14$0.281,048,576100.0%
Fireworksunknown$0.14$0.281,048,57696.4%
Novitafp8$0.14$0.281,048,57699.7%
Parasailfp8$0.14$0.281,048,57694.9%
Relacefp4$0.14$0.281,048,57699.5%
SiliconFlowfp8$0.14$0.281,048,57699.8%
Togetherunknown$0.14$0.281,048,57697.1%
Io Netfp8$0.15$0.32262,10099.8%
Veniceunknown$0.17$0.351,000,00097.0%
Phalaunknown$0.20$0.401,048,57699.3%
Mancer 2fp8$0.17$0.501,048,57697.6%
Waferunknown$0.28$0.561,048,57699.3%

Buying direct from DeepSeek costs $0.14 in and $0.28 out per million, cached input $0.003. That is 75% dearer than the cheapest routed endpoint on a 3:1 blend.

GPT-5.4 Nano

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.10$0.62400,00085.5%
Azureunknown$0.20$1.25400,000100.0%
OpenAIunknown$0.20$1.25400,00085.5%
Azureunknown$0.22$1.38400,000100.0%

GPT-5.4 Mini

5 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.38$2.25400,000100.0%
Azureunknown$0.75$4.50400,000100.0%
OpenAIunknown$0.75$4.50400,000100.0%
Azureunknown$0.83$4.95400,000100.0%
OpenAIunknown$1.50$9.00400,000100.0%

Qwen3.7 Plus

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.50$3.00
Alibaba / Qwenunknown$0.32$1.281,000,000100.0%

Buying direct from Alibaba / Qwen costs $0.50 in and $3.00 out per million, cached input $0.050. That is 101% dearer than the cheapest routed endpoint on a 3:1 blend.

GLM-5.1

18 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$1.40$4.40
StreamLakefp8$0.97$3.04200,00099.8%
Chutesfp8$0.98$3.08202,75274.6%
Waferfp4$1.00$3.20202,75295.5%
DeepInfrafp4$1.05$3.50202,752100.0%
DigitalOceanunknown$0.97$4.30163,84092.6%
SiliconFlowfp8$1.19$3.74204,80098.8%
AtlasCloudfp8$1.26$3.96202,752100.0%
Phalaunknown$1.21$4.20202,752
Crusoefp8$1.20$4.40202,752100.0%
Alibaba / Qwenfp8$1.33$4.18202,74599.5%
Novitafp8$1.38$4.40204,800100.0%
Baidufp8$1.40$4.40202,752100.0%
Friendliunknown$1.40$4.40202,752100.0%
GMICloudfp8$1.40$4.40202,75299.3%
Nebiusfp8$1.40$4.40202,75299.3%
Parasailfp8$1.40$4.40202,752100.0%
Z.ai / Zhipufp8$1.40$4.40202,75299.7%
Venicefp8$1.54$4.84200,00099.4%

Buying direct from Z.ai / Zhipu costs $1.40 in and $4.40 out per million, cached input $0.260. That is 45% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3.6 Plus

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.50$3.00
Alibaba / Qwenunknown$0.33$1.951,000,000100.0%

Buying direct from Alibaba / Qwen costs $0.50 in and $3.00 out per million, cached input $0.050. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3.6 27B

9 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.60$3.60
Chutesfp8$0.30$2.00262,14452.7%
Morphfp4$0.29$2.40131,072100.0%
Phalaunknown$0.32$2.70262,14495.5%
Alibaba / Qwenunknown$0.45$2.70262,144100.0%
Io Netfp8$0.39$2.8932,768100.0%
SiliconFlowfp8$0.30$3.20262,14499.9%
DeepInfrafp8$0.32$3.20262,14499.9%
Venicefp8$0.33$3.25256,000100.0%
CoreWeavefp8$0.60$3.60262,14499.8%

Buying direct from Alibaba / Qwen costs $0.60 in and $3.60 out per million. That is 86% dearer than the cheapest routed endpoint on a 3:1 blend.

MiniMax M2.7

11 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
MiniMax directfirst-party$0.30$1.20
Maraunknown$0.24$0.96196,608100.0%
DeepInfrafp8$0.25$1.00196,608100.0%
GMICloudfp8$0.27$1.08196,608100.0%
Novitafp8$0.27$1.08204,80098.2%
AtlasCloudfp8$0.30$1.20196,608
Fireworksunknown$0.30$1.20196,608
MiniMaxfp8$0.30$1.20204,80097.6%
DeepInfrafp8$0.38$1.70196,60899.5%
Groqunknown$0.60$1.80196,608100.0%
MiniMaxfp8$0.60$2.40204,800
SambaNovaunknown$0.60$2.40196,608

Buying direct from MiniMax costs $0.30 in and $1.20 out per million, cached input $0.060. That is 25% dearer than the cheapest routed endpoint on a 3:1 blend.

Inkling

3 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.95$4.05524,28899.8%
BaseTenfp8$1.00$4.051,048,57699.8%
Togetherunknown$1.00$4.05524,288100.0%

Not listed above: Thinking Machines directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GPT-5.1

5 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.62$5.00400,00099.6%
Azureunknown$1.25$10.00400,000100.0%
OpenAIunknown$1.25$10.00400,00099.6%
Azureunknown$1.38$11.00400,000
OpenAIunknown$2.50$20.00400,00099.6%

Gemini 3.5 Flash-Lite

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.15$1.251,048,57699.9%
Google AI Studiounknown$0.15$1.251,048,57699.6%
Googleunknown$0.30$2.501,048,57699.9%
Google AI Studiounknown$0.30$2.501,048,57699.6%
Googleunknown$0.33$2.751,048,576
Googleunknown$0.54$4.501,048,57699.9%
Google AI Studiounknown$0.54$4.501,048,57699.6%

Qwen3.5 397B A17B

11 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.60$3.60
DigitalOceanunknown$0.30$1.93131,072100.0%
Alibaba / Qwenunknown$0.39$2.34262,144100.0%
Chutesfp8$0.45$3.00262,14499.7%
DeepInfrafp8$0.45$3.00262,14499.8%
Parasailfp8$0.50$3.60262,14499.9%
AtlasCloudfp8$0.55$3.50262,14499.3%
Phalaunknown$0.55$3.50262,144100.0%
GMICloudfp8$0.60$3.60262,144
Novitaunknown$0.60$3.60262,14499.3%
StreamLakeunknown$0.60$3.60256,00096.2%
Veniceunknown$0.75$4.50128,000

Buying direct from Alibaba / Qwen costs $0.60 in and $3.60 out per million. That is 91% dearer than the cheapest routed endpoint on a 3:1 blend.

Mistral Medium 3.5

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$1.50$7.50
Mistralunknown$1.50$7.50262,144100.0%

Buying direct from Mistral costs $1.50 in and $7.50 out per million.

Kimi K2.5

10 providers · fp4, fp8, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Moonshot directfirst-party$0.60$3.00
DigitalOceanunknown$0.38$2.02262,14498.4%
DeepInfrafp4$0.45$2.25262,144100.0%
SiliconFlowint4$0.45$2.25262,144100.0%
AtlasCloudint4$0.49$2.50262,14499.8%
StreamLakefp8$0.54$2.70256,00099.9%
Novitaunknown$0.57$2.85262,14497.5%
Amazon Bedrockunknown$0.60$3.00262,14499.6%
Moonshotint4$0.60$3.00262,144100.0%
Phalaunknown$0.60$3.00262,144
Veniceunknown$0.56$3.50256,000100.0%

Buying direct from Moonshot costs $0.60 in and $3.00 out per million, cached input $0.100. That is 52% dearer than the cheapest routed endpoint on a 3:1 blend.

GLM 4.6

5 providers · bf16, fp4, fp8
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$0.60$2.20
Venicefp4$0.43$1.75198,000100.0%
DeepInfrafp4$0.50$2.00202,752
Novitabf16$0.55$2.20204,80098.5%
AtlasCloudfp8$0.60$2.20202,752100.0%
Z.ai / Zhipufp4$0.60$2.20202,75299.3%

Buying direct from Z.ai / Zhipu costs $0.60 in and $2.20 out per million, cached input $0.110. That is 32% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3.5-122B-A10B

5 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.40$3.20
Alibaba / Qwenunknown$0.26$2.08262,144100.0%
SiliconFlowfp8$0.26$2.08262,144100.0%
DeepInfrafp4$0.29$2.40262,14499.7%
AtlasCloudfp8$0.30$2.40262,14499.9%
Novitabf16$0.40$3.20262,14499.8%

Buying direct from Alibaba / Qwen costs $0.40 in and $3.20 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

LongCat 2.0

1 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
AtlasCloudfp8$0.30$1.201,048,756100.0%

Not listed above: Meituan directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GLM 4.7

9 providers · fp16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$0.60$2.20
DeepInfrafp4$0.40$1.75202,75299.7%
StreamLakefp8$0.48$1.76200,00096.4%
AtlasCloudfp8$0.52$1.85202,75296.7%
Novitafp8$0.54$1.98204,80095.4%
Googleunknown$0.60$2.20200,000100.0%
Z.ai / Zhipufp4$0.60$2.20202,75294.2%
Mancer 2fp4$0.60$2.50131,07298.4%
Venicefp4$0.55$2.65198,00098.9%
Cerebrasfp16$2.25$2.75131,072100.0%

Buying direct from Z.ai / Zhipu costs $0.60 in and $2.20 out per million, cached input $0.110. That is 36% dearer than the cheapest routed endpoint on a 3:1 blend.

DeepSeek V3.2

14 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
StreamLakefp8$0.21$0.32128,00099.4%
AtlasCloudfp8$0.26$0.38163,84099.8%
DeepInfrafp4$0.26$0.38163,84095.8%
SiliconFlowfp8$0.26$0.42163,84099.2%
Novitafp8$0.27$0.40163,84099.9%
Baidufp8$0.28$0.42131,072100.0%
GMICloudfp8$0.29$0.43163,840100.0%
Veniceunknown$0.33$0.48160,00091.8%
DigitalOceanunknown$0.25$0.80163,84098.4%
Alibaba / Qwenfp8$0.37$1.11131,07296.9%
Friendliunknown$0.50$1.50163,840100.0%
Googleunknown$0.56$1.68163,840100.0%
Phalaunknown$1.00$1.00163,840100.0%
SambaNovaunknown$3.00$4.5032,768

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

DeepSeek V3.1 Terminus

5 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp4$0.27$0.95163,840100.0%
Novitafp8$0.27$1.00131,072100.0%
SiliconFlowfp8$0.27$1.00163,84098.9%
AtlasCloudfp8$0.30$0.95131,072100.0%
StreamLakeunknown$0.34$1.03128,000100.0%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Gemma 4 31B

19 providers · bf16, fp16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenInferencebf16$0.08$0.35262,14494.7%
DeepInfrafp4$0.09$0.34262,144100.0%
CoreWeavebf16$0.10$0.34262,14497.0%
Venicebf16$0.12$0.36256,00096.5%
Chutesfp4$0.12$0.37131,0728.1%
DeepInfrafp8$0.13$0.38262,14499.1%
SiliconFlowfp8$0.13$0.40262,14498.9%
Crusoeunknown$0.14$0.40262,14493.9%
Friendliunknown$0.14$0.40262,14499.7%
Morphfp4$0.14$0.40175,00096.5%
Novitabf16$0.14$0.40262,14493.1%
Parasailfp8$0.15$0.40262,14487.6%
Phalaunknown$0.15$0.46262,14494.8%
DeepInfrafp8$0.27$0.76131,072100.0%
Togetherunknown$0.28$0.86262,144
Togetherunknown$0.39$0.97262,144
SambaNovaunknown$0.38$1.15262,14498.9%
ModelRunfp4$0.75$1.00262,144100.0%
Cerebrasfp16$0.99$1.49131,07299.9%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Ring-2.6-1T

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Novitaunknown$0.07$0.62262,144100.0%

Not listed above: InclusionAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Grok 4.3

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAI directfirst-party$1.25$2.50
xAIunknown$1.25$2.501,000,00099.9%
xAIunknown$1.25$2.501,000,00099.9%
xAIunknown$2.50$5.001,000,00099.9%
xAIunknown$2.50$5.001,000,00099.9%

Buying direct from xAI costs $1.25 in and $2.50 out per million, cached input $0.200.

Qwen3.6 35B A3B

9 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.25$1.49
Venicefp8$0.10$0.95256,000100.0%
DeepInfrafp8$0.10$0.95262,14495.8%
AkashMLfp8$0.14$1.00262,144100.0%
Parasailfp8$0.15$1.00262,14499.9%
AtlasCloudfp8$0.19$1.11262,144100.0%
Io Netfp8$0.19$1.19262,140100.0%
Phalaunknown$0.20$1.27262,14499.9%
CoreWeavefp8$0.25$1.25262,14499.9%
SiliconFlowfp8$0.20$1.60262,14499.9%

Buying direct from Alibaba / Qwen costs $0.25 in and $1.49 out per million. That is 79% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

o1

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$15.00$60.00200,000

Step 3.7 Flash

3 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfraunknown$0.20$1.15262,144100.0%
Novitafp8$0.20$1.15262,14499.1%
StepFunfp8$0.20$1.15256,00099.7%

Gemma 4 26B A4B

8 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.07$0.34262,14498.9%
Cloudflareunknown$0.10$0.30256,000100.0%
NextBitbf16$0.12$0.40262,14499.9%
SiliconFlowfp8$0.12$0.40262,14499.6%
Novitabf16$0.13$0.40262,14498.5%
Parasailbf16$0.13$0.40262,14499.3%
Venicebf16$0.13$0.40256,00099.0%
Googleunknown$0.15$0.60262,14498.3%

GPT-5

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$1.25$10.00400,000
OpenAIunknown$1.25$10.00400,00096.6%
Azureunknown$1.38$11.00400,000

Qwen3.5-35B-A3B

8 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.25$2.00
AkashMLfp8$0.14$1.00262,144100.0%
DeepInfrafp8$0.14$1.00262,144100.0%
Parasailfp8$0.15$1.00262,144100.0%
Alibaba / Qwenunknown$0.16$1.30262,14499.4%
CoreWeavefp8$0.25$1.25262,14488.4%
Veniceunknown$0.31$1.25256,00097.7%
AtlasCloudfp8$0.22$1.80262,144100.0%
SiliconFlowfp8$0.24$1.80262,144

Buying direct from Alibaba / Qwen costs $0.25 in and $2.00 out per million. That is 94% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3 Coder Next

4 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Parasailbf16$0.12$0.80262,144100.0%
StreamLakeunknown$0.18$0.90256,000
Novitafp8$0.20$1.50262,144100.0%
Alibaba / Qwenunknown$0.30$1.50262,144

Gemini 3.1 Flash Lite

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.12$0.751,048,57698.2%
Google AI Studiounknown$0.12$0.751,048,57699.7%
Googleunknown$0.25$1.501,048,57698.2%
Google AI Studiounknown$0.25$1.501,048,57699.7%
Googleunknown$0.28$1.651,048,576
Googleunknown$0.45$2.701,048,57698.2%
Google AI Studiounknown$0.45$2.701,048,57699.7%

Gemini 2.5 Pro

8 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.62$5.001,048,57699.8%
Google AI Studiounknown$0.62$5.001,048,57699.2%
Googleunknown$1.25$10.001,048,57699.8%
Googleunknown$1.25$10.001,048,576
Googleunknown$1.25$10.001,048,576
Google AI Studiounknown$1.25$10.001,048,57699.2%
Googleunknown$2.25$18.001,048,57699.8%
Google AI Studiounknown$2.25$18.001,048,57699.2%

Mercury 2

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Inceptionunknown$0.25$0.75128,000100.0%

gpt-oss-120b

20 providers · bf16, fp16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
CoreWeavefp4$0.03$0.17131,07294.2%
DeepInfrabf16$0.04$0.17131,07299.9%
Novitafp4$0.05$0.25131,07283.1%
DigitalOceanunknown$0.06$0.39128,000100.0%
SiliconFlowfp8$0.05$0.45131,07299.4%
AkashMLbf16$0.04$0.49131,07299.9%
Googleunknown$0.09$0.36131,07292.9%
Mancer 2fp8$0.08$0.50131,07299.3%
BaseTenfp4$0.10$0.50128,072100.0%
Amazon Bedrockunknown$0.15$0.60131,072100.0%
Amazon Bedrockunknown$0.15$0.60131,072
DeepInfrabf16$0.15$0.60131,07299.5%
Groqunknown$0.15$0.60131,072100.0%
Nebiusfp4$0.15$0.60131,07299.9%
Phalaunknown$0.15$0.60131,072100.0%
Togetherunknown$0.15$0.60131,07294.9%
Parasailfp4$0.10$0.75131,072100.0%
Maraunknown$0.15$0.75131,07286.2%
SambaNovaunknown$0.14$0.95131,07292.5%
Cerebrasfp16$0.35$0.75131,07299.8%

Not listed above: OpenAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3.5-9B

5 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.10$0.15262,14499.9%
SiliconFlowfp8$0.10$0.15262,14469.4%
Venicefp8$0.10$0.15256,000100.0%
Parasailbf16$0.10$0.25262,144100.0%
Togetherunknown$0.17$0.25262,14499.6%

Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Command A

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Cohereunknown$2.50$10.00256,000100.0%

Mistral Small 4

2 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$0.15$0.60
Mistralunknown$0.15$0.60262,144100.0%
Venicefp8$0.19$0.75256,000100.0%

Buying direct from Mistral costs $0.15 in and $0.60 out per million.

Trinity Large Thinking

2 providers · fp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Parasailfp4$0.22$0.85262,144100.0%
Arcee AIunknown$0.25$0.80262,144

Not listed above: Arcee Ai directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Ling-3.0-flash

2 providers · bf16, unknown
Served byQuant$/M in$/M outContextUptime 30m
Novitaunknown$0.02$0.06262,144100.0%
DeepInfrabf16$0.06$0.18131,07299.6%

Not listed above: InclusionAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Muse Spark 1.2

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Metaunknown$1.25$4.251,048,576100.0%

Qwen3.8 Max

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwenunknown$2.00$6.001,000,000100.0%

GPT-4o

2 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$2.50$10.00128,000100.0%
OpenAIunknown$2.50$10.00128,00099.8%

DeepSeek V3

3 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepSeek directfirst-party$0.14$0.28
StreamLakeunknown$0.26$1.03128,00099.8%
DeepInfrafp4$0.32$0.89163,84099.2%
Novitafp8$0.40$1.3064,000100.0%

Buying direct from DeepSeek costs $0.14 in and $0.28 out per million, cached input $0.003. That is 61% cheaper than the cheapest routed endpoint on a 3:1 blend.

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GPT-4 Turbo

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$10.00$30.00128,000100.0%

DeepSeek V3 0324

4 providers · bf16, fp4, fp8
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp4$0.24$0.90163,84097.8%
SiliconFlowfp8$0.25$1.00163,84099.6%
Novitafp8$0.27$1.12163,84099.9%
Crusoebf16$0.50$1.50163,840100.0%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

gpt-oss-20b

12 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
CoreWeavefp4$0.03$0.13131,07299.6%
DeepInfrabf16$0.03$0.14131,072100.0%
Parasailfp4$0.03$0.15131,07299.3%
Novitafp4$0.04$0.15131,072100.0%
Phalaunknown$0.04$0.15131,07296.6%
SiliconFlowfp8$0.04$0.18131,07295.3%
Togetherunknown$0.05$0.20131,072
Amazon Bedrockunknown$0.07$0.15131,072
Amazon Bedrockunknown$0.07$0.15131,07297.5%
Googleunknown$0.07$0.25131,07298.4%
Fireworksunknown$0.07$0.30131,07298.9%
Groqunknown$0.07$0.30131,07299.9%

Not listed above: OpenAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Mistral Medium 3.1

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistralunknown$0.40$2.00131,072100.0%

GPT-4.1 Mini

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$0.40$1.601,047,576100.0%
OpenAIunknown$0.40$1.601,047,57699.6%
Azureunknown$0.44$1.761,047,576100.0%

Llama 4 Maverick

5 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DigitalOceanunknown$0.20$0.70128,000100.0%
DeepInfrafp8$0.20$0.801,048,57699.6%
Novitafp8$0.27$0.851,048,57699.0%
Parasailfp8$0.35$1.00524,28899.9%
Googleunknown$0.35$1.15524,28898.0%

Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

o3 Mini

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.10$4.40200,000100.0%

Solar Pro 3

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Upstageunknown$0.15$0.60131,072

GPT-5 Mini

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.12$1.00400,00099.8%
Azureunknown$0.25$2.00400,000100.0%
OpenAIunknown$0.25$2.00400,00099.8%
Azureunknown$0.28$2.20400,000

Qwen3 32B

4 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.70$2.80
DeepInfrafp8$0.08$0.2840,960100.0%
Nebiusfp8$0.10$0.3040,96099.8%
SiliconFlowfp8$0.14$0.57131,072100.0%
Groqunknown$0.29$0.59131,072100.0%

Buying direct from Alibaba / Qwen costs $0.70 in and $2.80 out per million. That is 842% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3 14B

3 providers · fp8, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.35$1.40
NextBitint4$0.10$0.2240,96099.3%
DeepInfrafp8$0.12$0.2440,96099.9%
Alibaba / Qwenunknown$0.23$0.91131,072100.0%

Buying direct from Alibaba / Qwen costs $0.35 in and $1.40 out per million. That is 371% dearer than the cheapest routed endpoint on a 3:1 blend.

GPT-4

2 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$30.00$60.008,191100.0%
OpenAIunknown$30.00$60.008,191100.0%

GPT-4o-mini

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$0.15$0.60128,000100.0%
OpenAIunknown$0.15$0.60128,000100.0%
Azureunknown$0.17$0.66128,000

GPT-4.1 Nano

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$0.10$0.401,047,57699.9%
OpenAIunknown$0.10$0.401,047,57699.5%
Azureunknown$0.11$0.441,047,576

GPT-3.5 Turbo

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.50$1.5016,385100.0%

Granite 4.1 8B

1 providers · bf16
Served byQuant$/M in$/M outContextUptime 30m
CoreWeavebf16$0.05$0.10131,072100.0%

Not listed above: IBM directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3 8B

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.18$0.70
Alibaba / Qwenunknown$0.12$0.45131,072100.0%

Buying direct from Alibaba / Qwen costs $0.18 in and $0.70 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

Llama 4 Scout

4 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.10$0.30327,680100.0%
Groqunknown$0.11$0.34131,07299.9%
Novitabf16$0.18$0.59131,07299.9%
Googleunknown$0.25$0.701,310,72099.7%

Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Grok Build 0.1

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAI directfirst-party$1.00$2.00
xAIunknown$1.00$2.00256,000100.0%
xAIunknown$1.00$2.00256,000
xAIunknown$2.00$4.00256,000100.0%
xAIunknown$2.00$4.00256,000

Buying direct from xAI costs $1.00 in and $2.00 out per million, cached input $0.200.

Nemotron 3 Ultra 550B

4 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.50$2.50
DeepInfrafp4$0.50$2.20262,14491.2%
BaseTenfp4$0.60$2.40202,80098.8%
Venicefp8$0.62$3.12256,00087.4%
Togetherunknown$0.60$3.60512,28898.2%

Buying direct from Nvidia costs $0.50 in and $2.50 out per million, cached input $0.150. That is 8% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Nemotron 3 Super 120B

3 providers · bf16, fp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.20$0.80
DeepInfrabf16$0.08$0.40262,14492.8%
DigitalOceanunknown$0.17$0.361,000,00092.9%
Nebiusfp4$0.30$0.908,00098.0%

Buying direct from Nvidia costs $0.20 in and $0.80 out per million. That is 114% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

North Mini Code

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Cohere directfirst-party$0.00$0.00
Cohereunknown$0.00$0.00256,00096.3%

Buying direct from Cohere costs $0.00 in and $0.00 out per million.

Mistral Small 3.1

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Cloudflareunknown$0.35$0.55128,000100.0%

Not listed above: Mistral directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

DeepSeek R1

1 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
Novitafp8$0.70$2.5064,000100.0%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3 235B A22B 2507

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.70$2.80
Alibaba / Qwenunknown$0.45$1.82131,072100.0%

Buying direct from Alibaba / Qwen costs $0.70 in and $2.80 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

Mistral Large 3

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$0.50$1.50
Mistralunknown$2.00$6.00128,000100.0%

Buying direct from Mistral costs $0.50 in and $1.50 out per million. That is 75% cheaper than the cheapest routed endpoint on a 3:1 blend.

Nemotron 3 Nano 30B

3 providers · fp4, fp8
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.00$0.00
Crusoefp8$0.05$0.20262,14498.2%
DeepInfrafp4$0.05$0.20262,14496.3%
Novitafp4$0.05$0.20262,1440.0%

Buying direct from Nvidia costs $0.00 in and $0.00 out per million.

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Nemotron 3 Nano Omni 30B

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.00$0.00
Nvidiaunknown$0.00$0.00256,00099.2%

Buying direct from Nvidia costs $0.00 in and $0.00 out per million.

Mistral Small 3.2

3 providers · bf16, fp8
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$0.10$0.30
DeepInfrafp8$0.07$0.20128,00099.8%
Venicefp8$0.09$0.25256,00099.9%
Parasailbf16$0.09$0.30131,07298.8%

Buying direct from Mistral costs $0.10 in and $0.30 out per million. That is 41% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Mistral directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3 30B A3B 2507

2 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.12$0.5040,960100.0%
Alibaba / Qwenunknown$0.13$0.52131,072100.0%

Gemma 3 27B

5 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.08$0.16131,07299.8%
Novitabf16$0.12$0.2098,30499.3%
Nebiusfp8$0.10$0.30110,00099.4%
Parasailfp8$0.08$0.45131,07299.9%
Phalaunknown$0.15$0.46262,14489.2%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Gemma 3 12B

1 providers · bf16
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.05$0.15131,072100.0%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Gemma 3n E4B

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Togetherunknown$0.06$0.1232,768100.0%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Gemma 3 4B

1 providers · bf16
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.05$0.10131,072100.0%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

What it actually costs

Observed, not list price

Aggregates across opencode’s subscriber base: real cost per coding session after caching. The spread is far wider than list prices suggest, and the cache ratio explains much of it.

Model$ / sessionEffective $/MCache ratioTokens / sessionSessions observed
Xiaomi MiMo-V2.5$0.0722$0.0194.3%5.9M1,448,024
DeepSeek V4 Flash$0.1393$0.0196.4%14.6M12,096,796
Tencent Hy3$0.2951$0.1687.0%1.8M977
Xiaomi MiMo-V2.5-Pro$0.3802$0.0695.5%5.9M373,253
DeepSeek V4 Pro$0.4384$0.0596.8%8.1M3,949,114
Qwen3.7 Plus$0.5280$0.1794.6%3.0M645,425
MiniMax-M3$0.5434$0.0894.3%7.1M1,245,848
Kimi K2.6$0.8180$0.2891.2%2.9M314,084
Grok 4.5$0.8340$0.6786.7%1.2M79,708
Kimi K2.7 Code$1.0948$0.2594.2%4.3M720,175
Qwen3.7 Max$1.5594$0.9695.1%1.6M228,040
GLM-5.1$1.9516$0.6175.1%3.2M127,184
Kimi K3$2.0958$0.6992.0%3.0M426,856
GLM-5.2$2.5679$0.5180.6%5.0M1,750,878

Best value

Two ways to pay, two different answers

Value means something different depending on how you are billed, so this splits in two. On a subscription what matters is how much quality-weighted work a fixed monthly fee yields; on metered billing it is quality per unit of token spend. The same model can look strong under one and ordinary under the other.

On a subscription

Coding index times sessions per month, divided by what the plan costs. It answers a narrower question than it looks — how much quality-weighted work a subscription yields, assuming the cheap model is good enough for the job in hand.

#PlanModel$/moSessions/moCoding indexValue: index x sessions per $
1opencode GoXiaomi MiMo-V2.5$1083156.8██████████████████████████████████████████████ 4,720
2opencode GoDeepSeek V4 Flash$1043169.1█████████████████████████████ 2,976
3Qoder ProXiaomi MiMo-V2.5$2055456.8███████████████ 1,573
4GitHub MaxXiaomi MiMo-V2.5$1002,77056.8███████████████ 1,573
5Cursor UltraXiaomi MiMo-V2.5$2005,54056.8███████████████ 1,573
6Qoder Pro+Xiaomi MiMo-V2.5$601,66256.8███████████████ 1,573
7Qoder UltraXiaomi MiMo-V2.5$2005,54056.8███████████████ 1,573
8GitHub Pro+Xiaomi MiMo-V2.5$3997056.8██████████████ 1,412
9opencode GoTencent Hy3$1020358.8████████████ 1,196
10GitHub ProXiaomi MiMo-V2.5$1020856.8████████████ 1,180
11Qoder ProDeepSeek V4 Flash$2028769.1██████████ 992
12Qoder Pro+DeepSeek V4 Flash$6086169.1██████████ 992
13GitHub MaxDeepSeek V4 Flash$1001,43669.1██████████ 992
14Cursor UltraDeepSeek V4 Flash$2002,87269.1██████████ 992
Absence from this table is not a verdict. Only plans that publish a dollar budget can be ranked here, which is three vendors out of fourteen. The rest are excluded for being unmeasurable, not for being poor value — and several of them are deliberately generous. MiniMax is the clearest case: it publishes no quota at all, yet the figure reverse-engineered further down this page (about 32B tokens a month on the $120 tier) works out at roughly 4,200 sessions, which would place it third here — ahead of every GitHub and Cursor tier listed. Z.ai, Alibaba and Kimi meter in prompts or requests with no published token cost, so they cannot be placed at all. Read the table as the best of what can be measured, never as the best available.

Read it with the leaderboard beside you, too. A weak model on a generous plan can top this table while still being the wrong tool: the metric rewards volume, and quality enters only linearly. It is a guide to where the money goes furthest, not to what you should run for work that has to be right first time.

Pay per token

Coding index per dollar per million tokens, blended 3:1. The list column ranks what you are quoted; the observed column uses the effective rate actually paid in production once caching is working, which is the number that decides a monthly bill.

#ModelCoding index$/M listValue on list$/M observedValue observed
1Ling-3.0-flash · InclusionAI50.6$0.0321,606
2Solar Pro 4 · Upstage52.7$0.0521,004
3gpt-oss-120b · OpenAI30.4$0.065468
4DeepSeek V4 Flash · DeepSeek69.1$0.175395$0.016,910
5gpt-oss-20b · OpenAI20.7$0.055376
6Xiaomi MiMo-V2.5 · Xiaomi56.8$0.175325$0.015,680
7GPT-5.6 Luna · OpenAI71.4$0.225317
8Gemma 4 31B · Google43.4$0.160271
9Qwen3.5-9B · Alibaba / Qwen28.7$0.112255
10Tencent Hy3 · Tencent58.8$0.231255$0.16368
11Gemma 4 26B A4B · Google39.3$0.190207
12Ring-2.6-1T · InclusionAI42.8$0.212201
13Nemotron 3.5 Lightning · Nvidia26.8$0.138195
14Nemotron 3 Nano 30B · Nvidia14.4$0.087165
15Granite 4.1 8B · IBM9.5$0.062152
16DeepSeek V3.2 · DeepSeek44.2$0.302146
17Nex-N2-Pro · Nex AGI59.1$0.438135
18DeepSeek V4 Pro · DeepSeek68.8$0.544127$0.051,376
19Qwen3 Coder Next · Alibaba / Qwen36.2$0.290125
20GPT-5.4 Nano · OpenAI56.1$0.463121
21Qwen3 32B · Alibaba / Qwen15.3$0.130118
22Qwen3.6 35B A3B · Alibaba / Qwen41.9$0.362116
23KAT-Coder-Pro V2 · Kwaipilot59.5$0.525113
24MiniMax-M3 · MiniMax58.6$0.525112$0.08732
25Xiaomi MiMo-V2.5-Pro · Xiaomi60.2$0.544111$0.061,003
26Mistral Small 4 · Mistral26.6$0.262101
27MiniMax M2.7 · MiniMax52.6$0.525100
28Qwen3.7 Plus · Alibaba / Qwen55.9$0.560100$0.17329
29DeepSeek V3.1 Terminus · DeepSeek43.5$0.44099
30Mistral Small 3.2 · Mistral12.5$0.13394
31Qwen3 14B · Alibaba / Qwen13.8$0.15092
32Step 3.7 Flash · StepFun39.6$0.43891
33LongCat 2.0 · Meituan45.3$0.52586
34Mercury 2 · Inception31.1$0.37583
35Gemma 3 12B · Google5.8$0.07577
36Muse Glimmer · Meta49.0$0.63777
37Qwen3.6 Plus · Alibaba / Qwen54.5$0.73175
38Trinity Large Thinking · Arcee Ai25.8$0.37868
39Mistral Small 3.1 · Mistral26.3$0.40265
40GPT-4.1 Nano · OpenAI11.1$0.17563
41Solar Pro 3 · Upstage16.2$0.26262
42Gemini 3.1 Flash Lite · Google34.7$0.56262
43GLM 4.7 · Z.ai / Zhipu45.3$0.73861
44Kimi K2.6 · Moonshot61.8$1.01061$0.28221
45Qwen3.5-35B-A3B · Alibaba / Qwen37.0$0.61960
46Gemma 3 27B · Google10.1$0.17259
47Gemini 3.5 Flash-Lite · Google49.3$0.85058
48Qwen3 30B A3B 2507 · Alibaba / Qwen12.1$0.21556
49Qwen3.5-122B-A10B · Alibaba / Qwen45.7$0.81756
50Qwen3.5 397B A17B · Alibaba / Qwen48.2$0.87755
51Llama 4 Scout · Meta8.2$0.15055
52DeepSeek V3 · DeepSeek23.0$0.45051
53GLM 4.6 · Z.ai / Zhipu45.8$0.96348
54Llama 4 Maverick · Meta16.3$0.35047
55Gemini 3.6 Flash · Google69.2$1.50046
56Qwen3 8B · Alibaba / Qwen9.0$0.20245
57DeepSeek V3 0324 · DeepSeek21.2$0.48344
58GPT-4o-mini · OpenAI11.4$0.26243
59Gemma 3 4B · Google2.7$0.06243
60Kimi K2.7 Code · Moonshot60.8$1.40743$0.25243
61Gemma 3n E4B · Google3.2$0.07543
62Grok Build 0.1 · xAI51.5$1.25041
63Kimi K2.5 · Moonshot46.8$1.14041
64Qwen3.6 27B · Alibaba / Qwen53.7$1.35040
65GLM-5.2 · Z.ai / Zhipu68.8$1.82838$0.51135
66Nemotron 3 Ultra 550B · Nvidia49.3$1.35037
67Muse Spark 1.2 · Meta72.2$2.00036
68Muse Spark 1.1 · Meta71.3$2.00036
69GPT-5.6 Terra · OpenAI76.7$2.25034
70GPT-5.4 Mini · OpenAI56.1$1.68833
71Inkling · Thinking Machines52.1$1.72530
72Qwen3.7 Max · Alibaba / Qwen66.0$2.21330$0.9669
73GPT-4.1 Mini · OpenAI20.2$0.70029
74Qwen3 235B A22B 2507 · Alibaba / Qwen22.1$0.79628
75Grok 4.3 · xAI42.2$1.56227
76GLM-5.1 · Z.ai / Zhipu55.8$2.15026$0.6191
77Mistral Medium 3.1 · Mistral20.5$0.80026
78Grok 4.6 · xAI76.8$3.00026
79Grok 4.5 · xAI72.4$3.00024$0.67108
80Qwen3.8 Max · Alibaba / Qwen71.8$3.00024
81GPT-5 Mini · OpenAI15.6$0.68823
82DeepSeek R1 · DeepSeek24.6$1.15021
83Gemini 3.5 Flash · Google70.1$3.37521
84Claude Sonnet 5 · Anthropic71.5$4.00018
85Mistral Medium 3.5 · Mistral46.9$3.00016
86Gemini 3.1 Pro Preview · Google68.8$4.50015
87GPT-5.1 · OpenAI49.4$3.43814
88GPT-3.5 Turbo · OpenAI10.7$0.75014
89Kimi K3 · Moonshot76.2$6.00013$0.69110
90GPT-5.4 · OpenAI71.1$5.62513
91GPT-5 · OpenAI37.8$3.43811
92Claude Sonnet 4.6 · Anthropic63.0$6.00010
93Gemini 2.5 Pro · Google33.3$3.43810
94o3 Mini · OpenAI16.3$1.9258
95Claude Opus 5 · Anthropic78.0$10.0008
96Claude Opus 4.8 · Anthropic74.3$10.0007
97Claude Opus 4.7 · Anthropic73.6$10.0007
98GPT-5.6 Sol · OpenAI78.3$11.2507
99Mistral Large 3 · Mistral20.1$3.0007
100GPT-5.5 · OpenAI74.9$11.2507
101Command A · Cohere27.8$4.3756
102GPT-4o · OpenAI24.2$4.3756
103Claude Fable 5 · Anthropic76.5$20.0004
104o1 · OpenAI39.7$26.2502
105GPT-4 Turbo · OpenAI21.5$15.0001
106GPT-4 · OpenAI13.1$37.5000

The two columns do not agree, and that disagreement is the point: list price is a quote, the observed rate is a bill. A model with a poor cache ratio slides down the observed column even when its quoted price looked competitive. Observed rates come from opencode's subscriber base, so models it does not serve show a dash rather than a guess.

What a plan actually buys

Included usage divided by observed cost per session. Only dollar-denominated plans can be answered this precisely — plans metered in prompts, or with no published quota, cannot appear here at all.

Plan$/moIncludedXiaomi MiMo-V2.5DeepSeek V4 FlashTencent Hy3Xiaomi MiMo-V2.5-ProDeepSeek V4 ProQwen3.7 PlusMiniMax-M3Kimi K2.6Grok 4.5Kimi K2.7 CodeQwen3.7 MaxGLM-5.1Kimi K3GLM-5.2
opencode Go$10$6083143020315713611311073715438302823
GitHub Max$100$2002,7701,4356775264563783682442391821281029577
Cursor Ultra$200$4005,5402,8711,3551,052912757736488479365256204190155
Qoder Pro$20$4055428713510591757348473625201915
Qoder Pro+$60$1201,66286140631527322722014614310976615746
Qoder Ultra$200$4005,5402,8711,3551,052912757736488479365256204190155
GitHub Pro+$39$7096950223718415913212885836344353327
GitHub Pro$10$1520710750393428271817139775
Cursor Pro+$60$7096950223718415913212885836344353327
Venice Max$200$2253,1161,61576259151342641427526920514411510787
Venice Pro Plus$68$751,03853825419717114213891896848383529
Cursor Pro$20$202771436752453736242318121097
Augment Code Business$100$1001,3857173382632281891841221199164514738
Venice Pro$18$1137322111100000

Sessions per month. Heavier models burn a budget far faster; real mileage varies with session length and cache behaviour.

Subscription plans

Grouped by how verifiable the quota actually is

Most vendors do not publish a convertible quota. Effective $/token is computed only where a vendor publishes enough to do it honestly; everything else is left blank rather than guessed.

ProviderPlan$/monthQuota disclosureSee what's leftQuota converts to units?Checked
XiaomiMiMo Token Plan (Lite / Standard / Pro / Max)$6.00 / $16.00 / $50.00 / $100.00Convertible token quotaconsole onlyconverts2026-08-13
opencodeopencode Go$10.00Dollar usage budgetconsole onlyconverts2026-08-13
GitHubCopilot (Pro / Pro+ / Max)$0.00 / $10.00 / $39.00 / $100.00Dollar usage budgetconsole onlyconverts2026-08-13
CursorPro / Pro+ / Ultra$20.00 / $60.00 / $200.00Dollar usage budgetconsole onlyconverts2026-08-13
Alibaba / QwenCoding Plan (Pro)$50.00Requests / promptsconsole onlypartial2026-08-13
Alibaba / QwenToken Plan (Individual + Team)$6.00 / $18.00 / $68.00No numeric quotaconsole onlyunstated2026-08-13
Z.ai / ZhipuGLM Coding (Lite / Pro / Max)$18.00 / $72.00 / $160.00Requests / promptsconsole onlypartial2026-08-14
MiniMaxToken Plan (Plus / Max / Ultra)$20.00 / $50.00 / $120.00No numeric quotaAPIunstated2026-08-13
AnthropicClaude Pro / Max 5x / Max 20x$20.00 / $100.00 / $200.00No numeric quotaconsole onlyunstated2026-08-13
OpenAIChatGPT Plus / Pro 5x / Pro 20x$0.00 / $8.00 / $20.00 / $100.00 / $200.00No numeric quotaconsole onlypartial2026-08-13
GoogleAI Plus / Pro / Ultra$4.99 / $19.99 / $99.99 / $200.00No numeric quotaconsole onlyunstated2026-08-13
Moonshot / KimiModerato / Allegretto / Allegro / Vivace$19.00 / $39.00 / $99.00 / $199.00No numeric quotaAPIpartial2026-08-13
MistralLe Chat Free / Pro / Team (Mistral Vibe bundled)$0.00 / $14.99 / $24.99No numeric quotanowhereunstated2026-08-13
Windsurf (Cognition/Devin)Pro / Max / Teams$0.00 / $20.00 / $200.00No numeric quotaconsole onlyunstated2026-08-13
OllamaOllama Cloud (Free / Pro / Max)$0.00 / $20.00 / $100.00No numeric quotaconsole onlyunstated2026-08-13
VeniceVenice (Free / Pro / Pro Plus / Max)$0.00 / $18.00 / $68.00 / $200.00Dollar usage budgetconsole onlyconverts2026-08-13
AtlasCloudCoding Plan (Starter / Lite / Plus / Max / Ultra / Enterprise)$10.00 / $20.00 / $50.00 / $100.00 / $200.00 / $500.00Requests / promptsconsole onlyunstated2026-08-13
Io NetIO Intelligence (Standard / Professional / Developer)$0.00 / $15.00 / $150.00Requests / promptsconsole onlyunstated2026-08-13
ChutesPlus / Pro$10.00 / $20.00No numeric quotaconsole onlyunstated2026-08-13
MorphUsage-based (Free / Pay-per-token / Scale flat-rate)$0.00 / $200.00Dollar usage budgetconsole onlyconverts2026-08-13
QoderQoder (Free / Pro / Pro+ / Ultra)$0.00 / $20.00 / $60.00 / $200.00Dollar usage budgetconsole onlypartial2026-08-13
Agnes AIToken Plan (Starter / Plus / Pro)$4.00 / $10.00 / $50.00Requests / promptsconsole onlypartial2026-08-13
SyntheticSynthetic (Standard)$30.00Dollar usage budgetconsole onlypartial2026-08-13
Augment CodeAugment (Business / Enterprise) — Auggie CLI + Cosmos$100.00Dollar usage budgetconsole onlyconverts2026-08-13
One plan meters something else entirely. Ollama Cloud bills on “actual utilization of Ollama’s cloud infrastructure — primarily GPU time”, the only compute-metered subscription here. That inverts an assumption the rest of this page relies on: everywhere else a token costs the same whoever produced it, so model speed only affects your patience. On GPU time, speed is the quota — a model running at half the tokens per second burns roughly twice the allowance for identical output. Check the leaderboard's tok/s column before picking a model to run there.

These two columns are inverted. The vendors that let you read your quota programmatically do not tell you what you bought, and the ones that publish proper unit maths only give you a web page. MiniMax and Kimi expose an endpoint for a quota with no published units at all; Xiaomi, GitHub and opencode publish exact conversions but no API. Nobody does both. DeepSeek is the only one coherent on both axes, and only because it sells no subscription — there, remaining quota is simply money.

Only three providers document a remaining-quota endpoint, and all three are Chinese. Western billing APIs return consumed usage and never an allowance, which is why every community usage tool hardcodes the limit and does the subtraction itself. Endpoints for MiniMax, Z.ai, Kimi and DeepSeek were called directly against live keys for this table; provider coverage cross-checked against quota-sentinel.

What moves the price inside a plan

Bigger than the gap between plans

Choosing a subscription is the decision buyers agonise over, and it is usually the smaller one. Once inside a plan, what you pay per unit of work varies by model, by tier, by how you use the context and, at two vendors, by the time of day. At Qoder those axes compound into a 27-fold swing in what a credit is worth, far more than separates the plans themselves.

VendorVaries byRangeDetail
Qodermodel0.1x - 0.6xA published credit multiplier per model. They do not track open-market price: one unit of credit buys $5.60 of Qwen3.7-Plus but only $0.68 of GLM-5.2.
Qoderclockdown to 0.01xSelected models are discounted 14:00-00:00 UTC, weekends and holidays included. Qwen3.7-Max falls 0.5x to 0.1x; a promoted preview model reaches 0.01x. Credit price only -- model quality is explicitly unaffected.
Qodertierfree - 1.6xLite costs nothing, Efficient 0.3x, Auto 1.0x, Performance 1.1x, Ultimate 1.6x (currently halved to 0.8x). Sits on top of the model choice, so the two compound.
Z.ai / Zhipuclock2x - 3xThe mirror image: quota burns 3x during peak hours and 2x off-peak, rather than being discounted. A promotion holds off-peak at 1x until 30 September 2026, after which the usable plan roughly halves at an unchanged price.
Cursormechanics1.25x - 2xCache writes cost 1.25x input, and long context doubles above 200K tokens. Charges that depend on how you use the model rather than which one you pick.
Xiaomicache50x - 120xThe largest spread found anywhere, in the subscriber's favour: a cache hit costs 2 credits per token against 100 on mimo-v2.5, and 2.5 against 300 on the Pro model. Consumption also drops to 0.8x between 00:00 and 08:00 Beijing time.
opencodemodel2xKimi K3 draws twice the usage of every other model in the Go plan.
Anthropic / Google / OpenAItier2x - 20xSold as more usage rather than as a multiplier on consumption, and not comparable between vendors: Google's 4x is measured against its free tier while its 5x and 20x are against Pro, and Anthropic's multipliers are per session, not per month.
DeepSeekclockup to 12xFrom 2026-08-16 16:00 UTC DeepSeek moves to peak/off-peak API billing with item-specific multipliers, a sharp escalation of the flat peak surcharge it had run since mid-2026. Against the current flat price, V4 Pro cache-hit input rises 6x off-peak and 12x at peak ($0.0036 -> $0.022 / $0.044 per M), cache-miss input 1.5x / 3x ($0.435 -> $0.66 / $1.32) and output 2.25x / 4.5x ($0.87 -> $1.98 / $3.96); V4 Flash scales the same way ($0.14 miss -> $0.22 / $0.44, $0.28 output -> $0.66 / $1.32). Peak is 01:00-04:00 and 06:00-10:00 UTC (Beijing working hours 09:00-12:00 and 14:00-18:00), off-peak is everything else at half the peak rate. The cache-hit multiplier is the one that bites: a coding agent replays a large cached prefix every turn at ~92-95% cache-hit, so its effective bill tracks the 12x line, not the 3x cache-miss one. Confirmed on the official pricing page 2026-08-14; the flat rate a comparison table quotes is only the pre-Aug-16 number.

Two deserve naming. Qoder and Z.ai run the same mechanism in opposite directions, one discounting the quiet hours and the other surcharging the busy ones, and only a discount looks like a discount. And Xiaomi's cache multiplier is the one that runs in the customer's favour, at up to 120x, which is why cache behaviour decides more of a coding agent's bill than the choice of model does.

MiniMax publishes no token, credit or prompt figure for any tier, and its own API returns a percentage with the total zeroed out. But the two halves can be put together: the console reports tokens consumed, the API reports the percentage left. Divide one by the other and the denominator falls out.

Tier
Ultra ($120/month)
Weekly quota
7.4B tokens
Monthly
32B tokens
Effective
$0.0038 /M

Method. The console reports tokens consumed; the API reports the percentage of the weekly window still left. Dividing one by the other gives the denominator the vendor never publishes. Measured 2026-07-18 against the plan window 13-20 Jul: 7.68B tokens consumed over a rolling 7 days against 6% of the window remaining. Anyone can repeat this on their own account; it costs nothing and spends no quota.

Confidence. Derived from a SINGLE account; not an official figure. Robust to the API's integer-percentage resolution (worth about 0.05B) and to the rolling-versus-fixed window mismatch (7.2-7.7B). One caveat that cannot be resolved from outside: the internal weighting of cached input, fresh input and output has not been reverse-engineered, so this figure is valid at the observed ~92% cache-hit rate and should not be extrapolated to a very different one. Note also that the accounting has already changed once: cached tokens formerly did not count against the quota and now do, which raised the effective cost sharply and was not announced.

For scale: at list price that traffic would cost $0.525 per million blended, and about $0.06 per million once caching is working. The subscription lands near $0.004 — roughly 16x cheaper than pay-per-token even after caching, and two orders of magnitude under list.

Why this had to be measured. Cached tokens did not always count against this quota. At some point they began to, the effective cost rose sharply, and nothing was announced — the subscription simply started buying less. The internal weighting of cached input against fresh input and output remains un-reverse-engineered, so the figure above holds at the observed ~92% cache-hit rate and should not be stretched far beyond it. A number a vendor does not publish is also a number it can change quietly.

By provider

One page per lab: its models, what they cost, the endpoints serving them and at what precision, plus its own plan if it sells one.

Agnes AIrequests
provider
Token Plan (Starter / Plus / Pro) · from $4/mo
AkashML
serves 5
No subscription of its own
Alibaba / Qwenrequests
16 models · serves 18
Coding Plan (Pro) · from $50/mo
Amazon Bedrock
serves 26
No subscription of its own
Ambient
serves 3
No subscription of its own
Anthropicopaque
6 models · serves 9
Claude Pro / Max 5x / Max 20x · from $20/mo
Arcee AI
serves 1
No subscription of its own
Arcee Ai
1 model
No subscription of its own
AtlasCloudrequests
serves 20
Coding Plan (Starter / Lite / Plus / Max / Ultra / Enterprise) · from $10/mo
Augment Codedollars
provider
Augment (Business / Enterprise) — Auggie CLI + Cosmos · from $100/mo
Azure
serves 42
No subscription of its own
Baidu
serves 6
No subscription of its own
BaseTen
serves 9
No subscription of its own
Cerebras
serves 3
No subscription of its own
Chutesopaque
serves 6
Plus / Pro · from $10/mo
Claude Platform on AWS
serves 1
No subscription of its own
Cloudflare
serves 8
No subscription of its own
Cohere
2 models · serves 2
No subscription of its own
CoreWeave
serves 13
No subscription of its own
Crusoe
serves 6
No subscription of its own
Cursordollars
provider
Pro / Pro+ / Ultra · from $20/mo
Decart
serves 3
No subscription of its own
DeepInfra
serves 49
No subscription of its own
DeepSeek
7 models · serves 2
No subscription of its own
DigitalOcean
serves 12
No subscription of its own
Fireworks
serves 10
No subscription of its own
Friendli
serves 4
No subscription of its own
GMICloud
serves 11
No subscription of its own
GitHubdollars
provider
Copilot (Pro / Pro+ / Max) · from $0/mo
Googleopaque
12 models · serves 48
AI Plus / Pro / Ultra · from $5/mo
Google AI Studioopaque
serves 18
AI Plus / Pro / Ultra · from $5/mo
Groq
serves 5
No subscription of its own
IBM
1 model
No subscription of its own
Inception
1 model · serves 1
No subscription of its own
Inceptron
serves 4
No subscription of its own
InclusionAI
3 models
No subscription of its own
Io Netrequests
serves 3
IO Intelligence (Standard / Professional / Developer) · from $0/mo
Kwaipilot
1 model
No subscription of its own
Mancer 2
serves 3
No subscription of its own
Mara
serves 2
No subscription of its own
Meituan
1 model
No subscription of its own
Meta
5 models · serves 2
No subscription of its own
MiniMaxopaque
2 models · serves 3
Token Plan (Plus / Max / Ultra) · from $20/mo
Mistralopaque
7 models · serves 4
Le Chat Free / Pro / Team (Mistral Vibe bundled) · from $0/mo
Modal
serves 1
No subscription of its own
ModelRun
serves 3
No subscription of its own
Moonshotopaque
4 models · serves 5
Moderato / Allegretto / Allegro / Vivace · from $19/mo
Moonshot / Kimiopaque
provider
Moderato / Allegretto / Allegro / Vivace · from $19/mo
Morphdollars
serves 7
Usage-based (Free / Pay-per-token / Scale flat-rate) · from $0/mo
Nebius
serves 5
No subscription of its own
Nex AGI
1 model · serves 1
No subscription of its own
NextBit
serves 2
No subscription of its own
Novita
serves 33
No subscription of its own
Nvidia
5 models · serves 1
No subscription of its own
Ollamaopaque
provider
Ollama Cloud (Free / Pro / Max) · from $0/mo
OpenAIopaque
21 models · serves 35
ChatGPT Plus / Pro 5x / Pro 20x · from $0/mo
OpenInference
serves 2
No subscription of its own
Parasail
serves 20
No subscription of its own
Phala
serves 15
No subscription of its own
Qoderdollars
provider
Qoder (Free / Pro / Pro+ / Ultra) · from $0/mo
Relace
serves 1
No subscription of its own
Sail Research
serves 4
No subscription of its own
SambaNova
serves 4
No subscription of its own
SiliconFlow
serves 21
No subscription of its own
StepFun
1 model · serves 1
No subscription of its own
StreamLake
serves 14
No subscription of its own
Syntheticdollars
provider
Synthetic (Standard) · from $30/mo
Tencent
1 model · serves 1
No subscription of its own
Thinking Machines
1 model
No subscription of its own
Together
serves 15
No subscription of its own
Upstage
2 models · serves 2
No subscription of its own
Venicedollars
serves 21
Venice (Free / Pro / Pro Plus / Max) · from $0/mo
Wafer
serves 5
No subscription of its own
Windsurf (Cognition/Devin)opaque
provider
Pro / Max / Teams · from $0/mo
Xiaomitokens
2 models · serves 2
MiMo Token Plan (Lite / Standard / Pro / Max) · from $6/mo
Z.ai / Zhipurequests
4 models · serves 4
GLM Coding (Lite / Pro / Max) · from $18/mo
opencodedollars
provider
opencode Go · from $10/mo
xAI
4 models · serves 16
No subscription of its own

Gateways, resellers & inference hosts

Not the lab, and not the ranking

The places that sell a model without being the lab that made it, kept out of the leaderboard and the first-party price columns on purpose — because the distinction they blur is the useful part. An inference host serves open-weight models on its own GPUs, so its price is a real price for that model. A router forwards to providers with little markup. A reseller proxies someone else’s closed API (GPT, Claude) and reprices it — its number is the vendor’s minus a cut, and its “X% cheaper” is a claim against a list price, only meaningful if you can verify the effective rate before paying. The reseller/router class is detailed below; the full roster of every inference host we track (StreamLake, Together, Groq, SambaNova & 39 more) is at the end of the section.

ServiceOriginTypeServesBilling unitConverts to tokens?Verifiable pre-pay?Note
CometAPI🇭🇰 Hong Kong SARresellerGPT-5.6, Claude Sonnet 5 / Fable 5 / Opus 4.8, Gemini 3.6/3.5, DeepSeek V4 Pro, Kimi K3 (500+ models)tokens (official models) + per-image/clip for mediapartialpartialThe most transparent of the resellers here: the 0.8x formula is stated outright, so the effective token price is checkable against the vendor's own list. Still a proxy of closed APIs, so the price is theirs minus 20%, not a first-party price, and no monthly plan.
AIMLAPI🇪🇪 EstoniaresellerGPT-5.6 Sol, DeepSeek V4 Flash, Gemini 3.6 Flash, coder models (600+); Claude not listedtokens (per 1M)fullpartialPrices per million tokens directly, so no credit layer to decode — but it publishes no discount claim, and Anthropic models are absent, so it is not a full first-party substitute.
APIMart🇭🇰 Hong Kong SARresellerGPT-5, Claude Sonnet 4.5, DeepSeek V3.2, Gemini + heavy image/video (Sora, Midjourney, Flux)creditsnonefalseOpenAI-compatible (base api.apimart.ai/v1), but billed in credits with no published credit-to-token conversion, so the advertised saving cannot be verified before paying. This is the entry that prompted the section.
EURouter🇳🇱 NetherlandsrouterClaude, GPT, Mistral, DeepSeek, Qwen, Codestral (100+); EU data residencytokens + published markupfullyesA Netherlands-based router selling on EU data residency and GDPR rather than price. Notable as the transparent opposite of the credit resellers: it publishes its exact markup per tier and bills real tokens, so the effective price is checkable up front. Base www.eurouter.ai/api/v1.
ShareAIinference-hostopen-weight models only (Llama 4, GPT-OSS, Qwen, 150+) on a decentralized GPU gridtokensfullpartialNot a closed-API reseller but a decentralized inference marketplace: idle GPUs serve open weights, so its prices are real prices for those models. No frontier closed models (no GPT/Claude), so limited for coding against the leaders. Operator HeyShare SRL; country not confirmed.
APIMasterresellerOpenAI, Claude, DeepSeek via "fingerprint-verified channels"; works with Cursor / Claude Codetokens (shared balance across channels)nonefalseOpenAI-compatible, but the headline discounts are far beyond a normal reseller margin and no markup is published, only per-"channel" prices that vary. Discounts that large on a closed API usually mean grey-market or quota-farmed capacity, not a straight resale. Company location undisclosed. Treat the savings claim with caution.
TokenMix🇭🇰 Hong Kong SARresellerClaude Opus 4.7, GPT-5, DeepSeek V4, Qwen, Gemini, Llama 4 (multi-provider)tokens (per-model)partialpartialToken Limited (Hong Kong). Pay-per-token with per-model pricing and an OpenAI-compatible base (api.tokenmix.ai/v1). No markup formula published, but it bills tokens rather than opaque credits. $1 minimum top-up.
CrofAIinference-hostopen-weight coding models (Kimi, GLM, DeepSeek, Qwen, MiniMax, Gemma) at lowest ratetokens (pay-per-token)fullyesCheap OpenAI-compatible inference (crof.ai/v1) for open weights only, no markup, pure pay-per-token -- its own page now states \u201cno subscriptions, no minimums, only the tokens you actually burn\u201d (the tiered Hobby/Max plans some third-party listings still show are gone). One of the resellers this repo excluded from first-party price resolution earlier (as \u201ccrof\u201d); it belongs here, not in the leaderboard prices. Country not confirmed. checked 2026-07-22.
nano-gptrouterGPT, Claude, DeepSeek, Qwen + open models (1000+), OpenAI-compatibletokens (list price)fullyesA no-markup, privacy-first router: it passes provider list prices straight through (GPT-5.5 at $5/$30, same as OpenAI) with no deposit fee, and takes crypto from $0.10 with no KYC. The honest end of the spectrum -- the price is the vendor\u2019s, verifiable up front -- sold on privacy rather than a discount. Country not confirmed.
Fal.ai · Replicate · Wavespeed AI · PiAPI · EachLabsresellergenerative image / video / audio (Midjourney, Flux, Kling, etc.)per-second / per-run / creditsn/an/aNamed by APIMart as its competitors. All are media-generation gateways, not coding-LLM providers, so they are recorded here only to say they were checked and are out of scope — not to rank.
Featherless AI🇸🇬 Singaporeinference-hostopen-weight models only (40k+ pulled from Hugging Face) — DeepSeek V4 Pro/Flash, GLM-5.2, Qwen3, Kimi K2.x, Llamaflat monthly subscription by concurrency (not tokens)nonepartialServerless host of open weights (Singapore-founded by the RWKV team, AMD/Airbus-backed). Unusual model: you buy concurrency, not tokens — Premium/Chat $25/mo (4 units, 32K context), Agent $100/$200 (8 units, 256K). Overflow returns HTTP 429, no overage. It does serve DeepSeek V4 and GLM-5.2, but there is no per-token price to compare and no quantization disclosed, so cost per unit of real work is not derivable up front. A Feather Per-Request credit plan exists but its rate card is unpublished.
Nscale🇬🇧 United Kingdominference-hostopen-weight serverless — gpt-oss, Qwen2.5-Coder / Qwen3, Llama 4, Devstral, DeepSeek R1 distills (no full DeepSeek V4, no GLM, no Kimi)tokens (per 1M, pay-as-you-go)fullpartialUK-domiciled European "sovereign" AI cloud (London HQ; data centres in Norway / UK / Iceland). Two products: a GPU-rental hyperscaler (per-hour, its main business) and a smaller per-token serverless API. The serverless roster is mid-tier for coding — Qwen2.5-Coder, Devstral and gpt-oss — but the current open flagships (DeepSeek V4, GLM-5.2, Kimi) are absent, and DeepSeek is offered only as R1 distills. Quantization not disclosed; the full catalogue sits behind a signup gate.
OVHcloud AI Endpoints🇫🇷 Franceinference-hostopen-weight, EU-sovereign — Llama, Mistral (incl. Codestral / Small 3.2), Qwen3 Coder, gpt-oss, DeepSeek R1 distill (no DeepSeek V4, no GLM, no Kimi)tokens (per 1M, priced in EUR)fullyesFrance-based (OVHcloud, Roubaix), GDPR / EU data residency, served from Gravelines; states customer data is never used to train models. Open per-token catalogue in EUR (from ~€0.04/M). Relevant coding options are Qwen3 Coder 30B and gpt-oss-120b — the current open coding SOTA (DeepSeek V4, GLM-5.2, Kimi) is not onboarded. Prices are EUR, so the USD-normalised cost floats with FX. Quantization not disclosed.
Scaleway Generative APIs🇫🇷 Franceinference-hostopen-weight, EU-sovereign — GLM-5.2, Qwen3 Coder, Mistral (Devstral 2, Large 3, Small 3.2), Llama, gpt-oss, MiniMax-M2.5 (DeepSeek R1 distills only, no DeepSeek V4, no Kimi)tokens (per 1M, priced in EUR)fullyesFrance-based (Scaleway, Iliad group), served from Paris only; GDPR / EU residency and states it does not read or reuse prompt/output content. The most relevant of the HF-partner hosts for coding: it serves GLM-5.2 and Qwen3 Coder and — unusually — DISCLOSES quantization openly via model-id suffixes (:fp8, :bf16, :fp4, :int4, :awq). Prices are EUR (normalise to USD with FX). DeepSeek is only R1 distills, not V4. Single region (Paris), so no failover.
Public AIinference-hostpublic-good / sovereign open-weight models only — Apertus, SEA-LION, EuroLLM, Bielik, Olmo (no mainstream or coding models)free (donated compute)n/apartialA non-profit "public utility for AI" (Public AI Inference Utility), serving publicly-funded sovereign models on donated / partner clusters, also routed free via Hugging Face. Currently free with a ~20 req/min limit, but "free" is explicitly provisional and there is no price sheet or SLA. The roster is sovereign / multilingual general-instruct models (Apertus, SEA-LION, EuroLLM) with no DeepSeek / GLM / Qwen-base / Llama / gpt-oss and no dedicated coding model, so it is recorded for completeness, not as a coding option. Legal HQ not stated (Swiss / EU centre of gravity).
Hugging Face Inference Providers🇺🇸 United Statesrouteropen-weight models across partner hosts (Cerebras, Groq, Together, Fireworks, Novita, DeepInfra, Scaleway, Z.ai …) through one HF tokentokens (pass-through provider rate)fullyesA router, not a host: one HF token routes to leading inference providers, OpenAI-compatible (router.huggingface.co/v1), with no HF markup — you pay the partner's rate. The provider is chosen by policy suffix (:fastest default, :cheapest, :preferred, or an explicit :provider). A small monthly credit ($0.10 free / $2 on PRO at $9/mo) then pay-as-you-go. It serves whatever its partners serve (open weights only), so price and quantization come from the chosen partner, not HF. Closed models (GPT-5.x, Claude) are not routed.
NVIDIA Build🇺🇸 United Statesinference-host100+ open-weight models via an OpenAI-compatible API on NVIDIA's own GPUs — Nemotron, Llama, Qwen, gpt-oss, DeepSeek and GLM (Zhipu) familiesfree developer credits (production via NIM / per-token)fullpartialNVIDIA's hosted API catalogue (build.nvidia.com), OpenAI-compatible — a genuinely generous free developer tier over 100+ open models on NVIDIA's own hardware, no credit card. It is not a per-token storefront: the hosted catalogue is free for prototyping (rate-limited ~40 RPM), and production means self-hosting the NIM microservice under NVIDIA AI Enterprise (~$4,500/GPU/year, with a 90-day free eval) or a quoted per-token range of ~$0.10-$10/M. Serves recent open models — strong on its own Nemotron line — but not always the very latest frontier (newest DeepSeek/GLM). US company.
RunPodinference-hostPublic Endpoints — a LIMITED set of pre-deployed OpenAI-compatible models, image/video/audio-first (Flux, Qwen Image, WAN 2.5, Kling, SORA 2) with only a few text/code models (Qwen3 32B, IBM Granite, Moonshot Kimi). Any other open weight you self-deploy on Serverless (per-second GPU) or Pods (per-hour GPU).tokensfullpartialPrimarily a GPU-rental / serverless platform, but it DOES expose an OpenAI-compatible inference API via Public Endpoints (api.runpod.ai/v2/...), so a coding agent can point at it. Caveats for coding: the managed text catalogue is thin — Qwen3 32B is the main code-capable one, at a steep $10 per 1M tokens — and the platform is image/video-first. It is NOT an OpenRouter provider (no routed named-model prices); for arbitrary weights you self-deploy and pay GPU-time (Serverless per-second, Pods per-hour), not per-named-token. Sits between raw GPU rental (Vast/Lambda) and a full per-token host (DeepInfra/Together).

The open-weight inference hosts and the router are already priced per model in the endpoint tables above (the “where to buy” view per provider). This table exists for the reseller class, which those tables do not capture because a reseller’s repriced number is not a first-party price. Media-only gateways are listed to record that they were checked and are out of a coding site’s scope. Data in gateways.yaml.

What common models cost through every host

Input / output per million tokens for five reference models, against the vendor’s own list price — the resellers at the top (a fixed set we researched), then every OpenRouter host that serves them, with its real price and the precision it runs at. Two findings. “Gateway” does not mean “cheaper”: AIMLAPI and EURouter charge more than official, only CometAPI/TokenMix undercut it, and the credit/channel resellers can’t be placed on the scale at all. And for the open models the price is only half the story — the other half is quantization: the closed models (GPT-5.6, Claude) are first-party only, so hosts show “—”, while a cheaper DeepSeek here may be fp4, not fp8. Verified 2026-08-14.

ServiceDeepSeek V4 Pro · fp8 nativeDeepSeek V4 Flash · fp8 nativeGLM-5.2 · fp8 nativeGPT-5.6 SolClaude Sonnet 5Basis
Official (first-party)$0.435 / $0.87$0.14 / $0.28$1.4 / $4.4$5 / $30$2 / $10the list price
CometAPI$0.35 / $0.70$0.11 / $0.22$1.12 / $3.52$4 / $24$1.60 / $80.8x official; quant undisclosed
AIMLAPInot listed$0.18 / $0.36$6.50 / $39not servedper-model list; dearer on GPT, no Claude
APIMartopaqueopaqueopaqueopaqueopaquecredits, no conversion
EURouter+3-15%+3-15%+3-15%+3-15%+3-15%markup by tier on official
ShareAIopen, per-tokenopen, per-tokenopen, per-tokennot servednot servedopen weights; quant undisclosed
APIMasteropaqueopaqueopaqueopaqueopaque"channels", up to 90% off claimed
TokenMix$0.41 / $0.83$0.13 / $0.26$2.24 / $7.82$4.75 / $28.50Fable 5 onlyper-model; GLM (Fast Preview) dearer, quant undisclosed
CrofAI$0.35 / $0.80 · Q8_0$0.12 / $0.21 · Q4_0$0.30 / $1.05 · Q8_0not servednot servedopen weights, per-token; publishes quant (Q8_0 ≈ 8-bit, Q4_0 ≈ 4-bit)
nano-gpt$0.435 / $0.87$0.14 / $0.28$1.4 / $4.4$5 / $30$2 / $10list price, no markup; routes so quant varies
Featherless AIserved · flat sub, no per-tokenserved · flat sub, no per-tokenserved · flat sub, no per-tokennot servednot served$25-200/mo by concurrency, unlimited tokens; quant undisclosed
Nscalenot servednot served (R1 distills only)not servednot servednot servedopen-weight per-token; Qwen2.5-Coder 32B $0.06/$0.20, gpt-oss-120b $0.10/$0.40, Devstral Small $0.10/$0.30
OVHcloud AI Endpointsnot servednot served (R1 distill only)not servednot servednot servedopen-weight per-token EUR; Qwen3 Coder 30B ~€0.07/€0.26, gpt-oss-120b €0.08/€0.40, Mistral Small 3.2 €0.09/€0.28
Scaleway Generative APIsnot servednot served (R1 distills only)€1.80 / €5.50 · fp8not servednot servedopen-weight per-token EUR, quant disclosed; GLM-5.2 €1.80/€5.50, Qwen3 Coder 30B €0.20/€0.80, Mistral Small 3.2 €0.15/€0.35
Public AInot servednot servednot servednot servednot servedfree sovereign open models; no coding model, quant undisclosed
Hugging Face Inference Providerspartner rate (pass-through, if a partner serves it)partner rate (pass-through, if a partner serves it)partner rate (pass-through, e.g. via Scaleway / Z.ai)not served (open weights only)not served (open weights only)pass-through partner price, no markup; +$0.10/$2 monthly credit
NVIDIA Buildnot served (older DeepSeek gen)not servednot served (GLM-4 family)not servednot servedfree dev tier (credits, ~40 RPM); production via NIM / AI Enterprise
RunPodnot served (self-deploy on Serverless/Pods)not served (self-deploy)not served (self-deploy)not served (closed model)not served (closed model)Public Endpoint text — Qwen3 32B $10.00/1M tokens; image Flux Dev $0.02/megapixel; video WAN 2.5 $0.50/5s. Self-deploy = GPU-time (Pods per-hour, Serverless per-second).
AkashML$0.14 / $0.28fp8$0.77 / $2.42fp8OpenRouter · live
Alibaba / Qwen$0.97 / $3.04fp8OpenRouter · live
Amazon Bedrock$5.50 / $33.00?$2.00 / $10.00?OpenRouter · live
Ambient$0.14 / $0.28fp4$1.05 / $4.40fp8OpenRouter · live
Anthropic$2.00 / $10.00?OpenRouter · live
AtlasCloud$0.14 / $0.28fp4$1.26 / $3.96fp8OpenRouter · live
Azure$5.00 / $30.00?$2.00 / $10.00?OpenRouter · live
Baidu$0.08 / $0.16fp8$0.49 / $1.54fp8OpenRouter · live
BaseTen$1.32 / $3.96fp4$0.13 / $0.26fp8$1.40 / $4.40fp8OpenRouter · live
Cloudflare$1.32 / $3.96?$0.14 / $0.28?$1.40 / $4.40?OpenRouter · live
CoreWeave$0.13 / $0.28fp8$0.76 / $2.42fp4OpenRouter · live
Crusoe$1.40 / $4.40fp8OpenRouter · live
Decart$0.08 / $0.16fp4$0.77 / $2.56fp4OpenRouter · live
DeepInfra$0.08 / $0.18fp4$0.75 / $2.40fp4OpenRouter · live
DeepSeek$0.43 / $0.87?$0.14 / $0.28fp8OpenRouter · live
DigitalOcean$0.08 / $0.25?$0.63 / $1.98?OpenRouter · live
Fireworks$1.32 / $3.96?$0.14 / $0.28?$1.40 / $4.40?OpenRouter · live
Friendli$1.40 / $4.40?OpenRouter · live
GMICloud$1.22 / $2.44fp8$0.08 / $0.17fp8$0.74 / $2.33fp8OpenRouter · live
Google$2.00 / $10.00?OpenRouter · live
Inceptron$0.13 / $0.28fp4$0.75 / $2.90fp4OpenRouter · live
Io Net$0.15 / $0.32fp8OpenRouter · live
Mancer 2$0.17 / $0.50fp8OpenRouter · live
Morph$0.14 / $0.28bf16$1.10 / $4.10fp4OpenRouter · live
Novita$1.32 / $3.96fp8$0.14 / $0.28fp8$0.74 / $2.33fp8OpenRouter · live
OpenAI$2.50 / $15.00?OpenRouter · live
OpenInference$0.08 / $0.18fp4OpenRouter · live
Parasail$0.14 / $0.28fp8$1.40 / $4.40fp4OpenRouter · live
Phala$0.20 / $0.40?$1.13 / $3.00fp8OpenRouter · live
Relace$0.14 / $0.28fp4OpenRouter · live
Sail Research$0.09 / $0.18fp4$0.50 / $3.15fp8OpenRouter · live
SiliconFlow$1.32 / $3.96fp8$0.14 / $0.28fp8$1.19 / $3.74fp8OpenRouter · live
StreamLake$0.09 / $0.18fp8$0.75 / $2.37fp8OpenRouter · live
Together$0.14 / $0.28?$1.40 / $4.40?OpenRouter · live
Venice$0.17 / $0.35?$1.40 / $4.40fp8OpenRouter · live
Wafer$0.28 / $0.56?$1.26 / $3.96fp4OpenRouter · live
Z.ai / Zhipu$1.40 / $4.40fp8OpenRouter · live

Inference hosts & serving providers (44)

Every provider that serves a tracked model’s endpoint — open-weight hosts, cloud platforms, and labs serving their own models. These are priced per model in the “where to buy” tables on each provider page (from OpenRouter’s per-endpoint data, with quantization), which is why they are a roster here rather than a price grid: a host’s price depends on the specific model and precision, not a single rate. Click through for the per-model breakdown.

ProviderOriginSite
AkashML🇺🇸 United Statesakashml.com
Alibaba🇨🇳 Chinawww.alibabacloud.com/product/machine-learning
Amazon Bedrock🇺🇸 United Statesaws.amazon.com/bedrock/
Ambientwww.ambient.xyz
Anthropic🇺🇸 United Stateswww.anthropic.com
AtlasCloudwww.atlascloud.ai
Azure🇺🇸 United Statesazure.microsoft.com/products/ai-foundry
Baidu🇨🇳 Chinacloud.baidu.com
BaseTen🇺🇸 United Stateswww.baseten.co
Cerebras🇺🇸 United Stateswww.cerebras.ai
Chuteschutes.ai
Cloudflare🇺🇸 United Statesdevelopers.cloudflare.com/workers-ai/
Crusoe🇺🇸 United Statescrusoe.ai
Decart🇮🇱 Israelwww.decart.ai
DeepInfra🇺🇸 United Statesdeepinfra.com
DeepSeek🇨🇳 Chinawww.deepseek.com
DigitalOcean🇺🇸 United Stateswww.digitalocean.com/products/gradient
Fireworks🇺🇸 United Statesfireworks.ai
Friendli🇰🇷 South Koreafriendli.ai
GMICloud🇺🇸 United Stateswww.gmicloud.ai
Google🇺🇸 United Statescloud.google.com/vertex-ai
Groq🇺🇸 United Statesgroq.com
Inceptronwww.inceptron.io
Io Net🇺🇸 United Statesio.net
Ionstreamionstream.ai
Minimax🇨🇳 Chinawww.minimax.io
Mistral🇫🇷 Francemistral.ai
Moonshot AIwww.moonshot.ai
Morph🇺🇸 United Statesmorphllm.com
Nebius🇳🇱 Netherlandsnebius.com
Novitanovita.ai
Parasail🇺🇸 United Stateswww.parasail.io
Phalaphala.network
RunPodwww.runpod.io
SambaNova🇺🇸 United Statessambanova.ai
SiliconFlow🇨🇳 Chinawww.siliconflow.com
StreamLake🇨🇳 Chinawww.streamlake.ai
Together🇺🇸 United Stateswww.together.ai
Venice🇺🇸 United Statesvenice.ai
Waferwafer.ai
WandB🇺🇸 United Stateswandb.ai
xAI🇺🇸 United Statesx.ai
Xiaomi🇨🇳 Chinaplatform.xiaomimimo.com
Z.AI🇨🇳 Chinaz.ai

Pick of the week

What I actually run
2026-08-14

What I actually run at home

Qwen3.8-27B is here, but at Q3 on 16 GB it is not my upgrade yet · staying on the old Qwen until I get a bigger card

Last week I said we were not yet at productive local coding on the hardware most of us have, and this week put it to the test. Qwen3.8-27B dropped, a dense 27B under Apache-2.0, and the exciting part is speed: with MTP speculative decoding through llama-server it runs at 36 tok/s on my 16 GB card, twice what it does without spec-decode, and finally interactive. But here is the honest catch. To fit 16 GB you run it at Q3 (about 14 GB), and at Q3 it scores 27% on our BCB-Hard, which is actually below what I already run locally: the ternary Qwen3.6-27B rebuild is at 35%, and Gemma ranks higher still. The full model at good precision is 37 to 43%, so the 3-bit quant that fits my card costs about 10 points, and that gap is the whole story. So I am not switching yet. I will keep running my old Qwen until I have a card that fits Q4 or Q6, where the 27B actually pulls ahead. That is a spend not everyone can make. If you can, the dense 27B at Q4 or Q6 is what to put on the new card. If you cannot, the real hope is not this dense model but the small A3B MoE Qwen will ship next, which runs fast on the 16 GB you already own. Until one of those lands, I am staying on my old Qwen.

What to run locally

Your GPU, not a bill

VRAM figures are real GGUF bytes and coding numbers are our own first-party runs where the AA index has none. The catalogue keeps growing; treat it as a solid guide, not gospel, and corrections are welcome.

The best coding models that fit each VRAM budget — cumulative, because a bigger card also runs everything smaller. Sizes are real GGUF bytes at ~4-bit (measured where a build exists), plus the KV cache the context needs — weights alone are not the VRAM figure the review videos quote. Ranked on one unified scale so the pick is realistic at every budget: our measured BigCodeBench-Hard (shown as % BCB) for the models we actually ran on this GPU, and the Artificial-Analysis coding index (shown as idx) for the big ones we can’t run locally — mapped onto the same axis by our measured fit (r = 0.86). So a multi-GPU tier correctly points at a GLM-5.2 or Kimi, while at 16 GB our own tests put gpt-oss-20b and the ternary Bonsai 27B on top — the very coders the index scores low or not at all. Reasoning models that run out of token budget locally are ranked on that measured floor, not their optimistic index. Quality still isn’t proportional to size — a well-built 12B beats a 20B. Each model links to how to run it.

≤8 GB

#ModelCoderParamsWeights ~Q4Min VRAM @8KType
1Gemma 4 12B36% BCB12B7 GB8 GBdense
2Ternary Bonsai 27B36% BCB27B7 GB8 GBdense
3Granite 4.1 8B17% BCB8B5 GB7 GBdense

≤16 GB

Best pick unchanged from ≤8 GB: Gemma 4 12B still leads and needs only 8 GB. More VRAM here buys headroom and context, not a better coder — the larger open models score lower.

#ModelCoderParamsWeights ~Q4Min VRAM @8KType
1Gemma 4 12B36% BCB12B7 GB8 GBdense
2Ternary Bonsai 27B36% BCB27B7 GB8 GBdense
3Qwen3 14B31% BCB14B9 GB11 GBdense

≤32 GB

#ModelCoderParamsWeights ~Q4Min VRAM @8KType
1Qwen3.6 27Bidx 53.727B17 GB19 GBdense
2Gemma 4 26B A4Bidx 39.326B17 GB19 GBdense
3Nemotron Cascade 2 30B A3Bidx 25.332B25 GB28 GBMoE

≤80 GB

Best pick unchanged from ≤32 GB: Qwen3.6 27B still leads and needs only 19 GB. More VRAM here buys headroom and context, not a better coder — the larger open models score lower.

#ModelCoderParamsWeights ~Q4Min VRAM @8KType
1Qwen3.6 27Bidx 53.727B17 GB19 GBdense
2Gemma 4 26B A4Bidx 39.326B17 GB19 GBdense
3gpt-oss-120bidx 30.4117B63 GB70 GBMoE

Data centre

#ModelCoderParamsWeights ~Q4Min VRAM @8KType
1DeepSeek V4 Flashidx 69.1291B137 GB151 GBMoE
2GLM-5.2idx 68.8753B466 GB544 GBMoE
3DeepSeek V4 Proidx 68.8862B850 GB936 GBMoE

Every model, ranked

All 53 models in the catalogue by coding index. Filter the Fits or Type column to narrow to your hardware, or sort any column. HumanEval* and BCB* (BigCodeBench-Hard) are first-party pass@1 scores we ran ourselves — HumanEval on 12 models, the harder BCB-Hard on 23, on the real GPU. Note how HumanEval saturates near the top while BCB-Hard spreads the field out — that gap is why we run both. A dimmed ~NN% est in the BCB column is an estimate from the coding index for a model we haven’t run — only shown where the index is inside the range we fitted (r = 0.86), never extrapolated wildly. t/s is the measured local decode speed on an RTX 4060 Ti. Click a measured % for how it was run.

ModelCoding idxHumanEval*BCB*t/sParamsWeights ~Q4Min VRAMFitsType
DeepSeek V4 Flash69.1291B137 GB151 GBData centreMoE
GLM-5.268.8753B466 GB544 GBData centreMoE
DeepSeek V4 Pro68.8862B850 GB936 GBData centreMoE
Motif-3 Beta63.5315B196 GB221 GBData centreMoE
Kimi K2.661.81027B584 GB642 GBData centredense
Kimi K2.7 Code60.81027B495 GB545 GBData centredense
MiMo-V2.5-Pro60.21023B630 GB696 GBData centreMoE
Nex-N2-Pro59.1397B242 GB266 GBData centredense
Hy358.8299B182 GB203 GBData centreMoE
GLM-5.155.8754B465 GB521 GBData centreMoE
Qwen3.6 27B53.727B17 GB19 GB≤32 GBdense
Inkling52.9952B571 GB est629 GBData centredense
Nemotron 3 Ultra 550B49.3550B359 GB395 GBData centreMoE
Qwen3.5 397B A17B48.2397B488 GB537 GBData centredense
Mistral Medium 3.546.9128B75 GB82 GBData centredense
Qwen3.5 122B A10B45.7125B77 GB84 GBData centredense
Solar Open2 250B44.7250B136 GB151 GBData centreMoE
Gemma 4 26B A4B39.326B17 GB19 GB≤32 GBdense
Nemotron 3 Super 120B37.7120B83 GB92 GBData centreMoE
gpt-oss-120b30.4117B63 GB70 GB≤80 GBMoE
Nemotron Cascade 2 30B A3B25.332B25 GB28 GB≤32 GBMoE
EXAONE 4.5 33B23.633B20 GB22 GB≤32 GBdense
Qwen3 235B A22B 250722.1235B142 GB158 GBData centreMoE
Gemma 4 31B43.443%31B18 GB20 GB≤32 GBdense
Qwen3-Coder 30B A3B98%42%40 t/s31B19 GB21 GB≤32 GBMoE
Qwen3 Next 80B A3B17.4~37% est80B49 GB54 GB≤80 GBMoE
Gemma 4 12B31.063%36%31 t/s12B7 GB8 GB≤8 GBdense
Ling-3.0-flash50.636%124B155 GB176 GBData centreMoE
Ternary Bonsai 27B90%36%30 t/s27B7 GB8 GB≤8 GBdense
Nemotron 3.5 Lightning26.835%30B18 GB20 GB≤32 GBMoE
Ornith-1.0 35B35%35B22 GB24 GB≤32 GBdense
North Mini Code 1.036.535%30B18 GB est21 GB≤32 GBMoE
Laguna XS.234%33B21 GB24 GB≤32 GBMoE
Qwen3.6 35B A3B41.933%35B22 GB24 GB≤32 GBdense
Qwen3 32B15.3~33% est33B20 GB24 GB≤32 GBdense
Qwen3 14B13.863%31%25 t/s14B9 GB11 GB≤16 GBdense
JetBrains Mellum 2 12B28%12B8 GB9 GB≤16 GBMoE
Qwen3.8-27B27%36 t/s27B14 GB15 GB≤16 GBdense
gpt-oss-20b20.797%26%61 t/s22B12 GB14 GB≤16 GBMoE
Llama 3.3 70B11.9~25% est70B43 GB47 GB≤80 GBdense
Mistral Small 3.2 24B12.583%25%11 t/s24B14 GB16 GB≤16 GBdense
Nemotron Nano 9B V222%9B7 GB9 GB≤16 GBdense
Granite 4.1 30B10.4~22% est29B17 GB21 GB≤32 GBdense
Granite 4.1 8B9.567%17%46 t/s8B5 GB7 GB≤8 GBdense
Nanbeige 4.2-3B81%17%9 t/s4B3 GB4 GB≤8 GBdense
Ministral 3 14B14.463%14%30 t/s14B8 GB9 GB≤16 GBdense
Ornith-1.0 9B12%9B6 GB6 GB≤8 GBdense
Qwen3.5 9B28.740%12%44 t/s9B6 GB6 GB≤8 GBdense
Qwen3.5 4B22.633%12%71 t/s5B3 GB3 GB≤8 GBdense
Bonsai 8B63%6%156 t/s8B1 GB2 GB≤8 GBdense
Instella MoE 16B A3B16B10 GB13 GB≤16 GBMoE
KAT-Coder V2.5 35B35B21 GB24 GB≤32 GBdense
Laguna S 2.1118B73 GB82 GBData centreMoE

Each tier is everything that fits that budget, so the same strong small model can top several tiers — that repetition is the point. “Min VRAM” is weights plus KV cache at an 8K context; longer context or a heavier quant (Q8/fp16) needs more. A dash under coding index is unmeasured, never zero. Data in local.yaml.

It’s free

beta$0 to start

Beta. New view — free tiers change constantly, so amounts are approximate and “free for now” is a real caveat. Verify before relying on any of it.

Every way we track to run a capable model for nothing — a real free allowance, not a trial that needs a card. Ranked roughly by how much you get. The catch column is the honest part: rate limits, one-off vs monthly, or “free for now”.

Free API credits & allowances

Hosted, OpenAI-compatible APIs with a genuine free tier over open (and some closed) models.

ProviderWhat you get freeThe catch
Public AIunlimited, for nowdonated compute, ~20 req/min, sovereign models only, may end
OpenRouter15 :free routes today — gemma-4-26b-a4b-it, gemma-4-31b-it, gpt-oss-20b, laguna-s-2.1, laguna-xs-2.1, lfm-2.5-2.6b, nemotron-3-nano-30b-a3b, nemotron-3-nano-omni-30b-a3b-reasoning, nemotron-3-super-120b-a12b, nemotron-3-ultra-550b-a55b, nemotron-3.5-content-safety, nemotron-3.5-lightning, nemotron-nano-12b-v2-vl, nemotron-nano-9b-v2, north-mini-code~50 requests/day (1,000 once you have topped up $10+), 20/min; the set rotates without notice — snapshot as of 2026-08-14
opencode ZenRotating free set (Big Pickle, DeepSeek V4 Flash, MiMo-V2.5, Hy3, Laguna S 2.1, Nemotron 3 Ultra, Nemotron 3.5 Lightning)free through the opencode CLI; the set rotates without notice — as of 2026-08-14
Scaleway Generative APIs1,000,000 tokensone-off for new accounts, then per-token EUR
Google AI Studio (Gemini API)~1,000–1,500 requests/day on Gemini 3 Flash / 2.5 Flashno card, but free-tier prompts may be used to improve Google's models; ~5–15 req/min; Pro models removed from the free tier (Apr 2026)
Groq~500K tokens/day per model + ~14,400 requests/dayno card; 30 req/min and 6K tokens/min caps; resets midnight UTC
Mistral La Plateforme“Experiment” free tier — ~1B tokens/montheval/prototyping only, not production; phone verification and a data-sharing opt-in
Cohere1,000 API calls/month (trial keys)non-production; ~20 req/min; data may be used to improve models
NVIDIA Build1,000–5,000 API credits~40 req/min, dev use; production needs NIM / AI Enterprise
Cloudflare Workers AI10,000 Neurons/day (Free and Paid plans alike)no card; resets 00:00 UTC; on the Free plan you are blocked until reset once it is spent
Nscale~$5 signup creditone-off; then per-token
OVHcloud AI Endpointssign-up credit + a few always-free modelsnew Public Cloud projects
Hugging Face Inference Providers$0.10/mo free · $2/mo on PROmonthly; pass-through provider rates after
CerebrasFree trial ($5 credits) over all models (gpt-oss-120b, GLM-4.7, Gemma 4 31B)no card for the trial; ~8K-token context cap; the widely-cited “sunset 2026-08-17” applies only to PREVIEW models in the paid Developer tier, not the free tier
SambaNova CloudPersistent free tier + $5 signup creditthe $5 credit expires in 3 months; per-model daily caps (some ~20 req/day)
Nebius Token FactoryNo-card free trial credit (~$1) over open models (Llama-3.3-70B, Qwen3-235B)rebranded from Nebius AI Studio in 2026; the free trial credit is now small (~$1), and a card is needed for the $25 minimum top-up

No longer free (don’t bother): Together (signup credit retired), Fireworks (~$1, testing only), Chutes (free tier ended Mar 15 2026), and GitHub Models (retired 2026-07-30). Listed so you don’t go looking.

Free forever (nothing to burn)

Run it yourself. Every open-weight model in Run local is free once you have the GPU — a 12B fits an 8 GB card and codes well, with no rate limit and no rotation. The rate-limited-free routes above (OpenRouter :free, opencode Zen) are also $0 with no expiry — the trade there is the provider’s rate cap and a set that rotates without notice.

Free subscription tiers

Products with a permanent $0 plan (limited, but real).

ProviderFree tier
GitHubFree
Io NetStandard
MistralFree
MorphFree
OllamaFree
OpenAIFree
QoderFree
VeniceFree
Windsurf (Cognition/Devin)Free

Data from gateways.yaml and plans.yaml. “Free for now” means exactly that — promotional tiers get cut (Tencent Hy3’s free window already closed), so treat any allowance as this week’s, not forever.

What changed

Since the snapshot of 2026-08-13.

  • GPT-5.6 Sol — provider count 6 → 7
  • Claude Opus 5 — provider count 9 → 10
  • GPT-5.6 Terra — provider count 6 → 7
  • Kimi K3 — cost per session $2.0189 → $2.0958
  • Kimi K3 — cache ratio 92.1% → 92.0%
  • GPT-5.5 — provider count 6 → 7
  • Claude Opus 4.8 — provider count 8 → 10
  • Grok 4.5 — cost per session $0.8088 → $0.8340
  • Grok 4.5 — cache ratio 86.8% → 86.7%
  • Nemotron 3.5 Lightning — context 1,048,576 → 1,000,000
  • Claude Sonnet 5 — provider count 8 → 9
  • GPT-5.6 Luna — provider count 6 → 7
  • GPT-5.4 — provider count 6 → 7
  • Gemini 3.6 Flash — input price $1.50 → $0.75
  • Gemini 3.6 Flash — output price $7.50 → $3.75
  • GLM-5.2 — input price $0.63 → $1.19
  • GLM-5.2 — output price $1.98 → $3.74
  • GLM-5.2 — cost per session $2.5186 → $2.5679
  • GLM-5.2 — cache ratio 80.9% → 80.6%
  • Qwen3.7 Max — cost per session $1.5437 → $1.5594
  • Claude Sonnet 4.6 — provider count 8 → 9
  • Kimi K2.6 — input price $0.58 → $0.56
  • Kimi K2.6 — output price $2.44 → $2.36
  • Kimi K2.6 — cost per session $0.7972 → $0.8180
  • Kimi K2.7 Code — input price $0.67 → $0.71
  • Kimi K2.7 Code — output price $3.40 → $3.50
  • Kimi K2.7 Code — cost per session $1.0773 → $1.0948
  • Kimi K2.7 Code — cache ratio 94.3% → 94.2%
  • Xiaomi MiMo-V2.5-Pro — cost per session $0.3798 → $0.3802
  • DeepSeek V4 Pro — input price $1.17 → $0.43
  • DeepSeek V4 Pro — output price $2.34 → $0.87
  • DeepSeek V4 Pro — coding index 59.4 → 68.8
  • DeepSeek V4 Pro — cost per session $0.4328 → $0.4384
  • DeepSeek V4 Pro — cache ratio 96.7% → 96.8%
  • DeepSeek V4 Pro — provider count 18 → 7
  • Tencent Hy3 — cost per session $0.3111 → $0.2951
  • MiniMax-M3 — cost per session $0.5314 → $0.5434
  • MiniMax-M3 — provider count 11 → 12
  • Xiaomi MiMo-V2.5 — cost per session $0.0698 → $0.0722
  • DeepSeek V4 Flash — input price $0.08 → $0.14

Historical data

The previous Model Speed Arena — weekly measured throughput runs against live subscriptions — is preserved at https://oldmsa.millaguie.net. Those runs are no longer updated: sustaining a weekly benchmark fleet was not viable. They remain the better source for measured tok/s under a real subscription, which no aggregator publishes.

Where the data comes from

Every figure on this page is somebody else’s measurement or publication. Nothing here is re-benchmarked locally, so each dataset is credited with what is taken from it and what that source can and cannot tell you.

SourceWhat this site takes from itNature
OpenRouterList price, context window, and the per-provider quantization, uptime and endpoint listpublic API
models.devOfficial provider catalog, tiered and cache pricingpublic API
Artificial AnalysisIndependent quality evaluations — coding index (which sets the ranking here), general intelligence index, TerminalBench-Hard and long-context reasoning — plus the only measured speed figures (tok/s, time to first token). The coding index is AA’s own weighted composite: we cannot recompute it, which is why a model AA has not tested cannot be rankedindependent lab
opencodeObserved production usage across its subscriber base: cost per session, tokens per session, cache ratioaggregate telemetry
quota-sentinelWhich providers expose a readable quota endpointopen source
Inference providersEvery serving provider links to its own site, each checked to resolve. Where a referral link exists it is used and marked, and the ranking is computed from price and precision alone regardlessneutral link
Vendor documentationSubscription prices and quotas, reproduced as advertised. Each plan row links to its own source and carries the date it was checkedvendor-reported
Reseller gatewaysNamed and classified in the Gateways section, never priced into the leaderboard: a reseller reprices someone else’s API, so its number is not a first-party pricereseller
Third-party write-upsUsed only where a vendor publishes no figure at all. Kimi’s per-tier request counts are the case in point: its own page gives adjectives, so the numbers come from independent write-ups instead and the plan is marked low-confidence. Every such plan links the write-up it came fromunofficial
Logged-in consoleFigures visible only inside an account, so a reader cannot click through to verify them: MiniMax’s remaining-quota percentage (from which the hidden quota here was derived), Google AI Plus’s euro price, and DeepSeek’s peak-pricing notice — which its public pricing page still does not mentionnot publicly verifiable
Measured hereA small number of first-party runs, only for models no independent lab has scored: HumanEval, BigCodeBench-Hard and SWE-bench Verified. Own scaffold, so comparable between these runs and never to a public leaderboard. Code and redacted trajectories are in the repositoryfirst-party

Two things worth carrying. Quality figures are tagged vendor or independent throughout, because the flattering numbers are almost always the vendor’s own. And observed-usage figures describe opencode’s subscribers, not the whole market — a model barely used there will have a noisy cost per session.

Method & limits

This site is an aggregator: price, context, per-provider quantization and uptime come from OpenRouter; official catalog pricing from models.dev; independent quality and speed from Artificial Analysis; observed production cost from opencode. Plan quotas were researched from vendors’ own documentation and are reproduced as advertised. With one exception, stated rather than hidden: a handful of models have no independent score anywhere, and for those a first-party run was done here (HumanEval, BigCodeBench-Hard, SWE-bench Verified) purely so the column is not blank. Those figures use our own scaffold, are labelled wherever they appear, and are never mixed into the leaderboard — which stays entirely Artificial Analysis’ coding index.

Two honest limits. Quota comparison is mostly impossible: only one provider publishes a quota that converts to tokens, and several publish nothing numeric at all. And vendor multipliers are not comparable — some are measured against a free tier, others against a paid one, some per session rather than per month. Plan pricing moves fast; every figure here is dated.

◆ Some links on this page are referral links: if you sign up through them this site may earn a commission, at no extra cost to you. Rankings, prices and measurements are taken from the sources listed under Method, and are not influenced by whether a provider has a referral programme.