Model Spend Arena2ND ED.
131 MODELS · 24 LABS · 69 HOSTS · 26 PLANS · 19 GATEWAYSUPDATED 2026-10-03
Coding-model intelligence · Updated 2026-10-03

The honest map of what LLM coding actually costs.

Metered per token or bought as a subscription — price, provider precision and published quality, aggregated, not re-benchmarked. Everything below is sourced and dated, including the column no other comparison shows: what numeric precision you are actually buying.

Best value · idx per $
2Ling-3.0-flash638 idx/$
Best value · top 20% (idx ≥ 39)
Cheapest / session
Quality ceiling · intelligence
Cheapest open-weight · avg
Best plan deal · tok/$ @ $0.79/M

Editor’s picks

Hand-written · 15 on record

Written from the data alone; referral relationships carry no weight in what is picked or how it is described. Each pick keeps the date it was made and is not quietly rewritten later, so you can judge how well it aged.

Weekly notes

betathe run, in plain words

A short write-up each week on what changed and why — the same data as the tables, told as a story. Newest first.

2026-10-02

A new benchmark arrived, the old one stopped, and we fitted the bridge between them before throwing it away

What happened. Last week AA’s API had no Terminal-Bench 4.0 field. This week 141 models have a measured score on it, including models released three days ago. Meanwhile the coding index this site ranked on still holds 258 models whose newest release date is 11 September: AA computed it from Terminal-Bench 2.1 and no longer runs that. Two rulers, and neither covers the table — the new one reaches 60 of our 131 rows, the old one 114 but nothing recent.

What we changed. The rank now follows AA’s intelligence index, which covers all 131 rows. Both coding columns stay, measured only, side by side. This is not a retreat from coding: of the 50 models that carry both the new pass rate and the old index, the intelligence order agrees with Terminal-Bench 4.0 at Spearman 0.95, and the old coding index only at 0.93. The column we used to rank on is the worse predictor of the benchmark that replaced it.

Why it is the worse predictor: it saturates. Eight points of old index, from 74 to 82, hold real Terminal-Bench 4.0 scores from 12.6% to 59.1%. Fable 5.1 at 81.6 scores 52%; Kimi K3 at 76.2 scores 12.6%; GPT-5.6 Sol at 77.4 scores 39.9% while GPT-6 Astra at 76.9 scores 59.1%. The old index called those last two almost equal.

The bridge we fitted and discarded. The tempting move is to convert: 50 models have both scores, so map one onto the other and give every unmeasured row an estimate. Excluding the saturated top, the fit is tight — 1.3 points of average error below 68 index points, and 62 of the 65 unmeasured rows live in exactly that range. We discarded it for three reasons. A one-variable monotone fit preserves order, so it produces the ranking we already had and adds a number, not information. Its entire output range is 0% to 3.2%, and the benchmark is 66 tasks attempted three times each, so every value it can emit sits inside six solved attempts — about two tasks — and it cannot order the 65 rows it was built for. And it extrapolates into the gap in the wrong direction: Claude Opus 4.7, old index 73.6, comes out at 3.2% while the two measured models nearest to it on that index, Opus 4.8 at 74.3 and GLM-5.3 at 74.8, score 21.7% and 41.9%. We publish estimates elsewhere on this site and label them; we are not willing to put one in the same column as a measurement.

A trap in the API, for anyone building on it. terminalbench_v4_0 returns a plain 0 for a model AA has not run, while using real nulls elsewhere in the same record. Sixty-seven models carry that zero, Apodex 1.1 among them — it passes 69.7% of Terminal-Bench 2.1. Thirty-nine of our own 131 rows carry it, so treating those zeros as scores would have dropped almost a third of the table to the bottom. We hit the identical placeholder on their speed field last month, so we check both now.

AA also renamed almost everything. DeepSeek’s rows dropped “(Reasoning, Max Effort)” for “(Max)”; all nine Claude rows became “(Max)” or “(Max, Default Fallback)”, the second rename landing while we were building; and 123 more changed capitalisation only, (max) to (Max). Each rename silently empties a row matched by name. Under all that churn the list did not shrink: it went from 673 rows to 689, and every name that disappeared has a renamed counterpart. We added a guard that prints a loud warning when a tracked row stops matching any name in AA’s list, and we read it on every pass; it caught both renames. It does not stop the build — the row still publishes, without its numbers — so the warning is the whole mechanism.

The rest of the week. APIMart, which a week ago served no text models at all, sells them again with per-token prices 20% under list on Claude and GPT. RunPod’s public Kimi endpoint is gone, so it stops being a coding option. Weights & Biases Inference is now CoreWeave Forge. opencode added a $40 Go Plus tier, Alibaba a $10 entry plan, ChatGPT a $500 Pro tier, and Qoder extended Qwen3.8-Flash free with no end date. Krea, which served DeepSeek V4.1 Flash eight days ago, serves nothing and is retired.

2026-09-25

The ruler stopped growing: what we do when the index our table runs on freezes

The scoreboard froze, and we can prove it. On 19 September Artificial Analysis moved its Intelligence Index to methodology v4.3.2 and replaced Terminal-Bench 2.1 with 4.0. Its coding index — the column this site ranks on — is computed from 2.1, so nothing evaluated after that date gets one, and the newest model with a coding score was released on 11 September. We compared our two committed snapshots: 261 models had the index on the 19th and 259 have it today. No surviving model lost its index and no value changed; the two that are missing, Agnes 2.5 Pro Alpha and Beta, dropped out of AA’s listing altogether. It is a migration, not an outage. It also means five flagships landed in two days with nothing to rank them by.

What we did about it, and what we refused to do. We did not change what the ranking means, and we did not borrow a number from somewhere else. We looked first: BenchLM publishes a usable dataset, but it is CC BY-NC and it aggregates vendor-reported figures rather than running its own; Terminal-Bench's own leaderboard is the source AA is migrating to, but it lists fifteen model-and-agent pairs, not models. So the inclusion rule got one bounded extension instead. A model that OpenRouter serves and AA scores only on intelligence now appears in the table, at the end, with no rank and an empty coding column, if it was released after 10 September. Six rows today. They sort themselves in the day AA scores them. A row with no rank is a model we cannot rank, not a model that scored badly, and the table says so in place.

Our tests have moved to Yardstick. Everything we measure ourselves — BigCodeBench-Hard, BFCL v4, RULER, SWE-bench, speed and energy — now lives at yardstick.millaguie.net, with the harness that produces it. It publishes per card, which this page never did. Keeping two copies of the same tables was a guarantee that one of them would go stale, so this tab is now a signpost. Model pages here still show their own first-party scores, and the leaderboard keeps its three ᴹ columns.

Tool calling is saturated. With BFCL on 45 local models, the best twelve sit between 90.0% and 87.5% with about two points of margin each: a twelve-way tie that runs from a 2 GiB Granite 4.0 Micro to a 32B. The line worth keeping is Ternary Bonsai 27B, 88.3% from a 7.1 GiB file, tying the 27B it was built from. We will keep publishing the score, but as a filter, not as a ranking.

Long context still separates, and the quantisation still decides. Qwen3.8-27B scores 97.2% at 32K from 12.2 GiB, tied with Qwen3.5 9B’s 96.9% from 6.1 GiB. Gemma 4 12B scores 70.9% at 32K in BF16 and 40.1% at Q3_K_M — thirty points for the file you picked, on a test where the same step costs it far less in coding. The lower point carries no truncation guard recorded, so read it as a floor.

One number that moved because the binary moved. EXAONE 4.5 33B's three RULER points come from a newer llama.cpp than the rest of our runs, because the reference build produced looping garbage at 8K and 32K. Its 4K score fell from 86.2 to 78.5 with the new build: same data, same seed, temperature 0, only the binary changed. Its card says all of it.

The market, briefly. nano-gpt sells Claude Opus 5.5 at Anthropic's list price; AIMLAPI sells it above list. Kimi renamed its four tiers and withdrew the request counts we used to publish, so we withdrew them too. Qoder recalibrated its factors, making Kimi K3 75% dearer in credits. APIMart has no text models left at all. ShareAI now sells closed models, priced in euros. And KAT-Coder-Pro V2 and Devstral 2 left the table with no provider serving them.

On the numbers. Market figures come from the vendors' own pages, read this week and dated in plans.yaml and gateways.yaml. First-party figures are ours: BFCL non-live, one to three runs, about two to three points of margin; RULER N=5 per task except where the card says otherwise.

2026-09-18

A 32B that fits 24 GB ties our best score, and a 3B that calls tools like a 27B

What we measured this week. The new open models from the last two weeks all have full curves now, three runs per point. The headline is K2 Horizon 32B: 34.0% at Q4_K_M, a 19.6 GiB file that fits a 24 GB card, within the margin of the best we have ever measured. The best two it ties need around 85 GiB. Qwen3.8-27B found a better 16 GB build, UD-IQ3_S at 33.8% from 11.2 GiB. Ternary Bonsai 2 lifts the 7 GB ternary 27B from 24.3% to 27.0%, but only runs on PrismML's own llama.cpp fork. And Spark-X2.5-4B scores 27.7% at Q8_0 but 19.4% at Q4_K_M, so on that model the bigger file is the one to download. The new VRAM tables now show, for each budget, the best measured build that fits it, with its score and its size from the same file.

A small model for tool calls, not for code. We finished BFCL, the function-calling benchmark, on the small models we already had quantisation curves for. The one that surprised me is Granite 4.0 Micro. At Q4_K_M, a 2 GiB file, it scores 87.6% on the non-live set. Qwen3.8-27B at UD-Q3_K_XL scores 88.2%. Each figure carries two to three points of margin, so that is a tie. The two even fail in the same places: both drop in Java and JavaScript, where the small model is in fact slightly ahead. The catch is that tool calls and code are different skills: on BigCodeBench-Hard the same Granite scores 16.2%, against 31.8% for the 27B at full precision. If your agent mostly routes calls to tools, the small model does the job. If it writes the code itself, it does not.

Read the categories, not the headline. SmolLM3-3B scores 65.7% overall and 25% on irrelevance, the test where the right answer is to call nothing. Before publishing we checked that this is the model and not our harness: all 180 failures are valid calls that should not have been made. MiniCPM5-2B is the reverse: 87.9% on irrelevance, but 28.5% when it has to make several calls at once. Both numbers disappear inside the total.

Two bits is a cliff for small models, not for big ones. Last week we wrote that the cliff is at two bits. This week we collected three curves of 27-32B dense models that had been measured and never picked up, and they say that rule only holds for the small ones. Qwen3.6 27B scores 28.8% at UD-Q2_K_XL (11 GiB) and 28.4% at Q5_K_M. Qwen3 32B loses about one point at two bits, and Granite 4.1 30B under three. Qwen3.5-9B lost more than half going from Q4_K_M to two bits. On a big dense model, two bits is often a fair way to fit a smaller card. On a small one it is not. It is not a law, though: Qwen3.8-27B drops from 31.8% to 23.0% at UD-IQ2_S, so read the curve of the model you actually run.

Gemma 3 4B is not worth its size. Its full curve peaks at 11.5% (Q6_K). That is below Granite 4.0 Micro and MiniCPM5-2B, which are both smaller. And the packing format shows again: Q4_K_M beats Q4_0 on Gemma 3 4B (9.5% against 8.1%) and on Granite 4.1 30B (25.0% against 23.0%), with the same bit width.

What changed in the market. CrofAI is shutting down and refunding balances. Public AI stopped being free. Chutes now states its quota in dollars: a plan includes five times its price in pay-as-you-go usage. Solheim cut its 256K plan from €30 to €24 but took away its bursts. Weights & Biases now applies its $0.50-per-million surcharge to its ARIA agent only. On OpenRouter, three routes became free (DeepSeek V4 Flash 0731, Qwen3.8 27B, GLM 5.2) and the listed price of gpt-oss-120b went up fourfold on input.

On the numbers. Every BigCodeBench figure is 148 problems and three runs, so about three to four points of margin: 28.8% against 28.4% is a tie. BFCL figures are the non-live score (irrelevance is measured and reported, but not counted in the total), one run for the small models and three for Qwen3.8-27B, two to three points of margin. Market figures are read off the vendors' own pages this week, with the date, in plans.yaml and gateways.yaml.

2026-09-13

Not all four-bit builds are the same model, and other things six curves taught us

Six quantisation curves, and the one number that surprised me. We measured Qwen3.5 at 2B, 4B and 9B, plus SmolLM3-3B, Granite 4.0 Micro and MiniCPM5-2B — every quantisation level of each, three runs per point, 148 problems of BigCodeBench-Hard on one class of card. The finding I did not expect is that "four bits" is not one thing. On Granite 4.0 Micro the K-quant at four bits scores 16.2% and the legacy Q4_0 build of the same model at nearly the same file size scores 9.5%. That is a 40% drop bought by nothing but the packing format. If your runner defaults to Q4_0 for compatibility, that default is costing you a chunk of the model.

The cliff is at two bits, not four. From Q8 down to Q3 these curves are close to flat — on Qwen3.5-9B the whole range sits between 21.6% and 27.7%, a spread smaller than the gap between two neighbouring models. Paying for 8 bits over 4 buys almost nothing here. Two bits is where it collapses: the 9B drops to 12.2%, the 4B to 8.1%, and on the 2B the two-bit build produced nothing gradeable at all — our publication gate withdrew that point rather than print a zero, because a zero would read as "scores badly" when the truth is "emitted no usable code".

What quantisation damages, and what we cannot yet claim. We ran RULER on the same files for long-context recall. The coding collapse at two bits is not there: Qwen3.5-9B goes from 96.8% to 92.3% at 32K while its coding score more than halves. But be careful with that comparison, because we were: those RULER runs are N=5 per task, which puts a margin of 3 to 8 points on every figure, so the two overlap and we CANNOT say the context score is unaffected — only that no collapse is visible where coding clearly shows one. Settling it properly needs N=100, which is what our anchor runs use.

And size beats bits. The 9B at two bits (12.2%) beats the 4B at two bits (8.1%) and is two and a half times the 2B at full precision (4.7%). On a fixed VRAM budget, take the bigger model at a lower quant. The exception that proves how much the model itself matters: MiniCPM5-2B scores 21.6% at Q8_0, four times what Qwen3.5-2B manages at any level, from a file that is 1.6 GB at Q4_K_M.

What we could not measure, and why. Gemma 4 12B is absent because it cannot be measured, and the reason is not our rig: it falls into a deterministic repetition loop on roughly three runs in five, which reproduces at full precision across 48 seeds and is an open, acknowledged bug that Google confirmed in July. We had it filed here as "not measurable" with the cause guessed wrong; it is now filed with the cause, the rate and the link. If you are choosing a local model, a 40% chance of no answer is the kind of thing a quality index will never show you.

Plans and providers, re-read. All twenty-four subscription plans were checked against their own pages this week, and several moved in ways worth knowing: Z.ai now publishes the formula that converts its credits to tokens, Ollama switched from billing GPU time to billing tokens, and opencode Go replaced its single monthly cap with a per-model one. We also found twenty-five inference providers serving models on our board with no entry of their own; they now have one, and the build warns us if another appears.

On the numbers: BigCodeBench-Hard is hard, so these figures look low next to marketing ones — what matters is the shape of the curve, not its height. Everything is temperature 0, one request at a time, three runs per point; 23 of 24 points agreed to the decimal across runs, which is what temperature 0 should do.

2026-08-28

Two leaderboard rows were wrong, and a hybrid-attention model worth explaining

Two leaderboard rows were wrong, and I fixed them. Qwen3.8-27B, the model on our front page, was showing the quality score of its weakest variant (58.2) instead of the one that actually serves that route (68.1) — an internal mapping bug, my code grabbed the first text match without checking there were four variants sharing the same base name. And Nemotron 3 Super 120B showed no quality score at all, because of an extra "NVIDIA " in the name that meant it never matched anything. Both are fixed, and I checked all 112 leaderboard rows one by one so none of the others are silently broken the same way.

Qwen3.8-Flash-Next has an architecture worth explaining properly. It shipped August 26 as an open preview of Qwen4: 180B parameters stored in total, but only 6B active per token. The trick is how it spreads those 180B. The backbone is 48 layers in 12 identical blocks, and each block runs three layers of Gated DeltaNet followed by one of Qwen Sparse Attention. Gated DeltaNet is linear attention: instead of re-reading the whole conversation every time (the squared cost that makes ordinary attention slow), it compresses the history into a fixed state it updates with a learned gate. It "remembers" instead of re-reading. Sparse attention does the opposite: it groups the context into 2,048-token blocks and a learned indexer picks which ones deserve full attention. It "retrieves" precisely instead of recalling everything. Cheap memory almost always, expensive search only when needed — that is how it reaches a million tokens of context without the cost exploding. Alibaba says the sparse-attention layer gives up to 7.6x faster prefill and 4.9x faster decode than full attention at 1M context.

It also carries a 51B-parameter table with 20 million bigram and trigram entries, sitting in layer 2: a cheat sheet of common text patterns it looks up instead of computing, so it adds no real compute cost and can be offloaded to system RAM if VRAM is tight. The residual stream, which in most models is a single channel running through the whole network, is here four parallel branches with their own gate, kept in FP8; Alibaba says that is what lets it train in FP8 at this level of sparsity without the model destabilizing. The result: the capacity of a 180B model, the compute cost of a 6B one, and training at roughly a ninth of what Qwen3.7-Plus cost. Already running it on the machines I use for testing, and I have hit the other side of so much novelty: it needs a custom llama.cpp build because the architecture is not in the main branch yet, and unsloth's GGUF does not ship the multi-token-prediction module (a separate 4B, for faster generation) that the model does have out of the box. Right now that is 10.5 tok/s with experts in RAM and 21 tok/s with everything in unified memory. Quant curve in progress, for next week.

GLM-5.3-Flash is not a mystery any more. Last time I mentioned an unbranded model on OpenRouter that resolved 10 of 10 on our SWE-bench pool, and guessed from its tokenizer that it was from the GLM family. I was right: it is GLM-5.3-Flash, a 321B MoE under MIT, and Z.ai now serves it directly (mirrored by Novita) in fp8. It has a full leaderboard entry now.

A new model that sounds like a lot. Tencent released Hy4 preview: 770B total parameters, only 49B active per token, over 1M context. In their own internal evaluation it scores above GLM-5.3 and Kimi K3 on coding tasks. No Artificial Analysis index yet — it shipped today — so it is not on the leaderboard yet. Watching it for next week.

Two subscriptions got more expensive without announcing it. MiniMax's Token Plan rose about 10% across all three tiers — Plus $20 to $22, Max $50 to $55, Ultra $120 to $132 — confirmed against the raw pricing page. Z.ai's coding plan rose less in percentage but more in absolute terms on the tiers that matter: Pro $72 to $80, Max $160 to $168; the credit quotas underneath did not change. That one only shows on the real subscription page — the docs page still advertises the plan as "starting at $18" with no more detail.

Cursor did not raise its price, it deleted the number. Until this week, Pro/Pro+/Ultra included a stated dollar amount of API usage ($20/$70/$400). The docs now describe two usage "pools" per plan with no dollar or token figure at all, just "included". That is a real disclosure regression, not a price change to the plans themselves.

The free set rotated again. opencode Zen added Ling 3.0 Flash Fin (the free route for Ling reappearing, it had vanished) and Muse Spark 1.2 Contributor; it dropped Big Pickle and DeepSeek V4 Flash Free. And OpenRouter now has 18 public :free routes, with GLM-5.2 (68.8 coding index) and MiniMax-M3 (58.6) among the new ones — more free capacity than there has ever been.

Catalogue cleanup. I dropped three models from the table that no longer exist on OpenRouter under any name (Ring-2.6-1T, Ling 3.0 Tiny, Gemma 3n E4B) — they used to sit there as rows with no price. And a price I had mislabeled on Morph: what showed as GLM-5.2's price ($6/$22.50 per million) actually belonged to "Kimi K3 fast" — GLM-5.2's real price is $0.84/$3.14, far cheaper.

One note on Our tests. That tab keeps filling in daily as BCB-Hard, BFCL and RULER results land, not on a weekly schedule like the rest of the site. It has grown enough of its own shape, its own protocol, its own model and quant matrix, that it may end up as its own site rather than a tab here. Nothing changes for now, just flagging it before it happens.

As always, free sets rotate and close without notice, so treat any free allowance as this week's, not a standing plan.

2026-08-14

First weekly note: a free model that reasons, and prices that moved

I am starting a short note each week to explain what changed on the site, in plain words. Same data as the tables, just told as a story. This is the first one.

A new free model I could actually test. NVIDIA released Nemotron 3.5 Lightning this week, a 30B mixture-of-experts model with only 3B active parameters. The small active count is the interesting part: it means the model is cheap to run and it fits on modest hardware. It is also already free on opencode Zen, so I did not need to rent a GPU. I pointed our BigCodeBench-Hard harness at the free endpoint and measured it myself: 34.7% on BCB-Hard (42 of 121). It is a reasoning model, so it thinks before it answers and it is not fast, but for a free local-agent model that is a solid number. It is now both on the leaderboard (Artificial Analysis gives it a coding index of 26.8) and in the Run local catalogue with our own measurement.

Qwen3.8-27B landed, and I ran it. The Max is still datacenter-only (2.4T, 95B active), but the small dense 27B (Apache-2.0) that fits at home shipped this afternoon, so I tested it properly. On my 16 GB card it fits at Q3 (about 14 GB), and with MTP speculative decoding through llama-server it runs at 36 tok/s, twice the 18 it manages without spec-decode, because the model ships its own draft head and about 87% of the drafts get accepted. On our BCB-Hard it scores 27% at Q3 non-thinking; the full model at good precision (BF16/FP8) is around 37 to 43%, so the 3-bit quant you need to fit 16 GB costs about 10 points. It is not on the leaderboard yet, no Artificial Analysis index and not on OpenRouter, so for now it lives in Run local with our own numbers. It made me want a bigger GPU, but on the 16 GB card I actually have it does not beat what I already run, so I have not switched, the honest take is in the Run local pick. And it is dense, so if you cannot buy a card the one to wait for is the small A3B MoE Qwen will ship next, which runs fast on the hardware you already own.

The ceiling kept climbing. Grok 4.6 joined the board at a 76.8 coding index, near the top of everything tracked. Solar Pro 4 (52.7), Muse Glimmer 30B (49) and Nemotron 3.5 Lightning also went on. Qwen3.8-Max was already there at 71.8. DeepSeek also shipped a new revision of V4 Pro (0813) that jumped its coding index from 59.4 to 68.8, a big step for an already cheap open model.

GLM-5.3 dropped today too. Zhipu launched it on August 14, "built to code", and it looks strong: they claim +50% coding over GLM-5.2 and first place among open-source models on Terminal-Bench 3.0 (DeepSWE 1.1 66.9, CyberGym 84.5). The Z.ai coding plan already serves it, and GLM-5.2 and 5.1 requests now route to 5.3. But it is not on OpenRouter or the Artificial Analysis index yet, and the open weights are about two weeks out, so it does not go on the leaderboard yet. That is the rule here: a model only joins the board once it has a public coding index and a place to buy it. I will add it the moment it lands on both.

Prices moved in a few places. OpenAI turned on unlimited text with GPT-5.6 Luna for the Free and Go tiers, it went live this week. Note it is text only and the paid tiers still keep caps on Luna, the earlier claim that Luna was unlimited on Plus and Pro was wrong. Synthetic's $7 first-month promo ended, it is back to $30 a month. Mistral folded its old $5.99 Education tier into a student discount on the $14.99 Pro plan. Alibaba now shows a $2 new-user discount across its Individual token tiers.

DeepSeek is about to get more expensive. V4 Pro 0813 shipped at the old flat rate, but from August 16 (16:00 UTC) DeepSeek switches to peak/off-peak billing and the jumps are steep: cache-hit input rises up to 12x at peak, cache-miss up to 3x, output up to 4.5x, with peak on Beijing working hours (01:00-04:00 and 06:00-10:00 UTC). The cache-hit line is the one that stings, because a coding agent replays a big cached prefix every turn, so the effective bill tracks that 12x, not the headline flat price. If you lean on DeepSeek, front-load the heavy runs before the 16th, or push them to off-peak after.

Two corrections worth calling out. The "Cerebras free tier sunsets on Aug 17" note that went around is a false alarm: that date only kills preview models on the paid Developer tier, the free trial is not going away. And Morph quietly rebuilt its whole pricing to pure pay-per-token with 200 free requests a month, the old Free/Starter/Pro/Scale credit tiers are gone. Their page bot-walls plain requests, so I had to open it in a real browser to read it, now confirmed.

New this week. The free set on opencode Zen rotated: it added Hy3 and Nemotron 3.5 Lightning and dropped Ling-3.0-flash, LongCat-2.0 and North Mini Code. I added RunPod to the gateways, it has a real OpenAI-compatible inference API through its Public Endpoints, though the text catalogue is thin and it is mostly image and video. And I added this notes tab.

As always, free sets rotate and close without notice, so treat any free allowance as this week's, not a standing plan.

And one last thing: next week we bring a surprise. Stay tuned.

Leaderboard

Ranked by AA’s intelligence index. Value = index per $ (blended, 3:1 in:out).

The ranking metric changed on 2026-10-02, and here is why. Artificial Analysis froze its coding index on 2026-09-11: it was computed from Terminal-Bench 2.1, which AA no longer runs, so every model released after that date carries no coding score. Its per-model replacement is Terminal-Bench 4.0, a pass rate over 66 agentic tasks attempted three times each, and AA has run it on 60 of the 131 models here — not enough to rank the table. So the rank now follows AA’s intelligence index, which covers every row. That is not a convenience: among the 50 models that carry both, the intelligence order matches the Terminal-Bench 4.0 order more closely (Spearman 0.95) than the old coding index does (0.93). Both coding columns stay, measured only. We do not convert the frozen index into an estimated 4.0 score: a one-variable fit preserves the order, so it would buy no new ranking, and it extrapolates badly exactly where the data is missing — it puts Claude Opus 4.7 at 3.2% when the two measured models closest to it on the frozen index, Opus 4.8 and GLM-5.3, score 21.7% and 41.9%. One warning about AA’s API, since it caught us: it returns a plain 0 for a model it has not run on 4.0, while using a real null elsewhere. Apodex 1.1 passes 69.7% of Terminal-Bench 2.1 and shows 0 on 4.0. We read those zeros as missing, not as a score of zero.

#ModelIntelligenceTBench 4.0AA coding idxLong-ctxtok/s$/M blendedValue idx/$ContextBCB-Hard ᴹBFCL ᴹRULER ᴹ
1Claude Opus 5.5 · Anthropic57.659.6%—0.84797.3$8.00071,000,000———
2Claude Sonnet 5.5 · Anthropic56.063.6%—0.827137.5$4.000141,000,000———
3Claude Fable 5.1 · Anthropic53.452.0%81.60.85370.2$20.00031,000,000———
4GPT-6 Astra · OpenAI52.759.1%76.90.80755.6$20.00031,050,000———
5GPT-6.1 Sol · OpenAI51.856.1%—0.83058.8$4.000131,050,000———
6Claude Opus 5 · Anthropic50.849.0%78.00.793—$10.00051,000,000———
7Claude Fable 5 · Anthropic49.642.4%76.50.823—$20.00021,000,000———
8Muse Spark 1.3 · Meta48.133.3%75.80.830157.5$2.000241,048,576———
9GPT-6 Sol · OpenAI47.643.9%—0.837100.5$4.000121,050,000———
10GPT-5.6 Sol · OpenAI47.039.9%77.40.840—$4.000121,050,000———
11Grok 4.7 · xAI46.425.8%—0.76779.3$3.00015500,000———
12MiMo-V2.6-Pro · Xiaomi46.334.8%—0.86344.3$0.544851,050,000———
13Qwen3.8 Max · Alibaba / Qwen45.438.9%76.20.80337.4$3.000151,000,000———
14GLM-5.3 · Z.ai / Zhipu44.841.9%74.80.79774.0$2.150211,048,576———
15Grok 4.6 · xAI44.321.2%76.80.803—$3.00015500,000———
16Kimi K3 · Moonshot43.612.6%76.20.88735.2$5.40081,048,576———
17GPT-5.6 Terra · OpenAI42.135.4%76.70.830109.5$4.50091,050,000———
18Claude Opus 4.8 · Anthropic41.821.7%74.30.777—$10.00041,000,000———
19GLM-5.3-Flash · Z.ai / Zhipu41.832.8%71.50.80051.7$0.2371761,048,57619.6 ·5q86.9—
20Gemini 3.8 Flash · Google40.919.7%76.30.813239.1$1.500271,048,576———
21Claude Opus 4.7 · Anthropic40.7—73.60.787—$10.00041,000,000———
22Qwen3.8 2.4T A95B · Alibaba / Qwen39.911.1%71.90.80338.6$3.000131,048,576———
23Gemini 3.7 Flash · Google39.6—71.50.830—$1.500261,048,576———
24Muse Spark 1.2 · Meta39.67.1%72.20.790—$2.000201,048,576———
25DeepSeek V4.1 Flash · DeepSeek39.526.8%—0.840209.5$0.525751,048,576———
26GPT-5.4 · OpenAI39.0—71.10.820—$5.62571,050,000———
27Grok 4.5 · xAI38.810.6%72.40.793—$3.00013500,000———
28GPT-5.5 · OpenAI38.414.6%74.90.843—$11.25031,050,000———
29Claude Sonnet 5 · Anthropic38.214.1%71.50.820—$4.000101,000,000———
30GPT-6 Luna · OpenAI38.112.6%—0.833136.9$0.2001901,050,000———
31MiMo-V2.6-Flash · Xiaomi37.922.7%—0.74349.0$0.1752171,050,000———
32GPT-5.6 Luna · OpenAI37.311.6%71.40.837—$0.450831,050,000———
33DeepSeek V4 Pro · DeepSeek36.014.1%68.80.803108.4$0.990361,048,576———
34DeepSeek V4 Flash Vision · DeepSeek34.812.1%65.00.813216.3$0.3231081,048,576———
35DeepSeek V4 Flash · DeepSeek34.312.1%69.10.797—$0.3241061,048,57633.1 ·7q——
36Gemini 3.6 Flash · Google34.07.1%69.20.800—$1.500231,048,576———
37Muse Spark 1.1 · Meta33.76.1%71.30.777—$2.000171,048,576———
38GLM-5.2 · Z.ai / Zhipu33.71.0%68.80.783—$1.305261,048,576———
39Qwen3.8-27B · Alibaba / Qwen33.75.6%68.10.82045.8$1.065321,000,00031.8 ·12q88.195.4 @128K
40Gemini 3.5 Flash · Google33.6——0.743—$3.375101,048,576———
41Claude Sonnet 4.6 · Anthropic30.13.0%63.00.800—$6.00051,000,000———
42Gemini 3.1 Pro Preview · Google29.74.0%68.80.820123.6$4.50071,048,576———
43Qwen3.7 Max · Alibaba / Qwen29.51.5%66.00.790—$2.213131,000,000———
44MiniMax-M3 · MiniMax29.22.0%58.60.83074.3$0.525561,048,576———
45Solar Pro 4 · Upstage28.20.5%52.70.74085.3$0.158179524,288———
46Grok Build 0.1 · xAI27.2—51.50.747—$1.25022256,000———
47Kimi K2.6 · Moonshot27.00.5%61.80.810—$0.78334262,144———
48Qwen3.6 Plus · Alibaba / Qwen27.0—54.50.783—$0.731371,000,000———
49GLM-5.1 · Z.ai / Zhipu26.12.0%55.80.737—$2.15012204,800———
50Xiaomi MiMo-V2.5-Pro · Xiaomi26.0—60.20.797—$0.544481,050,000———
51Kimi K2.7 Code · Moonshot25.81.0%60.80.79384.6$1.34119262,144———
52Inkling Small · Thinking Machines25.71.0%52.90.757131.5$0.63740524,288———
53Tencent Hy3 · Tencent25.30.5%58.80.79087.7$0.231110262,144———
54Xiaomi MiMo-V2.5 · Xiaomi25.2—56.80.730—$0.1751441,050,000———
55Qwen3.7 Plus · Alibaba / Qwen25.21.0%55.90.73055.7$0.560451,000,000———
56Inkling · Thinking Machines25.01.0%52.10.773167.6$1.72514524,288———
57Grok 4.3 · xAI24.9—42.20.730—$1.562161,000,000———
58GPT-5.1 · OpenAI24.7—49.40.800—$3.4387400,000———
59Ling-3.0-flash-VL · InclusionAI24.6—57.00.783—$0.031790262,14425.088.2—
60GPT-5.4 Mini · OpenAI24.12.0%56.10.770—$1.68814400,000———
61Solar Mini 4 · Upstage24.11.0%—0.833192.7$0.087275524,288———
62Kimi K2.5 · Moonshot23.5—46.80.780—$0.90026262,144———
63GPT-5 · OpenAI23.0—37.80.782—$3.4387400,000———
64Nemotron 3 Ultra 550B · Nvidia22.90.5%49.30.793163.3$0.92525262,144———
65MiniMax M2.7 · MiniMax22.8—52.60.783—$0.36762204,80025.078.492.3 @32K
66Ling-3.0-flash-Fin · InclusionAI22.6—55.60.737354.9$0.062363262,14425.088.2—
67Gemini 3.5 Flash-Lite · Google22.21.0%49.30.760332.5$0.850261,048,576———
68GLM 4.7 · Z.ai / Zhipu22.2—45.30.710—$1.00022204,800———
69DeepSeek V3.2 · DeepSeek21.5—44.20.733—$0.31568163,840———
70Qwen3.6 27B · Alibaba / Qwen21.4—53.70.773—$1.04021262,14428.4 ·5q86.7—
71Qwen3.5 397B A17B · Alibaba / Qwen21.4——0.64385.5$1.28817262,144———
72GPT-5.4 Nano · OpenAI20.70.5%56.10.767—$0.46345400,000———
73GPT-5 Mini · OpenAI20.6——0.723—$0.68830400,000———
74Ling-3.0-flash · InclusionAI20.1—50.60.730354.3$0.032638262,14425.088.2—
75Step 3.7 Flash · StepFun19.5—39.60.737—$0.43845262,144———
76Qwen3.5-35B-A3B · Alibaba / Qwen19.3——0.720—$0.36253262,144———
77LongCat 2.0 · Meituan19.1—45.30.650—$0.525361,048,756———
78GLM 4.6 · Z.ai / Zhipu18.5—45.80.540—$0.76024204,800———
79Qwen3.6 35B A3B · Alibaba / Qwen18.2—41.90.717134.9$0.36250262,14427.782.3—
80Qwen3.5-122B-A10B · Alibaba / Qwen17.7—43.30.613143.3$0.71525262,144———
81Muse Glimmer · Meta17.50.5%49.00.833144.9$0.63727131,072———
82Gemma 4 26B A4B · Google16.7—39.30.657—$0.107156262,144—89.8—
83Gemini 2.5 Pro · Google16.1—33.30.690—$3.43851,048,576———
84Gemini 3.1 Flash Lite · Google15.60.5%34.70.743—$0.562281,048,576———
85o1 · OpenAI15.2—39.70.650—$26.2501200,000———
86DeepSeek V3.1 Terminus · DeepSeek14.8—43.50.693—$0.45333163,840———
87Gemma 4 31B · Google14.7—43.40.69735.5$0.15296262,14418.9 ·7q88.0—
88Mistral Medium 3.5 · Mistral14.2—46.90.693167.3$3.0005262,144———
89Mercury 2 · Inception13.8—31.10.437—$0.37537128,000———
90Qwen3.5-9B · Alibaba / Qwen13.3—23.50.46081.4$0.112118262,14429.0 ·8q81.596.8 @32K
91Command A · Cohere13.10.5%27.80.527211.2$4.3753256,000———
92Command A+ · Cohere13.10.5%27.80.527211.2$0.60022192,000———
93Nemotron 3.5 Lightning · Nvidia12.90.5%26.80.603290.2$0.087148262,14429.081.1—
94Nemotron 3 Super 120B · Nvidia12.8—37.70.657157.1$0.17274262,144———
95Qwen3 235B A22B 2507 · Alibaba / Qwen12.7—22.10.720—$0.79616131,072———
96o3 Mini · OpenAI12.5————$1.9256200,000———
97gpt-oss-120b · OpenAI11.6—30.40.520155.7$0.070165131,072—74.0—
98DeepSeek R1 · DeepSeek11.4—24.60.577—$1.1501064,000———
99Mistral Small 4 · Mistral11.3—26.60.497173.8$0.26243262,144———
100Granite 4.2 8B · IBM11.1—22.40.45056.1$0.107103131,072———
101Trinity Large Thinking · Arcee Ai10.80.5%25.80.380365.8$0.38728262,144———
102Nemotron 3 Nano Omni 30B · Nvidia10.3—13.80.397264.8$0.000—256,000———
103GPT-4.1 Mini · OpenAI10.2—20.20.440—$0.700151,047,576———
104gpt-oss-20b · OpenAI10.0——0.310220.1$0.036278131,07227.074.0—
105Llama 4 Maverick · Meta10.0—16.30.50078.5$0.304331,048,576———
106North Mini Code · Cohere9.90.5%36.50.37387.6$0.000—256,00023.679.9—
107Qwen3 30B A3B 2507 · Alibaba / Qwen9.8—12.10.613—$0.21546131,072———
108DeepSeek V3 0324 · DeepSeek9.7—21.20.407—$0.50219163,840———
109Mistral Large 3 · Mistral9.3—20.10.36079.3$0.75012262,144———
110Qwen3 Coder Next · Alibaba / Qwen9.2—36.20.47083.7$0.29032262,144———
111Mistral Medium 3.1 · Mistral9.2—20.50.427—$0.80011131,072———
112GPT-4o · OpenAI9.0————$4.3752128,000———
113Nemotron 3 Nano 30B · Nvidia8.9—14.40.380176.6$0.087102262,144———
114Qwen3 32B · Alibaba / Qwen8.6—15.30.000—$0.13066131,07223.9 ·4q89.3—
115Devstral 2 · Mistral8.6—31.30.323—$0.80011262,144———
116DeepSeek V3 · DeepSeek8.5—23.00.293—$0.45019163,840———
117LFM2.5-2.6B · Liquid8.4—7.70.057—$0.000—65,536———
118Qwen3 14B · Alibaba / Qwen8.2—13.80.000—$0.15055131,07223.088.687.4 @32K
119Mistral Small 3.2 · Mistral8.2—12.50.203—$0.13362256,00023.061.6—
120Llama 4 Scout · Meta8.1—8.20.277104.7$0.150541,310,720———
121Solar Pro 3 · Upstage7.8—16.20.323—$0.26230131,072———
122GPT-4.1 Nano · OpenAI7.8—11.10.203—$0.175451,047,576———
123Qwen3 8B · Alibaba / Qwen7.3—9.00.000—$0.20236131,07225.0—88.4 @32K
124Mistral Small 3.1 · Mistral7.1—26.30.223—$0.40218128,000———
125GPT-4 Turbo · OpenAI7.0—21.5——$15.0000128,000———
126GPT-4 · OpenAI6.7—13.1——$37.50008,191———
127GPT-4o-mini · OpenAI6.7—11.4——$0.26226128,000———
128GPT-3.5 Turbo · OpenAI5.5—10.7——$0.750716,385———
129Gemma 3 27B · Google4.9—10.10.073—$0.17228131,072———
130Gemma 3 4B · Google4.8—2.70.067—$0.06277131,07210.8 ·8q55.559.6 @32K
131Gemma 3 12B · Google3.8—5.80.083—$0.07551131,072———

The last three columns are ours, and they are not on the same scale as the rest. BCB-Hard is pass@1 over 148 hard programming problems — a 32 there and a 68 in the coding index are not comparable and subtracting them means nothing. They sit at the end of the row for that reason. BCB-Hard also shows how many quantisations we measured of that model (·7q), which is the part nobody else publishes: how much of the score is the model and how much is the file you downloaded. RULER carries the length it was measured at, because a score at 4K and the same score at 32K are different facts. Click the model to see the conditions — card, quantisation, how many runs, and the margin. Everything else here is third-party from Artificial Analysis — the only source publishing measured speed. A dash means the model has not been independently measured, not that it scored zero. Intelligence is AA’s general index, shown next to the coding one because they diverge: a model can code well and reason poorly, or the reverse, and only the coding column decides the ranking here. Long-ctx is AA’s long-context reasoning score — the column that matters on a large codebase, since Context only says how many tokens fit, not whether the model still reasons at that length.

Published quality

Mind the provenance

The flattering coding scores are vendor-reported. Where an independent lab measured the same benchmark the numbers agree closely — so the gap is about which benchmarks get published, not fabricated values.

ModelBenchmarkScoreProvenanceSource
MiniMax-M3SWE-bench Verified80.5%vendorminimax.io
MiniMax-M3SWE-bench Pro59.0%vendorminimax.io
MiniMax-M3Terminal-Bench 2.166.0%vendorminimax.io
MiniMax-M3Terminal-Bench 2.165.2%independentArtificial Analysis
MiniMax-M3TerminalBench-Hard42.4%independentArtificial Analysis
MiniMax-M3Aider polyglotnot listedindependentaider.chat
Xiaomi MiMo-V2.5-ProSWE-bench Verified78.9%vendormimo.xiaomi.com
Mistral Devstral Small 2SWE-bench Verified68.0%vendormistral.ai

The leaderboard ranks on Artificial Analysis’ intelligence index since 2026-10-02, and shows its coding scores where they exist. AA publishes any of them only for the models it chooses to test, and most models here have none: those are not benchmarks we can run. So they were measured on tests we can run, and placed next to models that appear on both scales.

Withdrawn 2026-08-23. The BigCodeBench-Hard bridge table that used to sit here was measured with our older protocol over 121 problems. It is not comparable with the numbers we publish now, and some models moved by more than 20 points, so leaving it up would have been worse than removing it. The current first-party scores are at yardstick.millaguie.net; we will rebuild the bridge once enough models have been re-measured.

SWE-bench Verified (run here, 25 of the 500 instances, chosen because their reference patch passes in our container; 80-step budget)

ModelScoreAA coding index
Tencent68% (17/25)58.8
Agnes AI52% (13/25)not measured

Providers & quantization

The column nobody else shows

The same model is served by many providers at nearly the same price but at different numeric precision. Cheapest is often the most aggressively quantized — and sometimes higher precision costs the same or less.

bf16fp8fp4unknown

GPT-5.6 Sol

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.00$5.001,050,000100.0%
OpenAIunknown$2.00$10.001,050,000100.0%
Azureunknown$4.00$20.001,050,000100.0%
OpenAIunknown$4.00$20.001,050,000—
Amazon Bedrockunknown$4.40$22.001,050,000—
Azureunknown$4.40$22.001,050,000—
Azureunknown$4.40$22.001,050,000—

Claude Opus 5

11 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Amazon Bedrockunknown$5.00$25.001,000,000100.0%
Anthropicunknown$5.00$25.001,000,000100.0%
Azureunknown$5.00$25.001,000,000—
Claude Platform on AWSunknown$5.00$25.001,000,000100.0%
Googleunknown$5.00$25.001,000,000100.0%
Amazon Bedrockunknown$5.50$27.501,000,000—
Amazon Bedrockunknown$5.50$27.501,000,000—
Azureunknown$5.50$27.501,000,000—
Googleunknown$5.50$27.501,000,000—
Googleunknown$5.50$27.501,000,000—
Anthropicunknown$10.00$50.001,000,000—

GPT-5.6 Terra

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.00$6.001,050,000—
Azureunknown$2.00$12.001,050,000100.0%
OpenAIunknown$2.00$12.001,050,000100.0%
Amazon Bedrockunknown$2.20$13.201,050,000—
Azureunknown$2.20$13.201,050,000—
Azureunknown$2.20$13.201,050,000—
OpenAIunknown$4.00$24.001,050,000—

Claude Fable 5

6 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$10.00$50.00——
Amazon Bedrockunknown$10.00$50.001,000,000—
Anthropicunknown$10.00$50.001,000,000100.0%
Azureunknown$10.00$50.001,000,000—
Claude Platform on AWSunknown$10.00$50.001,000,000—
Googleunknown$10.00$50.001,000,000—
Googleunknown$11.00$55.001,000,000—

Buying direct from Anthropic costs $10.00 in and $50.00 out per million, cached input $1.000.

Kimi K3

23 providers · bf16, fp4, fp8, mxfp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Moonshot directfirst-party$3.00$15.00——
InferenceNetfp4$0.99$13.001,048,576100.0%
Relacefp4$0.99$13.001,048,576100.0%
Phalaunknown$2.25$11.251,048,57699.2%
Waferunknown$1.44$14.001,048,576100.0%
Morphfp8$2.50$11.361,048,57693.4%
Makoraunknown$2.04$12.751,048,57699.2%
DigitalOceanunknown$2.55$12.951,048,576—
Togetherunknown$2.70$13.501,048,57699.9%
Sail Researchfp4$2.80$14.001,048,576—
Waferunknown$2.80$14.001,048,576100.0%
DeepInfrabf16$2.85$14.251,048,576100.0%
AkashMLfp4$3.00$15.001,048,576—
BaseTenfp8$3.00$15.001,048,57698.5%
Chutesmxfp4$3.00$15.001,048,57699.5%
Decartmxfp4$3.00$15.001,048,576100.0%
Fireworksunknown$3.00$15.001,048,57699.8%
InferenceNetfp4$3.00$15.00250,000100.0%
Modalmxfp4$3.00$15.001,048,576100.0%
Moonshotmxfp4$3.00$15.001,048,576100.0%
Parasailfp4$3.00$15.001,048,57699.9%
Alibaba / Qwenunknown$3.45$17.251,048,57698.7%
Fireworksunknown$4.50$22.501,048,576100.0%
Fireworksunknown$4.50$22.501,048,57699.5%

Buying direct from Moonshot costs $3.00 in and $15.00 out per million, cached input $0.300. That is 50% dearer than the cheapest routed endpoint on a 3:1 blend.

GPT-5.5

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$2.50$15.001,050,000—
Azureunknown$5.00$30.001,050,000100.0%
OpenAIunknown$5.00$30.001,050,000100.0%
Amazon Bedrockunknown$5.50$33.001,050,000—
Azureunknown$5.50$33.001,050,000—
Azureunknown$5.50$33.001,050,000—
OpenAIunknown$12.50$75.001,050,000—

Claude Opus 4.8

11 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$5.00$25.00——
Amazon Bedrockunknown$5.00$25.001,000,000—
Anthropicunknown$5.00$25.001,000,000100.0%
Azureunknown$5.00$25.001,000,000—
Claude Platform on AWSunknown$5.00$25.001,000,000100.0%
Googleunknown$5.00$25.001,000,000100.0%
Amazon Bedrockunknown$5.50$27.501,000,000—
Amazon Bedrockunknown$5.50$27.501,000,000—
Azureunknown$5.50$27.501,000,000—
Googleunknown$5.50$27.501,000,000—
Googleunknown$5.50$27.501,000,000—
Anthropicunknown$10.00$50.001,000,000100.0%

Buying direct from Anthropic costs $5.00 in and $25.00 out per million, cached input $0.500.

Claude Opus 4.7

8 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$5.00$25.00——
Amazon Bedrockunknown$5.00$25.001,000,000100.0%
Anthropicunknown$5.00$25.001,000,000—
Azureunknown$5.00$25.001,000,000—
Claude Platform on AWSunknown$5.00$25.001,000,000100.0%
Googleunknown$5.00$25.001,000,000—
Amazon Bedrockunknown$5.50$27.501,000,000—
Googleunknown$5.50$27.501,000,000—
Googleunknown$5.50$27.501,000,000—

Buying direct from Anthropic costs $5.00 in and $25.00 out per million, cached input $0.500.

Grok 4.5

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAI directfirst-party$2.00$6.00——
xAIunknown$2.00$6.00500,000100.0%
xAIunknown$2.00$6.00500,00099.8%
xAIunknown$4.00$12.00500,000—
xAIunknown$4.00$12.00500,000—

Buying direct from xAI costs $2.00 in and $6.00 out per million, cached input $0.300.

Grok 4.6

6 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAIunknown$2.00$6.00500,00099.7%
xAIunknown$2.00$6.00500,00099.7%
Amazon Bedrockunknown$2.20$6.60500,000—
xAIunknown$2.20$6.60500,000—
xAIunknown$4.00$12.00500,000—
xAIunknown$4.00$12.00500,000—

Muse Glimmer

3 providers · bf16, unknown
Served byQuant$/M in$/M outContextUptime 30m
Phalaunknown$0.30$1.10131,072100.0%
DeepInfrabf16$0.30$1.20131,072100.0%
Togetherunknown$0.35$1.50131,072100.0%

Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Solar Pro 4

2 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Upstageunknown$0.09$0.36524,288100.0%
Upstageunknown$0.09$0.36524,288100.0%

Nemotron 3.5 Lightning

5 providers · bf16, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Darkbloomint4$0.04$0.18262,14499.9%
DeepInfrabf16$0.06$0.16262,144100.0%
Io Netunknown$0.06$0.17262,144100.0%
CoreWeavebf16$0.07$0.20262,144100.0%
Phalaunknown$0.07$0.20262,144100.0%

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Claude Sonnet 5

10 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$2.00$10.00——
Amazon Bedrockunknown$2.00$10.001,000,00099.8%
Anthropicunknown$2.00$10.001,000,000100.0%
Azureunknown$2.00$10.001,000,000100.0%
Claude Platform on AWSunknown$2.00$10.001,000,000100.0%
Googleunknown$2.00$10.001,000,000100.0%
Amazon Bedrockunknown$2.20$11.001,000,000—
Amazon Bedrockunknown$2.20$11.001,000,000—
Azureunknown$2.20$11.001,000,000—
Googleunknown$2.20$11.001,000,000100.0%
Googleunknown$2.20$11.001,000,000100.0%

Buying direct from Anthropic costs $2.00 in and $10.00 out per million, cached input $0.200.

GPT-5.6 Luna

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.10$0.601,050,000100.0%
Azureunknown$0.20$1.201,050,00099.8%
OpenAIunknown$0.20$1.201,050,000100.0%
Amazon Bedrockunknown$0.22$1.321,050,000100.0%
Azureunknown$0.22$1.321,050,000—
Azureunknown$0.22$1.321,050,000100.0%
OpenAIunknown$0.40$2.401,050,000100.0%

Muse Spark 1.1

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Metaunknown$1.25$4.251,048,576—

GPT-5.4

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.25$7.501,050,000—
Azureunknown$2.50$15.001,050,000100.0%
OpenAIunknown$2.50$15.001,050,000100.0%
Amazon Bedrockunknown$2.75$16.501,050,000—
Azureunknown$2.75$16.501,050,000—
Azureunknown$2.75$16.501,050,000100.0%
OpenAIunknown$5.00$30.001,050,000—

Gemini 3.5 Flash

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.75$4.501,048,576—
Google AI Studiounknown$0.75$4.501,048,576100.0%
Googleunknown$1.50$9.001,048,57699.7%
Google AI Studiounknown$1.50$9.001,048,576100.0%
Googleunknown$1.65$9.901,048,576—
Googleunknown$2.70$16.201,048,576—
Google AI Studiounknown$2.70$16.201,048,576—

Gemini 3.6 Flash

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.38$1.881,048,576—
Google AI Studiounknown$0.38$1.881,048,576—
Googleunknown$0.75$3.751,048,576100.0%
Google AI Studiounknown$0.75$3.751,048,576100.0%
Googleunknown$0.83$4.121,048,576—
Googleunknown$1.35$6.751,048,576100.0%
Google AI Studiounknown$1.35$6.751,048,57699.4%

Gemini 3.7 Flash

6 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.38$1.881,048,576100.0%
Google AI Studiounknown$0.38$1.881,048,57699.9%
Googleunknown$0.75$3.751,048,57698.7%
Google AI Studiounknown$0.75$3.751,048,57699.9%
Googleunknown$1.35$6.751,048,576—
Google AI Studiounknown$1.35$6.751,048,576—

GLM-5.2

33 providers · fp4, fp8, mxfp4, nvfp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$1.40$4.40——
Baidufp8$0.14$0.441,048,576100.0%
Baidufp4$0.23$0.811,048,57699.9%
StreamLakefp8$0.28$0.881,024,00099.4%
InferenceNetunknown$0.10$2.201,048,576100.0%
DeepInfrafp4$0.56$1.801,048,57699.1%
Decartmxfp4$0.39$2.401,048,576100.0%
Morphfp8$0.17$3.071,048,57697.5%
Novitafp8$0.65$2.041,048,57699.7%
DigitalOceanunknown$0.70$2.20262,14498.2%
SiliconFlowfp8$0.70$2.201,048,576100.0%
Relaceunknown$0.10$4.001,048,576100.0%
CoreWeavefp4$0.76$2.421,048,576—
Waferunknown$0.41$3.991,048,57699.9%
AtlasCloudfp8$0.94$2.951,048,576100.0%
Alibaba / Qwenfp8$0.97$3.041,000,000100.0%
Phalafp8$1.26$3.001,048,576—
Cloudflareunknown$1.18$4.40262,144—
Inceptronfp4$1.25$4.391,048,576—
BaseTenfp8$1.40$4.401,048,576100.0%
BaseTenfp8$1.40$4.401,048,57699.4%
Fireworksunknown$1.40$4.401,048,5760.0%
Friendliunknown$1.40$4.401,048,576100.0%
GMICloudfp8$1.40$4.401,048,576—
Mistralnvfp4$1.40$4.401,048,576100.0%
Parasailfp4$1.40$4.40262,144100.0%
Togetherunknown$1.40$4.401,048,57599.8%
Venicefp8$1.40$4.401,000,000—
Z.ai / Zhipufp8$1.40$4.401,048,576100.0%
Mistralunknown$1.54$4.841,048,576—
BaseTenfp8$2.10$6.601,048,576100.0%
BaseTenfp8$2.10$6.601,048,576100.0%
Alibaba / Qwenfp8$2.31$7.261,000,000—
Decartfp4$2.25$8.001,048,576100.0%

Buying direct from Z.ai / Zhipu costs $1.40 in and $4.40 out per million, cached input $0.260. That is 890% dearer than the cheapest routed endpoint on a 3:1 blend.

GLM-5.3

40 providers · fp4, fp8, nvfp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Baidufp8$0.16$0.491,048,57699.5%
Novitafp8$0.42$1.321,048,57699.6%
Morphfp8$0.13$2.991,048,57699.2%
Rekaunknown$0.17$3.00262,14499.9%
Sail Researchfp8$0.20$3.401,048,57698.9%
Sail Researchfp8$0.20$3.401,048,57699.2%
InferenceNetunknown$0.22$3.391,048,576100.0%
Waferunknown$0.22$3.391,048,576100.0%
Waferunknown$0.22$3.391,048,576100.0%
DeepInfrafp4$0.56$2.501,048,57699.8%
SiliconFlowfp8$0.70$2.201,048,576100.0%
Relaceunknown$0.12$4.001,048,576100.0%
Makorafp4$0.24$3.93980,00099.8%
Phalaunknown$0.84$2.641,048,576100.0%
Inceptronfp4$0.60$3.391,048,576100.0%
DigitalOceanunknown$0.91$2.861,048,576—
GMICloudfp8$0.98$3.081,048,576—
AkashMLfp8$1.05$3.561,048,576100.0%
Alibaba / Qwenunknown$1.19$3.741,000,00099.8%
Decartfp4$1.19$3.741,048,576100.0%
Friendliunknown$1.26$3.961,048,576100.0%
AtlasCloudfp8$1.40$4.401,048,57661.4%
BaseTenfp4$1.40$4.401,048,57699.9%
BaseTenfp4$1.40$4.401,048,57699.6%
Cloudflareunknown$1.40$4.401,048,576100.0%
Crusoefp4$1.40$4.401,048,57699.9%
Fireworksunknown$1.40$4.401,048,576100.0%
Mistralnvfp4$1.40$4.401,048,57699.9%
Mistralnvfp4$1.40$4.401,048,57699.9%
Modalunknown$1.40$4.401,048,57699.9%
Nebiusfp4$1.40$4.401,024,000100.0%
Parasailfp8$1.40$4.401,048,576100.0%
PrimeIntellectunknown$1.40$4.401,048,576100.0%
Togetherunknown$1.40$4.401,048,57599.2%
Veniceunknown$1.40$4.401,000,000—
Z.ai / Zhipufp8$1.40$4.401,048,57699.9%
Mistralnvfp4$1.54$4.841,048,576—
BaseTenfp8$2.10$6.601,048,576100.0%
BaseTenfp8$2.10$6.601,048,57699.9%
Alibaba / Qwenunknown$2.80$8.801,000,000—

GLM-5.3-Flash

34 providers · fp4, fp8, nvfp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp4$0.07$0.251,048,576—
StreamLakefp8$0.08$0.281,024,00098.1%
Novitafp8$0.08$0.281,048,57699.1%
GMICloudfp8$0.09$0.301,048,57699.1%
Relaceunknown$0.04$0.501,048,57699.2%
Near AIfp8$0.10$0.351,048,57696.2%
DekaLLMunknown$0.10$0.401,048,57699.7%
Sail Researchfp4$0.04$0.601,048,57699.1%
InferenceNetfp4$0.05$0.601,048,57694.3%
Phalafp8$0.12$0.401,048,57699.0%
Waferunknown$0.10$0.501,048,57699.9%
Decartfp4$0.13$0.421,048,57699.9%
Io Netfp8$0.14$0.45262,124100.0%
Morphfp8$0.14$0.481,048,576100.0%
AtlasCloudfp8$0.15$0.501,048,57693.5%
BaseTenfp8$0.15$0.501,048,576100.0%
BaseTenfp8$0.15$0.501,048,576100.0%
CoreWeavenvfp4$0.15$0.501,048,57699.9%
Crusoefp4$0.15$0.501,048,57699.7%
DigitalOceanunknown$0.15$0.501,048,57697.6%
Fireworksunknown$0.15$0.501,048,57699.9%
Friendliunknown$0.15$0.501,048,57699.8%
Modalnvfp4$0.15$0.501,048,57699.9%
Parasailfp4$0.15$0.501,048,57699.1%
Rekaunknown$0.15$0.50262,14499.9%
SiliconFlowfp8$0.15$0.501,048,57697.6%
Togetherunknown$0.15$0.501,048,57599.9%
Veniceunknown$0.15$0.501,048,57698.0%
Z.ai / Zhipufp8$0.15$0.501,048,57699.9%
OpenInferencefp4$0.03$0.931,048,57699.3%
NextBitfp8$0.17$0.551,048,57699.9%
Inceptronfp8$0.22$0.451,048,57698.8%
Fireworksunknown$0.22$0.751,048,57699.7%
Cloudflareunknown$0.30$1.001,048,57699.9%

Gemini 3.1 Pro Preview

6 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$1.00$6.001,048,57691.5%
Google AI Studiounknown$1.00$6.001,048,576100.0%
Googleunknown$2.00$12.001,048,57699.1%
Google AI Studiounknown$2.00$12.001,048,57699.9%
Googleunknown$3.60$21.601,048,576—
Google AI Studiounknown$3.60$21.601,048,576—

Qwen3.7 Max

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$2.50$7.50——
Alibaba / Qwenunknown$1.48$4.421,000,000100.0%

Buying direct from Alibaba / Qwen costs $2.50 in and $7.50 out per million, cached input $0.500. That is 69% dearer than the cheapest routed endpoint on a 3:1 blend.

Claude Sonnet 4.6

9 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Anthropic directfirst-party$3.00$15.00——
Amazon Bedrockunknown$3.00$15.001,000,00097.1%
Anthropicunknown$3.00$15.001,000,000100.0%
Azureunknown$3.00$15.001,000,000—
Claude Platform on AWSunknown$3.00$15.001,000,000100.0%
Googleunknown$3.00$15.001,000,000100.0%
Amazon Bedrockunknown$3.30$16.501,000,000—
Amazon Bedrockunknown$3.30$16.501,000,000—
Googleunknown$3.30$16.501,000,000—
Googleunknown$3.30$16.501,000,000—

Buying direct from Anthropic costs $3.00 in and $15.00 out per million, cached input $0.300.

Kimi K2.6

18 providers · bf16, fp4, fp8, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Moonshot directfirst-party$0.95$4.00——
Baidufp4$0.43$1.83262,14499.9%
Inceptronint4$0.43$2.45262,144100.0%
DigitalOceanunknown$0.57$2.40262,144100.0%
Decartfp4$0.59$2.47262,14499.7%
StreamLakefp8$0.60$2.52256,00099.2%
Chutesint4$0.50$2.85262,14498.4%
CoreWeavefp4$0.65$3.41262,14499.9%
Crusoebf16$0.70$3.50262,144100.0%
SiliconFlowfp8$0.77$3.40262,144100.0%
DeepInfrafp4$0.75$3.50262,14491.7%
Parasailint4$0.75$3.50262,144100.0%
Veniceint4$0.75$3.50256,00074.1%
Novitaunknown$0.80$3.40262,14499.7%
GMICloudfp8$0.85$3.60262,144—
AtlasCloudint4$0.95$4.00262,144100.0%
Cloudflareunknown$0.95$4.00262,144100.0%
Moonshotint4$0.95$4.00262,144100.0%
Phalaunknown$1.09$4.60262,14499.6%

Buying direct from Moonshot costs $0.95 in and $4.00 out per million, cached input $0.160. That is 119% dearer than the cheapest routed endpoint on a 3:1 blend.

Kimi K2.7 Code

12 providers · fp4, fp8, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Moonshot directfirst-party$0.95$4.00——
StreamLakeunknown$0.71$3.00256,00098.5%
Inceptronint4$0.67$3.35262,14499.9%
CoreWeaveint4$0.71$3.50262,144100.0%
Veniceint4$0.75$3.50256,000—
SiliconFlowfp8$0.86$3.80262,144—
Novitaint4$0.91$3.84262,144100.0%
Alibaba / Qwenfp8$0.95$4.00262,144—
Cloudflareunknown$0.95$4.00262,144—
GMICloudfp8$0.95$4.00262,144—
Moonshotint4$0.95$4.00262,144100.0%
Nebiusfp4$0.95$4.00262,144—
Moonshotint4$1.90$8.00262,144—

Buying direct from Moonshot costs $0.95 in and $4.00 out per million, cached input $0.190. That is 33% dearer than the cheapest routed endpoint on a 3:1 blend.

Xiaomi MiMo-V2.5-Pro

6 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Xiaomi directfirst-party$0.43$0.87——
GMICloudbf16$0.30$0.611,050,00092.2%
AtlasCloudfp8$0.43$0.871,024,000—
Xiaomifp8$0.43$0.871,048,57699.1%
Novitaunknown$0.48$0.961,048,57699.3%
StreamLakeunknown$0.52$1.041,000,000—
DigitalOceanunknown$0.48$1.80262,144100.0%

Buying direct from Xiaomi costs $0.43 in and $0.87 out per million, cached input $0.004. That is 43% dearer than the cheapest routed endpoint on a 3:1 blend.

DeepSeek V4 Pro

20 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepSeek directfirst-party$0.66$1.98——
Baidufp8$0.13$0.401,048,576100.0%
Ionstreamunknown$0.21$1.481,048,57699.9%
DeepSeekunknown$0.66$1.981,048,576100.0%
StreamLakeunknown$0.66$1.981,024,00098.5%
Waferunknown$0.33$4.201,048,576100.0%
Relacefp4$0.90$2.701,048,576100.0%
Phalaunknown$0.96$2.881,048,576—
Novitafp8$0.99$2.971,048,576100.0%
GMICloudfp8$1.06$3.171,048,575100.0%
NextBitfp8$1.06$3.171,048,576—
DeepInfrafp8$1.30$2.601,048,576100.0%
Alibaba / Qwenunknown$1.12$3.371,000,00095.4%
CoreWeavefp8$1.31$3.961,048,576100.0%
AtlasCloudfp8$1.32$3.961,048,576—
Cloudflareunknown$1.32$3.961,048,57699.9%
DigitalOceanunknown$1.32$3.961,048,576—
Parasailfp8$1.32$3.961,048,576—
SiliconFlowfp8$1.32$3.961,048,576—
Togetherunknown$1.32$3.961,048,57697.9%
Veniceunknown$1.65$4.951,000,000—

Buying direct from DeepSeek costs $0.66 in and $1.98 out per million, cached input $0.022. That is 400% dearer than the cheapest routed endpoint on a 3:1 blend.

Tencent Hy3

6 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Tencent directfirst-party$0.15$0.59——
DeepInfrafp4$0.13$0.53262,14499.9%
Tencentfp8$0.13$0.53262,14499.8%
GMICloudbf16$0.14$0.58262,144100.0%
Novitaunknown$0.14$0.58262,14499.8%
Phalaunknown$0.15$0.64262,14499.7%
AtlasCloudfp8$0.20$0.80262,144100.0%

Buying direct from Tencent costs $0.15 in and $0.59 out per million, cached input $0.037. That is 13% dearer than the cheapest routed endpoint on a 3:1 blend. Priced in RMB: 1 / 4 / 0.25 per million (input / output / cached), converted at about 6.8 CNY to the dollar.

MiniMax-M3

13 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
MiniMax directfirst-party$0.30$1.20——
CoreWeavefp4$0.23$0.96262,14499.8%
GMICloudfp8$0.24$0.961,048,57699.2%
DeepInfrafp8$0.28$1.10524,28899.9%
AtlasCloudfp8$0.30$1.20524,30099.7%
MiniMaxfp8$0.30$1.20524,28899.6%
Novitafp8$0.30$1.201,000,000100.0%
Parasailfp8$0.30$1.201,048,57699.4%
StreamLakefp8$0.30$1.201,000,00097.7%
Togetherunknown$0.30$1.20524,28899.4%
Venicefp8$0.30$1.20524,28899.3%
Maraunknown$0.60$2.401,048,576—
SambaNovaunknown$0.60$2.401,048,576—
ModelRunfp4$0.75$3.001,048,57695.3%

Buying direct from MiniMax costs $0.30 in and $1.20 out per million, cached input $0.060. That is 27% dearer than the cheapest routed endpoint on a 3:1 blend.

Xiaomi MiMo-V2.5

5 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Xiaomi directfirst-party$0.14$0.28——
GMICloudfp8$0.12$0.241,050,00098.2%
Xiaomifp8$0.14$0.281,048,57698.8%
Novitafp8$0.17$0.341,048,57698.6%
StreamLakeunknown$0.17$0.341,000,00097.5%
Venicefp8$0.40$2.001,000,00098.4%

Buying direct from Xiaomi costs $0.14 in and $0.28 out per million, cached input $0.003. That is 18% dearer than the cheapest routed endpoint on a 3:1 blend.

DeepSeek V4 Flash

28 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepSeek directfirst-party$0.15$0.60——
StreamLakefp8$0.04$0.131,024,00099.9%
Sail Researchfp4$0.02$0.301,048,57699.9%
DeepInfrafp8$0.06$0.181,048,576100.0%
Sail Researchfp4$0.02$0.421,048,57699.9%
Rekaunknown$0.02$0.53262,14499.8%
DigitalOceanunknown$0.12$0.241,048,576100.0%
BaseTenfp8$0.13$0.261,048,57699.9%
BaseTenfp8$0.13$0.261,048,576100.0%
CoreWeavefp8$0.13$0.28262,144100.0%
Cohereunknown$0.14$0.281,048,57699.4%
Parasailfp8$0.14$0.281,048,57699.8%
Togetherunknown$0.14$0.281,048,57699.6%
Inceptronfp4$0.05$0.651,048,576100.0%
Morphbf16$0.14$0.401,048,576100.0%
Veniceunknown$0.17$0.351,000,00099.4%
OpenInferencefp8$0.00$1.041,048,57699.3%
Waferunknown$0.12$0.701,048,576100.0%
Mancer 2fp8$0.20$0.601,048,57698.8%
Relacefp4$0.01$1.281,048,57698.9%
SiliconFlowfp8$0.22$0.661,048,57699.4%
GMICloudfp8$0.29$0.861,048,575100.0%
Phalaunknown$0.31$0.921,048,576100.0%
Alibaba / Qwenunknown$0.35$1.061,000,00099.1%
NextBitfp8$0.35$1.061,048,576100.0%
Novitafp8$0.41$1.231,048,576100.0%
AtlasCloudfp4$0.44$1.321,048,576100.0%
Baidufp8$0.44$1.321,048,576100.0%
Cloudflareunknown$0.44$1.321,048,576100.0%

Buying direct from DeepSeek costs $0.15 in and $0.60 out per million, cached input $0.003. That is 298% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GPT-5.4 Nano

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.10$0.62400,000—
Azureunknown$0.20$1.25400,000100.0%
OpenAIunknown$0.20$1.25400,000100.0%
Azureunknown$0.22$1.38400,000—

GPT-5.4 Mini

5 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.38$2.25400,000100.0%
Azureunknown$0.75$4.50400,000100.0%
OpenAIunknown$0.75$4.50400,000100.0%
Azureunknown$0.83$4.95400,000—
OpenAIunknown$1.50$9.00400,000—

Qwen3.7 Plus

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.40$1.60——
Alibaba / Qwenunknown$0.32$1.281,000,000100.0%

Buying direct from Alibaba / Qwen costs $0.40 in and $1.60 out per million, cached input $0.040. That is 25% dearer than the cheapest routed endpoint on a 3:1 blend.

GLM-5.1

13 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$1.40$4.40——
Baidufp8$0.96$3.03202,75299.3%
StreamLakefp8$0.97$3.04200,00099.3%
Chutesfp8$0.98$3.08202,75292.3%
SiliconFlowfp8$1.19$3.74204,800100.0%
AtlasCloudfp8$1.26$3.96202,75298.5%
Phalaunknown$1.21$4.20202,75292.1%
Alibaba / Qwenfp8$1.33$4.18202,74599.3%
Novitafp8$1.38$4.40204,80099.0%
Friendliunknown$1.40$4.40202,752100.0%
GMICloudfp8$1.40$4.40202,75299.0%
Nebiusfp8$1.40$4.40202,752100.0%
Z.ai / Zhipufp8$1.40$4.40202,752100.0%
Venicefp8$1.40$4.40200,00095.1%

Buying direct from Z.ai / Zhipu costs $1.40 in and $4.40 out per million, cached input $0.260. That is 45% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3.6 Plus

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.50$3.00——
Alibaba / Qwenunknown$0.33$1.951,000,000100.0%

Buying direct from Alibaba / Qwen costs $0.50 in and $3.00 out per million, cached input $0.050. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3.6 27B

6 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.60$3.60——
Chutesfp8$0.30$2.00262,144100.0%
Phalaunknown$0.32$2.70262,14499.3%
Alibaba / Qwenunknown$0.45$2.70262,14499.3%
SiliconFlowfp8$0.30$3.20262,144—
DeepInfrafp8$0.32$3.20262,144100.0%
Venicefp8$0.33$3.25256,000—

Buying direct from Alibaba / Qwen costs $0.60 in and $3.60 out per million. That is 86% dearer than the cheapest routed endpoint on a 3:1 blend.

MiniMax M2.7

7 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
MiniMax directfirst-party$0.30$1.20——
GMICloudfp8$0.21$0.84196,60899.5%
Novitafp8$0.27$1.08204,80099.3%
AtlasCloudfp8$0.30$1.20196,60898.7%
MiniMaxfp8$0.30$1.20204,80099.9%
Groqunknown$0.60$1.80196,60897.0%
MiniMaxfp8$0.60$2.40204,800—
SambaNovaunknown$0.60$2.40196,608—

Buying direct from MiniMax costs $0.30 in and $1.20 out per million, cached input $0.060. That is 43% dearer than the cheapest routed endpoint on a 3:1 blend.

Inkling

2 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.95$4.05524,28898.6%
Togetherunknown$1.00$4.05524,288100.0%

Not listed above: Thinking Machines directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GPT-5.1

5 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.62$5.00400,000—
Azureunknown$1.25$10.00400,00099.9%
OpenAIunknown$1.25$10.00400,000100.0%
Azureunknown$1.38$11.00400,000—
OpenAIunknown$2.50$20.00400,000—

Gemini 3.5 Flash-Lite

8 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.15$1.251,048,576100.0%
Google AI Studiounknown$0.15$1.251,048,576100.0%
Googleunknown$0.30$2.501,048,576100.0%
Google AI Studiounknown$0.30$2.501,048,57699.9%
Googleunknown$0.33$2.751,048,576—
Googleunknown$0.33$2.751,048,576100.0%
Googleunknown$0.54$4.501,048,576100.0%
Google AI Studiounknown$0.54$4.501,048,576100.0%

Qwen3.5 397B A17B

10 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.60$3.60——
Alibaba / Qwenunknown$0.39$2.34262,14499.1%
DeepInfrafp8$0.45$3.00262,144100.0%
Parasailfp8$0.50$3.60262,14499.9%
AtlasCloudfp8$0.55$3.50262,14495.8%
DigitalOceanunknown$0.55$3.50131,07296.9%
Phalaunknown$0.55$3.50262,14498.7%
GMICloudfp8$0.60$3.60262,14499.4%
Novitaunknown$0.60$3.60262,14489.6%
StreamLakeunknown$0.60$3.60256,00097.5%
Veniceunknown$0.75$4.50128,00098.5%

Buying direct from Alibaba / Qwen costs $0.60 in and $3.60 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

Mistral Medium 3.5

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$1.50$7.50——
Mistralunknown$1.50$7.50262,144100.0%
Mistralunknown$1.50$7.50262,144100.0%
Mistralunknown$1.65$8.25262,144—

Buying direct from Mistral costs $1.50 in and $7.50 out per million, cached input $0.150.

Kimi K2.5

5 providers · int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
SiliconFlowint4$0.45$2.25262,144100.0%
AtlasCloudint4$0.49$2.50262,14499.6%
Novitaunknown$0.57$2.85262,14499.0%
Amazon Bedrockunknown$0.60$3.00262,144100.0%
Veniceunknown$0.53$3.32256,00097.1%

Not listed above: Moonshot directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GLM 4.6

4 providers · bf16, fp4
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$0.60$2.20——
Venicefp4$0.43$1.75198,00098.5%
DeepInfrafp4$0.50$2.00202,752100.0%
Novitabf16$0.55$2.20204,80099.7%
Z.ai / Zhipufp4$0.60$2.20202,75299.1%

Buying direct from Z.ai / Zhipu costs $0.60 in and $2.20 out per million, cached input $0.110. That is 32% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3.5-122B-A10B

4 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.40$3.20——
Alibaba / Qwenunknown$0.26$2.08262,14499.2%
SiliconFlowfp8$0.26$2.08262,14496.1%
AtlasCloudfp8$0.30$2.40262,14499.3%
Novitabf16$0.40$3.20262,14499.1%

Buying direct from Alibaba / Qwen costs $0.40 in and $3.20 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

LongCat 2.0

1 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
AtlasCloudfp8$0.30$1.201,048,756100.0%

Not listed above: Meituan directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GLM 4.7

7 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Z.ai / Zhipu directfirst-party$0.60$2.20——
DeepInfrafp4$0.40$1.75202,75299.8%
Venicefp4$0.40$1.93198,00099.8%
AtlasCloudfp8$0.52$1.85202,75286.4%
Novitafp8$0.54$1.98204,800100.0%
Googleunknown$0.60$2.20200,000100.0%
Z.ai / Zhipufp4$0.60$2.20202,75299.9%
Mancer 2fp4$0.70$2.50131,07298.5%

Buying direct from Z.ai / Zhipu costs $0.60 in and $2.20 out per million, cached input $0.110. That is 36% dearer than the cheapest routed endpoint on a 3:1 blend.

DeepSeek V3.2

13 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
GMICloudfp8$0.21$0.31163,84091.5%
AtlasCloudfp8$0.26$0.38163,84096.5%
DeepInfrafp4$0.26$0.38163,840100.0%
Veniceunknown$0.27$0.39160,000100.0%
SiliconFlowfp8$0.26$0.42163,84099.6%
Baidufp8$0.28$0.42131,072100.0%
DigitalOceanunknown$0.30$0.96163,84098.1%
Alibaba / Qwenfp8$0.37$1.11131,07296.9%
Friendliunknown$0.50$1.50163,840100.0%
Googleunknown$0.56$1.68163,84099.7%
Phalaunknown$1.00$1.00163,84098.3%
Maraunknown$3.00$4.5032,768—
SambaNovaunknown$3.00$4.5032,76898.7%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

DeepSeek V3.1 Terminus

3 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
SiliconFlowfp8$0.27$1.00163,84099.4%
AtlasCloudfp8$0.30$1.00131,07299.4%
StreamLakeunknown$0.34$1.03128,00099.1%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Gemma 4 31B

14 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp4$0.09$0.34262,14499.5%
DekaLLMunknown$0.10$0.33262,14496.5%
CoreWeavefp4$0.10$0.34262,14498.3%
Venicefp4$0.12$0.36256,00099.4%
Chutesfp4$0.12$0.37131,07277.3%
Crusoebf16$0.14$0.40262,14499.4%
Friendliunknown$0.14$0.40262,144100.0%
Novitabf16$0.14$0.40262,14476.7%
DeepInfrafp8$0.15$0.40262,14497.7%
Parasailfp8$0.15$0.40262,14496.9%
Io Netunknown$0.38$1.15262,144100.0%
SambaNovaunknown$0.38$1.15131,07298.0%
ModelRunfp4$0.75$1.00262,14499.7%
SiliconFlowfp8$0.75$1.00262,14439.4%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Grok 4.3

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAI directfirst-party$1.25$2.50——
xAIunknown$1.25$2.501,000,00099.8%
xAIunknown$1.25$2.501,000,00099.8%
xAIunknown$2.50$5.001,000,000—
xAIunknown$2.50$5.001,000,000—

Buying direct from xAI costs $1.25 in and $2.50 out per million, cached input $0.200.

Qwen3.6 35B A3B

9 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.25$1.49——
Darkbloomfp4$0.05$0.70262,144100.0%
AkashMLfp8$0.10$0.90262,144100.0%
DeepInfrafp8$0.10$0.95262,14499.9%
Venicefp8$0.10$1.00256,000100.0%
Parasailfp8$0.15$1.00262,144100.0%
AtlasCloudfp8$0.19$1.11262,14498.3%
Phalaunknown$0.20$1.27262,144100.0%
CoreWeavefp8$0.25$1.25262,144100.0%
SiliconFlowfp8$0.24$1.80262,144100.0%

Buying direct from Alibaba / Qwen costs $0.25 in and $1.49 out per million. That is 162% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

o1

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$15.00$60.00200,000—

Step 3.7 Flash

2 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
Novitafp8$0.20$1.15262,14499.8%
StepFunfp8$0.20$1.15256,00099.3%

Gemma 4 26B A4B

12 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Darkbloomunknown$0.04$0.22131,072100.0%
NextBitbf16$0.07$0.22262,144100.0%
DekaLLMbf16$0.06$0.33262,14499.9%
DeepInfrafp8$0.07$0.34262,14499.6%
Cloudflareunknown$0.10$0.30256,00099.9%
CoreWeavebf16$0.10$0.30262,144100.0%
Novitabf16$0.13$0.40262,14499.7%
Parasailbf16$0.13$0.40262,14499.9%
Venicebf16$0.13$0.40256,00099.8%
SiliconFlowfp8$0.14$0.40262,144100.0%
Io Netbf16$0.15$0.50262,142100.0%
Googleunknown$0.15$0.60262,14493.6%

GPT-5

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$1.25$10.00400,000100.0%
OpenAIunknown$1.25$10.00400,000100.0%
Azureunknown$1.38$11.00400,000—

Qwen3.5-35B-A3B

7 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.25$2.00——
Darkbloomfp4$0.08$0.75262,144100.0%
DeepInfrafp8$0.14$1.00262,144100.0%
Parasailfp8$0.15$1.00262,144100.0%
Alibaba / Qwenunknown$0.16$1.30262,14499.9%
Veniceunknown$0.31$1.25256,000100.0%
AtlasCloudfp8$0.22$1.80262,144—
SiliconFlowfp8$0.24$1.80262,144—

Buying direct from Alibaba / Qwen costs $0.25 in and $2.00 out per million. That is 178% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3 Coder Next

4 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Parasailbf16$0.12$0.80262,144100.0%
StreamLakeunknown$0.18$0.90256,00098.9%
Novitafp8$0.20$1.50262,144—
Alibaba / Qwenunknown$0.30$1.50262,144—

Gemini 3.1 Flash Lite

8 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.12$0.751,048,576100.0%
Google AI Studiounknown$0.12$0.751,048,576100.0%
Googleunknown$0.25$1.501,048,57699.9%
Google AI Studiounknown$0.25$1.501,048,57699.9%
Googleunknown$0.28$1.651,048,576100.0%
Googleunknown$0.28$1.651,048,576100.0%
Googleunknown$0.45$2.701,048,576100.0%
Google AI Studiounknown$0.45$2.701,048,576100.0%

Gemini 2.5 Pro

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Google AI Studiounknown$0.62$5.001,048,576—
Googleunknown$1.25$10.001,048,57699.3%
Googleunknown$1.25$10.001,048,57699.2%
Googleunknown$1.25$10.001,048,576—
Google AI Studiounknown$1.25$10.001,048,57699.7%
Googleunknown$2.25$18.001,048,576100.0%
Google AI Studiounknown$2.25$18.001,048,576—

Mercury 2

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Inceptionunknown$0.25$0.75128,000100.0%

gpt-oss-120b

23 providers · bf16, fp16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
CoreWeavefp4$0.03$0.17131,07298.8%
DekaLLMbf16$0.03$0.18131,07299.8%
DeepInfrabf16$0.04$0.17131,072100.0%
AkashMLbf16$0.04$0.19131,072100.0%
Crusoebf16$0.05$0.25131,072100.0%
Novitafp4$0.05$0.25131,07299.9%
Mancer 2fp8$0.04$0.28131,07296.5%
DigitalOceanunknown$0.06$0.42128,000100.0%
Googleunknown$0.09$0.36131,07251.9%
BaseTenfp4$0.10$0.50128,072100.0%
BaseTenfp4$0.10$0.50128,072100.0%
Amazon Bedrockunknown$0.15$0.60131,072100.0%
Amazon Bedrockunknown$0.15$0.60131,072100.0%
DeepInfrabf16$0.15$0.60131,072100.0%
Groqunknown$0.15$0.60131,07299.3%
Nebiusfp4$0.15$0.60131,072100.0%
Phalaunknown$0.15$0.60131,072100.0%
SiliconFlowfp8$0.15$0.60131,07276.7%
Togetherunknown$0.15$0.60131,072100.0%
Parasailfp4$0.10$0.75131,072100.0%
Maraunknown$0.15$0.75131,072100.0%
SambaNovaunknown$0.14$0.95131,072100.0%
Cerebrasfp16$0.35$0.75131,072100.0%

Not listed above: OpenAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3.5-9B

6 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Darkbloomfp4$0.08$0.13262,144100.0%
DeepInfrabf16$0.10$0.15262,144100.0%
SiliconFlowfp8$0.10$0.15262,14499.9%
Venicefp8$0.10$0.15256,000100.0%
Parasailbf16$0.10$0.25262,144100.0%
Togetherunknown$0.17$0.25262,14499.9%

Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Command A

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Cohereunknown$2.50$10.00256,000—

Mistral Small 4

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$0.15$0.60——
Mistralunknown$0.15$0.60262,144100.0%
Mistralunknown$0.15$0.60262,144100.0%
Mistralunknown$0.17$0.66262,144—
Mistralunknown$0.17$0.66262,144—

Buying direct from Mistral costs $0.15 in and $0.60 out per million, cached input $0.015.

Trinity Large Thinking

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Arcee AIunknown$0.25$0.80262,144—

Not listed above: Arcee Ai directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Ling-3.0-flash

2 providers · bf16, unknown
Served byQuant$/M in$/M outContextUptime 30m
Novitaunknown$0.02$0.06262,144100.0%
DeepInfrabf16$0.06$0.18131,07297.3%

Not listed above: InclusionAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Muse Spark 1.2

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Metaunknown$1.25$4.251,048,576100.0%

Qwen3.8 Max

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwenunknown$2.00$6.001,000,00099.8%

Qwen3.8-27B

16 providers · bf16, fp16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.15$1.88262,14499.7%
Phalaunknown$0.15$1.881,000,00098.8%
Darkbloomfp4$0.05$2.20262,144100.0%
AkashMLfp8$0.20$1.78262,144100.0%
Ionstreamfp8$0.09$2.20262,14499.6%
Chutesfp8$0.24$2.20262,14499.2%
Parasailfp8$0.24$2.20262,144100.0%
Mancer 2fp8$0.20$2.50262,14499.8%
DekaLLMunknown$0.05$3.00262,14499.8%
Alibaba / Qwenunknown$0.42$2.551,000,00099.8%
CoreWeavefp8$0.40$3.00262,144100.0%
Novitaunknown$0.42$3.001,000,00098.8%
Waferunknown$0.02$4.35262,144100.0%
Cerebrasfp16$0.99$1.4965,536100.0%
Cloudflareunknown$0.45$3.20262,14487.7%
Venicefp8$0.45$3.20262,14499.4%

GPT-4o

2 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$2.50$10.00128,00098.8%
OpenAIunknown$2.50$10.00128,000100.0%

DeepSeek V3

2 providers · fp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
StreamLakeunknown$0.26$1.03128,00099.2%
DeepInfrafp4$0.32$0.89163,84092.7%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GPT-4 Turbo

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$10.00$30.00128,000—

DeepSeek V3 0324

2 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
SiliconFlowfp8$0.25$1.00163,84099.9%
GMICloudfp8$0.29$1.14128,00099.1%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

gpt-oss-20b

12 providers · bf16, fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Darkbloomfp8$0.02$0.09131,07299.9%
AkashMLfp4$0.02$0.10131,07299.5%
CoreWeavefp4$0.03$0.13131,072100.0%
DekaLLMbf16$0.03$0.14131,07299.9%
DeepInfrabf16$0.03$0.14131,072100.0%
Parasailfp4$0.03$0.15131,072100.0%
Novitafp4$0.04$0.15131,072100.0%
SiliconFlowfp8$0.04$0.18131,07297.0%
Amazon Bedrockunknown$0.07$0.15131,072—
Amazon Bedrockunknown$0.07$0.15131,07299.7%
Googleunknown$0.07$0.25131,072100.0%
Groqunknown$0.07$0.30131,07295.9%

Not listed above: OpenAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Mistral Medium 3.1

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistralunknown$0.40$2.00131,072100.0%
Mistralunknown$0.40$2.00131,072100.0%
Mistralunknown$0.44$2.20131,072—

GPT-4.1 Mini

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$0.40$1.601,047,576100.0%
OpenAIunknown$0.40$1.601,047,57699.8%
Azureunknown$0.44$1.761,047,576—

Llama 4 Maverick

4 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DigitalOceanunknown$0.19$0.65128,00099.9%
Novitafp8$0.27$0.851,048,57698.9%
Parasailfp8$0.35$1.00524,28899.9%
Googleunknown$0.35$1.15524,288—

Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

o3 Mini

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.10$4.40200,000100.0%

Solar Pro 3

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Upstageunknown$0.15$0.60131,072100.0%

GPT-5 Mini

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.12$1.00400,000100.0%
Azureunknown$0.25$2.00400,00099.9%
OpenAIunknown$0.25$2.00400,000100.0%
Azureunknown$0.28$2.20400,000—

Qwen3 32B

2 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.70$2.80——
DeepInfrafp8$0.08$0.2840,960100.0%
SiliconFlowfp8$0.14$0.57131,072100.0%

Buying direct from Alibaba / Qwen costs $0.70 in and $2.80 out per million. That is 842% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Alibaba / Qwen directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3 14B

3 providers · fp8, int4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.35$1.40——
NextBitint4$0.10$0.2240,960100.0%
DeepInfrafp8$0.12$0.2440,96097.0%
Alibaba / Qwenunknown$0.23$0.91131,072100.0%

Buying direct from Alibaba / Qwen costs $0.35 in and $1.40 out per million. That is 371% dearer than the cheapest routed endpoint on a 3:1 blend.

GPT-4

2 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$30.00$60.008,191—
OpenAIunknown$30.00$60.008,191—

GPT-4o-mini

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$0.15$0.60128,000100.0%
OpenAIunknown$0.15$0.60128,00099.8%
Azureunknown$0.17$0.66128,000—

GPT-4.1 Nano

3 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Azureunknown$0.10$0.401,047,576100.0%
OpenAIunknown$0.10$0.401,047,576100.0%
Azureunknown$0.11$0.441,047,576—

GPT-3.5 Turbo

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.50$1.5016,385100.0%

Granite 4.2 8B

2 providers · bf16
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.06$0.25131,072100.0%
CoreWeavebf16$0.10$0.15131,072100.0%

Not listed above: IBM directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3 8B

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.18$0.70——
Alibaba / Qwenunknown$0.12$0.45131,072100.0%

Buying direct from Alibaba / Qwen costs $0.18 in and $0.70 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

Llama 4 Scout

3 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.10$0.30327,68099.8%
Novitabf16$0.18$0.59131,07299.9%
Googleunknown$0.25$0.701,310,720—

Not listed above: Meta directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Grok Build 0.1

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAI directfirst-party$1.00$2.00——
xAIunknown$1.00$2.00256,000100.0%
xAIunknown$1.00$2.00256,000—
xAIunknown$2.00$4.00256,000—
xAIunknown$2.00$4.00256,000—

Buying direct from xAI costs $1.00 in and $2.00 out per million, cached input $0.200.

Nemotron 3 Ultra 550B

4 providers · fp4, fp8
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.50$2.50——
DeepInfrafp4$0.50$2.20262,14499.7%
BaseTenfp4$0.60$2.40202,800100.0%
BaseTenfp4$0.60$2.40202,800100.0%
Venicefp8$0.62$3.12256,00099.4%

Buying direct from Nvidia costs $0.50 in and $2.50 out per million, cached input $0.150. That is 8% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Nemotron 3 Super 120B

2 providers · bf16, fp8
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.20$0.80——
DeepInfrabf16$0.08$0.40262,144100.0%
DekaLLMfp8$0.08$0.45262,14499.9%

Buying direct from Nvidia costs $0.20 in and $0.80 out per million. That is 114% dearer than the cheapest routed endpoint on a 3:1 blend.

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

North Mini Code

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Cohere directfirst-party$0.00$0.00——
Cohereunknown$0.00$0.00256,00095.4%

Buying direct from Cohere costs $0.00 in and $0.00 out per million.

Mistral Small 3.1

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Cloudflareunknown$0.35$0.55128,00099.1%

Not listed above: Mistral directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

DeepSeek R1

1 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
Novitafp8$0.70$2.5064,000100.0%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Qwen3 235B A22B 2507

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwen directfirst-party$0.70$2.80——
Alibaba / Qwenunknown$0.45$1.82131,07297.1%

Buying direct from Alibaba / Qwen costs $0.70 in and $2.80 out per million. That is 54% dearer than the cheapest routed endpoint on a 3:1 blend.

Nemotron 3 Nano 30B

4 providers · fp4, fp8
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.00$0.00——
Crusoefp8$0.05$0.20262,144100.0%
DeepInfrafp4$0.05$0.20262,14492.8%
Novitafp4$0.05$0.20262,144100.0%
Nebiusfp8$0.06$0.24262,144100.0%

Buying direct from Nvidia costs $0.00 in and $0.00 out per million.

Not listed above: Nvidia directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Nemotron 3 Nano Omni 30B

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Nvidia directfirst-party$0.00$0.00——
Nvidiaunknown$0.00$0.00256,00080.0%

Buying direct from Nvidia costs $0.00 in and $0.00 out per million.

Mistral Small 3.2

4 providers · bf16, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$0.10$0.30——
DeepInfrafp8$0.07$0.20128,000100.0%
Venicefp8$0.09$0.25256,00099.9%
Parasailbf16$0.09$0.30131,072100.0%
Mistralunknown$0.10$0.3032,768—

Buying direct from Mistral costs $0.10 in and $0.30 out per million. That is 41% dearer than the cheapest routed endpoint on a 3:1 blend.

Qwen3 30B A3B 2507

2 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.12$0.5040,960100.0%
Alibaba / Qwenunknown$0.13$0.52131,072100.0%

Gemma 3 27B

4 providers · bf16, fp8
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.08$0.16131,07299.4%
Novitabf16$0.12$0.2098,30499.9%
Nebiusfp8$0.10$0.30110,000100.0%
Parasailfp8$0.08$0.45131,072100.0%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Gemma 3 12B

1 providers · bf16
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.05$0.15131,072100.0%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Gemma 3 4B

1 providers · bf16
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrabf16$0.05$0.10131,072100.0%

Not listed above: Google directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

GPT-6 Astra

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$5.00$25.001,050,000100.0%
Azureunknown$10.00$50.001,050,000100.0%
OpenAIunknown$10.00$50.001,050,000100.0%
Amazon Bedrockunknown$11.00$55.001,050,000—
Azureunknown$11.00$55.001,050,000—
OpenAIunknown$20.00$100.001,050,000100.0%
OpenAIunknown$60.00$300.001,050,000—

Claude Fable 5.1

4 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Amazon Bedrockunknown$10.00$50.001,000,000—
Anthropicunknown$10.00$50.001,000,00099.9%
Azureunknown$10.00$50.001,000,000—
Googleunknown$10.00$50.001,000,000—

Gemini 3.8 Flash

6 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Googleunknown$0.38$1.881,048,57688.8%
Google AI Studiounknown$0.38$1.881,048,57699.9%
Googleunknown$0.75$3.751,048,57697.9%
Google AI Studiounknown$0.75$3.751,048,576100.0%
Googleunknown$1.35$6.751,048,576100.0%
Google AI Studiounknown$1.35$6.751,048,576100.0%

Muse Spark 1.3

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Metaunknown$1.25$4.251,048,57699.7%

Qwen3.8 2.4T A95B

7 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Alibaba / Qwenunknown$2.00$6.001,000,000100.0%
DeepInfrafp4$2.00$6.00262,144—
Modalunknown$2.00$6.001,000,000—
Novitaunknown$2.00$6.001,000,000100.0%
SiliconFlowfp8$2.00$6.001,048,576100.0%
Togetherunknown$2.00$6.001,010,000100.0%
Veniceunknown$2.00$6.00262,144—

DeepSeek V4 Flash Vision

4 providers · fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.22$0.651,048,57699.9%
GMICloudfp8$0.44$1.321,048,575100.0%
Novitaunknown$0.44$1.321,048,576100.0%
SiliconFlowfp8$0.44$1.321,048,576100.0%

Not listed above: DeepSeek directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Ling-3.0-flash-VL

2 providers · bf16, fp16
Served byQuant$/M in$/M outContextUptime 30m
Novitabf16$0.02$0.06262,14496.7%
DeepInfrafp16$0.06$0.18131,07299.9%

Not listed above: InclusionAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Devstral 2

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$0.40$2.00——
Mistralunknown$0.40$2.00262,144100.0%

Buying direct from Mistral costs $0.40 in and $2.00 out per million.

Mistral Large 3

2 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Mistral directfirst-party$0.50$1.50——
Mistralunknown$0.50$1.50262,144100.0%
Mistralunknown$0.55$1.65262,144—

Buying direct from Mistral costs $0.50 in and $1.50 out per million, cached input $0.050.

Command A+

1 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Cohereunknown$0.30$1.50192,00099.1%

Claude Sonnet 5.5

8 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Amazon Bedrockunknown$2.00$10.001,000,000100.0%
Anthropicunknown$2.00$10.001,000,000100.0%
Azureunknown$2.00$10.001,000,000100.0%
Claude Platform on AWSunknown$2.00$10.001,000,000100.0%
Googleunknown$2.00$10.001,000,000100.0%
Azureunknown$2.20$11.001,000,000—
Googleunknown$2.20$11.001,000,000—
Googleunknown$2.20$11.001,000,000100.0%

GPT-6.1 Sol

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.00$5.001,050,00099.6%
Azureunknown$2.00$10.001,050,000100.0%
OpenAIunknown$2.00$10.001,050,00099.9%
Amazon Bedrockunknown$2.20$11.001,050,000—
Azureunknown$2.20$11.001,050,000100.0%
Azureunknown$2.20$11.001,050,000100.0%
OpenAIunknown$4.00$20.001,050,000—

MiMo-V2.6-Flash

8 providers · bf16, fp4, fp8
Served byQuant$/M in$/M outContextUptime 30m
Darkbloomfp4$0.12$0.281,048,57698.8%
InferenceNetfp8$0.12$0.281,048,57699.5%
GMICloudbf16$0.13$0.271,050,00092.0%
DeepInfrafp8$0.14$0.281,048,57699.1%
Io Netfp8$0.14$0.281,048,32096.3%
Novitafp8$0.14$0.281,048,57699.1%
Xiaomifp8$0.14$0.281,048,57699.6%
Venicefp8$0.17$0.351,000,00098.9%

Solar Mini 4

2 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Upstageunknown$0.05$0.20524,288100.0%
Upstageunknown$0.05$0.20524,288100.0%

Claude Opus 5.5

11 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
Amazon Bedrockunknown$4.00$20.001,000,00099.8%
Anthropicunknown$4.00$20.001,000,000100.0%
Azureunknown$4.00$20.001,000,000100.0%
Claude Platform on AWSunknown$4.00$20.001,000,000100.0%
Googleunknown$4.00$20.001,000,00099.8%
Amazon Bedrockunknown$4.40$22.001,000,000—
Amazon Bedrockunknown$4.40$22.001,000,000100.0%
Azureunknown$4.40$22.001,000,000—
Googleunknown$4.40$22.001,000,000—
Googleunknown$4.40$22.001,000,000—
Anthropicunknown$8.00$40.001,000,000100.0%

GPT-6 Sol

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$1.00$5.001,050,000100.0%
Azureunknown$2.00$10.001,050,000100.0%
OpenAIunknown$2.00$10.001,050,000100.0%
Amazon Bedrockunknown$2.20$11.001,050,000—
Azureunknown$2.20$11.001,050,000100.0%
Azureunknown$2.20$11.001,050,000—
OpenAIunknown$4.00$20.001,050,000100.0%

GPT-6 Luna

7 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
OpenAIunknown$0.05$0.251,050,00099.8%
Azureunknown$0.10$0.501,050,000100.0%
OpenAIunknown$0.10$0.501,050,000100.0%
Amazon Bedrockunknown$0.11$0.551,050,000100.0%
Azureunknown$0.11$0.551,050,000100.0%
Azureunknown$0.11$0.551,050,000100.0%
OpenAIunknown$0.20$1.001,050,000100.0%

Grok 4.7

5 providers · unknown
Served byQuant$/M in$/M outContextUptime 30m
xAIunknown$2.00$6.00500,00099.8%
xAIunknown$2.00$6.00500,00098.9%
xAIunknown$2.20$6.60500,000—
xAIunknown$4.00$12.00500,000—
xAIunknown$4.00$12.00500,000—

MiMo-V2.6-Pro

4 providers · bf16, fp8
Served byQuant$/M in$/M outContextUptime 30m
GMICloudbf16$0.41$0.831,050,00098.7%
DeepInfrafp8$0.43$0.871,048,57699.9%
Novitafp8$0.43$0.871,048,57699.8%
Xiaomifp8$0.43$0.871,048,57699.8%

DeepSeek V4.1 Flash

31 providers · fp4, fp8, unknown
Served byQuant$/M in$/M outContextUptime 30m
Morphfp8$0.02$0.381,048,57699.7%
Decartfp4$0.09$0.181,048,576100.0%
Sail Researchfp4$0.08$0.401,048,57699.9%
Relaceunknown$0.02$0.601,048,57699.9%
InferenceNetunknown$0.04$0.601,040,00099.8%
OpenInferencefp4$0.01$0.681,048,57694.4%
Waferunknown$0.05$0.601,048,576100.0%
DekaLLMunknown$0.12$0.401,048,57699.7%
DeepInfrafp8$0.14$0.421,048,57699.9%
Baidufp8$0.14$0.561,048,57699.0%
AtlasCloudfp8$0.14$0.561,048,57699.0%
DeepSeekunknown$0.15$0.601,048,576100.0%
CoreWeavefp8$0.20$0.651,048,57699.2%
StreamLakefp8$0.18$0.721,024,00098.9%
DigitalOceanunknown$0.18$0.721,048,576100.0%
GMICloudfp8$0.18$0.721,048,57599.8%
Ionstreamunknown$0.10$1.151,048,576100.0%
NextBitfp8$0.21$0.841,048,576100.0%
Phalaunknown$0.21$0.841,048,57699.3%
Makorafp8$0.20$0.991,048,57699.9%
Novitafp8$0.24$0.961,048,57695.6%
Alibaba / Qwenunknown$0.30$1.201,000,000100.0%
BaseTenfp8$0.30$1.201,048,57699.8%
BaseTenfp8$0.30$1.201,048,57699.9%
Fireworksunknown$0.30$1.201,048,57699.9%
Modalunknown$0.30$1.201,048,576100.0%
Parasailfp8$0.30$1.201,048,57699.9%
SiliconFlowfp8$0.30$1.201,048,57699.9%
Togetherunknown$0.30$1.201,048,576100.0%
Venicefp8$0.38$1.501,000,00099.9%
Fireworksunknown$0.45$1.801,048,576100.0%

Ling-3.0-flash-Fin

2 providers · fp4, unknown
Served byQuant$/M in$/M outContextUptime 30m
Novitaunknown$0.04$0.12262,144100.0%
DeepInfrafp4$0.06$0.18262,144100.0%

Not listed above: InclusionAI directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

Inkling Small

1 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
DeepInfrafp8$0.45$1.20524,288100.0%

Not listed above: Thinking Machines directly. The lab sells this model from its own API but does not route it through the aggregator these figures come from, so its price and precision cannot be compared here.

LFM2.5-2.6B

1 providers · fp8
Served byQuant$/M in$/M outContextUptime 30m
Liquidfp8$0.00$0.0065,53699.5%

What it actually costs

Observed, not list price

Aggregates across opencode’s subscriber base: real cost per coding session after caching. The spread is far wider than list prices suggest, and the cache ratio explains much of it.

Model$ / sessionEffective $/MCache ratioTokens / sessionSessions observed
Xiaomi MiMo-V2.5$0.0032$0.0093.3%1.0M58,650,321
DeepSeek V4 Flash$0.0475$0.0194.9%7.3M37,030,883
Xiaomi MiMo-V2.5-Pro$0.0785$0.0296.5%3.6M426,008
MiniMax-M3$0.1033$0.0593.8%2.0M1,446,708
Qwen3.7 Plus$0.2007$0.1688.3%1.2M881,391
DeepSeek V4 Pro$0.3103$0.0396.9%11.7M1,457,532
Kimi K2.6$0.3773$0.2990.8%1.3M219,741
Grok 4.5$0.4024$0.7781.4%0.5M72,580
GLM-5.1$0.8457$0.5086.0%1.7M95,327
Kimi K2.7 Code$0.8896$0.2595.2%3.6M430,564
Qwen3.7 Max$0.9202$0.9188.8%1.0M87,289
Kimi K3$1.1803$0.8387.7%1.4M721,353
GLM-5.2$1.5959$0.4884.4%3.3M610,782

Best value

Two ways to pay, two different answers

Value means something different depending on how you are billed, so this splits in two. On a subscription what matters is how much quality-weighted work a fixed monthly fee yields; on metered billing it is quality per unit of token spend. The same model can look strong under one and ordinary under the other.

On a subscription

Intelligence index times sessions per month, divided by what the plan costs. It answers a narrower question than it looks — how much quality-weighted work a subscription yields, assuming the cheap model is good enough for the job in hand.

#PlanModel$/moSessions/moIntelligenceValue: index x sessions per $
1opencode Go PlusXiaomi MiMo-V2.5$4075,00025.2██████████████████████████████████████████████ 47,250
2Chutes PlusXiaomi MiMo-V2.5$1015,62525.2██████████████████████████████████████ 39,375
3Chutes ProXiaomi MiMo-V2.5$2031,25025.2██████████████████████████████████████ 39,375
4Ollama ProXiaomi MiMo-V2.5$2018,75025.2███████████████████████ 23,625
5Ollama MaxXiaomi MiMo-V2.5$10093,75025.2███████████████████████ 23,625
6GitHub MaxXiaomi MiMo-V2.5$10062,50025.2███████████████ 15,750
7Ollama TeamXiaomi MiMo-V2.5$500312,50025.2███████████████ 15,750
8Morph ScaleXiaomi MiMo-V2.5$200125,00025.2███████████████ 15,750
9GitHub Pro+Xiaomi MiMo-V2.5$3921,87525.2██████████████ 14,135
10opencode GoXiaomi MiMo-V2.5$104,68825.2████████████ 11,812
11GitHub ProXiaomi MiMo-V2.5$104,68825.2████████████ 11,812
12Qoder ProXiaomi MiMo-V2.5$208,31225.2██████████ 10,474
13Qoder UltraXiaomi MiMo-V2.5$20083,12525.2██████████ 10,474
14Qoder Pro+Xiaomi MiMo-V2.5$6024,93725.2██████████ 10,474
Absence from this table is not a verdict. Only plans that publish a dollar budget can be ranked here, which is three vendors out of fourteen. The rest are excluded for being unmeasurable, not for being poor value — and several of them are deliberately generous. MiniMax is the clearest case: it publishes no quota at all, yet the figure reverse-engineered further down this page (about 32B tokens a month on the $120 tier) works out at roughly 4,200 sessions, which would place it third here — ahead of every GitHub and Cursor tier listed. Z.ai, Alibaba and Kimi meter in prompts or requests with no published token cost, so they cannot be placed at all. Read the table as the best of what can be measured, never as the best available.

Read it with the leaderboard beside you, too. A weak model on a generous plan can top this table while still being the wrong tool: the metric rewards volume, and quality enters only linearly. It is a guide to where the money goes furthest, not to what you should run for work that has to be right first time.

Pay per token

Intelligence index per dollar per million tokens, blended 3:1. The list column ranks what you are quoted; the observed column uses the effective rate actually paid in production once caching is working, which is the number that decides a monthly bill.

#ModelIntelligence$/M listValue on list$/M observedValue observed
1Ling-3.0-flash-VL · InclusionAI24.6$0.031790——
2Ling-3.0-flash · InclusionAI20.1$0.032638——
3Ling-3.0-flash-Fin · InclusionAI22.6$0.062363——
4gpt-oss-20b · OpenAI10.0$0.036278——
5Solar Mini 4 · Upstage24.1$0.087275——
6MiMo-V2.6-Flash · Xiaomi37.9$0.175217——
7GPT-6 Luna · OpenAI38.1$0.200190——
8Solar Pro 4 · Upstage28.2$0.158179——
9GLM-5.3-Flash · Z.ai / Zhipu41.8$0.237176——
10gpt-oss-120b · OpenAI11.6$0.070165——
11Gemma 4 26B A4B · Google16.7$0.107156——
12Nemotron 3.5 Lightning · Nvidia12.9$0.087148——
13Xiaomi MiMo-V2.5 · Xiaomi25.2$0.175144$0.00—
14Qwen3.5-9B · Alibaba / Qwen13.3$0.112118——
15Tencent Hy3 · Tencent25.3$0.231110$0.00—
16DeepSeek V4 Flash Vision · DeepSeek34.8$0.323108——
17DeepSeek V4 Flash · DeepSeek34.3$0.324106$0.013,430
18Granite 4.2 8B · IBM11.1$0.107103——
19Nemotron 3 Nano 30B · Nvidia8.9$0.087102——
20Gemma 4 31B · Google14.7$0.15296——
21MiMo-V2.6-Pro · Xiaomi46.3$0.54485——
22GPT-5.6 Luna · OpenAI37.3$0.45083——
23Gemma 3 4B · Google4.8$0.06277——
24DeepSeek V4.1 Flash · DeepSeek39.5$0.52575——
25Nemotron 3 Super 120B · Nvidia12.8$0.17274——
26DeepSeek V3.2 · DeepSeek21.5$0.31568——
27Qwen3 32B · Alibaba / Qwen8.6$0.13066——
28MiniMax M2.7 · MiniMax22.8$0.36762——
29Mistral Small 3.2 · Mistral8.2$0.13362——
30MiniMax-M3 · MiniMax29.2$0.52556$0.05584
31Qwen3 14B · Alibaba / Qwen8.2$0.15055——
32Llama 4 Scout · Meta8.1$0.15054——
33Qwen3.5-35B-A3B · Alibaba / Qwen19.3$0.36253——
34Gemma 3 12B · Google3.8$0.07551——
35Qwen3.6 35B A3B · Alibaba / Qwen18.2$0.36250——
36Xiaomi MiMo-V2.5-Pro · Xiaomi26.0$0.54448$0.021,300
37Qwen3 30B A3B 2507 · Alibaba / Qwen9.8$0.21546——
38Qwen3.7 Plus · Alibaba / Qwen25.2$0.56045$0.16158
39GPT-5.4 Nano · OpenAI20.7$0.46345——
40GPT-4.1 Nano · OpenAI7.8$0.17545——
41Step 3.7 Flash · StepFun19.5$0.43845——
42Mistral Small 4 · Mistral11.3$0.26243——
43Inkling Small · Thinking Machines25.7$0.63740——
44Qwen3.6 Plus · Alibaba / Qwen27.0$0.73137——
45Mercury 2 · Inception13.8$0.37537——
46LongCat 2.0 · Meituan19.1$0.52536——
47DeepSeek V4 Pro · DeepSeek36.0$0.99036$0.031,200
48Qwen3 8B · Alibaba / Qwen7.3$0.20236——
49Kimi K2.6 · Moonshot27.0$0.78334$0.2993
50Llama 4 Maverick · Meta10.0$0.30433——
51DeepSeek V3.1 Terminus · DeepSeek14.8$0.45333——
52Qwen3 Coder Next · Alibaba / Qwen9.2$0.29032——
53Qwen3.8-27B · Alibaba / Qwen33.7$1.06532——
54GPT-5 Mini · OpenAI20.6$0.68830——
55Solar Pro 3 · Upstage7.8$0.26230——
56Gemma 3 27B · Google4.9$0.17228——
57Trinity Large Thinking · Arcee Ai10.8$0.38728——
58Gemini 3.1 Flash Lite · Google15.6$0.56228——
59Muse Glimmer · Meta17.5$0.63727——
60Gemini 3.8 Flash · Google40.9$1.50027——
61Gemini 3.7 Flash · Google39.6$1.50026——
62Gemini 3.5 Flash-Lite · Google22.2$0.85026——
63Kimi K2.5 · Moonshot23.5$0.90026——
64GLM-5.2 · Z.ai / Zhipu33.7$1.30526$0.4870
65GPT-4o-mini · OpenAI6.7$0.26226——
66Nemotron 3 Ultra 550B · Nvidia22.9$0.92525——
67Qwen3.5-122B-A10B · Alibaba / Qwen17.7$0.71525——
68GLM 4.6 · Z.ai / Zhipu18.5$0.76024——
69Muse Spark 1.3 · Meta48.1$2.00024——
70Gemini 3.6 Flash · Google34.0$1.50023——
71GLM 4.7 · Z.ai / Zhipu22.2$1.00022——
72Command A+ · Cohere13.1$0.60022——
73Grok Build 0.1 · xAI27.2$1.25022——
74GLM-5.3 · Z.ai / Zhipu44.8$2.15021——
75Qwen3.6 27B · Alibaba / Qwen21.4$1.04021——
76Muse Spark 1.2 · Meta39.6$2.00020——
77DeepSeek V3 0324 · DeepSeek9.7$0.50219——
78Kimi K2.7 Code · Moonshot25.8$1.34119$0.25103
79DeepSeek V3 · DeepSeek8.5$0.45019——
80Mistral Small 3.1 · Mistral7.1$0.40218——
81Muse Spark 1.1 · Meta33.7$2.00017——
82Qwen3.5 397B A17B · Alibaba / Qwen21.4$1.28817——
83Qwen3 235B A22B 2507 · Alibaba / Qwen12.7$0.79616——
84Grok 4.3 · xAI24.9$1.56216——
85Grok 4.7 · xAI46.4$3.00015——
86Qwen3.8 Max · Alibaba / Qwen45.4$3.00015——
87Grok 4.6 · xAI44.3$3.00015——
88GPT-4.1 Mini · OpenAI10.2$0.70015——
89Inkling · Thinking Machines25.0$1.72514——
90GPT-5.4 Mini · OpenAI24.1$1.68814——
91Claude Sonnet 5.5 · Anthropic56.0$4.00014——
92Qwen3.7 Max · Alibaba / Qwen29.5$2.21313$0.9132
93Qwen3.8 2.4T A95B · Alibaba / Qwen39.9$3.00013——
94GPT-6.1 Sol · OpenAI51.8$4.00013——
95Grok 4.5 · xAI38.8$3.00013$0.7750
96Mistral Large 3 · Mistral9.3$0.75012——
97GLM-5.1 · Z.ai / Zhipu26.1$2.15012$0.5052
98GPT-6 Sol · OpenAI47.6$4.00012——
99GPT-5.6 Sol · OpenAI47.0$4.00012——
100Mistral Medium 3.1 · Mistral9.2$0.80011——
101Devstral 2 · Mistral8.6$0.80011——
102Gemini 3.5 Flash · Google33.6$3.37510——
103DeepSeek R1 · DeepSeek11.4$1.15010——
104Claude Sonnet 5 · Anthropic38.2$4.00010——
105GPT-5.6 Terra · OpenAI42.1$4.5009——
106Kimi K3 · Moonshot43.6$5.4008$0.8353
107GPT-3.5 Turbo · OpenAI5.5$0.7507——
108Claude Opus 5.5 · Anthropic57.6$8.0007——
109GPT-5.1 · OpenAI24.7$3.4387——
110GPT-5.4 · OpenAI39.0$5.6257——
111GPT-5 · OpenAI23.0$3.4387——
112Gemini 3.1 Pro Preview · Google29.7$4.5007——
113o3 Mini · OpenAI12.5$1.9256——
114Claude Opus 5 · Anthropic50.8$10.0005——
115Claude Sonnet 4.6 · Anthropic30.1$6.0005——
116Mistral Medium 3.5 · Mistral14.2$3.0005——
117Gemini 2.5 Pro · Google16.1$3.4385——
118Claude Opus 4.8 · Anthropic41.8$10.0004——
119Claude Opus 4.7 · Anthropic40.7$10.0004——
120GPT-5.5 · OpenAI38.4$11.2503——
121Command A · Cohere13.1$4.3753——
122Claude Fable 5.1 · Anthropic53.4$20.0003——
123GPT-6 Astra · OpenAI52.7$20.0003——
124Claude Fable 5 · Anthropic49.6$20.0002——
125GPT-4o · OpenAI9.0$4.3752——
126o1 · OpenAI15.2$26.2501——
127GPT-4 Turbo · OpenAI7.0$15.0000——
128GPT-4 · OpenAI6.7$37.5000——

The two columns do not agree, and that disagreement is the point: list price is a quote, the observed rate is a bill. A model with a poor cache ratio slides down the observed column even when its quoted price looked competitive. Observed rates come from opencode's subscriber base, so models it does not serve show a dash rather than a guess.

What a plan actually buys

Included usage divided by observed cost per session. Only dollar-denominated plans can be answered this precisely — plans metered in prompts, or with no published quota, cannot appear here at all.

Plan$/moIncludedXiaomi MiMo-V2.5DeepSeek V4 FlashXiaomi MiMo-V2.5-ProMiniMax-M3Qwen3.7 PlusDeepSeek V4 ProKimi K2.6Grok 4.5GLM-5.1Kimi K2.7 CodeQwen3.7 MaxKimi K3GLM-5.2
opencode Go Plus$40$24075,0005,0523,0572,3231,195773636596283269260203150
Chutes Plus$10$5015,6251,0526364842491611321245956544231
Chutes Pro$20$10031,2502,1051,2739684983222652481181121088462
Ollama Pro$20$6018,7501,2637645802981931591497067655037
Ollama Max$100$30093,7506,3153,8212,9041,494966795745354337326254187
GitHub Max$100$20062,5004,2102,5471,936996644530497236224217169125
Ollama Team$500$1,000312,50021,05212,7389,6804,9823,2222,6502,4851,1821,1241,086847626
Morph Scale$200$400125,0008,4215,0953,8721,9931,2891,060994472449434338250
GitHub Pro+$39$7021,8751,4738916773482251851738278765943
opencode Go$10$154,68731519114574483937171616129
GitHub Pro$10$154,68731519114574483937171616129
Qoder Pro$20$278,3125603382571328570663129282216
Qoder Ultra$200$26683,1255,6003,3882,5751,325857705661314299289225166
Qoder Pro+$60$8024,9371,6801,0167723972572111989489866750
Venice Max$200$22570,3124,7362,8662,1781,121725596559266252244190140
Venice Pro Plus$68$7523,4371,5789557263732411981868884816346
Augment Code Standard$20$206,250421254193996453492322211612
Augment Code Business$100$10031,2502,1051,2739684983222652481181121088462
Venice Pro$18$131221129432211100

Sessions per month. Heavier models burn a budget far faster; real mileage varies with session length and cache behaviour.

Subscription plans

Grouped by how verifiable the quota actually is

Most vendors do not publish a convertible quota. Effective $/token is computed only where a vendor publishes enough to do it honestly; everything else is left blank rather than guessed.

ProviderPlan$/monthQuota disclosureSee what's leftQuota converts to units?Checked
XiaomiMiMo Token Plan (Lite / Standard / Pro / Max)$6.00 / $16.00 / $50.00 / $100.00Convertible token quotaconsole onlyconverts2026-10-02
opencodeopencode Go / Go Plus$10.00 / $40.00Dollar usage budgetconsole onlyconverts2026-10-02
GitHubCopilot (Pro / Pro+ / Max)$0.00 / $10.00 / $39.00 / $100.00Dollar usage budgetconsole onlyconverts2026-10-02
CursorPro / Pro+ / Ultra$20.00 / $60.00 / $200.00Dollar usage budgetconsole onlyunstated2026-10-02
Alibaba / QwenCoding Plan (Pro)$50.00Requests / promptsconsole onlypartial2026-10-02
Alibaba / QwenToken Plan (Individual + Team)$6.00 / $10.00 / $18.00 / $68.00No numeric quotaconsole onlyunstated2026-10-02
Z.ai / ZhipuGLM Coding (Lite / Pro / Max)$18.00 / $80.00 / $168.00Requests / promptsconsole onlyconverts2026-10-02
MiniMaxToken Plan (Plus / Max / Ultra)$22.00 / $55.00 / $132.00No numeric quotaAPIunstated2026-10-02
AnthropicClaude Pro / Max 5x / Max 20x$20.00 / $100.00 / $200.00No numeric quotaconsole onlyunstated2026-10-02
OpenAIChatGPT Plus / Pro 5x / Pro 20x$0.00 / $8.00 / $20.00 / $100.00 / $200.00 / $500.00No numeric quotaconsole onlypartial2026-10-02
GoogleAI Plus / Pro / Ultra$4.99 / $19.99 / $99.99 / $199.99No numeric quotaconsole onlyunstated2026-10-02
Moonshot / KimiAdagio (gratis) / Plus / Pro / Max / Ultra$19.00 / $39.00 / $99.00 / $199.00No numeric quotaAPIpartial2026-10-02
MistralFree / Pro / Team (Mistral Vibe)$0.00 / $14.99 / $24.99No numeric quotanowhereunstated2026-10-02
Windsurf (Cognition/Devin)Pro / Max / Teams$0.00 / $20.00 / $200.00No numeric quotaconsole onlyunstated2026-10-02
OllamaOllama Cloud (Free / Pro / Max)$0.00 / $20.00 / $100.00 / $500.00Dollar usage budgetconsole onlyconverts2026-10-02
VeniceVenice (Free / Pro / Pro Plus / Max)$0.00 / $18.00 / $68.00 / $200.00Dollar usage budgetconsole onlyconverts2026-10-02
AtlasCloudCoding Plan (Starter / Lite / Plus / Max / Ultra / Enterprise)$10.00 / $20.00 / $50.00 / $100.00 / $200.00 / $500.00Requests / promptsconsole onlypartial2026-10-02
Io NetIO Intelligence (Standard / Professional / Developer)not publishedRequests / promptsconsole onlyconverts2026-10-02
ChutesPlus / Pro$10.00 / $20.00Dollar usage budgetconsole onlyconverts2026-10-02
MorphUsage-based (Free / Pay-per-token / Scale flat-rate)$0.00 / $200.00Dollar usage budgetconsole onlyconverts2026-10-02
QoderQoder (Free / Pro / Pro+ / Ultra)$0.00 / $20.00 / $60.00 / $200.00Dollar usage budgetconsole onlypartial2026-10-02
Agnes AIToken Plan (Starter / Plus / Pro)$4.00 / $10.00 / $50.00Requests / promptsconsole onlypartial2026-10-02
SyntheticSynthetic (Standard)$30.00Requests / promptsconsole onlyconverts2026-10-02
Augment CodeAugment (Business / Enterprise) — Auggie CLI + Cosmos$20.00 / $100.00Dollar usage budgetconsole onlyconverts2026-10-02
SolheimCoder / Coder+ / Coder++$15.00 / $24.00 / $48.00No numeric quotaconsole onlyunstated2026-10-02
SolheimInstancias (VPL)$15.00 / $30.00 / $45.00Dollar usage budgetconsole onlyconverts2026-10-02
One plan meters something else entirely. Ollama Cloud bills on “actual utilization of Ollama’s cloud infrastructure — primarily GPU time”, the only compute-metered subscription here. That inverts an assumption the rest of this page relies on: everywhere else a token costs the same whoever produced it, so model speed only affects your patience. On GPU time, speed is the quota — a model running at half the tokens per second burns roughly twice the allowance for identical output. Check the leaderboard's tok/s column before picking a model to run there.

These two columns are inverted. The vendors that let you read your quota programmatically do not tell you what you bought, and the ones that publish proper unit maths only give you a web page. MiniMax and Kimi expose an endpoint for a quota with no published units at all; Xiaomi, GitHub and opencode publish exact conversions but no API. Nobody does both. DeepSeek is the only one coherent on both axes, and only because it sells no subscription — there, remaining quota is simply money.

Only three providers document a remaining-quota endpoint, and all three are Chinese. Western billing APIs return consumed usage and never an allowance, which is why every community usage tool hardcodes the limit and does the subtraction itself. Endpoints for MiniMax, Z.ai, Kimi and DeepSeek were called directly against live keys for this table; provider coverage cross-checked against quota-sentinel.

What moves the price inside a plan

Bigger than the gap between plans

Choosing a subscription is the decision buyers agonise over, and it is usually the smaller one. Once inside a plan, what you pay per unit of work varies by model, by tier, by how you use the context and, at two vendors, by the time of day. At Qoder those axes compound into a 27-fold swing in what a credit is worth, far more than separates the plans themselves.

VendorVaries byRangeDetail
Qodermodel0.1x - 0.6xA published credit multiplier per model. They do not track open-market price: one unit of credit buys $5.60 of Qwen3.7-Plus but only $0.68 of GLM-5.2.
Qoderclockdown to 0.01xSelected models are discounted 14:00-00:00 UTC, weekends and holidays included. Qwen3.7-Max falls 0.5x to 0.1x; a promoted preview model reaches 0.01x. Credit price only -- model quality is explicitly unaffected.
Qodertierfree - 1.6xLite costs nothing, Efficient 0.3x, Auto 1.0x, Performance 1.1x, Ultimate 1.6x (currently halved to 0.8x). Sits on top of the model choice, so the two compound.
Z.ai / Zhipuclock2x - 3xThe mirror image: quota burns 3x during peak hours and 2x off-peak, rather than being discounted. A promotion holds off-peak at 1x until 30 September 2026, after which the usable plan roughly halves at an unchanged price.
Cursormechanics1.25x - 2xCache writes cost 1.25x input, and long context doubles above 200K tokens. Charges that depend on how you use the model rather than which one you pick.
Xiaomicache50x - 120xThe largest spread found anywhere, in the subscriber's favour: a cache hit costs 2 credits per token against 100 on mimo-v2.5, and 2.5 against 300 on the Pro model. Consumption also drops to 0.8x between 00:00 and 08:00 Beijing time.
opencodemodel2xKimi K3 draws twice the usage of every other model in the Go plan.
Anthropic / Google / OpenAItier2x - 20xSold as more usage rather than as a multiplier on consumption, and not comparable between vendors: Google's 4x is measured against its free tier while its 5x and 20x are against Pro, and Anthropic's multipliers are per session, not per month.
DeepSeekclockup to 12xFrom 2026-08-16 16:00 UTC DeepSeek moves to peak/off-peak API billing with item-specific multipliers, a sharp escalation of the flat peak surcharge it had run since mid-2026. Against the current flat price, V4 Pro cache-hit input rises 6x off-peak and 12x at peak ($0.0036 -> $0.022 / $0.044 per M), cache-miss input 1.5x / 3x ($0.435 -> $0.66 / $1.32) and output 2.25x / 4.5x ($0.87 -> $1.98 / $3.96); V4 Flash scales the same way ($0.14 miss -> $0.22 / $0.44, $0.28 output -> $0.66 / $1.32). Peak is 01:00-04:00 and 06:00-10:00 UTC (Beijing working hours 09:00-12:00 and 14:00-18:00), off-peak is everything else at half the peak rate. The cache-hit multiplier is the one that bites: a coding agent replays a large cached prefix every turn at ~92-95% cache-hit, so its effective bill tracks the 12x line, not the 3x cache-miss one. From 2026-08-23 (per DeepSeek's email) the peak/off-peak split applies on WEEKDAYS only; on weekends (Sat + Sun Beijing time) every call bills at the off-peak rate all day. Confirmed on the official pricing page 2026-08-22; the flat rate a comparison table quotes is only the pre-Aug-16 number.

Two deserve naming. Qoder and Z.ai run the same mechanism in opposite directions, one discounting the quiet hours and the other surcharging the busy ones, and only a discount looks like a discount. And Xiaomi's cache multiplier is the one that runs in the customer's favour, at up to 120x, which is why cache behaviour decides more of a coding agent's bill than the choice of model does.

MiniMax publishes no token, credit or prompt figure for any tier, and its own API returns a percentage with the total zeroed out. But the two halves can be put together: the console reports tokens consumed, the API reports the percentage left. Divide one by the other and the denominator falls out.

Tier
Ultra ($120/month)
Weekly quota
7.4B tokens
Monthly
32B tokens
Effective
$0.0038 /M

Method. The console reports tokens consumed; the API reports the percentage of the weekly window still left. Dividing one by the other gives the denominator the vendor never publishes. Measured 2026-07-18 against the plan window 13-20 Jul: 7.68B tokens consumed over a rolling 7 days against 6% of the window remaining. Anyone can repeat this on their own account; it costs nothing and spends no quota.

Confidence. Derived from a SINGLE account; not an official figure. Robust to the API's integer-percentage resolution (worth about 0.05B) and to the rolling-versus-fixed window mismatch (7.2-7.7B). One caveat that cannot be resolved from outside: the internal weighting of cached input, fresh input and output has not been reverse-engineered, so this figure is valid at the observed ~92% cache-hit rate and should not be extrapolated to a very different one. Note also that the accounting has already changed once: cached tokens formerly did not count against the quota and now do, which raised the effective cost sharply and was not announced.

For scale: at list price that traffic would cost $0.525 per million blended, and about $0.06 per million once caching is working. The subscription lands near $0.004 — roughly 16x cheaper than pay-per-token even after caching, and two orders of magnitude under list.

Why this had to be measured. Cached tokens did not always count against this quota. At some point they began to, the effective cost rose sharply, and nothing was announced — the subscription simply started buying less. The internal weighting of cached input against fresh input and output remains un-reverse-engineered, so the figure above holds at the observed ~92% cache-hit rate and should not be stretched far beyond it. A number a vendor does not publish is also a number it can change quietly.

By provider

One page per lab: its models, what they cost, the endpoints serving them and at what precision, plus its own plan if it sells one.

Agnes AIrequests
provider
Token Plan (Starter / Plus / Pro) · from $4/mo
AkashML
serves 6
No subscription of its own
Alibaba / Qwenrequests
18 models · serves 26
Coding Plan (Pro) · from $50/mo
Amazon Bedrock
serves 35
No subscription of its own
Anthropicopaque
9 models · serves 12
Claude Pro / Max 5x / Max 20x · from $20/mo
Arcee AI
serves 1
No subscription of its own
Arcee Ai
1 model
No subscription of its own
AtlasCloudrequests
serves 21
Coding Plan (Starter / Lite / Plus / Max / Ultra / Enterprise) · from $10/mo
Augment Codedollars
provider
Augment (Business / Enterprise) — Auggie CLI + Cosmos · from $20/mo
Azure
serves 58
No subscription of its own
Baidu
serves 9
No subscription of its own
BaseTen
serves 19
No subscription of its own
Cerebras
serves 2
No subscription of its own
Chutesdollars
serves 6
Plus / Pro · from $10/mo
Claude Platform on AWS
serves 8
No subscription of its own
Cloudflare
serves 10
No subscription of its own
Cohere
3 models · serves 4
No subscription of its own
CoreWeave
serves 16
No subscription of its own
Crusoe
serves 6
No subscription of its own
Cursordollars
provider
Pro / Pro+ / Ultra · from $20/mo
Darkbloom
serves 8
No subscription of its own
Decart
serves 7
No subscription of its own
DeepInfra
serves 49
No subscription of its own
DeepSeek
9 models · serves 2
No subscription of its own
DekaLLM
serves 8
No subscription of its own
DigitalOcean
serves 13
No subscription of its own
Fireworks
serves 9
No subscription of its own
Friendli
serves 6
No subscription of its own
GMICloud
serves 20
No subscription of its own
GitHubdollars
provider
Copilot (Pro / Pro+ / Max) · from $0/mo
Googleopaque
13 models · serves 62
AI Plus / Pro / Ultra · from $5/mo
Google AI Studioopaque
serves 24
AI Plus / Pro / Ultra · from $5/mo
Groq
serves 3
No subscription of its own
IBM
1 model
No subscription of its own
Inception
1 model · serves 1
No subscription of its own
Inceptron
serves 6
No subscription of its own
InclusionAI
3 models
No subscription of its own
InferenceNet
serves 7
No subscription of its own
Io Netrequests
serves 5
No subscription of its own
Ionstream
serves 3
No subscription of its own
Liquid
1 model · serves 1
No subscription of its own
Makora
serves 3
No subscription of its own
Mancer 2
serves 4
No subscription of its own
Mara
serves 3
No subscription of its own
Meituan
1 model
No subscription of its own
Meta
6 models · serves 3
No subscription of its own
MiniMaxopaque
2 models · serves 3
Token Plan (Plus / Max / Ultra) · from $22/mo
Mistralopaque
7 models · serves 19
Free / Pro / Team (Mistral Vibe) · from $0/mo
Modal
serves 5
No subscription of its own
ModelRun
serves 2
No subscription of its own
Moonshotopaque
4 models · serves 4
Adagio (gratis) / Plus / Pro / Max / Ultra · from $19/mo
Moonshot / Kimiopaque
provider
Adagio (gratis) / Plus / Pro / Max / Ultra · from $19/mo
Morphdollars
serves 6
Usage-based (Free / Pay-per-token / Scale flat-rate) · from $0/mo
Near AI
serves 1
No subscription of its own
Nebius
serves 6
No subscription of its own
NextBit
serves 6
No subscription of its own
Novita
serves 38
No subscription of its own
Nvidia
5 models · serves 1
No subscription of its own
Ollamadollars
provider
Ollama Cloud (Free / Pro / Max) · from $0/mo
OpenAIopaque
25 models · serves 48
ChatGPT Plus / Pro 5x / Pro 20x · from $0/mo
OpenInference
serves 3
No subscription of its own
Parasail
serves 22
No subscription of its own
Phala
serves 18
No subscription of its own
PrimeIntellect
serves 1
No subscription of its own
Qoderdollars
provider
Qoder (Free / Pro / Pro+ / Ultra) · from $0/mo
Reka
serves 3
No subscription of its own
Relace
serves 7
No subscription of its own
Sail Research
serves 7
No subscription of its own
SambaNova
serves 5
No subscription of its own
SiliconFlow
serves 25
No subscription of its own
Solheimopaque
provider
Coder / Coder+ / Coder++ · from $15/mo
StepFun
1 model · serves 1
No subscription of its own
StreamLake
serves 15
No subscription of its own
Syntheticrequests
provider
Synthetic (Standard) · from $30/mo
Tencent
1 model · serves 1
No subscription of its own
Thinking Machines
2 models
No subscription of its own
Together
serves 13
No subscription of its own
Upstage
3 models · serves 5
No subscription of its own
Venicedollars
serves 27
Venice (Free / Pro / Pro Plus / Max) · from $0/mo
Wafer
serves 10
No subscription of its own
Windsurf (Cognition/Devin)opaque
provider
Pro / Max / Teams · from $0/mo
Xiaomitokens
4 models · serves 4
MiMo Token Plan (Lite / Standard / Pro / Max) · from $6/mo
Z.ai / Zhipurequests
6 models · serves 6
GLM Coding (Lite / Pro / Max) · from $18/mo
opencodedollars
provider
opencode Go / Go Plus · from $10/mo
xAI
5 models · serves 22
No subscription of its own

Gateways, resellers & inference hosts

Not the lab, and not the ranking

The places that sell a model without being the lab that made it, kept out of the leaderboard and the first-party price columns on purpose — because the distinction they blur is the useful part. An inference host serves open-weight models on its own GPUs, so its price is a real price for that model. A router forwards to providers with little markup. A reseller proxies someone else’s closed API (GPT, Claude) and reprices it — its number is the vendor’s minus a cut, and its “X% cheaper” is a claim against a list price, only meaningful if you can verify the effective rate before paying. The reseller/router class is detailed below; the full roster of every inference host we track (StreamLake, Together, Groq, SambaNova & 39 more) is at the end of the section.

ServiceOriginTypeServesBilling unitConverts to tokens?Verifiable pre-pay?Note
CometAPI🇭🇰 Hong Kong SARresellerGPT-6.1 Sol, GPT-6 Sol y Luna, Claude Sonnet 5.5 / Fable 5.5 / Haiku 5.5 / Opus 5.5, Gemini 4 Argon, Grok 5, DeepSeek V4 Pro, Kimi K3, MiniMax M3.1 Flash, MiMo V2.6 (321 modelos)tokens (official models) + per-image/clip for mediapartialpartialThe most transparent of the resellers here: the 0.8x formula is stated outright, so the effective token price is checkable against the vendor's own list. Still a proxy of closed APIs, so the price is theirs minus 20%, not a first-party price, and no monthly plan. 2026-09-13: la formula sigue intacta. Catalogo ampliado con GPT-6 Astra, Gemini 4, Claude Fable 5.1, GLM 5.3 Flash y Qwen3.8. Contradiccion suya sin resolver: la portada dice '500+ AI Models' y la pagina de modelos dice '295+'. El '20-40% mas barato' de portada es marketing sin formula detras; lo verificable es el 0,8x. 2026-09-18: la pagina de modelos dice ya '298 models available'. DeepSeek V4 Flash a $0,176/$0,528 contra un 'precio oficial' suyo de $0,22/$0,66, que no es el de DeepSeek ($0,14/$0,28): el 0,8x se aplica sobre su propia cifra, no sobre la del laboratorio. Y deepseek-v4-flash-vision-exp sale MAS CARO que su oficial ($0,352/$1,056 contra $0,22/$0,66), en contra del 'at least 20% below'. Kimi K3 a $2,4/$12. 2026-09-24: 314 modelos en su pagina de modelos (la portada sigue diciendo 500+). Vende ya cuatro de los cinco lanzamientos del 21-22 con precio de ENTRADA y el 0,8x clavado: GPT-6 Sol $1,6, GPT-6 Luna $0,08, Claude Opus 5.5 $3,2 y MiMo V2.6 Pro $0,348; Grok 4.7 esta en catalogo sin precio. En Opus 5.5 su 'oficial' si coincide con el de Anthropic. Sus ref_prices NO se han podido reverificar hoy: la tabla rota y no incluye ninguno, la ficha de DeepSeek V4 Pro da 404 y su API pide clave. 2026-10-02: 321 modelos (la portada sigue diciendo 500+). Vende los dos lanzamientos de la semana al 0,8x clavado: GPT-6.1 Sol y Claude Sonnet 5.5 a $1,60/$8 contra el oficial $2/$10. DeepSeek V4 Flash baja a $0,12/$0,48 porque su 'oficial' baja a $0,15/$0,60. Dato que faltaba: 1.000 creditos = $1, y hay credito de prueba gratis sin tarjeta. Sus ref_prices de DeepSeek V4 Pro y GPT-5.6 Sol siguen SIN reverificar por segunda semana: esas fichas ya no existen y redirigen al listado del proveedor. Y publica dos erratas suyas: Fable 5.5 y Haiku 5.5 a $60 por millon contra un oficial de $75, y Kling 4.0 a $60 por millon cuando se cobra por segundo.
AIMLAPI🇪🇪 Estoniareseller1000+ models. AHORA SI incluye Anthropic: Claude Sonnet 5, Opus 5, Fable 5.1 y 4.5 Sonnet, con precio. Tambien GLM-5.2 y 5.3, DeepSeek V4 Pro, Kimi K3 y K2.5/2.6/2.7 Code, Qwen 3.8 y Qwen3.5 Flash. Desde el 2026-10-02 tambien Claude Sonnet 5.5 y GPT-6.1 Sol.tokens (per 1M)fullpartialCORREGIDO 2026-09-13: la ficha decia 'Anthropic models are absent, so it is not a full first-party substitute' y llevaba siete semanas siendo falso. Hoy vende Claude con precio publicado: Sonnet 5 a $2,6 de entrada ($0,26 cacheado) y $13 de salida por millon. Sin descuento declarado sobre el precio oficial. 2026-09-18: confirmados en su pagina de precios GLM-5.2 ($1,82/$5,72) y DeepSeek V4 Pro ($0,5655/$1,131). GPT-5.6 Sol a secas ya no aparece, solo Sol Pro ($5,50/$27,51). 2026-09-24: vende los cinco lanzamientos del 21-22, con precio: GPT-6 Sol $2,6/$13, GPT-6 Luna $0,13/$0,65, Claude Opus 5.5 $5,2/$26, Grok 4.7 $2,6/$7,8 y MiMo-V2.6-Pro $0,598/$1,197. OJO: su Opus 5.5 esta POR ENCIMA del oficial ($4/$20), asi que aqui no hay descuento, hay recargo. Minimo de pago por uso: $20. Contradiccion suya: la pagina de precios dice 880 modelos y la portada 1000+. 2026-10-02: vende los dos lanzamientos de la semana, los dos POR ENCIMA del oficial: Sonnet 5.5 a $2,60/$13 y GPT-6.1 Sol a $2,60/$13, contra $2/$10 de lista. Un 30% de recargo, no un descuento. Y su DeepSeek V4 Pro MAS QUE DOBLA, de $0,5655/$1,131 a $1,17/$3,51. La contradiccion del recuento cambia de cifras: su pagina de precios dice 453 modelos y su API publica 976.
APIMart🇭🇰 Hong Kong SARreseller100+ modelos, y desde 2026-10 VUELVEN los LLM de texto con precio por token: Claude (Sonnet 5.5, Opus 5.5, Fable 5.1...), GPT (6.1 Sol, 6 Sol/Luna), DeepSeek, GLM, Kimi, Qwen, MiniMax, Grok, Doubao y Gemini, mas todo el catalogo de imagen y videocreditsfullpartialOpenAI-compatible (base api.apimart.ai/v1), but billed in credits with no published credit-to-token conversion, so the advertised saving cannot be verified before paying. This is the entry that prompted the section. 2026-09-13: `conversion` sube de 'none' a 'partial' — ahora ensenan el credito y el dolar al lado del precio oficial ('2Credits/M~$0.2/M'), de donde se deduce 1 credito ~ $0,10, aunque NO publican la tasa como tal. En cambio el catalogo de TEXTO no se ha movido: siguen GPT-5, GPT-4o, Claude Sonnet 4.5 y DeepSeek 'R1 y V3', sin DeepSeek V4, GLM-5.x, Kimi K3 ni Qwen3.8. Crecen en imagen y video. Se esta quedando atras para codigo. CAMBIO GRAVE 2026-09-24: se ha quedado SIN modelos de texto. Sus 48 fichas de catalogo y las 214 filas de precios son todas de imagen, video o audio; las unicas filas por millon de tokens son la entrada y salida de texto de modelos multimodales de imagen. La nota del 2026-09-13 decia que seguian GPT-5, Claude Sonnet 4.5 y DeepSeek: hoy es falso. Deja de ser una pasarela para codigo. Se mantiene el credito a ~$0,10 deducible y el 20% como norma. VUELTA ATRAS 2026-10-02: la semana pasada se habia quedado sin un solo modelo de texto y hoy los vuelve a vender con precio por token, en paginas por familia (apimart.ai/model/llm/claude, /gpt, /glm...). Sonnet 5.5 a $1,60/$8 y GPT-6.1 Sol a $1,60/$8, los dos el 20% por debajo del oficial, y su 'oficial' si coincide con el de Anthropic y OpenAI, asi que el descuento ya es verificable y la tasa deducible: 1 credito = $0,10. Nuevo: membresia Gold, Platinum y Diamond con un 0%, 5% y 10% adicional, que choca con su propia frase de 'pago por uso sin tramos'; y precio por franja horaria en DeepSeek V4.1 Flash, con hasta el 60% de ahorro en valle. NO se pudieron leer los precios de DeepSeek V4 Pro, V4 Flash ni GLM-5.2: su tabla solo pinta el modelo por defecto y hay que pulsar.
EURouter🇳🇱 Netherlandsrouter142 modelos, y desde el 2026-10-02 SI con cerrados y con precio publico: Claude por AWS Bedrock (Sonnet 5, Opus 5 y 5.5, Haiku 4.5) y GPT por Microsoft Foundry (5.6 Sol/Terra/Luna, 5.4, 4.1), mas pesos abiertos, Mistral, GreenPT y Apertustokens + published markupfullyesA Netherlands-based router selling on EU data residency and GDPR rather than price. Notable as the transparent opposite of the credit resellers: it publishes its exact markup per tier and bills real tokens, so the effective price is checkable up front. Base www.eurouter.ai/api/v1. 2026-09-13: cuotas por tramo, que no estaban: Free 10.000 peticiones al mes y 60 por minuto; Plus 1 millon y 100 por minuto; Pro 10 millones y 500 por minuto. Catalogo al dia, anuncian GLM-5.2, Kimi K3 y Opus 5 en vivo. AVISO 2026-09-24: dos cosas de esta ficha dejan de sostenerse. Su banner dice que GLM-5.2, Kimi K3 y Opus 5 estan vivos, y que enruta a GPT-5, Claude y Gemini, pero su propia tabla no los ensena: no digo que no los sirvan, digo que no es verificable. Y el markup del 3-15% sobre el precio oficial tampoco cuadra: publican lista propia por modelo, y su DeepSeek V4 Pro a $1,75/$3,50 es CUATRO VECES el oficial. Otros de su tabla: DeepSeek V4 Flash 0731 $0,13/$0,28, GLM 5.2 $1,10/$4,20, GLM 5.3 $1,47/$4,50, Kimi K3 $3,00/$15,00, Kimi K2.7 Code $0,75/$3,00, MiniMax M3 $0,40/$2,00. Ninguno de los cinco lanzamientos del 21-22. CORREGIDO 2026-10-02: dos avisos de esta ficha ya no valen. Primero, SI vende cerrados y con precio publico; su API (eurouter.ai/api/models, 142 entradas) los publica todos, asi que lo de 'no es verificable' se cae. Segundo, el markup del 3-15% no describe lo que hace: en los cerrados aplica el precio del proveedor por 1,10 (Sonnet 5 a $2,20/$11) y en GPT-5.6 Sol cobra $5,50/$33, que es un 37% y un 65% por encima del oficial. Entra Opus 5.5; Sonnet 5.5 y GPT-6.1 Sol todavia no. Kimi K2.7 Code sube de $3,00 a $3,50 de salida. Su web dice 142 modelos en un sitio y 120+ en otro.
ShareAI—inference-host199 modelos. Ya NO es solo pesos abiertos: vende cerrados de OpenAI y xAI (gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, grok-4.6, grok-4.3), muchos sin precio por token publicadotokensfullpartialNot a closed-API reseller but a decentralized inference marketplace: idle GPUs serve open weights, so its prices are real prices for those models. No GPT or Claude, so limited for coding against the leaders; since 2026-09-13 it does carry one closed model, Gemini 3.5 Flash ('Gemini 3.5 Flash is now available on ShareAI'). Operator HeyShare SRL; country not confirmed. CAMBIO GRAVE 2026-09-24: cruza la linea que describia esta ficha. Vende ya modelos cerrados y PUBLICA EN EUROS, no en dolares: GPT-6 Astra a 9,10 EUR de entrada y 45,50 de salida por millon hasta 272K, y el doble por encima. Matiz importante: muchos cerrados van por 'Token exchange' y su propia ficha dice que no hay precio por token publicado, GPT-5.6 Sol entre ellos. Gemini 3.5 Flash, que se anoto el 2026-09-18, hoy no aparece en su catalogo: no lo doy por retirado, pero no se confirma. Claude no se ha visto. 2026-10-02: sin cambios de precio. Nuevo con precio: gpt-5.4 a 2,21 EUR de entrada y 13,30 de salida por millon. GPT-5.6 Sol y Terra siguen en 'Token exchange' sin precio por token, y Claude sigue sin verse. Gemini 3.5 Flash sigue sin aparecer: tercera semana sin confirmar. Su propio catalogo dio 'We couldn't load the models' y solo pinta las primeras filas, asi que no es legible entero desde fuera.
APIMaster—resellerOpenAI, Claude, DeepSeek via "fingerprint-verified channels"; works with Cursor / Claude Codetokens (shared balance across channels)partialfalseOpenAI-compatible, but the headline discounts are far beyond a normal reseller margin and no markup is published, only per-"channel" prices that vary. Discounts that large on a closed API usually mean grey-market or quota-farmed capacity, not a straight resale. Company location undisclosed. Treat the savings claim with caution. 2026-09-13: Catalogo muy ampliado (GPT-6 Astra, Claude Sonnet 5, Kimi K3, GLM 5.3 Flash, y DeepSeek V4.1 Flash, ojo, no V4). Confirmado que NO hay creditos gratis: minimo de recarga $1. La cautela sigue justificada — 85-90% de descuento sobre API cerrada, sin markup publicado y con USDT entre los metodos de pago. 2026-09-18: la portada dice 'All discounts are on fingerprint-verified channels — same model, lower cost', y la palabra 'marketplace' no aparece. El minimo de recarga de $1 no se ve en la portada. CAMBIO 2026-09-24: la frase de esta ficha de que NO hay creditos gratis ya no vale. Su portada ofrece $20 de credito GPT en una prueba de 5 dias a precio oficial (GPT 6 Astra, GPT 5.6 y GPT 5.5), un 15% en la primera recarga de $10 y $20 por referido. Catalogo muy ampliado: Opus 5 y 5.5, GPT 6 Sol y Luna, Grok 4.7, MiMo v2.6 Pro, Kimi K2.8 Preview y K3, GLM 5.3 / Flash / FlashX, Gemini 3.8 Flash, Qwen 3.8 Max y Flash, Sonnet 5. Y la cautela va A MAS: su propia lista de canales verificados por huella sigue siendo la vieja (GPT-5.4, GPT-5.5, Sonnet 4.6, Opus 4.7/4.8, Fable 5, DeepSeek V4 Flash y Pro, MiniMax M3), asi que vende Opus 5.5 y los GPT-6 SIN verificar. La cita que usaba esta ficha ya no esta en su web; hoy dice que todo proveedor listado pasa su prueba de huella. 2026-10-02: ofrece ya GPT 6.1 Sol y Sonnet 5.5 en el selector. Su clasificador de huella se ha puesto al dia en parte: ya etiqueta los GPT-6 y el 6.1 Sol, y han caido deepseek-v4-pro y minimax-m3. Siguen SIN verificar Sonnet 5.5, Opus 5.5, Grok 4.7 y Kimi K3, y su propia FAQ de Claude no lista ni Sonnet 5.5 ni Opus 5.5 aunque la portada los venda. El 'hasta 90%' se queda corto: anuncia 96% en GPT-6 Astra ($0,45 por millon contra $10 de lista), lo que refuerza la cautela en vez de rebajarla. Publica precio por canal contra el oficial, asi que la conversion pasa de 'none' a 'partial'.
TokenMix🇭🇰 Hong Kong SARreseller198 modelos, 142 de chat: GPT-6 Sol y Luna, GPT-6 Astra, Claude Opus 5.5 / Sonnet 5 / Fable 5, DeepSeek V4, Kimi K3, GLM-5.2 / 5.3, Grok 4.6 y 4.5, Gemini 3.8 Flash, Qwen 3.8 Max / Flash / Omnitokens (per-model)partialpartialToken Limited (Hong Kong). Pay-per-token with per-model pricing and an OpenAI-compatible base (api.tokenmix.ai/v1). No markup formula published, but it bills tokens rather than opaque credits. $1 minimum top-up. 2026-09-13: Claude Sonnet 5 ya tiene precio publicado, $1,96/$9,80. Y ojo al desambiguar GLM: los $2,24/$7,82 de esta ficha son la variante Fast Preview; el GLM-5.2 normal y el GLM-5.3 estan a $1,117647/$3,911765. 2026-09-24: sin cambios de precio. Nuevos con precio: Kimi K3 $2,647/$13,235 y GPT-6 Astra $9,50/$47,50. NINGUNO de los cinco lanzamientos del 21-22. Sus ref_prices se reconfirman una a una. 2026-10-02: ya vende tres de los lanzamientos recientes, GPT-6 Sol ($1,90/$9,50), GPT-6 Luna ($0,095/$0,475) y Claude Opus 5.5 ($3,92/$19,60); Sonnet 5.5 y GPT-6.1 Sol dan 404 en sus URLs de modelo. El descuento ya esta declarado modelo a modelo. Y da creditos gratis al registrarse, que esta ficha no recogia. Minimo de $1, pero $20 si pagas en cripto.
CrofAI—inference-hostKimi, GLM, DeepSeek, Qwen, Gemma y MiMo-V2.5-Pro. Ya NO sirve MiniMax. Y desde 2026 tiene dos modelos PROPIOS de marca, greg-2-ultra y greg-2-super, asi que deja de ser solo pesos abiertos.tokens (pay-per-token)fullyesCheap OpenAI-compatible inference (crof.ai/v1) for open weights only, no markup, pure pay-per-token -- its own page now states \u201cno subscriptions, no minimums, only the tokens you actually burn\u201d (the tiered Hobby/Max plans some third-party listings still show are gone). One of the resellers this repo excluded from first-party price resolution earlier (as \u201ccrof\u201d); it belongs here, not in the leaderboard prices. Country not confirmed. checked 2026-07-22. 2026-09-13: DeepSeek V4 Flash baja de $0,12/$0,21 en Q4_0 a $0,07/$0,10 en Q8_0 — mas barato Y mejor cuantizado, que es lo raro. Declaran retencion cero explicita y no entrenar con las conversaciones. Los ToS dicen regirse por leyes de Estados Unidos, aunque la sede fisica sigue sin confirmarse. Contradiccion suya: la portada dice 'no subscriptions' y los ToS hablan de planes gratuitos y de pago. CERRANDO 2026-09-18: crof.ai redirige a nahcrof.com, que dice 'Nahcrof is in the process of shutting down. Currently we are working with Polar to process refunds to existing customers.' Hasta su propio /v1/models devuelve esa pagina. No compres aqui: los precios de abajo son historicos. 2026-09-24: sigue igual, cerrando y con los reembolsos en curso por Polar. Sin fecha de cierre publicada. En su cuenta de X: ya no se pueden comprar creditos y prometen una pagina para pedir el reembolso. 2026-10-02: sigue cerrando, con el mismo aviso de reembolsos por Polar y sin fecha. Un matiz: su /v1/models ya no devuelve la pagina de cierre, devuelve un 404 en JSON.
nano-gpt—routerGPT, Claude, DeepSeek, Qwen + open models (1000+), OpenAI-compatibletokens (list price)fullyesA no-markup, privacy-first router: it passes provider list prices straight through (GPT-5.5 at $5/$30, same as OpenAI) with no deposit fee, and takes crypto from $0.10 with no KYC. The honest end of the spectrum -- the price is the vendor\u2019s, verifiable up front -- sold on privacy rather than a discount. Country not confirmed. 2026-09-13: hay una suscripcion opcional de $12/mes que esta ficha no recogia. Y el 'sin markup' tiene dos excepciones declaradas: 5% por fijar proveedor a mano y 5% en BYOK. Precios que se mueven: DeepSeek V4 Pro de $0,435/$0,87 a $1,10/$2,20, y GLM-5.2 de $1,4/$4,4 a $0,42/$1,32. SIN CONFIRMAR el de GPT-5.6 Sol: su propia API lo da a $2/$10, identico byte a byte al de Claude Sonnet 5 incluida la escritura de cache, lo que apunta a un dato mal copiado en su catalogo. 2026-09-24: vende los cinco lanzamientos del 21-22, y Claude Opus 5.5 al precio OFICIAL ($4/$20); GPT-6 Sol $2/$10, GPT-6 Luna $0,1/$0,5, Grok 4.7 $1,6/$4,8, MiMo-V2.6-Pro $0,435/$0,87. Sus 617 modelos confirman los ref_prices sin un cambio. Sigue el dato sospechoso de GPT-5.6 Sol, byte a byte igual que Claude Sonnet 5. 2026-10-02: 634 modelos de texto, y vende los dos lanzamientos de la semana AL PRECIO OFICIAL: Claude Sonnet 5.5 a $2/$10 y GPT-6.1 Sol a $2/$10 con cache a $0,10. Sigue sin resolverse el dato sospechoso de GPT-5.6 Sol, que su API da a $2/$10, igual que Sonnet 5: tercera semana.
Fal.ai · Replicate · Wavespeed AI · PiAPI · EachLabs—resellergenerative image / video / audio (Midjourney, Flux, Kling, etc.)per-second / per-run / creditsn/an/aNamed by APIMart as its competitors. All are media-generation gateways, not coding-LLM providers, so they are recorded here only to say they were checked and are out of scope — not to rank.
Featherless AI🇸🇬 Singaporeinference-hostopen-weight models only (40k+ pulled from Hugging Face) — DeepSeek V4 Pro/Flash, GLM 5.3, Qwen3, Kimi K3, Llamasuscripcion mensual por concurrencia, Y desde 2026 tambien por token (plan Developer)partialpartialCORREGIDO 2026-09-13: la ficha decia que no hay precio por token que comparar, y ya lo hay — el plan Developer son $50/mes en creditos, 'billed per token', hasta 256K de contexto. Quedan Chat a $25/mes (4 unidades de concurrencia, 32K, tokens ilimitados), Developer, y Business a medida con GPU dedicadas. El plan Agent de $100/$200 ya no aparece. Anuncian 40.000+ modelos y destacan GLM 5.3 (753B) y Kimi K3 (2780B). 2026-09-24: sus precios por token si son publicos, en api.featherless.ai/v1/models, y de ahi salen ya los ref_prices: GLM-5.3 $1,40/$4,40, GLM-5.3-Flash $0,15/$0,50, Kimi K3 $3/$15, DeepSeek V4.1 Flash $0,30/$1,20, Qwen3.8-27B $0,40/$3. Developer trae 100 unidades de concurrencia y un entorno de agente. Contradiccion suya: su pagina de modelos dice 50.433 y la de precios 40.000+. 2026-10-02: precios y planes identicos, ref_prices reconfirmados uno a uno contra su API. Los creditos no caducan. Su contradiccion de recuento se mueve: la pagina de modelos dice 50.963 y la de precios sigue en 40.000+.
Nscale🇬🇧 United Kingdominference-hostopen-weight serverless — gpt-oss, Qwen2.5-Coder / Qwen3, Llama 4, Devstral, DeepSeek R1 distills (no full DeepSeek V4, no GLM, no Kimi)tokens (per 1M, pay-as-you-go)fullfalseCAMBIO GRAVE 2026-09-13: /pricing y /product/serverless devuelven 404 y el sitemap ya no tiene ninguna pagina de precios ni de serverless. El catalogo y las tarifas estan hoy TRAS LOGIN — el /v1/models publico devuelve 401 — y la web comercial se ha reorientado a venta consultiva ('Talk to an expert', 'Reserve GPUs'). El servicio NO esta muerto: los docs siguen y el alta por consola existe, pero ya no hay tarifa publica que verificar, por eso `verifiable` baja a 'no'. Y la ficha decia 'sin Kimi': es falso desde ABRIL, su propio blog anuncio Kimi K2.5 a $0,45/$2,20 el 2026-04-13, o sea que ya lo era en la revision del 22 de julio. DeepSeek V4 y GLM siguen sin confirmarse. Los '~5 $ de credito de alta' tampoco se confirman hoy. Sedes europeas: Noruega, Portugal (Sines), Islandia y Reino Unido; HQ en Londres, fuera de la UE. Sin ISO 27001, SOC 2, GDPR ni DORA publicados en esas paginas. 2026-10-02, a peor para verificar nada: /pricing sigue dando 404 y su /v1/models sigue pidiendo clave. /product/serverless ya no es 404, pero redirige a una pagina de servicios sin catalogo ni tarifa. Y OJO con la huella europea que decia esta ficha: hoy su portada lista tambien Texas, Virginia Occidental y Carolina del Norte. Servicio nuevo: ajuste fino con banco de pruebas de prompts. Sus ref_prices y su lista de modelos siguen sin poder reverificarse, segunda semana.
OVHcloud AI Endpoints🇫🇷 Franceinference-hostopen-weight, EU — Llama 3.3 70B, Qwen (3.8/3.6/3.5/2.5-VL) y gpt-oss. SIN DeepSeek V4, sin GLM y sin Kimi, y desde 2026 tampoco Mistral, Codestral, Devstral ni Qwen3 Coder.tokens (per 1M, priced in EUR)fullyesFrance-based (OVHcloud, Roubaix), 'a leading European provider of sovereign cloud services since 1999', con ISO 27000, SOC y certificaciones sanitarias en la plataforma, retencion cero y el compromiso de no entrenar con tus datos. AVISO 2026-09-13: la ficha afirmaba 'GDPR / EU data residency, served from Gravelines' y ninguna de esas dos cosas aparece ya en la pagina; pasan a no confirmadas. El motivo para descartarlo en codigo se sostiene, pero ha cambiado: antes era que faltaba el SOTA abierto; ahora ademas han caido Mistral entero, Qwen3 Coder 30B y el distill de DeepSeek, asi que la unica opcion de codigo que queda es gpt-oss-120b. Precios en EUROS, el coste en dolares flota con el cambio. Batch con 50% de descuento. Incoherencia del propio OVH: su FAQ comercial sigue diciendo 'Llama, Qwen, Deepseek y mas de 40 modelos' mientras el catalogo ensena 20 y ningun DeepSeek. 2026-09-18: el catalogo con precio se queda en ocho modelos (Qwen3.8-27B, Qwen3.6-27B, Qwen3.5-397B, Qwen3.5-9B, gpt-oss-120b y 20b, Llama 3.3 70B, Qwen2.5-VL-72B); salen Qwen3 Coder 30B, Qwen3 32B y Mistral Small 3.2. 2026-09-24: catalogo en 20 entradas, los ocho de texto con el MISMO precio que la semana pasada. Nuevos: Qwen3-Embedding-8B a 0,1 EUR y dos guardas en beta gratis (Qwen3Guard-Gen-8B y 0.6B), lo que respalda el 'algunos modelos gratis'. Sigue sin DeepSeek, GLM, Kimi ni Mistral, y GDPR y Gravelines siguen sin aparecer: segunda semana sin confirmar.
Scaleway Generative APIs🇫🇷 Franceinference-hostopen-weight, EU-sovereign — Llama, Mistral, Qwen3, gpt-oss, y DESDE 2026 deepseek-v4-flash-0731. Nuevos: qwen3.5-397b-a17b y gemma-4-26b-a4b-it.tokens (per 1M, priced in EUR)fullyesFrance-based, Paris only: 'all models are hosted in secure data centers in France', y declaran que no leen ni reutilizan tus prompts. CORREGIDO 2026-09-13: la ficha decia 'DeepSeek R1 distills only, no DeepSeek V4' y ya no es cierto — sirve deepseek-v4-flash-0731 a 0,40 EUR de entrada y 0,80 de salida (0,08 cacheado). Precios en EUROS, asi que el coste normalizado a dolares flota con el cambio. Batch con 50% de descuento, literal. SIN CONFIRMAR si siguen Devstral 2, Mistral Large 3 y MiniMax-M2.5: no salen en la tabla de precios de hoy, pero tampoco consta que se hayan retirado. 2026-09-24: precios identicos, y aparecen dos que faltaban en esta ficha: Gemma 4 26B A4B (0,25/0,50 EUR) y Pixtral 12B (0,20). El millon de tokens gratis se confirma literal. Devstral 2, Mistral Large 3 y MiniMax-M2.5 siguen sin salir: segunda semana sin confirmar. Y el -50% de Batch ya NO sale con cifra, solo como 'descuento adicional', asi que deja de ser literal. 2026-10-02: vuelta atras de la nota pasada, el -50% de Batch VUELVE a estar literal: 'All requests performed using Batches API are priced with a -50% discount'. Los 16 modelos con precio, identicos. Faltan en esta ficha dos embeddings con precio: qwen3-embedding-8b y bge-multilingual-gemma2, los dos a 0,10 EUR con salida gratis. Devstral 2, Mistral Large 3 y MiniMax-M2.5 siguen fuera: tercera semana sin confirmar.
Public AI—inference-hostpublic-good / sovereign open-weight models only — Apertus, SEA-LION, EuroLLM, Bielik, Olmo (no mainstream or coding models)tokens (wallet credits, per 1M)fullpartialA non-profit "public utility for AI" (Public AI Inference Utility), serving publicly-funded sovereign models on donated / partner clusters, also routed free via Hugging Face. Currently free with a ~20 req/min limit, but "free" is explicitly provisional and there is no price sheet or SLA. The roster is sovereign / multilingual general-instruct models (Apertus, SEA-LION, EuroLLM) with no DeepSeek / GLM / Qwen-base / Llama / gpt-oss and no dedicated coding model, so it is recorded for completeness, not as a coding option. Legal HQ not stated (Swiss / EU centre of gravity). MATIZ 2026-09-13: esto NO es residencia de datos en la UE. Hay una capa global de balanceo que enruta entre paises, con AWS y Australia entre los clusteres, asi que el 'centro de gravedad suizo/europeo' describe de donde sale el proyecto, no donde se procesa tu peticion. Ya nombran a los donantes del computo: Swiss National Supercomputing Centre, Intel, AWS, Exoscale, AI Singapore, Cudo Compute, NCI Australia y Juelich. Anadido ALIA (Barcelona Supercomputing Center) al catalogo; Apertus va por la 1.5, con comprension de imagen y cuatro veces mas contexto. CAMBIO 2026-09-18: YA NO ES GRATIS. Cobra por token contra un monedero: 'Token usage is billed against wallet credits… New accounts receive starter credits'. Precios en platform.publicai.co/models: Apertus 1.5 70B $0,82/$2,92, Gemma-SEA-LION-v4-27B $0,20/$0,40, Qwen-SEA-LION-v4-32B $0,25/$0,50, Bielik-11B v3 $0,40/$0,40. Limites por plan: Free 100 peticiones/min, Plus 200, Pro 300, Enterprise 10.000. EuroLLM y ALIA no salen en la tabla de precios de hoy. 2026-09-24: la cantidad del credito de inicio ya esta publicada, $2. El catalogo con precio pasa de 4 a 9 entradas, con Apertus 1.5 en 8B y 70B y sus variantes de razonamiento. EuroLLM y ALIA siguen fuera. Plus se consigue contribuyendo por OpenCollective con el mismo correo de la cuenta.
Hugging Face Inference Providers🇺🇸 United Statesrouter200+ modelos de pesos abiertos a traves de 18 proveedores: Baseten, Cerebras, Cohere, DeepInfra, Fal AI, Featherless AI, Fireworks, Groq, HF Inference, Novita, Nscale, OVHcloud AI Endpoints, Public AI, Replicate, Scaleway, Together, WaveSpeedAI y Z.aitokens (pass-through provider rate)fullyesA router, not a host: one HF token routes to leading inference providers, OpenAI-compatible (router.huggingface.co/v1), with no HF markup — you pay the partner's rate. The provider is chosen by policy suffix (:fastest default, :cheapest, :preferred, or an explicit :provider). A small monthly credit ($0.10 free / $2 on PRO at $9/mo) then pay-as-you-go. It serves whatever its partners serve (open weights only), so price and quantization come from the chosen partner, not HF. Closed models (GPT-5.x, Claude) are not routed. 2026-09-24: creditos y el 'sin markup' sin cambios. La lista de socios ha crecido a 18, y cuatro de ellos tienen ficha propia en este mismo fichero: Featherless, Nscale, OVHcloud y Public AI. 2026-10-02: sin cambios. Matiz nuevo: al usuario gratuito, agotado su credito de $0,10, el pago por uso le exige COMPRAR creditos antes.
NVIDIA Build🇺🇸 United Statesinference-host97 modelos de peso abierto en su web; su /v1/models publica 81. Ni un Qwen entre ellos (2026-10-02).peticiones por minuto (produccion via NIM / por token)fullpartialCORREGIDO 2026-09-13: la ficha decia que no siempre trae el frontier mas reciente y para DeepSeek es falso — hoy sirve deepseek-v4-pro-0813 (262K de contexto) y deepseek-v4-flash-0731, ademas de kimi-k3, gemma-4-31b-it y tres Nemotron 3. GLM: su /v1/models lista z-ai/glm-5.3 y z-ai/glm-5.3-flash (2026-09-18). PUNTO FLOJO DE ESTA FICHA: los '1.000-5.000 creditos' y los '40 peticiones por minuto' NO se verifican en build.nvidia.com, que solo dice 'free inference with leading models' sin cifras; una fuente secundaria llega a decir que los limites de credito se han quitado. OJO 2026-09-18: deepseek-v4-pro-0813 tiene ficha en build.nvidia.com pero NO sale en /v1/models (82 modelos), asi que no esta claro que hoy se pueda llamar por la API. 2026-09-24: el desajuste entre su web y su API se ensancha. En /v1/models (82 modelos) el unico DeepSeek de frontera es v4.1-flash: el v4-flash-0731 estaba la semana pasada y hoy no, y el v4-pro-0813 sigue fuera, aunque las dos fichas web existan. Y 'tres Nemotron 3' se queda corto: hoy hay seis (Ultra 550B, Super 120B, Nano 30B, Nano Omni 30B, 3.5 Lightning y Embed 1B). Nuevos tambien: Laguna XS 2.1, Muse Glimmer 30B, DiffusionGemma 26B y su propio Ising Calibration 1.5 31B. 2026-10-02: el punto flojo de esta ficha se arregla a medias. Los 40 por minuto YA son verificables, en su propia FAQ: 'The free tier runs on rate limits - up to 40 requests per minute (RPM) for most models, with no per-token billing'. En cambio los '1.000-5.000 creditos' se caen: la palabra credito no aparece ni una vez en su pagina de modelos. El catalogo baja de 103 a 97 y la API de 82 a 81. Dato llamativo: ni un solo modelo Qwen en la API. Nuevos: Nemotron 3.5 Content Safety y Nemotron Parse 2.0.
RunPod—inference-hostPublic Endpoints de texto, y desde el 2026-10-02 se queda en tres: Deep Cogito v2 Llama 70B, Qwen3 32B AWQ y IBM Granite 4.0 H Small. El endpoint de Moonshot Kimi HA DESAPARECIDO. Cualquier otro peso abierto lo despliegas tu en Serverless o en Pods.tokensfullpartialCORREGIDO 2026-09-13: la ficha lo marcaba como no apto para codigo porque 'el catalogo de texto es fino y Qwen3 32B es el principal'. Ya no: su endpoint de Kimi sirve K2.7 Code y K3, este a $3,00 de entrada y $15,00 de salida por millon, por API gestionada y por token. Qwen3 32B AWQ sigue a $10,00 por millon y Granite 4.0 H Small a $1,00. NO es un proxy de OpenRouter: salio de ahi y hoy no sirve ni un modelo de nuestra tabla por esa via, asi que esta ficha es lo unico que el sitio cuenta de el. Su tabla lista 'Deep Cogito v2 Llama 70B a $0.00001 por millon': es una errata casi segura y no se publica. El serverless lo cotizan hoy POR HORA (B300 $9,98, H200 $5,93, A100 $2,72) aunque la portada siga diciendo 'billed by the millisecond'. 2026-09-18: kimi-k2.6 y kimi-k2.7-code a $0,95/$4,00 por millon; el rango del endpoint de Kimi va de $0,95-$3,00 de entrada a $4,00-$15,00 de salida. 2026-09-24: Kimi identico (K2.6 y K2.7 Code a $0,95/$4,00, K3 a $3/$15) y sigue publicada la errata de Deep Cogito v2 a $0,00001 por millon. Dos matices: 'billed by the millisecond' ya no esta en el texto visible, solo en el JSON incrustado; y las horas de Serverless de esta ficha (B300 $9,98, H200 $5,93, A100 $2,72) no se han podido reconfirmar, porque lo que renderiza hoy /pricing es la tabla de Pods. CAMBIO GRAVE 2026-10-02: el endpoint publico de Kimi ya no existe. Las palabras kimi y moonshot aparecen cero veces en su pagina de precios, asi que la frase de esta ficha de que servia K2.7 Code y K3 se cae, y con ella el motivo para marcarlo apto para codigo: vuelve a `coding: false`. La errata de Deep Cogito a $0,00001 por millon sigue publicada. Sus horas de Serverless siguen sin reconfirmarse, y las de Pods han BAJADO: B300 $7,89, B200 $6,79, H200 $4,59, H100 SXM $3,49, A100 $1,59. Y 'billed by the millisecond' ha desaparecido tambien del JSON incrustado.
CoreWeave Forge (antes Weights & Biases Inference)—inference-host32 modelos de peso abierto por su propia API: DeepSeek V4 Pro/Flash y V4.1 Flash, GLM 5 / 5.2 / 5.3 Flash, Kimi K2.6 y K2.7 Code, MiniMax M3, Qwen3.8 27B, Gemma 4, Nemotron 3, Mellum2, gpt-oss.tokensfullpartialLeft OpenRouter, so it no longer appears anywhere in the provider table — checked 2026-09-13 against the endpoints of all 445 models on the listing, and it serves none of them. It still sells inference directly, which is why it belongs here: a reader can still buy from it. Two things to know before you do. Inference is billed SEPARATELY from the Enterprise and Pro licences, in arrears, and paid access requires prepayment. The $0.50 per million 'Agent Token Rate' is for the ARIA agent only ('ARIA runs on OpenAI models… This rate applies on top of model API pricing'), not for plain inference (as of 2026-09-18). Free, Pro and Academic plans include some serverless credits; when they run out the account has to switch to pay-as-you-go explicitly. You can cap it: a monthly budget stops requests once reached. 2026-09-18: tabla de modelos alojados con precio: DeepSeek V4-Pro $1,15/$2,55, V4-Flash $0,14/$0,28, GLM 5.2 $0,76/$2,42, GLM 5.3 Flash $0,15/$0,50, Kimi K2.7 Code $0,71/$3,50, Kimi K2.6 $0,65/$3,41, MiniMax M3 $0,23/$0,96, gpt-oss-120b $0,03/$0,17. Lo de 'in arrears' y el prepago no sale en esta pagina: sin confirmar. 2026-09-24: su tabla de modelos alojados pasa de 8 a 30. Nuevos con precio, entre otros: DeepSeek V4-Pro-0813 $1,31/$3,96, V4.1-Flash $0,20/$0,65, Qwen3.8 27B $0,40/$3,00, GLM 5 $1,00/$3,20, Gemma 4 31B $0,10/$0,34, Nemotron 3 Ultra $0,50/$2,15 y JetBrains Mellum2 $0,05/$0,10. Los ref_prices de esta ficha no se mueven. Y el recargo del agente ARIA, leido hoy, es mas amplio de lo que decia esta ficha: 'All plans include an Agent Token Rate of $0.50 per million tokens... for all token types'. Lo de la facturacion a mes vencido y el prepago sigue sin aparecer. CAMBIO DE IDENTIDAD 2026-10-02: la pagina ya no es de Weights & Biases. Se titula 'CoreWeave Forge Inference, Sandboxes, and ARIA pricing' y su tabla se llama 'CoreWeave Forge hosted models', aunque la URL de wandb.ai siga sirviendo. El catalogo pasa de 30 a 32. ARIA es ahora GRATIS por tiempo limitado, y despues pasara a cobro por consumo; el recargo de $0,50 por millon sigue descrito igual. Lo de la factura a mes vencido ya se apoya en algo ('Usage is then invoiced monthly'); el prepago sigue sin aparecer. Nuevo: CoreWeave Sandboxes, con CPU a $0,128 por nucleo y hora y memoria a $0,0212 por GiB y hora, y post-entrenamiento marcado por modelo.

The open-weight inference hosts and the router are already priced per model in the endpoint tables above (the “where to buy” view per provider). This table exists for the reseller class, which those tables do not capture because a reseller’s repriced number is not a first-party price. Media-only gateways are listed to record that they were checked and are out of a coding site’s scope. Data in gateways.yaml.

What common models cost through every host

Input / output per million tokens for five reference models, against the vendor’s own list price — the resellers at the top (a fixed set we researched), then every OpenRouter host that serves them, with its real price and the precision it runs at. Two findings. “Gateway” does not mean “cheaper”: AIMLAPI and EURouter charge more than official, only CometAPI/TokenMix undercut it, and the credit/channel resellers can’t be placed on the scale at all. And for the open models the price is only half the story — the other half is quantization: the closed models (GPT-5.6, Claude) are first-party only, so hosts show “—”, while a cheaper DeepSeek here may be fp4, not fp8. Verified 2026-10-03.

ServiceDeepSeek V4 Pro · fp8 nativeDeepSeek V4 Flash · fp8 nativeGLM-5.2 · fp8 nativeGPT-5.6 SolClaude Sonnet 5Claude Sonnet 5.5GPT-6.1 SolBasis
Official (first-party)$0.435 / $0.87$0.14 / $0.28$1.4 / $4.4$4 / $20$2 / $10$2.00 / $10.00$2.00 / $10.00the list price
CometAPI$0.35 / $0.70$0.12 / $0.48$1.12 / $3.528$4 / $24$1.60 / $80.8x of their own 'official' figure, which is not always the lab's; quant undisclosed
AIMLAPI$1.17 / $3.51$0.39 / $1.56$1.82 / $5.72Sol Pro only ($5.50 / $27.51)$2.60 / $13per-model list; no stated discount
APIMartopaqueopaqueopaqueopaqueopaquecredits, no conversion
EURouter+3-15%+3-15%+3-15%$5.50 / $33.00$2.20 / $11.00lista propia por modelo. En los cerrados el patron es el precio del proveedor x1,10 (gpt-4o $2,75 = $2,50x1,1); en los abiertos no, su DeepSeek V4 Pro sigue a 4x el oficial. Encima va el markup por tramo
ShareAIopen, per-tokenopen, per-tokenopen, per-tokenlistado, sin precio por token publicado (Token exchange)not servedopen weights; quant undisclosed
APIMasteropaqueopaqueopaqueopaqueopaque"channels", up to 90% off claimed
TokenMix$0.41 / $0.83$0.13 / $0.26$1.12 / $3.91$4.75 / $28.50$1.96 / $9.80per-model; quant undisclosed (GLM-5.2 Fast Preview is dearer, $2.24/$7.82)
CrofAIshutting downshutting downshutting downnot servednot servedshutting down 2026-09, refunds in progress; prices no longer valid
nano-gpt$1.10 / $2.20$0.14 / $0.28$0.42 / $1.32$5 / $30$2 / $10list price, no markup; routes so quant varies
Featherless AI$1.60 / $3.20$0.14 / $0.28$1.40 / $4.40not servednot servedprecio por millon publicado en su API (22.030 entradas); $25 Chat ilimitado / $50 Developer por token / Business a medida. OJO: su DeepSeek V4 Pro a $1,60/$3,20 es ~3,7x el oficial
Nscalenot servednot served (R1 distills only)not servednot servednot servedopen-weight per-token; Qwen2.5-Coder 32B $0.06/$0.20, gpt-oss-120b $0.10/$0.40, Devstral Small $0.10/$0.30
OVHcloud AI Endpointsnot servednot servednot servednot servednot servedopen-weight per-token EUR; gpt-oss-120b €0.08/€0.40, Qwen3.8 27B €0.40/€2.70
Scaleway Generative APIsnot served€0.40 / €0.80€1.80 / €5.50 · fp8not servednot servedopen-weight per-token EUR, quant disclosed; GLM-5.2 €1.80/€5.50, Qwen3 Coder 30B €0.20/€0.80, Mistral Small 3.2 €0.15/€0.35
Public AInot servednot servednot servednot servednot servedsovereign open models, per-token from wallet credits (Apertus 1.5 70B $0.82/$2.92); no coding model
Hugging Face Inference Providerspartner rate (pass-through, if a partner serves it)partner rate (pass-through, if a partner serves it)partner rate (pass-through, e.g. via Scaleway / Z.ai)not served (open weights only)not served (open weights only)pass-through partner price, no markup; +$0.10/$2 monthly credit
NVIDIA Buildlisted, not in /v1/models (unconfirmed)free tier (deepseek-v4.1-flash); el v4-flash-0731 tiene ficha web pero ya NO sale en /v1/modelsnot served (serves GLM-5.3 / 5.3 Flash)not servednot servedfree dev tier (credits, ~40 RPM); production via NIM / AI Enterprise
RunPodnot served (self-deploy on Serverless/Pods)not served (self-deploy)not served (self-deploy)not served (closed model)not served (closed model)Public Endpoint text — Qwen3 32B $10.00/1M tokens; image Flux Dev $0.02/megapixel; video WAN 2.5 $0.50/5s. Self-deploy = GPU-time (Pods per-hour, Serverless per-second).
CoreWeave Forge (antes Weights & Biases Inference)$1.15 / $2.55$0.14 / $0.28$0.76 / $2.42not servednot servedper-token list, hosted models; quant undisclosed
Alibaba / Qwen$1.12 / $3.37?$0.35 / $1.06?$0.97 / $3.04fp8——OpenRouter · live
Amazon Bedrock———$4.40 / $22.00?$2.00 / $10.00?OpenRouter · live
Anthropic————$2.00 / $10.00?OpenRouter · live
AtlasCloud$1.32 / $3.96fp8$0.44 / $1.32fp4$0.94 / $2.95fp8——OpenRouter · live
Azure———$4.00 / $20.00?$2.00 / $10.00?OpenRouter · live
Baidu$0.13 / $0.40fp8$0.44 / $1.32fp8$0.14 / $0.44fp8——OpenRouter · live
BaseTen—$0.13 / $0.26fp8$1.40 / $4.40fp8——OpenRouter · live
Claude Platform on AWS————$2.00 / $10.00?OpenRouter · live
Cloudflare$1.32 / $3.96?$0.44 / $1.32?$1.18 / $4.40?——OpenRouter · live
Cohere—$0.14 / $0.28?———OpenRouter · live
CoreWeave$1.31 / $3.96fp8$0.13 / $0.28fp8$0.76 / $2.42fp4——OpenRouter · live
Decart——$0.39 / $2.40mxfp4——OpenRouter · live
DeepInfra$1.30 / $2.60fp8$0.06 / $0.18fp8$0.56 / $1.80fp4——OpenRouter · live
DeepSeek$0.66 / $1.98?————OpenRouter · live
DigitalOcean$1.32 / $3.96?$0.12 / $0.24?$0.70 / $2.20?——OpenRouter · live
Fireworks——$1.40 / $4.40?——OpenRouter · live
Friendli——$1.40 / $4.40?——OpenRouter · live
GMICloud$1.06 / $3.17fp8$0.29 / $0.86fp8$1.40 / $4.40fp8——OpenRouter · live
Google————$2.00 / $10.00?OpenRouter · live
Inceptron—$0.05 / $0.65fp4$1.25 / $4.39fp4——OpenRouter · live
InferenceNet——$0.10 / $2.20?——OpenRouter · live
Ionstream$0.21 / $1.48?————OpenRouter · live
Mancer 2—$0.20 / $0.60fp8———OpenRouter · live
Mistral——$1.40 / $4.40nvfp4——OpenRouter · live
Morph—$0.14 / $0.40bf16$0.17 / $3.07fp8——OpenRouter · live
NextBit$1.06 / $3.17fp8$0.35 / $1.06fp8———OpenRouter · live
Novita$0.99 / $2.97fp8$0.41 / $1.23fp8$0.65 / $2.04fp8——OpenRouter · live
OpenAI———$1.00 / $5.00?—OpenRouter · live
OpenInference—$0.00 / $1.04fp8———OpenRouter · live
Parasail$1.32 / $3.96fp8$0.14 / $0.28fp8$1.40 / $4.40fp4——OpenRouter · live
Phala$0.96 / $2.88?$0.31 / $0.92?$1.26 / $3.00fp8——OpenRouter · live
Reka—$0.02 / $0.53?———OpenRouter · live
Relace$0.90 / $2.70fp4$0.01 / $1.28fp4$0.10 / $4.00?——OpenRouter · live
Sail Research—$0.02 / $0.30fp4———OpenRouter · live
SiliconFlow$1.32 / $3.96fp8$0.22 / $0.66fp8$0.70 / $2.20fp8——OpenRouter · live
StreamLake$0.66 / $1.98?$0.04 / $0.13fp8$0.28 / $0.88fp8——OpenRouter · live
Together$1.32 / $3.96?$0.14 / $0.28?$1.40 / $4.40?——OpenRouter · live
Venice$1.65 / $4.95?$0.17 / $0.35?$1.40 / $4.40fp8——OpenRouter · live
Wafer$0.33 / $4.20?$0.12 / $0.70?$0.41 / $3.99?——OpenRouter · live
Z.ai / Zhipu——$1.40 / $4.40fp8——OpenRouter · live

Inference hosts & serving providers (69)

Every provider that serves a tracked model’s endpoint — open-weight hosts, cloud platforms, and labs serving their own models. These are priced per model in the “where to buy” tables on each provider page (from OpenRouter’s per-endpoint data, with quantization), which is why they are a roster here rather than a price grid: a host’s price depends on the specific model and precision, not a single rate. Click through for the per-model breakdown.

ProviderOriginSite
AkashML🇺🇸 United Statesakashml.com
Alibaba🇨🇳 Chinawww.alibabacloud.com/product/machine-learning
Amazon Bedrock🇺🇸 United Statesaws.amazon.com/bedrock/
Anthropic🇺🇸 United Stateswww.anthropic.com
Arcee AI—www.arcee.ai
AtlasCloud—www.atlascloud.ai
Azure🇺🇸 United Statesazure.microsoft.com/products/ai-foundry
Baidu🇨🇳 Chinacloud.baidu.com
BaseTen🇺🇸 United Stateswww.baseten.co
Cerebras🇺🇸 United Stateswww.cerebras.ai
Chutes—chutes.ai
Claude Platform on AWS—aws.amazon.com/claude-platform/
Cloudflare🇺🇸 United Statesdevelopers.cloudflare.com/workers-ai/
Cohere🇨🇦 Canadacohere.com
CoreWeave—www.coreweave.com
Crusoe🇺🇸 United Statescrusoe.ai
Darkbloom—www.darkbloom.dev
Decart🇮🇱 Israelwww.decart.ai
DeepInfra🇺🇸 United Statesdeepinfra.com
DeepSeek🇨🇳 Chinawww.deepseek.com
DekaLLM—www.cloudeka.id
DigitalOcean🇺🇸 United Stateswww.digitalocean.com/products/gradient
Fireworks🇺🇸 United Statesfireworks.ai
Friendli🇰🇷 South Koreafriendli.ai
GMICloud🇺🇸 United Stateswww.gmicloud.ai
Google🇺🇸 United Statescloud.google.com/vertex-ai
Google AI Studio—aistudio.google.com
Groq🇺🇸 United Statesgroq.com
Inception🇺🇸 United Stateswww.inceptionlabs.ai
Inceptron—www.inceptron.io
InferenceNet—inference.net
Io Net🇺🇸 United Statesio.net
Ionstream—ionstream.ai
Liquid—www.liquid.ai
Makora—www.makora.com
Mancer 2—mancer.tech
Mara—www.mara.com
Meta🇺🇸 United Statesai.meta.com
Minimax🇨🇳 Chinawww.minimax.io
Mistral🇫🇷 Francemistral.ai
Modal—modal.com
ModelRun—www.modular.com
Moonshot AI—www.moonshot.ai
Morph🇺🇸 United Statesmorphllm.com
Near AI—near.ai
Nebius🇳🇱 Netherlandsnebius.com
NextBit—www.nextbit256.com
Novita—novita.ai
Nvidia—www.nvidia.com
OpenAI🇺🇸 United Statesopenai.com
OpenInference—www.openinference.ai
Parasail🇺🇸 United Stateswww.parasail.io
Phala—phala.network
PrimeIntellect—www.primeintellect.ai
Reka—reka.ai
Relace—www.relace.ai
Sail Research—www.sailresearch.com
SambaNova🇺🇸 United Statessambanova.ai
SiliconFlow🇨🇳 Chinawww.siliconflow.com
StepFun🇨🇳 Chinaplatform.stepfun.ai
StreamLake🇨🇳 Chinawww.streamlake.ai
Tencent🇨🇳 Chinawww.tencentcloud.com
Together🇺🇸 United Stateswww.together.ai
Upstage🇰🇷 South Koreawww.upstage.ai
Venice🇺🇸 United Statesvenice.ai
Wafer—wafer.ai
xAI🇺🇸 United Statesx.ai
Xiaomi🇨🇳 Chinaplatform.xiaomimimo.com
Z.AI🇨🇳 Chinaz.ai

Pick of the week

What I actually run5 on record

What to run locally

Your GPU, not a bill

VRAM figures are real GGUF bytes and coding numbers are our own first-party runs where the AA index has none. The catalogue keeps growing; treat it as a solid guide, not gospel, and corrections are welcome.

Our BigCodeBench-Hard runs were measured on the DIIC computing cluster (Departamento de Ingeniería de la Información y las Comunicaciones, Universidad de Murcia), plus two desktop GPUs at home. The tok/s we publish are from the home cards only — that is the speed a person actually sees. Full credits in Method.

The best coding models that fit each VRAM budget — cumulative, because a bigger card also runs everything smaller. Sizes are real GGUF bytes of the file behind the score: for a model we measured at several quantisations, each budget shows the best-scoring build that fits it, with its name next to the size, and only builds with three clean runs count. Otherwise the size is at ~4-bit (measured where a build exists). Add the KV cache the context needs — weights alone are not the VRAM figure the review videos quote. Ranked on one unified scale so the pick is realistic at every budget: our measured BigCodeBench-Hard (shown as % BCB) for the models we actually ran on this GPU, and the Artificial-Analysis coding index (shown as idx) for the big ones we can’t run locally. The two are not the same scale and we no longer pretend otherwise: until 2026-08-23 we converted between them with a fit trained on our older protocol, and the new measurements showed that fit was wrong by up to 18 points. So each row says where its number came from, and a tier can mix a model we measured with one only the index covers. Quality still isn’t proportional to size — a well-built 12B beats a 20B. Each model links to how to run it.

≤8 GB

#ModelCoderParamsWeightsMin VRAM @8KType
1Qwen3.5 9B28% BCB9B5 GB Q4_K_M6 GBdense
2Spark-X2.5-4B28% BCB4B4 GB Q8_05 GBdense
3Ternary Bonsai 2 27B27% BCB27B7 GB PQ2_08 GBdense

≤12 GB

#ModelCoderParamsWeightsMin VRAM @8KType
1Qwen3.8-27B28% BCB27B9 GB UD-Q2_K_XL11 GBdense
2Qwen3.5 9B28% BCB9B5 GB Q4_K_M6 GBdense
3Spark-X2.5-4B28% BCB4B4 GB Q8_05 GBdense

≤16 GB

Best pick unchanged from ≤12 GB: Qwen3.8-27B still leads and needs only 13 GB. More VRAM here buys headroom and context, not a better coder — the larger open models score lower.

#ModelCoderParamsWeightsMin VRAM @8KType
1Qwen3.8-27B34% BCB27B11 GB UD-IQ3_S13 GBdense
2Qwen3.6 27B30% BCB27B13 GB Q3_K_M15 GBdense
3Qwen3.5 9B28% BCB9B5 GB Q4_K_M6 GBdense

≤24 GB

#ModelCoderParamsWeightsMin VRAM @8KType
1K2 Horizon 32B34% BCB35B20 GB Q4_K_M24 GBdense
2Qwen3.8-27B34% BCB27B11 GB UD-IQ3_S13 GBdense
3Qwen3-Coder 30B A3B32% BCB31B19 GB21 GBMoE

≤32 GB

Best pick unchanged from ≤24 GB: K2 Horizon 32B still leads and needs only 24 GB. More VRAM here buys headroom and context, not a better coder — the larger open models score lower.

#ModelCoderParamsWeightsMin VRAM @8KType
1K2 Horizon 32B34% BCB35B20 GB Q4_K_M24 GBdense
2Qwen3.8-27B34% BCB27B11 GB UD-IQ3_S13 GBdense
3Qwen3-Coder 30B A3B32% BCB31B19 GB21 GBMoE

≤48 GB

Best pick unchanged from ≤32 GB: K2 Horizon 32B still leads and needs only 24 GB. More VRAM here buys headroom and context, not a better coder — the larger open models score lower.

#ModelCoderParamsWeightsMin VRAM @8KType
1K2 Horizon 32B34% BCB35B20 GB Q4_K_M24 GBdense
2Qwen3.8-27B34% BCB27B11 GB UD-IQ3_S13 GBdense
3Qwen3-Coder 30B A3B32% BCB31B19 GB21 GBMoE

≤96 GB

#ModelCoderParamsWeightsMin VRAM @8KType
1Qwen3.8-Flash-Next34% BCB180B84 GB UD-Q3_K_XL93 GBdense
2DeepSeek V4 Flash34% BCB291B85 GB UD-IQ2_M93 GBMoE
3K2 Horizon 32B34% BCB35B20 GB Q4_K_M24 GBdense

≤141 GB

Best pick unchanged from ≤96 GB: Qwen3.8-Flash-Next still leads and needs only 93 GB. More VRAM here buys headroom and context, not a better coder — the larger open models score lower.

#ModelCoderParamsWeightsMin VRAM @8KType
1Qwen3.8-Flash-Next34% BCB180B84 GB UD-Q3_K_XL93 GBdense
2DeepSeek V4 Flash34% BCB291B85 GB UD-IQ2_M93 GBMoE
3K2 Horizon 32B34% BCB35B20 GB Q4_K_M24 GBdense

Multi-GPU

Best pick unchanged from ≤141 GB: Qwen3.8-Flash-Next still leads and needs only 93 GB. More VRAM here buys headroom and context, not a better coder — the larger open models score lower.

#ModelCoderParamsWeightsMin VRAM @8KType
1Qwen3.8-Flash-Next34% BCB180B84 GB UD-Q3_K_XL93 GBdense
2DeepSeek V4 Flash34% BCB291B85 GB UD-IQ2_M93 GBMoE
3K2 Horizon 32B34% BCB35B20 GB Q4_K_M24 GBdense

Every model, ranked

All 90 models in the catalogue by AA’s coding index, the one it froze on 2026-09-11 — the only AA figure most of these small models ever got. Filter the Fits or Type column to narrow to your hardware, or sort any column. HumanEval* and BCB* (BigCodeBench-Hard) are first-party pass@1 scores we ran ourselves — HumanEval on 12 models, the harder BCB-Hard on 66, on the real GPU. Note how HumanEval saturates near the top while BCB-Hard spreads the field out — that gap is why we run both. An empty BCB cell means we have not run that model yet: we used to fill it with an estimate from the coding index and withdrew those on 2026-08-23, because the fit behind them was trained on a protocol we have since retired. t/s is measured on our own home cards, an RTX 4060 Ti 16GB or an R9700 32GB; the model page says which. Click a measured % for how it was run.

ModelCoding idxHumanEval*BCB*t/sParamsWeights ~Q4Min VRAMFitsType
DeepSeek V4 Flash69.1—33%—291B137 GB151 GBMulti-GPUMoE
Qwen3-Coder 30B A3B—98%32%159 t/s31B19 GB21 GB≤24 GBMoE
Qwen3.8-Flash-Next73.1—32%—180B94 GB104 GB≤141 GBdense
Qwen3.8-27B68.1—32%41 t/s27B13 GB16 GB≤16 GBdense
K2 Horizon 32B——32%—35B70 GB79 GB≤96 GBdense
Llama 3.3 70B11.9—30%—70B43 GB49 GB≤96 GBdense
Qwen3.8-3.6-27B-blend——30%—28B17 GB20 GB≤24 GBdense
Nemotron 3.5 Lightning26.8—29%—30B25 GB——MoE
Qwen3.5 9B28.740%29%72 t/s9B6 GB7 GB≤8 GBdense
Laguna XS.2——29%—33B21 GB23 GB≤24 GBMoE
Qwen3.8-35B-A3B-Distill——29%—35B22 GB24 GB≤32 GBdense
Qwen3.6 27B53.7—28%—27B17 GB20 GB≤24 GBdense
K2 Horizon 7B38.6—28%—9B18 GB21 GB≤24 GBdense
Qwen3.6 35B A3B41.9—28%—35B22 GB25 GB≤32 GBdense
Hemmingway-1——28%—27B17 GB20 GB≤24 GBdense
JetBrains Mellum 2 12B——27%—12B8 GB9 GB≤12 GBMoE
K2 Horizon MoVA 36B-A4B——27%—37B22 GB26 GB≤32 GBMoE
gpt-oss-20b20.797%27%159 t/s21B12 GB14 GB≤16 GBMoE
Ternary Bonsai 2 27B——27%—27B7 GB8 GB≤12 GBdense
Ornith-1.5-35B-A3B——27%—36B22 GB24 GB≤32 GBdense
Ornith-1.0 35B——26%107 t/s35B22 GB25 GB≤32 GBdense
Ornith-1.0 9B——26%72 t/s9B6 GB7 GB≤8 GBdense
Agnes-3.0-Flash——26%—33B20 GB23 GB≤24 GBdense
Spark-X2.5-4B——25%—4B3 GB3 GB≤8 GBdense
Granite 4.1 30B10.4—25%—29B17 GB21 GB≤24 GBdense
MiniMax M252.6—25%—229B138 GB154 GBMulti-GPUMoE
KAT-Coder V2.5 35B——25%—35B21 GB24 GB≤24 GBdense
Ling-3.0-flash57.0—25%—124B79 GB——MoE
Qwen3 8B9.0—25%—8B5 GB7 GB≤8 GBdense
Ternary Bonsai 27B—90%24%30 t/s27B7 GB8 GB≤12 GBdense
NeoHorse-1-9B——24%—9B6 GB7 GB≤8 GBdense
Ministral 3 14B14.463%24%30 t/s14B8 GB10 GB≤12 GBdense
Qwen3 32B15.3—24%—33B20 GB est24 GB≤24 GBdense
Qwen3-4B Instruct 2507——24%—4B2 GB4 GB≤8 GBdense
Ornith-1.5 9B——24%73 t/s9B7 GB8 GB≤8 GBdense
North Mini Code 1.036.5—24%—30B18 GB est20 GB≤24 GBMoE
Qwen3 14B13.863%23%25 t/s14B9 GB11 GB≤12 GBdense
Mistral Small 3.2 24B12.583%23%11 t/s24B14 GB17 GB≤24 GBdense
K2 Horizon 3.7B26.1—23%—5B10 GB12 GB≤16 GBdense
Qwen3.5 4B22.633%22%71 t/s5B3 GB3 GB≤8 GBdense
Qwen3.6-27B-A3B-Coder——22%—27B16 GB18 GB≤24 GBdense
R1-Distill-Qwen-32B——21%—33B20 GB24 GB≤24 GBdense
EXAONE 4.5 33B23.6—21%—33B20 GB24 GB≤32 GBdense
Nex-N2.5-mini——20%—35B22 GB25 GB≤32 GBdense
GLM-5.3-Flash71.5—20%—321B157 GB——dense
Granite 4.1 8B9.567%20%76 t/s8B5 GB7 GB≤8 GBdense
MiniCPM5-2B14.5—20%—2B2 GB2 GB≤8 GBdense
Gemma 4 31B43.4—19%—31B18 GB——dense
Xing4.0-29B-A4B——19%—31B20 GB27 GB≤32 GBMoE
Confucius4-T3PO——19%—15B11 GB13 GB≤16 GBdense
Nemotron Nano 9B V2——18%—9B7 GB——dense
MiMo-V2.6-Distill-Qwen-9B——18%—9B6 GB est6 GB≤8 GBdense
Llama 3.1 8B Instruct5.4—17%—8B5 GB est——dense
Nanbeige 4.2-3B—81%16%9 t/s4B3 GB4 GB≤8 GBdense
Granite 4.0 Micro——16%—3B2 GB3 GB≤8 GBdense
NeoHorse-1-4B——16%—4B3 GB3 GB≤8 GBdense
R1-Distill-Qwen-14B——16%—15B9 GB11 GB≤12 GBdense
Gemma 3 4B2.7—11%—4B2 GB3 GB≤8 GBdense
xLAM 2 3B——9%—3B2 GB2 GB≤8 GBdense
K2 Horizon 0.9B3.4—9%—1B2 GB3 GB≤8 GBdense
Bonsai 8B—63%7%156 t/s8B1 GB2 GB≤8 GBdense
Instella MoE 16B A3B——6%—16B10 GB12 GB≤12 GBMoE
Phi-3 mini 128k——5%—4B2 GB6 GB≤8 GBdense
Qwen3.5 2B2.9—5%—2B1 GB2 GB≤8 GBdense
SmolLM3-3B——4%—3B2 GB3 GB≤8 GBdense
CobrIX-1.5-Coder-Flash-33B-A13B——3%—33B20 GB23 GB≤24 GBMoE
GLM-5.268.8———753B466 GB514 GBMulti-GPUMoE
DeepSeek V4 Pro68.8———862B517 GB est569 GBMulti-GPUMoE
Motif-3 Beta63.5———315B196 GB217 GBMulti-GPUMoE
Kimi K2.661.8———1027B616 GB est692 GBMulti-GPUdense
Kimi K2.7 Code60.8———1027B495 GB546 GBMulti-GPUdense
MiMo-V2.5-Pro60.2———1023B630 GB693 GBMulti-GPUMoE
Qwen3.5 397B A17B59.1———397B244 GB270 GBMulti-GPUdense
Nex-N2-Pro59.1———397B242 GB267 GBMulti-GPUdense
Hy358.8———299B182 GB203 GBMulti-GPUMoE
GLM-5.155.8———754B465 GB512 GBMulti-GPUMoE
Inkling52.9———952B571 GB est631 GBMulti-GPUdense
Nemotron 3 Ultra 550B49.3———550B359 GB——MoE
Mistral Medium 3.546.9———128B75 GB85 GB≤96 GBdense
Qwen3.5 122B A10B45.7———125B78 GB87 GB≤96 GBdense
Solar Open2 250B45.0———250B150 GB est——MoE
Gemma 4 26B A4B39.3———26B17 GB——dense
Nemotron 3 Super 120B37.7———120B83 GB——MoE
Gemma 4 12B31.063%—31 t/s12B7 GB——dense
gpt-oss-120b30.4———117B63 GB69 GB≤96 GBMoE
Nemotron Cascade 2 30B A3B25.3———32B25 GB——MoE
Qwen3 235B A22B 250722.1———235B142 GB158 GBMulti-GPUMoE
Qwen3 Next 80B A3B17.4———80B49 GB54 GB≤96 GBMoE
Laguna S 2.1————118B73 GB81 GB≤96 GBMoE
Needle 3————121M0 GB0 GB≤8 GBdense

Each tier is everything that fits that budget, so the same strong small model can top several tiers — that repetition is the point. “Min VRAM” is weights plus KV cache at an 8K context; longer context or a heavier quant (Q8/fp16) needs more. A dash under coding index is unmeasured, never zero. Data in local.yaml.

It’s free

beta$0 to start

Beta. New view — free tiers change constantly, so amounts are approximate and “free for now” is a real caveat. Verify before relying on any of it.

Every way we track to run a capable model for nothing — a real free allowance, not a trial that needs a card. Ranked roughly by how much you get. The catch column is the honest part: rate limits, one-off vs monthly, or “free for now”.

Free API credits & allowances

Hosted, OpenAI-compatible APIs with a genuine free tier over open (and some closed) models.

ProviderWhat you get freeThe catch
OpenRouter17 :free routes today — apodex-1.1-mini, dots-3-note-preview, gemma-4-26b-a4b-it, gemma-4-31b-it, inkling, inkling-small, laguna-s-2.1, laguna-xs-2.1, lfm-2.5-2.6b, ling-3.0-flash-sante, nemotron-3-nano-omni-30b-a3b-reasoning, nemotron-3-super-120b-a12b, nemotron-3-ultra-550b-a55b, nemotron-3.5-content-safety, nemotron-3.5-lightning, north-mini-code, qwen3.8-27b~50 requests/day (1,000 once you have topped up $10+), 20/min; the set rotates without notice — snapshot as of 2026-10-03
opencode ZenRotating free set (DeepSeek V4 Flash, MiMo-V2.5, Hy3, Laguna S 2.1, Nemotron 3 Ultra, Nemotron 3.5 Lightning, Muse Spark 1.2) plus the Ox Alpha stealth model, free this week (1M context, multimodal)free through the opencode CLI; the set rotates without notice — as of 2026-10-03
Scaleway Generative APIs1,000,000 tokens + 60 minutos de transcripcion de audiouna vez, cuentas nuevas
Google AI Studio (Gemini API)~1,000–1,500 requests/day on Gemini 3 Flash / 2.5 Flashno card, but free-tier prompts may be used to improve Google's models; ~5–15 req/min; Pro models removed from the free tier (Apr 2026)
Groq~500K tokens/day per model + ~14,400 requests/dayno card; 30 req/min and 6K tokens/min caps; resets midnight UTC
Mistral La Plateforme“Experiment” free tier — ~1B tokens/montheval/prototyping only, not production; phone verification and a data-sharing opt-in
Cohere1,000 API calls/month (trial keys)non-production; ~20 req/min; data may be used to improve models
Cloudflare Workers AI10,000 Neurons/day (Free and Paid plans alike)no card; resets 00:00 UTC; on the Free plan you are blocked until reset once it is spent
APIMaster$20 de credito GPTprueba de 5 dias, a precio oficial y solo para modelos GPT; mas 15% en la primera recarga de $10 y $20 por referido
Nscaleno confirmadoni eso: la frase de los creditos para primeros usuarios ya no aparece en su web (2026-10-02). Lo unico con cifra son $2 de credito, y es de ajuste fino, no de inferencia
OVHcloud AI EndpointsUS$200 de credito Public Cloud (en dolares) + varios modelos siempre gratisproyectos nuevos de Public Cloud
Public AI$2 de credito de iniciouna vez, cuentas nuevas; limites de 100 peticiones/min en Free, 200 en Plus, 300 en Pro
Hugging Face Inference Providers$0.10/mo free · $2/mo on PROmonthly; pass-through provider rates after
CerebrasFree trial ($5 credits) over all models (gpt-oss-120b, GLM-4.7, Gemma 4 31B)no card for the trial; ~8K-token context cap; the widely-cited “sunset 2026-08-17” applies only to PREVIEW models in the paid Developer tier, not the free tier
SambaNova CloudPersistent free tier + $5 signup creditthe $5 credit expires in 3 months; per-model daily caps (some ~20 req/day)
Nebius Token FactoryNo-card free trial credit (~$1) over open models (Llama-3.3-70B, Qwen3-235B)rebranded from Nebius AI Studio in 2026; the free trial credit is now small (~$1), and a card is needed for the $25 minimum top-up
NVIDIA Buildsin credito publicado: el tramo gratis es por limite de peticioneshasta 40 peticiones por minuto en la mayoria de modelos, sin facturacion por token; produccion via NIM / AI Enterprise

No longer free (don’t bother): Together (signup credit retired), Fireworks (~$1, testing only), Chutes (free tier ended Mar 15 2026), and GitHub Models (retired 2026-07-30). Listed so you don’t go looking.

Free forever (nothing to burn)

Run it yourself. Every open-weight model in Run local is free once you have the GPU — a 12B fits an 8 GB card and codes well, with no rate limit and no rotation. The rate-limited-free routes above (OpenRouter :free, opencode Zen) are also $0 with no expiry — the trade there is the provider’s rate cap and a set that rotates without notice.

Free subscription tiers

Products with a permanent $0 plan (limited, but real).

ProviderFree tier
GitHubFree
MistralFree
MorphFree
OllamaFree
OpenAIFree
QoderFree
VeniceFree
Windsurf (Cognition/Devin)Free

Data from gateways.yaml and plans.yaml. “Free for now” means exactly that — promotional tiers get cut (Tencent Hy3’s free window already closed), so treat any allowance as this week’s, not forever.

Our tests

Moved to Yardstick

We have moved. Everything we measure ourselves — BigCodeBench-Hard, BFCL v4, RULER, SWE-bench, speed and energy — now lives at yardstick.millaguie.net, together with the harness that produces it. It is published per card, which this page never was.

This site keeps what it is for: what the models cost, who serves them and at what precision. Where a model of ours has a first-party score, its own page still shows it, and the leaderboard still carries the three columns marked ᴹ.

What changed

Since the snapshot of 2026-10-02.

  • Kimi K3 — cost per session $1.1813 → $1.1803
  • Nemotron 3.5 Lightning — input price $0.06 → $0.06
  • Nemotron 3.5 Lightning — output price $0.16 → $0.17
  • GLM-5.2 — cost per session $1.5970 → $1.5959
  • Qwen3.7 Max — cost per session $0.9207 → $0.9202
  • Qwen3.7 Max — cache ratio 88.9% → 88.8%
  • Kimi K2.6 — cost per session $0.3774 → $0.3773
  • Kimi K2.7 Code — cost per session $0.8900 → $0.8896
  • Tencent Hy3 — input price $0.08 → $0.13
  • Tencent Hy3 — output price $0.33 → $0.53
  • DeepSeek V4 Flash — input price $0.01 → $0.01
  • Qwen3.7 Plus — cost per session $0.2010 → $0.2007
  • GLM-5.1 — input price $0.96 → $1.40
  • GLM-5.1 — output price $3.03 → $4.40
  • GLM-5.1 — cost per session $0.8461 → $0.8457
  • DeepSeek V3.1 Terminus — input price $0.30 → $0.27
  • Gemma 4 26B A4B — provider count 13 → 12
  • MiMo-V2.6-Flash — quantizations bf16, fp4, fp8, mxfp4 → bf16, fp4, fp8
  • DeepSeek V4.1 Flash — input price $0.02 → $0.30
  • DeepSeek V4.1 Flash — output price $0.60 → $1.20
  • DeepSeek V4.1 Flash — quantizations fp32, fp4, fp8, unknown → fp4, fp8, unknown

Historical data

The previous Model Speed Arena — weekly measured throughput runs against live subscriptions — is preserved at https://oldmsa.millaguie.net. Those runs are no longer updated: sustaining a weekly benchmark fleet was not viable. They remain the better source for measured tok/s under a real subscription, which no aggregator publishes.

Where the data comes from

Every figure on this page is somebody else’s measurement or publication. Nothing here is re-benchmarked locally, so each dataset is credited with what is taken from it and what that source can and cannot tell you.

Heads-up (2026-08-19): Stripe agreed to acquire OpenRouter, this site’s main source for list price, context and per-provider precision. Both sides stress OpenRouter stays “neutral infrastructure” and it keeps operating for now, but pricing and API terms could shift under new ownership — worth watching, since much of this page depends on it.

SourceWhat this site takes from itNature
OpenRouterList price, context window, and the per-provider quantization, uptime and endpoint listpublic API
models.devOfficial provider catalog, tiered and cache pricingpublic API
Artificial AnalysisIndependent quality evaluations — the intelligence index (which sets the ranking here since 2026-10-02), Terminal-Bench 4.0, the old coding index AA froze on 2026-09-11, TerminalBench-Hard and long-context reasoning — plus the only measured speed figures (tok/s, time to first token). These are AA’s own runs and weightings: we cannot recompute them, which is why a model AA has not tested cannot be rankedindependent lab
opencodeObserved production usage across its subscriber base: cost per session, tokens per session, cache ratioaggregate telemetry
quota-sentinelWhich providers expose a readable quota endpointopen source
Inference providersEvery serving provider links to its own site, each checked to resolve. Where a referral link exists it is used and marked, and the ranking is computed from price and precision alone regardlessneutral link
Vendor documentationSubscription prices and quotas, reproduced as advertised. Each plan row links to its own source and carries the date it was checkedvendor-reported
Reseller gatewaysNamed and classified in the Gateways section, never priced into the leaderboard: a reseller reprices someone else’s API, so its number is not a first-party pricereseller
Third-party write-upsUsed only where a vendor publishes no figure at all. Kimi’s per-tier request counts are the case in point: its own page gives adjectives, so the numbers come from independent write-ups instead and the plan is marked low-confidence. Every such plan links the write-up it came fromunofficial
Logged-in consoleFigures visible only inside an account, so a reader cannot click through to verify them: MiniMax’s remaining-quota percentage (from which the hidden quota here was derived), Google AI Plus’s euro price, and DeepSeek’s peak-pricing notice — which its public pricing page still does not mentionnot publicly verifiable
Measured hereOur own runs of BigCodeBench-Hard and SWE-bench Verified, on the official protocol: 148 problems, one attempt each, temperature 0, served with llama.cpp one request at a time. Comparable between our runs; not comparable with a public leaderboard, because we serve quantised GGUF files and they do not. Every number and how it was produced. Code and redacted trajectories are in the repositoryfirst-party

Two things worth carrying. Quality figures are tagged vendor or independent throughout, because the flattering numbers are almost always the vendor’s own. And observed-usage figures describe opencode’s subscribers, not the whole market — a model barely used there will have a noisy cost per session.

Method & limits

This site is an aggregator: price, context, per-provider quantization and uptime come from OpenRouter; official catalog pricing from models.dev; independent quality and speed from Artificial Analysis; observed production cost from opencode. Plan quotas were researched from vendors’ own documentation and are reproduced as advertised. And we run our own tests, stated rather than hidden: BigCodeBench-Hard on the official protocol (148 problems, one attempt each, temperature 0, llama.cpp serving one request at a time) and SWE-bench Verified with mini-swe-agent — the SWE runs so far cover a 30-instance pool and not the full 500, which is why that table has counts and no percentage. They are collected at yardstick.millaguie.net, labelled wherever they appear, and kept out of the leaderboard, which stays entirely Artificial Analysis’ coding index. One place does mix the two scales on purpose and says so: the VRAM-tier picks under Run local, where a model we measured has to be rankable against a big one we cannot run. Scores measured before 2026-08-23 used an earlier protocol over 121 problems; they have been withdrawn rather than converted, because they were not comparable.

Two honest limits. Quota comparison is mostly impossible: only one provider publishes a quota that converts to tokens, and several publish nothing numeric at all. And vendor multipliers are not comparable — some are measured against a free tier, others against a paid one, some per session rather than per month. Plan pricing moves fast; every figure here is dated.

Acknowledgements. The first-party local-model runs in this project were carried out on the DIIC computing cluster (Departamento de Ingeniería de la Información y las Comunicaciones, Universidad de Murcia). This publication has also benefited from the use of the GaiaLab infrastructure, as well as the infrastructure deployed for the Gaia6G (TSI-064100-2023-0013) and CoCoNet (TSI-064100-2023-0013) as part of the UNICO 6G R&D programme, funded by NextGenerationEU.