Model Spend Arena2ND ED.
provider Β· 2026-08-14

Ollama

Official site Β· πŸ‡ΊπŸ‡Έ United States

Plan

Ollama Cloud (Free / Pro / Max)

Tier$/monthIncluded usage
Free$0.00β€”
Pro$20.00β€”
Max$100.00β€”

console only units unstated

As advertised. 'Light usage' / '50x more than Free' / '5x more than Pro'; 5-hour session limits and weekly limits

What that actually means. The only plan surveyed that meters COMPUTE rather than tokens: billing follows "actual utilization of Ollama's cloud infrastructure - primarily GPU time". No hours, tokens or request figures are published for any tier, only relative multipliers. Concurrency is stated: 1 cloud model on Free, 3 on Pro, 10 on Max.

Where to check what is left. usage page in the account; an email lands at 90% of the limit

Models covered. deepseek-v4-pro, deepseek-v4-flash, kimi-k2.7-code, kimi-k2.6, minimax-m2.5, mistral-large-3, gpt-oss 20b/120b, nemotron-3-ultra

GPU-time metering inverts an assumption the rest of this site can take for granted. Everywhere else a token costs the same whoever generates it, so model speed affects only your patience. Here speed IS the quota: a model running at half the tokens per second consumes roughly twice the GPU time for identical output, so the slow models elsewhere on this page are the expensive ones here. Annual billing is $200 for Pro.

checked 2026-08-13 Β· high (prices, structure and metering basis) / high (that no numeric quota is published) Β· source

β—† Some links on this page are referral links: if you sign up through them this site may earn a commission, at no extra cost to you. Rankings, prices and measurements are taken from the sources listed under Method, and are not influenced by whether a provider has a referral programme.