LLM API pricing comparison
AI model pricing for 40 text and chat LLMs across 8 providers — input, output, cached-input and batch rates per 1M tokens, side by side. Sort, filter, open any model for its full price matrix, or use the token-cost calculator to rank models on your own workload. List prices, USD, no signup.
40 models
| Provider | Tier | ||||||
|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flashoff-peak | DeepSeek | Budget | 1M | $0.14 | $0.28 | $0.0028 | $0.21 |
| GPT-5 nanolegacy | OpenAI | Budget | — | $0.05 | $0.40 | $0.005 | $0.225 |
| Gemini 2.5 Flash-Lite | Budget | 1M | $0.10 | $0.40 | $0.01 | $0.25 | |
| Qwen-FlashModel Studio | Alibaba (Qwen) | Budget | 1M | $0.10 | $0.40 | — | $0.25 |
| Mistral Small 4 | Mistral | Budget | — | $0.15 | $0.60 | — | $0.375 |
| Llama 4 Scoutvia Together AI | Meta (Llama) | Budget | 1M | $0.18 | $0.59 | — | $0.385 |
| Llama 4 Maverickvia Together AI | Meta (Llama) | Mid-range | 1.05M | $0.27 | $0.85 | — | $0.56 |
| DeepSeek V4 Prooff-peak | DeepSeek | Flagship | 1M | $0.435 | $0.87 | $0.003625 | $0.6525 |
| GPT-5.6 Luna | OpenAI | Budget | 1.05M | $0.20 | $1.20 | $0.02 | $0.70 |
| GPT-5.4 nano | OpenAI | Budget | — | $0.20 | $1.25 | $0.02 | $0.725 |
| Gemini 3.1 Flash-Lite | Budget | 1M | $0.25 | $1.50 | $0.025 | $0.875 | |
| Mistral Large 3 | Mistral | Mid-range | — | $0.50 | $1.50 | — | $1.00 |
| Qwen-PlusModel Studio | Alibaba (Qwen) | Mid-range | 256K | $0.40 | $1.60 | — | $1.00 |
| Gemini 2.5 Flash | Mid-range | 1M | $0.30 | $2.50 | $0.03 | $1.40 | |
| Gemini 3.5 Flash-Lite | Budget | 1M | $0.30 | $2.50 | $0.03 | $1.40 | |
| Grok 4.3 | xAI | Mid-range | 1M | $1.25 | $2.50 | $0.20 | $1.88 |
| GPT-5.4 mini | OpenAI | Budget | — | $0.75 | $4.50 | $0.075 | $2.63 |
| o4-mini | OpenAI | Reasoning | 200K | $1.10 | $4.40 | $0.275 | $2.75 |
| Claude Haiku 4.5 | Anthropic | Budget | 200K | $1.00 | $5.00 | $0.10 | $3.00 |
| Magistral Medium | Mistral | Reasoning | — | $2.00 | $5.00 | — | $3.50 |
| Grok 4.5long-ctx | xAI | Flagship | 500K | $2.00 | $6.00 | $0.30 | $4.00 |
| Gemini 3.6 Flash | Mid-range | 1M | $1.50 | $7.50 | $0.15 | $4.50 | |
| Mistral Medium 3.5 | Mistral | Flagship | — | $1.50 | $7.50 | — | $4.50 |
| o3 | OpenAI | Reasoning | 200K | $2.00 | $8.00 | $0.50 | $5.00 |
| Qwen-MaxModel Studio | Alibaba (Qwen) | Flagship | — | $2.50 | $7.50 | — | $5.00 |
| Gemini 3.5 Flash | Mid-range | 1M | $1.50 | $9.00 | $0.15 | $5.25 | |
| GPT-5legacy | OpenAI | Flagship | 400K | $1.25 | $10.00 | $0.125 | $5.63 |
| Gemini 2.5 Prolong-ctx | Flagship | 1M | $1.25 | $10.00 | $0.125 | $5.63 | |
| Claude Sonnet 5intro | Anthropic | Mid-range | 1M | $2.00 | $10.00 | $0.20 | $6.00 |
| GPT-5.6 Terra | OpenAI | Mid-range | 1.05M | $2.00 | $12.00 | $0.20 | $7.00 |
| Gemini 3.1 Propreviewlong-ctx | Flagship | 1M | $2.00 | $12.00 | $0.20 | $7.00 | |
| GPT-5.4 | OpenAI | Mid-range | 1.05M | $2.50 | $15.00 | $0.25 | $8.75 |
| Claude Sonnet 4.6legacy | Anthropic | Mid-range | 1M | $3.00 | $15.00 | $0.30 | $9.00 |
| Claude Opus 4.8legacy | Anthropic | Flagship | 1M | $5.00 | $25.00 | $0.50 | $15.00 |
| Claude Opus 5 | Anthropic | Flagship | 1M | $5.00 | $25.00 | $0.50 | $15.00 |
| GPT-5.5 | OpenAI | Flagship | 1.05M | $5.00 | $30.00 | $0.50 | $17.50 |
| GPT-5.6 Sol | OpenAI | Flagship | 1.05M | $5.00 | $30.00 | $0.50 | $17.50 |
| Claude Fable 5 | Anthropic | Flagship | 1M | $10.00 | $50.00 | $1.00 | $30.00 |
| o3-pro | OpenAI | Pro | — | $20.00 | $80.00 | — | $50.00 |
| GPT-5.5 Pro | OpenAI | Pro | — | $30.00 | $180.00 | — | $105.00 |
List prices, USD, directional. Rates are provider list prices per 1M tokens and are meant for comparison, not billing. Batch rates shown as 50% off are derived where a provider offers batch but does not publish a separate figure. Preview, promo, intro, peak/off-peak, long-context, and third-party-host prices are labeled where they apply. Token counts vary by tokenizer, so per-token price is not always a like-for-like cost. Always confirm with the provider before relying on a number. Prices as of 2026-08-03.
Token-cost calculator
Enter your per-request token counts, cache rate and monthly volume to rank every model by real monthly cost. All math runs in your browser over the list prices below — nothing is sent anywhere.
Cheapest for this workload: Gemini 2.5 Flash-Lite at $180.00/mo ($0.0018/request)
Ranking uses current headline rates — intro and off-peak prices (flagged) may change.
| # | Model | Provider | Cost / request | Monthly cost | Applied |
|---|---|---|---|---|---|
| 1 | Gemini 2.5 Flash-Lite | $0.0018 | $180.00 | ||
| 2 | Qwen-Flash | Alibaba (Qwen) | $0.0018 | $180.00 | |
| 3 | DeepSeek V4 Flashoff-peak | DeepSeek | $0.0020 | $196.00 | |
| 4 | Mistral Small 4 | Mistral | $0.0027 | $270.00 | |
| 5 | Llama 4 Scout | Meta (Llama) | $0.0030 | $298.00 | |
| 6 | GPT-5.6 Luna | OpenAI | $0.0044 | $440.00 | |
| 7 | Llama 4 Maverick | Meta (Llama) | $0.0044 | $440.00 | |
| 8 | GPT-5.4 nano | OpenAI | $0.0045 | $450.00 | |
| 9 | Gemini 3.1 Flash-Lite | $0.0055 | $550.00 | ||
| 10 | DeepSeek V4 Prooff-peak | DeepSeek | $0.0061 | $609.00 | |
| 11 | Qwen-Plus | Alibaba (Qwen) | $0.0072 | $720.00 | |
| 12 | Gemini 2.5 Flash | $0.0080 | $800.00 | ||
| 13 | Gemini 3.5 Flash-Lite | $0.0080 | $800.00 | ||
| 14 | Mistral Large 3 | Mistral | $0.0080 | $800.00 | |
| 15 | GPT-5.4 mini | OpenAI | $0.0165 | $1,650.00 | |
| 16 | Grok 4.3 | xAI | $0.0175 | $1,750.00 | |
| 17 | o4-mini | OpenAI | $0.0198 | $1,980.00 | |
| 18 | Claude Haiku 4.5 | Anthropic | $0.0200 | $2,000.00 | |
| 19 | Gemini 3.6 Flash | $0.0300 | $3,000.00 | ||
| 20 | Mistral Medium 3.5 | Mistral | $0.0300 | $3,000.00 | |
| 21 | Magistral Medium | Mistral | $0.0300 | $3,000.00 | |
| 22 | Grok 4.5 | xAI | $0.0320 | $3,200.00 | |
| 23 | Gemini 2.5 Pro | $0.0325 | $3,250.00 | ||
| 24 | Gemini 3.5 Flash | $0.0330 | $3,300.00 | ||
| 25 | o3 | OpenAI | $0.0360 | $3,600.00 | |
| 26 | Claude Sonnet 5intro | Anthropic | $0.0400 | $4,000.00 | |
| 27 | Qwen-Max | Alibaba (Qwen) | $0.0400 | $4,000.00 | |
| 28 | GPT-5.6 Terra | OpenAI | $0.0440 | $4,400.00 | |
| 29 | GPT-5.4 | OpenAI | $0.0550 | $5,500.00 | |
| 30 | Claude Opus 5 | Anthropic | $0.1000 | $10,000.00 | |
| 31 | GPT-5.5 | OpenAI | $0.1100 | $11,000.00 | |
| 32 | GPT-5.6 Sol | OpenAI | $0.1100 | $11,000.00 | |
| 33 | Claude Fable 5 | Anthropic | $0.2000 | $20,000.00 | |
| 34 | o3-pro | OpenAI | $0.3600 | $36,000.00 | |
| 35 | GPT-5.5 Pro | OpenAI | $0.6600 | $66,000.00 |
Costs are directional. Batch applies only where a provider offers it; the cached rate applies to the cached share of input only where the provider publishes one. Reasoning models generate hiddenreasoning tokens billed as output — measure real usage before trusting a per-request figure.
Browse by provider
Anthropic
Anthropic makes the Claude family — frontier models tuned for reliable tool use, long-context work, and coding.
OpenAI
OpenAI ships the GPT-5 series plus the o-series reasoning models.
Google’s Gemini API pairs 1M-token context windows with some of the lowest budget-tier prices.
xAI
xAI’s Grok models compete on output price and large context windows.
DeepSeek
DeepSeek offers open-weight-derived models at aggressive prices, with deep prompt-cache discounts.
Mistral
Mistral’s European lineup runs from the flagship Medium to the very cheap Small, plus the Magistral reasoning model.
Meta (Llama)
Meta publishes the open-weight Llama 4 models but does not sell tokens directly.
Alibaba (Qwen)
Alibaba’s Qwen models are served through Model Studio.
How LLM API pricing works
LLM APIs bill per token, quoted in dollars per 1M tokens, split into an inputrate (everything you send) and a higher output rate (what the model generates). Beyond the sticker rate, real cost hinges onprompt caching, the Batch API,reasoning tokens, and each model'stokenizer — two models at the same per-token price can cost different amounts on the same text.
This explorer labels the honest complications competitors gloss over: preview models, promos, intro pricing, off-peak multipliers,long-context surcharges, and third-party hosting. New to the terms? Start with the LLM pricing glossary. Looking at cloud compute too? See the AWS EC2 Pricing Explorer.