LLM Pricing: API Prices for 50+ Models
50 models from 10 providers · prices checked Oct 2, 2026
NewGemini 4 Argon cost calculatorEnter token counts to add a cost column to the table. Open the full calculator →
| Model | Provider | Tier | Checked | ||||
|---|---|---|---|---|---|---|---|
| GPT-6 Astra OpenAI · 1.05M ctx | OpenAI | Flagship | $10* | $50* | $1 | 1.05M | Oct 1, 2026 |
| GPT-5.6 Sol OpenAI · 1.05M ctx · Promo price Promotional price, available at least through Nov 21, 2026. | OpenAI | Flagship | $4* | $20* | $0.40 | 1.05M | Oct 1, 2026 |
| GPT-5.5 OpenAI · 1.05M ctx | OpenAI | Flagship | $5* | $30* | $0.50 | 1.05M | Oct 1, 2026 |
| GPT-6.1 Sol OpenAI · 1.05M ctx | OpenAI | Mid | $2* | $10* | $0.10 | 1.05M | Oct 1, 2026 |
| GPT-6 Sol OpenAI · 1.05M ctx | OpenAI | Mid | $2* | $10* | $0.20 | 1.05M | Oct 1, 2026 |
| GPT-5.6 Terra OpenAI · 1.05M ctx | OpenAI | Mid | $2* | $12* | $0.20 | 1.05M | Oct 1, 2026 |
| GPT-6 Luna OpenAI · 1.05M ctx | OpenAI | Small | $0.10* | $0.50* | $0.01 | 1.05M | Oct 1, 2026 |
| GPT-5.6 Luna OpenAI · 1.05M ctx | OpenAI | Small | $0.20* | $1.20* | $0.02 | 1.05M | Oct 1, 2026 |
| Claude Opus 5.5 Anthropic · 1M ctx | Anthropic | Flagship | $4 | $20 | $0.20 | 1M | Oct 1, 2026 |
| Claude Fable 5.1 Anthropic · 1M ctx | Anthropic | Flagship | $10 | $50 | $0.25 | 1M | Oct 1, 2026 |
| Claude Sonnet 5.5 Anthropic · 1M ctx | Anthropic | Mid | $2 | $10 | $0.20 | 1M | Oct 1, 2026 |
| Claude Haiku 4.5 Anthropic · 200K ctx | Anthropic | Small | $1 | $5 | $0.10 | 200K | Oct 1, 2026 |
| Gemini 3.1 Pro Preview Google · 1.05M ctx | Flagship | $2* | $12* | $0.20 | 1.05M | Oct 1, 2026 | |
| Gemini 4 Argon Google · — ctx · Announced Announced, not yet in the API. Introductory price; then $4 / $20. | Flagship | $2 | $10 | $0.10 | — | Oct 2, 2026 | |
| Gemini 3.8 Flash Google · 1.05M ctx · Promo price Promotional price until Dec 31, 2026; then $1.50 / $7.50. | Mid | $0.75 | $3.75 | $0.075 | 1.05M | Oct 1, 2026 | |
| Gemini 3.7 Flash Google · 1.05M ctx · Promo price Promotional price until Dec 31, 2026; then $1.50 / $7.50. | Mid | $0.75 | $3.75 | $0.075 | 1.05M | Oct 1, 2026 | |
| Gemini 3.6 Flash Google · 1.05M ctx · Promo price Promotional price until Dec 31, 2026; then $1.50 / $7.50. | Mid | $0.75 | $3.75 | $0.075 | 1.05M | Oct 1, 2026 | |
| Gemini 3.5 Flash-Lite Google · 1.05M ctx | Small | $0.30 | $2.50 | $0.03 | 1.05M | Oct 1, 2026 | |
| Gemini 3.1 Flash-Lite Google · 1.05M ctx | Small | $0.25 | $1.50 | $0.025 | 1.05M | Oct 1, 2026 | |
| DeepSeek-V4-Pro DeepSeek · 1M ctx · Peak-hour price Peak-hour price; 50% off outside peak hours. | DeepSeek | Flagship | $1.32 | $3.96 | $0.044 | 1M | Oct 1, 2026 |
| DeepSeek-V4.1-Flash DeepSeek · 1M ctx · Peak-hour price Peak-hour price; 50% off outside peak hours. | DeepSeek | Small | $0.30 | $1.20 | $0.006 | 1M | Oct 1, 2026 |
| Grok 4.7 xAI · 500K ctx | xAI | Flagship | $2* | $6* | $0.50 | 500K | Oct 1, 2026 |
| Grok 4.6 xAI · 500K ctx | xAI | Flagship | $2* | $6* | $0.50 | 500K | Oct 1, 2026 |
| Grok 4.5 xAI · 500K ctx | xAI | Flagship | $2* | $6* | $0.30 | 500K | Oct 1, 2026 |
| Grok 4.3 xAI · 1M ctx | xAI | Mid | $1.25* | $2.50* | $0.20 | 1M | Oct 1, 2026 |
| Grok 4.20 xAI · 1M ctx | xAI | Mid | $1.25* | $2.50* | $0.20 | 1M | Oct 1, 2026 |
| Grok Build 0.1Preview xAI · 256K ctx | xAI | Small | $1* | $2* | $0.20 | 256K | Oct 1, 2026 |
| Qwen3.8-Max Alibaba (Qwen) · 1M ctx | Alibaba (Qwen) | Flagship | $2 | $6 | $0.25 | 1M | Oct 1, 2026 |
| Qwen3.7-Plus Alibaba (Qwen) · 1M ctx | Alibaba (Qwen) | Mid | $0.40* | $1.60* | $0.08 | 1M | Oct 1, 2026 |
| Qwen3.8-Flash Alibaba (Qwen) · 1M ctx | Alibaba (Qwen) | Small | $0.15 | $0.47 | $0.016 | 1M | Oct 1, 2026 |
| Qwen3.8-27B Alibaba (Qwen) · 1M ctx | Alibaba (Qwen) | Small | $0.50 | $3 | $0.10 | 1M | Oct 1, 2026 |
| Qwen3.7-Flash Alibaba (Qwen) · 1M ctx | Alibaba (Qwen) | Small | $0.03* | $0.13* | $0.006 | 1M | Oct 1, 2026 |
| GLM-5.3 Zhipu (GLM) · 1M ctx | Zhipu (GLM) | Flagship | $1.40 | $4.40 | $0.26 | 1M | Oct 1, 2026 |
| GLM-5.2 Zhipu (GLM) · 1M ctx | Zhipu (GLM) | Flagship | $1.40 | $4.40 | $0.26 | 1M | Oct 1, 2026 |
| GLM-5.3-Flash Zhipu (GLM) · 1M ctx | Zhipu (GLM) | Small | $0.15 | $0.50 | $0.03 | 1M | Oct 1, 2026 |
| GLM-5.3-FlashX Zhipu (GLM) · 1M ctx | Zhipu (GLM) | Small | $0.37 | $1.25 | $0.075 | 1M | Oct 1, 2026 |
| Kimi K3 Moonshot (Kimi) · 1.05M ctx | Moonshot (Kimi) | Flagship | $3 | $15 | $0.30 | 1.05M | Oct 1, 2026 |
| Kimi K2.6 Moonshot (Kimi) · 262K ctx | Moonshot (Kimi) | Mid | $0.95 | $4 | $0.16 | 262K | Oct 1, 2026 |
| Kimi K2.7 Code Moonshot (Kimi) · 262K ctx | Moonshot (Kimi) | Mid | $0.95 | $4 | $0.19 | 262K | Oct 1, 2026 |
| Kimi K2.7 Code HighSpeed Moonshot (Kimi) · 262K ctx | Moonshot (Kimi) | Mid | $1.90 | $8 | $0.38 | 262K | Oct 1, 2026 |
| MiniMax-M3 MiniMax · 1M ctx | MiniMax | Flagship | $0.30* | $1.20* | $0.06 | 1M | Oct 1, 2026 |
| MiniMax-M2.7 MiniMax · 205K ctx | MiniMax | Mid | $0.30 | $1.20 | $0.06 | 205K | Oct 1, 2026 |
| MiniMax-M2.7-highspeed MiniMax · 205K ctx | MiniMax | Mid | $0.60 | $2.40 | $0.06 | 205K | Oct 1, 2026 |
| Mistral Medium 3.5 Mistral · 256K ctx | Mistral | Flagship | $1.50 | $7.50 | $0.15 | 256K | Oct 1, 2026 |
| Mistral Large 3 Mistral · 256K ctx | Mistral | Mid | $0.50 | $1.50 | $0.05 | 256K | Oct 1, 2026 |
| Mistral Small 4 Mistral · 256K ctx | Mistral | Small | $0.15 | $0.60 | $0.015 | 256K | Oct 1, 2026 |
| Ministral 3 14B Mistral · 256K ctx | Mistral | Small | $0.20 | $0.20 | $0.02 | 256K | Oct 1, 2026 |
| Ministral 3 8B Mistral · 256K ctx | Mistral | Small | $0.15 | $0.15 | $0.015 | 256K | Oct 1, 2026 |
| Ministral 3 3B Mistral · 256K ctx | Mistral | Small | $0.10 | $0.10 | $0.01 | 256K | Oct 1, 2026 |
| Codestral Mistral · 128K ctx | Mistral | Small | $0.30 | $0.90 | $0.03 | 128K | Oct 1, 2026 |
Standard pay-as-you-go rates. * Higher rates apply above a context-length threshold. Click a date to open the official pricing page it was checked against.
Tools & guides
- LLM Price Comparison
Pick two to four models and see input, output and cached rates, discounts and context limits in one view. Use this side-by-side model price check before switching providers.
- Cheapest LLM APIs
The lowest-priced models overall, among flagships, with 1M-token context, with image input and with open weights. Browse the cheapest LLM provider list when budget comes first.
- Token Cost Calculator
Enter a token count or paste text and see what it costs on every model. Handy for pricing a single prompt or a document.
- LLM Cost Calculator
Choose a workload, set requests per month and compare the bill with caching and batch discounts. Start here to estimate your monthly API spend.
How LLM API pricing works
What is a token?
A token is the unit LLM APIs bill by: a word fragment of about four characters in English, so 1,000 words come to roughly 1,300 tokens. Prices are quoted per million tokens.
At GPT-6.1 Sol's $10 per 1M output tokens, generating a 1,000-word article costs about $0.013.
Why output costs more than input
The model reads your prompt in one pass but writes its answer one token at a time, so output takes more compute and costs more.
Claude Opus 5.5 charges $4 for input and $20 for output, 5 times as much. DeepSeek-V4-Pro charges $1.32 and $3.96, 3 times as much. For long answers, compare output prices first.
Cached input pricing
When you resend the same prefix, such as a system prompt or a document, providers can serve it from a cache and bill it at a lower cached input rate.
GPT-6.1 Sol bills cached input at $0.10 instead of $2; Claude Opus 5.5 at $0.20 instead of $4. Some providers also charge for writing to the cache, at $5 per 1M tokens for Claude Opus 5.5.
Batch discounts
If you can wait minutes or hours for results, batch APIs process requests asynchronously at a discount. 26 of the 50 models in the table offer 50% off this way; a few providers offer smaller discounts or none.
Long-context tiered pricing
Some models charge more once a request passes a context-length threshold, and the higher rate usually applies to the whole request. These models are marked with * in the table.
GPT-6.1 Sol goes from $2 / $10 to $4 / $15 above 272K input tokens; Gemini 3.1 Pro Preview to $4 / $18 above 200K; Qwen3.7-Plus to $1.20 / $4.80 above 256K.
FAQ
Prices in our table run from $0.03 to $10 per million input tokens and from $0.10 to $50 per million output tokens. For chat requests with 1,000 input and 300 output tokens, 1,000 requests cost between $0.069 and $25.00, depending on the model.
Providers quote prices per million tokens. A token is a chunk of text, about four characters of English. If a model charges $2 per 1M input tokens, a 2,000-token prompt costs $0.004.
About 1,300 tokens for English text. Code and languages such as Chinese or Japanese use more tokens for the same length.
It depends on your mix. Qwen3.7-Flash has the lowest input price at $0.03 per million tokens, while Ministral 3 3B has the lowest output price at $0.10. Prompt-heavy work and output-heavy work have different winners. See the full rankings.
The input and output columns are standard pay-as-you-go rates. The cached input column is the price for prompt tokens served from cache. Batch discounts, off-peak pricing and long-context surcharges are not included in the table; promotional or peak-hour prices are marked under the model name.
Every price is copied from the provider's official pricing page, and each row links to its source and shows the date it was checked. We recheck the major providers every week. How we check prices.