- Home
- LLM Price Comparison
LLM Price Comparison
Put up to four models side by side and see per-token rates, discounts, context limits and what the same workload costs on each. It starts with four flagship models; swap in any of the 50 models we track.
| GPT-6 Astra | Claude Opus 5.5 | Gemini 3.1 Pro Preview | DeepSeek-V4-Pro | |
|---|---|---|---|---|
| Your monthly cost | $450 | $180 | $100 | $46.20 |
| Input $/1M | $10 | $4 | $2 | $1.32 |
| Output $/1M | $50 | $20 | $12 | $3.96 |
| Context window | 1.05M | 1M | 1.05M | 1M |
| vs cheapest | 9.7× | 3.9× | 2.2× | Cheapest |
| Provider | OpenAI | Anthropic | DeepSeek | |
| Tier | Flagship | Flagship | Flagship | Flagship |
| Cached input $/1M | $1 | $0.20 | $0.20 | $0.044 |
| Batch discount | 50% off | 50% off | 50% off | — |
| Long-context pricing | >272K: $20 / $75 | Flat | >200K: $4 / $18 | Flat |
| Max output | 128K | 128K | 66K | 384K |
| Image input | Yes | Yes | Yes | No |
| Open weights | No | No | No | Yes |
| Released | Sep 3, 2026 | Sep 22, 2026 | Feb 19, 2026 | Aug 13, 2026 |
Popular comparisons
Claude Opus 5.5 vs GPT-6 Astra
Claude Opus 5.5: $4 / $20 · GPT-6 Astra: $10 / $50 per 1M tokens. Claude Opus 5.5 is 2.5× cheaper on a 3:1 input/output mix.
Claude Sonnet 5.5 vs GPT-6.1 Sol
Claude Sonnet 5.5: $2 / $10 · GPT-6.1 Sol: $2 / $10 per 1M tokens. About the same blended price.
Gemini 3.8 Flash vs GPT-6 Luna
Gemini 3.8 Flash: $0.75 / $3.75 · GPT-6 Luna: $0.10 / $0.50 per 1M tokens. GPT-6 Luna is 7.5× cheaper on a 3:1 input/output mix.
DeepSeek-V4-Pro vs Qwen3.8-Max
DeepSeek-V4-Pro: $1.32 / $3.96 · Qwen3.8-Max: $2 / $6 per 1M tokens. DeepSeek-V4-Pro is 1.5× cheaper on a 3:1 input/output mix.
Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite
Claude Haiku 4.5: $1 / $5 · Gemini 3.5 Flash-Lite: $0.30 / $2.50 per 1M tokens. Gemini 3.5 Flash-Lite is 2.4× cheaper on a 3:1 input/output mix.
Grok 4.7 vs Gemini 3.1 Pro Preview
Grok 4.7: $2 / $6 · Gemini 3.1 Pro Preview: $2 / $12 per 1M tokens. Grok 4.7 is 1.5× cheaper on a 3:1 input/output mix.
Reading an LLM price comparison
Input and output are priced separately. Output usually costs 3–6 times more than input, so a model that looks cheap on input can still be expensive for long answers.
Caching changes the picture. If you resend the same system prompt or documents, the cached input rate matters more than the list price.
Watch the context threshold. Some models switch the whole request to a higher rate above 200K or 272K input tokens.
FAQ
Compare input and output rates separately, then price a realistic request. A model with cheap input but expensive output can cost more for chat or code generation, where replies are long. The comparison table above prices the same request on every selected model.
Several providers charge higher rates once a request passes a context threshold, such as 200K or 272K input tokens, and apply that rate to the whole request. The long-context row shows where this applies.
No. They are standard pay-as-you-go rates. Cached input and batch discounts are listed separately because how much you save depends on your traffic.
Related tools
- LLM API pricing tableEvery model's input, output and cached rates in one sortable table.
- Cheapest LLM APIsRanked lists of the lowest-priced models overall, among flagships, with 1M-token context, with image input and with open weights.
- Token Cost CalculatorEnter a token count and see what that many input and output tokens cost on every model.
- LLM Cost CalculatorPick a workload and request volume, and compare the bill across models with caching and batch discounts.