Skip to main content

LLM Pricing: API Prices for 50+ Models

50 models from 10 providers · prices checked Oct 2, 2026

NewGemini 4 Argon cost calculator

Enter token counts to add a cost column to the table. Open the full calculator →

50 models · USD per 1M tokens
ModelProviderTierChecked
GPT-6 Astra
OpenAI · 1.05M ctx
OpenAIFlagship$10*$50*$11.05MOct 1, 2026
GPT-5.6 Sol
OpenAI · 1.05M ctx · Promo price
Promotional price, available at least through Nov 21, 2026.
OpenAIFlagship$4*$20*$0.401.05MOct 1, 2026
GPT-5.5
OpenAI · 1.05M ctx
OpenAIFlagship$5*$30*$0.501.05MOct 1, 2026
GPT-6.1 Sol
OpenAI · 1.05M ctx
OpenAIMid$2*$10*$0.101.05MOct 1, 2026
GPT-6 Sol
OpenAI · 1.05M ctx
OpenAIMid$2*$10*$0.201.05MOct 1, 2026
GPT-5.6 Terra
OpenAI · 1.05M ctx
OpenAIMid$2*$12*$0.201.05MOct 1, 2026
GPT-6 Luna
OpenAI · 1.05M ctx
OpenAISmall$0.10*$0.50*$0.011.05MOct 1, 2026
GPT-5.6 Luna
OpenAI · 1.05M ctx
OpenAISmall$0.20*$1.20*$0.021.05MOct 1, 2026
Claude Opus 5.5
Anthropic · 1M ctx
AnthropicFlagship$4$20$0.201MOct 1, 2026
Claude Fable 5.1
Anthropic · 1M ctx
AnthropicFlagship$10$50$0.251MOct 1, 2026
Claude Sonnet 5.5
Anthropic · 1M ctx
AnthropicMid$2$10$0.201MOct 1, 2026
Claude Haiku 4.5
Anthropic · 200K ctx
AnthropicSmall$1$5$0.10200KOct 1, 2026
Gemini 3.1 Pro Preview
Google · 1.05M ctx
GoogleFlagship$2*$12*$0.201.05MOct 1, 2026
Gemini 4 Argon
Google · — ctx · Announced
Announced, not yet in the API. Introductory price; then $4 / $20.
GoogleFlagship$2$10$0.10—Oct 2, 2026
Gemini 3.8 Flash
Google · 1.05M ctx · Promo price
Promotional price until Dec 31, 2026; then $1.50 / $7.50.
GoogleMid$0.75$3.75$0.0751.05MOct 1, 2026
Gemini 3.7 Flash
Google · 1.05M ctx · Promo price
Promotional price until Dec 31, 2026; then $1.50 / $7.50.
GoogleMid$0.75$3.75$0.0751.05MOct 1, 2026
Gemini 3.6 Flash
Google · 1.05M ctx · Promo price
Promotional price until Dec 31, 2026; then $1.50 / $7.50.
GoogleMid$0.75$3.75$0.0751.05MOct 1, 2026
Gemini 3.5 Flash-Lite
Google · 1.05M ctx
GoogleSmall$0.30$2.50$0.031.05MOct 1, 2026
Gemini 3.1 Flash-Lite
Google · 1.05M ctx
GoogleSmall$0.25$1.50$0.0251.05MOct 1, 2026
DeepSeek-V4-Pro
DeepSeek · 1M ctx · Peak-hour price
Peak-hour price; 50% off outside peak hours.
DeepSeekFlagship$1.32$3.96$0.0441MOct 1, 2026
DeepSeek-V4.1-Flash
DeepSeek · 1M ctx · Peak-hour price
Peak-hour price; 50% off outside peak hours.
DeepSeekSmall$0.30$1.20$0.0061MOct 1, 2026
Grok 4.7
xAI · 500K ctx
xAIFlagship$2*$6*$0.50500KOct 1, 2026
Grok 4.6
xAI · 500K ctx
xAIFlagship$2*$6*$0.50500KOct 1, 2026
Grok 4.5
xAI · 500K ctx
xAIFlagship$2*$6*$0.30500KOct 1, 2026
Grok 4.3
xAI · 1M ctx
xAIMid$1.25*$2.50*$0.201MOct 1, 2026
Grok 4.20
xAI · 1M ctx
xAIMid$1.25*$2.50*$0.201MOct 1, 2026
Grok Build 0.1Preview
xAI · 256K ctx
xAISmall$1*$2*$0.20256KOct 1, 2026
Qwen3.8-Max
Alibaba (Qwen) · 1M ctx
Alibaba (Qwen)Flagship$2$6$0.251MOct 1, 2026
Qwen3.7-Plus
Alibaba (Qwen) · 1M ctx
Alibaba (Qwen)Mid$0.40*$1.60*$0.081MOct 1, 2026
Qwen3.8-Flash
Alibaba (Qwen) · 1M ctx
Alibaba (Qwen)Small$0.15$0.47$0.0161MOct 1, 2026
Qwen3.8-27B
Alibaba (Qwen) · 1M ctx
Alibaba (Qwen)Small$0.50$3$0.101MOct 1, 2026
Qwen3.7-Flash
Alibaba (Qwen) · 1M ctx
Alibaba (Qwen)Small$0.03*$0.13*$0.0061MOct 1, 2026
GLM-5.3
Zhipu (GLM) · 1M ctx
Zhipu (GLM)Flagship$1.40$4.40$0.261MOct 1, 2026
GLM-5.2
Zhipu (GLM) · 1M ctx
Zhipu (GLM)Flagship$1.40$4.40$0.261MOct 1, 2026
GLM-5.3-Flash
Zhipu (GLM) · 1M ctx
Zhipu (GLM)Small$0.15$0.50$0.031MOct 1, 2026
GLM-5.3-FlashX
Zhipu (GLM) · 1M ctx
Zhipu (GLM)Small$0.37$1.25$0.0751MOct 1, 2026
Kimi K3
Moonshot (Kimi) · 1.05M ctx
Moonshot (Kimi)Flagship$3$15$0.301.05MOct 1, 2026
Kimi K2.6
Moonshot (Kimi) · 262K ctx
Moonshot (Kimi)Mid$0.95$4$0.16262KOct 1, 2026
Kimi K2.7 Code
Moonshot (Kimi) · 262K ctx
Moonshot (Kimi)Mid$0.95$4$0.19262KOct 1, 2026
Kimi K2.7 Code HighSpeed
Moonshot (Kimi) · 262K ctx
Moonshot (Kimi)Mid$1.90$8$0.38262KOct 1, 2026
MiniMax-M3
MiniMax · 1M ctx
MiniMaxFlagship$0.30*$1.20*$0.061MOct 1, 2026
MiniMax-M2.7
MiniMax · 205K ctx
MiniMaxMid$0.30$1.20$0.06205KOct 1, 2026
MiniMax-M2.7-highspeed
MiniMax · 205K ctx
MiniMaxMid$0.60$2.40$0.06205KOct 1, 2026
Mistral Medium 3.5
Mistral · 256K ctx
MistralFlagship$1.50$7.50$0.15256KOct 1, 2026
Mistral Large 3
Mistral · 256K ctx
MistralMid$0.50$1.50$0.05256KOct 1, 2026
Mistral Small 4
Mistral · 256K ctx
MistralSmall$0.15$0.60$0.015256KOct 1, 2026
Ministral 3 14B
Mistral · 256K ctx
MistralSmall$0.20$0.20$0.02256KOct 1, 2026
Ministral 3 8B
Mistral · 256K ctx
MistralSmall$0.15$0.15$0.015256KOct 1, 2026
Ministral 3 3B
Mistral · 256K ctx
MistralSmall$0.10$0.10$0.01256KOct 1, 2026
Codestral
Mistral · 128K ctx
MistralSmall$0.30$0.90$0.03128KOct 1, 2026

Standard pay-as-you-go rates. * Higher rates apply above a context-length threshold. Click a date to open the official pricing page it was checked against.

Tools & guides

How LLM API pricing works

What is a token?

A token is the unit LLM APIs bill by: a word fragment of about four characters in English, so 1,000 words come to roughly 1,300 tokens. Prices are quoted per million tokens.

At GPT-6.1 Sol's $10 per 1M output tokens, generating a 1,000-word article costs about $0.013.

Why output costs more than input

The model reads your prompt in one pass but writes its answer one token at a time, so output takes more compute and costs more.

Claude Opus 5.5 charges $4 for input and $20 for output, 5 times as much. DeepSeek-V4-Pro charges $1.32 and $3.96, 3 times as much. For long answers, compare output prices first.

Cached input pricing

When you resend the same prefix, such as a system prompt or a document, providers can serve it from a cache and bill it at a lower cached input rate.

GPT-6.1 Sol bills cached input at $0.10 instead of $2; Claude Opus 5.5 at $0.20 instead of $4. Some providers also charge for writing to the cache, at $5 per 1M tokens for Claude Opus 5.5.

Batch discounts

If you can wait minutes or hours for results, batch APIs process requests asynchronously at a discount. 26 of the 50 models in the table offer 50% off this way; a few providers offer smaller discounts or none.

Long-context tiered pricing

Some models charge more once a request passes a context-length threshold, and the higher rate usually applies to the whole request. These models are marked with * in the table.

GPT-6.1 Sol goes from $2 / $10 to $4 / $15 above 272K input tokens; Gemini 3.1 Pro Preview to $4 / $18 above 200K; Qwen3.7-Plus to $1.20 / $4.80 above 256K.

FAQ