Skip to main content

Cheapest LLM APIs (October 2026)

The cheapest LLM API right now is Qwen3.7-Flash at $0.055 per 1M tokens blended ($0.03 input, $0.13 output). Below are the five lowest-priced models in each group, from 50 models we track.

Rankings use a blended price that assumes 3 input tokens for every output token: (3 × input + output) ÷ 4. If your ratio is different, put your own numbers into the LLM Cost Calculator.

Cheapest overall

Every model available in the API, regardless of size.

#ModelProviderBlended $/1MInput $/1MOutput $/1MContext
1Qwen3.7-Flash
Alibaba (Qwen) · $0.03 in / $0.13 out
Alibaba (Qwen)$0.055$0.03$0.131M
2Ministral 3 3B
Mistral · $0.10 in / $0.10 out
Mistral$0.10$0.10$0.10256K
3Ministral 3 8B
Mistral · $0.15 in / $0.15 out
Mistral$0.15$0.15$0.15256K
4GPT-6 Luna
OpenAI · $0.10 in / $0.50 out
OpenAI$0.20$0.10$0.501.05M
5Ministral 3 14B
Mistral · $0.20 in / $0.20 out
Mistral$0.20$0.20$0.20256K

Cheapest flagship models

Each provider's top tier, by the provider's own naming.

#ModelProviderBlended $/1MInput $/1MOutput $/1MContext
1MiniMax-M3
MiniMax · $0.30 in / $1.20 out
MiniMax$0.525$0.30$1.201M
2DeepSeek-V4-Pro
DeepSeek · $1.32 in / $3.96 out
DeepSeek$1.98$1.32$3.961M
3GLM-5.3
Zhipu (GLM) · $1.40 in / $4.40 out
Zhipu (GLM)$2.15$1.40$4.401M
4GLM-5.2
Zhipu (GLM) · $1.40 in / $4.40 out
Zhipu (GLM)$2.15$1.40$4.401M
5Qwen3.8-Max
Alibaba (Qwen) · $2 in / $6 out
Alibaba (Qwen)$3$2$61M

Cheapest with a 1M-token context window

Models that accept roughly one million tokens or more.

#ModelProviderBlended $/1MInput $/1MOutput $/1MContext
1Qwen3.7-Flash
Alibaba (Qwen) · $0.03 in / $0.13 out
Alibaba (Qwen)$0.055$0.03$0.131M
2GPT-6 Luna
OpenAI · $0.10 in / $0.50 out
OpenAI$0.20$0.10$0.501.05M
3Qwen3.8-Flash
Alibaba (Qwen) · $0.15 in / $0.47 out
Alibaba (Qwen)$0.23$0.15$0.471M
4GLM-5.3-Flash
Zhipu (GLM) · $0.15 in / $0.50 out
Zhipu (GLM)$0.2375$0.15$0.501M
5GPT-5.6 Luna
OpenAI · $0.20 in / $1.20 out
OpenAI$0.45$0.20$1.201.05M

Cheapest with image input

Models that accept images in the prompt.

#ModelProviderBlended $/1MInput $/1MOutput $/1MContext
1Qwen3.7-Flash
Alibaba (Qwen) · $0.03 in / $0.13 out
Alibaba (Qwen)$0.055$0.03$0.131M
2Ministral 3 3B
Mistral · $0.10 in / $0.10 out
Mistral$0.10$0.10$0.10256K
3Ministral 3 8B
Mistral · $0.15 in / $0.15 out
Mistral$0.15$0.15$0.15256K
4GPT-6 Luna
OpenAI · $0.10 in / $0.50 out
OpenAI$0.20$0.10$0.501.05M
5Ministral 3 14B
Mistral · $0.20 in / $0.20 out
Mistral$0.20$0.20$0.20256K

Cheapest open-weight models

Models whose weights are published, priced on the provider’s own API.

#ModelProviderBlended $/1MInput $/1MOutput $/1MContext
1Ministral 3 3B
Mistral · $0.10 in / $0.10 out
Mistral$0.10$0.10$0.10256K
2Ministral 3 8B
Mistral · $0.15 in / $0.15 out
Mistral$0.15$0.15$0.15256K
3Ministral 3 14B
Mistral · $0.20 in / $0.20 out
Mistral$0.20$0.20$0.20256K
4Qwen3.8-Flash
Alibaba (Qwen) · $0.15 in / $0.47 out
Alibaba (Qwen)$0.23$0.15$0.471M
5GLM-5.3-Flash
Zhipu (GLM) · $0.15 in / $0.50 out
Zhipu (GLM)$0.2375$0.15$0.501M

Input-heavy vs output-heavy work

The cheapest model depends on what you send and what you get back. Lowest input price: Qwen3.7-Flash at $0.03 per 1M tokens, which suits long prompts, retrieval and classification. Lowest output price: Ministral 3 3B at $0.10 per 1M tokens, which matters more for long answers and generated code.

Prices are each provider's standard pay-as-you-go rates. Cached input, batch discounts and off-peak pricing can lower them further; long-context surcharges can raise them.

FAQ

Related tools