- Home
- Cheapest LLM APIs
Cheapest LLM APIs (October 2026)
The cheapest LLM API right now is Qwen3.7-Flash at $0.055 per 1M tokens blended ($0.03 input, $0.13 output). Below are the five lowest-priced models in each group, from 50 models we track.
Rankings use a blended price that assumes 3 input tokens for every output token: (3 × input + output) ÷ 4. If your ratio is different, put your own numbers into the LLM Cost Calculator.
Cheapest overall
Every model available in the API, regardless of size.
| # | Model | Provider | Blended $/1M | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Qwen3.7-Flash Alibaba (Qwen) · $0.03 in / $0.13 out | Alibaba (Qwen) | $0.055 | $0.03 | $0.13 | 1M |
| 2 | Ministral 3 3B Mistral · $0.10 in / $0.10 out | Mistral | $0.10 | $0.10 | $0.10 | 256K |
| 3 | Ministral 3 8B Mistral · $0.15 in / $0.15 out | Mistral | $0.15 | $0.15 | $0.15 | 256K |
| 4 | GPT-6 Luna OpenAI · $0.10 in / $0.50 out | OpenAI | $0.20 | $0.10 | $0.50 | 1.05M |
| 5 | Ministral 3 14B Mistral · $0.20 in / $0.20 out | Mistral | $0.20 | $0.20 | $0.20 | 256K |
Cheapest flagship models
Each provider's top tier, by the provider's own naming.
| # | Model | Provider | Blended $/1M | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | MiniMax-M3 MiniMax · $0.30 in / $1.20 out | MiniMax | $0.525 | $0.30 | $1.20 | 1M |
| 2 | DeepSeek-V4-Pro DeepSeek · $1.32 in / $3.96 out | DeepSeek | $1.98 | $1.32 | $3.96 | 1M |
| 3 | GLM-5.3 Zhipu (GLM) · $1.40 in / $4.40 out | Zhipu (GLM) | $2.15 | $1.40 | $4.40 | 1M |
| 4 | GLM-5.2 Zhipu (GLM) · $1.40 in / $4.40 out | Zhipu (GLM) | $2.15 | $1.40 | $4.40 | 1M |
| 5 | Qwen3.8-Max Alibaba (Qwen) · $2 in / $6 out | Alibaba (Qwen) | $3 | $2 | $6 | 1M |
Cheapest with a 1M-token context window
Models that accept roughly one million tokens or more.
| # | Model | Provider | Blended $/1M | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Qwen3.7-Flash Alibaba (Qwen) · $0.03 in / $0.13 out | Alibaba (Qwen) | $0.055 | $0.03 | $0.13 | 1M |
| 2 | GPT-6 Luna OpenAI · $0.10 in / $0.50 out | OpenAI | $0.20 | $0.10 | $0.50 | 1.05M |
| 3 | Qwen3.8-Flash Alibaba (Qwen) · $0.15 in / $0.47 out | Alibaba (Qwen) | $0.23 | $0.15 | $0.47 | 1M |
| 4 | GLM-5.3-Flash Zhipu (GLM) · $0.15 in / $0.50 out | Zhipu (GLM) | $0.2375 | $0.15 | $0.50 | 1M |
| 5 | GPT-5.6 Luna OpenAI · $0.20 in / $1.20 out | OpenAI | $0.45 | $0.20 | $1.20 | 1.05M |
Cheapest with image input
Models that accept images in the prompt.
| # | Model | Provider | Blended $/1M | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Qwen3.7-Flash Alibaba (Qwen) · $0.03 in / $0.13 out | Alibaba (Qwen) | $0.055 | $0.03 | $0.13 | 1M |
| 2 | Ministral 3 3B Mistral · $0.10 in / $0.10 out | Mistral | $0.10 | $0.10 | $0.10 | 256K |
| 3 | Ministral 3 8B Mistral · $0.15 in / $0.15 out | Mistral | $0.15 | $0.15 | $0.15 | 256K |
| 4 | GPT-6 Luna OpenAI · $0.10 in / $0.50 out | OpenAI | $0.20 | $0.10 | $0.50 | 1.05M |
| 5 | Ministral 3 14B Mistral · $0.20 in / $0.20 out | Mistral | $0.20 | $0.20 | $0.20 | 256K |
Cheapest open-weight models
Models whose weights are published, priced on the provider’s own API.
| # | Model | Provider | Blended $/1M | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Ministral 3 3B Mistral · $0.10 in / $0.10 out | Mistral | $0.10 | $0.10 | $0.10 | 256K |
| 2 | Ministral 3 8B Mistral · $0.15 in / $0.15 out | Mistral | $0.15 | $0.15 | $0.15 | 256K |
| 3 | Ministral 3 14B Mistral · $0.20 in / $0.20 out | Mistral | $0.20 | $0.20 | $0.20 | 256K |
| 4 | Qwen3.8-Flash Alibaba (Qwen) · $0.15 in / $0.47 out | Alibaba (Qwen) | $0.23 | $0.15 | $0.47 | 1M |
| 5 | GLM-5.3-Flash Zhipu (GLM) · $0.15 in / $0.50 out | Zhipu (GLM) | $0.2375 | $0.15 | $0.50 | 1M |
Input-heavy vs output-heavy work
The cheapest model depends on what you send and what you get back. Lowest input price: Qwen3.7-Flash at $0.03 per 1M tokens, which suits long prompts, retrieval and classification. Lowest output price: Ministral 3 3B at $0.10 per 1M tokens, which matters more for long answers and generated code.
Prices are each provider's standard pay-as-you-go rates. Cached input, batch discounts and off-peak pricing can lower them further; long-context surcharges can raise them.
FAQ
Qwen3.7-Flash from Alibaba (Qwen) has the lowest blended price in our table at $0.055 per 1M tokens ($0.03 input, $0.13 output). If your requests are mostly output, check the output column too: the ranking changes.
MiniMax-M3 is the lowest-priced flagship model at $0.30 per 1M input tokens and $1.20 per 1M output tokens.
This page ranks price only, not quality. Small models handle classification, extraction and short answers well; harder reasoning and long code changes often need a larger model. Run a sample of your own prompts through two or three candidates before switching.
Some providers offer free tiers with rate limits, for example the Google Gemini API. Free tiers usually come with lower limits and different data-use terms than paid usage, so read the terms before sending production data.
Related tools
- LLM API pricing tableEvery model's input, output and cached rates in one sortable table.
- LLM Price ComparisonPut two to four models side by side: per-token rates, context window, discounts and what the same job costs on each.
- Token Cost CalculatorEnter a token count and see what that many input and output tokens cost on every model.
- LLM Cost CalculatorPick a workload and request volume, and compare the bill across models with caching and batch discounts.