- Home
- LLM Cost Calculator
LLM Cost Calculator
Estimate the monthly bill for an LLM workload and see how it changes from one model to the next. Choose a preset or enter your own request size and volume; every model below is priced with the same numbers, cheapest first.
Discounts are off by default. Models without a cached rate or batch discount are billed at their standard price.
Monthly cost by model
| # | Model | Provider | Per month | Per request |
|---|---|---|---|---|
| 1 | Qwen3.7-Flash Alibaba (Qwen) | Alibaba (Qwen) | $0.54 | $0.000054 |
| 2 | Ministral 3 3B Mistral | Mistral | $0.80 | $0.00008 |
| 3 | Ministral 3 8B Mistral | Mistral | $1.20 | $0.00012 |
| 4 | Ministral 3 14B Mistral | Mistral | $1.60 | $0.00016 |
| 5 | GPT-6 Luna OpenAI | OpenAI | $2.00 | $0.0002 |
| 6 | Qwen3.8-Flash Alibaba (Qwen) | Alibaba (Qwen) | $2.16 | $0.00022 |
| 7 | GLM-5.3-Flash Zhipu (GLM) | Zhipu (GLM) | $2.25 | $0.00022 |
| 8 | Mistral Small 4 Mistral | Mistral | $2.55 | $0.00026 |
| 9 | Codestral Mistral | Mistral | $4.20 | $0.00042 |
| 10 | GPT-5.6 Luna OpenAI | OpenAI | $4.60 | $0.00046 |
| 11 | DeepSeek-V4.1-Flash DeepSeek | DeepSeek | $5.10 | $0.00051 |
| 12 | MiniMax-M3 MiniMax | MiniMax | $5.10 | $0.00051 |
| 13 | MiniMax-M2.7 MiniMax | MiniMax | $5.10 | $0.00051 |
| 14 | GLM-5.3-FlashX Zhipu (GLM) | Zhipu (GLM) | $5.60 | $0.00056 |
| 15 | Gemini 3.1 Flash-Lite Google | $5.75 | $0.00057 |
How to estimate your monthly LLM cost
Request size. Count the input tokens of a typical request, including the system prompt, chat history and any retrieved context, plus the tokens in a typical reply. The presets are starting points: a chatbot turn is small, a coding agent step carries a lot of context.
Volume. Multiply by requests per month. If you only know daily traffic, multiply by 30.
Discounts. Prompt caching bills repeated input at a fraction of the normal rate, and batch processing trades latency for a lower price. Both are off by default so the estimate never looks cheaper than your first invoice.
Monthly cost examples
Four common workloads, priced on a small model (Gemini 3.1 Flash-Lite), a mid-range model (Claude Sonnet 5.5) and a flagship (GPT-6 Astra). Each uses the matching preset in the calculator above.
Customer support chatbot: 50,000 requests a month
About 1,700 short exchanges a day costs $28.75 a month on Gemini 3.1 Flash-Lite, $200 on Claude Sonnet 5.5 and $1,000 on GPT-6 Astra. That is a 35× spread for the same traffic, so test whether a small model answers well enough before paying for a larger one.
Internal RAG assistant: 20,000 requests a month
Each question carries retrieved passages, so input dominates. On Claude Sonnet 5.5 the bill is $260 a month, and 62% of it is input. Retrieving fewer, better passages is the most direct way to cut it. Caching helps only with the fixed part of the prompt, such as instructions, because the passages change with every question. The same workload costs $35.00 on Gemini 3.1 Flash-Lite and $1,300 on GPT-6 Astra.
Coding agent: 5,000 steps a month
An agent resends a long, mostly unchanged context on every step, which is what prompt caching is for. Without caching, Claude Sonnet 5.5 costs $400 a month and GPT-6 Astra $2,000. With 80% of input served from the cache, they drop to $184 and $920. Turn on caching in the calculator and set the cached share to match your agent.
Overnight summaries: 100,000 documents a month
Nobody waits on these results, so batch processing is an easy saving. At standard rates the job costs $290 a month on Gemini 3.1 Flash-Lite and $2,200 on Claude Sonnet 5.5. Through the Batch API it comes to $145 and $1,100.
Cost per user
If you charge for your product, divide the monthly bill by active users to see your margin. The chatbot example is 1,000 users sending 50 messages each a month, so every user costs $0.029 on Gemini 3.1 Flash-Lite, $0.20 on Claude Sonnet 5.5 and $1.00 on GPT-6 Astra. Your heaviest users can send many times the average, so check the cost of the top 10% of users as well before setting a price or a usage limit.
What the estimate leaves out
The calculator prices input and output tokens at list rates. A few charges sit outside that, so leave some headroom in your budget and compare the estimate with your first invoice.
- Reasoning tokens. Models that think before they answer bill that thinking as output. For reasoning models, set output tokens per request well above the length of the visible reply.
- Cache writes. Some providers charge a premium the first time a prompt prefix is written to the cache. The calculator applies only the cheaper cached read rate.
- Retries. Timeouts, errors you retry and answers your users regenerate are all billed as new requests.
- Tools and storage. Built-in web search, code execution and file storage are usually priced per call or per day, separately from tokens.
- Processing tiers. Priority or fast modes and regional data-residency endpoints cost more on some providers.
- Tax. Prices are in US dollars before tax.
How to lower your LLM bill
- Route by difficulty. Send simple requests such as classification, extraction and short replies to a small model, and only the hard ones to a flagship. Prices between the two differ by an order of magnitude or more.
- Cap the reply. Set a maximum output length and ask for concise answers. Output is the more expensive side on almost every model.
- Trim the context. Summarize long chat history, retrieve fewer and better passages, and drop tool definitions a request does not need.
- Keep the prefix stable. Put instructions and reference documents at the start of the prompt and leave them unchanged between requests, so they are served from the cache.
- Batch what can wait. Evaluations, backfills and nightly reports rarely need an answer within seconds.
- Set a spend limit. Most provider consoles let you cap spend and send alerts before a runaway loop turns into a large invoice.
FAQ
It depends on three numbers: tokens in, tokens out and how many requests you send. A chatbot sending 10,000 short requests a month can cost well under $10 on a small model and several hundred dollars on a flagship. Enter your own numbers above to see the range.
Add up everything you send: system prompt, conversation history, retrieved documents and the user message. Then estimate the reply length. Most provider dashboards show average input and output tokens per request once you have some traffic.
Cached input usually costs 2–25% of the normal input rate, depending on the provider. If most of each request is a repeated system prompt or document, turning on caching in the calculator shows the saving directly.
If results can wait minutes or hours, yes. Most providers take 50% off batch requests; a few offer smaller discounts or none. The batch switch applies each model’s own discount.
Chatbot: 500 input and 300 output tokens per request; RAG Q&A: 4,000 input and 500 output tokens per request; Coding agent: 30,000 input and 2,000 output tokens per request; Summarization: 8,000 input and 600 output tokens per request. They are typical starting points. Once you have real traffic, replace them with the averages from your provider’s usage dashboard.
It applies each provider’s published list prices to the numbers you enter, so the estimate is as good as your request size and volume. Real bills differ when traffic changes from month to month, when reasoning tokens or tool calls add usage, or when you have a negotiated discount.
The bill grows in line with requests, but request size tends to grow too as you add features: longer system prompts, more retrieved context, more agent steps. Run the estimate with the request size and volume you expect in three months, not only today’s.
Related tools
- LLM API pricing tableEvery model's input, output and cached rates in one sortable table.
- LLM Price ComparisonPut two to four models side by side: per-token rates, context window, discounts and what the same job costs on each.
- Cheapest LLM APIsRanked lists of the lowest-priced models overall, among flagships, with 1M-token context, with image input and with open weights.
- Token Cost CalculatorEnter a token count and see what that many input and output tokens cost on every model.