- Home
- Token Cost Calculator
Token Cost Calculator
Enter how many input and output tokens you need, or paste text to estimate it, and see what they cost on 50 models.
Token counts from pasted text are estimates at about 4 characters per token. Each model's tokenizer counts a little differently, so treat them as a guide.
Token cost on every model
| # | Model | Provider | Total | Input | Output |
|---|---|---|---|---|---|
| 1 | Qwen3.7-Flash Alibaba (Qwen) | Alibaba (Qwen) | $0.00016 | $0.00003 | $0.00013 |
| 2 | Ministral 3 3B Mistral | Mistral | $0.0002 | $0.0001 | $0.0001 |
| 3 | Ministral 3 8B Mistral | Mistral | $0.0003 | $0.00015 | $0.00015 |
| 4 | Ministral 3 14B Mistral | Mistral | $0.0004 | $0.0002 | $0.0002 |
| 5 | GPT-6 Luna OpenAI | OpenAI | $0.0006 | $0.0001 | $0.0005 |
| 6 | Qwen3.8-Flash Alibaba (Qwen) | Alibaba (Qwen) | $0.00062 | $0.00015 | $0.00047 |
| 7 | GLM-5.3-Flash Zhipu (GLM) | Zhipu (GLM) | $0.00065 | $0.00015 | $0.0005 |
| 8 | Mistral Small 4 Mistral | Mistral | $0.00075 | $0.00015 | $0.0006 |
| 9 | Codestral Mistral | Mistral | $0.0012 | $0.0003 | $0.0009 |
| 10 | GPT-5.6 Luna OpenAI | OpenAI | $0.0014 | $0.0002 | $0.0012 |
| 11 | DeepSeek-V4.1-Flash DeepSeek | DeepSeek | $0.0015 | $0.0003 | $0.0012 |
| 12 | MiniMax-M3 MiniMax | MiniMax | $0.0015 | $0.0003 | $0.0012 |
| 13 | MiniMax-M2.7 MiniMax | MiniMax | $0.0015 | $0.0003 | $0.0012 |
| 14 | GLM-5.3-FlashX Zhipu (GLM) | Zhipu (GLM) | $0.0016 | $0.00037 | $0.0013 |
| 15 | Gemini 3.1 Flash-Lite Google | $0.0018 | $0.00025 | $0.0015 |
How token pricing works
LLM APIs bill per token and quote prices per million tokens, with separate rates for the tokens you send (input) and the tokens the model writes back (output).
For example, GPT-6.1 Sol charges $2 per 1M input tokens and $10 per 1M output tokens. A request with 2,000 input tokens and 500 output tokens costs 2,000 × $2 ÷ 1M + 500 × $10 ÷ 1M = $0.009.
Cached input, batch discounts and long-context surcharges change the rate per token. This calculator uses each model's standard rates.
Token cost examples
What common jobs cost in tokens on a low-cost model, a mid-range model and a flagship, at standard rates with no discounts.
| Job | Tokens in / out | DeepSeek-V4.1-Flash | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|---|---|
| Summarize a 10-page report | 6,500 / 500 | $0.0025 | $0.018 | $0.036 |
| Translate a 1,000-word article | 1,300 / 1,300 | $0.0019 | $0.016 | $0.031 |
| Write a 2,000-word blog post | 200 / 2,600 | $0.0032 | $0.026 | $0.053 |
| Ask one question about a 300-page book | 160,000 / 500 | $0.049 | $0.33 | $0.65 |
| Label 1,000 support tickets | 1,000 × 250 / 5 | $0.081 | $0.55 | $1.10 |
The same job can cost 13.4× as much on Claude Opus 5.5 as on DeepSeek-V4.1-Flash. Model choice moves the price far more than trimming a few words from a prompt.
The question about a book costs more than writing a whole blog post, even though the answer is short: on Claude Opus 5.5 it comes to $0.65, against $0.053 for the post. Long inputs are priced by the input rate, long answers by the output rate, so check which one dominates your job before you compare models.
Input tokens vs output tokens
Output tokens cost more than input tokens on almost every model, but by very different amounts. Across the 50 models we track, Ministral 3 14B charges the same rate for both, while Gemini 3.5 Flash-Lite charges 8.3× its input rate.
That gap matters for output-heavy jobs. Grok 4.7 and GPT-6.1 Sol both charge $2 per 1M input tokens, but $6 and $10 per 1M output tokens. For the 2,000-word blog post above, Grok 4.7 costs $0.016 and GPT-6.1 Sol $0.026.
As a rule of thumb: for summarizing, classifying and answering questions about long documents, compare input prices first. For writing, translation and code generation, compare output prices first. The calculator above shows both parts of the cost for every model, so you can see which side drives the total.
How many tokens is my text?
- 1 token ≈ 4 characters ≈ ¾ of an English word.
- A 280-character social post ≈ 70 tokens.
- A 200-word email ≈ 260 tokens.
- 1,000 words ≈ 1,300 tokens; a 10-page report ≈ 6,500 tokens.
- A 90,000-word novel ≈ 120,000 tokens.
- Code, tables and languages such as Chinese or Japanese use more tokens per character.
What counts as a token
Token cost covers more than the text you type. Everything in the request is billed as input, and everything the model generates is billed as output, including parts you never see.
- Instructions and history. The system prompt, tool and function definitions, and every earlier turn of a chat are sent again with each new message. A long conversation gets more expensive with every turn because the whole history is billed as input again.
- Retrieved documents. Passages pulled from a search index or knowledge base are input tokens too, and they often make up most of the prompt.
- Images and files. Images and PDF pages are turned into tokens and billed at the input rate.
- Reasoning. Models that think before they answer bill those reasoning tokens as output, so a short answer can use far more output tokens than its length suggests.
- Formatting. Spaces, punctuation, JSON brackets and markup all count. Asking for plain text instead of verbose JSON can cut output tokens noticeably.
Each model family splits text into tokens its own way, so the same prompt can come to 10–20% more tokens on one model than on another. When two models are close on price, that difference can decide which one is actually cheaper.
Price per 1K tokens vs per 1M tokens
Providers now quote prices per million tokens, but older price lists and some SDKs use per thousand. To convert, divide the per-1M price by 1,000.
GPT-6.1 Sol at $2 per 1M input tokens is $0.002 per 1K input tokens, and $10 per 1M output tokens is $0.01 per 1K. A single output token costs $0.00001.
Prices this small are why a single request usually costs a fraction of a cent. The cost adds up through volume and through long prompts, so price the token counts you actually send, not just the question you type.
FAQ
Anywhere from a few cents to $50. Small models charge well under $1 per million input tokens, while the most expensive flagships charge $10 per million input tokens and $50 per million output tokens. The table above prices your exact token count on every model.
About 1,300 tokens for English prose. One token is roughly four characters or three quarters of a word. Code, numbers and non-English text usually take more tokens per word.
The model generates output one token at a time, which takes far more compute than reading the prompt. Providers usually price output at 3 to 6 times the input rate.
No. Counts from pasted text use about four characters per token. Each model family has its own tokenizer, so the real count can differ by 10–20%, more for code and non-English text. Type an exact count into the input field if you have one.
Yes. The system prompt, tool definitions and any earlier turns of the conversation are sent with every request and billed as input tokens each time. A 2,000-token system prompt on a one-line question still costs 2,000 input tokens.
Yes. Models with image input convert each image into tokens and bill them at the input rate. Depending on the provider and the image size, one image counts as a few hundred to a couple of thousand tokens.
Yes. Models that think before they answer bill those hidden reasoning tokens as output tokens, even though you never see them. For reasoning models, enter more output tokens than the visible reply you expect.
Send a few real requests. Every API response reports the input and output tokens it used, and some providers also offer a token-counting endpoint you can call before sending. Those figures are exact for that model; the estimate here is for quick comparisons.
Usually because the estimate left something out: the system prompt and chat history resent with every message, retrieved documents, hidden reasoning tokens, or retried requests. Compare the token counts in your usage dashboard with the numbers you entered here.
Related tools
- LLM API pricing tableEvery model's input, output and cached rates in one sortable table.
- LLM Price ComparisonPut two to four models side by side: per-token rates, context window, discounts and what the same job costs on each.
- Cheapest LLM APIsRanked lists of the lowest-priced models overall, among flagships, with 1M-token context, with image input and with open weights.
- LLM Cost CalculatorPick a workload and request volume, and compare the bill across models with caching and batch discounts.