Skip to main content

Token Cost Calculator

Enter how many input and output tokens you need, or paste text to estimate it, and see what they cost on 50 models.

Token counts from pasted text are estimates at about 4 characters per token. Each model's tokenizer counts a little differently, so treat them as a guide.

Token cost on every model

#ModelProviderTotalInputOutput
1Qwen3.7-Flash
Alibaba (Qwen)
Alibaba (Qwen)$0.00016$0.00003$0.00013
2Ministral 3 3B
Mistral
Mistral$0.0002$0.0001$0.0001
3Ministral 3 8B
Mistral
Mistral$0.0003$0.00015$0.00015
4Ministral 3 14B
Mistral
Mistral$0.0004$0.0002$0.0002
5GPT-6 Luna
OpenAI
OpenAI$0.0006$0.0001$0.0005
6Qwen3.8-Flash
Alibaba (Qwen)
Alibaba (Qwen)$0.00062$0.00015$0.00047
7GLM-5.3-Flash
Zhipu (GLM)
Zhipu (GLM)$0.00065$0.00015$0.0005
8Mistral Small 4
Mistral
Mistral$0.00075$0.00015$0.0006
9Codestral
Mistral
Mistral$0.0012$0.0003$0.0009
10GPT-5.6 Luna
OpenAI
OpenAI$0.0014$0.0002$0.0012
11DeepSeek-V4.1-Flash
DeepSeek
DeepSeek$0.0015$0.0003$0.0012
12MiniMax-M3
MiniMax
MiniMax$0.0015$0.0003$0.0012
13MiniMax-M2.7
MiniMax
MiniMax$0.0015$0.0003$0.0012
14GLM-5.3-FlashX
Zhipu (GLM)
Zhipu (GLM)$0.0016$0.00037$0.0013
15Gemini 3.1 Flash-Lite
Google
Google$0.0018$0.00025$0.0015

How token pricing works

LLM APIs bill per token and quote prices per million tokens, with separate rates for the tokens you send (input) and the tokens the model writes back (output).

For example, GPT-6.1 Sol charges $2 per 1M input tokens and $10 per 1M output tokens. A request with 2,000 input tokens and 500 output tokens costs 2,000 × $2 ÷ 1M + 500 × $10 ÷ 1M = $0.009.

Cached input, batch discounts and long-context surcharges change the rate per token. This calculator uses each model's standard rates.

Token cost examples

What common jobs cost in tokens on a low-cost model, a mid-range model and a flagship, at standard rates with no discounts.

JobTokens in / outDeepSeek-V4.1-FlashGPT-6.1 SolClaude Opus 5.5
Summarize a 10-page report6,500 / 500$0.0025$0.018$0.036
Translate a 1,000-word article1,300 / 1,300$0.0019$0.016$0.031
Write a 2,000-word blog post200 / 2,600$0.0032$0.026$0.053
Ask one question about a 300-page book160,000 / 500$0.049$0.33$0.65
Label 1,000 support tickets1,000 × 250 / 5$0.081$0.55$1.10

The same job can cost 13.4× as much on Claude Opus 5.5 as on DeepSeek-V4.1-Flash. Model choice moves the price far more than trimming a few words from a prompt.

The question about a book costs more than writing a whole blog post, even though the answer is short: on Claude Opus 5.5 it comes to $0.65, against $0.053 for the post. Long inputs are priced by the input rate, long answers by the output rate, so check which one dominates your job before you compare models.

Input tokens vs output tokens

Output tokens cost more than input tokens on almost every model, but by very different amounts. Across the 50 models we track, Ministral 3 14B charges the same rate for both, while Gemini 3.5 Flash-Lite charges 8.3× its input rate.

That gap matters for output-heavy jobs. Grok 4.7 and GPT-6.1 Sol both charge $2 per 1M input tokens, but $6 and $10 per 1M output tokens. For the 2,000-word blog post above, Grok 4.7 costs $0.016 and GPT-6.1 Sol $0.026.

As a rule of thumb: for summarizing, classifying and answering questions about long documents, compare input prices first. For writing, translation and code generation, compare output prices first. The calculator above shows both parts of the cost for every model, so you can see which side drives the total.

How many tokens is my text?

  • 1 token ≈ 4 characters ≈ ¾ of an English word.
  • A 280-character social post ≈ 70 tokens.
  • A 200-word email ≈ 260 tokens.
  • 1,000 words ≈ 1,300 tokens; a 10-page report ≈ 6,500 tokens.
  • A 90,000-word novel ≈ 120,000 tokens.
  • Code, tables and languages such as Chinese or Japanese use more tokens per character.

What counts as a token

Token cost covers more than the text you type. Everything in the request is billed as input, and everything the model generates is billed as output, including parts you never see.

  • Instructions and history. The system prompt, tool and function definitions, and every earlier turn of a chat are sent again with each new message. A long conversation gets more expensive with every turn because the whole history is billed as input again.
  • Retrieved documents. Passages pulled from a search index or knowledge base are input tokens too, and they often make up most of the prompt.
  • Images and files. Images and PDF pages are turned into tokens and billed at the input rate.
  • Reasoning. Models that think before they answer bill those reasoning tokens as output, so a short answer can use far more output tokens than its length suggests.
  • Formatting. Spaces, punctuation, JSON brackets and markup all count. Asking for plain text instead of verbose JSON can cut output tokens noticeably.

Each model family splits text into tokens its own way, so the same prompt can come to 10–20% more tokens on one model than on another. When two models are close on price, that difference can decide which one is actually cheaper.

Price per 1K tokens vs per 1M tokens

Providers now quote prices per million tokens, but older price lists and some SDKs use per thousand. To convert, divide the per-1M price by 1,000.

GPT-6.1 Sol at $2 per 1M input tokens is $0.002 per 1K input tokens, and $10 per 1M output tokens is $0.01 per 1K. A single output token costs $0.00001.

Prices this small are why a single request usually costs a fraction of a cent. The cost adds up through volume and through long prompts, so price the token counts you actually send, not just the question you type.

FAQ

Related tools