Skip to main content

Gemini 4 Argon Cost Calculator

Estimate what Gemini 4 Argon costs per request, per day and per month. It is priced at $2 per 1M input tokens and $10 per 1M output tokens. Pick a workload or enter your own token counts; the table below runs the same numbers on every other model. For specs and pricing rules, see the Gemini 4 Argon overview.

Gemini 4 Argon has been announced but is not yet available in the API. Estimates use the announced introductory price; after the introductory period it rises to $4 / $20 per 1M input / output tokens.

Discounts are off by default. Models without a cached rate or batch discount are billed at their standard price.

Gemini 4 Argon estimate

Per request
$0.004
Per day
$1.33
Per month
$40.00

Monthly split: Input $10.00 · Cached input $0 · Output $30.00

The same workload on other models

#ModelProviderPer monthPer request
1Qwen3.7-Flash
Alibaba (Qwen)
Alibaba (Qwen)$0.54$0.000054
2Ministral 3 3B
Mistral
Mistral$0.80$0.00008
3Ministral 3 8B
Mistral
Mistral$1.20$0.00012
4Ministral 3 14B
Mistral
Mistral$1.60$0.00016
5GPT-6 Luna
OpenAI
OpenAI$2.00$0.0002
6Qwen3.8-Flash
Alibaba (Qwen)
Alibaba (Qwen)$2.16$0.00022
7GLM-5.3-Flash
Zhipu (GLM)
Zhipu (GLM)$2.25$0.00022
8Mistral Small 4
Mistral
Mistral$2.55$0.00026
9Codestral
Mistral
Mistral$4.20$0.00042
10GPT-5.6 Luna
OpenAI
OpenAI$4.60$0.00046
11DeepSeek-V4.1-Flash
DeepSeek
DeepSeek$5.10$0.00051
12MiniMax-M3
MiniMax
MiniMax$5.10$0.00051
13MiniMax-M2.7
MiniMax
MiniMax$5.10$0.00051
14GLM-5.3-FlashX
Zhipu (GLM)
Zhipu (GLM)$5.60$0.00056
15Gemini 3.1 Flash-Lite
Google
Google$5.75$0.00057
42Gemini 4 Argon
Google
Google$40.00$0.004

How the estimate is calculated

Cost per request = (uncached input tokens × input rate + cached input tokens × cached rate + output tokens × output rate) ÷ 1,000,000.

The batch switch applies the provider's batch discount to every token type. Requests above a long-context threshold use the higher long-context rates for the whole request. A month is 30 days.

Worked example: Gemini 4 Argon for a RAG assistant

Here is how the Gemini 4 Argon cost calculator gets its numbers, using the RAG Q&A preset. Each request sends 4,000 input tokens (instructions, retrieved passages and the question) and gets a 500-token answer, and the assistant handles 20,000 requests a month.

  1. Input: 4,000 × $2 ÷ 1M = $0.008
  2. Output: 500 × $10 ÷ 1M = $0.005
  3. One request: $0.013. One month: $0.013 × 20,000 = $260

That uses the introductory price. At the later price of $4 input and $20 output per 1M tokens, the same month costs $520.

If the first 1,000 tokens are a fixed system prompt, they can be served from the prompt cache at $0.10 per 1M instead of $2. That brings the month to $222.

Gemini 4 Argon vs Claude Opus 5.5: cost

Gemini 4 Argon launches at $2 input and $10 output per 1M tokens, rising to $4 and $20 after the introductory period. Claude Opus 5.5 costs $4 and $20, with cached input at $0.20. Monthly cost at 10,000 requests, standard rates, no caching:

WorkloadGemini 4 Argon (intro)Gemini 4 Argon (later)Claude Opus 5.5
Chatbot$40.00$80.00$80.00
RAG Q&A$130$260$260
Coding agent$800$1,600$1,600
Summarization$220$440$440

At the introductory price, Gemini 4 Argon costs 50% less than Claude Opus 5.5 for the same workload. After the introductory period the per-token prices are the same, so the bill comes down to discounts and to how many tokens each model needs for the same task.

Claude Opus 5.5 is available in the API today, with a 50% batch discount. Gemini 4 Argon is not yet, and Google has not announced its batch pricing.

This compares price only. How many tokens and retries each model needs for your task decides the real bill, so test both on your own prompts before you switch.

Estimate your own Gemini 4 Argon workload

The presets are a starting point. To estimate your own Gemini 4 Argon bill, replace them with numbers from your app:

  1. Tokens per request. Every API response reports the input and output tokens it used. Log them for a day of real traffic and enter the averages. Before launch, count the system prompt, any chat history or documents you send, and a typical reply.
  2. Requests per month. Multiply daily requests by 30. For an agent, count each model call, not each task: one task can take dozens of steps.
  3. Cached share. Work out how much of each prompt is identical from one request to the next, such as instructions and reference documents, and enter that share with caching turned on.
  4. Batch. Turn on the batch switch for the part of your traffic that can wait for results.
  5. Headroom. Retries, reasoning tokens and traffic spikes all add to the bill, so budget above the estimate rather than exactly at it.

What changes your Gemini 4 Argon bill

Reply length

Output costs 5 times as much as input on Gemini 4 Argon. Every extra 100 output tokens per request adds $10.00 per 10,000 requests, so asking for shorter answers and setting a maximum output length pays off quickly.

Prompt caching

Cached input costs $0.10 per 1M tokens, 5% of the normal input rate. Keep instructions and reference documents at the start of the prompt and unchanged between requests so they are served from the cache.

Batch processing

Google has not announced batch pricing for Gemini 4 Argon yet, so the batch switch has no effect on it for now.

Long prompts

Google has not announced a context window or long-context pricing for Gemini 4 Argon yet. Until it does, the calculator bills every prompt length at the same rates.

Is Gemini 4 Argon the cheapest option for this job?

For the RAG example above, Gemini 4 Argon comes to $260 a month. Among similar models, Grok 4.7 would cost $220 and GPT-6 Astra $1,300 for the same traffic. The cheapest of the 49 models available in the API for this workload is Qwen3.7-Flash, at $3.70 a month.

Price is only half the decision: a cheaper model that needs a second attempt or a longer prompt to get the answer right can end up costing more. Run your own token counts through the table above, which prices the same workload on every model we track, and test the two or three cheapest candidates on real requests before you switch.

FAQ

Related tools