- Home
- Gemini 4 Argon
- Cost calculator
Gemini 4 Argon Cost Calculator
Estimate what Gemini 4 Argon costs per request, per day and per month. It is priced at $2 per 1M input tokens and $10 per 1M output tokens. Pick a workload or enter your own token counts; the table below runs the same numbers on every other model. For specs and pricing rules, see the Gemini 4 Argon overview.
Gemini 4 Argon has been announced but is not yet available in the API. Estimates use the announced introductory price; after the introductory period it rises to $4 / $20 per 1M input / output tokens.
Discounts are off by default. Models without a cached rate or batch discount are billed at their standard price.
Gemini 4 Argon estimate
- Per request
- $0.004
- Per day
- $1.33
- Per month
- $40.00
Monthly split: Input $10.00 · Cached input $0 · Output $30.00
The same workload on other models
| # | Model | Provider | Per month | Per request |
|---|---|---|---|---|
| 1 | Qwen3.7-Flash Alibaba (Qwen) | Alibaba (Qwen) | $0.54 | $0.000054 |
| 2 | Ministral 3 3B Mistral | Mistral | $0.80 | $0.00008 |
| 3 | Ministral 3 8B Mistral | Mistral | $1.20 | $0.00012 |
| 4 | Ministral 3 14B Mistral | Mistral | $1.60 | $0.00016 |
| 5 | GPT-6 Luna OpenAI | OpenAI | $2.00 | $0.0002 |
| 6 | Qwen3.8-Flash Alibaba (Qwen) | Alibaba (Qwen) | $2.16 | $0.00022 |
| 7 | GLM-5.3-Flash Zhipu (GLM) | Zhipu (GLM) | $2.25 | $0.00022 |
| 8 | Mistral Small 4 Mistral | Mistral | $2.55 | $0.00026 |
| 9 | Codestral Mistral | Mistral | $4.20 | $0.00042 |
| 10 | GPT-5.6 Luna OpenAI | OpenAI | $4.60 | $0.00046 |
| 11 | DeepSeek-V4.1-Flash DeepSeek | DeepSeek | $5.10 | $0.00051 |
| 12 | MiniMax-M3 MiniMax | MiniMax | $5.10 | $0.00051 |
| 13 | MiniMax-M2.7 MiniMax | MiniMax | $5.10 | $0.00051 |
| 14 | GLM-5.3-FlashX Zhipu (GLM) | Zhipu (GLM) | $5.60 | $0.00056 |
| 15 | Gemini 3.1 Flash-Lite Google | $5.75 | $0.00057 | |
| 42 | Gemini 4 Argon Google | $40.00 | $0.004 |
How the estimate is calculated
Cost per request = (uncached input tokens × input rate + cached input tokens × cached rate + output tokens × output rate) ÷ 1,000,000.
The batch switch applies the provider's batch discount to every token type. Requests above a long-context threshold use the higher long-context rates for the whole request. A month is 30 days.
Worked example: Gemini 4 Argon for a RAG assistant
Here is how the Gemini 4 Argon cost calculator gets its numbers, using the RAG Q&A preset. Each request sends 4,000 input tokens (instructions, retrieved passages and the question) and gets a 500-token answer, and the assistant handles 20,000 requests a month.
- Input: 4,000 × $2 ÷ 1M = $0.008
- Output: 500 × $10 ÷ 1M = $0.005
- One request: $0.013. One month: $0.013 × 20,000 = $260
That uses the introductory price. At the later price of $4 input and $20 output per 1M tokens, the same month costs $520.
If the first 1,000 tokens are a fixed system prompt, they can be served from the prompt cache at $0.10 per 1M instead of $2. That brings the month to $222.
Gemini 4 Argon vs Claude Opus 5.5: cost
Gemini 4 Argon launches at $2 input and $10 output per 1M tokens, rising to $4 and $20 after the introductory period. Claude Opus 5.5 costs $4 and $20, with cached input at $0.20. Monthly cost at 10,000 requests, standard rates, no caching:
| Workload | Gemini 4 Argon (intro) | Gemini 4 Argon (later) | Claude Opus 5.5 |
|---|---|---|---|
| Chatbot | $40.00 | $80.00 | $80.00 |
| RAG Q&A | $130 | $260 | $260 |
| Coding agent | $800 | $1,600 | $1,600 |
| Summarization | $220 | $440 | $440 |
At the introductory price, Gemini 4 Argon costs 50% less than Claude Opus 5.5 for the same workload. After the introductory period the per-token prices are the same, so the bill comes down to discounts and to how many tokens each model needs for the same task.
Claude Opus 5.5 is available in the API today, with a 50% batch discount. Gemini 4 Argon is not yet, and Google has not announced its batch pricing.
This compares price only. How many tokens and retries each model needs for your task decides the real bill, so test both on your own prompts before you switch.
Estimate your own Gemini 4 Argon workload
The presets are a starting point. To estimate your own Gemini 4 Argon bill, replace them with numbers from your app:
- Tokens per request. Every API response reports the input and output tokens it used. Log them for a day of real traffic and enter the averages. Before launch, count the system prompt, any chat history or documents you send, and a typical reply.
- Requests per month. Multiply daily requests by 30. For an agent, count each model call, not each task: one task can take dozens of steps.
- Cached share. Work out how much of each prompt is identical from one request to the next, such as instructions and reference documents, and enter that share with caching turned on.
- Batch. Turn on the batch switch for the part of your traffic that can wait for results.
- Headroom. Retries, reasoning tokens and traffic spikes all add to the bill, so budget above the estimate rather than exactly at it.
What changes your Gemini 4 Argon bill
Reply length
Output costs 5 times as much as input on Gemini 4 Argon. Every extra 100 output tokens per request adds $10.00 per 10,000 requests, so asking for shorter answers and setting a maximum output length pays off quickly.
Prompt caching
Cached input costs $0.10 per 1M tokens, 5% of the normal input rate. Keep instructions and reference documents at the start of the prompt and unchanged between requests so they are served from the cache.
Batch processing
Google has not announced batch pricing for Gemini 4 Argon yet, so the batch switch has no effect on it for now.
Long prompts
Google has not announced a context window or long-context pricing for Gemini 4 Argon yet. Until it does, the calculator bills every prompt length at the same rates.
Is Gemini 4 Argon the cheapest option for this job?
For the RAG example above, Gemini 4 Argon comes to $260 a month. Among similar models, Grok 4.7 would cost $220 and GPT-6 Astra $1,300 for the same traffic. The cheapest of the 49 models available in the API for this workload is Qwen3.7-Flash, at $3.70 a month.
Price is only half the decision: a cheaper model that needs a second attempt or a longer prompt to get the answer right can end up costing more. Run your own token counts through the table above, which prices the same workload on every model we track, and test the two or three cheapest candidates on real requests before you switch.
FAQ
Not yet. Google has announced Gemini 4 Argon and its price, but the model is not yet open to API customers. The calculator uses the announced price so you can budget before launch.
Multiply input tokens by $2 and output tokens by $10, then divide by one million. Cached input tokens cost $0.10 per million instead. Add the two for the cost of one request, and multiply by your request count.
At 10,000 requests a month, a chatbot sending 500 input and 300 output tokens per request costs $40.00 on Gemini 4 Argon. A coding agent step with 30,000 input and 2,000 output tokens costs $800 for the same number of requests. Enter your own numbers in the calculator for your workload.
$2 for 1 million input tokens and $10 for 1 million output tokens. At a typical mix of three input tokens for every output token, 1 million tokens cost $4. That is the introductory price; afterwards it rises to $4 input and $20 output.
Gemini 4 Argon charges $2 input and $10 output per 1M tokens, against $4 and $20 for Claude Opus 5.5. Cached input costs $0.10 versus $0.20. For 20,000 RAG requests a month, that comes to $260 on Gemini 4 Argon and $520 on Claude Opus 5.5.
Google has not announced long-context pricing for Gemini 4 Argon yet. Until it does, the calculator applies the same rates at every prompt length.
Only if you count them in. Reasoning tokens are billed as output at $10 per 1M even though they are not shown in the reply. Set output tokens per request to include them; your API usage dashboard reports the real figure.
Reuse long prompts so they hit the prompt cache, keep outputs short since output tokens cost more, and send work that can wait through the Batch API if Google offers batch pricing for it.
Related tools
- LLM API pricing tableEvery model's input, output and cached rates in one sortable table.
- LLM Price ComparisonPut two to four models side by side: per-token rates, context window, discounts and what the same job costs on each.
- Cheapest LLM APIsRanked lists of the lowest-priced models overall, among flagships, with 1M-token context, with image input and with open weights.
- Token Cost CalculatorEnter a token count and see what that many input and output tokens cost on every model.
- LLM Cost CalculatorPick a workload and request volume, and compare the bill across models with caching and batch discounts.