FREE COST TOOL • GEMINI API PRICING
Gemini API Pricing
Calculator
Estimate paid Gemini Developer API cost using model-specific Standard or Batch rates. Include input tokens, thinking-output tokens, cached input, cache storage, the Gemini 3.1 Pro 200K-token price boundary, and Google Search grounding overage in one monthly scenario.
PAID API WORKLOAD
Editable example only—this page is not connected to a billing account and cannot charge you. Replace the preset workload with your own numbers before using the estimate.
Checked July 24, 2026: Gemini 3.6 Flash Standard uses $1.5 input, $0.15 cached input, and $7.5 output per 1M tokens. Cache storage is $1 per 1M token-hours.
$0.0105 per request after entered uplift
$3 input + $7.5 output · 2,000,000 in / 1,000,000 out
$0 storage + $0 Search grounding
Compared with billing every entered input token at the uncached rate, after cache storage.
This selected model has no separate 200K price band in the supported table.
Queries left before the entered scenario reaches the shared 5,000-query monthly allowance.
Compare the same workload across Gemini models
| Model | Tier | Context band | Monthly estimate | Per request |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Standard | Standard | $3.1 | $0.0031 |
| Gemini 3.6 Flash | Standard | Standard | $10.5 | $0.0105 |
| Gemini 3.5 Flash | Standard | Standard | $12 | $0.012 |
| Gemini 3.1 Pro Preview | Standard | Standard | $16 | $0.016 |
The lowest cost is not automatically the right model. Compare output quality, latency, model status, context behavior, rate limits, tool support, and data requirements before choosing.
AVOID THE COMMON BILLING MISS
Model more than token multiplication.
- Thinking tokens: include them in output because Google's output rates include thinking-token consumption.
- Cache storage: a lower cached-input rate does not remove the separate token-hour storage charge.
- Search grounding: enter search queries, not prompts; one customer request can trigger more than one billable query.
- Long prompts: Gemini 3.1 Pro Preview changes input, output, and cached-input rates above 200K prompt tokens.
- Batch: run it only for asynchronous workloads that can accept Batch behavior and limits.
Frequently asked questions
How much does the Gemini API cost?
Gemini API cost depends on the model, service tier, input tokens, output tokens including thinking tokens, cached input, cache storage, and optional tools such as Google Search grounding. This calculator applies the paid-tier list rates checked on July 24, 2026 to the workload you enter.
Is Gemini Batch API cheaper than Standard?
Google describes Batch API as a 50% cost reduction for paid usage. The exact cached-input price is model-specific, so this calculator uses the published Batch row for each supported model rather than applying one blanket multiplier.
Does Gemini charge for thinking tokens?
Yes. Google's pricing tables describe output prices as including thinking tokens. Enter the total billable output shown by your representative workload, not only the visible answer text.
When does Gemini 3.1 Pro long-context pricing apply?
The official Gemini 3.1 Pro Preview pricing table uses a higher price band when a prompt exceeds 200,000 input tokens. The calculator activates that band automatically when input tokens per request are above 200,000.
How is Gemini context caching calculated?
Cached input is billed at its published per-million-token rate. Explicit cache storage is a separate token-hour charge, so the calculator asks for both the cached share of monthly input and the average stored cache size multiplied by storage hours.
How much does Google Search grounding cost?
For the supported Gemini 3 models in this calculator, the paid pricing page listed 5,000 queries per month at no additional Search grounding charge, then $14 per 1,000 search queries. One prompt may produce more than one query, so enter queries rather than prompts.
Independent estimate; not affiliated with Google. This calculator models paid Gemini Developer API list rates for the supported text and multimodal input models. It excludes free-tier limits, Flex, Priority, Maps, Live API, image or video output, agent loops, enterprise discounts, prepaid-credit rules, rate limits, and charges outside the fields shown. Verify the official pricing page and your invoice before purchase or launch.