Gemini API Cost Calculator With 2026 Pricing Tiers

Gemini API pricing calculator covering 3.1 Pro's 200K context tier, 3.5 Flash, Flash-Lite, storage-billed caching, and batch discounts, updated weekly.

Model$ / 1M in · outCost / requestCost / 1K requestsCost / month
GPT-5.6 Sol OpenAI$5.00 · $30.00$0.017750$17.750$1775.00
GPT-5.6 Terra OpenAI$2.50 · $15.00$0.008875$8.875$887.50
GPT-5.6 Luna OpenAI$1.00 · $6.00$0.003550$3.550$355.00
GPT-5.4 mini OpenAI$0.75 · $4.50$0.002662$2.663$266.25
GPT-5.4 nano OpenAI$0.20 · $1.25$0.000735$0.735$73.50
Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3×$10.00 · $50.00$0.031125$31.125$3112.50
Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3×$5.00 · $25.00$0.015562$15.562$1556.25
Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled$2.00 · $10.00$0.009338$9.338$933.75
Claude Haiku 4.5 Anthropic (Claude)$1.00 · $5.00$0.003112$3.112$311.25
Gemini 3.1 Pro (Preview) Google (Gemini API)$2.00 · $12.00$0.007100$7.100$710.00
Gemini 3.5 Flash Google (Gemini API)$1.50 · $9.00$0.005325$5.325$532.50
Gemini 2.5 Flash-Lite Google (Gemini API)$0.10 · $0.40$0.000255$0.255$25.50
DeepSeek V4 Pro DeepSeek$0.43 · $0.87$0.000654$0.654$65.43
DeepSeek V4 Flash DeepSeek$0.14 · $0.28$0.000211$0.211$21.14
Mistral Medium 3.5 Mistral$1.50 · $7.50$0.005250$5.250$525.00
Mistral Large 3 Mistral$0.50 · $1.50$0.001025$1.025$102.50
Mistral Small 4 Mistral$0.15 · $0.60$0.000382$0.382$38.25
Grok 4.5 xAI (Grok)$2.00 · $6.00$0.004250$4.250$425.00
Grok 4.3 xAI (Grok)$1.25 · $2.50$0.001975$1.975$197.50
Llama 4 Scout Meta Llama (hosted)$0.10 · $0.30$0.000250$0.250$25.00
Llama 4 Maverick Meta Llama (hosted)$0.15 · $0.60$0.000450$0.450$45.00
Amazon Nova Pro Amazon (Nova on Bedrock)$0.80 · $3.20$0.002400$2.400$240.00
Amazon Nova Lite Amazon (Nova on Bedrock)$0.06 · $0.24$0.000180$0.180$18.00
Amazon Nova Micro Amazon (Nova on Bedrock)$0.04 · $0.14$0.000105$0.105$10.50

The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.

How Google (Gemini API) pricing works

  • Prompt caching: Cached tokens read cheaply, but cache STORAGE is billed per 1M tokens per hour ($1.00 to $4.50 per 1M per hour by model). It is not modeled per-request and appears as a footnote
  • Batch API: −50% on asynchronous workloads.
  • Gemini 3.1 Pro doubles input pricing above 200K input tokens, the only two-tier context pricing in this comparison.
  • Real free tier exists (data may be used for training).
  • Source: official pricing page, details last verified 2026-07-11.

The 200K cliff, worked through a $1.209 request

Gemini 3.1 Pro is the only model in this 24-model comparison with two-tier context pricing, charging $2 input / $12 output per million tokens for requests up to 200K input tokens and $4/$18 above that line. The higher tier applies to the entire request, not the overage. Send a 300,000-token prompt with a 500-token reply and the bill is $1.209, which is all 300K input tokens at $4 ($1.20) plus the output at $18 ($0.009).

Split the same material into two 150K-token requests and each stays in tier one, at $0.30 of input apiece plus $0.006 for each 500-token reply. That comes to $0.612 in total, 49% less for the same tokens. Chunking carries engineering costs of its own (retrieval logic, answer merging, lost cross-document attention), but at 2x the input rate and 1.5x the output rate above the line, the incentive is explicit. The agent-task figure of $0.252 in the calculator assumes a 150K-token context precisely because it stays under the threshold.

Staying in tier one

  1. Cap retrieval so prompt plus documents stay below 200K input tokens per call.
  2. Summarize long conversation history instead of replaying it verbatim.
  3. Split multi-document analysis into per-document calls and merge the results.

For genuinely indivisible long contexts, note that Claude and the GPT-5.6 family both price 1M-token context flat, with no tier jump at any length.

Storage-billed caching works like rent, not a toll

Gemini’s context caching combines discounted reads with a storage fee of $1.00 to $4.50 per million tokens per hour the cache exists. Holding a 150K-token agent context cached therefore costs $0.15 to $0.675 per hour whether one request reads it or a thousand do. Anthropic’s alternative charges a one-time write premium of 1.25-2x input and nothing while the cache sits idle. The consequences follow directly. Gemini’s model wins on sustained, high-frequency traffic against a shared context, where hourly rent amortizes across many reads, and loses on spiky or overnight-idle workloads where the meter runs against zero requests. If traffic against a given context falls to occasional reuse, delete the cache and re-send the prompt. The per-request figures in the calculator assume discounted cache reads only. Storage rent depends on how long you hold each cache, so model it against your actual request rate.

Flash-Lite is the cheapest serious model in the comparison

Gemini 2.5 Flash-Lite at $0.10/$0.40 prices a standard chat exchange at $0.000255, the lowest of the 24 models compared above. The next cheapest, DeepSeek V4 Pro, costs $0.000654 for the identical exchange, and Mistral Large 3 costs $0.001025. One million Flash-Lite chats a month run $255, against $5,325 on Gemini 3.5 Flash and $7,100 on 3.1 Pro. That roughly 21x gap between Flash-Lite and Flash is the widest intra-provider spread in the comparison, which argues for evaluating Flash-Lite first on classification, extraction, and routing tasks and reserving the bigger models for steps where evals show it failing.

3.5 Flash and 3.1 Pro per workload

Gemini 3.5 Flash ($1.50/$9, with thinking tokens included in the output rate) costs $0.005325 per chat exchange, $0.01017 per RAG query, and $0.189 per reasoning-heavy agent task. Gemini 3.1 Pro costs $0.0071, $0.01356, and $0.252 on the same anchors. Against the rest of the market at identical token counts, Pro’s $0.0071 chat cost undercuts GPT-5.6 Terra’s $0.008875 and sits above Claude Sonnet 5’s introductory $0.006225, a gap that inverts when Sonnet rises to $3/$15 on September 1, 2026, taking its exchange cost to about $0.00934. At 1,000 agent tasks a month, Pro is $252 and Flash $189, versus $315 on Terra and $254 on Sonnet 5’s intro rate. Gemini’s 50% batch discount applies on top for asynchronous work, taking Flash agent tasks to $0.0945 each.

The free tier, and why stale calculators mislead

Gemini’s free tier is a genuine $0 tier, rare among the providers compared here, with the documented caveat that free-tier prompts and outputs may be used for model training, which rules it out for confidential data while leaving it useful for prototyping and public content. A second data-quality warning applies when comparing tools. Many Gemini calculators still show 1.5-era prices, a generation whose rates bear no relation to 3.1 Pro’s tiered pricing or 3.5 Flash’s thinking-inclusive output billing. The figures on this page refresh weekly, and the last-verified date displayed above the calculator shows exactly how current they are.

Frequently asked questions

A standard chat exchange (1,000 input tokens, 500 output tokens, 50% cache hits) costs $0.0071 on Gemini 3.1 Pro, $0.005325 on Gemini 3.5 Flash, and $0.000255 on Gemini 2.5 Flash-Lite. At one million exchanges a month, that spans $255 on Flash-Lite to $7,100 on 3.1 Pro.

Gemini 3.1 Pro charges $2/$12 per million tokens for requests up to 200K input tokens and $4/$18 for requests above that line, applied to the whole request rather than just the overage. A single 300,000-token request with a 500-token reply costs $1.209, while the same material split into two 150K requests costs $0.612. It is the only two-tier context pricing among the 24 models in this comparison.

Gemini bills caching as storage. Reads are discounted, but the cache itself costs $1.00 to $4.50 per million tokens per hour it exists, whether or not requests hit it. Holding a 150K-token context cached costs $0.15 to $0.675 per hour. This favors sustained high-frequency traffic against a shared context and penalizes spiky or idle workloads.

Yes, Google offers a genuine $0 free tier for the Gemini API, which most competitors do not. The caveat is that prompts and outputs on the free tier may be used for model training, so it suits prototypes and public data but not confidential material. Production workloads with sensitive data belong on the paid tier.

Among the 24 models in this comparison, yes. Gemini 2.5 Flash-Lite at $0.10/$0.40 per million tokens prices a standard chat exchange at $0.000255, below DeepSeek V4 Pro at $0.000654 and Mistral Large 3 at $0.001025. It is roughly 21x cheaper per exchange than Gemini 3.5 Flash.

Many fee calculators still display Gemini 1.5-era pricing, a model generation whose rates do not apply to the current 3-series models. Gemini 3.1 Pro’s two-tier context pricing and 3.5 Flash’s thinking-inclusive output rate are both recent structures that stale tools miss. The pricing behind this page is re-verified weekly against Google’s official pricing documentation, with the last-verified date shown above the calculator.