

Gemini API
Gemini API pricing charges nothing on the free tier and per million tokens on the paid one, both on the same endpoint.
Updated on:
Gemini API pricing: the free tier is real, and you pay for it in data
Gemini API pricing charges nothing on the free tier and per million tokens on the paid one, both on the same endpoint. The free tier isn't a trial with an expiry. It's a permanent plan where the price is your data: Google uses free-tier prompts and responses to improve its products, and paid-tier ones it doesn't. Every rate is published. The free tier's limits are not.
Key takeaways
Google uses free-tier content to improve its products and doesn't use paid-tier content, and it says so on every model row of the pricing page.
Paid rates match Vertex AI at global endpoints: Gemini 3.1 Pro Preview runs $2.00 per million input and $12.00 output on both, and Vertex adds 10% on non-global endpoints.
Batch and Flex each cost 50% of standard, Priority costs 1.8x, and cached input reads cost 0.1x.
Prepay credits expire after 12 months, aren't refundable unless you switch to Postpay, and stop every API key on the billing account once the balance hits $0.
Gemini API pricing in 2026
Model | Input /1M | Output /1M | Cached input /1M | Free tier
|
|---|---|---|---|---|
Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | Yes |
Gemini 3.7 Flash | $0.75 | $3.75 | $0.075 | Yes |
Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.03 | Yes |
Gemini 3.1 Pro Preview | $2.00 | $12.00 | $0.20 | No |
Gemini Embedding 2 (text) | $0.20 | n/a | n/a | Yes |
What the Gemini API actually meters
Tokens, counted four ways: input, output including thinking tokens, cached token count, and cached token storage duration. Google's billing FAQ lists those four and nothing else.
Two qualifiers matter more than the headline rate. Prompts above 200,000 tokens on Gemini 3.1 Pro Preview move to a second card, where input doubles to $4.00 and output rises to $18.00. Input climbs 100% and output 50%, so the long-context penalty isn't one multiplier. Thinking tokens bill as output at the full rate, so on Flash models reasoning costs five times what the prompt did.
Caching runs implicitly, on by default for Gemini 2.5 and newer. Cache reads bill at 0.1x base input, and a hit needs a 4,096-token prefix on every Gemini 3 model. Storage bills per million tokens per hour, $4.50 on 3.1 Pro Preview and $0.50 on Flash.
Grounding meters per request. Gemini 3.x models share 5,000 free Google Search requests a month, then bill $14 per 1,000. One user prompt can fire several searches, and Google charges for each. Google Maps grounding carries the same allowance and the same $14.
How credits work
Gemini API credits are a prepaid balance, not a usage allowance. Paid accounts default to Prepay: load a minimum of $5 and a maximum of $5,000, and usage draws down in near real time.
The mechanics are strict. Credits expire 12 months after purchase. Refunds don't exist on Prepay except when you switch to Postpay, which closes the account and returns the balance; close it any other way and the balance is forfeited. Deduction order runs promotional Google Cloud credits first, then prepaid funds, but only while the prepaid balance is positive. At $0 the Cloud credits stop being consumed too, so promotional credit alone can't keep an account alive.
Auto-reload tops the balance up at a threshold you set, and a monthly auto-charge limit caps how much it spends per cycle before switching itself off.
What happens when you hit the limit
Everything stops, and Google built four separate stops. A $0 Prepay balance kills every API key on that billing account at once. A project spend cap pauses one project. A billing account tier cap pauses them all until the 1st of the next month, at $250 on Tier 1, $2,000 on Tier 2 and $20,000 to $100,000 on Tier 3. A spend-based rate limit returns 429 on a rolling 10-minute window, at $10, $50 and $200 across those tiers.
None of these are exact, and Google says so. Billing data lags about 10 minutes, so batch jobs and agent sessions can overrun a cap before the system catches up. Free-tier rate limits aren't published as numbers in the docs at all: you're told to read them in AI Studio.
How Gemini API pricing has changed across all these years
Date | Milestone | Source
|
|---|---|---|
1 Jan 2027 | Flash introductory pricing ends, $0.75 and $3.75 become $1.50 and $7.50 | Vendor |
13 Aug 2026 | Gemini 3.7 Flash ships GA at an introductory price through 31 Dec 2026 | Vendor |
23 Mar 2026 | Prepay and Postpay billing plans roll out, Prepay becomes the default | Vendor |
16 Mar 2026 | Usage tiers revamped, billing account spend caps introduced | Vendor |
12 Mar 2026 | Project-level spend caps added in AI Studio | Vendor |
5 Dec 2025 | Google announces Gemini 3 Grounding with Google Search will start billing on 5 Jan 2026 | Vendor |
4 Nov 2025 | Gemini 2.5 Flash Image input drops from 1,290 to 258 tokens per image, cutting edit costs | Vendor |
Customer
Sentiment Highlights
As usual for something so simple, Google's docs seem unclear
Developer comparing Gemini image token pricing, Hacker News, August 2026
Explore other providers

Devin
Developer Tool
Devin pricing meters an Agent Compute Unit, and the useful thing about it is how precisely Cognition documents what does and does not consume one.

ElevenLabs
AI Voice
ElevenLabs pricing sells a monthly credit quota, and the credit is the old character unit renamed.

Deepgram
AI Voice
Deepgram pricing meters audio by the second and refuses to round.
Is the Gemini API free?
How much does the Gemini API cost?
Is the Gemini API cheaper than Vertex AI?
Do Gemini API credits expire?
























