

Vertex AI
Vertex AI pricing sells the same Gemini model at four different per-token rates and a fifth rate that's measured in reserved capacity per hour.
Updated on:
Google Vertex AI pricing: four rates for one model, plus reserved capacity
Vertex AI pricing sells the same Gemini model at four different per-token rates and a fifth rate that's measured in reserved capacity per hour. Standard, Priority, Flex and Batch all bill tokens at different multipliers, and Provisioned Throughput bills generative AI scale units by the hour regardless of traffic. Google publishes every figure, which makes the platform unusually legible and unusually easy to mis-buy.
Key takeaways
Priority costs 1.8x the standard rate and Flex and Batch each cost 50% of it, so a Gemini 3.1 Pro input token under 200K ranges from $1.00 to $3.60 per million, and $2.00 to $7.20 above it.
Provisioned Throughput bills per GSU per hour, from $7.14 on a one-week commit down to $2.74 on a one-year commit at global endpoints.
When you exceed reserved throughput, Google processes the request at pay-as-you-go rates by default. Setting the header X-Vertex-AI-LLM-Request-Type: dedicated turns that into a 429 instead.
Implicit context caching is on by default and cuts cached input tokens by 90% with no storage charge.
Vertex AI pricing in 2026
Consumption mode | Rate versus standard | Commitment | Built for
|
|---|---|---|---|
Standard PayGo | 1.0x | None | General production traffic |
Priority PayGo | 1.8x | None | Latency-sensitive work |
Flex PayGo | 0.5x | None | Synchronous but latency-tolerant jobs, preview, global endpoint only |
Batch | 0.5x | None | Asynchronous bulk work, roughly 24-hour turnaround |
Provisioned Throughput | Per GSU per hour | 1 week to 1 year | Guaranteed throughput for real-time apps |
What Vertex AI actually meters
Tokens, counted per model and per modality, with three qualifiers that change the bill more than the headline rate does.
First, price steps up at 200K input tokens. Gemini 3.1 Pro charges $2.00 per million input below that line and $4.00 above it, so a long-context application pays a different unit price than a short-prompt one on the same model.
Second, Google charges only for requests returning a 200 response code. Failed requests cost nothing on input or output, which is rarer than it sounds.
Third, reserved capacity meters something else. A GSU buys a fixed throughput allowance, and each model publishes burndown rates converting its modalities into it. On gemini-2.0-flash an input text token burns down as 1, an input audio token as 7, an output text token as 4. Google's worked example turns 10 queries per second into 57,000 adjusted tokens per second, divides by 3,360 per GSU, and lands on 17 GSUs.
Grounding meters separately. Gemini 3 includes 5,000 Grounding Queries a month free then bills $14 per 1,000. Gemini 2.x uses a different unit, the Grounding Prompt, with daily allowances and a $35 per 1,000 rate, or $45 for Web Grounding for Enterprise.
How credits and caching discounts work
Vertex AI has no credit wallet. What Google calls credits are billing credits applied to an invoice, and the current one is a 50% Provisioned Throughput promotion on Gemini 3.8, 3.7 and 3.6 Flash running 13 August to 31 December 2026. Those credits issue on the 7th of the following month and expire 30 days later. There's no rollover, pooling or top-up purchase.
The real discount mechanic is caching. Implicit caching is on by default on every project and takes 90% off cached input tokens with no storage charge. Explicit caching lets you control what gets cached and for how long, at 90% off on Gemini 2.5 and later and 75% on Gemini 2.0, but adds a storage charge per cached token per hour, $0.0000045 on Gemini 3.1 Pro and $0.000001 on Flash. Default TTL is 60 minutes and minimum cache sizes run 2,048 to 6,144 tokens. Cache and Batch discounts don't stack: the 90% cache hit wins.
What happens when you hit the limit
Nothing blocks by default, which is the most consequential line in the docs. A request exceeding your remaining Provisioned Throughput quota gets processed in full as an on-demand request at the pay-as-you-go rate, appearing as spillover on the monitoring dashboard.
You change that with a header. X-Vertex-AI-LLM-Request-Type: dedicated restricts you to reserved capacity and returns a 429 past it. shared bypasses the reservation entirely. The choice between a hard cap and an uncapped bill is a per-request header, not an account setting, and the default is the uncapped one.
How Vertex AI pricing has changed across all these years
Date | Milestone | Source
|
|---|---|---|
1 Jan 2027 | Gemini 3.8, 3.7 and 3.6 Flash introductory pricing ends, $0.75/$3.75 becomes $1.50/$7.50 | Vendor |
13 Aug 2026 | 50% Provisioned Throughput billing credit on Gemini 3.x Flash models, through 31 Dec 2026 | Vendor |
1 Jul 2026 | Non-global endpoint premium takes effect for Gemini 3 and later | Vendor |
22 Apr 2026 | Vertex AI becomes Gemini Enterprise Agent Platform. Names change, rates don't | Vendor |
5 Jan 2026 | Gemini 3 grounding switches to the Grounding Query unit at $14 per 1,000 after 5,000 free per month | Vendor |
15 Jul 2025 | Provisioned Throughput purchases restricted to GA endpoints, existing preview orders must migrate | Vendor |
17 Jun 2025 | Gemini 2.5 Flash GA repriced with unified output tokens: thinking output cheaper, non-thinking output dearer | Vendor |
5 May 2025 | Web Grounding for Enterprise starts billing at $45 per 1,000 prompts | Vendor |
1 Aug 2023 | TensorBoard moves off a $300 per user per month licence to $10 per GiB per month of log storage | Vendor |
Customer
Sentiment Highlights
Our GCP sales rep keeps recommending \"provisioned throughput,\" but it's both expensive, and doesn't fit our workload type.
film42, Hacker News, August 2025
something like 6x more expensive with 3.1 flash lite
data-ottawa, on BigQuery workflows, Hacker News, July 2026
Explore other providers

Gemini API
API
Gemini API pricing charges nothing on the free tier and per million tokens on the paid one, both on the same endpoint.

ElevenLabs
AI Voice
ElevenLabs pricing sells a monthly credit quota, and the credit is the old character unit renamed.

GitHub Copilot
Developer Tool
GitHub Copilot pricing stopped counting requests in June 2026 and started counting money.
How much does Vertex AI cost?
What is a GSU in Vertex AI?
Does Vertex AI have a free tier?
What happens when you exceed Provisioned Throughput on Vertex AI?
























