Vertex AI

Vertex AI pricing sells the same Gemini model at four different per-token rates and a fifth rate that's measured in reserved capacity per hour.

Pricing Model:

Pricing Model:

Usage-based, with optional reserved capacity

Usage-based, with optional reserved capacity

Usage-based, with optional reserved capacity

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

Google Cloud billing credits only. No wallet, no top-ups

Google Cloud billing credits only. No wallet, no top-ups

Google Cloud billing credits only. No wallet, no top-ups

Updated on:

Google Vertex AI pricing: four rates for one model, plus reserved capacity

Vertex AI pricing sells the same Gemini model at four different per-token rates and a fifth rate that's measured in reserved capacity per hour. Standard, Priority, Flex and Batch all bill tokens at different multipliers, and Provisioned Throughput bills generative AI scale units by the hour regardless of traffic. Google publishes every figure, which makes the platform unusually legible and unusually easy to mis-buy.

Key takeaways

  • Priority costs 1.8x the standard rate and Flex and Batch each cost 50% of it, so a Gemini 3.1 Pro input token under 200K ranges from $1.00 to $3.60 per million, and $2.00 to $7.20 above it.

  • Provisioned Throughput bills per GSU per hour, from $7.14 on a one-week commit down to $2.74 on a one-year commit at global endpoints.

  • When you exceed reserved throughput, Google processes the request at pay-as-you-go rates by default. Setting the header X-Vertex-AI-LLM-Request-Type: dedicated turns that into a 429 instead.

  • Implicit context caching is on by default and cuts cached input tokens by 90% with no storage charge.

Vertex AI pricing in 2026

Consumption mode

Rate versus standard

Commitment

Built for

 

Standard PayGo

1.0x

None

General production traffic

Priority PayGo

1.8x

None

Latency-sensitive work

Flex PayGo

0.5x

None

Synchronous but latency-tolerant jobs, preview, global endpoint only

Batch

0.5x

None

Asynchronous bulk work, roughly 24-hour turnaround

Provisioned Throughput

Per GSU per hour

1 week to 1 year

Guaranteed throughput for real-time apps

What Vertex AI actually meters

Tokens, counted per model and per modality, with three qualifiers that change the bill more than the headline rate does.

First, price steps up at 200K input tokens. Gemini 3.1 Pro charges $2.00 per million input below that line and $4.00 above it, so a long-context application pays a different unit price than a short-prompt one on the same model.

Second, Google charges only for requests returning a 200 response code. Failed requests cost nothing on input or output, which is rarer than it sounds.

Third, reserved capacity meters something else. A GSU buys a fixed throughput allowance, and each model publishes burndown rates converting its modalities into it. On gemini-2.0-flash an input text token burns down as 1, an input audio token as 7, an output text token as 4. Google's worked example turns 10 queries per second into 57,000 adjusted tokens per second, divides by 3,360 per GSU, and lands on 17 GSUs.

Grounding meters separately. Gemini 3 includes 5,000 Grounding Queries a month free then bills $14 per 1,000. Gemini 2.x uses a different unit, the Grounding Prompt, with daily allowances and a $35 per 1,000 rate, or $45 for Web Grounding for Enterprise.

How credits and caching discounts work

Vertex AI has no credit wallet. What Google calls credits are billing credits applied to an invoice, and the current one is a 50% Provisioned Throughput promotion on Gemini 3.8, 3.7 and 3.6 Flash running 13 August to 31 December 2026. Those credits issue on the 7th of the following month and expire 30 days later. There's no rollover, pooling or top-up purchase.

The real discount mechanic is caching. Implicit caching is on by default on every project and takes 90% off cached input tokens with no storage charge. Explicit caching lets you control what gets cached and for how long, at 90% off on Gemini 2.5 and later and 75% on Gemini 2.0, but adds a storage charge per cached token per hour, $0.0000045 on Gemini 3.1 Pro and $0.000001 on Flash. Default TTL is 60 minutes and minimum cache sizes run 2,048 to 6,144 tokens. Cache and Batch discounts don't stack: the 90% cache hit wins.

What happens when you hit the limit

Nothing blocks by default, which is the most consequential line in the docs. A request exceeding your remaining Provisioned Throughput quota gets processed in full as an on-demand request at the pay-as-you-go rate, appearing as spillover on the monitoring dashboard.

You change that with a header. X-Vertex-AI-LLM-Request-Type: dedicated restricts you to reserved capacity and returns a 429 past it. shared bypasses the reservation entirely. The choice between a hard cap and an uncapped bill is a per-request header, not an account setting, and the default is the uncapped one.

How Vertex AI pricing has changed across all these years

Date

Milestone

Source

 

1 Jan 2027

Gemini 3.8, 3.7 and 3.6 Flash introductory pricing ends, $0.75/$3.75 becomes $1.50/$7.50

Vendor

13 Aug 2026

50% Provisioned Throughput billing credit on Gemini 3.x Flash models, through 31 Dec 2026

Vendor

1 Jul 2026

Non-global endpoint premium takes effect for Gemini 3 and later

Vendor

22 Apr 2026

Vertex AI becomes Gemini Enterprise Agent Platform. Names change, rates don't

Vendor

5 Jan 2026

Gemini 3 grounding switches to the Grounding Query unit at $14 per 1,000 after 5,000 free per month

Vendor

15 Jul 2025

Provisioned Throughput purchases restricted to GA endpoints, existing preview orders must migrate

Vendor

17 Jun 2025

Gemini 2.5 Flash GA repriced with unified output tokens: thinking output cheaper, non-thinking output dearer

Vendor

5 May 2025

Web Grounding for Enterprise starts billing at $45 per 1,000 prompts

Vendor

1 Aug 2023

TensorBoard moves off a $300 per user per month licence to $10 per GiB per month of log storage

Vendor

Flexprice’s Take

Vertex AI is the best-documented pricing in this index, and it's still hard to buy correctly.

Google publishes every Vertex AI rate, and that's not the problem. The discount tiers are simple multipliers, the GSU maths comes with a worked example, and future price rises carry a date. Charging nothing for failed requests is the fair choice and almost nobody else makes it.

The problem is how many dials you set at once. Four per-token rates, a capacity unit, two caching modes, two grounding units and a price step at 200K context compound into a bill you can't predict even with every number.

Spillover as the default is the one real mistake. Overflow bills at on-demand rates unless you send the dedicated header, so the protective behaviour is the one you opt into.

Best For

Teams with predictable high-volume traffic who'll actually do the GSU sizing work.

Watch Out For

Default spillover quietly billing overflow at on-demand rates.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling Gemini capacity and need cost and margin per customer per model?

Flexprice tracks it at that grain.

Flexprice’s Take

Vertex AI is the best-documented pricing in this index, and it's still hard to buy correctly.

Google publishes every Vertex AI rate, and that's not the problem. The discount tiers are simple multipliers, the GSU maths comes with a worked example, and future price rises carry a date. Charging nothing for failed requests is the fair choice and almost nobody else makes it.

The problem is how many dials you set at once. Four per-token rates, a capacity unit, two caching modes, two grounding units and a price step at 200K context compound into a bill you can't predict even with every number.

Spillover as the default is the one real mistake. Overflow bills at on-demand rates unless you send the dedicated header, so the protective behaviour is the one you opt into.

Best For

Teams with predictable high-volume traffic who'll actually do the GSU sizing work.

Watch Out For

Default spillover quietly billing overflow at on-demand rates.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling Gemini capacity and need cost and margin per customer per model?

Flexprice tracks it at that grain.

Customer
Sentiment Highlights

Our GCP sales rep keeps recommending \"provisioned throughput,\" but it's both expensive, and doesn't fit our workload type.

film42, Hacker News, August 2025

something like 6x more expensive with 3.1 flash lite

data-ottawa, on BigQuery workflows, Hacker News, July 2026

Frequently Asked Questions

Frequently Asked Questions

How much does Vertex AI cost?

What is a GSU in Vertex AI?

Does Vertex AI have a free tier?

What happens when you exceed Provisioned Throughput on Vertex AI?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack