Explanation
Why the waste happens and who it affects.
It is bought in generative AI scale units (GSUs) for a 1-week, 1-month, 3-month or 1-year term, and the term fee applies regardless of actual usage. Throughput that is not used in a given period does not accumulate or carry over.
Orders are often sized for a launch, a forecast or a peak hour and then never revisited, so average GSU utilization settles well below what was bought. Bursty traffic that spikes in business hours and drops overnight makes it worse, as the FinOps Foundation notes when it describes the resulting idle allocated capacity. Orders can also be left auto-renewing on a model the application no longer calls, because term fees keep applying even if the model is discontinued.
Billing model
The pricing dimensions that drive this cost.
- Generative AI scale unit (GSU)
- The unit of reserved throughput for a model and region, billed at a fixed price per GSU for the chosen term, with lower prices for longer terms
- Commitment term
- 1 week (Google models only), 1 month, 3 months or 1 year; orders cannot be cancelled mid-term and are billed once active
- Spillover
- Requests that exceed the reserved quota are processed and billed at pay-as-you-go rates by default
- GSU changes
- Increases apply immediately on approval, while decreases take effect only at auto-renewal for the next term
How to detect
4 checks to find it in your estate.
- Open the Provisioned Throughput page Utilization summary tab, which shows per model the GSUs owned, peak throughput usage in GSUs, average GSU utilization and how often the limit was reached
- In Cloud Monitoring, compare aiplatform.googleapis.com/publisher/online_serving/consumed_token_throughput filtered to request_type dedicated with dedicated_token_limit or dedicated_gsu_limit on the PublisherModel resource
- Flag orders whose peak usage rarely approaches the limit and that show no spillover, and orders with near-zero dedicated traffic because the application moved to another model, version or region
- List active orders with their term, end date and auto-renewal setting, and register Essential Contacts so expiration and auto-renewal notices, sent two weeks ahead for monthly and longer terms, reach the owner
How to fix
5 ways to remove the waste.
- Decrease GSUs on auto-renewing orders to match the sustained baseline; the reduction applies from the next term, and traffic above it spills over to pay-as-you-go by default
- Turn off auto-renewal for orders that are no longer needed, keeping in mind that an active order cannot be changed in the last five days before expiry unless it auto-renews
- If the application switched models, change the order's model or model version (model changes are limited to the same publisher), or its region, rather than paying for throughput nobody calls
- Consolidate traffic for the same model and region onto one order, and route development or experimental traffic to pay-as-you-go with the X-Vertex-AI-LLM-Request-Type header set to shared so it does not consume reserved capacity
- Use shorter terms while demand is uncertain and size new orders with the GSU estimator and measured traffic, not a launch forecast
Documentation
Vendor references for pricing and configuration.