# Underutilized Vertex AI Provisioned Throughput

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-vertex-ai-provisioned-throughput

Provisioned Throughput on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, is a fixed-cost, fixed-term subscription that...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Provisioned Throughput on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, is a fixed-cost, fixed-term subscription that reserves throughput for a specific generative AI model in a specific region.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Vertex AI](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-vertex-ai)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0512

Type

Underutilized Commitment

## Explanation

Why the waste happens and who it affects.

It is bought in generative AI scale units (GSUs) for a 1-week, 1-month, 3-month or 1-year term, and the term fee applies regardless of actual usage. Throughput that is not used in a given period does not accumulate or carry over.

Orders are often sized for a launch, a forecast or a peak hour and then never revisited, so average GSU utilization settles well below what was bought. Bursty traffic that spikes in business hours and drops overnight makes it worse, as the FinOps Foundation notes when it describes the resulting idle allocated capacity. Orders can also be left auto-renewing on a model the application no longer calls, because term fees keep applying even if the model is discontinued.

## Billing model

The pricing dimensions that drive this cost.

Generative AI scale unit (GSU)

The unit of reserved throughput for a model and region, billed at a fixed price per GSU for the chosen term, with lower prices for longer terms

Commitment term

1 week (Google models only), 1 month, 3 months or 1 year; orders cannot be cancelled mid-term and are billed once active

Spillover

Requests that exceed the reserved quota are processed and billed at pay-as-you-go rates by default

GSU changes

Increases apply immediately on approval, while decreases take effect only at auto-renewal for the next term

## How to detect

4 checks to find it in your estate.

- Open the Provisioned Throughput page Utilization summary tab, which shows per model the GSUs owned, peak throughput usage in GSUs, average GSU utilization and how often the limit was reached

- In Cloud Monitoring, compare aiplatform.googleapis.com/publisher/online\_serving/consumed\_token\_throughput filtered to request\_type dedicated with dedicated\_token\_limit or dedicated\_gsu\_limit on the PublisherModel resource

- Flag orders whose peak usage rarely approaches the limit and that show no spillover, and orders with near-zero dedicated traffic because the application moved to another model, version or region

- List active orders with their term, end date and auto-renewal setting, and register Essential Contacts so expiration and auto-renewal notices, sent two weeks ahead for monthly and longer terms, reach the owner

## How to fix

5 ways to remove the waste.

- Decrease GSUs on auto-renewing orders to match the sustained baseline; the reduction applies from the next term, and traffic above it spills over to pay-as-you-go by default

- Turn off auto-renewal for orders that are no longer needed, keeping in mind that an active order cannot be changed in the last five days before expiry unless it auto-renews

- If the application switched models, change the order's model or model version (model changes are limited to the same publisher), or its region, rather than paying for throughput nobody calls

- Consolidate traffic for the same model and region onto one order, and route development or experimental traffic to pay-as-you-go with the X-Vertex-AI-LLM-Request-Type header set to shared so it does not consume reserved capacity

- Use shorter terms while demand is uncertain and size new orders with the GSU estimator and measured traffic, not a launch forecast

## Documentation

Vendor references for pricing and configuration.

- [Provisioned Throughput overview  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/provisioned-throughput)

- [Purchase Provisioned Throughput  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/provisioned-throughput/purchase-provisioned-throughput)

- [Use Provisioned Throughput  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/provisioned-throughput/use-provisioned-throughput)

- [Calculate Provisioned Throughput requirements  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/provisioned-throughput/measure-provisioned-throughput)

- [Agent Platform Pricing  cloud.google.com](https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing)

- [Navigating GenAI Capacity Options  finops.org](https://www.finops.org/wg/genai-capacity-options/)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Vertex AI  CER-0267

### [Using High-Cost Models for Low-Complexity Tasks in Vertex AI](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-models-for-low-complexity-tasks-bec92)

Vertex AI workloads often include low-complexity tasks such as classification, routing, keyword extraction, metadata parsing, document triage, or summarization of short and simple text. These operations do not require the advanced...

AI

- GCP Vertex AI  CER-0556

### [Vertex AI Workbench Instances With Idle Shutdown Disabled](https://www.pointfive.co/efficiency-hub/inefficiencies/vertex-ai-workbench-instances-with-idle-shutdown-disabled)

Vertex AI Workbench instances, now documented as Agent Platform Workbench under the Gemini Enterprise Agent Platform name, are notebook VMs that bill for CPU, memory and any attached GPUs whenever they are running. To help manage costs,...

AI

- GCP Vertex AI  CER-0567

### [Idle Vertex AI Endpoints With Deployed Models](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-vertex-ai-endpoints-with-deployed-models)

When a model is deployed to a Vertex AI endpoint for online inference on dedicated resources (the service is now documented as Agent Platform Inference under the Gemini Enterprise Agent Platform name), each replica is a VM that bills...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

