# Idle Vertex AI Endpoints With Deployed Models

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/idle-vertex-ai-endpoints-with-deployed-models

When a model is deployed to a Vertex AI endpoint for online inference on dedicated resources (the service is now documented as Agent Platform Inference...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

When a model is deployed to a Vertex AI endpoint for online inference on dedicated resources (the service is now documented as Agent Platform Inference under the Gemini Enterprise Agent Platform name), each replica is a VM that bills node-hours while it waits for requests, not only while it serves them.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Vertex AI](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-vertex-ai)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0567

Type

Idle or Unused Resource

## Explanation

Why the waste happens and who it affects.

A deployment must keep at least one replica by default, so a model that receives no traffic still pays for at least one node, and for its GPUs if it has them, around the clock.

Idle deployments accumulate from experiments, demos, A/B tests whose losing variant was never removed, older model versions left behind after a new one took the traffic, and staging endpoints copied from production sizing. Because endpoints are managed resources rather than VMs in the Compute Engine console, they are easy to overlook. Google's AI and ML cost guidance asks teams to find idle or underutilized VMs and GPUs and shut them down or rightsize them, and the pricing page defines a node hour as including time spent waiting in an active state on an endpoint with a model deployed.

## Billing model

The pricing dimensions that drive this cost.

Online inference on dedicated resources is billed per node-hour for each replica.

Node hour

The time a VM spends running inference or waiting in an active state on an endpoint with one or more models deployed, charged in 30-second increments

Minimum replicas

DedicatedResources.minReplicaCount must be at least 1 unless the Scale To Zero Preview feature is used, so an idle deployment bills at least one node

Management fees

Agent Platform Inference adds management fees on top of the underlying machine and accelerator cost

Undeployed models

Models that are not deployed or failed to deploy are not charged; empty endpoints do not bill node-hours

## How to detect

4 checks to find it in your estate.

- List endpoints and their deployed models with gcloud ai endpoints list and describe, recording machine type, accelerator type and count, and minReplicaCount for each deployment

- Chart aiplatform.googleapis.com/prediction/online/request\_count per deployed model over 7 to 30 days and flag deployments with no or negligible traffic

- Check aiplatform.googleapis.com/prediction/online/cpu/utilization and accelerator/duty\_cycle for deployments that receive some traffic but keep more replicas or larger machines than they use

- Look for endpoints with several deployed models where the traffic split sends 0 percent to one or more of them, which usually marks superseded versions

## How to fix

5 ways to remove the waste.

- Undeploy models that no longer serve traffic and delete empty endpoints; charges stop only when the model is undeployed

- For deployments with long daily or weekly idle periods, consider Scale To Zero (Preview) by setting min\_replica\_count to 0; it is not available on shared public endpoints or multi-host deployments, the first request after scale-down gets a 429 while replicas start, and models scaled to zero for more than 30 days are undeployed automatically

- Lower minReplicaCount and tune autoscaling targets on low-traffic deployments with mutateDeployedModel, which changes replica settings without redeploying

- Move non-interactive scoring to batch inference, which runs only for the duration of the job, and co-host small models on shared resources where supported

- Tag endpoints with an owner and purpose and review deployments without traffic on a regular schedule

## Documentation

Vendor references for pricing and configuration.

- [Scale inference nodes by using autoscaling  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/machine-learning/predictions/autoscaling)

- [Gemini Enterprise Agent Platform pricing  cloud.google.com](https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing)

- [AI and ML perspective: Cost optimization  docs.cloud.google.com](https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/cost-optimization)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Vertex AI  CER-0267

### [Using High-Cost Models for Low-Complexity Tasks in Vertex AI](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-models-for-low-complexity-tasks-bec92)

Vertex AI workloads often include low-complexity tasks such as classification, routing, keyword extraction, metadata parsing, document triage, or summarization of short and simple text. These operations do not require the advanced...

AI

- GCP Vertex AI  CER-0512

### [Underutilized Vertex AI Provisioned Throughput](https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-vertex-ai-provisioned-throughput)

Provisioned Throughput on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, is a fixed-cost, fixed-term subscription that reserves throughput for a specific generative AI model in a specific region. It is bought in...

AI

- GCP Vertex AI  CER-0556

### [Vertex AI Workbench Instances With Idle Shutdown Disabled](https://www.pointfive.co/efficiency-hub/inefficiencies/vertex-ai-workbench-instances-with-idle-shutdown-disabled)

Vertex AI Workbench instances, now documented as Agent Platform Workbench under the Gemini Enterprise Agent Platform name, are notebook VMs that bill for CPU, memory and any attached GPUs whenever they are running. To help manage costs,...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

