Skip to content
Cloud Efficiency Hub

Idle Vertex AI Endpoints With Deployed Models

The short version

When a model is deployed to a Vertex AI endpoint for online inference on dedicated resources (the service is now documented as Agent Platform Inference under the Gemini Enterprise Agent Platform name), each replica is a VM that bills node-hours while it waits for requests, not only while it serves them.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP Vertex AI
Category
AI
Reference
CER-0567
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

A deployment must keep at least one replica by default, so a model that receives no traffic still pays for at least one node, and for its GPUs if it has them, around the clock.

Idle deployments accumulate from experiments, demos, A/B tests whose losing variant was never removed, older model versions left behind after a new one took the traffic, and staging endpoints copied from production sizing. Because endpoints are managed resources rather than VMs in the Compute Engine console, they are easy to overlook. Google's AI and ML cost guidance asks teams to find idle or underutilized VMs and GPUs and shut them down or rightsize them, and the pricing page defines a node hour as including time spent waiting in an active state on an endpoint with a model deployed.

Billing model

The pricing dimensions that drive this cost.

Online inference on dedicated resources is billed per node-hour for each replica.

Node hour
The time a VM spends running inference or waiting in an active state on an endpoint with one or more models deployed, charged in 30-second increments
Minimum replicas
DedicatedResources.minReplicaCount must be at least 1 unless the Scale To Zero Preview feature is used, so an idle deployment bills at least one node
Management fees
Agent Platform Inference adds management fees on top of the underlying machine and accelerator cost
Undeployed models
Models that are not deployed or failed to deploy are not charged; empty endpoints do not bill node-hours

How to detect

4 checks to find it in your estate.

  • List endpoints and their deployed models with gcloud ai endpoints list and describe, recording machine type, accelerator type and count, and minReplicaCount for each deployment
  • Chart aiplatform.googleapis.com/prediction/online/request_count per deployed model over 7 to 30 days and flag deployments with no or negligible traffic
  • Check aiplatform.googleapis.com/prediction/online/cpu/utilization and accelerator/duty_cycle for deployments that receive some traffic but keep more replicas or larger machines than they use
  • Look for endpoints with several deployed models where the traffic split sends 0 percent to one or more of them, which usually marks superseded versions

How to fix

5 ways to remove the waste.

  • Undeploy models that no longer serve traffic and delete empty endpoints; charges stop only when the model is undeployed
  • For deployments with long daily or weekly idle periods, consider Scale To Zero (Preview) by setting min_replica_count to 0; it is not available on shared public endpoints or multi-host deployments, the first request after scale-down gets a 429 while replicas start, and models scaled to zero for more than 30 days are undeployed automatically
  • Lower minReplicaCount and tune autoscaling targets on low-traffic deployments with mutateDeployedModel, which changes replica settings without redeploying
  • Move non-interactive scoring to batch inference, which runs only for the duration of the job, and co-host small models on shared resources where supported
  • Tag endpoints with an owner and purpose and review deployments without traffic on a regular schedule

Documentation

Vendor references for pricing and configuration.