Explanation
Why the waste happens and who it affects.
A deployment must keep at least one replica by default, so a model that receives no traffic still pays for at least one node, and for its GPUs if it has them, around the clock.
Idle deployments accumulate from experiments, demos, A/B tests whose losing variant was never removed, older model versions left behind after a new one took the traffic, and staging endpoints copied from production sizing. Because endpoints are managed resources rather than VMs in the Compute Engine console, they are easy to overlook. Google's AI and ML cost guidance asks teams to find idle or underutilized VMs and GPUs and shut them down or rightsize them, and the pricing page defines a node hour as including time spent waiting in an active state on an endpoint with a model deployed.
Billing model
The pricing dimensions that drive this cost.
Online inference on dedicated resources is billed per node-hour for each replica.
- Node hour
- The time a VM spends running inference or waiting in an active state on an endpoint with one or more models deployed, charged in 30-second increments
- Minimum replicas
- DedicatedResources.minReplicaCount must be at least 1 unless the Scale To Zero Preview feature is used, so an idle deployment bills at least one node
- Management fees
- Agent Platform Inference adds management fees on top of the underlying machine and accelerator cost
- Undeployed models
- Models that are not deployed or failed to deploy are not charged; empty endpoints do not bill node-hours
How to detect
4 checks to find it in your estate.
- List endpoints and their deployed models with gcloud ai endpoints list and describe, recording machine type, accelerator type and count, and minReplicaCount for each deployment
- Chart aiplatform.googleapis.com/prediction/online/request_count per deployed model over 7 to 30 days and flag deployments with no or negligible traffic
- Check aiplatform.googleapis.com/prediction/online/cpu/utilization and accelerator/duty_cycle for deployments that receive some traffic but keep more replicas or larger machines than they use
- Look for endpoints with several deployed models where the traffic split sends 0 percent to one or more of them, which usually marks superseded versions
How to fix
5 ways to remove the waste.
- Undeploy models that no longer serve traffic and delete empty endpoints; charges stop only when the model is undeployed
- For deployments with long daily or weekly idle periods, consider Scale To Zero (Preview) by setting min_replica_count to 0; it is not available on shared public endpoints or multi-host deployments, the first request after scale-down gets a 429 while replicas start, and models scaled to zero for more than 30 days are undeployed automatically
- Lower minReplicaCount and tune autoscaling targets on low-traffic deployments with mutateDeployedModel, which changes replica settings without redeploying
- Move non-interactive scoring to batch inference, which runs only for the duration of the job, and co-host small models on shared resources where supported
- Tag endpoints with an owner and purpose and review deployments without traffic on a regular schedule
Documentation
Vendor references for pricing and configuration.