Explanation
Why the waste happens and who it affects.
Even with no customer workloads, every cluster pays the flat hourly cluster management fee, and a Standard cluster also pays for every node VM in its node pools, plus any persistent disks and load balancers the workloads created.
Idle clusters are easy to miss because the cost is spread across Compute Engine, Kubernetes Engine and networking SKUs rather than showing up as one line item, and clusters without clear owner labels tend to be left alone. GKE's Idle GKE cluster recommender (google.container.DiagnosisRecommender) flags clusters that show no meaningful activity for 30 days and recommends deleting them.
Billing model
The pricing dimensions that drive this cost.
- Cluster management fee
- A flat per-cluster hourly fee for every GKE cluster in any mode; a monthly free tier credit per billing account covers one zonal Standard or Autopilot cluster
- Standard node pools
- Node VMs are billed at Compute Engine rates until the nodes are deleted, whether or not Pods are running
- Autopilot Pods
- Billed on Pod resource requests, so an empty Autopilot cluster mainly pays the management fee
- Attached resources
- Persistent disks and load balancers created by workloads bill separately and can outlive the cluster
How to detect
5 checks to find it in your estate.
- Review idle cluster insights and recommendations on the Clusters page Cost Optimization tab, with gcloud, or through the Recommender API (recommender google.container.DiagnosisRecommender, recommendation subtype CLUSTER_IDLE)
- CLUSTER_IDLE_NO_RUNNING_PODS: zero Pods in the Running state outside the kube-system and gmp-system namespaces over 30 days
- CLUSTER_IDLE_NO_NODES: zero nodes or zero node pools over 30 days, where the cluster pays only the management fee
- CLUSTER_IDLE_LOW_CPU_UTILIZATION: CPU utilization averages under 7 percent in every hour over 30 days while the active Pod count stays unchanged
- For clusters younger than 30 days, which GKE does not evaluate, check owner labels, creation purpose and last deployment time
How to fix
5 ways to remove the waste.
- Confirm with the owner that the cluster is not intentionally idle, for example as failover capacity, before acting on the recommendation
- Export any manifests, configuration or data that must be kept, then delete the cluster
- Delete Services of type LoadBalancer before deleting the cluster, since GKE might not remove every load balancer resource for clusters with many Services
- After deletion, find and remove persistent disks that are no longer needed, because GKE retains persistent disk volumes during cluster deletion
- For clusters that must stay but are rarely used, scale Standard node pools down or run the workloads in Autopilot, where empty clusters do not pay for idle nodes
Documentation
Vendor references for pricing and configuration.