Skip to content
Cloud Efficiency Hub

Balanced Cluster Autoscaler Profile on Cost-Sensitive GKE Clusters

The short version

The GKE cluster autoscaler has autoscaling profiles that decide how eagerly it removes nodes.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP GKE
Category
Compute
Reference
CER-0498
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

Standard clusters default to the balanced profile, which Google describes as prioritizing keeping more resources readily available for incoming Pods. The optimize-utilization profile instead prioritizes utilization over spare capacity: the autoscaler removes more nodes, removes them faster, and prefers scheduling Pods onto nodes that are already allocated.

Because balanced is the default and the profile is set per cluster, batch, CI, data processing and development clusters inherit a setting designed to keep headroom for fast Pod start-up. On those clusters, nodes that finished their work linger as underused VMs before the autoscaler reclaims them, and Pods stay spread across more nodes than necessary. Standard node pools bill for every node VM until it is deleted, so slower scale-down shows up directly as Compute Engine spend.

Billing model

The pricing dimensions that drive this cost.

Standard node pool compute
Nodes are billed as Compute Engine instances, per second with a one-minute minimum, until the nodes are deleted
Balanced profile
The default for Standard clusters; keeps more spare capacity so incoming Pods start sooner, at the cost of more billed nodes
Optimize-utilization profile
Scales down more aggressively and packs Pods onto already allocated nodes, reducing billed node-hours

How to detect

5 checks to find it in your estate.

  • Read autoscaling.autoscalingProfile for each Standard cluster (for example with gcloud container clusters describe) and list clusters that are on BALANCED or unset
  • Prioritize clusters that run batch jobs, CI runners, scheduled data processing or development workloads, where Pod start-up latency is not user facing
  • Check the GKE Clusters page Notifications column and the Recommendations page for cost savings hints on allocated versus requested resources
  • Query the container.googleapis.com/cluster-autoscaler-visibility log for scaleDown and noScaleDown events to see how quickly nodes are removed and what blocks removal
  • Compare node allocatable capacity with the sum of Pod requests per node pool over time, using Cloud Monitoring or GKE cost allocation, to find persistently underused nodes

How to fix

5 ways to remove the waste.

  • Switch cost-sensitive and latency-tolerant clusters to the optimize-utilization profile with gcloud container clusters update CLUSTER_NAME --autoscaling-profile optimize-utilization
  • Expect the tradeoff: fewer spare nodes means some new Pods wait for a node to be provisioned, so keep balanced on clusters whose serving workloads need fast scale-up
  • Run batch workloads on dedicated node pools using labels, taints and tolerations so empty nodes can be removed as soon as jobs finish, as Google's GKE cost guidance recommends
  • Move Pods that block scale-down, such as kube-system Pods without a PodDisruptionBudget, Pods without a controller, or Pods annotated safe-to-evict false, onto separate node pools or fix their configuration
  • If a buffer is needed for spikes, use low-priority pause Pods sized for the expected burst instead of relying on the balanced profile to keep idle nodes around

Documentation

Vendor references for pricing and configuration.