Explanation
Why the waste happens and who it affects.
A user node pool created with a fixed node count stays sized for its peak even after pods scale in at night, on weekends or after a batch run. When the cluster autoscaler is later disabled on a pool, AKS does not remove nodes, so the pool keeps whatever count it had at that moment.
Enabling the autoscaler is not always enough. Scale-down is governed by a cluster-wide autoscaler profile: by default a node must be unneeded for 10 minutes before removal, scale-down evaluation pauses for 10 minutes after any scale-up, and a node only counts as unneeded when both its summed CPU requests and its summed memory requests are below half of its allocatable capacity. Clusters with frequent small scale-ups, restrictive PodDisruptionBudgets, standalone pods, pods annotated as not safe to evict, or skip-nodes-with-local-storage set to true can keep underused nodes for long periods. Because the profile applies to every autoscaler-enabled pool, one conservative setting chosen for a single workload affects the whole cluster.
Billing model
The pricing dimensions that drive this cost.
AKS charges for the VMs, disks and networking that back the cluster's node pools, plus a per-cluster fee on the Standard and Premium tiers.
- Agent node compute
- Billed for the duration and type of every VM that AKS launches in a node pool, whether or not pods use its capacity
- Fixed node count
- A pool without the cluster autoscaler keeps its configured node count until someone scales it manually
- Scale-down delay
- An underused node keeps billing until it has been unneeded for scale-down-unneeded-time (default 10 minutes), and no node is removed within scale-down-delay-after-add of a scale-up
- Cluster autoscaler
- No extra charge; it adds and removes nodes within each pool's min-count and max-count
How to detect
5 checks to find it in your estate.
- Run an Azure Resource Graph query over microsoft.containerservice/managedclusters, expand properties.agentPoolProfiles and list pools where enableAutoScaling is false or minCount equals maxCount, along with their count and vmSize
- For pools without autoscaling, compare node count over time with summed pod CPU and memory requests (for example in AKS cost analysis) to find pools that stay at peak size during quiet periods
- Review the cluster's autoscalerProfile for raised scale-down-unneeded-time or scale-down-delay-after-add, a lowered scale-down-utilization-threshold, or skip-nodes-with-local-storage set to true, all of which delay scale-down
- Enable the controlplane-cluster-autoscaler metrics target and watch cluster_autoscaler_unneeded_nodes_count, or query the cluster-autoscaler log category, to find nodes marked unneeded that are not removed
- Check Azure Advisor for the AKS cost recommendation 'Fine-tune the cluster autoscaler profile for rapid scale down and cost savings'
How to fix
5 ways to remove the waste.
- Enable the cluster autoscaler on user node pools with az aks nodepool update --enable-cluster-autoscaler and a --min-count and --max-count based on observed demand (a user pool can use --min-count 0 so it may autoscale to zero; system pools always keep running nodes), and stop changing node counts manually on autoscaled pools
- Apply a cost-oriented profile: lower scale-down-unneeded-time and scale-down-delay-after-add, raise scale-down-utilization-threshold, and keep skip-nodes-with-local-storage at its default of false where workloads allow. Microsoft advises against aggressive scale-down for clusters that scale out and in within short intervals, because it can lengthen node provisioning times
- Remove scale-down blockers: standalone pods not backed by a controller, PodDisruptionBudgets that never allow a pod to move, and safe-to-evict false annotations that are no longer needed
- Separate long-running and bursty workloads into different node pools so one profile does not have to serve conflicting patterns
- For new clusters or clusters with many pool shapes, consider node auto-provisioning or AKS Automatic, which select VM sizes from pending pod requests and bin-pack workloads
Documentation
Vendor references for pricing and configuration.
- Use the Cluster Autoscaler in Azure Kubernetes Service (AKS)learn.microsoft.com
- Cluster autoscaling in Azure Kubernetes Service (AKS) overviewlearn.microsoft.com
- Manually Scale Nodes in an Azure Kubernetes Service (AKS) Clusterlearn.microsoft.com
- Best Practices for Cost Optimization in Azure Kubernetes Service (AKS)learn.microsoft.com
- Cost recommendations - Azure Advisorlearn.microsoft.com
- Azure Kubernetes Service (AKS) pricingazure.microsoft.com