Explanation
Why the waste happens and who it affects.
The Kubernetes scheduler places pods according to their CPU and memory requests, and the cluster autoscaler adds nodes when pending pods cannot fit on existing ones. When requests are set well above real consumption, nodes look full on paper while their CPUs and memory sit mostly idle, and the cluster keeps more nodes than the workload needs.
Inflated requests are common because they are copied from templates or Helm chart defaults, set generously during an incident and never lowered, or sized for a peak that rarely happens. The same inflation also blocks scale-down: the AKS cluster autoscaler only considers a node for removal when the sum of requests on it falls below its scale-down utilization threshold. Microsoft AKS cost guidance states that requests and limits higher than actual usage result in overprovisioned workloads and wasted resources, and Azure Advisor has a dedicated recommendation to enable Vertical Pod Autoscaler recommendation mode to rightsize them.
Billing model
The pricing dimensions that drive this cost.
AKS costs come from the node pools, whose size is driven by requested rather than used resources.
- Node VM hours
- Each node in a node pool is billed as a virtual machine for as long as it runs, whether its capacity is used or not
- Request-based scheduling
- Pods are placed on nodes by their CPU and memory requests, so reserved but unused capacity still occupies billed nodes
- Scale-down threshold
- The cluster autoscaler considers a node for removal only when the larger of its summed CPU or memory requests divided by allocatable is below scale-down-utilization-threshold (default 0.5)
How to detect
5 checks to find it in your estate.
- Review Azure Advisor cost recommendations for AKS, including Enable Vertical Pod Autoscaler recommendation mode to rightsize resource requests and limits
- Deploy VPA objects with updateMode Off for major workloads and compare the recommended target requests with the requests in the deployment manifests; the VPA recommender keeps up to eight days of history, so review recommendations after a representative period
- Compare actual container CPU and memory usage from Container insights or managed Prometheus (or kubectl top pods) with configured requests over several weeks, looking for workloads whose usage stays far below their requests
- Check node-level allocation with kubectl describe node: nodes whose allocated requests are high while measured utilization is low indicate request inflation rather than real demand
- Enable the AKS cost analysis add-on (Standard or Premium tier) to see cost by namespace and idle charges, and prioritize the namespaces with the largest spend for rightsizing
How to fix
5 ways to remove the waste.
- Lower CPU and memory requests toward VPA or observed usage, keeping headroom for peaks, and roll the changes out through the normal deployment process
- Where workloads tolerate it, let VPA apply recommendations using Initial, Recreate or, on AKS 1.34 and later, InPlaceOrRecreate mode; do not combine VPA with an HPA that scales on the same CPU or memory metrics, and note that VPA does not support JVM-based workloads or Windows containers
- Use HPA or KEDA to add replicas for load instead of sizing every replica for peak demand
- Set namespace resource quotas and LimitRanges, or use deployment safeguards, so new workloads start with reasonable requests
- After requests come down, confirm that the cluster autoscaler actually removes the freed nodes, and review node pool minimum counts that could keep them running
Documentation
Vendor references for pricing and configuration.
- Best Practices for Cost Optimization in Azure Kubernetes Service (AKS)learn.microsoft.com
- Vertical Pod Autoscaling in Azure Kubernetes Service (AKS)learn.microsoft.com
- Cost recommendations - Azure Advisorlearn.microsoft.com
- Use the Cluster Autoscaler in Azure Kubernetes Service (AKS)learn.microsoft.com
- Resource management best practices for Azure Kubernetes Service (AKS)learn.microsoft.com
- Azure Kubernetes Service (AKS) cost analysislearn.microsoft.com