Skip to content
Cloud Efficiency Hub

Overprovisioned Pod Resource Requests in AKS

The short version

In AKS, the node VMs are what you pay for, and how many nodes a cluster needs is decided by pod resource requests, not by what pods actually use.

PointFive Research

Cloud cost research at PointFive

Azure service
Azure AKS
Category
Compute
Reference
CER-0557
Type
Overprovisioned Resource

Explanation

Why the waste happens and who it affects.

The Kubernetes scheduler places pods according to their CPU and memory requests, and the cluster autoscaler adds nodes when pending pods cannot fit on existing ones. When requests are set well above real consumption, nodes look full on paper while their CPUs and memory sit mostly idle, and the cluster keeps more nodes than the workload needs.

Inflated requests are common because they are copied from templates or Helm chart defaults, set generously during an incident and never lowered, or sized for a peak that rarely happens. The same inflation also blocks scale-down: the AKS cluster autoscaler only considers a node for removal when the sum of requests on it falls below its scale-down utilization threshold. Microsoft AKS cost guidance states that requests and limits higher than actual usage result in overprovisioned workloads and wasted resources, and Azure Advisor has a dedicated recommendation to enable Vertical Pod Autoscaler recommendation mode to rightsize them.

Billing model

The pricing dimensions that drive this cost.

AKS costs come from the node pools, whose size is driven by requested rather than used resources.

Node VM hours
Each node in a node pool is billed as a virtual machine for as long as it runs, whether its capacity is used or not
Request-based scheduling
Pods are placed on nodes by their CPU and memory requests, so reserved but unused capacity still occupies billed nodes
Scale-down threshold
The cluster autoscaler considers a node for removal only when the larger of its summed CPU or memory requests divided by allocatable is below scale-down-utilization-threshold (default 0.5)

How to detect

5 checks to find it in your estate.

  • Review Azure Advisor cost recommendations for AKS, including Enable Vertical Pod Autoscaler recommendation mode to rightsize resource requests and limits
  • Deploy VPA objects with updateMode Off for major workloads and compare the recommended target requests with the requests in the deployment manifests; the VPA recommender keeps up to eight days of history, so review recommendations after a representative period
  • Compare actual container CPU and memory usage from Container insights or managed Prometheus (or kubectl top pods) with configured requests over several weeks, looking for workloads whose usage stays far below their requests
  • Check node-level allocation with kubectl describe node: nodes whose allocated requests are high while measured utilization is low indicate request inflation rather than real demand
  • Enable the AKS cost analysis add-on (Standard or Premium tier) to see cost by namespace and idle charges, and prioritize the namespaces with the largest spend for rightsizing

How to fix

5 ways to remove the waste.

  • Lower CPU and memory requests toward VPA or observed usage, keeping headroom for peaks, and roll the changes out through the normal deployment process
  • Where workloads tolerate it, let VPA apply recommendations using Initial, Recreate or, on AKS 1.34 and later, InPlaceOrRecreate mode; do not combine VPA with an HPA that scales on the same CPU or memory metrics, and note that VPA does not support JVM-based workloads or Windows containers
  • Use HPA or KEDA to add replicas for load instead of sizing every replica for peak demand
  • Set namespace resource quotas and LimitRanges, or use deployment safeguards, so new workloads start with reasonable requests
  • After requests come down, confirm that the cluster autoscaler actually removes the freed nodes, and review node pool minimum counts that could keep them running

Documentation

Vendor references for pricing and configuration.