Skip to content
Cloud Efficiency Hub

Non-Production AKS Clusters Running Outside Working Hours

The short version

Development, test and sandbox AKS clusters are usually needed only during business hours, but they are created once and left running around the clock.

PointFive Research

Cloud cost research at PointFive

Azure service
Azure AKS
Category
Compute
Reference
CER-0571
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

Outside working hours they often run nothing but system components, yet every node in every node pool keeps billing as a running VM. Scaling user node pools to zero helps, but the System node pool must keep running nodes (at least two, per the AKS scaling guidance) for as long as the cluster is running.

AKS can stop an entire cluster, which stops the control plane and all agent nodes and removes their compute cost while the cluster state is kept for when it starts again. Because the stop and start have to be triggered by an external schedule, and stopped clusters have restrictions (standalone pods are deleted, the API server IP can change, clusters using node auto-provisioning cannot be stopped), many non-production clusters are never put on one. The cost multiplies in organizations that give each team, branch or test run its own cluster.

Billing model

The pricing dimensions that drive this cost.

Agent node compute
Billed per VM for the duration each node in a system or user pool runs
System node pool
Must keep running nodes while the cluster runs, so scaling user pools to zero still leaves a billed baseline
Stopped cluster
Control plane and agent nodes are stopped, saving the compute costs while objects other than standalone pods are preserved
Stored cluster state
Kept for up to 12 months while stopped; after that the state cannot be recovered

How to detect

5 checks to find it in your estate.

  • List clusters by environment tag or naming convention and check powerState (az aks show or Azure Resource Graph); non-production clusters that always report Running are candidates
  • Compare node counts across nights and weekends with business hours; flat 24x7 node counts on dev or test clusters indicate no stop schedule and no scale-in
  • Check API server and workload activity (for example kube-audit logs or deployment activity) outside working hours to confirm the cluster is not used then
  • Confirm whether any automation already runs az aks stop and az aks start (runbook, pipeline or Logic App) for each candidate cluster
  • Exclude clusters that use node auto-provisioning, which cannot be stopped, and route them to user pool scale-in instead

How to fix

5 ways to remove the waste.

  • Schedule az aks stop at the end of the working day and az aks start before it begins, using Azure Automation, a pipeline or a Logic App; wait 15-30 minutes after a stop before starting again
  • Before adopting stop and start, move any standalone pods into Deployments or Jobs, since stopping drains all nodes and deletes standalone pods, and narrow admission webhooks with wildcard rules that can make the stop operation fail
  • Expect the API server IP address to change on start, and for private clusters verify where AKS recreates the API server private endpoint
  • Where a full stop is not possible (for example clusters using node auto-provisioning), scale user node pools to zero off-hours with az aks nodepool scale --node-count 0 (the autoscaler must be disabled on that pool first), or set an autoscaler min-count of 0, which allows but does not force scale-in to zero
  • Do not stop mission-critical clusters: Microsoft warns that in capacity-constrained regions a stopped cluster may not be able to start again

Documentation

Vendor references for pricing and configuration.