Explanation
Why the waste happens and who it affects.
AKS supports Spot node pools backed by Azure Spot Virtual Machine Scale Sets, which use unused Azure capacity at a significant discount but can be evicted when Azure needs the capacity back.
Spot has to be opted into deliberately. A Spot pool can only be a secondary pool, its priority and max price cannot be changed after creation, and AKS applies a NoSchedule taint to Spot nodes, so workloads only land there when teams add a matching toleration and node affinity. Without that step, interruptible work keeps running at full price. The opportunity is limited to workloads that are stateless, retry-safe or checkpointed; anything that needs an SLA should stay on regular nodes.
Billing model
The pricing dimensions that drive this cost.
- Regular node pool
- Nodes billed at pay-as-you-go VM rates (or covered by reservations or savings plans) for as long as they run
- Spot node pool
- Nodes billed at variable Spot prices by region and VM size, which Microsoft cites as up to 90% below pay-as-you-go
- Spot max price
- With -1 nodes are not evicted on price and pay the lower of the current Spot or standard price
- Eviction
- Spot nodes have no SLA and are removed when Azure needs capacity; with the Delete policy the nodes are deleted
How to detect
4 checks to find it in your estate.
- Run the FinOps toolkit Azure Resource Graph query 'AKS clusters without Spot VMs', which lists agent pools where enableAutoScaling is true and scaleSetPriority is null
- Check Azure Advisor for the AKS cost recommendation 'Consider Spot nodes for workloads that can handle interruptions'
- Inventory workloads on regular pools that are Jobs, CronJobs, CI runners, queue workers or other controllers that can be rescheduled, and exclude StatefulSets and latency-critical services
- Use AKS cost analysis to size the spend on node pools that host those interruption-tolerant workloads
How to fix
5 ways to remove the waste.
- Add a secondary Spot node pool with az aks nodepool add --priority Spot --eviction-policy Delete --spot-max-price -1 and enable the cluster autoscaler on it, which replaces evicted nodes when capacity returns
- Add a toleration for kubernetes.azure.com/scalesetpriority=spot:NoSchedule and node affinity on the kubernetes.azure.com/scalesetpriority=spot label to the workloads that should move
- Keep a regular node pool available as fallback for capacity shortages, and use a priority expander in the autoscaler so Spot pools are tried first when both can serve a pod
- Make moved workloads eviction-safe with graceful shutdown handling, retries or checkpointing and appropriate PodDisruptionBudgets
- Prefer the Delete eviction policy: nodes left in stopped-deallocated state under the Deallocate policy count against compute quota and can interfere with scaling and upgrades
Documentation
Vendor references for pricing and configuration.
- Add an Azure Spot node pool to an Azure Kubernetes Service (AKS) clusterlearn.microsoft.com
- Best Practices for Cost Optimization in Azure Kubernetes Service (AKS)learn.microsoft.com
- FinOps best practices for computelearn.microsoft.com
- Cost recommendations - Azure Advisorlearn.microsoft.com
- Cluster autoscaling in Azure Kubernetes Service (AKS) overviewlearn.microsoft.com
- Linux Virtual Machine Scale Sets pricingazure.microsoft.com