# Fault-Tolerant AKS Workloads Not Using Spot Node Pools

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/fault-tolerant-aks-workloads-not-using-spot-node-pools

Many AKS clusters run batch jobs, CI runners, queue consumers, rendering and dev or test services on regular pay-as-you-go node pools, even though...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Many AKS clusters run batch jobs, CI runners, queue consumers, rendering and dev or test services on regular pay-as-you-go node pools, even though these workloads can tolerate a node disappearing.

PointFive Research

Cloud cost research at PointFive

Azure service

[Azure AKS](https://www.pointfive.co/efficiency-hub/cloud-services/azure-aks)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0572

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

AKS supports Spot node pools backed by Azure Spot Virtual Machine Scale Sets, which use unused Azure capacity at a significant discount but can be evicted when Azure needs the capacity back.

Spot has to be opted into deliberately. A Spot pool can only be a secondary pool, its priority and max price cannot be changed after creation, and AKS applies a NoSchedule taint to Spot nodes, so workloads only land there when teams add a matching toleration and node affinity. Without that step, interruptible work keeps running at full price. The opportunity is limited to workloads that are stateless, retry-safe or checkpointed; anything that needs an SLA should stay on regular nodes.

## Billing model

The pricing dimensions that drive this cost.

Regular node pool

Nodes billed at pay-as-you-go VM rates (or covered by reservations or savings plans) for as long as they run

Spot node pool

Nodes billed at variable Spot prices by region and VM size, which Microsoft cites as up to 90% below pay-as-you-go

Spot max price

With -1 nodes are not evicted on price and pay the lower of the current Spot or standard price

Eviction

Spot nodes have no SLA and are removed when Azure needs capacity; with the Delete policy the nodes are deleted

## How to detect

4 checks to find it in your estate.

- Run the FinOps toolkit Azure Resource Graph query 'AKS clusters without Spot VMs', which lists agent pools where enableAutoScaling is true and scaleSetPriority is null

- Check Azure Advisor for the AKS cost recommendation 'Consider Spot nodes for workloads that can handle interruptions'

- Inventory workloads on regular pools that are Jobs, CronJobs, CI runners, queue workers or other controllers that can be rescheduled, and exclude StatefulSets and latency-critical services

- Use AKS cost analysis to size the spend on node pools that host those interruption-tolerant workloads

## How to fix

5 ways to remove the waste.

- Add a secondary Spot node pool with az aks nodepool add --priority Spot --eviction-policy Delete --spot-max-price -1 and enable the cluster autoscaler on it, which replaces evicted nodes when capacity returns

- Add a toleration for kubernetes.azure.com/scalesetpriority=spot:NoSchedule and node affinity on the kubernetes.azure.com/scalesetpriority=spot label to the workloads that should move

- Keep a regular node pool available as fallback for capacity shortages, and use a priority expander in the autoscaler so Spot pools are tried first when both can serve a pod

- Make moved workloads eviction-safe with graceful shutdown handling, retries or checkpointing and appropriate PodDisruptionBudgets

- Prefer the Delete eviction policy: nodes left in stopped-deallocated state under the Deallocate policy count against compute quota and can interfere with scaling and upgrades

## Documentation

Vendor references for pricing and configuration.

- [Add an Azure Spot node pool to an Azure Kubernetes Service (AKS) cluster  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/aks/spot-node-pool)

- [Best Practices for Cost Optimization in Azure Kubernetes Service (AKS)  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/aks/best-practices-cost)

- [FinOps best practices for compute  learn.microsoft.com](https://learn.microsoft.com/en-us/cloud-computing/finops/best-practices/compute)

- [Cost recommendations - Azure Advisor  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/advisor/advisor-reference-cost-recommendations)

- [Cluster autoscaling in Azure Kubernetes Service (AKS) overview  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/aks/cluster-autoscaler-overview)

- [Linux Virtual Machine Scale Sets pricing  azure.microsoft.com](https://azure.microsoft.com/en-us/pricing/details/virtual-machine-scale-sets/linux/)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Azure AKS  CER-0099

### [Orphaned Kubernetes Resources in AKS](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-kubernetes-resources-230dd)

Kubernetes environments often accumulate unused resources over time as applications evolve. Common examples include Persistent Volume Claims (PVCs) backed by Azure Disks, Services that trigger load balancer provisioning, or stale...

Compute

- Azure AKS  CER-0100

### [Orphaned and Overprovisioned Resources in AKS Clusters](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-and-overprovisioned-resources-in-aks-clusters)

Clusters often accumulate unused components when applications are terminated or environments are cloned. These include PVCs backed by Managed Disks, Services that still front Azure Load Balancers, and test namespaces that are no longer...

Compute

- Azure AKS  CER-0557

### [Overprovisioned Pod Resource Requests in AKS](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-pod-resource-requests-in-aks)

In AKS, the node VMs are what you pay for, and how many nodes a cluster needs is decided by pod resource requests, not by what pods actually use. The Kubernetes scheduler places pods according to their CPU and memory requests, and the...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

