Skip to content
Cloud Efficiency Hub

Underuse of Spot VMs for Fault-Tolerant GKE Workloads

The short version

Batch jobs, CI runners, rendering, data processing and other restart-tolerant workloads on GKE frequently run on standard node pools or standard Autopilot Pods at regular prices.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP GKE
Category
Compute
Reference
CER-0499
Type
Suboptimal Pricing Model

Explanation

Why the waste happens and who it affects.

GKE supports Spot VMs in Standard node pools and Spot Pods in Autopilot, which run the same machine types at a deep discount in exchange for no availability guarantee: Compute Engine can preempt the capacity at any time.

Node pools are often created once from a template and never split by workload type, and a blanket no-Spot rule adopted for stateful or latency-sensitive services ends up applied to every workload. The result is fault-tolerant work that could absorb a preemption and retry paying full price. Google's GKE cost guidance names Spot as a lever for batch and fault-tolerant jobs, while warning against it for stateful and serving workloads that are not designed for it.

Billing model

The pricing dimensions that drive this cost.

Standard node pool VMs
Billed at Compute Engine on-demand rates for every node until it is deleted
Spot VMs
The same machine types billed at variable Spot prices that Google states are discounted up to 91 percent, with no availability guarantee
Autopilot Spot Pods
Pod requests billed at Spot rates that the GKE pricing page lists alongside regular Autopilot vCPU and memory prices

How to detect

4 checks to find it in your estate.

  • List node pools and check whether they use Spot (the node pool spot setting, or the cloud.google.com/gke-spot=true label on nodes); flag clusters where no pool uses Spot
  • Identify Jobs, CronJobs, CI runners and stateless batch Deployments scheduled on standard capacity, and in Autopilot, Pods that do not request cloud.google.com/gke-spot
  • Confirm each candidate tolerates interruption: it handles SIGTERM within the preemption grace period, checkpoints or is idempotent, and has retries configured
  • Use GKE cost allocation or usage metering to size spend by namespace and label for batch and non-production workloads, which are the usual Spot candidates

How to fix

5 ways to remove the waste.

  • Create Spot node pools with the --spot flag in Standard clusters, taint them, and add matching tolerations and node selectors or affinity only to fault-tolerant workloads
  • In Autopilot, request Spot Pods with a nodeSelector or node affinity on cloud.google.com/gke-spot: "true"; Autopilot provisions Spot capacity and adds the taints and tolerations
  • Keep a standard node pool as fallback so work continues when Spot capacity is reclaimed or unavailable, as Google recommends mixing Spot and standard node pools
  • Design for the short shutdown window: GKE gives Spot VM nodes a 30-second graceful termination period by default (15 seconds for user Pods) and Spot Pods a maximum 15-second grace period
  • Leave stateful, serving and other critical workloads on standard capacity unless they are explicitly built to survive preemption, because data on Spot nodes is deleted on preemption and PodDisruptionBudgets might not be respected

Documentation

Vendor references for pricing and configuration.