# Underuse of Spot VMs for Fault-Tolerant GKE Workloads

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/underuse-of-spot-vms-for-fault-tolerant-gke-workloads

Batch jobs, CI runners, rendering, data processing and other restart-tolerant workloads on GKE frequently run on standard node pools or standard...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Batch jobs, CI runners, rendering, data processing and other restart-tolerant workloads on GKE frequently run on standard node pools or standard Autopilot Pods at regular prices.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP GKE](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-gke)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0499

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

GKE supports Spot VMs in Standard node pools and Spot Pods in Autopilot, which run the same machine types at a deep discount in exchange for no availability guarantee: Compute Engine can preempt the capacity at any time.

Node pools are often created once from a template and never split by workload type, and a blanket no-Spot rule adopted for stateful or latency-sensitive services ends up applied to every workload. The result is fault-tolerant work that could absorb a preemption and retry paying full price. Google's GKE cost guidance names Spot as a lever for batch and fault-tolerant jobs, while warning against it for stateful and serving workloads that are not designed for it.

## Billing model

The pricing dimensions that drive this cost.

Standard node pool VMs

Billed at Compute Engine on-demand rates for every node until it is deleted

Spot VMs

The same machine types billed at variable Spot prices that Google states are discounted up to 91 percent, with no availability guarantee

Autopilot Spot Pods

Pod requests billed at Spot rates that the GKE pricing page lists alongside regular Autopilot vCPU and memory prices

## How to detect

4 checks to find it in your estate.

- List node pools and check whether they use Spot (the node pool spot setting, or the cloud.google.com/gke-spot=true label on nodes); flag clusters where no pool uses Spot

- Identify Jobs, CronJobs, CI runners and stateless batch Deployments scheduled on standard capacity, and in Autopilot, Pods that do not request cloud.google.com/gke-spot

- Confirm each candidate tolerates interruption: it handles SIGTERM within the preemption grace period, checkpoints or is idempotent, and has retries configured

- Use GKE cost allocation or usage metering to size spend by namespace and label for batch and non-production workloads, which are the usual Spot candidates

## How to fix

5 ways to remove the waste.

- Create Spot node pools with the --spot flag in Standard clusters, taint them, and add matching tolerations and node selectors or affinity only to fault-tolerant workloads

- In Autopilot, request Spot Pods with a nodeSelector or node affinity on cloud.google.com/gke-spot: "true"; Autopilot provisions Spot capacity and adds the taints and tolerations

- Keep a standard node pool as fallback so work continues when Spot capacity is reclaimed or unavailable, as Google recommends mixing Spot and standard node pools

- Design for the short shutdown window: GKE gives Spot VM nodes a 30-second graceful termination period by default (15 seconds for user Pods) and Spot Pods a maximum 15-second grace period

- Leave stateful, serving and other critical workloads on standard capacity unless they are explicitly built to survive preemption, because data on Spot nodes is deleted on preemption and PodDisruptionBudgets might not be respected

## Documentation

Vendor references for pricing and configuration.

- [Spot VMs  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/spot-vms)

- [Run fault-tolerant workloads at lower costs in Spot Pods  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/autopilot-spot-pods)

- [Best practices for running cost-optimized Kubernetes applications on GKE  docs.cloud.google.com](https://docs.cloud.google.com/architecture/best-practices-for-running-cost-effective-kubernetes-applications-on-gke)

- [Pricing  cloud.google.com](https://cloud.google.com/spot-vms/pricing)

- [Google Kubernetes Engine pricing  cloud.google.com](https://cloud.google.com/kubernetes-engine/pricing)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP GKE  CER-0287

### [Spot-Only GKE Capacity Without Standard Fallback](https://www.pointfive.co/efficiency-hub/inefficiencies/spot-only-gke-capacity-without-standard-fallback)

Workloads are constrained to run only on Spot-based capacity with no viable path to standard nodes when Spot capacity is reclaimed or unavailable. While Spot reduces unit cost, rigid dependence can create hidden costs by requiring standby...

Compute

- GCP GKE  CER-0193

### [Orphaned and Overprovisioned Resources in GKE Clusters](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-and-overprovisioned-resources-in-gke-clusters)

As environments scale, GKE clusters tend to accumulate artifacts from ephemeral workloads, dev environments, or incomplete job execution. PVCs can continue to retain Persistent Disks, Services may continue to expose public IPs and...

Compute

- GCP GKE  CER-0270

### [Orphaned Kubernetes Resources in GKE](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-kubernetes-resources)

In GKE environments, it is common for unused Kubernetes resources to accumulate over time. Examples include Persistent Volume Claims (PVCs) that retain provisioned Persistent Disks, or Services of type LoadBalancer that continue to front...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

