Skip to content
Cloud Efficiency Hub

Missed Use of Spot VMs for Fault-Tolerant Compute Engine Workloads

The short version

Spot VMs run the same machine types as standard VMs at discounts of up to 91% for many machine types, GPUs, TPUs and Local SSDs, in exchange for no availability guarantee: Compute Engine can preempt them at any time.

PointFive Research

Cloud cost research at PointFive

Category
Compute
Reference
CER-0468
Type
Suboptimal Pricing Model

Explanation

Why the waste happens and who it affects.

Batch processing, CI and build runners, rendering, simulations, queue workers, and other stateless or checkpointed jobs tolerate that well. Yet these workloads often run on the standard provisioning model because the instance template was copied from a production service, or because nobody revisited the choice after the workload became automated.

Google's Well-Architected cost guidance recommends Spot VMs for non-critical or fault-tolerant workloads. Work that can be retried or resumed pays full on-demand rates on the standard model when it could be discounted by up to 91%, and large autoscaled worker fleets in managed instance groups can absorb preemption because the group recreates preempted VMs.

Billing model

The pricing dimensions that drive this cost.

Standard provisioning model
VCPU, memory, GPUs and Local SSD billed at on-demand rates, eligible for sustained use and committed use discounts
Spot VMs
Billed at variable Spot prices, up to 91% below the corresponding on-demand price for many machine types
Discount interaction
Compute Engine discount types cannot be combined, so Spot usage does not also receive committed use or sustained use discounts
Preemption
Preempted VMs stop billing for vCPU and memory; with the STOP termination action, disks remain and keep billing

How to detect

4 checks to find it in your estate.

  • List VMs and instance templates with their provisioning model (scheduling.provisioningModel STANDARD or SPOT) and map them to workload type using labels, MIG names or job metadata
  • Flag standard-model MIGs and VMs running batch, CI runners, rendering, ETL or queue consumers, especially those that scale up and down frequently
  • Check whether those workloads already retry or checkpoint (job schedulers, queues with redelivery, Batch job retries); if they do, preemption is already survivable
  • Exclude workloads that need live migration, the Compute Engine SLA, or uninterrupted runs longer than the application can checkpoint

How to fix

4 ways to remove the waste.

  • Create fault-tolerant capacity with --provisioning-model=SPOT in instance templates, and run it in managed instance groups, which recreate preempted VMs when capacity returns
  • Handle preemption: watch the preempted metadata value or the ACPI shutdown signal and use the up to 30 seconds of shutdown time (or a 120-second preemption notice, in Preview) to checkpoint or requeue work; choose STOP or DELETE as the termination action to match the workload
  • For batch jobs, use Batch or other schedulers with Spot provisioning and retries, and spread across zones and machine types to reduce the impact of capacity shortages
  • Keep a standard-VM baseline (optionally covered by CUDs) for deadline-bound work, since Spot capacity is not guaranteed

Documentation

Vendor references for pricing and configuration.