# Missed Use of Spot VMs for Fault-Tolerant Compute Engine Workloads

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/missed-use-of-spot-vms-for-fault-tolerant-compute-engine-workloads

Spot VMs run the same machine types as standard VMs at discounts of up to 91% for many machine types, GPUs, TPUs and Local SSDs, in exchange for no...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Spot VMs run the same machine types as standard VMs at discounts of up to 91% for many machine types, GPUs, TPUs and Local SSDs, in exchange for no availability guarantee: Compute Engine can preempt them at any time.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Compute Engine](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-compute-engine)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0468

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

Batch processing, CI and build runners, rendering, simulations, queue workers, and other stateless or checkpointed jobs tolerate that well. Yet these workloads often run on the standard provisioning model because the instance template was copied from a production service, or because nobody revisited the choice after the workload became automated.

Google's Well-Architected cost guidance recommends Spot VMs for non-critical or fault-tolerant workloads. Work that can be retried or resumed pays full on-demand rates on the standard model when it could be discounted by up to 91%, and large autoscaled worker fleets in managed instance groups can absorb preemption because the group recreates preempted VMs.

## Billing model

The pricing dimensions that drive this cost.

Standard provisioning model

VCPU, memory, GPUs and Local SSD billed at on-demand rates, eligible for sustained use and committed use discounts

Spot VMs

Billed at variable Spot prices, up to 91% below the corresponding on-demand price for many machine types

Discount interaction

Compute Engine discount types cannot be combined, so Spot usage does not also receive committed use or sustained use discounts

Preemption

Preempted VMs stop billing for vCPU and memory; with the STOP termination action, disks remain and keep billing

## How to detect

4 checks to find it in your estate.

- List VMs and instance templates with their provisioning model (scheduling.provisioningModel STANDARD or SPOT) and map them to workload type using labels, MIG names or job metadata

- Flag standard-model MIGs and VMs running batch, CI runners, rendering, ETL or queue consumers, especially those that scale up and down frequently

- Check whether those workloads already retry or checkpoint (job schedulers, queues with redelivery, Batch job retries); if they do, preemption is already survivable

- Exclude workloads that need live migration, the Compute Engine SLA, or uninterrupted runs longer than the application can checkpoint

## How to fix

4 ways to remove the waste.

- Create fault-tolerant capacity with --provisioning-model=SPOT in instance templates, and run it in managed instance groups, which recreate preempted VMs when capacity returns

- Handle preemption: watch the preempted metadata value or the ACPI shutdown signal and use the up to 30 seconds of shutdown time (or a 120-second preemption notice, in Preview) to checkpoint or requeue work; choose STOP or DELETE as the termination action to match the workload

- For batch jobs, use Batch or other schedulers with Spot provisioning and retries, and spread across zones and machine types to reduce the impact of capacity shortages

- Keep a standard-VM baseline (optionally covered by CUDs) for deadline-bound work, since Spot capacity is not guaranteed

## Documentation

Vendor references for pricing and configuration.

- [Spot VMs  docs.cloud.google.com](https://docs.cloud.google.com/compute/docs/instances/spot)

- [Optimize resource usage  docs.cloud.google.com](https://docs.cloud.google.com/architecture/framework/cost-optimization/optimize-resource-usage)

- [VM instance pricing  cloud.google.com](https://cloud.google.com/products/compute/pricing)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Compute Engine  CER-0282

### [Missed Use of Committed Use Discounts for Compute Engine](https://www.pointfive.co/efficiency-hub/inefficiencies/missed-use-of-committed-use-discounts-for-compute-engine)

Workloads with predictable, long-running compute usage continue to run entirely on on-demand pricing instead of leveraging Committed Use Discounts. For stable environments, such as production services or continuously running batch...

Compute

- GCP Compute Engine  CER-0089

### [Underutilized VM Commitments Due to Architectural Drift](https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-vm-commitments-due-to-architectural-drift)

VM-based Committed Use Discounts in GCP offer cost savings for predictable workloads, but they are rigid: they apply only to specified VM types, quantities, and regions. When organizations evolve their architecture - such as moving to GKE...

Compute

- GCP Compute Engine  CER-0095

### [Missing Scheduled Shutdown for Non-Production Compute Engine Instances](https://www.pointfive.co/efficiency-hub/inefficiencies/missing-scheduled-shutdown-for-non-production-compute-engine-instances)

Development and test environments on Compute Engine are commonly provisioned and left running around the clock, even if only used during business hours. This results in wasteful spend on compute time that could be eliminated by scheduling...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

