Explanation
Why the waste happens and who it affects.
Google charges for the reserved resources at the same on-demand rate as running VMs, including any applicable discounts, for as long as the reservation exists. When VMs consume the reservation, that charge is simply the VMs' normal cost; when they do not, the unconsumed slots are paid for with nothing running on them.
Reservations are typically created ahead of a launch, a migration, a peak event or a period of capacity scarcity for GPUs or large machine types, and then left in place after the event or after the workload moves to another machine type, zone or to GKE. A reservation sized for peak also keeps billing for the full VM count while autoscaling runs well below it.
Billing model
The pricing dimensions that drive this cost.
- Reserved resources
- VCPUs, memory, GPUs and Local SSD in a reservation are billed at on-demand rates, including applicable discounts, from creation until deletion, whether consumed or not
- Consumed capacity
- VMs that match and consume a reservation are not billed twice; only the unconsumed portion is waste
- Commitment-attached reservations
- Reservations attached to resource-based commitments receive the commitment discount and are tied to the commitment
- Auto-delete
- A reservation can be set to delete itself at a specified time, whether or not it is fully consumed
How to detect
4 checks to find it in your estate.
- Review the Idle reservations recommender (google.compute.IdleResourceRecommender, "Delete unused resource reservations"); it flags reservations that accrue cost but were not consumed by any active resource over the observation period (7 days by default, configurable up to 30 days)
- Review the Underutilized reservations recommender (google.compute.RightSizeResourceRecommender, "Right-size underutilized reservations"); the default utilization threshold is 80% and can be changed
- Chart compute.googleapis.com/reservation/reserved against reservation/used, or compare specificReservation.count with inUseCount from gcloud compute reservations describe
- Check reservations attached to CUDs and reservations for TPU VMs separately, since the recommenders do not cover them
How to fix
4 ways to remove the waste.
- Delete reservations that are no longer needed with gcloud compute reservations delete, after confirming no planned event depends on the capacity
- Resize underutilized reservations down to the steady consumed count (gcloud compute reservations update --vm-count), keeping headroom only where capacity risk justifies it
- Set auto-delete on reservations created for a bounded event so they end automatically, and use future reservations for planned capacity needs instead of holding capacity open ahead of time
- Make sure workloads actually target the reservation (automatic consumption or matching specific reservation affinity); a VM with a different machine type, zone or affinity will not consume it
Documentation
Vendor references for pricing and configuration.