Explanation
Why the waste happens and who it affects.
Google states that an HA-configured instance costs twice as much as a standalone instance, including CPU, RAM and storage. The standby cannot serve read queries, so the extra spend buys only zonal failover.
Dev, test, QA and staging instances are frequently created from the same Terraform modules, scripts or clone operations as production and inherit HA. Google's own instance settings guidance recommends HA for production and describes zonal (single zone) availability as recommended for test and development. In non-production, HA doubles the instance and storage bill for automatic zonal failover that is rarely needed.
Billing model
The pricing dimensions that drive this cost.
Cloud SQL publishes separate HA rates; prices vary by region and edition.
- HA vCPUs and memory
- Regional instances are billed at HA vCPU and HA memory rates, which are double the standalone rates
- HA storage
- SSD, HDD and Hyperdisk Balanced capacity for regional instances is billed at HA storage rates, double the zonal rates
- Backups
- Backup storage is billed at the same rate for HA and zonal instances
- Zonal availability
- Instance and backups in a single zone at standalone rates, with no automatic failover
How to detect
4 checks to find it in your estate.
- List instances with settings.availabilityType set to REGIONAL (gcloud sql instances list --format with settings.availabilityType) and join them with project, label or naming conventions that mark dev, test, QA or staging
- In the Cloud Billing export, look for HA vCPU, HA memory and HA storage SKUs in non-production projects
- Review which environments have a documented availability requirement; instances without one are candidates for zonal availability
- Check provisioning templates and modules for availability_type = REGIONAL defaults that apply to every environment
How to fix
4 ways to remove the waste.
- Switch non-production instances to zonal availability with gcloud sql instances patch INSTANCE --availability-type ZONAL, or select Single zone in the console; the instance restarts, which typically takes a few minutes but can take up to an hour for instances with a large disk or load
- Accept the tradeoff explicitly: after a zonal outage a zonal instance has to be recovered manually, which is usually fine for dev and test
- Make ZONAL the default in non-production Terraform modules and require an explicit override for HA
- Keep HA for production and for non-production instances that are used to test failover behavior
Documentation
Vendor references for pricing and configuration.