Explanation
Why the waste happens and who it affects.
Microsoft's Well-Architected guidance states plainly that zone-redundant or zonal configurations double instance costs. That is justified for production databases that need a 99.99% or 99.95% uptime SLA and automatic failover.
Development, test, QA and staging servers are often created from the same templates or portal defaults as production, so they carry a standby they don't need. Non-production workloads can usually tolerate the restart-based recovery of a server without HA, which still keeps three copies of data on locally redundant storage and automatically restarts or relocates a failed server, or a point-in-time restore from backups. HA also prevents moving the server to the Burstable tier, which doesn't support it, so these servers often stay on larger General Purpose or Memory Optimized compute as well.
Billing model
The pricing dimensions that drive this cost.
- Standby replica
- A second server with the same compute and storage as the primary, billed while HA is enabled
- Redundancy cost
- Zone-redundant or same-zone HA doubles the instance cost, per the Well-Architected guidance
- Server without HA
- A single billed server with locally redundant storage and built-in restart and relocation, with a lower uptime SLA
- Burstable tier
- Lowest-cost compute tier; high availability isn't supported on it
How to detect
4 checks to find it in your estate.
- List flexible servers with Azure Resource Graph (type microsoft.dbforpostgresql/flexibleservers) and filter where properties.highAvailability.mode is ZoneRedundant or SameZone
- Join the results with environment tags, subscription names or resource group naming to find HA-enabled servers in development, test, QA and staging scopes
- For HA-enabled non-production servers, check CPU and connection metrics; low-utilization servers are also candidates for the Burstable tier once HA is removed
- Confirm with owners that no documented availability objective requires automatic failover for the server
How to fix
5 ways to remove the waste.
- Disable high availability on non-production servers from the High availability page in the portal (set Zonal resiliency to Disabled) or with az postgres flexible-server update; Microsoft documents enabling and disabling HA as an online operation that doesn't change networking, firewall, parameter or backup settings
- After HA is removed, evaluate moving low-utilization non-production servers to the Burstable tier and stopping them outside working hours with the start/stop feature
- Rely on automated backups and point-in-time restore for non-production recovery
- Enforce the standard with Azure Policy or infrastructure-as-code defaults so non-production servers are deployed without HA unless an exception is approved
- Re-enable HA manually after operations that disable it automatically, such as in-place major version upgrades, only on servers that need it
Documentation
Vendor references for pricing and configuration.