Explanation
Why the waste happens and who it affects.
Each such instance reserves 10 capacity units, and those units are billed every hour the gateway is active whether or not any traffic uses them. Teams often choose a high manual count or a high autoscale minimum at launch for safety, or copy production settings into every environment, and never revisit it once real traffic is known.
The result is that Estimated Billed Capacity Units sits flat at the Fixed Billable Capacity Units level while Current Capacity Units stay far below it. Some headroom is deliberate: new instances take three to five minutes to provision, and Microsoft suggests keeping the minimum well above zero for gateways with sharp traffic bursts. The waste is the gap between that justified buffer and a minimum sized for a peak that never arrives, which is most common on gateways left in manual mode and on non-production gateways configured like production.
Billing model
The pricing dimensions that drive this cost.
v2 gateways are billed as a fixed hourly charge plus capacity units, computed hourly.
- Capacity unit
- The highest of compute units, persistent connections (2,500 per CU) and throughput (1 GB per hour per CU) consumed in the hour
- Reserved capacity units
- 10 per instance of the manual count or autoscale minimum, billed while the gateway is active regardless of consumption
- Estimated billed capacity units
- The greater of current capacity units and fixed billable capacity units, which is what the variable charge is based on
- Maximum instance count
- An upper limit only; raising it does not add cost because only consumed capacity is billed
How to detect
4 checks to find it in your estate.
- Chart the gateway metrics Fixed Billable Capacity Units, Current Capacity Units and Estimated Billed Capacity Units over the last 30 days; when estimated billed units equal fixed billable units most of the time, the reservation from the instance count is driving the bill
- List v2 gateways and their scaling mode with Azure Resource Graph (type microsoft.network/applicationgateways, properties.sku.capacity for manual mode and properties.autoscaleConfiguration.minCapacity for autoscaling) and flag high manual counts or minimums
- Compare the configured minimum with Microsoft's sizing guidance: take the peak Current Compute Units over the past month and divide by 10 to estimate the instances needed
- Flag non-production gateways whose minimum or manual instance count matches production
How to fix
5 ways to remove the waste.
- Switch gateways from manual mode to autoscaling so capacity follows traffic instead of being fixed at the provisioned count
- Lower the autoscale minimum to a level derived from observed compute units plus the buffer your burst profile needs, remembering that scale-out takes three to five minutes
- Set a minimum of 0 for low-traffic and non-production gateways where brief latency during scale-out is acceptable; the fixed hourly charge still applies
- Set the maximum instance count high (up to 125, subnet size permitting) instead of relying on a high minimum for protection, since the maximum does not add cost
- Add alerts on Current Compute Units and Current Capacity Units so the minimum can be raised deliberately if sustained traffic grows
Documentation
Vendor references for pricing and configuration.
- Understanding pricing - Azure Application Gatewaylearn.microsoft.com
- Scaling and Zone-redundant Application Gateway v2learn.microsoft.com
- Application Gateway high traffic volume supportlearn.microsoft.com
- Architecture Best Practices for Azure Application Gateway v2learn.microsoft.com
- Application Gateway pricingazure.microsoft.com