Explanation
Why the waste happens and who it affects.
Many apps are nevertheless deployed with minReplicas set to 1 or more to avoid cold starts, copied from production templates into every environment, or pinned to a minimum during an incident and never reset. For internal tools, admin portals, webhooks, low-traffic APIs and dev or test apps that are called a few times an hour or less, those replicas run and bill every second of the day.
Replicas held at the minimum can qualify for a reduced idle rate, but only when they receive no HTTP requests, use under 0.01 vCPU and receive under 1,000 bytes per second; background loops, polling or sidecars that keep a replica busy push it back to the active rate. Serverless GPU apps never get the idle rate. In multiple-revision mode, older revisions stay active with their own scale settings, so a superseded revision with a nonzero minimum keeps running as well. Microsoft's Well-Architected guidance for Container Apps recommends scale-to-zero rules for apps that don't need to run continuously.
Billing model
The pricing dimensions that drive this cost.
The Consumption plan bills each running replica per second for the vCPU and memory allocated to it, plus HTTP requests.
- Resource consumption
- Billed in vCPU-seconds and GiB-seconds for every running replica, at the active rate by default
- Scaled to zero
- A revision with zero replicas incurs no resource consumption charges
- Idle rate
- Reduced rate for replicas held at a minimum above zero that meet all idle conditions; never applies to serverless GPU apps
- Monthly free grant
- The first 180,000 vCPU-seconds, 360,000 GiB-seconds and 2 million HTTP requests per subscription each month are free
How to detect
5 checks to find it in your estate.
- List container apps and their scale settings (properties.template.scale.minReplicas in az containerapp list or Azure Resource Graph type microsoft.app/containerapps) and flag Consumption-plan apps with minReplicas of 1 or more
- For flagged apps, compare the Requests metric over 30 days with the Replica Count metric; apps with few requests but a replica count that never drops are candidates
- Check for active non-latest revisions in multiple-revision mode that still have minReplicas above zero and receive no traffic
- Break down Container Apps charges in Cost analysis by meter to see how much is billed at active versus idle rates for each app
- Review environment and team standards or templates that set a default minimum replica count for all apps
How to fix
5 ways to remove the waste.
- Set minReplicas to 0 for intermittent HTTP apps and rely on the HTTP scale rule to start a replica on the first request; accept the cold-start delay or reduce it with smaller images
- For queue- or event-driven workers, use KEDA custom scale rules (for example Service Bus or Storage Queue length) with minReplicas 0 so replicas run only when there is work, keeping in mind the 300-second default cool down before scaling to zero
- Keep a minimum above zero only where a latency objective requires a warm replica, and use one replica rather than several where availability requirements allow
- Deactivate old revisions that no longer receive traffic, or use single-revision mode where traffic splitting isn't needed
- Rightsize CPU and memory per replica, since the minimum replicas are billed for their full allocation whether idle or active. Apps without ingress need a custom scale rule before minReplicas is set to 0, or they won't start again
Documentation
Vendor references for pricing and configuration.