Skip to content
Cloud Efficiency Hub

Container Apps With Minimum Replicas Above Zero for Intermittent Workloads

The short version

On the Container Apps Consumption plan, a revision scaled to zero replicas incurs no resource consumption charges, and the default scale rule allows HTTP apps to scale to zero.

PointFive Research

Cloud cost research at PointFive

Category
Compute
Reference
CER-0408
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

Many apps are nevertheless deployed with minReplicas set to 1 or more to avoid cold starts, copied from production templates into every environment, or pinned to a minimum during an incident and never reset. For internal tools, admin portals, webhooks, low-traffic APIs and dev or test apps that are called a few times an hour or less, those replicas run and bill every second of the day.

Replicas held at the minimum can qualify for a reduced idle rate, but only when they receive no HTTP requests, use under 0.01 vCPU and receive under 1,000 bytes per second; background loops, polling or sidecars that keep a replica busy push it back to the active rate. Serverless GPU apps never get the idle rate. In multiple-revision mode, older revisions stay active with their own scale settings, so a superseded revision with a nonzero minimum keeps running as well. Microsoft's Well-Architected guidance for Container Apps recommends scale-to-zero rules for apps that don't need to run continuously.

Billing model

The pricing dimensions that drive this cost.

The Consumption plan bills each running replica per second for the vCPU and memory allocated to it, plus HTTP requests.

Resource consumption
Billed in vCPU-seconds and GiB-seconds for every running replica, at the active rate by default
Scaled to zero
A revision with zero replicas incurs no resource consumption charges
Idle rate
Reduced rate for replicas held at a minimum above zero that meet all idle conditions; never applies to serverless GPU apps
Monthly free grant
The first 180,000 vCPU-seconds, 360,000 GiB-seconds and 2 million HTTP requests per subscription each month are free

How to detect

5 checks to find it in your estate.

  • List container apps and their scale settings (properties.template.scale.minReplicas in az containerapp list or Azure Resource Graph type microsoft.app/containerapps) and flag Consumption-plan apps with minReplicas of 1 or more
  • For flagged apps, compare the Requests metric over 30 days with the Replica Count metric; apps with few requests but a replica count that never drops are candidates
  • Check for active non-latest revisions in multiple-revision mode that still have minReplicas above zero and receive no traffic
  • Break down Container Apps charges in Cost analysis by meter to see how much is billed at active versus idle rates for each app
  • Review environment and team standards or templates that set a default minimum replica count for all apps

How to fix

5 ways to remove the waste.

  • Set minReplicas to 0 for intermittent HTTP apps and rely on the HTTP scale rule to start a replica on the first request; accept the cold-start delay or reduce it with smaller images
  • For queue- or event-driven workers, use KEDA custom scale rules (for example Service Bus or Storage Queue length) with minReplicas 0 so replicas run only when there is work, keeping in mind the 300-second default cool down before scaling to zero
  • Keep a minimum above zero only where a latency objective requires a warm replica, and use one replica rather than several where availability requirements allow
  • Deactivate old revisions that no longer receive traffic, or use single-revision mode where traffic splitting isn't needed
  • Rightsize CPU and memory per replica, since the minimum replicas are billed for their full allocation whether idle or active. Apps without ingress need a custom scale rule before minReplicas is set to 0, or they won't start again

Documentation

Vendor references for pricing and configuration.