# Container Apps With Minimum Replicas Above Zero for Intermittent Workloads

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/container-apps-with-minimum-replicas-above-zero-for-intermittent-workloads

On the Container Apps Consumption plan, a revision scaled to zero replicas incurs no resource consumption charges, and the default scale rule allows...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

On the Container Apps Consumption plan, a revision scaled to zero replicas incurs no resource consumption charges, and the default scale rule allows HTTP apps to scale to zero.

PointFive Research

Cloud cost research at PointFive

Azure service

[Azure Container Apps](https://www.pointfive.co/efficiency-hub/cloud-services/azure-container-apps)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0408

Type

Inefficient Configuration

## Explanation

Why the waste happens and who it affects.

Many apps are nevertheless deployed with minReplicas set to 1 or more to avoid cold starts, copied from production templates into every environment, or pinned to a minimum during an incident and never reset. For internal tools, admin portals, webhooks, low-traffic APIs and dev or test apps that are called a few times an hour or less, those replicas run and bill every second of the day.

Replicas held at the minimum can qualify for a reduced idle rate, but only when they receive no HTTP requests, use under 0.01 vCPU and receive under 1,000 bytes per second; background loops, polling or sidecars that keep a replica busy push it back to the active rate. Serverless GPU apps never get the idle rate. In multiple-revision mode, older revisions stay active with their own scale settings, so a superseded revision with a nonzero minimum keeps running as well. Microsoft's Well-Architected guidance for Container Apps recommends scale-to-zero rules for apps that don't need to run continuously.

## Billing model

The pricing dimensions that drive this cost.

The Consumption plan bills each running replica per second for the vCPU and memory allocated to it, plus HTTP requests.

Resource consumption

Billed in vCPU-seconds and GiB-seconds for every running replica, at the active rate by default

Scaled to zero

A revision with zero replicas incurs no resource consumption charges

Idle rate

Reduced rate for replicas held at a minimum above zero that meet all idle conditions; never applies to serverless GPU apps

Monthly free grant

The first 180,000 vCPU-seconds, 360,000 GiB-seconds and 2 million HTTP requests per subscription each month are free

## How to detect

5 checks to find it in your estate.

- List container apps and their scale settings (properties.template.scale.minReplicas in az containerapp list or Azure Resource Graph type microsoft.app/containerapps) and flag Consumption-plan apps with minReplicas of 1 or more

- For flagged apps, compare the Requests metric over 30 days with the Replica Count metric; apps with few requests but a replica count that never drops are candidates

- Check for active non-latest revisions in multiple-revision mode that still have minReplicas above zero and receive no traffic

- Break down Container Apps charges in Cost analysis by meter to see how much is billed at active versus idle rates for each app

- Review environment and team standards or templates that set a default minimum replica count for all apps

## How to fix

5 ways to remove the waste.

- Set minReplicas to 0 for intermittent HTTP apps and rely on the HTTP scale rule to start a replica on the first request; accept the cold-start delay or reduce it with smaller images

- For queue- or event-driven workers, use KEDA custom scale rules (for example Service Bus or Storage Queue length) with minReplicas 0 so replicas run only when there is work, keeping in mind the 300-second default cool down before scaling to zero

- Keep a minimum above zero only where a latency objective requires a warm replica, and use one replica rather than several where availability requirements allow

- Deactivate old revisions that no longer receive traffic, or use single-revision mode where traffic splitting isn't needed

- Rightsize CPU and memory per replica, since the minimum replicas are billed for their full allocation whether idle or active. Apps without ingress need a custom scale rule before minReplicas is set to 0, or they won't start again

## Documentation

Vendor references for pricing and configuration.

- [Billing in Azure Container Apps  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/container-apps/billing)

- [Scaling in Azure Container Apps  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/container-apps/scale-app)

- [Architecture Best Practices for Azure Container Apps  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/well-architected/service-guides/azure-container-apps)

- [Azure Container Apps pricing  azure.microsoft.com](https://azure.microsoft.com/en-us/pricing/details/container-apps/)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Azure Virtual Machine Scale Sets  CER-0314

### [Fixed Instance Count on Virtual Machine Scale Set Without Autoscaling](https://www.pointfive.co/efficiency-hub/inefficiencies/fixed-instance-count-on-virtual-machine-scale-set-without-autoscaling)

Azure Virtual Machine Scale Sets can operate in two modes: manual scaling with a fixed instance count, or autoscaling with dynamic instance counts that respond to demand. When a scale set is configured with manual scaling, it maintains the...

Compute

- Azure App Service  CER-0306

### [Non-Production App Service Plans Running Higher Tiers During Off-Hours](https://www.pointfive.co/efficiency-hub/inefficiencies/non-production-app-service-plans-running-higher-tiers-during-off-hours)

Azure App Service Plans define the compute resources allocated to web applications and are billed continuously based on their pricing tier - regardless of whether the hosted apps are actively serving traffic. In non-production environments...

Compute

- Azure Virtual Machines  CER-0144

### [Missing Scheduled Shutdown for Non-Production Azure Virtual Machines](https://www.pointfive.co/efficiency-hub/inefficiencies/missing-scheduled-shutdown-for-non-production-azure-virtual-machines)

Non-production Azure VMs are frequently left running during off-hours despite being used only during business hours. When these instances remain active overnight or on weekends, they generate unnecessary compute spend. Azure offers...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

