# Excessive Minimum Instances in Cloud Run Services

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/excessive-minimum-instances-in-cloud-run-services

Cloud Run scales to zero by default, but the minimum instances setting keeps a number of instances started and warm so requests avoid cold starts.

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Cloud Run scales to zero by default, but the minimum instances setting keeps a number of instances started and warm so requests avoid cold starts.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Cloud Run](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-cloud-run)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0471

Type

Idle or Unused Resource

## Explanation

Why the waste happens and who it affects.

Those instances are billed whether or not traffic arrives. Minimums are often raised for a launch or a latency incident and never lowered, copied into every environment including dev and staging, or set higher than the number of instances the service needs at its quietest hours.

A less visible source is revision-level minimum instances combined with traffic tags. When minimums are set on the revision, every tagged revision is started and kept active even when it receives no requests, so old tagged revisions used for testing or rollback keep billing indefinitely. The cost scales with the number of services, revisions and environments rather than with traffic.

## Billing model

The pricing dimensions that drive this cost.

How warm instances are charged depends on the service's billing setting; rates vary by region.

Idle minimum instance time

Under request-based billing, minimum instances not serving requests are billed at an idle rate; the published idle CPU rate is lower than the active rate, while idle memory is billed at the same rate as active memory

Instance-based billing

Minimum instances are billed at the full rate for their entire lifetime, the same as any other instance

Zero minimum instances

With request-based billing and minimum instances set to 0, idle instances are not charged

Tagged revisions

With revision-level minimums, each tagged revision keeps its minimum instances running even without traffic

## How to detect

4 checks to find it in your estate.

- List services and revisions whose minimum instances setting (run.googleapis.com/minScale at service level, autoscaling.knative.dev/minScale at revision level) is greater than zero, and flag non-production projects first

- Chart run.googleapis.com/container/instance\_count split by the state label (active or idle) and look for services where idle instances make up most of the instance count for long periods

- Compare the minimum with the active instance count during the lowest-traffic hours of the week; a minimum well above that baseline is paying for warm capacity nobody uses

- Find tagged revisions that receive no traffic but carry a revision-level minimum instances value

## How to fix

4 ways to remove the waste.

- Set minimum instances close to the number of instances needed for typical low traffic, as Google recommends, or to 0 where cold-start latency is acceptable (for example in dev and test)

- Use service-level minimum instances instead of revision-level minimums so capacity follows traffic splits, and remove tags from revisions that are no longer needed

- Clear a minimum with gcloud run services update SERVICE --min default, or lower it with --min; service-level changes take effect without a new deployment

- For minimums that must stay (latency-sensitive production), cover the predictable always-on usage with a committed use discount; keep in mind that lowering minimums increases cold starts, which startup CPU boost and smaller container images can offset

## Documentation

Vendor references for pricing and configuration.

- [Set minimum instances for services  docs.cloud.google.com](https://docs.cloud.google.com/run/docs/configuring/min-instances)

- [Cloud Run pricing  cloud.google.com](https://cloud.google.com/run/pricing)

- [Billing settings for services  docs.cloud.google.com](https://docs.cloud.google.com/run/docs/configuring/billing-settings)

- [Google Cloud metrics: P through Z  docs.cloud.google.com](https://docs.cloud.google.com/monitoring/api/metrics_gcp_p_z)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Cloud Run  CER-0179

### [Overprovisioned Memory Allocation in Cloud Run Services](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-memory-allocation-in-cloud-run-services)

In Cloud Run, each revision is deployed with a fixed memory allocation (e.g., 512MiB, 1GiB, 2GiB, etc.). These settings are often overestimated during initial development or copied from templates. Unlike auto-scaling platforms that adapt...

Compute

- GCP Cloud Run  CER-0470

### [Suboptimal Billing Mode for Cloud Run Services](https://www.pointfive.co/efficiency-hub/inefficiencies/suboptimal-billing-mode-for-cloud-run-services)

Cloud Run services run under one of two billing settings. Request-based billing, the default, charges CPU and memory only while an instance is starting, shutting down or processing at least one request, and adds a per-request fee....

Compute

- GCP Dataflow  CER-0244

### [Idle Dataflow Workers Running After Pipeline Failure](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-dataflow-workers-running-after-pipeline-failure-f6b1a)

When a Dataflow pipeline fails - often due to dependency issues, misconfigurations, or data format mismatches-its worker instances may remain active temporarily until the service terminates them. In some cases, misconfigured jobs, stuck...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

