Skip to content
Cloud Efficiency Hub

Excessive Minimum Instances in Cloud Run Services

The short version

Cloud Run scales to zero by default, but the minimum instances setting keeps a number of instances started and warm so requests avoid cold starts.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP Cloud Run
Category
Compute
Reference
CER-0471
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

Those instances are billed whether or not traffic arrives. Minimums are often raised for a launch or a latency incident and never lowered, copied into every environment including dev and staging, or set higher than the number of instances the service needs at its quietest hours.

A less visible source is revision-level minimum instances combined with traffic tags. When minimums are set on the revision, every tagged revision is started and kept active even when it receives no requests, so old tagged revisions used for testing or rollback keep billing indefinitely. The cost scales with the number of services, revisions and environments rather than with traffic.

Billing model

The pricing dimensions that drive this cost.

How warm instances are charged depends on the service's billing setting; rates vary by region.

Idle minimum instance time
Under request-based billing, minimum instances not serving requests are billed at an idle rate; the published idle CPU rate is lower than the active rate, while idle memory is billed at the same rate as active memory
Instance-based billing
Minimum instances are billed at the full rate for their entire lifetime, the same as any other instance
Zero minimum instances
With request-based billing and minimum instances set to 0, idle instances are not charged
Tagged revisions
With revision-level minimums, each tagged revision keeps its minimum instances running even without traffic

How to detect

4 checks to find it in your estate.

  • List services and revisions whose minimum instances setting (run.googleapis.com/minScale at service level, autoscaling.knative.dev/minScale at revision level) is greater than zero, and flag non-production projects first
  • Chart run.googleapis.com/container/instance_count split by the state label (active or idle) and look for services where idle instances make up most of the instance count for long periods
  • Compare the minimum with the active instance count during the lowest-traffic hours of the week; a minimum well above that baseline is paying for warm capacity nobody uses
  • Find tagged revisions that receive no traffic but carry a revision-level minimum instances value

How to fix

4 ways to remove the waste.

  • Set minimum instances close to the number of instances needed for typical low traffic, as Google recommends, or to 0 where cold-start latency is acceptable (for example in dev and test)
  • Use service-level minimum instances instead of revision-level minimums so capacity follows traffic splits, and remove tags from revisions that are no longer needed
  • Clear a minimum with gcloud run services update SERVICE --min default, or lower it with --min; service-level changes take effect without a new deployment
  • For minimums that must stay (latency-sensitive production), cover the predictable always-on usage with a committed use discount; keep in mind that lowering minimums increases cold starts, which startup CPU boost and smaller container images can offset

Documentation

Vendor references for pricing and configuration.