Skip to content
Cloud Efficiency Hub

Suboptimal Billing Mode for Cloud Run Services

The short version

Cloud Run services run under one of two billing settings.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP Cloud Run
Category
Compute
Reference
CER-0470
Type
Suboptimal Pricing Model

Explanation

Why the waste happens and who it affects.

Request-based billing, the default, charges CPU and memory only while an instance is starting, shutting down or processing at least one request, and adds a per-request fee. Instance-based billing charges for the entire lifetime of every instance, even when no requests arrive, but at a lower per-second rate and with no per-request fee. Which one is cheaper depends on traffic shape, and the setting is rarely revisited after the first deploy.

Services with steady, high-concurrency traffic often stay on request-based billing and pay the higher active rate plus request fees for time the instances would have been busy anyway. The opposite also happens: a service switched to instance-based billing to run background work or monitoring agents keeps paying for full instance lifetimes after its traffic becomes sporadic. Google's own guidance is that request-based billing fits sporadic, bursty or spiky traffic and instance-based billing fits steady, slowly varying traffic.

Billing model

The pricing dimensions that drive this cost.

Rates are per vCPU-second and GiB-second of billable instance time, rounded up to the nearest 100 ms, and vary by region.

Request-based billing
CPU and memory billed only while an instance starts, shuts down or processes at least one request, plus a charge per million requests
Instance-based billing
CPU and memory billed for the full instance lifetime from start to termination, with a 1-minute minimum and no per-request charge, at a lower unit rate than request-based active time
Idle minimum instances
Under request-based billing, instances kept warm by minimum instances are billed at a separate idle rate when not serving requests
Committed use discounts
Cloud Run CUDs and Compute Flexible CUDs lower the rate for continuous usage under either setting

How to detect

4 checks to find it in your estate.

  • Review the Cloud Run CPU allocation recommender (google.run.service.CostRecommender, recommendation "Switch to CPU always allocated"); it looks at the past month of traffic and flags services where instance-based billing would be cheaper
  • Note that the recommender only suggests moving from request-based to instance-based billing; services already on instance-based billing need a manual check
  • For services on instance-based billing, compare run.googleapis.com/container/billable_instance_time with request count and concurrency over time to find long idle periods between bursts
  • For services on request-based billing, look for a consistently high number of concurrent requests and instance counts that rarely drop, which is the pattern Google says suits instance-based billing

How to fix

4 ways to remove the waste.

  • Switch steady-traffic services to instance-based billing with gcloud run services update --no-cpu-throttling, or in the console under the service's billing setting, after confirming with the recommender or a cost model
  • Move sporadic or spiky services back to request-based billing with --cpu-throttling, unless they depend on CPU outside of requests for background work
  • Before switching to request-based billing, check that the code does not rely on goroutines, async tasks, threads or agents that keep working after a response is returned; under request-based billing that work is throttled
  • For services that stay continuously busy under either mode, cover the baseline with a Cloud Run or Compute Flexible committed use discount

Documentation

Vendor references for pricing and configuration.