Explanation
Why the waste happens and who it affects.
Request-based billing, the default, charges CPU and memory only while an instance is starting, shutting down or processing at least one request, and adds a per-request fee. Instance-based billing charges for the entire lifetime of every instance, even when no requests arrive, but at a lower per-second rate and with no per-request fee. Which one is cheaper depends on traffic shape, and the setting is rarely revisited after the first deploy.
Services with steady, high-concurrency traffic often stay on request-based billing and pay the higher active rate plus request fees for time the instances would have been busy anyway. The opposite also happens: a service switched to instance-based billing to run background work or monitoring agents keeps paying for full instance lifetimes after its traffic becomes sporadic. Google's own guidance is that request-based billing fits sporadic, bursty or spiky traffic and instance-based billing fits steady, slowly varying traffic.
Billing model
The pricing dimensions that drive this cost.
Rates are per vCPU-second and GiB-second of billable instance time, rounded up to the nearest 100 ms, and vary by region.
- Request-based billing
- CPU and memory billed only while an instance starts, shuts down or processes at least one request, plus a charge per million requests
- Instance-based billing
- CPU and memory billed for the full instance lifetime from start to termination, with a 1-minute minimum and no per-request charge, at a lower unit rate than request-based active time
- Idle minimum instances
- Under request-based billing, instances kept warm by minimum instances are billed at a separate idle rate when not serving requests
- Committed use discounts
- Cloud Run CUDs and Compute Flexible CUDs lower the rate for continuous usage under either setting
How to detect
4 checks to find it in your estate.
- Review the Cloud Run CPU allocation recommender (google.run.service.CostRecommender, recommendation "Switch to CPU always allocated"); it looks at the past month of traffic and flags services where instance-based billing would be cheaper
- Note that the recommender only suggests moving from request-based to instance-based billing; services already on instance-based billing need a manual check
- For services on instance-based billing, compare run.googleapis.com/container/billable_instance_time with request count and concurrency over time to find long idle periods between bursts
- For services on request-based billing, look for a consistently high number of concurrent requests and instance counts that rarely drop, which is the pattern Google says suits instance-based billing
How to fix
4 ways to remove the waste.
- Switch steady-traffic services to instance-based billing with gcloud run services update --no-cpu-throttling, or in the console under the service's billing setting, after confirming with the recommender or a cost model
- Move sporadic or spiky services back to request-based billing with --cpu-throttling, unless they depend on CPU outside of requests for background work
- Before switching to request-based billing, check that the code does not rely on goroutines, async tasks, threads or agents that keep working after a response is returned; under request-based billing that work is throttled
- For services that stay continuously busy under either mode, cover the baseline with a Cloud Run or Compute Flexible committed use discount
Documentation
Vendor references for pricing and configuration.