# Suboptimal Billing Mode for Cloud Run Services

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/suboptimal-billing-mode-for-cloud-run-services

Cloud Run services run under one of two billing settings. Request-based billing, the default, charges CPU and memory only while an instance is...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Cloud Run services run under one of two billing settings.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Cloud Run](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-cloud-run)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0470

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

Request-based billing, the default, charges CPU and memory only while an instance is starting, shutting down or processing at least one request, and adds a per-request fee. Instance-based billing charges for the entire lifetime of every instance, even when no requests arrive, but at a lower per-second rate and with no per-request fee. Which one is cheaper depends on traffic shape, and the setting is rarely revisited after the first deploy.

Services with steady, high-concurrency traffic often stay on request-based billing and pay the higher active rate plus request fees for time the instances would have been busy anyway. The opposite also happens: a service switched to instance-based billing to run background work or monitoring agents keeps paying for full instance lifetimes after its traffic becomes sporadic. Google's own guidance is that request-based billing fits sporadic, bursty or spiky traffic and instance-based billing fits steady, slowly varying traffic.

## Billing model

The pricing dimensions that drive this cost.

Rates are per vCPU-second and GiB-second of billable instance time, rounded up to the nearest 100 ms, and vary by region.

Request-based billing

CPU and memory billed only while an instance starts, shuts down or processes at least one request, plus a charge per million requests

Instance-based billing

CPU and memory billed for the full instance lifetime from start to termination, with a 1-minute minimum and no per-request charge, at a lower unit rate than request-based active time

Idle minimum instances

Under request-based billing, instances kept warm by minimum instances are billed at a separate idle rate when not serving requests

Committed use discounts

Cloud Run CUDs and Compute Flexible CUDs lower the rate for continuous usage under either setting

## How to detect

4 checks to find it in your estate.

- Review the Cloud Run CPU allocation recommender (google.run.service.CostRecommender, recommendation "Switch to CPU always allocated"); it looks at the past month of traffic and flags services where instance-based billing would be cheaper

- Note that the recommender only suggests moving from request-based to instance-based billing; services already on instance-based billing need a manual check

- For services on instance-based billing, compare run.googleapis.com/container/billable\_instance\_time with request count and concurrency over time to find long idle periods between bursts

- For services on request-based billing, look for a consistently high number of concurrent requests and instance counts that rarely drop, which is the pattern Google says suits instance-based billing

## How to fix

4 ways to remove the waste.

- Switch steady-traffic services to instance-based billing with gcloud run services update --no-cpu-throttling, or in the console under the service's billing setting, after confirming with the recommender or a cost model

- Move sporadic or spiky services back to request-based billing with --cpu-throttling, unless they depend on CPU outside of requests for background work

- Before switching to request-based billing, check that the code does not rely on goroutines, async tasks, threads or agents that keep working after a response is returned; under request-based billing that work is throttled

- For services that stay continuously busy under either mode, cover the baseline with a Cloud Run or Compute Flexible committed use discount

## Documentation

Vendor references for pricing and configuration.

- [Cloud Run pricing  cloud.google.com](https://cloud.google.com/run/pricing)

- [Billing settings for services  docs.cloud.google.com](https://docs.cloud.google.com/run/docs/configuring/billing-settings)

- [Optimize with Recommender  docs.cloud.google.com](https://docs.cloud.google.com/run/docs/recommender)

- [Recommenders  docs.cloud.google.com](https://docs.cloud.google.com/recommender/docs/recommenders)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Cloud Run  CER-0179

### [Overprovisioned Memory Allocation in Cloud Run Services](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-memory-allocation-in-cloud-run-services)

In Cloud Run, each revision is deployed with a fixed memory allocation (e.g., 512MiB, 1GiB, 2GiB, etc.). These settings are often overestimated during initial development or copied from templates. Unlike auto-scaling platforms that adapt...

Compute

- GCP Cloud Run  CER-0471

### [Excessive Minimum Instances in Cloud Run Services](https://www.pointfive.co/efficiency-hub/inefficiencies/excessive-minimum-instances-in-cloud-run-services)

Cloud Run scales to zero by default, but the minimum instances setting keeps a number of instances started and warm so requests avoid cold starts. Those instances are billed whether or not traffic arrives. Minimums are often raised for a...

Compute

- GCP Compute Engine  CER-0282

### [Missed Use of Committed Use Discounts for Compute Engine](https://www.pointfive.co/efficiency-hub/inefficiencies/missed-use-of-committed-use-discounts-for-compute-engine)

Workloads with predictable, long-running compute usage continue to run entirely on on-demand pricing instead of leveraging Committed Use Discounts. For stable environments, such as production services or continuously running batch...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

