# Overprovisioned Pod Resource Requests in GKE Autopilot

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-pod-resource-requests-in-gke-autopilot

GKE Autopilot bills general-purpose Pods on the CPU, memory and ephemeral storage they request, not on what they use, and not on the nodes underneath.

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

GKE Autopilot bills general-purpose Pods on the CPU, memory and ephemeral storage they request, not on what they use, and not on the nodes underneath.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP GKE](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-gke)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0573

Type

Overprovisioned Resource

## Explanation

Why the waste happens and who it affects.

Every vCPU and GiB requested above real need is billed for as long as the Pod runs. Requests copied from Standard clusters or Helm chart defaults, generous safety margins, and containers with no requests at all are the common causes.

Autopilot also changes requests on its own. Containers without requests receive default values (for the general-purpose platform, 0.5 vCPU and 2 GiB of memory), requests below the minimum are raised, and when the CPU to memory ratio falls outside the allowed range Autopilot increases the smaller resource. Init containers without requests are given the sum of the application containers' requests, and are not adjusted when the app containers are later resized.

## Billing model

The pricing dimensions that drive this cost.

General-purpose Autopilot workloads use a Pod-based billing model.

Pod-based billing

Charged in one-second increments for the CPU, memory and ephemeral storage that running Pods request, with no minimum duration

Effective resource request

The larger of the largest single init container request and the sum of all application container requests in the Pod

Default and minimum requests

Applied automatically when requests are missing or too small, and billed like any other request

CPU to memory ratio adjustment

If a Pod's ratio is outside the compute class range, Autopilot raises the smaller resource, increasing the billed request

## How to detect

4 checks to find it in your estate.

- Review GKE workload rightsizing insights with subtype WORKLOAD\_OVERPROVISIONED, raised when CPU or memory utilization is under 50 percent for at least 90 percent of the time over the last 15 days, through the gcloud CLI or Recommender API; recommendations include a projected monthly saving where possible

- Create VerticalPodAutoscaler objects in Off mode for key workloads and compare their recommended requests with the current Pod specs

- Find containers with no requests set, which receive the compute class defaults, and Pods whose requests were raised by minimum or ratio enforcement by comparing the deployed spec with the manifest in source control

- Check init containers whose requests no longer match the application containers after a resize, since the effective request is the larger of the two

## How to fix

5 ways to remove the waste.

- Set explicit CPU and memory requests for every container based on observed usage plus headroom, as Google recommends instead of relying on defaults

- Use vertical Pod autoscaling, which is enabled by default in Autopilot clusters, in Initial, Recreate or InPlaceOrRecreate mode for workloads that tolerate it, or keep it in Off mode and apply its recommendations manually

- Choose requests whose CPU to memory ratio fits the compute class, or move the workload to a class whose ratio matches, so Autopilot does not raise one resource to fix the ratio

- After resizing application containers, update or remove init container requests so Autopilot recomputes them

- Where the cluster supports bursting, set requests to steady-state needs and limits higher to absorb start-up or traffic peaks instead of sizing requests for the peak

## Documentation

Vendor references for pricing and configuration.

- [Resource requests in Autopilot  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/autopilot-resource-requests)

- [Identify underprovisioned and overprovisioned workloads  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/optimize-workload-resource-utilization)

- [Vertical Pod autoscaling  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/verticalpodautoscaler)

- [Google Kubernetes Engine pricing  cloud.google.com](https://cloud.google.com/kubernetes-engine/pricing)

- [Best practices for running cost-optimized Kubernetes applications on GKE  docs.cloud.google.com](https://docs.cloud.google.com/architecture/best-practices-for-running-cost-effective-kubernetes-applications-on-gke)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP GKE  CER-0287

### [Spot-Only GKE Capacity Without Standard Fallback](https://www.pointfive.co/efficiency-hub/inefficiencies/spot-only-gke-capacity-without-standard-fallback)

Workloads are constrained to run only on Spot-based capacity with no viable path to standard nodes when Spot capacity is reclaimed or unavailable. While Spot reduces unit cost, rigid dependence can create hidden costs by requiring standby...

Compute

- GCP GKE  CER-0193

### [Orphaned and Overprovisioned Resources in GKE Clusters](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-and-overprovisioned-resources-in-gke-clusters)

As environments scale, GKE clusters tend to accumulate artifacts from ephemeral workloads, dev environments, or incomplete job execution. PVCs can continue to retain Persistent Disks, Services may continue to expose public IPs and...

Compute

- GCP GKE  CER-0270

### [Orphaned Kubernetes Resources in GKE](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-kubernetes-resources)

In GKE environments, it is common for unused Kubernetes resources to accumulate over time. Examples include Persistent Volume Claims (PVCs) that retain provisioned Persistent Disks, or Services of type LoadBalancer that continue to front...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

