Skip to content
Cloud Efficiency Hub

Overprovisioned Pod Resource Requests in GKE Autopilot

The short version

GKE Autopilot bills general-purpose Pods on the CPU, memory and ephemeral storage they request, not on what they use, and not on the nodes underneath.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP GKE
Category
Compute
Reference
CER-0573
Type
Overprovisioned Resource

Explanation

Why the waste happens and who it affects.

Every vCPU and GiB requested above real need is billed for as long as the Pod runs. Requests copied from Standard clusters or Helm chart defaults, generous safety margins, and containers with no requests at all are the common causes.

Autopilot also changes requests on its own. Containers without requests receive default values (for the general-purpose platform, 0.5 vCPU and 2 GiB of memory), requests below the minimum are raised, and when the CPU to memory ratio falls outside the allowed range Autopilot increases the smaller resource. Init containers without requests are given the sum of the application containers' requests, and are not adjusted when the app containers are later resized.

Billing model

The pricing dimensions that drive this cost.

General-purpose Autopilot workloads use a Pod-based billing model.

Pod-based billing
Charged in one-second increments for the CPU, memory and ephemeral storage that running Pods request, with no minimum duration
Effective resource request
The larger of the largest single init container request and the sum of all application container requests in the Pod
Default and minimum requests
Applied automatically when requests are missing or too small, and billed like any other request
CPU to memory ratio adjustment
If a Pod's ratio is outside the compute class range, Autopilot raises the smaller resource, increasing the billed request

How to detect

4 checks to find it in your estate.

  • Review GKE workload rightsizing insights with subtype WORKLOAD_OVERPROVISIONED, raised when CPU or memory utilization is under 50 percent for at least 90 percent of the time over the last 15 days, through the gcloud CLI or Recommender API; recommendations include a projected monthly saving where possible
  • Create VerticalPodAutoscaler objects in Off mode for key workloads and compare their recommended requests with the current Pod specs
  • Find containers with no requests set, which receive the compute class defaults, and Pods whose requests were raised by minimum or ratio enforcement by comparing the deployed spec with the manifest in source control
  • Check init containers whose requests no longer match the application containers after a resize, since the effective request is the larger of the two

How to fix

5 ways to remove the waste.

  • Set explicit CPU and memory requests for every container based on observed usage plus headroom, as Google recommends instead of relying on defaults
  • Use vertical Pod autoscaling, which is enabled by default in Autopilot clusters, in Initial, Recreate or InPlaceOrRecreate mode for workloads that tolerate it, or keep it in Off mode and apply its recommendations manually
  • Choose requests whose CPU to memory ratio fits the compute class, or move the workload to a class whose ratio matches, so Autopilot does not raise one resource to fix the ratio
  • After resizing application containers, update or remove init container requests so Autopilot recomputes them
  • Where the cluster supports bursting, set requests to steady-state needs and limits higher to absorb start-up or traffic peaks instead of sizing requests for the peak

Documentation

Vendor references for pricing and configuration.