Explanation
Why the waste happens and who it affects.
Every vCPU and GiB requested above real need is billed for as long as the Pod runs. Requests copied from Standard clusters or Helm chart defaults, generous safety margins, and containers with no requests at all are the common causes.
Autopilot also changes requests on its own. Containers without requests receive default values (for the general-purpose platform, 0.5 vCPU and 2 GiB of memory), requests below the minimum are raised, and when the CPU to memory ratio falls outside the allowed range Autopilot increases the smaller resource. Init containers without requests are given the sum of the application containers' requests, and are not adjusted when the app containers are later resized.
Billing model
The pricing dimensions that drive this cost.
General-purpose Autopilot workloads use a Pod-based billing model.
- Pod-based billing
- Charged in one-second increments for the CPU, memory and ephemeral storage that running Pods request, with no minimum duration
- Effective resource request
- The larger of the largest single init container request and the sum of all application container requests in the Pod
- Default and minimum requests
- Applied automatically when requests are missing or too small, and billed like any other request
- CPU to memory ratio adjustment
- If a Pod's ratio is outside the compute class range, Autopilot raises the smaller resource, increasing the billed request
How to detect
4 checks to find it in your estate.
- Review GKE workload rightsizing insights with subtype WORKLOAD_OVERPROVISIONED, raised when CPU or memory utilization is under 50 percent for at least 90 percent of the time over the last 15 days, through the gcloud CLI or Recommender API; recommendations include a projected monthly saving where possible
- Create VerticalPodAutoscaler objects in Off mode for key workloads and compare their recommended requests with the current Pod specs
- Find containers with no requests set, which receive the compute class defaults, and Pods whose requests were raised by minimum or ratio enforcement by comparing the deployed spec with the manifest in source control
- Check init containers whose requests no longer match the application containers after a resize, since the effective request is the larger of the two
How to fix
5 ways to remove the waste.
- Set explicit CPU and memory requests for every container based on observed usage plus headroom, as Google recommends instead of relying on defaults
- Use vertical Pod autoscaling, which is enabled by default in Autopilot clusters, in Initial, Recreate or InPlaceOrRecreate mode for workloads that tolerate it, or keep it in Off mode and apply its recommendations manually
- Choose requests whose CPU to memory ratio fits the compute class, or move the workload to a class whose ratio matches, so Autopilot does not raise one resource to fix the ratio
- After resizing application containers, update or remove init container requests so Autopilot recomputes them
- Where the cluster supports bursting, set requests to steady-state needs and limits higher to absorb start-up or traffic peaks instead of sizing requests for the peak
Documentation
Vendor references for pricing and configuration.
- Resource requests in Autopilotdocs.cloud.google.com
- Identify underprovisioned and overprovisioned workloadsdocs.cloud.google.com
- Vertical Pod autoscalingdocs.cloud.google.com
- Google Kubernetes Engine pricingcloud.google.com
- Best practices for running cost-optimized Kubernetes applications on GKEdocs.cloud.google.com