Skip to content
Cloud Efficiency Hub

Underutilized GPU Nodes in GKE Without GPU Sharing

The short version

Kubernetes Pods request GPUs in whole numbers, so by default an entire physical GPU is allocated to one container even when that container needs only a fraction of it.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP GKE
Category
AI
Reference
CER-0501
Type
Overprovisioned Resource

Explanation

Why the waste happens and who it affects.

Small inference services, notebooks, development environments and light batch jobs each end up holding a full accelerator, and the GPU sits mostly idle while the node that carries it bills at GPU rates.

The waste compounds when GPU node pools are sized for peak demand and never scale back, or when a fixed minimum node count keeps GPU nodes running with no GPU Pods scheduled. GKE offers three GPU sharing strategies - multi-instance GPU, GPU time-sharing and NVIDIA MPS - that let several containers use one GPU, and Google's documentation positions them as a way to minimize underutilized capacity and save running costs. Platform teams that serve many small AI workloads are the most affected.

Billing model

The pricing dimensions that drive this cost.

GPU capacity in GKE is paid for per node, not per unit of GPU work done.

Standard GPU node pools
The node's Compute Engine VM and attached GPUs are billed until the node is deleted, whether or not GPU Pods are running
Autopilot accelerator workloads
Use node-based billing, charged at Compute Engine pricing for the whole node Autopilot creates plus an Autopilot management premium
Whole-GPU allocation
Without a sharing strategy, each container that requests a GPU is allocated at least one full physical GPU

How to detect

4 checks to find it in your estate.

  • Chart container/accelerator/duty_cycle and container/accelerator/memory_used against memory_total in Cloud Monitoring, or DCGM metrics such as DCGM_FI_DEV_GPU_UTIL and DCGM_FI_DEV_FB_USED, per node and per container
  • Find GPUs whose duty cycle and memory use stay low for days while one small Pod holds the whole device
  • List GPU node pools and check whether a GPU sharing strategy is configured (Standard node pool accelerator settings, or cloud.google.com/gke-gpu-sharing-strategy node selectors in Autopilot)
  • Check GPU node pools for autoscaling minimums above zero and for GPU nodes that have no GPU Pods scheduled

How to fix

5 ways to remove the waste.

  • Use multi-instance GPU to split a supported GPU into up to seven hardware-isolated partitions for parallel inference workloads that need predictable quality of service
  • Use GPU time-sharing for bursty or interactive workloads with idle periods, accepting that it enforces no memory limits between shared workloads and adds context-switch overhead
  • Use NVIDIA MPS for cooperative CUDA workloads such as small batch jobs, where the workloads can tolerate its memory protection and error containment limits
  • Enable autoscaling on GPU node pools with --min-nodes 0 so they scale down when unused, and keep CPU-only Pods off GPU nodes with the nvidia.com/gpu taint
  • Match the GPU type and count per node to what the workload actually uses, and consider Spot VMs for interruption-tolerant GPU jobs

Documentation

Vendor references for pricing and configuration.