# Underutilized GPU Nodes in GKE Without GPU Sharing

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-gpu-nodes-in-gke-without-gpu-sharing

Kubernetes Pods request GPUs in whole numbers, so by default an entire physical GPU is allocated to one container even when that container needs only a...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Kubernetes Pods request GPUs in whole numbers, so by default an entire physical GPU is allocated to one container even when that container needs only a fraction of it.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP GKE](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-gke)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0501

Type

Overprovisioned Resource

## Explanation

Why the waste happens and who it affects.

Small inference services, notebooks, development environments and light batch jobs each end up holding a full accelerator, and the GPU sits mostly idle while the node that carries it bills at GPU rates.

The waste compounds when GPU node pools are sized for peak demand and never scale back, or when a fixed minimum node count keeps GPU nodes running with no GPU Pods scheduled. GKE offers three GPU sharing strategies - multi-instance GPU, GPU time-sharing and NVIDIA MPS - that let several containers use one GPU, and Google's documentation positions them as a way to minimize underutilized capacity and save running costs. Platform teams that serve many small AI workloads are the most affected.

## Billing model

The pricing dimensions that drive this cost.

GPU capacity in GKE is paid for per node, not per unit of GPU work done.

Standard GPU node pools

The node's Compute Engine VM and attached GPUs are billed until the node is deleted, whether or not GPU Pods are running

Autopilot accelerator workloads

Use node-based billing, charged at Compute Engine pricing for the whole node Autopilot creates plus an Autopilot management premium

Whole-GPU allocation

Without a sharing strategy, each container that requests a GPU is allocated at least one full physical GPU

## How to detect

4 checks to find it in your estate.

- Chart container/accelerator/duty\_cycle and container/accelerator/memory\_used against memory\_total in Cloud Monitoring, or DCGM metrics such as DCGM\_FI\_DEV\_GPU\_UTIL and DCGM\_FI\_DEV\_FB\_USED, per node and per container

- Find GPUs whose duty cycle and memory use stay low for days while one small Pod holds the whole device

- List GPU node pools and check whether a GPU sharing strategy is configured (Standard node pool accelerator settings, or cloud.google.com/gke-gpu-sharing-strategy node selectors in Autopilot)

- Check GPU node pools for autoscaling minimums above zero and for GPU nodes that have no GPU Pods scheduled

## How to fix

5 ways to remove the waste.

- Use multi-instance GPU to split a supported GPU into up to seven hardware-isolated partitions for parallel inference workloads that need predictable quality of service

- Use GPU time-sharing for bursty or interactive workloads with idle periods, accepting that it enforces no memory limits between shared workloads and adds context-switch overhead

- Use NVIDIA MPS for cooperative CUDA workloads such as small batch jobs, where the workloads can tolerate its memory protection and error containment limits

- Enable autoscaling on GPU node pools with --min-nodes 0 so they scale down when unused, and keep CPU-only Pods off GPU nodes with the nvidia.com/gpu taint

- Match the GPU type and count per node to what the workload actually uses, and consider Spot VMs for interruption-tolerant GPU jobs

## Documentation

Vendor references for pricing and configuration.

- [About GPU sharing strategies in GKE  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/timesharing-gpus)

- [About GPUs in Google Kubernetes Engine (GKE)  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/concepts/gpus)

- [Run GPUs in GKE Standard node pools  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/gpus)

- [Collect and view DCGM metrics  docs.cloud.google.com](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/dcgm-metrics)

- [Google Kubernetes Engine pricing  cloud.google.com](https://cloud.google.com/kubernetes-engine/pricing)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP GKE  CER-0573

### [Overprovisioned Pod Resource Requests in GKE Autopilot](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-pod-resource-requests-in-gke-autopilot)

GKE Autopilot bills general-purpose Pods on the CPU, memory and ephemeral storage they request, not on what they use, and not on the nodes underneath. Every vCPU and GiB requested above real need is billed for as long as the Pod runs....

Compute

- GCP GKE  CER-0287

### [Spot-Only GKE Capacity Without Standard Fallback](https://www.pointfive.co/efficiency-hub/inefficiencies/spot-only-gke-capacity-without-standard-fallback)

Workloads are constrained to run only on Spot-based capacity with no viable path to standard nodes when Spot capacity is reclaimed or unavailable. While Spot reduces unit cost, rigid dependence can create hidden costs by requiring standby...

Compute

- GCP GKE  CER-0193

### [Orphaned and Overprovisioned Resources in GKE Clusters](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-and-overprovisioned-resources-in-gke-clusters)

As environments scale, GKE clusters tend to accumulate artifacts from ephemeral workloads, dev environments, or incomplete job execution. PVCs can continue to retain Persistent Disks, Services may continue to expose public IPs and...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

