Skip to content
Cloud Efficiency Hub

Vertex AI Workbench Instances With Idle Shutdown Disabled

The short version

Vertex AI Workbench instances, now documented as Agent Platform Workbench under the Gemini Enterprise Agent Platform name, are notebook VMs that bill for CPU, memory and any attached GPUs whenever they are running.

PointFive Research

Cloud cost research at PointFive

GCP service
GCP Vertex AI
Category
AI
Reference
CER-0556
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

To help manage costs, Google enables idle shutdown by default and stops an instance after 180 minutes without kernel activity. Users can turn that off or raise the timeout.

The usual reason is that idle shutdown watches kernel activity, not CPU: a long computation that prints nothing can look idle, so data scientists disable the feature or set the timeout to a day to protect long runs. Those instances then keep running through nights, weekends and holidays after the work is done, often with GPUs attached. Google's AI and ML cost guidance asks teams to identify idle or underutilized VMs and GPUs and shut them down or rightsize them.

Billing model

The pricing dimensions that drive this cost.

Compute and accelerators
VCPU, memory and GPU are billed while the instance is in STARTING, PROVISIONING, ACTIVE and other running states, plus Workbench management fees
Stopped instances
No CPU or GPU usage charges while shut down, except scheduled executions that run during the shutdown
Disk storage
Boot and data disks keep billing while the instance is stopped or suspended

How to detect

5 checks to find it in your estate.

  • List Workbench instances and read the idle-timeout-seconds metadata key; an empty value means idle shutdown is off, and values far above the 180-minute default deserve review
  • On the Instance details page, check the Software and security tab for Enable Idle Shutdown and the Time of inactivity before shutdown setting
  • Find instances in ACTIVE state outside working hours, and prioritize those with GPUs attached
  • Check whether guest attributes are disabled or Jupyter was moved off port 8080, since idle shutdown depends on guest attributes and looks for kernel activity on port 8080
  • Flag instances that have been stopped for a long time with large disks, which still bill for storage

How to fix

4 ways to remove the waste.

  • Re-enable idle shutdown or lower the timeout with gcloud workbench instances update INSTANCE_NAME --metadata=idle-timeout-seconds=SECONDS, or in the console, choosing a value that fits how the team works
  • Bake idle-timeout-seconds into Terraform modules and provisioning templates so new instances cannot be created without it
  • Move long-running work out of interactive sessions: use scheduled notebook executions, which run even while the instance is shut down, or managed training jobs instead of keeping a notebook VM up
  • Detach GPUs or switch to a smaller machine type for exploratory work, and delete abandoned instances and their disks after saving notebooks to source control

Documentation

Vendor references for pricing and configuration.