Explanation
Why the waste happens and who it affects.
To help manage costs, Google enables idle shutdown by default and stops an instance after 180 minutes without kernel activity. Users can turn that off or raise the timeout.
The usual reason is that idle shutdown watches kernel activity, not CPU: a long computation that prints nothing can look idle, so data scientists disable the feature or set the timeout to a day to protect long runs. Those instances then keep running through nights, weekends and holidays after the work is done, often with GPUs attached. Google's AI and ML cost guidance asks teams to identify idle or underutilized VMs and GPUs and shut them down or rightsize them.
Billing model
The pricing dimensions that drive this cost.
- Compute and accelerators
- VCPU, memory and GPU are billed while the instance is in STARTING, PROVISIONING, ACTIVE and other running states, plus Workbench management fees
- Stopped instances
- No CPU or GPU usage charges while shut down, except scheduled executions that run during the shutdown
- Disk storage
- Boot and data disks keep billing while the instance is stopped or suspended
How to detect
5 checks to find it in your estate.
- List Workbench instances and read the idle-timeout-seconds metadata key; an empty value means idle shutdown is off, and values far above the 180-minute default deserve review
- On the Instance details page, check the Software and security tab for Enable Idle Shutdown and the Time of inactivity before shutdown setting
- Find instances in ACTIVE state outside working hours, and prioritize those with GPUs attached
- Check whether guest attributes are disabled or Jupyter was moved off port 8080, since idle shutdown depends on guest attributes and looks for kernel activity on port 8080
- Flag instances that have been stopped for a long time with large disks, which still bill for storage
How to fix
4 ways to remove the waste.
- Re-enable idle shutdown or lower the timeout with gcloud workbench instances update INSTANCE_NAME --metadata=idle-timeout-seconds=SECONDS, or in the console, choosing a value that fits how the team works
- Bake idle-timeout-seconds into Terraform modules and provisioning templates so new instances cannot be created without it
- Move long-running work out of interactive sessions: use scheduled notebook executions, which run even while the instance is shut down, or managed training jobs instead of keeping a notebook VM up
- Detach GPUs or switch to a smaller machine type for exploratory work, and delete abandoned instances and their disks after saving notebooks to source control
Documentation
Vendor references for pricing and configuration.