# Vertex AI Workbench Instances With Idle Shutdown Disabled

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/vertex-ai-workbench-instances-with-idle-shutdown-disabled

Vertex AI Workbench instances, now documented as Agent Platform Workbench under the Gemini Enterprise Agent Platform name, are notebook VMs that bill...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Vertex AI Workbench instances, now documented as Agent Platform Workbench under the Gemini Enterprise Agent Platform name, are notebook VMs that bill for CPU, memory and any attached GPUs whenever they are running.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Vertex AI](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-vertex-ai)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0556

Type

Inefficient Configuration

## Explanation

Why the waste happens and who it affects.

To help manage costs, Google enables idle shutdown by default and stops an instance after 180 minutes without kernel activity. Users can turn that off or raise the timeout.

The usual reason is that idle shutdown watches kernel activity, not CPU: a long computation that prints nothing can look idle, so data scientists disable the feature or set the timeout to a day to protect long runs. Those instances then keep running through nights, weekends and holidays after the work is done, often with GPUs attached. Google's AI and ML cost guidance asks teams to identify idle or underutilized VMs and GPUs and shut them down or rightsize them.

## Billing model

The pricing dimensions that drive this cost.

Compute and accelerators

VCPU, memory and GPU are billed while the instance is in STARTING, PROVISIONING, ACTIVE and other running states, plus Workbench management fees

Stopped instances

No CPU or GPU usage charges while shut down, except scheduled executions that run during the shutdown

Disk storage

Boot and data disks keep billing while the instance is stopped or suspended

## How to detect

5 checks to find it in your estate.

- List Workbench instances and read the idle-timeout-seconds metadata key; an empty value means idle shutdown is off, and values far above the 180-minute default deserve review

- On the Instance details page, check the Software and security tab for Enable Idle Shutdown and the Time of inactivity before shutdown setting

- Find instances in ACTIVE state outside working hours, and prioritize those with GPUs attached

- Check whether guest attributes are disabled or Jupyter was moved off port 8080, since idle shutdown depends on guest attributes and looks for kernel activity on port 8080

- Flag instances that have been stopped for a long time with large disks, which still bill for storage

## How to fix

4 ways to remove the waste.

- Re-enable idle shutdown or lower the timeout with gcloud workbench instances update INSTANCE\_NAME --metadata=idle-timeout-seconds=SECONDS, or in the console, choosing a value that fits how the team works

- Bake idle-timeout-seconds into Terraform modules and provisioning templates so new instances cannot be created without it

- Move long-running work out of interactive sessions: use scheduled notebook executions, which run even while the instance is shut down, or managed training jobs instead of keeping a notebook VM up

- Detach GPUs or switch to a smaller machine type for exploratory work, and delete abandoned instances and their disks after saving notebooks to source control

## Documentation

Vendor references for pricing and configuration.

- [Idle shutdown  docs.cloud.google.com](https://docs.cloud.google.com/gemini-enterprise-agent-platform/notebooks/workbench/instances/idle-shutdown)

- [Gemini Enterprise Agent Platform pricing  cloud.google.com](https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing)

- [AI and ML perspective: Cost optimization  docs.cloud.google.com](https://docs.cloud.google.com/architecture/framework/perspectives/ai-ml/cost-optimization)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Vertex AI  CER-0267

### [Using High-Cost Models for Low-Complexity Tasks in Vertex AI](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-models-for-low-complexity-tasks-bec92)

Vertex AI workloads often include low-complexity tasks such as classification, routing, keyword extraction, metadata parsing, document triage, or summarization of short and simple text. These operations do not require the advanced...

AI

- GCP Vertex AI  CER-0512

### [Underutilized Vertex AI Provisioned Throughput](https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-vertex-ai-provisioned-throughput)

Provisioned Throughput on Vertex AI, now documented under the Gemini Enterprise Agent Platform name, is a fixed-cost, fixed-term subscription that reserves throughput for a specific generative AI model in a specific region. It is bought in...

AI

- GCP Vertex AI  CER-0567

### [Idle Vertex AI Endpoints With Deployed Models](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-vertex-ai-endpoints-with-deployed-models)

When a model is deployed to a Vertex AI endpoint for online inference on dedicated resources (the service is now documented as Agent Platform Inference under the Gemini Enterprise Agent Platform name), each replica is a VM that bills...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

