# Idle Dataproc Cluster Without Scheduled Deletion

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/idle-dataproc-cluster-without-scheduled-deletion

Dataproc clusters (now branded Managed Service for Apache Spark) bill for their Compute Engine VMs and Persistent Disks plus a per-vCPU management fee...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Dataproc clusters (now branded Managed Service for Apache Spark) bill for their Compute Engine VMs and Persistent Disks plus a per-vCPU management fee for as long as the cluster is running, whether or not any job is executing.

PointFive Research

Cloud cost research at PointFive

GCP service

[GCP Dataproc](https://www.pointfive.co/efficiency-hub/cloud-services/gcp-dataproc)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0491

Type

Idle or Unused Resource

## Explanation

Why the waste happens and who it affects.

Clusters created for ad hoc analysis, notebooks, one-off backfills or pipeline development are often left up after the work finishes, overnight or for weeks. Clusters in an error state also keep their VMs active and keep accruing charges until they are deleted.

Google provides scheduled deletion and scheduled stop specifically to avoid charges for inactive clusters, but both have to be set explicitly on each cluster; a cluster created without them runs until someone deletes it.

## Billing model

The pricing dimensions that drive this cost.

Cluster pricing is in addition to the Compute Engine price of each VM; rates vary by region.

Compute Engine VMs and disks

Master, primary worker and secondary worker VMs and their Persistent Disks billed at Compute Engine rates while they exist

Management fee

Billed per vCPU-hour across all master, worker and secondary worker nodes, per second with a 1-minute minimum, while the cluster runs

Error-state clusters

Cluster VMs remain active and charges continue until the cluster is deleted

Stopped clusters

Stopping a cluster stops its VMs, but associated resources such as Persistent Disks keep billing

## How to detect

4 checks to find it in your estate.

- List clusters (gcloud dataproc clusters list) and inspect config.lifecycleConfig for idleDeleteTtl, autoDeleteTime or autoDeleteTtl; clusters with none of these never delete themselves

- Find running clusters with no Dataproc jobs and no YARN applications over a representative period, using the jobs list and YARN application history

- Flag clusters in ERROR state, which keep billing until deleted

- Attribute spend per cluster from the Cloud Billing export (management fee plus the Compute Engine resources labeled for the cluster) to prioritize the largest idle clusters

## How to fix

4 ways to remove the waste.

- Set scheduled deletion on ephemeral and interactive clusters at create time or by updating existing clusters: --delete-max-idle (idle period, 5 minutes to 14 days), --delete-expiration-time or --delete-max-age

- Idle time counts both YARN and Dataproc Jobs API activity by default (dataproc:dataproc.cluster-ttl.consider-yarn-activity); keep it on so long-running YARN sessions are not treated as idle

- For clusters that must persist, use scheduled stop (--stop-max-idle, --stop-expiration-time or --stop-max-age), accepting that disks keep billing while stopped and that clusters with secondary workers, local SSDs or flexible VMs cannot be stopped

- Run batch work on job-scoped ephemeral clusters or on the serverless deployment model, which bills per second for the compute units a workload consumes instead of for a standing cluster, and delete error-state clusters promptly

## Documentation

Vendor references for pricing and configuration.

- [Cluster Scheduled Deletion  docs.cloud.google.com](https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/scheduled-deletion)

- [Cluster Scheduled Stop  docs.cloud.google.com](https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/scheduled-stop)

- [Stop and start clusters  docs.cloud.google.com](https://docs.cloud.google.com/managed-spark/docs/guides/start-stop)

- [Managed Service for Apache Spark (formerly Dataproc) pricing  cloud.google.com](https://cloud.google.com/products/managed-service-for-apache-spark/pricing)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- GCP Dataflow  CER-0244

### [Idle Dataflow Workers Running After Pipeline Failure](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-dataflow-workers-running-after-pipeline-failure-f6b1a)

When a Dataflow pipeline fails - often due to dependency issues, misconfigurations, or data format mismatches-its worker instances may remain active temporarily until the service terminates them. In some cases, misconfigured jobs, stuck...

Compute

- GCP GKE  CER-0193

### [Orphaned and Overprovisioned Resources in GKE Clusters](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-and-overprovisioned-resources-in-gke-clusters)

As environments scale, GKE clusters tend to accumulate artifacts from ephemeral workloads, dev environments, or incomplete job execution. PVCs can continue to retain Persistent Disks, Services may continue to expose public IPs and...

Compute

- GCP GKE  CER-0270

### [Orphaned Kubernetes Resources in GKE](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-kubernetes-resources)

In GKE environments, it is common for unused Kubernetes resources to accumulate over time. Examples include Persistent Volume Claims (PVCs) that retain provisioned Persistent Disks, or Services of type LoadBalancer that continue to front...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

