Explanation
Why the waste happens and who it affects.
Clusters created for ad hoc analysis, notebooks, one-off backfills or pipeline development are often left up after the work finishes, overnight or for weeks. Clusters in an error state also keep their VMs active and keep accruing charges until they are deleted.
Google provides scheduled deletion and scheduled stop specifically to avoid charges for inactive clusters, but both have to be set explicitly on each cluster; a cluster created without them runs until someone deletes it.
Billing model
The pricing dimensions that drive this cost.
Cluster pricing is in addition to the Compute Engine price of each VM; rates vary by region.
- Compute Engine VMs and disks
- Master, primary worker and secondary worker VMs and their Persistent Disks billed at Compute Engine rates while they exist
- Management fee
- Billed per vCPU-hour across all master, worker and secondary worker nodes, per second with a 1-minute minimum, while the cluster runs
- Error-state clusters
- Cluster VMs remain active and charges continue until the cluster is deleted
- Stopped clusters
- Stopping a cluster stops its VMs, but associated resources such as Persistent Disks keep billing
How to detect
4 checks to find it in your estate.
- List clusters (gcloud dataproc clusters list) and inspect config.lifecycleConfig for idleDeleteTtl, autoDeleteTime or autoDeleteTtl; clusters with none of these never delete themselves
- Find running clusters with no Dataproc jobs and no YARN applications over a representative period, using the jobs list and YARN application history
- Flag clusters in ERROR state, which keep billing until deleted
- Attribute spend per cluster from the Cloud Billing export (management fee plus the Compute Engine resources labeled for the cluster) to prioritize the largest idle clusters
How to fix
4 ways to remove the waste.
- Set scheduled deletion on ephemeral and interactive clusters at create time or by updating existing clusters: --delete-max-idle (idle period, 5 minutes to 14 days), --delete-expiration-time or --delete-max-age
- Idle time counts both YARN and Dataproc Jobs API activity by default (dataproc:dataproc.cluster-ttl.consider-yarn-activity); keep it on so long-running YARN sessions are not treated as idle
- For clusters that must persist, use scheduled stop (--stop-max-idle, --stop-expiration-time or --stop-max-age), accepting that disks keep billing while stopped and that clusters with secondary workers, local SSDs or flexible VMs cannot be stopped
- Run batch work on job-scoped ephemeral clusters or on the serverless deployment model, which bills per second for the compute units a workload consumes instead of for a standing cluster, and delete error-state clusters promptly
Documentation
Vendor references for pricing and configuration.