Skip to content
Cloud Efficiency Hub

Idle Instances Retained in Databricks Instance Pools

The short version

Databricks instance pools keep a set of ready cloud instances so classic clusters can start and scale faster.

PointFive Research

Cloud cost research at PointFive

Databricks service
Databricks Compute
Category
Compute
Reference
CER-0523
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

While an instance sits idle in a pool, Databricks charges no DBUs for it, but the cloud provider still bills the virtual machine. A pool's Minimum Idle Instances setting keeps that many instances running at all times, and the idle instance auto termination timer only removes idle instances above that minimum, so any non-zero minimum is warm capacity paid for around the clock.

The waste is easy to miss because it never appears in Databricks usage: the cost shows up on the AWS, Azure or Google Cloud bill as ordinary instances tagged with the pool ID. Pools are commonly created with a minimum idle count to shave start time off scheduled jobs and then left unchanged through nights, weekends and quiet periods, or kept in non-production workspaces where start latency does not matter. Long idle auto termination times add to this by holding released instances well past the next job that could reuse them. Databricks' own pool best practices recommend setting Min Idle to 0 to avoid paying for running instances that are not doing work.

Billing model

The pricing dimensions that drive this cost.

Idle pool instances
No DBUs are charged while an instance is idle in the pool, but the cloud provider bills the instance for as long as it runs
Minimum Idle Instances
Instances the pool always keeps idle; auto termination does not remove them
Idle instance auto termination
Minutes an instance above the minimum can stay idle before the pool terminates it; cloud billing continues until then
Instances in use
Once attached to a cluster, instances are billed as normal cluster compute, DBUs plus cloud instance charges

How to detect

5 checks to find it in your estate.

  • List pools with the Instance Pools API (GET /api/2.0/instance-pools/list) and flag pools with min_idle_instances greater than 0, especially in development and test workspaces
  • Check stats.idle_count against stats.used_count over time for each pool; pools that mostly hold idle instances are paying for capacity no cluster is using
  • Compare idle_instance_autotermination_minutes with the actual gap between the jobs that use the pool; timers far longer than the gap keep instances running without reuse
  • In the cloud provider bill, filter instances by the DatabricksInstancePoolId tag and compare instance hours with DBU usage for the same pool in system.billing.usage (usage_metadata.instance_pool_id) to estimate idle instance hours
  • Identify pools that no cluster, job or policy references any more but that still hold minimum idle instances

How to fix

5 ways to remove the waste.

  • Set Minimum Idle Instances to 0 wherever a few minutes of instance acquisition time is acceptable, which Databricks recommends to avoid paying for idle instances
  • Where warm capacity is needed for a latency-sensitive schedule, keep the minimum only during that window, for example by raising and lowering min_idle_instances through the Instance Pools API on a schedule, or pre-populate the pool with a starter job
  • Shorten idle instance auto termination to a small buffer above the typical gap between jobs that reuse the pool
  • Delete pools that are no longer referenced by any cluster or job, and set Max Capacity where budgets or quotas need a hard ceiling
  • Use spot instances for worker pools where interruptions are acceptable and apply chargeback tags to pools, since Databricks tags propagate to the cloud instances and make idle pool cost visible to owners

Documentation

Vendor references for pricing and configuration.