Explanation
Why the waste happens and who it affects.
While an instance sits idle in a pool, Databricks charges no DBUs for it, but the cloud provider still bills the virtual machine. A pool's Minimum Idle Instances setting keeps that many instances running at all times, and the idle instance auto termination timer only removes idle instances above that minimum, so any non-zero minimum is warm capacity paid for around the clock.
The waste is easy to miss because it never appears in Databricks usage: the cost shows up on the AWS, Azure or Google Cloud bill as ordinary instances tagged with the pool ID. Pools are commonly created with a minimum idle count to shave start time off scheduled jobs and then left unchanged through nights, weekends and quiet periods, or kept in non-production workspaces where start latency does not matter. Long idle auto termination times add to this by holding released instances well past the next job that could reuse them. Databricks' own pool best practices recommend setting Min Idle to 0 to avoid paying for running instances that are not doing work.
Billing model
The pricing dimensions that drive this cost.
- Idle pool instances
- No DBUs are charged while an instance is idle in the pool, but the cloud provider bills the instance for as long as it runs
- Minimum Idle Instances
- Instances the pool always keeps idle; auto termination does not remove them
- Idle instance auto termination
- Minutes an instance above the minimum can stay idle before the pool terminates it; cloud billing continues until then
- Instances in use
- Once attached to a cluster, instances are billed as normal cluster compute, DBUs plus cloud instance charges
How to detect
5 checks to find it in your estate.
- List pools with the Instance Pools API (GET /api/2.0/instance-pools/list) and flag pools with min_idle_instances greater than 0, especially in development and test workspaces
- Check stats.idle_count against stats.used_count over time for each pool; pools that mostly hold idle instances are paying for capacity no cluster is using
- Compare idle_instance_autotermination_minutes with the actual gap between the jobs that use the pool; timers far longer than the gap keep instances running without reuse
- In the cloud provider bill, filter instances by the DatabricksInstancePoolId tag and compare instance hours with DBU usage for the same pool in system.billing.usage (usage_metadata.instance_pool_id) to estimate idle instance hours
- Identify pools that no cluster, job or policy references any more but that still hold minimum idle instances
How to fix
5 ways to remove the waste.
- Set Minimum Idle Instances to 0 wherever a few minutes of instance acquisition time is acceptable, which Databricks recommends to avoid paying for idle instances
- Where warm capacity is needed for a latency-sensitive schedule, keep the minimum only during that window, for example by raising and lowering min_idle_instances through the Instance Pools API on a schedule, or pre-populate the pool with a starter job
- Shorten idle instance auto termination to a small buffer above the typical gap between jobs that reuse the pool
- Delete pools that are no longer referenced by any cluster or job, and set Max Capacity where budgets or quotas need a hard ceiling
- Use spot instances for worker pools where interruptions are acceptable and apply chargeback tags to pools, since Databricks tags propagate to the cloud instances and make idle pool cost visible to owners
Documentation
Vendor references for pricing and configuration.