Explanation
Why the waste happens and who it affects.
Every UPDATE, DELETE, MERGE, overwrite and OPTIMIZE writes new files and marks the old ones as removed in the transaction log, but the old files stay in cloud object storage so earlier table versions remain readable. They are only physically deleted when VACUUM runs and removes unreferenced files older than the retention threshold, 7 days by default. Databricks notes that deleting unused data files reduces cloud storage costs.
On tables that are rewritten often - MERGE-heavy CDC targets, tables fully overwritten on each run, tables compacted by frequent OPTIMIZE - unreferenced files accumulate with every rewrite and can exceed the size of the current table if VACUUM never runs. Predictive optimization runs VACUUM automatically, but only on Unity Catalog managed tables where it is enabled, so external tables, legacy Hive metastore tables and managed tables with predictive optimization disabled are the usual sources. The cost lands on the cloud provider's storage bill rather than as a Databricks line item, which is why it tends to go unnoticed. Tables using deletion vectors add a further case: deleted rows are only marked in metadata until the files are rewritten, so VACUUM alone does not reclaim them.
Billing model
The pricing dimensions that drive this cost.
- Object storage
- Every data file under the table path, referenced or not, is billed by the cloud provider per GB-month until it is deleted
- Retention threshold
- VACUUM removes only unreferenced files older than the retention period, 7 days by default and configurable with delta.deletedFileRetentionDuration
- Maintenance compute
- VACUUM run manually uses the cluster or warehouse it runs on; predictive optimization runs it on serverless compute billed under a serverless jobs SKU
How to detect
5 checks to find it in your estate.
- Compare DESCRIBE DETAIL sizeInBytes, which reflects the current table version, with the total size of objects under the table's storage location from a cloud storage inventory report; a large gap indicates unreferenced files
- Run DESCRIBE HISTORY on high-churn tables and check for VACUUM START and VACUUM END operations; tables with frequent MERGE, UPDATE, DELETE or OPTIMIZE operations and no recent VACUUM are the main candidates
- Check whether predictive optimization is enabled for the catalogs and schemas holding managed tables (DESCRIBE CATALOG EXTENDED or DESCRIBE SCHEMA EXTENDED) and review system.storage.predictive_optimization_operations_history to see which tables it is vacuuming
- List external tables and Hive metastore tables, which predictive optimization does not cover, and confirm each has a scheduled VACUUM
- Look for table properties that set delta.deletedFileRetentionDuration far above what time travel and recovery needs require
How to fix
5 ways to remove the waste.
- Enable predictive optimization for Unity Catalog managed tables (ALTER CATALOG or ALTER SCHEMA ... ENABLE PREDICTIVE OPTIMIZATION) so VACUUM, OPTIMIZE and ANALYZE run automatically
- Schedule VACUUM for external and legacy tables, using VACUUM ... DRY RUN first to preview deletions, and VACUUM LITE (Databricks Runtime 16.4 LTS and above) on very large tables where listing the whole directory is slow
- For tables with deletion vectors, run REORG TABLE ... APPLY (PURGE) before VACUUM so soft-deleted rows are rewritten out of data files and can be reclaimed
- Set delta.deletedFileRetentionDuration to the time travel window the business actually needs; VACUUM removes the ability to query versions older than the retention period, so confirm recovery and audit requirements first
- Size VACUUM compute as Databricks recommends - modest autoscaling workers with a larger driver for tables with many files - rather than running it on a large general-purpose cluster
Documentation
Vendor references for pricing and configuration.