Explanation
Why the waste happens and who it affects.
In triggered mode, the default, an update refreshes all tables from the data available when it starts and then stops. In continuous mode, the pipeline keeps processing new data as it arrives. Databricks' pipeline-mode documentation states the cost difference directly: triggered pipelines reduce resource consumption and expense because the cluster runs only long enough to update the pipeline, while continuous pipelines require an always-running cluster, which is more expensive.
Pipelines are switched to continuous for convenience, during development, or because a source is a stream, even when every consumer of the output tables reads them on an hourly or daily schedule. Once set, the pipeline's compute runs 24x7 whether or not any consumer needs that freshness. This applies to pipelines on both classic and serverless compute; Databricks' broader cost guidance to balance always-on and triggered streaming makes the same point for streaming workloads generally.
Billing model
The pricing dimensions that drive this cost.
- Pipeline compute
- Billed in DBUs for as long as the pipeline's compute runs, plus cloud instance charges on classic compute; the product tier (dlt_tier CORE, PRO or ADVANCED) sets the classic DBU rate
- Triggered mode
- Compute runs only for the duration of each update, so cost follows the update schedule
- Continuous mode
- Compute stays up between arrivals of new data, so the pipeline bills around the clock
How to detect
4 checks to find it in your estate.
- Query the latest record per pipeline in system.lakeflow.pipelines (Public Preview) for settings.continuous = true and settings.development flags, and list the owners
- Join to system.billing.usage where billing_origin_product = 'DLT' by usage_metadata.dlt_pipeline_id to confirm 24x7 DBU consumption and rank by cost
- For each continuous pipeline, check how its target tables are consumed - dashboard refresh schedules, downstream jobs, query history on the tables - and flag pipelines whose consumers read on an hourly or longer cadence
- Review the Pipeline mode setting in the pipeline UI or the continuous field in pipeline JSON and bundle definitions to find pipelines set to continuous without a documented latency requirement
How to fix
5 ways to remove the waste.
- Switch batch-tolerant pipelines to Triggered in the Pipeline mode setting and schedule updates with a job at the freshness consumers need; streaming tables still process only new data incrementally on each triggered update
- For serverless triggered pipelines, also consider clearing Performance optimized in the schedule to use standard performance mode where a start of about four to six minutes is acceptable
- Where continuous processing is truly required, follow Databricks' recommendation to run the pipeline with a continuous job rather than the pipeline's built-in continuous mode, which enables serverless performance modes
- For continuous pipelines that stay on, set pipelines.trigger.interval on individual flows to process less frequently where near-real-time is not needed
- Record the freshness requirement on each pipeline and revisit continuous pipelines when their consumers change
Documentation
Vendor references for pricing and configuration.