Skip to content
Cloud Efficiency Hub

Always-On Structured Streaming Jobs for Batch-Tolerant Databricks Workloads

The short version

Structured Streaming jobs are often deployed as 24x7 jobs - a continuous job trigger or a run that never ends - so that ingestion from Kafka, Auto Loader or Delta sources is processed as soon as data lands.

PointFive Research

Cloud cost research at PointFive

Databricks service
Databricks Workflows
Category
Compute
Reference
CER-0528
Type
Inefficient Architecture

Explanation

Why the waste happens and who it affects.

That keeps job compute running around the clock. Databricks' cost-optimization best practices point out that many use cases built on a continuous stream of events do not need the events in the analytics data set immediately; if freshness every few hours or once a day is enough, a few runs per day meet the requirement at a much lower cost. Databricks recommends Structured Streaming with the AvailableNow trigger for incremental workloads without low-latency requirements.

The pattern persists because streaming is the default design for event data, and the stream's latency is seldom compared with what its consumers actually need: dashboards refreshed hourly, daily reports, ML features recomputed nightly. A second cost comes from the trigger itself: by default Structured Streaming uses a processingTime trigger of 0 ms and checks for new data every few milliseconds, which Databricks warns can generate a high volume of cloud storage API calls and unexpected charges from the cloud provider. Low-volume streams that stay up all day pay for both idle compute and constant polling.

Billing model

The pricing dimensions that drive this cost.

Always-on job compute
DBUs per second plus, on classic compute, cloud instance charges for every hour the streaming job runs, whether or not data arrives
Scheduled incremental runs
With AvailableNow, the job processes all available data and stops, so compute is billed only for the length of each run
Cloud storage API requests
List and get requests made by the stream when checking for new data are billed by the cloud provider, and the default 0 ms trigger makes them frequent

How to detect

4 checks to find it in your estate.

  • Query system.lakeflow.jobs for trigger_type = 'CONTINUOUS' and system.lakeflow.job_run_timeline for runs that span most hours of the day, then map them to job DBUs in system.billing.usage by usage_metadata.job_id
  • For each always-on streaming job, find how the output tables are consumed: dashboard refresh schedules, downstream job schedules and query history on the sink tables; flag streams whose consumers read on an hourly or daily cadence
  • Check the streaming query metrics (numInputRows and inputRowsPerSecond in the Spark UI Structured Streaming tab or a StreamingQueryListener) for streams that process little data most of the time
  • Review stream code for writeStream calls without an explicit trigger, which fall back to the 0 ms processingTime default, and review the cloud provider bill for storage request charges on the source buckets

How to fix

5 ways to remove the waste.

  • Convert batch-tolerant streams to scheduled incremental runs: use .trigger(availableNow=True) and run the job on a CRON or file-arrival trigger at the freshness consumers need; checkpoints preserve exactly-once progress between runs
  • Replace Trigger.Once with AvailableNow, since Trigger.Once is deprecated in Databricks Runtime 11.3 LTS and above
  • For streams that must stay on, set an explicit processingTime interval (for example 10 seconds or longer) that matches latency needs, to reduce storage API calls and micro-batch overhead
  • Right-size the cluster for scheduled runs, or run them on serverless jobs compute, instead of keeping a cluster sized for a continuous stream
  • Agree freshness targets with data consumers and document them on the job, so always-on streaming is kept only where a real low-latency requirement exists

Documentation

Vendor references for pricing and configuration.