Explanation
Why the waste happens and who it affects.
When scale to zero is not enabled, the endpoint never drops below its minimum provisioned concurrency, and Databricks' pricing FAQ states that with no traffic the minimum charge depends on the minimum provisioned concurrency of the chosen range. An endpoint serving a handful of requests a day, or none, therefore bills for its floor capacity every hour.
This is typical for development, staging and demo endpoints, endpoints for models that have been superseded, and low-traffic internal tools. Scale to zero is optional and off unless selected, and teams often raise the minimum concurrency or pick a GPU workload type during testing and never revisit it. GPU endpoints make the idle floor expensive, since GPU serving is billed per GPU instance per hour. Scale to zero brings cold starts and no guarantee of GPU capacity when scaling back up, and Databricks advises against it for production workloads that need consistent uptime, so the fix applies mainly to non-production and low-traffic endpoints.
Billing model
The pricing dimensions that drive this cost.
Model serving usage appears on the bill under the Serverless Real-time Inference SKU.
- CPU serving
- Billed in DBUs based on the compute nodes used for the endpoint's provisioned concurrency, with the cloud instance cost included in the DBU price
- GPU serving
- Billed per GPU instance per hour at a DBU rate set by GPU configuration, listed on the Model Serving pricing page
- Minimum provisioned concurrency
- Without scale to zero, the endpoint bills at least for the minimum concurrency of its configured range even when it receives no requests
- Scale to zero
- After 30 minutes with no requests the endpoint scales to zero and is not charged; charges resume when a new request triggers scale-up
How to detect
5 checks to find it in your estate.
- Query system.billing.usage where billing_origin_product = 'MODEL_SERVING' and group by usage_metadata.endpoint_name and day; flat daily DBUs over several weeks indicate an endpoint running at its floor
- Use product_features.serving_type (MODEL or GPU_MODEL) in the same table to separate CPU and GPU custom model endpoints and prioritize GPU ones
- List endpoint configurations through the Serving Endpoints API and flag served entities with scale_to_zero_enabled = false, a large workload_size, or a min_provisioned_concurrency above what traffic needs, outside production
- Check the endpoint health metrics in the Serving UI (request rate, latency, CPU and memory usage for the last 14 days) to confirm near-zero request rates on endpoints with steady DBUs
- Look for synthetic health checks or monitoring probes that send periodic requests and keep a scale-to-zero endpoint from ever reaching 30 minutes of inactivity
How to fix
5 ways to remove the waste.
- Enable scale to zero on development, test and low-traffic endpoints by updating the served entity configuration with scale_to_zero_enabled = true, accepting cold starts of typically 10 to 20 seconds and sometimes minutes on the first request
- Lower workload_size or min_provisioned_concurrency (multiples of 4) to what peak traffic needs, sizing with Databricks' guidance that provisioned concurrency equals queries per second times model execution time
- Stop custom model endpoints that must keep their configuration but are not needed for a period (POST /api/2.0/serving-endpoints/{name}/config:stop, or the Stop button); a stopped endpoint cannot serve queries and returns a 400 error until started
- Delete endpoints for retired or superseded models; deletion removes all data associated with the endpoint and cannot be undone
- Move CPU-capable models off GPU workload types, and remove or slow down synthetic traffic that prevents scale to zero
Documentation
Vendor references for pricing and configuration.