# Databricks Model Serving Endpoints Running Without Scale to Zero

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/databricks-model-serving-endpoints-running-without-scale-to-zero

Custom model serving endpoints in Databricks run on serverless CPU or GPU compute sized by provisioned concurrency.

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Custom model serving endpoints in Databricks run on serverless CPU or GPU compute sized by provisioned concurrency.

PointFive Research

Cloud cost research at PointFive

Databricks service

[Databricks Model Serving](https://www.pointfive.co/efficiency-hub/cloud-services/databricks-model-serving)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0524

Type

Idle or Unused Resource

## Explanation

Why the waste happens and who it affects.

When scale to zero is not enabled, the endpoint never drops below its minimum provisioned concurrency, and Databricks' pricing FAQ states that with no traffic the minimum charge depends on the minimum provisioned concurrency of the chosen range. An endpoint serving a handful of requests a day, or none, therefore bills for its floor capacity every hour.

This is typical for development, staging and demo endpoints, endpoints for models that have been superseded, and low-traffic internal tools. Scale to zero is optional and off unless selected, and teams often raise the minimum concurrency or pick a GPU workload type during testing and never revisit it. GPU endpoints make the idle floor expensive, since GPU serving is billed per GPU instance per hour. Scale to zero brings cold starts and no guarantee of GPU capacity when scaling back up, and Databricks advises against it for production workloads that need consistent uptime, so the fix applies mainly to non-production and low-traffic endpoints.

## Billing model

The pricing dimensions that drive this cost.

Model serving usage appears on the bill under the Serverless Real-time Inference SKU.

CPU serving

Billed in DBUs based on the compute nodes used for the endpoint's provisioned concurrency, with the cloud instance cost included in the DBU price

GPU serving

Billed per GPU instance per hour at a DBU rate set by GPU configuration, listed on the Model Serving pricing page

Minimum provisioned concurrency

Without scale to zero, the endpoint bills at least for the minimum concurrency of its configured range even when it receives no requests

Scale to zero

After 30 minutes with no requests the endpoint scales to zero and is not charged; charges resume when a new request triggers scale-up

## How to detect

5 checks to find it in your estate.

- Query system.billing.usage where billing\_origin\_product = 'MODEL\_SERVING' and group by usage\_metadata.endpoint\_name and day; flat daily DBUs over several weeks indicate an endpoint running at its floor

- Use product\_features.serving\_type (MODEL or GPU\_MODEL) in the same table to separate CPU and GPU custom model endpoints and prioritize GPU ones

- List endpoint configurations through the Serving Endpoints API and flag served entities with scale\_to\_zero\_enabled = false, a large workload\_size, or a min\_provisioned\_concurrency above what traffic needs, outside production

- Check the endpoint health metrics in the Serving UI (request rate, latency, CPU and memory usage for the last 14 days) to confirm near-zero request rates on endpoints with steady DBUs

- Look for synthetic health checks or monitoring probes that send periodic requests and keep a scale-to-zero endpoint from ever reaching 30 minutes of inactivity

## How to fix

5 ways to remove the waste.

- Enable scale to zero on development, test and low-traffic endpoints by updating the served entity configuration with scale\_to\_zero\_enabled = true, accepting cold starts of typically 10 to 20 seconds and sometimes minutes on the first request

- Lower workload\_size or min\_provisioned\_concurrency (multiples of 4) to what peak traffic needs, sizing with Databricks' guidance that provisioned concurrency equals queries per second times model execution time

- Stop custom model endpoints that must keep their configuration but are not needed for a period (POST /api/2.0/serving-endpoints/{name}/config:stop, or the Stop button); a stopped endpoint cannot serve queries and returns a 400 error until started

- Delete endpoints for retired or superseded models; deletion removes all data associated with the endpoint and cannot be undone

- Move CPU-capable models off GPU workload types, and remove or slow down synthetic traffic that prevents scale to zero

## Documentation

Vendor references for pricing and configuration.

- [Custom models overview  docs.databricks.com](https://docs.databricks.com/aws/en/machine-learning/model-serving/custom-models)

- [Create custom model serving endpoints  docs.databricks.com](https://docs.databricks.com/aws/en/machine-learning/model-serving/create-manage-serving-endpoints)

- [Manage model serving endpoints  docs.databricks.com](https://docs.databricks.com/aws/en/machine-learning/model-serving/manage-serving-endpoints)

- [Monitor model quality and endpoint health  docs.databricks.com](https://docs.databricks.com/aws/en/machine-learning/model-serving/monitor-diagnose-endpoints)

- [Billable usage system table reference  docs.databricks.com](https://docs.databricks.com/aws/en/admin/system-tables/billing)

- [Model Serving Pricing  databricks.com](https://www.databricks.com/product/pricing/model-serving)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Databricks AI Search  CER-0525

### [Unused or Unconsolidated Databricks AI Search Endpoints](https://www.pointfive.co/efficiency-hub/inefficiencies/unused-or-unconsolidated-databricks-ai-search-endpoints)

Databricks AI Search (formerly Databricks Vector Search) serves vector indexes from endpoints that are billed per hour in capacity units. Once an endpoint hosts at least one index it bills for its provisioned units continuously, and...

AI

- AWS SageMaker  CER-0333

### [Idle SageMaker Notebook Instances Left Running Continuously](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-sagemaker-notebook-instances-left-running-continuously)

SageMaker notebook instances are billed continuously while in an active state - and critically, they do not automatically shut down when idle. Closing a browser tab, shutting down a Jupyter kernel, or simply walking away does not stop the...

AI

- AWS S3  CER-0330

### [Orphaned MLflow Training Artifacts and Model Checkpoints in Object Storage](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-mlflow-training-artifacts-and-model-checkpoints-in-object-storage)

Machine learning experimentation workflows - particularly those managed through experiment tracking platforms - generate large volumes of artifacts in object storage. Every training run produces model checkpoints, evaluation outputs,...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

