# Unused or Unconsolidated Databricks AI Search Endpoints

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/unused-or-unconsolidated-databricks-ai-search-endpoints

Databricks AI Search (formerly Databricks Vector Search) serves vector indexes from endpoints that are billed per hour in capacity units.

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Databricks AI Search (formerly Databricks Vector Search) serves vector indexes from endpoints that are billed per hour in capacity units.

PointFive Research

Cloud cost research at PointFive

Databricks service

[Databricks AI Search](https://www.pointfive.co/efficiency-hub/cloud-services/databricks-ai-search)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0525

Type

Idle or Unused Resource

## Explanation

Why the waste happens and who it affects.

Once an endpoint hosts at least one index it bills for its provisioned units continuously, and Databricks states that endpoints with indexes that receive no query traffic still incur serving costs. There is no scale to zero: the pricing FAQ says both Standard and Storage Optimized packages scale down only to one unit. Charges stop only 24 hours after the last index is deleted from the endpoint.

The waste accumulates from RAG prototypes, hackathon projects and evaluations whose indexes are never deleted, and from teams creating a separate endpoint for every index or application even though one endpoint can serve up to 50 indexes. Each extra endpoint carries its own minimum unit cost, so many lightly used endpoints cost more than one shared endpoint. Delta Sync indexes left on Continuous Sync add a streaming ingestion cost on top, even when the source table changes rarely or the index is no longer queried.

## Billing model

The pricing dimensions that drive this cost.

Endpoint units

Billed per hour per unit; a Standard unit covers about 2 million 768-dimension vectors (4.00 DBU per hour per unit on the pricing page) and a Storage Optimized unit about 64 million (18.29 DBU per hour per unit)

Minimum capacity

An endpoint with at least one index bills at least one unit around the clock; it no longer incurs charges 24 hours after its last index is deleted

Index sync

Delta Sync ingestion is charged while syncing at serverless jobs compute rates; Continuous Sync keeps a streaming pipeline running, Triggered Sync runs only when invoked

Storage

Index storage is billed per GB-month; the Standard package includes the first 30 GB free

## How to detect

4 checks to find it in your estate.

- Query system.billing.usage where billing\_origin\_product = 'VECTOR\_SEARCH', grouped by usage\_metadata.endpoint\_name and day, to rank endpoints by steady daily DBUs

- Use Databricks' 'Identify unused AI Search endpoints' query on system.access.audit (service\_name = 'vectorSearch'): compare indexes created (createVectorIndex) and not deleted (deleteVectorIndex) with indexes that received queryVectorIndex or scanVectorIndex calls in the last 30 days, then aggregate to endpoint level

- List endpoints and their indexes and flag endpoints hosting a single small, low-traffic index that could share an endpoint with others

- Check the sync mode of Delta Sync indexes and flag Continuous Sync on indexes whose source tables update infrequently or that have no recent queries

## How to fix

5 ways to remove the waste.

- Delete indexes with no query traffic and then delete the endpoint if it is empty; the endpoint stops incurring charges 24 hours after its last index is removed

- Consolidate low-QPS indexes from multiple endpoints onto a shared endpoint (up to 50 indexes per endpoint), keeping separate endpoints only where isolation, QPS or latency requirements justify them

- Switch Delta Sync indexes to Triggered Sync where near real-time freshness is not needed, which Databricks recommends to reduce streaming costs, and trigger syncs from the pipeline that updates the source table

- For large vector counts that tolerate about 250 ms of added query latency, compare Storage Optimized endpoints, which Databricks positions as the better TCO for large datasets

- Tag endpoints with owners and add unused-endpoint checks to regular reviews so prototype endpoints are removed when projects end

## Documentation

Vendor references for pricing and configuration.

- [AI Search cost management guide  docs.databricks.com](https://docs.databricks.com/aws/en/ai-search/cost-management)

- [Identify unused AI Search endpoints  docs.databricks.com](https://docs.databricks.com/aws/en/ai-search/unused-endpoints)

- [Databricks AI Search  docs.databricks.com](https://docs.databricks.com/aws/en/ai-search/ai-search)

- [Billable usage system table reference  docs.databricks.com](https://docs.databricks.com/aws/en/admin/system-tables/billing)

- [AI Search  databricks.com](https://www.databricks.com/product/pricing/ai-search)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Databricks Model Serving  CER-0524

### [Databricks Model Serving Endpoints Running Without Scale to Zero](https://www.pointfive.co/efficiency-hub/inefficiencies/databricks-model-serving-endpoints-running-without-scale-to-zero)

Custom model serving endpoints in Databricks run on serverless CPU or GPU compute sized by provisioned concurrency. When scale to zero is not enabled, the endpoint never drops below its minimum provisioned concurrency, and Databricks'...

AI

- AWS SageMaker  CER-0333

### [Idle SageMaker Notebook Instances Left Running Continuously](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-sagemaker-notebook-instances-left-running-continuously)

SageMaker notebook instances are billed continuously while in an active state - and critically, they do not automatically shut down when idle. Closing a browser tab, shutting down a Jupyter kernel, or simply walking away does not stop the...

AI

- AWS S3  CER-0330

### [Orphaned MLflow Training Artifacts and Model Checkpoints in Object Storage](https://www.pointfive.co/efficiency-hub/inefficiencies/orphaned-mlflow-training-artifacts-and-model-checkpoints-in-object-storage)

Machine learning experimentation workflows - particularly those managed through experiment tracking platforms - generate large volumes of artifacts in object storage. Every training run produces model checkpoints, evaluation outputs,...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

