Skip to content
Cloud Efficiency Hub

Unused or Unconsolidated Databricks AI Search Endpoints

The short version

Databricks AI Search (formerly Databricks Vector Search) serves vector indexes from endpoints that are billed per hour in capacity units.

PointFive Research

Cloud cost research at PointFive

Databricks service
Databricks AI Search
Category
AI
Reference
CER-0525
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

Once an endpoint hosts at least one index it bills for its provisioned units continuously, and Databricks states that endpoints with indexes that receive no query traffic still incur serving costs. There is no scale to zero: the pricing FAQ says both Standard and Storage Optimized packages scale down only to one unit. Charges stop only 24 hours after the last index is deleted from the endpoint.

The waste accumulates from RAG prototypes, hackathon projects and evaluations whose indexes are never deleted, and from teams creating a separate endpoint for every index or application even though one endpoint can serve up to 50 indexes. Each extra endpoint carries its own minimum unit cost, so many lightly used endpoints cost more than one shared endpoint. Delta Sync indexes left on Continuous Sync add a streaming ingestion cost on top, even when the source table changes rarely or the index is no longer queried.

Billing model

The pricing dimensions that drive this cost.

Endpoint units
Billed per hour per unit; a Standard unit covers about 2 million 768-dimension vectors (4.00 DBU per hour per unit on the pricing page) and a Storage Optimized unit about 64 million (18.29 DBU per hour per unit)
Minimum capacity
An endpoint with at least one index bills at least one unit around the clock; it no longer incurs charges 24 hours after its last index is deleted
Index sync
Delta Sync ingestion is charged while syncing at serverless jobs compute rates; Continuous Sync keeps a streaming pipeline running, Triggered Sync runs only when invoked
Storage
Index storage is billed per GB-month; the Standard package includes the first 30 GB free

How to detect

4 checks to find it in your estate.

  • Query system.billing.usage where billing_origin_product = 'VECTOR_SEARCH', grouped by usage_metadata.endpoint_name and day, to rank endpoints by steady daily DBUs
  • Use Databricks' 'Identify unused AI Search endpoints' query on system.access.audit (service_name = 'vectorSearch'): compare indexes created (createVectorIndex) and not deleted (deleteVectorIndex) with indexes that received queryVectorIndex or scanVectorIndex calls in the last 30 days, then aggregate to endpoint level
  • List endpoints and their indexes and flag endpoints hosting a single small, low-traffic index that could share an endpoint with others
  • Check the sync mode of Delta Sync indexes and flag Continuous Sync on indexes whose source tables update infrequently or that have no recent queries

How to fix

5 ways to remove the waste.

  • Delete indexes with no query traffic and then delete the endpoint if it is empty; the endpoint stops incurring charges 24 hours after its last index is removed
  • Consolidate low-QPS indexes from multiple endpoints onto a shared endpoint (up to 50 indexes per endpoint), keeping separate endpoints only where isolation, QPS or latency requirements justify them
  • Switch Delta Sync indexes to Triggered Sync where near real-time freshness is not needed, which Databricks recommends to reduce streaming costs, and trigger syncs from the pipeline that updates the source table
  • For large vector counts that tolerate about 250 ms of added query latency, compare Storage Optimized endpoints, which Databricks positions as the better TCO for large datasets
  • Tag endpoints with owners and add unused-endpoint checks to regular reviews so prototype endpoints are removed when projects end

Documentation

Vendor references for pricing and configuration.