Skip to content
Cloud Efficiency Hub

Overprovisioned Replicas and Partitions in Azure AI Search

The short version

In the Dedicated pricing model, an Azure AI Search service is billed every hour for its search units, the number of replicas multiplied by the number of partitions, at the rate of its tier.

PointFive Research

Cloud cost research at PointFive

Azure service
Azure AI Search
Category
Databases
Reference
CER-0455
Type
Overprovisioned Resource

Explanation

Why the waste happens and who it affects.

Capacity is set manually, so replicas added for a load test, a launch or a large indexing run, and partitions added for an index that was later trimmed, often stay in place long after they are needed. Because the bill is the product of replicas and partitions, going from 1 x 1 to 2 x 2 quadruples the cost, and Microsoft notes that doubling capacity more than doubles costs on the same tier.

A Dedicated service cannot be paused: resources stay allocated for the life of the service, and the only way to stop billing is to delete it. Non-production services are frequently configured like production, with the multiple replicas that SLAs require, and services are often created on a higher tier than their index size and query volume need. Microsoft cost guidance says partitions should be added only when index size or ingestion throughput requires it, and replicas only for query volume, throttling or high availability.

Billing model

The pricing dimensions that drive this cost.

Dedicated Azure AI Search services are billed on provisioned capacity, not on queries.

Search units
Replicas multiplied by partitions, billed at the prorated hourly rate of the service tier
Always-on capacity
Dedicated resources are allocated for the lifetime of the service and are not billed per query, so idle capacity costs the same as busy capacity
Deleting to stop billing
A Dedicated service cannot be shut down temporarily, and deleting it also deletes its data
Premium features
Semantic ranker, agentic retrieval and AI enrichment are billed separately from search units

How to detect

5 checks to find it in your estate.

  • Compare replica count with the SearchQueriesPerSecond, SearchLatency and ThrottledSearchQueriesPercentage metrics over at least a month; Microsoft recommends using QPS, latency and throttling to decide when to add or remove replicas, and sustained low QPS with no throttling suggests extra replicas
  • Compare partition count with index storage usage against the tier's per-partition limits, and with indexing throughput needs; partitions beyond what index size or ingestion requires are candidates for removal
  • Flag non-production services with two or more replicas, which Microsoft ties to SLA requirements (two for read SLAs, three for read-write SLAs) that dev and test services rarely need
  • Identify services with no queries and no index updates over the metric retention window, which may be unused altogether
  • Check services created before April or May 2024 for eligibility for the one-time upgrade to newer infrastructure with larger partitions, which can reduce the number of partitions needed

How to fix

5 ways to remove the waste.

  • Remove replicas and partitions that query volume, throttling, high availability requirements or index size do not justify, using the Scale page, Azure CLI, PowerShell or the Management REST API; scaling runs in the background and can take from minutes to several hours
  • Scale up for resource-intensive operations such as bulk indexing and then scale back down for regular query load, and automate this for predictable patterns as Microsoft suggests
  • Move to the lightest tier that meets the workload; existing Dedicated services can switch between Basic, S1, S2 and S3 as long as the current configuration fits the target tier's limits and the region has capacity
  • Upgrade eligible older services to newer infrastructure at no extra cost to get larger partitions, then reduce partition count if the index now fits in fewer
  • Delete unused services after exporting what is needed, and treat the Serverless pricing model as an option only with care: it is in preview without an SLA, is not recommended for production, is available in specific regions, and an existing Dedicated service cannot be converted, so moving means creating a new service and reindexing

Documentation

Vendor references for pricing and configuration.