# Overprovisioned Replicas and Partitions in Azure AI Search

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-replicas-and-partitions-in-azure-ai-search

In the Dedicated pricing model, an Azure AI Search service is billed every hour for its search units, the number of replicas multiplied by the number...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

In the Dedicated pricing model, an Azure AI Search service is billed every hour for its search units, the number of replicas multiplied by the number of partitions, at the rate of its tier.

PointFive Research

Cloud cost research at PointFive

Azure service

[Azure AI Search](https://www.pointfive.co/efficiency-hub/cloud-services/azure-ai-search)

Category

[Databases](https://www.pointfive.co/efficiency-hub/service-category/databases)

Reference

CER-0455

Type

Overprovisioned Resource

## Explanation

Why the waste happens and who it affects.

Capacity is set manually, so replicas added for a load test, a launch or a large indexing run, and partitions added for an index that was later trimmed, often stay in place long after they are needed. Because the bill is the product of replicas and partitions, going from 1 x 1 to 2 x 2 quadruples the cost, and Microsoft notes that doubling capacity more than doubles costs on the same tier.

A Dedicated service cannot be paused: resources stay allocated for the life of the service, and the only way to stop billing is to delete it. Non-production services are frequently configured like production, with the multiple replicas that SLAs require, and services are often created on a higher tier than their index size and query volume need. Microsoft cost guidance says partitions should be added only when index size or ingestion throughput requires it, and replicas only for query volume, throttling or high availability.

## Billing model

The pricing dimensions that drive this cost.

Dedicated Azure AI Search services are billed on provisioned capacity, not on queries.

Search units

Replicas multiplied by partitions, billed at the prorated hourly rate of the service tier

Always-on capacity

Dedicated resources are allocated for the lifetime of the service and are not billed per query, so idle capacity costs the same as busy capacity

Deleting to stop billing

A Dedicated service cannot be shut down temporarily, and deleting it also deletes its data

Premium features

Semantic ranker, agentic retrieval and AI enrichment are billed separately from search units

## How to detect

5 checks to find it in your estate.

- Compare replica count with the SearchQueriesPerSecond, SearchLatency and ThrottledSearchQueriesPercentage metrics over at least a month; Microsoft recommends using QPS, latency and throttling to decide when to add or remove replicas, and sustained low QPS with no throttling suggests extra replicas

- Compare partition count with index storage usage against the tier's per-partition limits, and with indexing throughput needs; partitions beyond what index size or ingestion requires are candidates for removal

- Flag non-production services with two or more replicas, which Microsoft ties to SLA requirements (two for read SLAs, three for read-write SLAs) that dev and test services rarely need

- Identify services with no queries and no index updates over the metric retention window, which may be unused altogether

- Check services created before April or May 2024 for eligibility for the one-time upgrade to newer infrastructure with larger partitions, which can reduce the number of partitions needed

## How to fix

5 ways to remove the waste.

- Remove replicas and partitions that query volume, throttling, high availability requirements or index size do not justify, using the Scale page, Azure CLI, PowerShell or the Management REST API; scaling runs in the background and can take from minutes to several hours

- Scale up for resource-intensive operations such as bulk indexing and then scale back down for regular query load, and automate this for predictable patterns as Microsoft suggests

- Move to the lightest tier that meets the workload; existing Dedicated services can switch between Basic, S1, S2 and S3 as long as the current configuration fits the target tier's limits and the region has capacity

- Upgrade eligible older services to newer infrastructure at no extra cost to get larger partitions, then reduce partition count if the index now fits in fewer

- Delete unused services after exporting what is needed, and treat the Serverless pricing model as an option only with care: it is in preview without an SLA, is not recommended for production, is available in specific regions, and an existing Dedicated service cannot be converted, so moving means creating a new service and reindexing

## Documentation

Vendor references for pricing and configuration.

- [Plan and Manage Costs - Azure AI Search  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/search/search-sku-manage-costs)

- [Estimate Capacity for Query and Index Workloads - Azure AI Search  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/search/search-capacity-planning)

- [Monitor Queries - Azure AI Search  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/search/search-monitor-queries)

- [Foundry IQ pricing  azure.microsoft.com](https://azure.microsoft.com/en-us/pricing/details/search/)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Azure Cache for Redis  CER-0310

### [Overprovisioned Azure Cache for Redis Instance](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-azure-cache-for-redis-instance)

Azure Cache for Redis is billed at a fixed rate determined entirely by the provisioned tier and cache size - not by actual utilization. A cache instance that consumes only a fraction of its available memory and throughput incurs the same...

Databases

- Azure SQL  CER-0028

### [Overprovisioned Compute Tier in Azure SQL Database](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-compute-tier-in-azure-sql-database)

Azure SQL Database resources are frequently overprovisioned due to default configurations, conservative sizing, or legacy requirements that no longer apply. This inefficiency appears across all deployment models: Single Databases may be...

Databases

- Azure Database for PostgreSQL - Flexible Server  CER-0148

### [Overprovisioned Azure Database for PostgreSQL Flexible Server](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-azure-database-for-postgresql-flexible-server)

Azure Database for PostgreSQL - Flexible Server often defaults to general-purpose D-series VMs, which may be oversized for many production or development workloads. Many PostgreSQL workloads do not need high sustained CPU and can run on...

Databases

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

