Explanation
Why the waste happens and who it affects.
The types differ in where prompts and responses are processed: Global may use any Azure region, Data Zone stays within the US, EU or APAC data zone, and Standard stays within the chosen Azure geography. Data at rest stays in the resource's geography for all of them. Microsoft's guidance is to start with Global Standard, which it describes as having the lowest price, the broadest region coverage and the earliest access to new models, and to move to another type only for a specific reason such as data residency.
Deployments created from older templates, or by teams that assume regional processing is required, often use regional Standard or Data Zone Standard without a documented residency requirement. On the Azure OpenAI pricing page (East US, USD, checked September 2026) the Data Zone and Regional per-token rates for the GPT-4.1 family are 10% higher than Global (GPT-4.1 input $2.20 vs $2 per 1M tokens), Data Zone rates for GPT-5 and GPT-5.4 carry a similar premium, and regional rows show no Batch or Priority processing option. Where no residency obligation exists, the premium buys nothing.
Billing model
The pricing dimensions that drive this cost.
Standard deployment types are billed per million input, cached input and output tokens, at a rate that depends on the deployment type.
- Global Standard
- SKU GlobalStandard, pay-per-token, may process in any Azure region, lowest listed per-token price
- Data Zone Standard
- SKU DataZoneStandard, processing kept within the US, EU or APAC data zone, priced above Global
- Standard (regional)
- SKU Standard, processing kept within the Azure geography, priced above Global and with fewer models and lower quota
- Data at rest
- Stays in the designated Azure geography for every deployment type
How to detect
4 checks to find it in your estate.
- List deployments and their sku.name for every Azure OpenAI or Foundry resource (for example with az cognitiveservices account deployment list) and flag Standard and DataZoneStandard deployments
- For each flagged deployment, confirm with the owner whether a documented data residency, sovereignty or contractual requirement exists; deployments without one are candidates
- Compare the per-token rates for the same model and deployment type on the Azure OpenAI pricing page for the resource's region, and multiply the difference by the deployment's monthly token volume from Azure Monitor metrics or Cost analysis
- Check whether regional deployments are blocking use of Batch or Priority processing, which the pricing page lists only for Global and Data Zone types
How to fix
5 ways to remove the waste.
- Create a Global Standard deployment of the same model and version, move traffic to it by updating the deployment name used by clients, then delete the old regional deployment
- Where only zone-level residency is required, use Data Zone Standard rather than a regional deployment so the workload can also use Data Zone Batch and Priority processing
- Keep regional Standard only for workloads with a documented geography requirement, and record the reason on the deployment or resource tags
- Use the Azure Policy definition documented for Foundry deployment types to restrict which sku.name values can be created in subscriptions that have no residency requirement
- Before switching, check quota for the target deployment type and model and review the high availability guidance, since Global and Data Zone deployments are tied to a primary region
Documentation
Vendor references for pricing and configuration.