# Latency-Tolerant Azure OpenAI Workloads Not Using Batch Deployments

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/latency-tolerant-azure-openai-workloads-not-using-batch-deployments

Many Azure OpenAI workloads have no user waiting on the answer: nightly document summarization, classification and extraction over large datasets,...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Many Azure OpenAI workloads have no user waiting on the answer: nightly document summarization, classification and extraction over large datasets, embedding or enrichment backfills, evaluation runs and bulk content generation.

PointFive Research

Cloud cost research at PointFive

Azure service

[Azure Cognitive Services](https://www.pointfive.co/efficiency-hub/cloud-services/azure-cognitive-services)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0564

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

These jobs are often written as loops of synchronous calls against a Global Standard or Standard deployment, because that is how the application was prototyped, and they pay full real-time token rates.

Azure OpenAI Global Batch and Data Zone Batch deployments process the same Chat Completions or Responses requests asynchronously from a JSONL file, with a 24-hour target turnaround at 50% less cost than Global Standard. Batch requests draw on a separate enqueued-token quota, so moving bulk work to Batch also stops it from consuming the real-time quota that interactive traffic depends on. Microsoft's deployment type guidance recommends Global Batch or Data Zone Batch for large asynchronous jobs that are not time-sensitive.

## Billing model

The pricing dimensions that drive this cost.

Standard and Global Standard

Per-token pricing for synchronous requests at the real-time rate

Global Batch and Data Zone Batch

Per-token pricing at 50% less than Global Standard for requests submitted as batch jobs

Enqueued token quota

Separate quota for batch jobs that doesn't reduce the real-time tokens-per-minute quota

Cancelled or long-running jobs

Jobs are not expired if they exceed 24 hours; cancelling returns completed work, which is billed

## How to detect

4 checks to find it in your estate.

- List deployments per Azure OpenAI or Foundry resource by sku.name and check whether any GlobalBatch or DataZoneBatch deployments exist; resources with heavy token volume and no batch deployment are candidates

- Chart the Azure OpenAI Requests and Processed Inference Tokens (TokenTransaction) metrics split by ModelDeploymentName; deployments with large spikes at fixed times, such as nightly or weekly, usually serve scheduled jobs rather than users

- Trace high-volume callers (application identity, API key, pipeline or job name) to confirm whether the results are needed within seconds or can wait up to 24 hours

- Check for workloads hitting real-time 429 rate limits during bulk runs, a sign that offline work is competing with interactive traffic for the same quota

## How to fix

5 ways to remove the waste.

- Create a Global Batch deployment (or Data Zone Batch where data zone processing is required) of the same model, write requests to a JSONL file with custom\_id, method and url fields, upload it and submit a batch job

- Keep Standard, Global Standard or provisioned deployments for interactive traffic only, and route scheduled and bulk jobs to the batch deployment

- Enable dynamic quota on batch deployments, as Microsoft recommends, and use the documented retry-with-backoff pattern for very large jobs that exceed the enqueued token quota

- For delay-tolerant work that must stay synchronous, evaluate Flex processing (preview) on Global Standard deployments, priced at 50% of Standard token rates; it supports only listed models (gpt-5.6-sol at the time of writing), returns HTTP 400 for unsupported models and HTTP 429 when Flex capacity is unavailable, with no automatic fallback to Standard

- Confirm the model and region are listed for Global Batch or Data Zone Batch before migrating, since not every model supports every deployment type

## Documentation

Vendor references for pricing and configuration.

- [How to use global batch processing with Azure OpenAI in Microsoft Foundry Models  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/batch)

- [Understanding deployment types in Microsoft Foundry Models  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/deployment-types)

- [Use Flex processing with Azure OpenAI in Foundry Models (preview)  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/flex-processing)

- [Monitoring data reference for Azure OpenAI  learn.microsoft.com](https://learn.microsoft.com/en-us/azure/foundry/openai/monitor-openai-reference)

- [Azure OpenAI pricing  azure.microsoft.com](https://azure.microsoft.com/en-us/pricing/details/cognitive-services/openai-service/)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Azure Cognitive Services  CER-0248

### [Non-Production Azure OpenAI Deployments Using PTUs Instead of PAYG](https://www.pointfive.co/efficiency-hub/inefficiencies/non-production-azure-openai-deployments-using-ptus-instead-of-payg-9fe81)

Development, testing, QA, and sandbox environments rarely have the steady, predictable traffic patterns needed to justify PTU deployments. These workloads often run intermittently, with lower throughput and shorter usage windows. When PTUs...

AI

- Azure Cognitive Services  CER-0246

### [Missing Reserved PTUs for Steady-State Azure OpenAI Workloads](https://www.pointfive.co/efficiency-hub/inefficiencies/missing-reserved-ptus-for-steady-state-azure-openai-workloads-f78d4)

Many production Azure OpenAI workloads - such as chatbots, inference services, and retrieval-augmented generation (RAG) pipelines-use PTUs consistently throughout the day. When usage stabilizes after initial experimentation, continuing to...

AI

- Azure Cognitive Services  CER-0510

### [Regional Standard Azure OpenAI Deployments Without a Data Residency Requirement](https://www.pointfive.co/efficiency-hub/inefficiencies/regional-standard-azure-openai-deployments-without-a-data-residency-requirement)

Azure OpenAI models in Microsoft Foundry (formerly Azure AI services) can be deployed as Global, Data Zone or regional (Standard) pay-per-token deployments. The types differ in where prompts and responses are processed: Global may use any...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

