Skip to content
Cloud Efficiency Hub

Latency-Tolerant Azure OpenAI Workloads Not Using Batch Deployments

The short version

Many Azure OpenAI workloads have no user waiting on the answer: nightly document summarization, classification and extraction over large datasets, embedding or enrichment backfills, evaluation runs and bulk content generation.

PointFive Research

Cloud cost research at PointFive

Category
AI
Reference
CER-0564
Type
Suboptimal Pricing Model

Explanation

Why the waste happens and who it affects.

These jobs are often written as loops of synchronous calls against a Global Standard or Standard deployment, because that is how the application was prototyped, and they pay full real-time token rates.

Azure OpenAI Global Batch and Data Zone Batch deployments process the same Chat Completions or Responses requests asynchronously from a JSONL file, with a 24-hour target turnaround at 50% less cost than Global Standard. Batch requests draw on a separate enqueued-token quota, so moving bulk work to Batch also stops it from consuming the real-time quota that interactive traffic depends on. Microsoft's deployment type guidance recommends Global Batch or Data Zone Batch for large asynchronous jobs that are not time-sensitive.

Billing model

The pricing dimensions that drive this cost.

Standard and Global Standard
Per-token pricing for synchronous requests at the real-time rate
Global Batch and Data Zone Batch
Per-token pricing at 50% less than Global Standard for requests submitted as batch jobs
Enqueued token quota
Separate quota for batch jobs that doesn't reduce the real-time tokens-per-minute quota
Cancelled or long-running jobs
Jobs are not expired if they exceed 24 hours; cancelling returns completed work, which is billed

How to detect

4 checks to find it in your estate.

  • List deployments per Azure OpenAI or Foundry resource by sku.name and check whether any GlobalBatch or DataZoneBatch deployments exist; resources with heavy token volume and no batch deployment are candidates
  • Chart the Azure OpenAI Requests and Processed Inference Tokens (TokenTransaction) metrics split by ModelDeploymentName; deployments with large spikes at fixed times, such as nightly or weekly, usually serve scheduled jobs rather than users
  • Trace high-volume callers (application identity, API key, pipeline or job name) to confirm whether the results are needed within seconds or can wait up to 24 hours
  • Check for workloads hitting real-time 429 rate limits during bulk runs, a sign that offline work is competing with interactive traffic for the same quota

How to fix

5 ways to remove the waste.

  • Create a Global Batch deployment (or Data Zone Batch where data zone processing is required) of the same model, write requests to a JSONL file with custom_id, method and url fields, upload it and submit a batch job
  • Keep Standard, Global Standard or provisioned deployments for interactive traffic only, and route scheduled and bulk jobs to the batch deployment
  • Enable dynamic quota on batch deployments, as Microsoft recommends, and use the documented retry-with-backoff pattern for very large jobs that exceed the enqueued token quota
  • For delay-tolerant work that must stay synchronous, evaluate Flex processing (preview) on Global Standard deployments, priced at 50% of Standard token rates; it supports only listed models (gpt-5.6-sol at the time of writing), returns HTTP 400 for unsupported models and HTTP 429 when Flex capacity is unavailable, with no automatic fallback to Standard
  • Confirm the model and region are listed for Global Batch or Data Zone Batch before migrating, since not every model supports every deployment type

Documentation

Vendor references for pricing and configuration.