# Latency-Tolerant OpenAI API Workloads Not Using Batch or Flex Processing

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/latency-tolerant-openai-api-workloads-not-using-batch-or-flex-processing

Evaluations, dataset classification, data enrichment, embedding of content repositories and other asynchronous jobs are frequently run against the...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Evaluations, dataset classification, data enrichment, embedding of content repositories and other asynchronous jobs are frequently run against the OpenAI API at Standard processing rates, because they were built as loops over the synchronous Responses or Chat Completions endpoints.

PointFive Research

Cloud cost research at PointFive

OpenAI service

[OpenAI API](https://www.pointfive.co/efficiency-hub/cloud-services/openai-api)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0507

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

OpenAI offers two cheaper ways to run the same models: the Batch API, which processes a file of requests within a 24-hour window at a 50 percent discount, and Flex processing, which keeps the synchronous API but accepts slower responses and occasional resource unavailability, with tokens priced at Batch rates.

The inverse mistake also happens: background jobs sent with service\_tier set to priority or fast, the tier renamed Fast mode on July 30, 2026, which is priced above Standard. OpenAI's cost optimization guide names Batch and Flex as the ways to lower cost for work that can wait; for gpt-6-astra, Standard input and output tokens cost twice the Batch rate and Fast mode four times it.

## Billing model

The pricing dimensions that drive this cost.

OpenAI bills tokens per million at the rate of the processing tier used for each request.

Standard processing

The default per-token rate for input, cached input, cache writes and output

Batch API

A 50 percent discount compared with the synchronous APIs, for requests submitted as a file and completed within 24 hours

Flex processing

Synchronous requests priced at Batch API rates in exchange for slower responses and possible 429 Resource Unavailable errors, which are not charged

Fast mode

Service\_tier priority or fast, priced above Standard; for gpt-6-astra the pricing page lists $20 input and $100 output per 1M short-context tokens against $10 and $50 Standard

## How to detect

4 checks to find it in your estate.

- Query the Admin Usage API completions endpoint (GET /v1/organization/usage/completions) grouped by project\_id, api\_key\_id, model, batch and service\_tier to see how much volume runs outside Batch and Flex

- Identify projects and API keys whose traffic comes from schedulers, pipelines, evaluation harnesses or backfills rather than interactive users

- Flag Fast mode (priority or fast service\_tier) usage on keys that serve non-interactive workloads

- Check client code for long synchronous loops with custom concurrency and retry logic over large datasets, a sign that batch work is running at Standard rates

## How to fix

4 ways to remove the waste.

- Move offline jobs to the Batch API: upload a .jsonl file with one request per line and a unique custom\_id, create the batch with a 24h completion window, and map results by custom\_id because output order can differ; input files are limited to one model each and output files are deleted 30 days after completion

- Set service\_tier to flex for latency-tolerant synchronous calls on supported models, raise client timeouts, and retry 429 Resource Unavailable responses with exponential backoff, or fall back to Standard with service\_tier auto when completion matters more than cost

- Remove priority or fast service\_tier from background and batch-like workloads, keeping Fast mode for interactive paths where latency justifies the premium

- Combine Batch or Flex with prompt caching, since Flex tokens get the additional prompt caching discount

## Documentation

Vendor references for pricing and configuration.

- [Batch API  developers.openai.com](https://developers.openai.com/api/docs/guides/batch)

- [Flex processing  developers.openai.com](https://developers.openai.com/api/docs/guides/flex-processing)

- [Pricing  developers.openai.com](https://developers.openai.com/api/docs/pricing)

- [Cost optimization  developers.openai.com](https://developers.openai.com/api/docs/guides/cost-optimization)

- [Completions  developers.openai.com](https://developers.openai.com/api/reference/typescript/resources/admin/subresources/organization/subresources/usage/methods/completions)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- OpenAI API  CER-0508

### [Low Prompt Cache Hit Rate from Unstable Prompt Prefixes in the OpenAI API](https://www.pointfive.co/efficiency-hub/inefficiencies/low-prompt-cache-hit-rate-from-unstable-prompt-prefixes-in-the-openai-api)

Prompt caching is enabled by default for supported OpenAI models, and reused prefix tokens are billed at a cached-input rate discounted by up to 90 percent. Reuse only happens when the entire rendered prefix matches: hidden system content,...

AI

- OpenAI API  CER-0509

### [Using High-Cost OpenAI Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-openai-models-for-low-complexity-tasks)

Classification, routing, triage, simple data extraction and small scoped edits are often sent to OpenAI's most capable model because a single model name is configured for the whole application. The per-token gap between tiers is large: at...

AI

- Azure Cognitive Services  CER-0248

### [Non-Production Azure OpenAI Deployments Using PTUs Instead of PAYG](https://www.pointfive.co/efficiency-hub/inefficiencies/non-production-azure-openai-deployments-using-ptus-instead-of-payg-9fe81)

Development, testing, QA, and sandbox environments rarely have the steady, predictable traffic patterns needed to justify PTU deployments. These workloads often run intermittently, with lower throughput and shorter usage windows. When PTUs...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

