# OpenAI API Optimization | Cloud Efficiency Hub

Canonical: https://www.pointfive.co/efficiency-hub/cloud-services/openai-api

3 documented OpenAI API cost inefficiencies, each with how to detect it and how to fix it.

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

3 documented OpenAI API cost inefficiencies, each with how to detect it and how to fix it.

## 3 inefficiencies

[All services](https://www.pointfive.co/efficiency-hub#browse-by-service)

- OpenAI API  CER-0507

### [Latency-Tolerant OpenAI API Workloads Not Using Batch or Flex Processing](https://www.pointfive.co/efficiency-hub/inefficiencies/latency-tolerant-openai-api-workloads-not-using-batch-or-flex-processing)

Evaluations, dataset classification, data enrichment, embedding of content repositories and other asynchronous jobs are frequently run against the OpenAI API at Standard processing rates, because they were built as loops over the...

AI

- OpenAI API  CER-0508

### [Low Prompt Cache Hit Rate from Unstable Prompt Prefixes in the OpenAI API](https://www.pointfive.co/efficiency-hub/inefficiencies/low-prompt-cache-hit-rate-from-unstable-prompt-prefixes-in-the-openai-api)

Prompt caching is enabled by default for supported OpenAI models, and reused prefix tokens are billed at a cached-input rate discounted by up to 90 percent. Reuse only happens when the entire rendered prefix matches: hidden system content,...

AI

- OpenAI API  CER-0509

### [Using High-Cost OpenAI Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-openai-models-for-low-complexity-tasks)

Classification, routing, triage, simple data extraction and small scoped edits are often sent to OpenAI's most capable model because a single model name is configured for the whole application. The per-token gap between tiers is large: at...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

