# Low Prompt Cache Hit Rate from Unstable Prompt Prefixes in the OpenAI API

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/low-prompt-cache-hit-rate-from-unstable-prompt-prefixes-in-the-openai-api

Prompt caching is enabled by default for supported OpenAI models, and reused prefix tokens are billed at a cached-input rate discounted by up to 90...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Prompt caching is enabled by default for supported OpenAI models, and reused prefix tokens are billed at a cached-input rate discounted by up to 90 percent.

PointFive Research

Cloud cost research at PointFive

OpenAI service

[OpenAI API](https://www.pointfive.co/efficiency-hub/cloud-services/openai-api)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0508

Type

Inefficient Configuration

## Explanation

Why the waste happens and who it affects.

Reuse only happens when the entire rendered prefix matches: hidden system content, tool definitions and schemas, developer messages and conversation history, in that order. Anything that changes early in that sequence, such as a timestamp or user ID in the developer message, tools added, removed or reordered per request, or a switch in structured output schema, reasoning effort or verbosity, forces the rest of the prompt to be billed at the full uncached rate.

Because caching is automatic, a broken prefix raises no error; it only shows up as low cached\_tokens in the usage data. On GPT-5.6 and later the stakes are higher in both directions: cache writes are billed at 1.25 times the uncached input rate, so a prefix that is written on every request but never reused costs more than no caching at all, while a well-structured prefix is read at 0.1 times the input rate. Short shared prefixes below the minimum cacheable length of 1,024 visible tokens are never cached. OpenAI documents these as the main levers in its prompt caching guide and provides a dashboard and diagnostics tool for cache misses.

## Billing model

The pricing dimensions that drive this cost.

Input tokens are billed at one of three rates, not as additive fees.

Uncached input

The model's standard per-million input rate, for tokens that are neither read from nor written to the cache

Cached input

Reused prefix tokens, discounted up to 90 percent; on GPT-5.6 and later, 0.1 times the uncached input rate

Cache writes

On GPT-5.6 and later, 1.25 times the uncached input rate for newly cached prefix tokens; earlier models have no separate write charge

Minimum cacheable prefix

1,024 visible input tokens on GPT-5.6 and later; shorter prefixes are billed uncached

## How to detect

5 checks to find it in your estate.

- Track usage.input\_tokens\_details.cached\_tokens and cache\_write\_tokens against input\_tokens per request, and compute the token cache-hit rate (cached tokens divided by total input tokens) by application, user or day

- Use the Admin Usage API completions endpoint grouped by project\_id, api\_key\_id and model to compare input\_cached\_tokens, input\_cache\_write\_tokens and input\_uncached\_tokens across workloads

- Review the Prompt Caching Dashboard for low read hit rates and run the Prompt Cache Diagnostics tool on affected requests to see where the prefix diverged

- Inspect request templates for dynamic values early in the prompt, per-request tool lists or ordering, varying text.format schemas, reasoning.effort or text.verbosity, and shared prefixes just under the minimum cacheable length

- Flag workloads with high cache\_write\_tokens but low cached\_tokens, which pay the write premium without reuse

## How to fix

5 ways to remove the waste.

- Put static content first (tool definitions, developer instructions, examples, reference material) and move timestamps, IDs and per-request context to the end of the prompt

- Keep tool definitions, ordering and schemas identical across requests; use tool\_choice none or allowed\_tools to change which tools are callable instead of removing definitions

- On GPT-5.6 and later, place an explicit prompt\_cache\_breakpoint after the stable prefix and consider explicit-only mode so changing suffixes are not written to the cache at the 1.25x write rate

- On earlier models, send a stable prompt\_cache\_key for requests that share a prefix to improve cache routing, since cached state lives on individual machines and traffic above 15 requests per minute can be overflow-routed to machines without the entry

- Where a shared prefix falls just below the minimum cacheable length, weigh expanding it with useful stable instructions or examples, and measure whether reuse offsets the extra tokens

## Documentation

Vendor references for pricing and configuration.

- [Prompt caching  developers.openai.com](https://developers.openai.com/api/docs/guides/prompt-caching)

- [Pricing  developers.openai.com](https://developers.openai.com/api/docs/pricing)

- [Cost optimization  developers.openai.com](https://developers.openai.com/api/docs/guides/cost-optimization)

- [Completions  developers.openai.com](https://developers.openai.com/api/reference/typescript/resources/admin/subresources/organization/subresources/usage/methods/completions)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- OpenAI API  CER-0507

### [Latency-Tolerant OpenAI API Workloads Not Using Batch or Flex Processing](https://www.pointfive.co/efficiency-hub/inefficiencies/latency-tolerant-openai-api-workloads-not-using-batch-or-flex-processing)

Evaluations, dataset classification, data enrichment, embedding of content repositories and other asynchronous jobs are frequently run against the OpenAI API at Standard processing rates, because they were built as loops over the...

AI

- OpenAI API  CER-0509

### [Using High-Cost OpenAI Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-openai-models-for-low-complexity-tasks)

Classification, routing, triage, simple data extraction and small scoped edits are often sent to OpenAI's most capable model because a single model name is configured for the whole application. The per-token gap between tiers is large: at...

AI

- AWS SageMaker  CER-0383

### [SageMaker Studio Applications Without Idle Shutdown](https://www.pointfive.co/efficiency-hub/inefficiencies/sagemaker-studio-applications-without-idle-shutdown)

In SageMaker Studio, each running JupyterLab or Code Editor application runs on an instance that is billed for as long as the application is running, whether or not the user is doing anything. Data scientists routinely leave spaces running...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

