Skip to content
Cloud Efficiency Hub

Low Prompt Cache Hit Rate from Unstable Prompt Prefixes in the OpenAI API

The short version

Prompt caching is enabled by default for supported OpenAI models, and reused prefix tokens are billed at a cached-input rate discounted by up to 90 percent.

PointFive Research

Cloud cost research at PointFive

OpenAI service
OpenAI API
Category
AI
Reference
CER-0508
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

Reuse only happens when the entire rendered prefix matches: hidden system content, tool definitions and schemas, developer messages and conversation history, in that order. Anything that changes early in that sequence, such as a timestamp or user ID in the developer message, tools added, removed or reordered per request, or a switch in structured output schema, reasoning effort or verbosity, forces the rest of the prompt to be billed at the full uncached rate.

Because caching is automatic, a broken prefix raises no error; it only shows up as low cached_tokens in the usage data. On GPT-5.6 and later the stakes are higher in both directions: cache writes are billed at 1.25 times the uncached input rate, so a prefix that is written on every request but never reused costs more than no caching at all, while a well-structured prefix is read at 0.1 times the input rate. Short shared prefixes below the minimum cacheable length of 1,024 visible tokens are never cached. OpenAI documents these as the main levers in its prompt caching guide and provides a dashboard and diagnostics tool for cache misses.

Billing model

The pricing dimensions that drive this cost.

Input tokens are billed at one of three rates, not as additive fees.

Uncached input
The model's standard per-million input rate, for tokens that are neither read from nor written to the cache
Cached input
Reused prefix tokens, discounted up to 90 percent; on GPT-5.6 and later, 0.1 times the uncached input rate
Cache writes
On GPT-5.6 and later, 1.25 times the uncached input rate for newly cached prefix tokens; earlier models have no separate write charge
Minimum cacheable prefix
1,024 visible input tokens on GPT-5.6 and later; shorter prefixes are billed uncached

How to detect

5 checks to find it in your estate.

  • Track usage.input_tokens_details.cached_tokens and cache_write_tokens against input_tokens per request, and compute the token cache-hit rate (cached tokens divided by total input tokens) by application, user or day
  • Use the Admin Usage API completions endpoint grouped by project_id, api_key_id and model to compare input_cached_tokens, input_cache_write_tokens and input_uncached_tokens across workloads
  • Review the Prompt Caching Dashboard for low read hit rates and run the Prompt Cache Diagnostics tool on affected requests to see where the prefix diverged
  • Inspect request templates for dynamic values early in the prompt, per-request tool lists or ordering, varying text.format schemas, reasoning.effort or text.verbosity, and shared prefixes just under the minimum cacheable length
  • Flag workloads with high cache_write_tokens but low cached_tokens, which pay the write premium without reuse

How to fix

5 ways to remove the waste.

  • Put static content first (tool definitions, developer instructions, examples, reference material) and move timestamps, IDs and per-request context to the end of the prompt
  • Keep tool definitions, ordering and schemas identical across requests; use tool_choice none or allowed_tools to change which tools are callable instead of removing definitions
  • On GPT-5.6 and later, place an explicit prompt_cache_breakpoint after the stable prefix and consider explicit-only mode so changing suffixes are not written to the cache at the 1.25x write rate
  • On earlier models, send a stable prompt_cache_key for requests that share a prefix to improve cache routing, since cached state lives on individual machines and traffic above 15 requests per minute can be overflow-routed to machines without the entry
  • Where a shared prefix falls just below the minimum cacheable length, weigh expanding it with useful stable instructions or examples, and measure whether reuse offsets the extra tokens

Documentation

Vendor references for pricing and configuration.