Explanation
Why the waste happens and who it affects.
Without prompt caching, every one of those tokens is billed at the full input rate on every call, even though only the last user message changed. On the Claude API caching is not applied implicitly; it has to be requested with a cache_control field, either once at the top level of the request (automatic caching) or on individual content blocks (explicit breakpoints).
Even teams that enable caching often get few cache hits. The cache matches an exact prefix in the order tools, system, messages, so a timestamp or request ID early in the system prompt, tool definitions that change between calls, JSON keys serialized in random order, or a breakpoint placed on a block that changes every request all prevent reads. Prompts shorter than the model's minimum cacheable length are silently not cached. Anthropic lists prompt caching among its cost optimization strategies and, in its own measurements, calls it the largest lever for agent workloads.
Billing model
The pricing dimensions that drive this cost.
Prompt caching is priced as multipliers on the model's base input token rate.
- Base input tokens
- Uncached input billed per million tokens at the model's input rate
- Cache write
- 1.25 times the base input rate for the 5-minute cache and 2 times for the 1-hour cache
- Cache read
- 0.1 times the base input rate on most models, lower on some newer models, so a 5-minute cache pays off after a single read
- Stacking
- Caching multipliers stack with the Batch API discount and data residency pricing
How to detect
4 checks to find it in your estate.
- Pull the Usage and Cost Admin API usage report (/v1/organizations/usage_report/messages) grouped by model, workspace or API key and compare uncached input, cache creation and cache read tokens; high uncached input with near-zero cache reads marks missing or broken caching; Anthropic reports that agent loops read a median 84 percent of input from cache and that below about 80 percent something is usually breaking the cache
- In application logs, check the usage block of each response: if cache_creation_input_tokens and cache_read_input_tokens are both 0 the prompt was not cached, and repeated cache_creation with no later cache_read means the prefix changes between calls
- Review high-volume request templates for large identical leading content, such as tool schemas, system prompts, few-shot examples and documents, sent without any cache_control
- Look for prefix breakers: timestamps or per-request IDs in the system prompt, tool definitions or tool_choice that vary per call, toggling images, thinking or effort settings between turns, and unstable JSON key ordering; the API's cache diagnostics can compare consecutive requests and report which part of the prompt diverged
How to fix
5 ways to remove the waste.
- Enable automatic caching with a top-level cache_control field for multi-turn conversations, or place explicit breakpoints on the last block that stays identical across requests, such as the end of the tools and system prompt
- Move all static content to the start of the prompt and all per-request content, including timestamps and the incoming message, after the last breakpoint
- Use up to four breakpoints for sections that change at different rates, and add an intermediate breakpoint when a conversation can grow by 20 or more blocks per turn, since the lookback only checks 20 positions
- Choose the TTL by request cadence: the 5-minute cache is refreshed at no cost on each hit, while the 1-hour cache costs 2 times base input to write and suits gaps of 5 to 60 minutes
- Confirm the cached prefix meets the model's minimum cacheable length, and track cache read share over time using the usage fields or the Usage and Cost API
Documentation
Vendor references for pricing and configuration.