# Missing Prompt Caching on the Claude API | Cloud Efficiency Hub

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/missing-prompt-caching-on-the-claude-api

Agents, coding assistants, RAG applications and multi-turn chat on the Claude API resend the same long prefix on every request: tool definitions, a...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Agents, coding assistants, RAG applications and multi-turn chat on the Claude API resend the same long prefix on every request: tool definitions, a system prompt, reference documents and the conversation so far.

PointFive Research

Cloud cost research at PointFive

Anthropic service

[Claude API](https://www.pointfive.co/efficiency-hub/cloud-services/claude-api)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0503

Type

Inefficient Configuration

## Explanation

Why the waste happens and who it affects.

Without prompt caching, every one of those tokens is billed at the full input rate on every call, even though only the last user message changed. On the Claude API caching is not applied implicitly; it has to be requested with a cache\_control field, either once at the top level of the request (automatic caching) or on individual content blocks (explicit breakpoints).

Even teams that enable caching often get few cache hits. The cache matches an exact prefix in the order tools, system, messages, so a timestamp or request ID early in the system prompt, tool definitions that change between calls, JSON keys serialized in random order, or a breakpoint placed on a block that changes every request all prevent reads. Prompts shorter than the model's minimum cacheable length are silently not cached. Anthropic lists prompt caching among its cost optimization strategies and, in its own measurements, calls it the largest lever for agent workloads.

## Billing model

The pricing dimensions that drive this cost.

Prompt caching is priced as multipliers on the model's base input token rate.

Base input tokens

Uncached input billed per million tokens at the model's input rate

Cache write

1.25 times the base input rate for the 5-minute cache and 2 times for the 1-hour cache

Cache read

0.1 times the base input rate on most models, lower on some newer models, so a 5-minute cache pays off after a single read

Stacking

Caching multipliers stack with the Batch API discount and data residency pricing

## How to detect

4 checks to find it in your estate.

- Pull the Usage and Cost Admin API usage report (/v1/organizations/usage\_report/messages) grouped by model, workspace or API key and compare uncached input, cache creation and cache read tokens; high uncached input with near-zero cache reads marks missing or broken caching; Anthropic reports that agent loops read a median 84 percent of input from cache and that below about 80 percent something is usually breaking the cache

- In application logs, check the usage block of each response: if cache\_creation\_input\_tokens and cache\_read\_input\_tokens are both 0 the prompt was not cached, and repeated cache\_creation with no later cache\_read means the prefix changes between calls

- Review high-volume request templates for large identical leading content, such as tool schemas, system prompts, few-shot examples and documents, sent without any cache\_control

- Look for prefix breakers: timestamps or per-request IDs in the system prompt, tool definitions or tool\_choice that vary per call, toggling images, thinking or effort settings between turns, and unstable JSON key ordering; the API's cache diagnostics can compare consecutive requests and report which part of the prompt diverged

## How to fix

5 ways to remove the waste.

- Enable automatic caching with a top-level cache\_control field for multi-turn conversations, or place explicit breakpoints on the last block that stays identical across requests, such as the end of the tools and system prompt

- Move all static content to the start of the prompt and all per-request content, including timestamps and the incoming message, after the last breakpoint

- Use up to four breakpoints for sections that change at different rates, and add an intermediate breakpoint when a conversation can grow by 20 or more blocks per turn, since the lookback only checks 20 positions

- Choose the TTL by request cadence: the 5-minute cache is refreshed at no cost on each hit, while the 1-hour cache costs 2 times base input to write and suits gaps of 5 to 60 minutes

- Confirm the cached prefix meets the model's minimum cacheable length, and track cache read share over time using the usage fields or the Usage and Cost API

## Documentation

Vendor references for pricing and configuration.

- [Prompt caching - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)

- [Pricing - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/about-claude/pricing)

- [Usage and Cost API - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/manage-claude/usage-cost-api)

- [Optimizing for cost and intelligence - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Claude API  CER-0504

### [Latency-Tolerant Claude API Workloads Not Using the Message Batches API](https://www.pointfive.co/efficiency-hub/inefficiencies/latency-tolerant-claude-api-workloads-not-using-the-message-batches-api)

Many Claude API workloads have no user waiting on the answer: evaluation suites, content moderation backlogs, bulk classification and tagging, data extraction, summarization of document sets and generation of product descriptions. When...

AI

- Claude API  CER-0505

### [Outdated Claude Model Versions on the Claude API](https://www.pointfive.co/efficiency-hub/inefficiencies/outdated-claude-model-versions-on-the-claude-api)

Claude API requests are billed at the rate of the model ID they name, and applications usually pin an ID in code or configuration. When Anthropic ships a newer model in the same family, pinned workloads keep running on the older one....

AI

- Claude API  CER-0506

### [Using High-Cost Claude Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-claude-models-for-low-complexity-tasks)

Routing, classification, tagging, entity extraction, short summarization and similar high-volume tasks are often sent to the same top-tier Claude model that powers an application's hardest agentic work, simply because one model ID is...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

