# Missing Prompt Caching for Repeated Context in Amazon Bedrock

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/missing-prompt-caching-for-repeated-context-in-amazon-bedrock

Many Bedrock applications send the same large block of context on every request: a long system prompt, tool definitions, a reference document in a...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Many Bedrock applications send the same large block of context on every request: a long system prompt, tool definitions, a reference document in a document Q&A flow, or an agent's growing conversation history.

PointFive Research

Cloud cost research at PointFive

AWS service

[AWS Bedrock](https://www.pointfive.co/efficiency-hub/cloud-services/aws-bedrock)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0565

Type

Inefficient Configuration

## Explanation

Why the waste happens and who it affects.

Bedrock prompt caching lets supported models reuse a previously processed prompt prefix, and tokens read from the cache are billed at the model's much lower cache-read rate. When caching is not set up, or the prompt is structured so the prefix never matches, every one of those repeated tokens is billed again at the full input rate.

Anthropic models and Amazon Nova on Bedrock also attempt Implicit Prompt Caching, but AWS describes it as best effort with variable hit rates, so explicit cache checkpoints are the reliable way to capture the saving. Common causes of missed savings are not adding checkpoints, placing dynamic content such as timestamps, user IDs or retrieved chunks before the static content, changing tool definitions between calls (which invalidates the system and messages caches after them), prefixes shorter than the model's minimum checkpoint size, and request gaps longer than the cache TTL. The AWS Well-Architected Generative AI Lens (GENCOST03-BP03) lists prompt caching as a cost optimization best practice.

## Billing model

The pricing dimensions that drive this cost.

Prompt caching changes how input tokens are priced on on-demand inference; it is not supported with the batch inference API.

Standard input tokens

Billed per million tokens at the model's normal input rate for tokens not read from or written to the cache

Cache write tokens

Billed when a prefix is written to the cache, on some models above the standard input rate (Sonnet 4.6 lists $3.75 for 5-minute and $6.00 for 1-hour writes)

Cache read tokens

Billed at a reduced rate, for example $0.30 versus $3.00 per million input tokens for Claude Sonnet 4.6 (US East Ohio, global cross-Region inference)

Cache TTL

Cached prefixes expire after the TTL (5 minutes by default, 1 hour on supported models), which resets on each cache hit

## How to detect

5 checks to find it in your estate.

- Compare the AWS/Bedrock CloudWatch metrics CacheReadInputTokenCount and CacheWriteInputTokenCount with InputTokenCount by ModelId; models that support caching with high input volume and near-zero cache reads are candidates

- In Converse responses or invocation logs, check cacheReadInputTokens and cacheWriteInputTokens; remember that inputTokens then counts only uncached tokens

- Identify applications whose requests share a large static prefix, such as the same system prompt, tool schemas or reference documents, across many calls

- Review prompt templates for dynamic values placed before static content, and for changing tool definitions, both of which prevent prefix matches

- Watch for high cache writes with few reads, which means the cache is expiring before reuse or the prefix is changing, and caching is adding cost instead of saving it

## How to fix

5 ways to remove the waste.

- Add explicit cache checkpoints (cachePoint in the Converse API, cache\_control for Claude in InvokeModel) at the end of the static content in the tools, system and messages fields, up to the model's maximum (4 for Claude); for Claude models a single checkpoint at the end of the static content is enough for Bedrock's simplified cache management to find the longest matching prefix

- Reorder prompts so stable content comes first (tools, then system, then messages) and variable content last, and keep tool definitions and system prompts byte-identical across calls

- Make sure the cached prefix meets the model's minimum tokens per checkpoint (for example 1,024 for Claude Sonnet 4.6 and 4,096 for Claude Haiku 4.5), otherwise the request succeeds but nothing is cached

- Use the 1-hour TTL on supported models only for prefixes reused at intervals longer than 5 minutes, since 1-hour cache writes cost more than 5-minute writes

- Measure savings after rollout by comparing cache read volume with cache write volume and total input cost, and remove checkpoints that only generate writes

## Documentation

Vendor references for pricing and configuration.

- [Prompt caching for faster model inference  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html)

- [GENCOST03-BP03 Implement prompt caching to reduce token costs  docs.aws.amazon.com](https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/gencost03-bp03.html)

- [Amazon Bedrock Pricing  aws.amazon.com](https://aws.amazon.com/bedrock/pricing/)

- [Monitor bedrock-runtime inference using CloudWatch metrics  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/monitoring-runtime-metrics.html)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- AWS Bedrock  CER-0276

### [Using High-Cost Bedrock Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-bedrock-models-for-low-complexity-tasks-734ca)

Many Bedrock workloads involve low-complexity tasks such as tagging, classification, routing, entity extraction, keyword detection, document triage, or lightweight summarization. These tasks do not require the advanced reasoning or...

AI

- AWS Bedrock  CER-0256

### [Suboptimal Bedrock Inference Profile Model](https://www.pointfive.co/efficiency-hub/inefficiencies/suboptimal-bedrock-inference-profile-model-08c00)

AWS frequently updates Bedrock with improved foundation models, offering higher quality and better cost efficiency. When workloads remain tied to older model versions, token consumption may increase, latency may be higher, and output...

AI

- AWS Bedrock  CER-0502

### [Unoptimized Prompt and Response Length in Bedrock Inference](https://www.pointfive.co/efficiency-hub/inefficiencies/unoptimized-prompt-and-response-length-in-bedrock-inference)

Amazon Bedrock on-demand inference bills every input token sent and every output token generated. Applications often send much more than a task needs: the full conversation history on every turn, entire documents instead of the relevant...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

