Explanation
Why the waste happens and who it affects.
Bedrock prompt caching lets supported models reuse a previously processed prompt prefix, and tokens read from the cache are billed at the model's much lower cache-read rate. When caching is not set up, or the prompt is structured so the prefix never matches, every one of those repeated tokens is billed again at the full input rate.
Anthropic models and Amazon Nova on Bedrock also attempt Implicit Prompt Caching, but AWS describes it as best effort with variable hit rates, so explicit cache checkpoints are the reliable way to capture the saving. Common causes of missed savings are not adding checkpoints, placing dynamic content such as timestamps, user IDs or retrieved chunks before the static content, changing tool definitions between calls (which invalidates the system and messages caches after them), prefixes shorter than the model's minimum checkpoint size, and request gaps longer than the cache TTL. The AWS Well-Architected Generative AI Lens (GENCOST03-BP03) lists prompt caching as a cost optimization best practice.
Billing model
The pricing dimensions that drive this cost.
Prompt caching changes how input tokens are priced on on-demand inference; it is not supported with the batch inference API.
- Standard input tokens
- Billed per million tokens at the model's normal input rate for tokens not read from or written to the cache
- Cache write tokens
- Billed when a prefix is written to the cache, on some models above the standard input rate (Sonnet 4.6 lists $3.75 for 5-minute and $6.00 for 1-hour writes)
- Cache read tokens
- Billed at a reduced rate, for example $0.30 versus $3.00 per million input tokens for Claude Sonnet 4.6 (US East Ohio, global cross-Region inference)
- Cache TTL
- Cached prefixes expire after the TTL (5 minutes by default, 1 hour on supported models), which resets on each cache hit
How to detect
5 checks to find it in your estate.
- Compare the AWS/Bedrock CloudWatch metrics CacheReadInputTokenCount and CacheWriteInputTokenCount with InputTokenCount by ModelId; models that support caching with high input volume and near-zero cache reads are candidates
- In Converse responses or invocation logs, check cacheReadInputTokens and cacheWriteInputTokens; remember that inputTokens then counts only uncached tokens
- Identify applications whose requests share a large static prefix, such as the same system prompt, tool schemas or reference documents, across many calls
- Review prompt templates for dynamic values placed before static content, and for changing tool definitions, both of which prevent prefix matches
- Watch for high cache writes with few reads, which means the cache is expiring before reuse or the prefix is changing, and caching is adding cost instead of saving it
How to fix
5 ways to remove the waste.
- Add explicit cache checkpoints (cachePoint in the Converse API, cache_control for Claude in InvokeModel) at the end of the static content in the tools, system and messages fields, up to the model's maximum (4 for Claude); for Claude models a single checkpoint at the end of the static content is enough for Bedrock's simplified cache management to find the longest matching prefix
- Reorder prompts so stable content comes first (tools, then system, then messages) and variable content last, and keep tool definitions and system prompts byte-identical across calls
- Make sure the cached prefix meets the model's minimum tokens per checkpoint (for example 1,024 for Claude Sonnet 4.6 and 4,096 for Claude Haiku 4.5), otherwise the request succeeds but nothing is cached
- Use the 1-hour TTL on supported models only for prefixes reused at intervals longer than 5 minutes, since 1-hour cache writes cost more than 5-minute writes
- Measure savings after rollout by comparing cache read volume with cache write volume and total input cost, and remove checkpoints that only generate writes
Documentation
Vendor references for pricing and configuration.