Explanation
Why the waste happens and who it affects.
Applications often send much more than a task needs: the full conversation history on every turn, entire documents instead of the relevant passages, long system prompts with instructions and examples that no longer change the result, and verbose tool or JSON schemas. On the output side, requests are frequently made with no response length limit and no instruction to be brief, so the model returns long explanations when the caller only uses a label, a number or a short answer.
Because the cost is per token and repeats on every call, a few hundred unnecessary tokens per request become a large line item at production volume, and chat-style applications grow input per turn as history accumulates. Output tokens cost more than input: every Anthropic Claude model on the Bedrock pricing page for US East (Ohio) lists output at five times the input rate. The AWS Well-Architected Generative AI Lens lists optimizing prompt token length (GENCOST03-BP01) and controlling model response length (GENCOST03-BP02) as cost optimization best practices.
Billing model
The pricing dimensions that drive this cost.
On-demand and batch inference on Bedrock is priced per million tokens, with separate rates per model.
- Input tokens
- Billed per million tokens for everything sent to the model, including system prompt, history, retrieved context and tool definitions
- Output tokens
- Billed per million tokens generated, typically at a higher rate than input tokens
- Provisioned Throughput
- Billed per model unit hour, where longer prompts and responses consume more of the purchased capacity
How to detect
4 checks to find it in your estate.
- Chart the AWS/Bedrock CloudWatch metrics InputTokenCount and OutputTokenCount divided by Invocations, by ModelId, to track average tokens per request and spot growth over time
- Enable model invocation logging and use CloudWatch Logs Insights on input.inputTokenCount and output.outputTokenCount, grouped by identity.arn or requestMetadata, to find the applications and callers with the largest prompts and responses
- Sample logged request bodies for full conversation histories, whole documents, repeated instructions or unused tool definitions that could be trimmed or retrieved selectively
- Check whether callers set a response length limit (for example maxTokens in the Converse inferenceConfig) and compare typical output length with what downstream code actually consumes
How to fix
5 ways to remove the waste.
- Trim prompts: remove redundant instructions and examples, keep only the conversation turns needed (or summarize older turns), and retrieve only the relevant chunks instead of whole documents; Amazon Bedrock Prompt Optimization can help rewrite verbose prompts
- Set a response length limit on every call and ask for concise or structured output, such as a label, a key or a short JSON object, where the caller does not need prose
- Re-test quality after each reduction with a representative evaluation set, since overly aggressive trimming can lower accuracy and cause retries that cost more than the savings
- For large static prefixes that cannot be removed, use prompt caching where the model supports it rather than paying the full input rate on every call
- Track tokens per request as a KPI per application so regressions are caught when prompts or templates change
Documentation
Vendor references for pricing and configuration.
- Amazon Bedrock Pricingaws.amazon.com
- GENCOST03-BP01 Optimize prompt token lengthdocs.aws.amazon.com
- GENCOST03-BP02 Control model response lengthdocs.aws.amazon.com
- Monitor bedrock-runtime inference using CloudWatch metricsdocs.aws.amazon.com
- Monitor model invocation using CloudWatch Logs and Amazon S3docs.aws.amazon.com