# Unoptimized Prompt and Response Length in Bedrock Inference

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/unoptimized-prompt-and-response-length-in-bedrock-inference

Amazon Bedrock on-demand inference bills every input token sent and every output token generated.

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Amazon Bedrock on-demand inference bills every input token sent and every output token generated.

PointFive Research

Cloud cost research at PointFive

AWS service

[AWS Bedrock](https://www.pointfive.co/efficiency-hub/cloud-services/aws-bedrock)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0502

Type

Excessive Ingestion or Processing

## Explanation

Why the waste happens and who it affects.

Applications often send much more than a task needs: the full conversation history on every turn, entire documents instead of the relevant passages, long system prompts with instructions and examples that no longer change the result, and verbose tool or JSON schemas. On the output side, requests are frequently made with no response length limit and no instruction to be brief, so the model returns long explanations when the caller only uses a label, a number or a short answer.

Because the cost is per token and repeats on every call, a few hundred unnecessary tokens per request become a large line item at production volume, and chat-style applications grow input per turn as history accumulates. Output tokens cost more than input: every Anthropic Claude model on the Bedrock pricing page for US East (Ohio) lists output at five times the input rate. The AWS Well-Architected Generative AI Lens lists optimizing prompt token length (GENCOST03-BP01) and controlling model response length (GENCOST03-BP02) as cost optimization best practices.

## Billing model

The pricing dimensions that drive this cost.

On-demand and batch inference on Bedrock is priced per million tokens, with separate rates per model.

Input tokens

Billed per million tokens for everything sent to the model, including system prompt, history, retrieved context and tool definitions

Output tokens

Billed per million tokens generated, typically at a higher rate than input tokens

Provisioned Throughput

Billed per model unit hour, where longer prompts and responses consume more of the purchased capacity

## How to detect

4 checks to find it in your estate.

- Chart the AWS/Bedrock CloudWatch metrics InputTokenCount and OutputTokenCount divided by Invocations, by ModelId, to track average tokens per request and spot growth over time

- Enable model invocation logging and use CloudWatch Logs Insights on input.inputTokenCount and output.outputTokenCount, grouped by identity.arn or requestMetadata, to find the applications and callers with the largest prompts and responses

- Sample logged request bodies for full conversation histories, whole documents, repeated instructions or unused tool definitions that could be trimmed or retrieved selectively

- Check whether callers set a response length limit (for example maxTokens in the Converse inferenceConfig) and compare typical output length with what downstream code actually consumes

## How to fix

5 ways to remove the waste.

- Trim prompts: remove redundant instructions and examples, keep only the conversation turns needed (or summarize older turns), and retrieve only the relevant chunks instead of whole documents; Amazon Bedrock Prompt Optimization can help rewrite verbose prompts

- Set a response length limit on every call and ask for concise or structured output, such as a label, a key or a short JSON object, where the caller does not need prose

- Re-test quality after each reduction with a representative evaluation set, since overly aggressive trimming can lower accuracy and cause retries that cost more than the savings

- For large static prefixes that cannot be removed, use prompt caching where the model supports it rather than paying the full input rate on every call

- Track tokens per request as a KPI per application so regressions are caught when prompts or templates change

## Documentation

Vendor references for pricing and configuration.

- [Amazon Bedrock Pricing  aws.amazon.com](https://aws.amazon.com/bedrock/pricing/)

- [GENCOST03-BP01 Optimize prompt token length  docs.aws.amazon.com](https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/gencost03-bp01.html)

- [GENCOST03-BP02 Control model response length  docs.aws.amazon.com](https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/gencost03-bp02.html)

- [Monitor bedrock-runtime inference using CloudWatch metrics  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/monitoring-runtime-metrics.html)

- [Monitor model invocation using CloudWatch Logs and Amazon S3  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/model-invocation-logging.html)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- AWS Bedrock  CER-0276

### [Using High-Cost Bedrock Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-bedrock-models-for-low-complexity-tasks-734ca)

Many Bedrock workloads involve low-complexity tasks such as tagging, classification, routing, entity extraction, keyword detection, document triage, or lightweight summarization. These tasks do not require the advanced reasoning or...

AI

- AWS Bedrock  CER-0256

### [Suboptimal Bedrock Inference Profile Model](https://www.pointfive.co/efficiency-hub/inefficiencies/suboptimal-bedrock-inference-profile-model-08c00)

AWS frequently updates Bedrock with improved foundation models, offering higher quality and better cost efficiency. When workloads remain tied to older model versions, token consumption may increase, latency may be higher, and output...

AI

- AWS Bedrock  CER-0558

### [Underutilized Bedrock Provisioned Throughput](https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-bedrock-provisioned-throughput)

Amazon Bedrock Provisioned Throughput reserves dedicated inference capacity for a base or custom model at a fixed hourly price, whether or not any requests arrive. Capacity is bought in model units, each delivering a set number of input...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

