Skip to content
Cloud Efficiency Hub

Unoptimized Prompt and Response Length in Bedrock Inference

The short version

Amazon Bedrock on-demand inference bills every input token sent and every output token generated.

PointFive Research

Cloud cost research at PointFive

AWS service
AWS Bedrock
Category
AI
Reference
CER-0502
Type
Excessive Ingestion or Processing

Explanation

Why the waste happens and who it affects.

Applications often send much more than a task needs: the full conversation history on every turn, entire documents instead of the relevant passages, long system prompts with instructions and examples that no longer change the result, and verbose tool or JSON schemas. On the output side, requests are frequently made with no response length limit and no instruction to be brief, so the model returns long explanations when the caller only uses a label, a number or a short answer.

Because the cost is per token and repeats on every call, a few hundred unnecessary tokens per request become a large line item at production volume, and chat-style applications grow input per turn as history accumulates. Output tokens cost more than input: every Anthropic Claude model on the Bedrock pricing page for US East (Ohio) lists output at five times the input rate. The AWS Well-Architected Generative AI Lens lists optimizing prompt token length (GENCOST03-BP01) and controlling model response length (GENCOST03-BP02) as cost optimization best practices.

Billing model

The pricing dimensions that drive this cost.

On-demand and batch inference on Bedrock is priced per million tokens, with separate rates per model.

Input tokens
Billed per million tokens for everything sent to the model, including system prompt, history, retrieved context and tool definitions
Output tokens
Billed per million tokens generated, typically at a higher rate than input tokens
Provisioned Throughput
Billed per model unit hour, where longer prompts and responses consume more of the purchased capacity

How to detect

4 checks to find it in your estate.

  • Chart the AWS/Bedrock CloudWatch metrics InputTokenCount and OutputTokenCount divided by Invocations, by ModelId, to track average tokens per request and spot growth over time
  • Enable model invocation logging and use CloudWatch Logs Insights on input.inputTokenCount and output.outputTokenCount, grouped by identity.arn or requestMetadata, to find the applications and callers with the largest prompts and responses
  • Sample logged request bodies for full conversation histories, whole documents, repeated instructions or unused tool definitions that could be trimmed or retrieved selectively
  • Check whether callers set a response length limit (for example maxTokens in the Converse inferenceConfig) and compare typical output length with what downstream code actually consumes

How to fix

5 ways to remove the waste.

  • Trim prompts: remove redundant instructions and examples, keep only the conversation turns needed (or summarize older turns), and retrieve only the relevant chunks instead of whole documents; Amazon Bedrock Prompt Optimization can help rewrite verbose prompts
  • Set a response length limit on every call and ask for concise or structured output, such as a label, a key or a short JSON object, where the caller does not need prose
  • Re-test quality after each reduction with a representative evaluation set, since overly aggressive trimming can lower accuracy and cause retries that cost more than the savings
  • For large static prefixes that cannot be removed, use prompt caching where the model supports it rather than paying the full input rate on every call
  • Track tokens per request as a KPI per application so regressions are caught when prompts or templates change

Documentation

Vendor references for pricing and configuration.