# Real-Time Bedrock Inference Used for Non-Urgent Batch Workloads

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/real-time-bedrock-inference-used-for-non-urgent-batch-workloads

Offline jobs such as document enrichment, classification and tagging backfills, summarizing archives, generating embeddings and running model...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Offline jobs such as document enrichment, classification and tagging backfills, summarizing archives, generating embeddings and running model evaluations are often implemented as loops of synchronous InvokeModel or Converse calls.

PointFive Research

Cloud cost research at PointFive

AWS service

[AWS Bedrock](https://www.pointfive.co/efficiency-hub/cloud-services/aws-bedrock)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0566

Type

Suboptimal Pricing Model

## Explanation

Why the waste happens and who it affects.

Those calls are billed at Standard on-demand rates, even though nobody is waiting for the individual responses. Amazon Bedrock offers two cheaper options for this kind of work: batch inference, which AWS prices at 50 percent below on-demand for select models, and the Flex service tier, which the Bedrock pricing page lists at a 50 percent discount to Standard for models that support it.

The pattern is common because synchronous calls are the first thing developers build, and moving a pipeline to asynchronous S3-based batch jobs takes extra work. It is compounded when latency-tolerant traffic is sent with the Priority tier, which the pricing page lists at a 75 percent premium to Standard. The AWS Well-Architected Generative AI Lens (GENCOST02) asks teams to select a cost-effective pricing model among provisioned, on-demand, hosted and batch options.

## Billing model

The pricing dimensions that drive this cost.

Bedrock text model inference is billed per million input and output tokens, with the rate depending on how the request is served.

Standard on-demand tokens

Default rate for synchronous InvokeModel and Converse requests without a service\_tier setting

Batch inference tokens

For select models, billed at 50 percent below on-demand rates for asynchronous jobs that read from and write to Amazon S3

Flex tier tokens

50 percent discount to Standard for supported models, in exchange for longer processing times

Priority tier tokens

75 percent premium to Standard for the fastest response times

## How to detect

4 checks to find it in your estate.

- Identify scheduled jobs, backfills, ETL steps and evaluation runs that call InvokeModel or Converse in loops, for example by grouping model invocation logs by identity.arn or requestMetadata and looking for bursts from non-interactive roles

- Look for token spikes in the AWS/Bedrock InputTokenCount and OutputTokenCount metrics that line up with batch schedules rather than user traffic

- Check the ServiceTier and ResolvedServiceTier CloudWatch dimensions to find latency-tolerant workloads sent with the Priority tier or left on Standard where Flex is supported

- Confirm on the Bedrock pricing page and model cards that the model in use offers batch pricing or the Flex tier, since not every model does (for example, the pricing page lists batch as N/A for some of the newest Claude models)

## How to fix

5 ways to remove the waste.

- Move offline jobs to batch inference: write requests as JSONL records in InvokeModel or Converse format to S3, submit a batch inference job, and read results from the S3 output location

- Use the Flex tier by setting service\_tier to flex for latency-tolerant online or agentic traffic on supported models where batch jobs are impractical

- Reserve Standard and Priority for interactive, user-facing paths that need fast responses

- Check batch inference limitations before migrating: it is not supported for provisioned models, does not support tool calling or structured output, processes each record independently, and prompt caching is not available with the batch API

- Use EventBridge job state notifications instead of polling so pipelines can pick up batch results automatically

## Documentation

Vendor references for pricing and configuration.

- [Amazon Bedrock Pricing  aws.amazon.com](https://aws.amazon.com/bedrock/pricing/)

- [Process multiple prompts with batch inference  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/batch-inference.html)

- [Service tiers for optimizing performance and cost  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/service-tiers-inference.html)

- [Generative AI pricing model  docs.aws.amazon.com](https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/gencost02.html)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- AWS Bedrock  CER-0276

### [Using High-Cost Bedrock Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-bedrock-models-for-low-complexity-tasks-734ca)

Many Bedrock workloads involve low-complexity tasks such as tagging, classification, routing, entity extraction, keyword detection, document triage, or lightweight summarization. These tasks do not require the advanced reasoning or...

AI

- AWS Bedrock  CER-0256

### [Suboptimal Bedrock Inference Profile Model](https://www.pointfive.co/efficiency-hub/inefficiencies/suboptimal-bedrock-inference-profile-model-08c00)

AWS frequently updates Bedrock with improved foundation models, offering higher quality and better cost efficiency. When workloads remain tied to older model versions, token consumption may increase, latency may be higher, and output...

AI

- AWS Bedrock  CER-0502

### [Unoptimized Prompt and Response Length in Bedrock Inference](https://www.pointfive.co/efficiency-hub/inefficiencies/unoptimized-prompt-and-response-length-in-bedrock-inference)

Amazon Bedrock on-demand inference bills every input token sent and every output token generated. Applications often send much more than a task needs: the full conversation history on every turn, entire documents instead of the relevant...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

