Skip to content
Cloud Efficiency Hub

Real-Time Bedrock Inference Used for Non-Urgent Batch Workloads

The short version

Offline jobs such as document enrichment, classification and tagging backfills, summarizing archives, generating embeddings and running model evaluations are often implemented as loops of synchronous InvokeModel or Converse calls.

PointFive Research

Cloud cost research at PointFive

AWS service
AWS Bedrock
Category
AI
Reference
CER-0566
Type
Suboptimal Pricing Model

Explanation

Why the waste happens and who it affects.

Those calls are billed at Standard on-demand rates, even though nobody is waiting for the individual responses. Amazon Bedrock offers two cheaper options for this kind of work: batch inference, which AWS prices at 50 percent below on-demand for select models, and the Flex service tier, which the Bedrock pricing page lists at a 50 percent discount to Standard for models that support it.

The pattern is common because synchronous calls are the first thing developers build, and moving a pipeline to asynchronous S3-based batch jobs takes extra work. It is compounded when latency-tolerant traffic is sent with the Priority tier, which the pricing page lists at a 75 percent premium to Standard. The AWS Well-Architected Generative AI Lens (GENCOST02) asks teams to select a cost-effective pricing model among provisioned, on-demand, hosted and batch options.

Billing model

The pricing dimensions that drive this cost.

Bedrock text model inference is billed per million input and output tokens, with the rate depending on how the request is served.

Standard on-demand tokens
Default rate for synchronous InvokeModel and Converse requests without a service_tier setting
Batch inference tokens
For select models, billed at 50 percent below on-demand rates for asynchronous jobs that read from and write to Amazon S3
Flex tier tokens
50 percent discount to Standard for supported models, in exchange for longer processing times
Priority tier tokens
75 percent premium to Standard for the fastest response times

How to detect

4 checks to find it in your estate.

  • Identify scheduled jobs, backfills, ETL steps and evaluation runs that call InvokeModel or Converse in loops, for example by grouping model invocation logs by identity.arn or requestMetadata and looking for bursts from non-interactive roles
  • Look for token spikes in the AWS/Bedrock InputTokenCount and OutputTokenCount metrics that line up with batch schedules rather than user traffic
  • Check the ServiceTier and ResolvedServiceTier CloudWatch dimensions to find latency-tolerant workloads sent with the Priority tier or left on Standard where Flex is supported
  • Confirm on the Bedrock pricing page and model cards that the model in use offers batch pricing or the Flex tier, since not every model does (for example, the pricing page lists batch as N/A for some of the newest Claude models)

How to fix

5 ways to remove the waste.

  • Move offline jobs to batch inference: write requests as JSONL records in InvokeModel or Converse format to S3, submit a batch inference job, and read results from the S3 output location
  • Use the Flex tier by setting service_tier to flex for latency-tolerant online or agentic traffic on supported models where batch jobs are impractical
  • Reserve Standard and Priority for interactive, user-facing paths that need fast responses
  • Check batch inference limitations before migrating: it is not supported for provisioned models, does not support tool calling or structured output, processes each record independently, and prompt caching is not available with the batch API
  • Use EventBridge job state notifications instead of polling so pipelines can pick up batch results automatically

Documentation

Vendor references for pricing and configuration.