# Underutilized Bedrock Provisioned Throughput

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-bedrock-provisioned-throughput

Amazon Bedrock Provisioned Throughput reserves dedicated inference capacity for a base or custom model at a fixed hourly price, whether or not any...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Amazon Bedrock Provisioned Throughput reserves dedicated inference capacity for a base or custom model at a fixed hourly price, whether or not any requests arrive.

PointFive Research

Cloud cost research at PointFive

AWS service

[AWS Bedrock](https://www.pointfive.co/efficiency-hub/cloud-services/aws-bedrock)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0558

Type

Overprovisioned Resource

## Explanation

Why the waste happens and who it affects.

Capacity is bought in model units, each delivering a set number of input and output tokens per minute. Teams often size a purchase for a launch, a load test or an expected peak, or buy it for a custom model that ends up being called only occasionally, and the reserved capacity then sits mostly idle while the hourly charge continues.

The commitment options make the waste sticky. A Provisioned Throughput by Model Units with a 1-month or 6-month commitment cannot be deleted until the term ends, only its name, tags and custom model association can be edited after purchase, it renews automatically at the end of each term, and billing continues until it is deleted. The AWS Well-Architected Generative AI Lens (GENCOST02-BP01) advises validating scaling requirements with shorter commitments to avoid over-provisioning. Idle Provisioned Throughput is most common on custom models, because a customized model needs Provisioned Throughput unless it qualifies for on-demand custom model deployment.

## Billing model

The pricing dimensions that drive this cost.

Model unit hours

Each Provisioned Throughput is billed hourly per model unit at a rate set by the model, whether or not it serves requests

Commitment term

No commitment, 1 month or 6 months, with lower hourly rates for longer terms and no deletion before a committed term ends

Fixed capacity

Model units cannot be changed on a Provisioned Throughput by Model Units, so resizing means buying a new one

On-demand alternative

Per-token billing for base models and for supported custom model deployments, with no charge when idle

## How to detect

4 checks to find it in your estate.

- List Provisioned Throughputs with ListProvisionedModelThroughputs and record model, model units, commitment duration, commitment expiration time and status for each

- For each provisioned model ARN, chart the AWS/Bedrock metrics Invocations, InputTokenCount and OutputTokenCount per minute and compare peak and average tokens per minute with the purchased model unit capacity, which AWS provides through your account team

- Flag Provisioned Throughputs with zero or near-zero invocations over 7 to 30 days, including those attached to custom models that are no longer used by any application

- Review commitment expiration dates ahead of time, since AWS documents that Provisioned Throughput renews at the end of each commitment term and billing continues until it is deleted

## How to fix

5 ways to remove the waste.

- Delete idle no-commitment Provisioned Throughputs, and plan deletion of committed ones for their expiration date so they do not auto-renew; for Provisioned Throughput by Tokens, cancel auto renew or reduce the configured tokens per minute. Deleting a Provisioned Throughput does not delete the custom model

- Right-size by purchasing a new Provisioned Throughput with fewer model units and moving traffic to it, since model units on an existing purchase cannot be reduced

- Move low or bursty traffic to on-demand inference, including on-demand custom model deployments where supported (Amazon Nova Lite, Nova 2 Lite, Micro and Pro in US East (N. Virginia) and Llama 3.3 70B Instruct in US West (Oregon), for models customized on or after July 16, 2025)

- Validate sizing with a no-commitment Provisioned Throughput before buying 1-month or 6-month terms, and only buy the longer term when sustained utilization justifies it

- Reassign a Provisioned Throughput for a custom model to another custom model derived from the same base model instead of buying new capacity when a model is replaced

## Documentation

Vendor references for pricing and configuration.

- [Increase model invocation capacity with Provisioned Throughput in Amazon Bedrock  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/prov-throughput.html)

- [Modify a Provisioned Throughput  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/prov-thru-edit.html)

- [Delete a Provisioned Throughput or cancel auto renew  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/prov-thru-delete.html)

- [Deploy a custom model for on-demand inference  docs.aws.amazon.com](https://docs.aws.amazon.com/bedrock/latest/userguide/deploy-custom-model-on-demand.html)

- [GENCOST02-BP01 Balance cost and performance when selecting inference paradigms  docs.aws.amazon.com](https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/gencost02-bp01.html)

- [Amazon Bedrock Pricing  aws.amazon.com](https://aws.amazon.com/bedrock/pricing/)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- AWS Bedrock  CER-0276

### [Using High-Cost Bedrock Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-bedrock-models-for-low-complexity-tasks-734ca)

Many Bedrock workloads involve low-complexity tasks such as tagging, classification, routing, entity extraction, keyword detection, document triage, or lightweight summarization. These tasks do not require the advanced reasoning or...

AI

- AWS Bedrock  CER-0256

### [Suboptimal Bedrock Inference Profile Model](https://www.pointfive.co/efficiency-hub/inefficiencies/suboptimal-bedrock-inference-profile-model-08c00)

AWS frequently updates Bedrock with improved foundation models, offering higher quality and better cost efficiency. When workloads remain tied to older model versions, token consumption may increase, latency may be higher, and output...

AI

- AWS Bedrock  CER-0502

### [Unoptimized Prompt and Response Length in Bedrock Inference](https://www.pointfive.co/efficiency-hub/inefficiencies/unoptimized-prompt-and-response-length-in-bedrock-inference)

Amazon Bedrock on-demand inference bills every input token sent and every output token generated. Applications often send much more than a task needs: the full conversation history on every turn, entire documents instead of the relevant...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

