# Overprovisioned MSK Provisioned Brokers | Cloud Efficiency Hub

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-msk-provisioned-brokers

Amazon MSK Provisioned clusters are sized up front: the team picks a broker size, a broker count and a storage volume per broker, usually from a...

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Amazon MSK Provisioned clusters are sized up front: the team picks a broker size, a broker count and a storage volume per broker, usually from a peak-throughput estimate or a sizing spreadsheet, and then leaves them in place.

PointFive Research

Cloud cost research at PointFive

AWS service

[AWS MSK](https://www.pointfive.co/efficiency-hub/cloud-services/aws-msk)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0392

Type

Overprovisioned Resource

## Explanation

Why the waste happens and who it affects.

When actual traffic turns out lower than planned, or a workload moves off the cluster, the brokers keep running at low CPU and the provisioned storage stays mostly empty, while every broker-hour and every provisioned GB is still billed.

Storage is the stickier part of the problem. MSK lets you increase broker storage but not decrease it, and storage is often raised during an incident or a retention change and never revisited. Brokers are also frequently kept at a size or count chosen for a launch event or a migration that has since finished. Platform teams running shared Kafka clusters for many producers are the most affected, because no single application team owns the capacity decision.

## Billing model

The pricing dimensions that drive this cost.

MSK Provisioned clusters with Standard brokers are billed on provisioned capacity, not on the traffic the cluster carries.

Broker instance-hours

An hourly rate per broker based on broker size, billed at one-second resolution for every broker in the cluster

Broker storage

Billed per GB-month of storage provisioned per broker, whether or not it holds data

Provisioned storage throughput

Optional additional EBS throughput billed per MB/s provisioned per month

Tiered storage

Billed per GB-month of data held in the low-cost tier plus a per-GB charge for data retrieved from it, an alternative to holding long retention on broker volumes

## How to detect

5 checks to find it in your estate.

- Build a CloudWatch metric math expression of CpuUser + CpuSystem per broker and review it over several weeks; AWS recommends keeping this under 60% to retain headroom for broker failures and rolling updates, so brokers sitting far below that level across the whole period are candidates for a smaller size or fewer brokers

- Compare the PartitionCount metric per broker (which includes replicas) with the recommended partitions per broker for the current broker size in the MSK best practices; a broker size recommended for thousands of partitions that hosts only a few hundred is a sizing signal

- Review KafkaDataLogsDiskUsed per broker against the provisioned volume; consistently low percentages mean storage was provisioned well beyond what current retention needs

- Review BytesInPerSec and BytesOutPerSec per broker against the throughput the broker size and count were planned for

- Check whether provisioned storage throughput is enabled on clusters whose volume metrics (VolumeReadBytes, VolumeWriteBytes at PER\_BROKER level) show low disk activity

## How to fix

5 ways to remove the waste.

- Move to a smaller broker size with update-broker-type, which runs as a rolling update while the cluster keeps serving traffic; AWS recommends trying the smaller size on a test cluster first, and the update is blocked if partitions per broker exceed the documented maximum for the target size. Supported Standard broker moves are M5 or T3 to M7g, T3 to M5 and M7g to M5; moving down to T3 sizes is not supported

- Reduce the broker count with update-broker-count after moving all user partitions off the brokers to be removed (kafka-reassign-partitions.sh or Cruise Control) and confirming UserPartitionExists is 0 for them; removal is supported on M5 and M7g clusters on Kafka 2.8.1 and later, and the target count must be a multiple of the number of Availability Zones

- Because broker storage can only be increased, reclaim overprovisioned storage by removing brokers or by moving topics to a new, right-sized cluster; set retention.ms or retention.bytes per topic and delete unused topics so the new size holds

- For topics that need long retention, enable tiered storage so older segments move to the lower-cost tier instead of sizing broker volumes for the full retention period

- Move M5 brokers to the equivalent M7g (Graviton) size, and consider Express brokers, which bill storage for the GB used rather than provisioned, or MSK Serverless for spiky or unpredictable throughput

## Documentation

Vendor references for pricing and configuration.

- [Amazon MSK pricing  aws.amazon.com](https://aws.amazon.com/msk/pricing/)

- [Best practices for Standard brokers  docs.aws.amazon.com](https://docs.aws.amazon.com/msk/latest/developerguide/bestpractices.html)

- [Update the Amazon MSK cluster broker size  docs.aws.amazon.com](https://docs.aws.amazon.com/msk/latest/developerguide/msk-update-broker-type.html)

- [Remove a broker from an Amazon MSK cluster  docs.aws.amazon.com](https://docs.aws.amazon.com/msk/latest/developerguide/msk-remove-broker.html)

- [Manual scaling for Standard brokers  docs.aws.amazon.com](https://docs.aws.amazon.com/msk/latest/developerguide/manually-expand-storage.html)

- [Amazon MSK metrics for monitoring Standard brokers with CloudWatch  docs.aws.amazon.com](https://docs.aws.amazon.com/msk/latest/developerguide/metrics-details.html)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- AWS EKS  CER-0280

### [Fargate Resource Rounding and Per-Pod Overhead Driving Step-Up Costs](https://www.pointfive.co/efficiency-hub/inefficiencies/fargate-resource-rounding-and-per-pod-overhead-driving-step-up-costs)

Pod resource requests - often inflated by sidecar containers-push total memory or CPU just over a Fargate sizing boundary. Because Fargate adds mandatory system overhead and only supports fixed resource combinations, small incremental...

Compute

- Amazon WorkSpaces Applications  CER-0050

### [Underutilized or Overprovisioned WorkSpaces Applications Instances](https://www.pointfive.co/efficiency-hub/inefficiencies/underutilized-or-overprovisioned-appstream-instances)

Amazon WorkSpaces Applications (formerly AppStream 2.0) fleets often default to instance types designed for worst-case or peak usage scenarios, even when average workloads are significantly lighter. This leads to consistently low...

Compute

- AWS Lambda  CER-0103

### [Overprovisioned Memory Allocation for Lambda Functions](https://www.pointfive.co/efficiency-hub/inefficiencies/overprovisioned-memory-allocation-for-lambda-functions)

Each Lambda function must be configured with a memory setting, which indirectly controls the amount of CPU and networking performance allocated. In many environments, memory settings are defined arbitrarily or left unchanged as functions...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

