Explanation
Why the waste happens and who it affects.
When actual traffic turns out lower than planned, or a workload moves off the cluster, the brokers keep running at low CPU and the provisioned storage stays mostly empty, while every broker-hour and every provisioned GB is still billed.
Storage is the stickier part of the problem. MSK lets you increase broker storage but not decrease it, and storage is often raised during an incident or a retention change and never revisited. Brokers are also frequently kept at a size or count chosen for a launch event or a migration that has since finished. Platform teams running shared Kafka clusters for many producers are the most affected, because no single application team owns the capacity decision.
Billing model
The pricing dimensions that drive this cost.
MSK Provisioned clusters with Standard brokers are billed on provisioned capacity, not on the traffic the cluster carries.
- Broker instance-hours
- An hourly rate per broker based on broker size, billed at one-second resolution for every broker in the cluster
- Broker storage
- Billed per GB-month of storage provisioned per broker, whether or not it holds data
- Provisioned storage throughput
- Optional additional EBS throughput billed per MB/s provisioned per month
- Tiered storage
- Billed per GB-month of data held in the low-cost tier plus a per-GB charge for data retrieved from it, an alternative to holding long retention on broker volumes
How to detect
5 checks to find it in your estate.
- Build a CloudWatch metric math expression of CpuUser + CpuSystem per broker and review it over several weeks; AWS recommends keeping this under 60% to retain headroom for broker failures and rolling updates, so brokers sitting far below that level across the whole period are candidates for a smaller size or fewer brokers
- Compare the PartitionCount metric per broker (which includes replicas) with the recommended partitions per broker for the current broker size in the MSK best practices; a broker size recommended for thousands of partitions that hosts only a few hundred is a sizing signal
- Review KafkaDataLogsDiskUsed per broker against the provisioned volume; consistently low percentages mean storage was provisioned well beyond what current retention needs
- Review BytesInPerSec and BytesOutPerSec per broker against the throughput the broker size and count were planned for
- Check whether provisioned storage throughput is enabled on clusters whose volume metrics (VolumeReadBytes, VolumeWriteBytes at PER_BROKER level) show low disk activity
How to fix
5 ways to remove the waste.
- Move to a smaller broker size with update-broker-type, which runs as a rolling update while the cluster keeps serving traffic; AWS recommends trying the smaller size on a test cluster first, and the update is blocked if partitions per broker exceed the documented maximum for the target size. Supported Standard broker moves are M5 or T3 to M7g, T3 to M5 and M7g to M5; moving down to T3 sizes is not supported
- Reduce the broker count with update-broker-count after moving all user partitions off the brokers to be removed (kafka-reassign-partitions.sh or Cruise Control) and confirming UserPartitionExists is 0 for them; removal is supported on M5 and M7g clusters on Kafka 2.8.1 and later, and the target count must be a multiple of the number of Availability Zones
- Because broker storage can only be increased, reclaim overprovisioned storage by removing brokers or by moving topics to a new, right-sized cluster; set retention.ms or retention.bytes per topic and delete unused topics so the new size holds
- For topics that need long retention, enable tiered storage so older segments move to the lower-cost tier instead of sizing broker volumes for the full retention period
- Move M5 brokers to the equivalent M7g (Graviton) size, and consider Express brokers, which bill storage for the GB used rather than provisioned, or MSK Serverless for spiky or unpredictable throughput
Documentation
Vendor references for pricing and configuration.
- Amazon MSK pricingaws.amazon.com
- Best practices for Standard brokersdocs.aws.amazon.com
- Update the Amazon MSK cluster broker sizedocs.aws.amazon.com
- Remove a broker from an Amazon MSK clusterdocs.aws.amazon.com
- Manual scaling for Standard brokersdocs.aws.amazon.com
- Amazon MSK metrics for monitoring Standard brokers with CloudWatchdocs.aws.amazon.com