Skip to content
Cloud Efficiency Hub

Overprovisioned Shards in Kinesis Data Streams

The short version

In provisioned mode, a Kinesis data stream's capacity is the sum of its shards, each providing up to 1 MB per second or 1,000 records per second of writes and 2 MB per second of reads, and every shard is billed per shard-hour whether it carries data or not.

PointFive Research

Cloud cost research at PointFive

AWS service
AWS Kinesis
Category
Other
Reference
CER-0390
Type
Overprovisioned Resource

Explanation

Why the waste happens and who it affects.

Streams are commonly sized for a launch, a forecast peak or last year's traffic and then left at that shard count, and uneven partition keys create cold shards that carry little traffic while still billing in full.

The same pattern applies to consumers: each enhanced fan-out consumer registered on a provisioned stream adds a charge per consumer per shard-hour on top of its data retrieval charge, so consumers that were registered for a test or a retired application keep adding cost across every shard. AWS cost optimization guidance for Kinesis Data Streams notes that any unused or idle shard still incurs costs and recommends adjusting shard counts to the ingestion rate, merging cold shards and avoiding enhanced fan-out where it is not needed.

Billing model

The pricing dimensions that drive this cost.

Provisioned-mode streams are billed on the following dimensions (rates from the pricing page for US East, N. Virginia).

Shard-hour
Each shard is billed per hour it exists, listed at $0.015 per shard-hour, regardless of the data it carries
PUT payload units
Writes are billed per million 25 KB PUT payload units, listed at $0.014 per million
Enhanced fan-out
Billed per consumer per shard-hour (listed at $0.015) plus per GB of data retrieved (listed at $0.013)
Extended retention
Retention beyond 24 hours up to 7 days is billed per shard-hour (listed at $0.020), so it also scales with shard count

How to detect

5 checks to find it in your estate.

  • For each provisioned stream, compare the Sum of the stream-level IncomingBytes and IncomingRecords metrics per second with the stream's shard count multiplied by 1 MB per second and 1,000 records per second; streams running far below this at their peaks are overprovisioned
  • Check that the stream has no sustained WriteProvisionedThroughputExceeded or ReadProvisionedThroughputExceeded throttling before planning a reduction
  • To find cold shards, temporarily enable enhanced shard-level monitoring (which incurs CloudWatch charges) and compare IncomingBytes per ShardId
  • List enhanced fan-out consumers with ListStreamConsumers and check SubscribeToShard.Success and SubscribeToShardEvent.Bytes by ConsumerName; consumers with no active subscriptions are billed per shard-hour for nothing
  • Review the stream's retention period, since retention above 24 hours adds a per-shard-hour charge

How to fix

5 ways to remove the waste.

  • Reduce shard count with UpdateShardCount; by default a call cannot go below half the current shard count and a stream can be rescaled at most ten times per rolling 24 hours, so large reductions take several steps
  • Merge adjacent cold shards with MergeShards, and fix partition key distribution so load spreads evenly and fewer shards are needed
  • Automate shard scaling for provisioned streams with Application Auto Scaling, or switch streams with unpredictable traffic to on-demand mode, which scales shards automatically
  • Deregister enhanced fan-out consumers that are no longer used with DeregisterStreamConsumer, and use shared-throughput GetRecords consumers where dedicated throughput is not required
  • Use record aggregation in producers (for example the Kinesis Producer Library) to reduce records per second when the record-rate limit, rather than bytes, is what drives the shard count

Documentation

Vendor references for pricing and configuration.