Explanation
Why the waste happens and who it affects.
When a cluster grows past its tier's default storage, Atlas bills the full configured custom storage amount, not just the increase over the default. After a large deletion, a data migration, an archive job or a one-time bulk load, the disk stays at its high-water mark and keeps billing for capacity the data no longer needs until someone reduces it manually.
Oversized disk can also lock in a larger tier. Atlas enforces disk-to-RAM ratios per tier, cannot auto-scale a cluster down to a tier smaller than its current disk configuration, and may raise the cluster tier automatically to accommodate storage growth. A cluster whose disk grew during a spike may therefore be unable to return to a cheaper tier even when compute load allows it. The extra capacity is also billed while the cluster is paused.
Billing model
The pricing dimensions that drive this cost.
- Default storage
- Included in the tier's per-hour rate for each data-bearing node
- Custom storage capacity
- When storage exceeds the tier default, Atlas charges for the full custom amount and does not deduct the default storage cost
- Storage auto-scaling
- Raises capacity when disk used reaches 90% on any node, targeting 70% used afterwards; it never scales storage down
- Paused clusters
- Continue to be charged for storage while compute charges stop
How to detect
4 checks to find it in your estate.
- Compare each cluster's configured storage with Disk Space Used in Atlas metrics over recent weeks and flag clusters where used space sits far below the 90% threshold that triggered past growth
- Review the project Activity Feed for storage auto-scaling events and cross-check them against later deletions, TTL cleanup, Online Archive moves or migrations that reduced data size
- Check whether the configured disk forces a larger tier than compute needs, using the disk-to-RAM ratio limits (60:1 for M10 to M40, 120:1 above) and the auto-scaling rule that a cluster cannot scale to a minimum tier smaller than its disk configuration
- Use Cost Explorer or the invoice to separate storage from compute for clusters with custom storage and rank by storage spend
How to fix
4 ways to remove the waste.
- Reduce configured storage from the Edit Cluster page or the Administration API to fit current data plus index size and headroom below the auto-scaling threshold
- Plan the reduction as a maintenance operation: AWS and Azure do not shrink volumes in place, so Atlas provisions new volumes and syncs data to them, with downtime on each node during its sync; on large clusters each node can take several hours, and AWS limits EBS volume modifications to 4 per 24 hours
- After reducing disk, revisit the tier and the auto-scaling minimum, since the smaller disk may now permit a lower tier
- Control growth at the source with TTL indexes and Online Archive for aging data so storage does not auto-scale again, and keep storage auto-scaling enabled as a safety net once capacity is right-sized
Documentation
Vendor references for pricing and configuration.