Explanation
Why the waste happens and who it affects.
In the Atlas UI, auto-scaling and scale-down are enabled by default for eligible clusters. Through the Atlas Administration API, auto-scaling is not selected by default and must be enabled explicitly, so clusters created by scripts, Terraform or other infrastructure-as-code often have it off and stay at their launch tier permanently.
Even where auto-scaling is on, two settings limit its effect on cost. Scale-down can be disabled, in which case a cluster that scales up for a spike never comes back down. And the UI sets the minimum cluster size to the current tier by default, so a cluster launched at a generous tier can scale up but never below where it started. MongoDB recommends enabling compute and storage auto-scaling for staging and production and bounding it with minimum and maximum sizes to control cost; its guidance for development and test is the opposite - leave auto-scaling off there to avoid growth in non-production spend.
Billing model
The pricing dimensions that drive this cost.
- Current tier rate
- Each data-bearing node is billed per hour at whatever tier the cluster is on, so time spent at a higher tier than needed is paid in full
- Reactive scale-down
- Happens only when all nodes have Normalized System CPU below 45% and projected memory below 60% at the lower tier over the last 10 minutes and 4 hours, at most once per 24 hours
- Tier bounds
- MinInstanceSize and maxInstanceSize limit how far auto-scaling can move the cluster; the UI defaults the minimum to the current tier
How to detect
4 checks to find it in your estate.
- List clusters with the Administration API or Terraform state and check replicationSpecs[n].regionConfigs[m].autoScaling.compute: flag enabled = false and scaleDownEnabled = false on staging and production clusters
- Flag clusters where minInstanceSize equals the current tier, which prevents any scale-down below the launch size
- Compare the tier history in the cluster's activity feed with its Normalized System CPU and System Memory metrics; clusters with clear daily or weekly cycles whose tier never changes are candidates
- Check eligibility: reactive auto-scaling applies to General and Low-CPU class clusters of M10 and above, not Local NVMe SSD clusters, so flag ineligible clusters for manual review instead
How to fix
5 ways to remove the waste.
- Enable compute auto-scaling with scaleDownEnabled = true on staging and production clusters, setting minInstanceSize and maxInstanceSize explicitly in API calls and Terraform modules so new clusters do not start with it off
- Lower minInstanceSize to the smallest tier that meets off-peak load and latency needs, so Atlas can scale down outside peaks; keep maxInstanceSize as a cost ceiling
- Keep predictive auto-scaling enabled where eligible (M30 and above, General or Low-CPU class, active for at least two weeks) so scale-ups for predictable cycles happen before load arrives rather than by oversizing the minimum
- Leave compute auto-scaling off on development and test clusters, as MongoDB recommends, and size them manually at a small tier
- Review the settings after major workload changes, since scale-down waits 24 hours after the last scale-down, provisioning or unpause
Documentation
Vendor references for pricing and configuration.