Skip to content
Cloud Efficiency Hub

Compute Auto-Scaling Disabled or Bounded at Launch Tier on MongoDB Atlas Clusters

The short version

Atlas compute auto-scaling moves a dedicated cluster between a minimum and maximum tier based on sustained CPU and memory usage, so clusters with daily or weekly load cycles can run on a smaller tier outside their peaks.

PointFive Research

Cloud cost research at PointFive

MongoDB Atlas service
MongoDB Atlas
Category
Databases
Reference
CER-0532
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

In the Atlas UI, auto-scaling and scale-down are enabled by default for eligible clusters. Through the Atlas Administration API, auto-scaling is not selected by default and must be enabled explicitly, so clusters created by scripts, Terraform or other infrastructure-as-code often have it off and stay at their launch tier permanently.

Even where auto-scaling is on, two settings limit its effect on cost. Scale-down can be disabled, in which case a cluster that scales up for a spike never comes back down. And the UI sets the minimum cluster size to the current tier by default, so a cluster launched at a generous tier can scale up but never below where it started. MongoDB recommends enabling compute and storage auto-scaling for staging and production and bounding it with minimum and maximum sizes to control cost; its guidance for development and test is the opposite - leave auto-scaling off there to avoid growth in non-production spend.

Billing model

The pricing dimensions that drive this cost.

Current tier rate
Each data-bearing node is billed per hour at whatever tier the cluster is on, so time spent at a higher tier than needed is paid in full
Reactive scale-down
Happens only when all nodes have Normalized System CPU below 45% and projected memory below 60% at the lower tier over the last 10 minutes and 4 hours, at most once per 24 hours
Tier bounds
MinInstanceSize and maxInstanceSize limit how far auto-scaling can move the cluster; the UI defaults the minimum to the current tier

How to detect

4 checks to find it in your estate.

  • List clusters with the Administration API or Terraform state and check replicationSpecs[n].regionConfigs[m].autoScaling.compute: flag enabled = false and scaleDownEnabled = false on staging and production clusters
  • Flag clusters where minInstanceSize equals the current tier, which prevents any scale-down below the launch size
  • Compare the tier history in the cluster's activity feed with its Normalized System CPU and System Memory metrics; clusters with clear daily or weekly cycles whose tier never changes are candidates
  • Check eligibility: reactive auto-scaling applies to General and Low-CPU class clusters of M10 and above, not Local NVMe SSD clusters, so flag ineligible clusters for manual review instead

How to fix

5 ways to remove the waste.

  • Enable compute auto-scaling with scaleDownEnabled = true on staging and production clusters, setting minInstanceSize and maxInstanceSize explicitly in API calls and Terraform modules so new clusters do not start with it off
  • Lower minInstanceSize to the smallest tier that meets off-peak load and latency needs, so Atlas can scale down outside peaks; keep maxInstanceSize as a cost ceiling
  • Keep predictive auto-scaling enabled where eligible (M30 and above, General or Low-CPU class, active for at least two weeks) so scale-ups for predictable cycles happen before load arrives rather than by oversizing the minimum
  • Leave compute auto-scaling off on development and test clusters, as MongoDB recommends, and size them manually at a small tier
  • Review the settings after major workload changes, since scale-down waits 24 hours after the last scale-down, provisioning or unpause

Documentation

Vendor references for pricing and configuration.