Explanation
Why the waste happens and who it affects.
A replica set pays for every node in its replication factor, and a sharded cluster pays for replication factor times the number of shards, so a tier that is one or two sizes too large is multiplied across the whole cluster for every hour it runs.
Tiers are often picked at launch for expected peak load or future growth, or raised during an incident or a slow-query problem, and not revisited afterwards, leaving clusters whose CPU and memory stay low for weeks. MongoDB's own billing optimization guidance lists underutilized clusters as a cost driver and recommends auto-scaling the tier, scaling dedicated clusters down to a lower tier, and optimizing slow queries that otherwise push workloads onto higher tiers. Sizing the whole cluster for occasional analytical queries is another common cause.
Billing model
The pricing dimensions that drive this cost.
- Cluster tier
- A per-hour rate for each data-bearing node, covering default RAM, storage capacity and storage speed for that tier
- Replica set nodes
- The number of billed data-bearing nodes equals the replication factor
- Sharded clusters
- Billed nodes equal replication factor times shards, and config servers are charged at a separate rate
- Analytics nodes
- Optional extra nodes that can run a different tier from the operational nodes
How to detect
5 checks to find it in your estate.
- Review Atlas cluster metrics over at least several weeks - Normalized System CPU, System Memory, WiredTiger cache usage, Opcounters and Connections - and flag clusters whose peaks stay well below the tier's capacity
- As a reference point, Atlas reactive auto-scaling scales down only when Normalized System CPU stays below 45% over the last 10 minutes and 4 hours and projected memory at the lower tier stays below 60%; clusters that would meet those conditions most of the time are candidates
- Use Cost Explorer and invoices to rank clusters by tier cost, multiplying by replication factor and shard count, to find where one tier step matters most
- Check the Performance Advisor and Query Profiler for slow or unindexed queries whose resource use is driving the current tier choice
- Identify clusters sized up for analytical or reporting queries that could run on separate analytics nodes instead
How to fix
5 ways to remove the waste.
- Scale down one tier at a time from the Edit Cluster page or the Administration API, watching CPU, memory and cache metrics after each step
- Enable compute auto-scaling with scale-down and set the minimum cluster size below the current tier so Atlas can move the cluster down during quiet periods (the auto-scaling configuration itself is covered in a separate entry)
- Create the indexes recommended by Performance Advisor and remove unused ones before sizing down, since inefficient queries otherwise force a higher tier
- For memory-bound workloads that are not CPU-bound, consider the Low-CPU class, which MongoDB describes as providing half the vCPUs of the General tier of the same size at lower cost
- Move aggregation and reporting workloads to analytics nodes on their own tier, as MongoDB suggests, instead of sizing every operational node for them
Documentation
Vendor references for pricing and configuration.