Explanation
Why the waste happens and who it affects.
Development, test, QA and demo clusters are typically used during working hours or for the length of a project, yet they run through nights, weekends and the weeks between releases. For clusters of tier M10 and above, Atlas supports pausing: a paused cluster is charged only for storage, with no charges for compute, other services or data transfer.
Pausing is manual. The Atlas pause documentation describes no scheduled pause, a paused cluster is automatically resumed after 30 days, and a resumed cluster must run for at least 60 minutes before it can be paused again, so off-hours pausing has to be done by hand or scripted. MongoDB's billing optimization guidance recommends pausing dedicated clusters that will not be used for an extended period.
Billing model
The pricing dimensions that drive this cost.
- Running cluster
- Per-hour tier rate for each data-bearing node, charged continuously while the cluster runs
- Paused cluster
- Only storage is charged, including the default storage that is otherwise bundled into the tier's hourly rate; no compute, other services or data transfer charges
- Pause limits
- Available for M10 and above, not Flex, free or NVMe clusters; Atlas resumes a paused cluster after 30 days
- Terminated cluster
- No further cluster charges; retained snapshots, if kept, continue to bill as backup storage
How to detect
4 checks to find it in your estate.
- Tag or name-filter clusters by environment and list those that are not production using the Atlas UI, Atlas CLI or Administration API
- Review Connections and Opcounters metrics for those clusters by hour of day and day of week; near-zero activity outside working hours marks candidates for a pause schedule
- Flag non-production clusters with no connections or operations for many consecutive days as candidates for pausing or termination
- Use Cost Explorer filtered by project and cluster to quantify hourly spend of non-production clusters, and check which of them have never been paused
How to fix
5 ways to remove the waste.
- Schedule pause and resume for non-production M10+ clusters outside working hours using the Atlas CLI (atlas clusters pause and atlas clusters start) or the Administration API (set paused to true or false) from an external scheduler such as a CI job or cloud function
- Account for the platform rules in the schedule: re-pause clusters that Atlas auto-resumes after 30 days, and allow at least 60 minutes of running time after a resume before pausing again
- Note the tradeoffs: a paused cluster cannot be read or written, no new backups are taken while paused, and Search Nodes data is deleted and rebuilt on resume
- Terminate clusters for finished projects, turning on Keep existing snapshots after termination first if the data may be needed later
- For small, lightly used development databases, compare the cost of a Flex cluster, which cannot be paused but is billed on usage with a monthly cap, and keep dedicated non-production tiers small
Documentation
Vendor references for pricing and configuration.