Skip to content
Cloud Efficiency Hub

Cold Data Kept on MongoDB Atlas Cluster Storage Instead of Online Archive

The short version

Event, log, audit, order-history and time-series collections in MongoDB tend to grow without limit, while applications query only recent documents.

PointFive Research

Cloud cost research at PointFive

MongoDB Atlas service
MongoDB Atlas
Category
Storage
Reference
CER-0536
Type
Excessive Data Retention

Explanation

Why the waste happens and who it affects.

When older documents stay on the cluster, they occupy disk on every data-bearing node, enlarge every backup snapshot, and add to the working set and indexes the tier must support. Because custom storage is billed for the full configured amount and storage auto-scaling only grows disks, retained history steadily raises storage and backup cost and can push the cluster to a larger tier.

MongoDB's billing optimization guidance recommends using Online Archive or TTL indexes to move older data from more expensive hot storage to less expensive cold storage. Online Archive, available on M10 and larger clusters, moves documents matching a date-based or custom rule to Atlas-managed cloud object storage, where they remain queryable read-only through Atlas Data Federation. TTL indexes delete documents automatically after a set age where the data does not need to be kept at all.

Billing model

The pricing dimensions that drive this cost.

Cluster storage
Default storage is included in the tier's hourly rate; custom storage is billed for the full configured amount
Backup storage
Snapshot storage is billed per GB-month, so data kept on the cluster is also carried in retained snapshots
Online Archive storage
Archived data is billed by GB stored and days retained, at rates that vary by cloud provider and region
Online Archive queries
$5.00 per TB of archived data processed, with a 10 MB minimum per query, plus data transferred and returned

How to detect

4 checks to find it in your estate.

  • Identify the largest collections by storage size and index size in Data Explorer or with db.collection.stats(), and check whether they hold time-stamped documents far older than the application's access window
  • Use the Query Profiler or Performance Advisor to confirm that queries on those collections filter on recent date ranges and rarely touch older documents
  • Flag collections with no TTL index (check db.collection.getIndexes() for expireAfterSeconds) and no Online Archive rule on the cluster's Online Archive tab
  • Estimate the storage, backup and tier impact of the cold portion by comparing document counts or sizes by date range, and check whether storage auto-scaling events coincide with the growth

How to fix

5 ways to remove the waste.

  • Configure Online Archive on M10+ clusters with a date-based rule (date field plus days to keep on the cluster) or a custom query, and index the date field, which MongoDB recommends for archiving performance
  • Query archived data through the federated connection string that Atlas creates, and use query limits and explain() to control the per-TB processing charges if archived data is still read
  • Set a Data Retention Period on the archive where regulations allow, so Atlas deletes archived data after 7 to 9125 days; deleted archive data cannot be recovered
  • Use TTL indexes (expireAfterSeconds on a date field) for data that can simply expire, such as sessions, logs and ephemeral events
  • After archiving or expiring large volumes, reduce the cluster's configured storage and revisit the tier, since freed space does not shrink the disk automatically

Documentation

Vendor references for pricing and configuration.