Skip to content
Cloud Efficiency Hub

Datadog APM Span Ingestion and Indexing Beyond Host Allotments

The short version

Datadog APM bills per APM host, and each host brings a monthly allotment of ingested span volume and indexed spans.

PointFive Research

Cloud cost research at PointFive

Datadog service
Datadog APM
Category
Other
Reference
CER-0542
Type
Excessive Ingestion or Processing

Explanation

Why the waste happens and who it affects.

Usage above the pooled allotment is billed separately per GB ingested and per million spans indexed. Overage usually comes from a few high-throughput services: a sampling rule or DD_TRACE_SAMPLE_RATE set to 100% during an investigation and never reverted, a root service whose sampling decision pulls in every downstream span of the trace, or custom retention filters that index a large share of all spans at a high retention rate.

Much of this data is not needed. Datadog computes request, error and duration (RED) trace metrics from 100% of application traffic regardless of sampling, and the default Intelligent Retention Filter keeps a representative set of traces without counting toward indexed span usage. Ingesting and indexing far beyond that pays for spans that are rarely searched, and because ingestion and indexing are metered separately from the host fee, the overage often goes unnoticed until the invoice.

Billing model

The pricing dimensions that drive this cost.

APM combines a per-host fee with two volume dimensions that are only charged above the allotment pooled across all APM hosts.

APM host
Billed per underlying host per month on the 99th percentile of hourly counts ($31 per host per month on the published APM billing page), including 150 GB of ingested spans and 1 million indexed spans
Ingested spans
Billed per GB of span data ingested per month above the pooled allotment ($0.10 per GB on the published APM billing page)
Indexed spans
Billed per million spans indexed by retention filters per month above the pooled allotment ($1.70 per million with default 15-day retention)
Intelligent Retention Filter
Diversity sampling and 1% flat sampling whose indexed spans do not count toward indexed span usage

How to detect

5 checks to find it in your estate.

  • On the APM Ingestion Control page, check the Monthly Ingestion panel: it shows ingestion so far this month, projected end-of-month ingestion and the monthly allotment, and flags when the projection exceeds the allotment
  • Sort the Ingestion Control service table by Downstream Bytes/s to find the root services whose sampling decisions drive most ingested volume, and check the Configuration column for APM Local or APM Remote rules that set a high fixed rate
  • Chart datadog.estimated_usage.apm.ingested_bytes by service and env, and datadog.estimated_usage.apm.indexed_spans by retention filter, against the allotment (150 GB and 1 million indexed spans per APM host)
  • On the Retention Filters settings page, review the Retention Rate and Spans Indexed columns for custom filters that index broad queries at or near 100%
  • Open the service ingestion summary for the top services and check the Ingestion reasons breakdown and the Sampling rates by resource table for rules at or near 100%, such as a global DD_TRACE_SAMPLE_RATE of 1.0 left from a debugging session

How to fix

5 ways to remove the waste.

  • Move high-throughput services to adaptive sampling with a monthly ingestion target, or set resource-based sampling rates so health checks and high-volume endpoints are sampled lower than business-critical ones
  • Lower the Agent head-based target with DD_APM_TARGET_TPS (10 traces per second per Agent by default), reduce DD_APM_ERROR_TPS, or keep the rare sampler disabled when the default mechanism drives most volume
  • Narrow custom retention filters to the queries that matter, such as errors, high latency and key endpoints, and lower their retention rates; rely on the Intelligent Retention Filter for general sampling
  • Build dashboards and monitors on trace.* metrics, which are computed from 100% of traffic, or on metrics generated from spans, instead of indexing spans only to count them; span-based metrics are billed as custom metrics, so avoid high-cardinality group-bys
  • Tradeoff: reducing ingestion lowers the number of traces available for investigation and affects count-based metrics from spans and trace analytics monitors, so check those before lowering rates

Documentation

Vendor references for pricing and configuration.