Explanation
Why the waste happens and who it affects.
Exporters such as kube-state-metrics or node exporters expose hundreds of metrics by default, and they are often scraped in full at short intervals. Labels with many distinct values (pod names, instance IDs, request paths, user or tenant IDs) multiply the number of time series, so sample volume grows with cluster size and churn rather than with what anyone actually looks at.
Google states that the number of samples ingested is the primary contributor to Managed Service for Prometheus cost. Metrics that are never queried, charted or alerted on are billed the same as ones that are, and scrape intervals copied from examples are often shorter than dashboards and alerts need, so sample volume grows as clusters are added without any gain in visibility.
Billing model
The pricing dimensions that drive this cost.
Prometheus-format metrics in Cloud Monitoring are metered by samples, not bytes.
- Samples ingested
- Billed per million samples in decreasing volume tiers, with samples counted per billing account
- Scalar sample count
- Each point written to a time series counts as 1 sample
- Distribution sample count
- Each histogram point counts as 2 samples plus samples for its non-zero buckets (1 per bucket for explicit buckets)
- Scrape interval
- Samples per series scale linearly with scrape frequency; Google notes that moving from 10-second to 30-second sampling cuts sample volume by 66%
How to detect
5 checks to find it in your estate.
- Open the Cloud Monitoring Metrics Management page and sort metrics by billable samples ingested to find the largest prometheus.googleapis.com metrics
- Use the Metrics Management view of unused billable metrics - active metrics not queried in the last 30 days and not used in any custom dashboard or alerting policy
- Check label cardinality per metric on the Metrics Management page and flag labels such as pod, instance, path or ID values that create many series
- In Metrics Explorer, chart the Metric Ingestion Attribution resource's Samples written by attribution id metric grouped by attribution_dimension and metric_type to attribute volume to namespaces
- Review the interval field in PodMonitoring and ClusterPodMonitoring resources (or scrape_interval in self-deployed configs) for intervals shorter than dashboards and alerts require
How to fix
5 ways to remove the waste.
- Drop metrics nobody uses with metricRelabeling rules in PodMonitoring or ClusterPodMonitoring (keep action as an allowlist, drop action as a denylist), or with metric-exclusion rules on the Metrics Management page; excluded metrics are not billed
- Lengthen scrape intervals where alerting and dashboard resolution allow, for example 10 seconds to 30 or 60 seconds; this reduces resolution for short-lived spikes
- For self-deployed collection, aggregate away high-cardinality labels such as instance with recording rules and export only the aggregates using the --export.match filter
- Set sample_limit in scrape configs as a guard against a misconfigured exporter suddenly emitting a high-cardinality metric; Google recommends it as protection, not as the main control
- Remove or bucket unbounded label values (user IDs, full URLs) at the instrumentation level so series counts stay bounded
Documentation
Vendor references for pricing and configuration.