Explanation
Why the waste happens and who it affects.
Google's memory management guidance is explicit that instance capacity is the amount of memory you provision and what you are charged for, so an instance holding 4 GiB of keys in a 30 GiB instance pays for all 30 GiB, every second it runs.
Oversizing is common because capacity is chosen up front from rough estimates with generous safety margins, copied from production into staging, or set for a data set that later shrank after TTLs, eviction policies or application changes. Standard Tier adds a replica and cross-zone replication on top, which is often unnecessary for non-production caches. Google's memory management guidance recommends right-sizing instances rather than creating over-provisioned ones, and scaling preserves data. Instances with no traffic at all are a separate idle-resource case.
Billing model
The pricing dimensions that drive this cost.
Memorystore for Redis charges depend on how the instance is provisioned, not how much memory it uses:
- Provisioned capacity
- Billed per GiB-hour in 1-second increments for the full provisioned capacity
- Capacity tiers
- Instances fall into tiers M1 to M5 by size; the per-GiB rate is lower in larger tiers, and larger tiers give more network throughput
- Service tier
- Basic Tier is a standalone instance; Standard Tier adds automatic cross-zone replication and failover at a higher price, plus optional read replicas
- Committed use discounts
- 1-year and 3-year commitments discount the per-GiB rate for M2 and larger tiers; the pricing page lists no CUD rate for M1
How to detect
5 checks to find it in your estate.
- Compare redis.googleapis.com/stats/memory/usage (used memory) and stats/memory/usage_ratio with the provisioned capacity over at least 30 days, including peak periods, to find instances whose used memory stays far below capacity
- Check stats/memory/system_memory_usage_ratio to confirm there is headroom; Google flags values above 80% as memory pressure, so only instances well below that are downsizing candidates
- Review stats/cache_hit_ratio and stats/evicted_keys; a high hit ratio with few or no evictions at low memory usage indicates capacity beyond what the working set needs
- Flag Standard Tier instances and instances with read replicas in development, test and staging projects where high availability is not required
- Rank instances by Memorystore cost in the Cloud Billing export to focus on the largest
How to fix
5 ways to remove the waste.
- Scale instance capacity down to observed peak usage plus headroom that keeps the system memory usage ratio below 80%; for Standard Tier the new size must be greater than the data currently stored, and Standard Tier reserves 10% of capacity as a replication buffer
- Scale during low write traffic and export the data first as Google recommends; scaling preserves data and the IP address but causes a brief connection reset, so clients need retry logic
- Compare total cost across capacity tiers, since a smaller instance can drop into a tier with a higher per-GiB rate
- For non-production caches that do not need failover, create a Basic Tier instance and migrate with export and import, because Memorystore cannot change an instance between Basic and Standard Tier
- Apply committed use discounts only to capacity that remains after rightsizing
Documentation
Vendor references for pricing and configuration.