Explanation
Why the waste happens and who it affects.
Between peaks, all of those nodes keep running and billing even when YARN has little or nothing scheduled. Unlike an idle cluster, the cluster is in use, so auto-termination does not help: the waste is the gap between provisioned and needed capacity over the day.
EMR managed scaling, available from EMR 5.30.0 (except 6.0.0; some newer Regions require 6.14.0 or later) for YARN applications such as Spark, Hive, Hadoop and Flink, adds and removes core and task capacity based on workload, and AWS describes it as optimizing clusters for cost and speed. Clusters created from older templates or infrastructure-as-code modules often run without any scaling policy.
Billing model
The pricing dimensions that drive this cost.
- Node charges
- Every node is billed for EC2 and attached EBS for as long as the cluster runs, whether or not YARN is using it
- EMR charge
- A per-second EMR price per instance, with a one-minute minimum, on top of the EC2 and EBS price
- Scaled capacity
- With managed scaling, nodes added for peaks are removed when demand drops, so they bill only while needed
How to detect
4 checks to find it in your estate.
- Call get-managed-scaling-policy for each long-running cluster and flag those with no managed scaling policy and no custom automatic scaling policy on their instance groups
- Review EMR CloudWatch metrics such as YARNMemoryAvailablePercentage, ContainerPending and AppsRunning over several days; sustained high available memory with no pending containers indicates capacity above demand
- Compare node counts over time with job schedules: clusters that keep the same node count through nights and weekends while jobs run only during business hours are candidates
- Confirm the cluster runs YARN-based applications and a supported release; managed scaling does not support non-YARN applications such as Presto and HBase
How to fix
5 ways to remove the waste.
- Enable managed scaling with a MinimumCapacityUnits sized for the steady baseline and a MaximumCapacityUnits sized for peaks, instead of provisioning the peak permanently
- Set MaximumCoreCapacityUnits to keep HDFS-bearing core capacity small and let most scaling happen on task nodes, and set MaximumOnDemandCapacityUnits so scaled capacity can use Spot
- Keep Spark dynamic resource allocation enabled (the default), since disabling it can cause managed scaling to scale up more than needed
- Use a recent EMR release, as AWS recommends, to benefit from fixes and features such as shuffle-data-aware scale-down
- For workloads that run as discrete jobs, consider transient clusters with auto-termination or EMR Serverless instead of a permanently running cluster
Documentation
Vendor references for pricing and configuration.