Explanation
Why the waste happens and who it affects.
If a task node running on a Spot Instance is interrupted, no data is lost and the effect on the cluster is minimal, which is why the EMR guidelines list Spot (or an instance-fleet mix) for task nodes in every application scenario they describe, including long-running clusters and data-critical workloads. Many clusters nevertheless run all task capacity, and sometimes the entire cluster, on On-Demand Instances because the cluster template was written with On-Demand defaults and never revisited.
On-Demand task capacity pays the full EC2 rate for the most interruption-tolerant part of the cluster. The waste is largest on clusters that add many task nodes for peak processing, and on transient, test, or cost-driven clusters where AWS guidance goes further and suggests running primary, core and task nodes all on Spot.
Billing model
The pricing dimensions that drive this cost.
- EC2 and EBS charges
- Each cluster node is billed at its EC2 purchasing option price (On-Demand or Spot) plus attached EBS volumes
- EMR charge
- A per-second EMR price per instance, with a one-minute minimum, added on top of the EC2 and EBS price
- Spot Instances
- Spare EC2 capacity that AWS lists at up to a 90 percent discount compared to On-Demand prices, which can be interrupted
How to detect
4 checks to find it in your estate.
- List instance groups with list-instance-groups and flag TASK groups whose Market is ON_DEMAND
- For instance fleets, use list-instance-fleets and flag TASK fleets with TargetOnDemandCapacity greater than zero and little or no TargetSpotCapacity
- Identify transient, test and development clusters (from tags, naming or auto-termination settings) where primary and core nodes are also On-Demand even though interruptions are acceptable
- For clusters with managed scaling, check whether MaximumOnDemandCapacityUnits is left at its default, which equals MaximumCapacityUnits and lets all scaled capacity be On-Demand
How to fix
5 ways to remove the waste.
- Run task capacity on Spot: add a Spot task instance group (and remove the On-Demand one, since the purchasing option of an existing group cannot be changed), or set TargetSpotCapacity on the task instance fleet
- Use instance fleets with several instance types and the price-capacity-optimized Spot allocation strategy, which AWS recommends for higher chance of getting capacity and lower interruption rates
- With managed scaling, set MaximumOnDemandCapacityUnits so baseline capacity stays On-Demand and additional scaled capacity is provisioned as Spot
- Keep primary and core nodes On-Demand where HDFS data or cluster completion matters; for transient, test and cost-driven clusters consider Spot for all node types
- On EMR 6.x and later, where YARN node labels are disabled by default, enable node labels so application master processes run only on core nodes and a Spot task node interruption does not fail the whole job
Documentation
Vendor references for pricing and configuration.