# EMR Clusters Without Managed Scaling | Cloud Efficiency Hub

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/emr-clusters-without-managed-scaling

Long-running Amazon EMR clusters are often provisioned with a fixed number of core and task nodes sized for the heaviest job or the busiest time of day.

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Long-running Amazon EMR clusters are often provisioned with a fixed number of core and task nodes sized for the heaviest job or the busiest time of day.

PointFive Research

Cloud cost research at PointFive

AWS service

[AWS EMR](https://www.pointfive.co/efficiency-hub/cloud-services/aws-emr)

Category

[Compute](https://www.pointfive.co/efficiency-hub/service-category/compute)

Reference

CER-0396

Type

Overprovisioned Resource

## Explanation

Why the waste happens and who it affects.

Between peaks, all of those nodes keep running and billing even when YARN has little or nothing scheduled. Unlike an idle cluster, the cluster is in use, so auto-termination does not help: the waste is the gap between provisioned and needed capacity over the day.

EMR managed scaling, available from EMR 5.30.0 (except 6.0.0; some newer Regions require 6.14.0 or later) for YARN applications such as Spark, Hive, Hadoop and Flink, adds and removes core and task capacity based on workload, and AWS describes it as optimizing clusters for cost and speed. Clusters created from older templates or infrastructure-as-code modules often run without any scaling policy.

## Billing model

The pricing dimensions that drive this cost.

Node charges

Every node is billed for EC2 and attached EBS for as long as the cluster runs, whether or not YARN is using it

EMR charge

A per-second EMR price per instance, with a one-minute minimum, on top of the EC2 and EBS price

Scaled capacity

With managed scaling, nodes added for peaks are removed when demand drops, so they bill only while needed

## How to detect

4 checks to find it in your estate.

- Call get-managed-scaling-policy for each long-running cluster and flag those with no managed scaling policy and no custom automatic scaling policy on their instance groups

- Review EMR CloudWatch metrics such as YARNMemoryAvailablePercentage, ContainerPending and AppsRunning over several days; sustained high available memory with no pending containers indicates capacity above demand

- Compare node counts over time with job schedules: clusters that keep the same node count through nights and weekends while jobs run only during business hours are candidates

- Confirm the cluster runs YARN-based applications and a supported release; managed scaling does not support non-YARN applications such as Presto and HBase

## How to fix

5 ways to remove the waste.

- Enable managed scaling with a MinimumCapacityUnits sized for the steady baseline and a MaximumCapacityUnits sized for peaks, instead of provisioning the peak permanently

- Set MaximumCoreCapacityUnits to keep HDFS-bearing core capacity small and let most scaling happen on task nodes, and set MaximumOnDemandCapacityUnits so scaled capacity can use Spot

- Keep Spark dynamic resource allocation enabled (the default), since disabling it can cause managed scaling to scale up more than needed

- Use a recent EMR release, as AWS recommends, to benefit from fixes and features such as shuffle-data-aware scale-down

- For workloads that run as discrete jobs, consider transient clusters with auto-termination or EMR Serverless instead of a permanently running cluster

## Documentation

Vendor references for pricing and configuration.

- [Using managed scaling in Amazon EMR  docs.aws.amazon.com](https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-managed-scaling.html)

- [Configuring Amazon EMR cluster instance types and best practices for Spot instances  docs.aws.amazon.com](https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-plan-instances-guidelines.html)

- [Monitoring Amazon EMR metrics with CloudWatch  docs.aws.amazon.com](https://docs.aws.amazon.com/emr/latest/ManagementGuide/UsingEMR_ViewingMetrics.html)

- [Amazon EMR pricing  aws.amazon.com](https://aws.amazon.com/emr/pricing/)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- AWS EMR  CER-0057

### [Idle EMR Cluster Without Auto-Termination Policy](https://www.pointfive.co/efficiency-hub/inefficiencies/idle-emr-cluster-without-auto-termination-policy)

Amazon EMR clusters often run on large, multi-node EC2 fleets, making them costly to leave running unnecessarily. If a cluster becomes idle-no longer processing jobs - but is not terminated, it continues accruing EC2 and EMR service...

Compute

- AWS EMR  CER-0395

### [On-Demand Task Nodes in EMR Clusters](https://www.pointfive.co/efficiency-hub/inefficiencies/on-demand-task-nodes-in-emr-clusters)

Amazon EMR task nodes run processing work but do not store persistent data in HDFS. If a task node running on a Spot Instance is interrupted, no data is lost and the effect on the cluster is minimal, which is why the EMR guidelines list...

Compute

- AWS EKS  CER-0280

### [Fargate Resource Rounding and Per-Pod Overhead Driving Step-Up Costs](https://www.pointfive.co/efficiency-hub/inefficiencies/fargate-resource-rounding-and-per-pod-overhead-driving-step-up-costs)

Pod resource requests - often inflated by sidecar containers-push total memory or CPU just over a Fargate sizing boundary. Because Fargate adds mandatory system overhead and only supports fixed resource combinations, small incremental...

Compute

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

