Skip to content
Cloud Efficiency Hub

Glue Spark Jobs With Fixed Worker Counts Instead of Auto Scaling

The short version

AWS Glue Spark jobs are commonly configured with a fixed NumberOfWorkers sized for the heaviest stage, the largest expected input or a generous guess.

PointFive Research

Cloud cost research at PointFive

AWS service
AWS Glue
Category
Compute
Reference
CER-0394
Type
Overprovisioned Resource

Explanation

Why the waste happens and who it affects.

Without Auto Scaling, every allocated worker is billed for the whole run, including phases when most of them are idle: while the driver lists large numbers of files in S3, during stages that run with only a few executors, and when data skew concentrates work on a small part of the cluster.

Glue Auto Scaling, available for Glue 3.0 and later, adds workers when the Spark application requests more executors and removes executors and their workers when they sit idle, treating the configured worker count as a maximum. AWS lists over-provisioned stages and driver-side file listing among the scenarios where Auto Scaling helps with cost and utilization, yet many jobs, particularly those created from older templates or upgraded from earlier Glue versions, still run with a fixed worker count.

Billing model

The pricing dimensions that drive this cost.

DPU-hours
Glue Spark jobs are billed for the DPUs used over the run, per second with a 1-minute minimum; each worker type maps to a fixed number of DPUs, for example 1 DPU for G.1X and 2 DPUs for G.2X
Fixed workers
Without Auto Scaling, all requested workers are allocated for the run and billed whether or not executors are busy
Auto Scaling usage
With Auto Scaling, workers are added and removed during the run; Glue 4.0 and later batch jobs report actual usage as DPUSeconds on the job run

How to detect

3 checks to find it in your estate.

  • List Glue jobs and flag Spark jobs on Glue 3.0 or later with supported worker types that do not set --enable-auto-scaling to true in their default arguments
  • On Glue 4.0 and later, enable job observability metrics and review glue.driver.workerUtilization, the percentage of allocated workers actually used; AWS notes that when it is low, Auto Scaling can help
  • With job metrics enabled, compare glue.driver.ExecutorAllocationManager.executors.numberMaxNeededExecutors with the number of executors allocated to see how much of the fixed cluster the job actually needed

How to fix

5 ways to remove the waste.

  • Enable Auto Scaling with the Automatically scale the number of workers option in Glue Studio or the --enable-auto-scaling job argument, and treat NumberOfWorkers as the maximum
  • Set the maximum number of workers from a realistic DPU estimate rather than an extreme value, as AWS advises against very large maximums for low-volume data
  • If fixed spark.sql.shuffle.partitions or spark.default.parallelism values are needed, set them explicitly, because Auto Scaling derives them from the maximum DPU setting
  • After enabling, compare DPUSeconds and DPU hours per run before and after the change, and revisit worker type selection using observability metrics
  • Upgrade jobs on Glue 2.0 or earlier to a current Glue version so they can use Auto Scaling

Documentation

Vendor references for pricing and configuration.