Explanation
Why the waste happens and who it affects.
Without Auto Scaling, every allocated worker is billed for the whole run, including phases when most of them are idle: while the driver lists large numbers of files in S3, during stages that run with only a few executors, and when data skew concentrates work on a small part of the cluster.
Glue Auto Scaling, available for Glue 3.0 and later, adds workers when the Spark application requests more executors and removes executors and their workers when they sit idle, treating the configured worker count as a maximum. AWS lists over-provisioned stages and driver-side file listing among the scenarios where Auto Scaling helps with cost and utilization, yet many jobs, particularly those created from older templates or upgraded from earlier Glue versions, still run with a fixed worker count.
Billing model
The pricing dimensions that drive this cost.
- DPU-hours
- Glue Spark jobs are billed for the DPUs used over the run, per second with a 1-minute minimum; each worker type maps to a fixed number of DPUs, for example 1 DPU for G.1X and 2 DPUs for G.2X
- Fixed workers
- Without Auto Scaling, all requested workers are allocated for the run and billed whether or not executors are busy
- Auto Scaling usage
- With Auto Scaling, workers are added and removed during the run; Glue 4.0 and later batch jobs report actual usage as DPUSeconds on the job run
How to detect
3 checks to find it in your estate.
- List Glue jobs and flag Spark jobs on Glue 3.0 or later with supported worker types that do not set --enable-auto-scaling to true in their default arguments
- On Glue 4.0 and later, enable job observability metrics and review glue.driver.workerUtilization, the percentage of allocated workers actually used; AWS notes that when it is low, Auto Scaling can help
- With job metrics enabled, compare glue.driver.ExecutorAllocationManager.executors.numberMaxNeededExecutors with the number of executors allocated to see how much of the fixed cluster the job actually needed
How to fix
5 ways to remove the waste.
- Enable Auto Scaling with the Automatically scale the number of workers option in Glue Studio or the --enable-auto-scaling job argument, and treat NumberOfWorkers as the maximum
- Set the maximum number of workers from a realistic DPU estimate rather than an extreme value, as AWS advises against very large maximums for low-volume data
- If fixed spark.sql.shuffle.partitions or spark.default.parallelism values are needed, set them explicitly, because Auto Scaling derives them from the maximum DPU setting
- After enabling, compare DPUSeconds and DPU hours per run before and after the change, and revisit worker type selection using observability metrics
- Upgrade jobs on Glue 2.0 or earlier to a current Glue version so they can use Auto Scaling
Documentation
Vendor references for pricing and configuration.