Explanation
Why the waste happens and who it affects.
Glue also offers a Flex execution class that runs jobs on spare compute capacity at a lower DPU-hour rate. AWS positions Flex for non-urgent work such as pre-production jobs, testing, one-time data loads, backfills, and nightly or weekend processing where variable start and end times are acceptable.
Many such jobs still run as standard because standard is the default and the execution class is rarely revisited. Development and test environments, historical reprocessing and batch jobs with generous completion windows then pay the standard rate for capacity that does not need guaranteed start times.
Billing model
The pricing dimensions that drive this cost.
Glue publishes separate standard and Flex DPU-hour rates for Glue 3.0 to 5.1 and for Glue 6.0 and later.
- Standard execution
- Spark jobs are billed per DPU-hour, per second with a 1-minute minimum; the Glue pricing examples use $0.44 per DPU-hour
- Flex execution
- Billed per worker at the lower Flex DPU-hour rate for the time each worker actually ran; the pricing page examples use $0.29 per DPU-hour
- Reclaimed workers
- Flex workers can be reclaimed during a run, and billing for reclaimed capacity stops while the job continues on the remaining workers
How to detect
4 checks to find it in your estate.
- List Glue jobs with GetJobs and flag Spark jobs on Glue 3.0 or later whose ExecutionClass is STANDARD (or unset) and that use G.1X or G.2X workers
- Prioritize jobs in development, test and staging accounts, jobs tagged as non-production, and one-time or backfill jobs
- Review job schedules and downstream dependencies to find batch jobs whose completion window is much longer than their typical run time
- Estimate the saving from each job's DPU-hours per month (from job run history or the DPU hours column in the Glue Studio monitoring view) multiplied by the rate difference
How to fix
4 ways to remove the waste.
- Set the execution class to FLEX on eligible jobs in Glue Studio or with the ExecutionClass parameter of CreateJob, UpdateJob or StartJobRun
- Keep STANDARD for jobs with tight SLAs, streaming jobs and jobs that feed time-critical pipelines, since Flex start and run times can vary
- Review job timeouts and retries for Flex jobs so longer or delayed runs do not fail unnecessarily, and consider job run queuing for non-time-sensitive batch jobs
- Note that Flex requires Glue 3.0 or later and G.1X or G.2X workers; the newer G.12X, G.16X and R worker types do not support Flex
Documentation
Vendor references for pricing and configuration.