Explanation
Why the waste happens and who it affects.
Dataflow queues the job and starts it within six hours of creation, and runs it on a mix of preemptible and regular VMs with Dataflow Shuffle. In return, vCPU and memory are billed at a uniform discounted rate that Google describes as about 40% below regular Dataflow prices, regardless of which worker type ran the work.
Many batch pipelines have no need to start the moment they are submitted: nightly and weekly aggregations, backfills, reprocessing jobs and exports that only need to finish by a deadline hours away. They still run with default scheduling and pay full batch rates because FlexRS has to be requested explicitly per job.
Billing model
The pricing dimensions that drive this cost.
Dataflow bills worker resources per second, per job; rates vary by region.
- Batch worker vCPU and memory
- Billed per vCPU-hour and GiB-hour at regular batch rates
- FlexRS vCPU and memory
- Billed at a uniform discounted rate, about 40% below regular Dataflow prices, regardless of worker type
- Dataflow Shuffle
- Billed per GiB of data processed during shuffle; FlexRS enables Shuffle automatically and Shuffle is not discounted
- Persistent Disk
- FlexRS workers use 25 GB of Persistent Disk each, billed at normal rates
How to detect
4 checks to find it in your estate.
- List batch jobs and check whether the FlexRS goal was set (flexResourceSchedulingGoal in the job environment, or the flexRSGoal / flexrs_goal pipeline option); jobs without it run on regular pricing
- Identify scheduled batch jobs whose completion deadline leaves more than six hours of slack after submission
- Rank candidate jobs by vCPU and memory spend from the Cloud Billing export, since the discount applies only to those resources
- Exclude jobs that need GPUs, Compute Engine reservations, specific zones, custom autoscaling or M2, M3 or H3 machine types, which FlexRS does not support
How to fix
4 ways to remove the waste.
- Launch eligible batch jobs with --flexRSGoal=COST_OPTIMIZED (Java) or --flexrs_goal=COST_OPTIMIZED (Python and Go), including in Flex and classic templates
- Move job submission earlier in the schedule so the up-to-six-hour queue still meets downstream deadlines
- Set maxNumWorkers (max_num_workers) to bound cost, since numWorkers only sets the initial worker count under FlexRS
- Keep time-critical or interactive batch jobs on standard scheduling, and use Apache Beam SDK 2.12.0 or later
Documentation
Vendor references for pricing and configuration.