Explanation
Why the waste happens and who it affects.
When minvCpus is above zero, Batch keeps enough EC2 instances running to hold that many vCPUs even when no jobs are queued or running. Teams often set a nonzero minimum to shave startup time off the first job, copy it from an example, or leave it in place after a busy period, and the capacity then runs around the clock between batch windows.
The minimum also applies when the compute environment is disabled: AWS documents that minvCpus is maintained even in the DISABLED state and that disabled environments can keep incurring charges for a nonzero minimum. Idle scale-down also keeps the instance size, so an environment that launched a large instance may keep that instance at the minimum rather than shrinking to a smaller type. A configured scale-down delay adds further idle, billable time after each job finishes.
Billing model
The pricing dimensions that drive this cost.
AWS Batch has no charge of its own; the EC2 instances it launches are billed at normal EC2 rates.
- EC2 instance hours
- Instances launched by a managed compute environment are billed per second or hour at On-Demand or Spot rates while running
- Minimum vCPUs floor
- Capacity equal to minvCpus stays running and billed even with an empty queue or a DISABLED environment
- Scale-down delay
- Instances kept by minScaleDownDelayMinutes after their jobs finish are billable at standard EC2 pricing
- Attached storage
- EBS volumes on those instances are billed while the instances exist
How to detect
4 checks to find it in your estate.
- Run describe-compute-environments and list managed EC2 and SPOT environments where computeResources.minvCpus is greater than 0, including those with state DISABLED
- Compare each environment's desiredvCpus over time with job activity; long periods with desiredvCpus equal to minvCpus and no RUNNABLE or RUNNING jobs in attached queues indicate idle floor capacity
- Find EC2 instances launched by Batch (through the environment's tags or its Auto Scaling group) that stay running between job windows
- Check scalingPolicy.minScaleDownDelayMinutes on each environment and whether the retained idle time is justified by job arrival patterns
How to fix
5 ways to remove the waste.
- Set minvCpus to 0 with update-compute-environment; this is a scaling update that does not replace existing instances, and the AWS CDK reference for managed EC2 compute environments states that setting minvCpus to 0 ensures you are not billed for unused capacity
- Accept a short cold start for the first job after an idle period, or keep a small minimum only for queues with latency-sensitive jobs and document why
- Remove or shorten the scale-down delay unless jobs arrive in tight bursts that benefit from reusing warm instances
- Delete compute environments that are no longer attached to active job queues rather than leaving them disabled, since a disabled environment still maintains its minimum
- Use Spot or Fargate compute environments where jobs allow it, since Fargate environments do not use minvCpus at all
Documentation
Vendor references for pricing and configuration.