Explanation
Why the waste happens and who it affects.
Services default to spreading tasks across Availability Zones, and many teams also spread by instance ID or use random placement. These strategies distribute tasks thinly, leaving a fragment of unused CPU and memory on every instance. The cluster then needs more EC2 instances for the same set of tasks, and scale-in rarely frees a whole instance because every instance still hosts a few tasks.
The binpack strategy places each task on the instance with the least remaining CPU or memory that can still fit it, which AWS describes as minimizing the number of container instances in use, and the ECS cluster auto scaling documentation calls it the most efficient strategy in terms of capacity. Because container instances are billed as ordinary EC2 instances whether they are full or nearly empty, fragmentation translates directly into extra instance-hours. This applies to the EC2 launch type and EC2 capacity providers, not to Fargate.
Billing model
The pricing dimensions that drive this cost.
- Container instances
- Billed as regular EC2 instances per second of run time, regardless of how many tasks or how much reserved CPU and memory they host
- Cluster size
- The number of instances is driven by how tasks are packed; unused fragments on each instance still count as paid capacity
How to detect
4 checks to find it in your estate.
- List services and standalone task launches and inspect placementStrategy; flag those with no binpack entry, or with spread on instanceId (or host) or random
- Compare the cluster-level CloudWatch metrics CPUReservation and MemoryReservation (AWS/ECS namespace) with the instance count; low reservation spread across many instances indicates capacity that packing could free
- For clusters with EC2 capacity providers and managed scaling, check the CapacityProviderReservation metric and the targetCapacity setting, and whether scale-in actually removes instances
- Check per-instance remainingResources from DescribeContainerInstances to see whether free CPU and memory is scattered across many instances rather than concentrated on a few
How to fix
5 ways to remove the waste.
- Use spread on attribute:ecs.availability-zone followed by binpack on memory (or cpu, whichever is the binding resource), which keeps tasks balanced across Availability Zones while packing instances within each zone
- With managed scaling and a targetCapacity below 100 percent, place binpack before spread in the strategy order, as the ECS documentation requires, so the capacity provider does not keep scaling out to give tasks their own instances
- Turn on managed scaling and managed termination protection on the EC2 capacity provider so instances freed by packing are actually terminated
- Keep instance-level spread only for services that genuinely need one task per instance for resilience, and consider the daemon scheduling strategy where exactly one task per instance is the intent
- Right-size task CPU and memory reservations first, since inflated reservations limit how many tasks fit on an instance no matter which strategy is used
Documentation
Vendor references for pricing and configuration.
- Use strategies to define Amazon ECS task placementdocs.aws.amazon.com
- Example Amazon ECS task placement strategiesdocs.aws.amazon.com
- Automatically manage Amazon ECS capacity with cluster auto scalingdocs.aws.amazon.com
- Monitor Amazon ECS using CloudWatchdocs.aws.amazon.com
- EC2 On-Demand Instance Pricingaws.amazon.com