Explanation
Why the waste happens and who it affects.
When that size was chosen for an anticipated peak, copied from a template or never revisited after the application changed, each instance carries more vCPU and memory than the workload uses. Because the size is multiplied by the instance count, an oversized SKU in a scale set wastes that difference on every instance.
Scale sets themselves add no charge; the bill is the compute, storage and networking of the instances they run. Autoscale can reduce the number of instances but never changes their size, so a scale set that scales correctly on count can still run consistently low CPU and memory on every node. Azure Advisor evaluates scale sets separately from standalone VMs with its High-impact recommendation 'Right-size or shutdown underutilized virtual machine scale sets'.
Billing model
The pricing dimensions that drive this cost.
Scale set cost is the sum of what each instance consumes.
- No scale set charge
- There is no extra cost for the scale set; charges come from the compute, network and storage resources it uses
- Per-instance compute
- Each instance is billed at the rate of the model's VM size for the time it runs, so cost scales with size times instance count
- Deallocated instances
- Compute stops while instances are deallocated, but managed disks and networking continue to incur charges
How to detect
5 checks to find it in your estate.
- Review Azure Advisor Cost recommendations for 'Right-size or shutdown underutilized virtual machine scale sets'; for SKU changes Advisor aggregates metrics using the maximum across instances and targets P95 CPU and outbound network at 40 percent or lower and P99 memory at 60 percent or lower on the new SKU for user-facing workloads
- Extend the Advisor lookback period from the default 7 days to 30, 60 or 90 days in Advisor Configuration so weekly and monthly peaks are included
- Chart Percentage CPU and Available Memory Percentage for the scale set split by VMName to confirm that no instance approaches the capacity of the current size
- Check whether autoscale already holds the instance count at its minimum while per-instance utilization stays low, which indicates the size rather than the count is too large
- Exclude scale sets where the size is dictated by other constraints Advisor does not measure, such as GPU, disk throughput, Accelerated Networking or Premium Storage requirements
How to fix
5 ways to remove the waste.
- Change the VM size in the scale set model to a smaller SKU in the same family, a newer version or a different family that still meets CPU, memory, network and disk requirements
- Roll the size change out according to the upgrade policy: with Rolling mode instances are updated in batches, with Manual mode existing instances keep the old size until you update them; a size change requires restarting or redeploying each instance
- Re-tune autoscale rules after resizing, since smaller instances reach thresholds sooner and may need a different minimum count
- Before changing families, check reservations and savings plans: Advisor savings are based on retail rates and do not account for existing commitments, and a cross-series move can raise cost if it leaves a reservation unused
- Shut down or delete scale sets that Advisor flags with no meaningful CPU or outbound network use over the lookback period
Documentation
Vendor references for pricing and configuration.
- Optimize virtual machine (VM) or virtual machine scale set (VMSS) spend by resizing or shutting down underutilized instanceslearn.microsoft.com
- Cost recommendations - Azure Advisorlearn.microsoft.com
- Upgrade policy modes for Virtual Machine Scale Setslearn.microsoft.com
- Azure Virtual Machine Scale Sets overviewlearn.microsoft.com
- Supported metrics - Microsoft.Compute/virtualmachineScaleSetslearn.microsoft.com