Explanation
Why the waste happens and who it affects.
When the minimum is set above zero, usually to avoid waiting a few minutes for nodes to start, that many nodes stay allocated and billed around the clock even when no jobs run. Microsoft states that any value larger than 0 keeps that number of nodes running even if they are not in use.
The cost is highest on GPU clusters and persists quietly because the cluster looks healthy and ready. Clusters created for a project or a demo often keep their warm nodes long after the work ends.
Billing model
The pricing dimensions that drive this cost.
Cluster nodes are billed as the underlying VMs while they are allocated.
- Allocated nodes
- Each node is billed per hour at its VM size rate while allocated, whether or not a job is running on it
- Minimum nodes
- The cluster never scales below min_instances, so those nodes bill continuously; with 0, idle nodes are deallocated and stop billing
- Idle time before scale down
- Nodes above the minimum stay allocated for this idle period after jobs finish, 120 seconds by default
- Load balancer
- Every 50 nodes of a compute cluster carry one standard load balancer billed per day, until the cluster is deleted
How to detect
5 checks to find it in your estate.
- List compute clusters with az ml compute list --type AmlCompute (or in studio under Compute > Compute clusters) and flag those whose min_instances is greater than 0
- Use az ml compute list-nodes on flagged clusters to see allocated nodes, and compare with job history to find nodes that sit idle between jobs or for days
- Prioritize clusters with GPU VM sizes and dedicated tier, where idle minimum nodes cost the most
- Check idle_time_before_scale_down on rarely used clusters; long values keep scaled-up nodes billed after each job
- In Cost Management, review the VM costs attributed to each cluster over periods with no training runs
How to fix
5 ways to remove the waste.
- Set the minimum nodes to 0 with az ml compute update --min-instances 0, the SDK or studio, so Azure Machine Learning can deallocate nodes when no jobs run
- Tune idle time before scale down to the team's iteration pattern: shorter for occasional jobs, longer only for rapid dev/test iteration where constant scale-up and scale-down would cost more
- Where start-up latency matters, accept a few minutes of node provisioning instead of paying for always-on nodes, or move ad hoc jobs to Azure Machine Learning serverless compute, which removes cluster lifecycle management
- Use low-priority tier for interruptible, checkpointed training and batch jobs; since March 31, 2026 these nodes are allocated as Spot VMs at variable Spot rates
- Delete clusters that are no longer used, which also releases their quota and load balancer charges
Documentation
Vendor references for pricing and configuration.