Explanation
Why the waste happens and who it affects.
Microsoft notes that when you create a compute instance the VM stays on so it's available for your work, and it is billed per hour for as long as it runs. Data scientists routinely leave instances running overnight, over weekends and between projects, and GPU instances in particular cost the same whether a notebook is running or the kernel has been idle for days.
Idle shutdown and start/stop schedules are per-instance settings, and an instance created without them runs until someone stops it; the idle time of an existing instance also cannot be changed from the CLI or Python SDK, only from studio or the REST API. Because each data scientist typically owns one or more instances, the waste scales with team size, and it is easy to miss because instances appear as Azure Machine Learning compute rather than in the usual virtual machine reviews.
Billing model
The pricing dimensions that drive this cost.
Compute instances bill like VMs while running, plus a few charges that continue while stopped.
- Running compute instance
- Billed per hour at the rate of the chosen VM size (CPU or GPU) for as long as it is running, whether or not it is in use
- Stopped compute instance
- No VM compute charge, but the instance's P10 OS disk (120 GB) continues to bill
- Load balancer
- One load balancer per compute instance is billed per day, including while it is stopped, until the compute instance is deleted
- No service surcharge
- There is no additional charge to use Azure Machine Learning itself; the charges are for compute and the other Azure services consumed
How to detect
5 checks to find it in your estate.
- In Azure Machine Learning studio under Compute > Compute instances, or with az ml compute list, find running instances and check each for an idle shutdown setting (idleTimeBeforeShutdown) and start/stop schedules
- Assign the built-in Azure Policy definition Azure Machine Learning Compute Instance should have idle shutdown in Audit mode to list non-compliant instances across subscriptions
- Review instances by VM size and owner, prioritizing GPU sizes, and check whether they are running outside working hours or for days without jobs
- In Cost analysis, filter Service name to Virtual Machines for the Azure Machine Learning workspaces' scope and group by resource to see which compute instances accrue the most hours
- Identify stopped instances that have not been started for weeks; they still pay for the OS disk and load balancer
How to fix
5 ways to remove the waste.
- Enable idle shutdown on every compute instance; the idle period can be set between 15 minutes and three days, and an instance counts as idle only with no Jupyter kernels or terminals, no runs, no VS Code connection and no custom applications running
- Add start and stop schedules (up to four per instance) that match working hours, and use Azure Policy to append a default shutdown schedule when none exists
- Assign the built-in idle shutdown policy in Deny mode for new instances once teams are ready, so instances cannot be created without it
- If the workspace uses a managed identity, grant it contributor access to the workspace, since otherwise idle shutdown does not trigger
- Delete compute instances that are no longer used rather than leaving them stopped, to remove the disk and load balancer charges; note that VM size cannot be changed after creation, so recreate oversized instances at a smaller size
Documentation
Vendor references for pricing and configuration.
- Manage and optimize costs - Azure Machine Learninglearn.microsoft.com
- Create a compute instance - Azure Machine Learninglearn.microsoft.com
- Plan to manage costs - Azure Machine Learninglearn.microsoft.com
- Built-in policy definitions for Azure Machine Learninglearn.microsoft.com
- Pricing - Azure Machine Learningazure.microsoft.com