Explanation
Why the waste happens and who it affects.
Each deployment under an endpoint runs on its own dedicated VM instances, chosen by instance type and instance count, and Microsoft states that costs apply to the virtual machines assigned to the deployment. Those VMs bill for as long as the deployment exists, whether it receives heavy traffic, a trickle of requests, or none at all.
Deployments accumulate quietly. Blue/green rollouts leave the previous deployment in place with 0 percent of traffic, test and demo endpoints outlive the project, deployments for retired model versions are never removed, and failed deployments can still incur charges if they got as far as creating compute. Endpoints serving GPU models are the most expensive case, because each idle instance is a GPU VM.
Billing model
The pricing dimensions that drive this cost.
Managed online endpoints are billed for the compute and networking they use, with no Azure Machine Learning surcharge.
- Deployment instances
- Each deployment is billed for its VM instance type multiplied by its instance count for as long as the deployment exists, including deployments that receive 0 percent of traffic
- Quota reservation
- For many VM SKUs an extra 20 percent of quota is reserved for upgrades; it does not incur cost unless system operations use it
- Failed deployments
- A failed deployment that passed the compute creation stage incurs charges until it is deleted
- Managed virtual network
- If outbound traffic is secured with a managed virtual network, Private Link and FQDN outbound rules are charged separately
How to detect
5 checks to find it in your estate.
- List endpoints and deployments (az ml online-endpoint list and az ml online-deployment list, or studio > Endpoints) and review each endpoint's traffic allocation; deployments with 0 percent traffic and no mirrored traffic are candidates for deletion
- Chart the endpoint RequestsPerMinute metric split by the deployment dimension over 14 to 30 days to find deployments with little or no traffic
- Check the deployment metrics DeploymentCapacity and CpuUtilizationPercentage or GpuUtilizationPercentage for instances that are provisioned but nearly idle
- Find deployments in a failed provisioning state that still exist
- In Cost Analysis, filter to the workspace resource and use the azuremlendpoint and azuremldeployment tags to see the cost of each endpoint and deployment
How to fix
5 ways to remove the waste.
- Delete deployments that no longer receive traffic, such as the old side of a completed blue/green rollout, and delete endpoints used only for tests or demos
- Delete failed deployments once debugging is finished
- For deployments that must stay, lower the instance count and configure Azure Monitor autoscale with metric rules and schedule-based profiles (for example fewer instances on weekends), weighing lower minimums against availability; Microsoft recommends at least 3 instances for high availability
- Move latency-tolerant or periodic scoring to batch endpoints, which run as jobs on compute clusters that deallocate when the job completes and can scale to zero
- Right-size the instance type using utilization metrics, and move CPU-sufficient models off GPU SKUs
Documentation
Vendor references for pricing and configuration.
- Online endpoints for real-time inference - Azure Machine Learninglearn.microsoft.com
- Manage and optimize costs - Azure Machine Learninglearn.microsoft.com
- View costs for managed online endpoints - Azure Machine Learninglearn.microsoft.com
- Autoscale online endpoints - Azure Machine Learninglearn.microsoft.com
- Azure Machine Learning monitoring data referencelearn.microsoft.com
- What are batch endpoints? - Azure Machine Learninglearn.microsoft.com