Skip to content
Cloud Efficiency Hub

Azure Machine Learning Compute Clusters With Nonzero Minimum Nodes

The short version

Azure Machine Learning compute clusters (AmlCompute) scale up when training or batch inference jobs are submitted and scale back down to the configured minimum node count when jobs finish.

PointFive Research

Cloud cost research at PointFive

Category
AI
Reference
CER-0457
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

When the minimum is set above zero, usually to avoid waiting a few minutes for nodes to start, that many nodes stay allocated and billed around the clock even when no jobs run. Microsoft states that any value larger than 0 keeps that number of nodes running even if they are not in use.

The cost is highest on GPU clusters and persists quietly because the cluster looks healthy and ready. Clusters created for a project or a demo often keep their warm nodes long after the work ends.

Billing model

The pricing dimensions that drive this cost.

Cluster nodes are billed as the underlying VMs while they are allocated.

Allocated nodes
Each node is billed per hour at its VM size rate while allocated, whether or not a job is running on it
Minimum nodes
The cluster never scales below min_instances, so those nodes bill continuously; with 0, idle nodes are deallocated and stop billing
Idle time before scale down
Nodes above the minimum stay allocated for this idle period after jobs finish, 120 seconds by default
Load balancer
Every 50 nodes of a compute cluster carry one standard load balancer billed per day, until the cluster is deleted

How to detect

5 checks to find it in your estate.

  • List compute clusters with az ml compute list --type AmlCompute (or in studio under Compute > Compute clusters) and flag those whose min_instances is greater than 0
  • Use az ml compute list-nodes on flagged clusters to see allocated nodes, and compare with job history to find nodes that sit idle between jobs or for days
  • Prioritize clusters with GPU VM sizes and dedicated tier, where idle minimum nodes cost the most
  • Check idle_time_before_scale_down on rarely used clusters; long values keep scaled-up nodes billed after each job
  • In Cost Management, review the VM costs attributed to each cluster over periods with no training runs

How to fix

5 ways to remove the waste.

  • Set the minimum nodes to 0 with az ml compute update --min-instances 0, the SDK or studio, so Azure Machine Learning can deallocate nodes when no jobs run
  • Tune idle time before scale down to the team's iteration pattern: shorter for occasional jobs, longer only for rapid dev/test iteration where constant scale-up and scale-down would cost more
  • Where start-up latency matters, accept a few minutes of node provisioning instead of paying for always-on nodes, or move ad hoc jobs to Azure Machine Learning serverless compute, which removes cluster lifecycle management
  • Use low-priority tier for interruptible, checkpointed training and batch jobs; since March 31, 2026 these nodes are allocated as Spot VMs at variable Spot rates
  • Delete clusters that are no longer used, which also releases their quota and load balancer charges

Documentation

Vendor references for pricing and configuration.