Skip to content
Cloud Efficiency Hub

Idle Dedicated Nodes in Azure Batch Pools Without Autoscaling

The short version

Azure Batch pools are often created with a fixed targetDedicatedNodes value so that jobs can start immediately.

PointFive Research

Cloud cost research at PointFive

Azure service
Azure Batch
Category
Compute
Reference
CER-0410
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

Between jobs, and overnight or on weekends, those nodes sit idle but remain allocated VMs with managed OS disks, so they keep billing at the full VM rate. Microsoft's own pool lifetime guidance notes that pre-created pools start tasks faster but that nodes might sit idle waiting for work.

Idle nodes are not the only leak. Nodes can also end up in states where they cannot run tasks at all, such as unusable or starttaskfailed, and Microsoft states that allocated nodes normally still incur costs. Pools that run one task per node on large VMs, or that spread tasks across many partly used nodes, also need more nodes than the work requires, especially when a pool is treated as a standing cluster rather than elastic capacity.

Billing model

The pricing dimensions that drive this cost.

Azure Batch itself is free; the cost comes from the resources that pools allocate.

Pool node VM
Each allocated dedicated node is billed as an Azure VM for as long as it is in the pool, busy or idle
Managed OS disk
Each Virtual Machine Configuration node has a managed OS disk billed alongside the VM unless ephemeral OS disks are used
Spot node
Uses surplus capacity at a lower hourly price than dedicated nodes but can be preempted
Autoscale evaluation
Batch re-evaluates the pool's autoscale formula every 15 minutes by default (minimum five minutes)

How to detect

5 checks to find it in your estate.

  • In Batch metrics, compare Idle Node Count with Running Node Count per pool over several weeks; pools with sustained idle nodes are candidates
  • List pools where enableAutoScale is false and targetDedicatedNodes is fixed above zero
  • Monitor node state counts for unusable and starttaskfailed nodes and investigate their cause
  • Check each pool's taskSlotsPerNode and taskSchedulingPolicy; one slot per node with the spread fill type on multi-core VMs often means low node utilization
  • Use Cost analysis with a Resource filter on the Batch account and pool name (or the poolname tag in user subscription mode) to size the spend per pool

How to fix

5 ways to remove the waste.

  • Enable automatic scaling with a formula driven by $PendingTasks (or $ActiveTasks) that sets $TargetDedicatedNodes to zero when no tasks are queued, and use $NodeDeallocationOption = taskcompletion so running tasks finish before nodes are removed
  • For job-shaped workloads, use autopools or create a pool per job and delete it when the job completes, accepting the node allocation delay at job start
  • Move preemption-tolerant work to Spot nodes by setting $TargetLowPriorityNodes, keeping dedicated nodes only for work that cannot be interrupted
  • Increase taskSlotsPerNode and set taskSchedulingPolicy to pack so tasks fill nodes before new ones are used and empty nodes can be scaled away
  • Fix or remove nodes stuck in unusable or starttaskfailed states, and use ephemeral OS disks where the VM size supports them. Node VM size cannot be changed after pool creation, so rightsizing means creating a new pool

Documentation

Vendor references for pricing and configuration.