Explanation
Why the waste happens and who it affects.
From that moment, every node in the instance is billed per vCore-hour, prorated by the minute, until the pool pauses. Microsoft states that billing runs from pool start until the idle timeout, whether or not the nodes are doing work.
When automatic pause is disabled, or set to a long idle timeout, nodes keep running after notebooks and pipeline jobs finish. Interactive development makes this worse: Synapse Studio sends keep-alive messages to hold sessions open, so a forgotten notebook can keep a pool alive for hours. Fixed-size pools without autoscale add further idle capacity during light stages of a job. Azure Advisor flags both conditions with 'Consider enabling automatic pause feature on spark compute' and 'Consider enabling autoscale feature on spark compute'.
Billing model
The pricing dimensions that drive this cost.
Spark pool charges depend only on how long instances run and how many nodes they use.
- Pool definition
- Creating a Spark pool is free; charges start only when a Spark instance is instantiated for a job or session
- vCore-hours
- Running instances are charged per vCore-hour for every node, prorated by the minute
- Idle timeout
- Billing continues after the last activity until the automatic pause delay expires
- Autoscale
- No extra charge; scaling up or down changes node count and increases pool runtime while scaling
How to detect
5 checks to find it in your estate.
- Review Azure Advisor Cost recommendations for 'Consider enabling automatic pause feature on spark compute' and 'Consider enabling autoscale feature on spark compute' on Synapse workspaces
- List Spark pools with az synapse spark pool list and flag pools where autoPause.enabled is false or autoPause.delayInMinutes is long relative to how the pool is used
- Flag pools where autoScale.enabled is false and nodeCount is well above the minimum of three nodes
- Compare the bigDataPools metrics vCores allocated (BigDataPoolAllocatedCores) and Active Apache Spark applications (BigDataPoolApplicationsActive) to find periods where cores stay allocated with no active applications
- Review Apache Spark pool vCore-hour cost in Cost Management by pool to find pools whose runtime far exceeds scheduled job durations
How to fix
5 ways to remove the waste.
- Enable automatic pause on every Spark pool with a short idle delay (az synapse spark pool update --enable-auto-pause true --delay <minutes>, or the pool's Additional settings in the portal); active sessions must be restarted for the change to apply
- Enable autoscale with a minimum and maximum node count so light workloads run on fewer nodes; the minimum cannot be below three nodes
- Ask developers to stop notebook sessions when they finish and to shorten the Synapse Studio session timeout, since keep-alive messages hold sessions open
- Create separate small pool definitions for development and validation and reserve larger node sizes for performance testing and production, since pool definitions cost nothing
- When updating an existing pool, note that forcing new settings terminates all running Spark sessions; schedule the change outside active job windows
Documentation
Vendor references for pricing and configuration.
- Plan to manage costs for Azure Synapse Analyticslearn.microsoft.com
- Apache Spark pool conceptslearn.microsoft.com
- Automatically scale Apache Spark instanceslearn.microsoft.com
- Cost recommendations - Azure Advisorlearn.microsoft.com
- az synapse spark poollearn.microsoft.com
- Pricing - Azure Synapse Analyticsazure.microsoft.com