Explanation
Why the waste happens and who it affects.
Many training workloads tolerate interruption: experiments, scheduled retraining, tuning sweeps and any job that can resume from a checkpoint. Running all of them on on-demand capacity means paying the full rate for work that has no hard start or finish time.
Managed spot training runs the job on Amazon EC2 Spot capacity and handles interruptions, and AWS states it can optimize training cost by up to 90% over on-demand instances. It is often left off because job templates default to on-demand or because checkpointing was never added to custom training scripts. The Well-Architected Machine Learning Lens lists using exclusively on-demand instances for training jobs as an anti-pattern under MLCOST04-BP06.
Billing model
The pricing dimensions that drive this cost.
Training jobs are billed for the instances they run on, and managed spot changes both the rate and the billable time.
- On-demand training
- Charged for the chosen instance type based on the duration of use, billed per second, at the on-demand rate
- Managed spot training
- Runs on EC2 Spot capacity; the job's BillableTimeInSeconds is what is billed, and savings equal (1 - BillableTimeInSeconds / TrainingTimeInSeconds) x 100
- Wait time
- MaxWaitTimeInSeconds sets how long SageMaker waits for spot capacity and must be larger than MaxRuntimeInSeconds
- Non-checkpointing algorithms
- Built-in and Marketplace algorithms that do not checkpoint are limited to a MaxWaitTimeInSeconds of 3600 seconds
How to detect
5 checks to find it in your estate.
- List recent training jobs (list-training-jobs, then describe-training-job) and find those where EnableManagedSpotTraining is false, prioritizing long-running, repeated or scheduled jobs and hyperparameter tuning jobs
- For jobs that do use spot, compare BillableTimeInSeconds with TrainingTimeInSeconds to confirm the expected savings are being realized
- Check whether training jobs already set CheckpointConfig (an S3Uri with the default local path /opt/ml/checkpoints) or use frameworks and built-in algorithms that checkpoint without script changes; these are the easiest to move to spot
- Review SageMaker training spend in Cost Explorer by instance type to find the largest on-demand training workloads, especially GPU instances
- Identify jobs with hard deadlines or very short runtimes where spot waiting time would not be acceptable, and exclude them from the candidate list
How to fix
5 ways to remove the waste.
- Set EnableManagedSpotTraining to true on eligible training and tuning jobs and set MaxWaitTimeInSeconds larger than MaxRuntimeInSeconds to give SageMaker time to obtain spot capacity and resume after interruptions
- Add checkpointing for any job that is not short: write checkpoints to /opt/ml/checkpoints and set CheckpointConfig so SageMaker syncs them to S3 and resumes from the last checkpoint instead of restarting
- For distributed training, configure distinct checkpoint file names or paths per instance in the training script, because the SageMaker checkpoint configuration uses a single S3 location
- Keep on-demand capacity for jobs with strict completion deadlines, and cover the remaining steady on-demand training baseline with SageMaker AI Savings Plans if it is large and predictable
- Make managed spot the default in shared training templates and pipelines so new jobs opt out deliberately rather than opt in
Documentation
Vendor references for pricing and configuration.