Explanation
Why the waste happens and who it affects.
Compute Savings Plans and EC2 Instance Savings Plans do not apply to SageMaker AI, so if the organization has not purchased a SageMaker AI Savings Plan, that entire baseline is billed at on-demand rates even when the rest of the account is well covered by commitments.
The gap often exists because commitment purchasing is run by a team focused on EC2, while SageMaker spend belongs to ML teams and grows quickly from a small base. SageMaker AI Savings Plans apply regardless of instance family, size, Region and component, so a plan bought while training dominates also covers inference if the workload shifts.
Billing model
The pricing dimensions that drive this cost.
SageMaker AI instance usage is billed at on-demand rates unless a SageMaker AI Savings Plan applies.
- On-demand ML instance usage
- Each eligible SageMaker instance is billed at the on-demand rate for its instance type and Region for the duration of use
- SageMaker AI Savings Plans
- A dollar-per-hour commitment for a 1- or 3-year term, up to 64% off on-demand, applied across instance family, size, Region and component
- Eligible usage
- ML instance usage in Studio notebooks, on-demand notebook instances, Processing, Data Wrangler, Training, Real-Time Inference and Batch Transform
- Hourly commitment
- Each hour's commitment is used only within that hour, and usage above it is charged at on-demand rates
How to detect
5 checks to find it in your estate.
- Review the Trusted Advisor check AWS Savings Plans purchase recommendations for Amazon SageMaker AI, sourced from Cost Optimization Hub, which flags accounts with an identified SageMaker AI savings action
- Open Cost Explorer Savings Plans recommendations with the SageMaker Savings Plans type and a 30- or 60-day lookback, and compare the recommended hourly commitment with current coverage
- Use the Savings Plans coverage report filtered to SageMaker to see what share of eligible SageMaker spend is on-demand
- Identify the steady part of the baseline, such as endpoints that have been in service continuously for the lookback period and training or processing jobs that run on a fixed schedule
- Before sizing a commitment, exclude idle endpoints, notebooks and Studio apps that should be shut down, and usage you plan to move to managed spot training or serverless inference
How to fix
5 ways to remove the waste.
- Remove idle SageMaker resources and right-size endpoints first, then purchase a SageMaker AI Savings Plan sized to the remaining minimum steady hourly spend rather than to peaks
- Use a 1-year term or No Upfront payment where the ML roadmap is uncertain, and a 3-year term for long-lived production inference with a stable baseline
- Rely on the plan's flexibility across instance family, size, Region and component when workloads move between training and inference or to newer instance types, rather than delaying the purchase
- Keep interruptible training on managed spot training, and cover only the on-demand baseline with the plan
- Review coverage and utilization monthly, and add further plans in increments as the steady baseline grows, since a plan's terms cannot be changed after purchase
Documentation
Vendor references for pricing and configuration.