Skip to content
Cloud Efficiency Hub

Idle SageMaker Real-Time Inference Endpoint

The short version

A SageMaker real-time inference endpoint keeps its instances running and billed for as long as the endpoint is in service, whether or not any application calls it.

PointFive Research

Cloud cost research at PointFive

AWS service
AWS SageMaker
Category
AI
Reference
CER-0559
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

Endpoints are created for model evaluations, demos, A/B tests, proofs of concept and previous model versions, and they are often left in place after the traffic moves to a new endpoint or the project ends. Because models are frequently hosted on GPU or accelerator instances, and production variants may run several instances for availability, a single forgotten endpoint can be expensive.

The waste is easy to overlook because an idle endpoint reports no errors and nothing breaks. It is common in teams that deploy endpoints from notebooks or pipelines without a teardown step, and in accounts where each model version gets its own endpoint. AWS Compute Optimizer includes SageMaker endpoints in its idle resource recommendations, using zero invocations over the lookback period as the idle signal.

Billing model

The pricing dimensions that drive this cost.

Real-time inference is billed for the instances behind the endpoint, not for the requests it serves.

Instance usage
Each instance behind a production variant is charged for the instance type chosen for as long as the endpoint is running
Data processed
Data processed in and out of the endpoint is charged separately per GB
Scale to zero
Endpoints that host inference components can scale in to zero instances, removing instance charges while there is no traffic
Serverless inference
On-demand serverless inference is billed only for the compute used to process requests, by the millisecond, plus data processed, so idle time is not charged (Provisioned Concurrency adds a charge for keeping capacity warm)

How to detect

5 checks to find it in your estate.

  • Review AWS Compute Optimizer idle recommendations for Amazon SageMaker endpoints, which flag endpoints with zero invocations over the 14-day lookback period
  • In CloudWatch (AWS/SageMaker namespace), check the Sum of Invocations per EndpointName and VariantName over at least 14 to 30 days; no datapoints or a sum of zero means the endpoint served no requests
  • Check instance metrics in /aws/sagemaker/Endpoints (CPUUtilization, GPUUtilization, MemoryUtilization) to confirm the instances are doing no work, and note the instance type and count to size the waste
  • List endpoints with list-endpoints and compare creation time and last update with active model versions and application configuration to find superseded or experimental endpoints
  • Confirm with the model owner that no scheduled or low-frequency caller, such as a monthly batch or a failover path, depends on the endpoint

How to fix

5 ways to remove the waste.

  • Delete endpoints that are no longer used with delete-endpoint; the endpoint configuration and model can be kept so the endpoint can be recreated later if needed
  • For endpoints with intermittent traffic, move the model to inference components and allow scale in to zero by setting MinInstanceCount to 0 in ManagedInstanceScaling and registering each inference component with a minimum capacity of 0; add a step scaling policy triggered by the NoCapacityInvocationFailures metric, and accept that scaling out from zero takes several minutes during which requests fail
  • Move spiky workloads that can tolerate variable p99 latency to serverless inference, and latency-insensitive workloads to asynchronous inference, which can scale down to zero
  • Consolidate several lightly used endpoints onto a multi-model or multi-container endpoint so they share instances
  • Add an owner and expiry tag to endpoints created for experiments, and include endpoint deletion in the teardown step of notebooks and pipelines

Documentation

Vendor references for pricing and configuration.