Explanation
Why the waste happens and who it affects.
An endpoint is billed per inference unit per second from the moment it starts until it is deleted, even if no documents are analyzed, and it cannot be provisioned below one inference unit. Endpoints created for a proof of concept, a model evaluation or a one-off labeling task are often left running after the work ends.
The same waste appears when an endpoint is used for work that does not need real-time answers, such as nightly batch classification, while it sits idle the rest of the day, or when it is provisioned with far more inference units than its traffic consumes. AWS Trusted Advisor has a dedicated check, Amazon Comprehend underutilized endpoints, that flags endpoints not used for real-time inference for 15 consecutive days.
Billing model
The pricing dimensions that drive this cost.
- Endpoint inference units
- Billed per inference unit per second ($0.0005 per IU-second on the Comprehend pricing page), with one IU providing 100 characters per second of throughput
- Billing duration
- Charged in one-second increments with a 60-second minimum, from endpoint start until deletion, regardless of requests
- Asynchronous analysis jobs
- Billed by text processed in 100-character units with a 300-character minimum per request, and nothing when no jobs run
- Custom model management
- Custom models are billed a monthly management fee separately from any endpoint
How to detect
4 checks to find it in your estate.
- Review the Trusted Advisor cost optimization check Amazon Comprehend underutilized endpoints (check ID Cm24dfsM12), which flags active endpoints with no real-time inference requests in the past 15 days
- List endpoints with list-endpoints and compare the CloudWatch metrics ConsumedInferenceUnits and InferenceUtilization (EndpointArn dimension) with ProvisionedInferenceUnits over 15 to 30 days
- Flag endpoints where ConsumedInferenceUnits is zero, or where InferenceUtilization stays low relative to provisioned IUs outside short bursts
- Identify endpoints whose only callers are scheduled batch processes that could run as asynchronous analysis jobs instead
How to fix
4 ways to remove the waste.
- Delete endpoints that are no longer used; the custom model is kept and a new endpoint can be created later when real-time inference is needed
- Move batch and backlog processing to asynchronous classification or entity detection jobs, which bill only for the text processed
- For endpoints that are used intermittently, configure Application Auto Scaling with target tracking or scheduled scaling to reduce inference units outside busy periods, keeping in mind that an endpoint always has at least one IU
- Right-size DesiredInferenceUnits with update-endpoint based on observed ConsumedInferenceUnits, leaving headroom to avoid throttling
Documentation
Vendor references for pricing and configuration.