Skip to content
Cloud Efficiency Hub

Underutilized Comprehend Custom Endpoint

The short version

Amazon Comprehend custom classification and custom entity recognition models can serve real-time requests through an endpoint provisioned with a number of inference units.

PointFive Research

Cloud cost research at PointFive

AWS service
AWS Comprehend
Category
AI
Reference
CER-0385
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

An endpoint is billed per inference unit per second from the moment it starts until it is deleted, even if no documents are analyzed, and it cannot be provisioned below one inference unit. Endpoints created for a proof of concept, a model evaluation or a one-off labeling task are often left running after the work ends.

The same waste appears when an endpoint is used for work that does not need real-time answers, such as nightly batch classification, while it sits idle the rest of the day, or when it is provisioned with far more inference units than its traffic consumes. AWS Trusted Advisor has a dedicated check, Amazon Comprehend underutilized endpoints, that flags endpoints not used for real-time inference for 15 consecutive days.

Billing model

The pricing dimensions that drive this cost.

Endpoint inference units
Billed per inference unit per second ($0.0005 per IU-second on the Comprehend pricing page), with one IU providing 100 characters per second of throughput
Billing duration
Charged in one-second increments with a 60-second minimum, from endpoint start until deletion, regardless of requests
Asynchronous analysis jobs
Billed by text processed in 100-character units with a 300-character minimum per request, and nothing when no jobs run
Custom model management
Custom models are billed a monthly management fee separately from any endpoint

How to detect

4 checks to find it in your estate.

  • Review the Trusted Advisor cost optimization check Amazon Comprehend underutilized endpoints (check ID Cm24dfsM12), which flags active endpoints with no real-time inference requests in the past 15 days
  • List endpoints with list-endpoints and compare the CloudWatch metrics ConsumedInferenceUnits and InferenceUtilization (EndpointArn dimension) with ProvisionedInferenceUnits over 15 to 30 days
  • Flag endpoints where ConsumedInferenceUnits is zero, or where InferenceUtilization stays low relative to provisioned IUs outside short bursts
  • Identify endpoints whose only callers are scheduled batch processes that could run as asynchronous analysis jobs instead

How to fix

4 ways to remove the waste.

  • Delete endpoints that are no longer used; the custom model is kept and a new endpoint can be created later when real-time inference is needed
  • Move batch and backlog processing to asynchronous classification or entity detection jobs, which bill only for the text processed
  • For endpoints that are used intermittently, configure Application Auto Scaling with target tracking or scheduled scaling to reduce inference units outside busy periods, keeping in mind that an endpoint always has at least one IU
  • Right-size DesiredInferenceUnits with update-endpoint based on observed ConsumedInferenceUnits, leaving headroom to avoid throttling

Documentation

Vendor references for pricing and configuration.