A CPU-based model will serve real-time predictions with intermittent traffic during business hours and idle periods after hours. Which SageMaker hosting option is the most cost-effective for this pattern?
Choose an answer
Tap an option to check your answer.
Correct answer: Deploy the model to a SageMaker Serverless Inference endpoint. Configure increased provisioned concurrency during business hours..
Why this is the answer
SageMaker Serverless Inference is the most cost-effective for intermittent traffic because it automatically scales compute resources based on demand, and you only pay for the inference duration and data processed. This eliminates the cost of idle resources during off-peak hours. Configuring increased provisioned concurrency during business hours ensures low latency during peak times. Deploying to a SageMaker real-time endpoint with a schedule-based auto scaling policy is less cost-effective because even with scaling, there's a minimum instance count that incurs cost during idle periods. Asynchronous Inference is for large payloads or long processing times, not real-time, low-latency predictions. Activating a real-time endpoint with Lambda is overly complex and still incurs costs for the minimum instance count when active.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed