A custom forecasting model must be hosted; the workload is predictable and sustained during the same 2-hour window each day with many invocations needing fast responses. The company wants AWS to manage the infrastructure and autoscaling. Which hosting option meets these requirements?
Choose an answer
Tap an option to check your answer.
Correct answer: Use Amazon SageMaker Serverless Inference and configure provisioned concurrency..
Why this is the answer
Amazon SageMaker Serverless Inference with provisioned concurrency is ideal for predictable, sustained workloads requiring fast responses and managed infrastructure. Serverless Inference handles infrastructure management and autoscaling. Provisioned concurrency ensures that a pre-warmed set of invocations are ready, eliminating cold starts and guaranteeing low latency during the 2-hour peak window. Scheduling a SageMaker batch transform job is incorrect because batch transform is for offline processing of large datasets, not real-time, low-latency inference. An EC2 Auto Scaling group with scheduled scaling would provide managed infrastructure and autoscaling, but it requires more operational overhead compared to SageMaker Serverless Inference and might not achieve the same level of rapid scaling for bursty, low-latency needs without careful tuning. Running on Amazon EKS on EC2 with pod autoscaling offers flexibility but involves significant operational overhead for cluster management, which contradicts the requirement for AWS to manage the infrastructure.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed