An ML model will score customers in an online application to determine which product to show. The engineer needs to minimize response latency. How should the model be deployed in SageMaker to meet low-latency requirements?
Choose an answer
Tap an option to check your answer.
Correct answer: Configure a real-time inference endpoint..
Why this is the answer
A real-time inference endpoint is the correct choice because it provides immediate predictions for individual requests, which is essential for low-latency requirements in online applications. Batch transform is for processing large datasets offline, not for real-time scoring. Serverless inference endpoints offer automatic scaling and cost-effectiveness but can introduce cold start latencies that might not meet strict low-latency demands. Asynchronous inference endpoints are designed for large payloads or long-running inferences where immediate responses aren't critical, making them unsuitable for this scenario.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed