A company will deploy a model for production inference on a SageMaker endpoint. Typical request payloads are between 100 MB and 300 MB, and each request must complete within 60 minutes. Which SageMaker inference option should be used?
Choose an answer
Tap an option to check your answer.
Correct answer: Asynchronous inference.
Why this is the answer
Asynchronous inference is the correct choice because it is designed for large payload sizes (up to 1 GB) and long processing times (up to 1 hour), aligning perfectly with the given requirements of 100-300 MB payloads and 60-minute completion times. Serverless inference has payload limits of 4 MB and timeout limits of 300 seconds, making it unsuitable. Real-time inference has even stricter limits (5 MB payload, 60-second timeout). Batch transform is for offline processing of entire datasets, not for individual, on-demand requests to an endpoint.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed