An airline runs a SageMaker real-time endpoint to adjust ticket prices based on demand. Previous deployments failed to scale quickly enough when website traffic rose. The ML engineer must configure target-tracking auto scaling to respond rapidly to sudden traffic spikes. Which configuration will be most responsive?
Choose an answer
Tap an option to check your answer.
Correct answer: Use the SageMaker InvocationsPerInstance metric with high-resolution 10-second intervals and keep the default 300-second scale-in cooldown..
Why this is the answer
To respond rapidly to sudden traffic spikes, the most responsive configuration uses high-resolution metrics and a shorter scale-in cooldown. High-resolution 10-second intervals for the InvocationsPerInstance metric allow Auto Scaling to detect changes in demand much faster than standard 1-minute intervals. A shorter scale-in cooldown (like the default 300 seconds) means that once the spike subsides, instances can be removed more quickly, optimizing cost. A longer 600-second scale-in cooldown would delay the removal of unneeded instances, making it less responsive to cost optimization after a spike. Standard metrics with 1-minute resolution are less responsive to sudden spikes.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed