Before training, the dataset’s class imbalance must be resolved with minimal operational effort. Which option best accomplishes this?
Choose an answer
Tap an option to check your answer.
Correct answer: Use the Amazon SageMaker Data Wrangler balance data operation to oversample the minority class..
Why this is the answer
The Amazon SageMaker Data Wrangler balance data operation directly addresses class imbalance by allowing you to oversample the minority class or undersample the majority class with minimal operational effort. This is a built-in feature designed for data preparation within the SageMaker ecosystem. Using Amazon Athena to discover patterns and then manually adjusting the dataset is a cumbersome and inefficient approach, requiring significant manual intervention. SageMaker Studio Classic built-in algorithms might handle imbalanced datasets during training, but Data Wrangler offers a dedicated, pre-processing step for data balancing. AWS Glue DataBrew can perform data transformations, but SageMaker Data Wrangler is specifically optimized for ML data preparation workflows and offers the 'balance data' operation as a direct solution for this problem.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed