An ML engineer is training a binary health-risk model (positive = likely at risk, negative = unlikely) with ages 30–60 included as a feature. The difference in proportions of labels (DPL) for the positive class is +0.9 for the 40–45 age group versus other ages, indicating overrepresentation. How should the engineer correct this imbalance?
Choose an answer
Tap an option to check your answer.
Correct answer: Undersample the positive examples in the 40–45 age group..
Why this is the answer
The problem states that the difference in proportions of labels (DPL) for the positive class is +0.9 for the 40–45 age group, indicating an overrepresentation of positive examples in this specific group. To correct an overrepresentation, you should reduce the number of instances of the overrepresented class. Therefore, undersampling the positive examples in the 40–45 age group is the correct approach. Oversampling would exacerbate the existing imbalance. Undersampling other groups or oversampling the negative class for other groups would not directly address the overrepresentation within the 40-45 age group.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed