A dataset has one column with 30% missing values. You believe other columns can help reconstruct the missing entries while preserving dataset integrity. Which imputation method should you use?
Choose an answer
Tap an option to check your answer.
Correct answer: Use multiple imputation to estimate and replace missing values..
Why this is the answer
Multiple imputation is the most robust method for handling missing data in this scenario because it estimates missing values multiple times, creating several complete datasets. This approach accounts for the uncertainty of the imputations, leading to more accurate statistical inferences and preserving the dataset's integrity by leveraging relationships with other columns. Deleting records (listwise deletion) would result in a significant loss of data (30%), potentially introducing bias and reducing statistical power. Carrying the last observed value forward is suitable for time-series data but not general tabular data, and it assumes the missing value is the same as the previous one, which is often not true. Replacing missing values with the column mean is a simple method but reduces variance, distorts relationships between variables, and underestimates standard errors, making it less ideal when other columns can inform the imputation.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed