A national census collects roughly 500 responses per person. Which pair of algorithms would be appropriate to gain insights from this high-dimensional survey data? (Choose two.)
Choose an answer
Tap an option to check your answer.
Correct answer: Principal Component Analysis (PCA) algorithm, k-means clustering algorithm.
Why this is the answer
PCA is suitable for high-dimensional data because it reduces dimensionality while retaining most of the variance, making the data more manageable for subsequent analysis. This helps to identify underlying patterns in the 500 responses per person. K-means clustering is appropriate for grouping similar respondents based on their survey answers, which is a common goal in census data analysis. It can identify distinct segments within the population. Latent Dirichlet Allocation (LDA) is primarily for topic modeling in text data, not numerical survey responses. Factorization Machines (FM) are used for recommender systems and predicting interactions between features, which isn't the primary goal here. Random Cut Forest (RCF) is an anomaly detection algorithm, not for general insights or clustering.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed