A company has a large unstructured dataset containing many duplicate records across key attributes. Which AWS solution will identify duplicates with the least amount of custom coding?
Choose an answer
Tap an option to check your answer.
Correct answer: Use the AWS Glue FindMatches transform to identify duplicate records..
Why this is the answer
The AWS Glue FindMatches transform is specifically designed for identifying duplicate or matching records in datasets, even when exact matches don't exist (fuzzy matching). It uses machine learning to learn patterns for matching records, requiring minimal custom coding. Amazon Mechanical Turk is a crowdsourcing service, which, while capable of identifying duplicates, involves significant human effort and management, making it less automated than the FindMatches transform. Amazon QuickSight ML Insights focuses on anomaly detection and forecasting within structured data for business intelligence, not general-purpose data deduplication. Amazon SageMaker Data Wrangler is a data preparation tool that can be used for preprocessing, but it doesn't offer a built-in, low-code solution for fuzzy duplicate detection like AWS Glue FindMatches. While you could potentially build a custom deduplication model in SageMaker, it would require significantly more custom coding than using FindMatches.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed