An ML engineer is building a fraud detection model. Data sources include transaction logs and customer profiles in S3 plus tables from an on-premises MySQL database. Which AWS service or capability can centrally aggregate these varied data sources for ML use?
Choose an answer
Tap an option to check your answer.
Correct answer: Use AWS Lake Formation to centralize and manage aggregated data across sources..
Why this is the answer
AWS Lake Formation is the correct choice because it is designed to build, secure, and manage data lakes, which are ideal for aggregating diverse data sources like S3 objects and on-premises databases. It simplifies the process of collecting, cleaning, and cataloging data, making it readily available for ML. Running Amazon EMR Spark jobs could process and aggregate data, but EMR itself isn't a central management layer for the data lake; it's a processing engine. You'd still need a service like Lake Formation to catalog and secure the aggregated data. Amazon Kinesis Data Streams is for real-time data ingestion and processing, not for centralizing and managing historical, varied data sources for a data lake. Amazon DynamoDB is a NoSQL database, suitable for specific types of application data, but not for building a comprehensive data lake that aggregates diverse formats from multiple sources for ML.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed