A company wants to run ML analytics on data in an S3 data lake. They have two transformation needs: (1) perform scheduled daily transformations on 300 GB of various-format data that arrives at a set time, and (2) run a one-time transformation on terabytes of archived data in the lake. Amazon MWAA DAGs will orchestrate the workflows. Which two tasks should the DAGs schedule to meet these needs most cost-effectively? (Choose two.)
Choose an answer
Tap an option to check your answer.
Correct answer: For daily incoming data, use AWS Glue crawlers to scan and identify the schema., For daily and archived data, use Amazon EMR to perform data transformations..
Why this is the answer
AWS Glue crawlers are the most cost-effective solution for automatically scanning and inferring schemas from various data formats in S3, making them ideal for the scheduled daily transformations. Amazon EMR is a powerful and flexible big data processing service that can handle both the scheduled daily transformations on 300 GB and the one-time transformation on terabytes of archived data. Its ability to scale and process large datasets efficiently makes it cost-effective for these diverse needs. Amazon Athena is a query service, not designed for schema inference, making it unsuitable for scanning and identifying schemas. Amazon Redshift is a data warehouse, optimized for analytical queries on structured data, not for general-purpose data transformations on varied formats in a data lake. Amazon SageMaker is primarily for machine learning model training and deployment, not for large-scale data transformation tasks.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed