A data engineer must maintain a central metadata repository that is accessible from Amazon EMR and Amazon Athena queries and must import existing Apache Hive metadata into it with minimal development effort. Which solution should they use?
Choose an answer
Tap an option to check your answer.
Correct answer: Use the AWS Glue Data Catalog..
Why this is the answer
The AWS Glue Data Catalog is the correct choice because it provides a fully managed, central metadata repository compatible with Apache Hive metastores. It seamlessly integrates with Amazon EMR and Amazon Athena, allowing both services to query the same metadata. Importing existing Apache Hive metadata into the AWS Glue Data Catalog is straightforward, often requiring minimal effort due to its Hive Metastore compatibility. Using Amazon EMR and Apache Ranger is incorrect because Apache Ranger is primarily for authorization and access control, not a central metadata repository. A Hive metastore on an EMR cluster is not central or fully managed; it's tied to a specific cluster and requires manual management. A metastore on an Amazon RDS for MySQL DB instance would require significant development effort for integration and management with EMR and Athena, and it wouldn't be a fully managed, purpose-built metadata service like the AWS Glue Data Catalog.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed