A company is migrating on-premises Hadoop clusters to Amazon EMR and needs to move an existing Apache Hive metastore into a persistent, serverless solution for the data catalog. Which approach is the most cost-effective and meets the serverless requirement?
Choose an answer
Tap an option to check your answer.
Correct answer: Configure a Hive metastore in Amazon EMR. Migrate the existing on-premises Hive metastore into Amazon EMR. Use AWS Glue Data Catalog to store the company's data catalog as an external data catalog..
Why this is the answer
The correct option leverages AWS Glue Data Catalog, which is a serverless, persistent metadata repository compatible with Apache Hive metastores. Migrating the existing on-premises Hive metastore into Amazon EMR and then integrating it with AWS Glue Data Catalog provides a cost-effective and serverless solution for the data catalog. The incorrect options are: Migrating to Amazon S3 and scanning with Glue Data Catalog is not a direct migration of a Hive metastore; S3 is object storage, not a metastore. Using Amazon Aurora MySQL to store the data catalog would be a persistent solution but not serverless, as Aurora requires managing database instances. Configuring a new Hive metastore in Amazon EMR and using it directly as the company's data catalog would not be serverless, as EMR clusters are transient and require management.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed