A company is building a data lake on Amazon S3 and must enforce row-level and column-level access controls for teams that will query via Amazon Athena, Redshift Spectrum, and Apache Hive on EMR. Which approach minimizes operational overhead?
Choose an answer
Tap an option to check your answer.
Correct answer: Use Amazon S3 for data lake storage. Use AWS Lake Formation to restrict data access by rows and columns. Provide data access through AWS Lake Formation..
Why this is the answer
The correct answer is to use Amazon S3 for data lake storage and AWS Lake Formation to restrict data access by rows and columns, providing access through Lake Formation. AWS Lake Formation is specifically designed to simplify building, securing, and managing data lakes on S3. It offers fine-grained access control (row and column-level) that integrates natively with Athena, Redshift Spectrum, and Apache Hive on EMR, minimizing operational overhead compared to managing separate security policies for each service. Using S3 access policies alone is insufficient for row and column-level control. Apache Ranger on EMR is specific to EMR and doesn't cover Athena or Redshift Spectrum natively for a unified approach. Redshift is a data warehouse, not ideal for raw data lake storage, and its security policies wouldn't apply directly to Athena or EMR querying S3.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed