A company performs quarterly analyses of data in a data lake to run inventory checks. A data engineer uses AWS Glue DataBrew to find any personally identifiable information (PII) about customers in the data. The company's privacy rules include some custom PII categories that are not part of DataBrew's built-in data quality rules. The engineer must update the process to scan for these custom PII categories across many datasets in the data lake. Which approach meets the requirement with the least operational overhead?
Choose an answer
Tap an option to check your answer.
Correct answer: Create custom data quality rules in DataBrew and apply those rules across the datasets..
Why this is the answer
Creating custom data quality rules in DataBrew and applying them across datasets is the most efficient solution. DataBrew is designed for data quality and cleansing, offering built-in PII detection and the flexibility to define custom rules. This approach leverages DataBrew's managed service capabilities, minimizing operational overhead compared to managing custom scripts or manual reviews. Manually reviewing datasets is impractical for large-scale data lakes and prone to human error. While custom Python scripts could work, integrating and managing them within DataBrew adds complexity and operational overhead that custom rules avoid. Using regular expressions during the ETL process is a valid technique for PII detection, but DataBrew's custom rules provide a more centralized and manageable solution for data quality checks, especially when dealing with evolving custom PII categories and multiple datasets.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed