A dataset repository on S3 contains personally identifiable information that must be redacted before model training. Which option minimizes development effort to remove PII from the data?
Choose an answer
Tap an option to check your answer.
Correct answer: Use AWS Glue DataBrew to detect and redact PII from the dataset..
Why this is the answer
AWS Glue DataBrew offers built-in transformations specifically designed for PII detection and redaction, making it the most efficient and low-effort solution for this task. It provides a visual interface and pre-built functions that simplify the process without requiring extensive coding. Amazon SageMaker Data Wrangler also offers data preparation capabilities, but its strength lies more in feature engineering and data transformation for machine learning workflows. While custom transformations are possible, DataBrew's dedicated PII features are more direct for this specific use case. Creating a custom AWS Lambda function would involve significant development effort to write, test, and maintain the PII detection and redaction logic from scratch. Using an AWS Glue development endpoint with custom code in a notebook offers flexibility but requires writing and managing the PII redaction logic manually, which is more development effort than DataBrew's pre-built functionality.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed