An ML engineer must build a pipeline that discovers and removes PII from petabytes of unstructured data, and then use the cleaned data to train SageMaker models. Which solution meets this requirement at scale?
Choose an answer
Tap an option to check your answer.
Correct answer: Use the Apache Spark–based serverless engine in AWS Glue interactive sessions and apply the Detect PII transform to find and remove PII..
Why this is the answer
The correct solution leverages AWS Glue interactive sessions with its Apache Spark-based serverless engine and the built-in Detect PII transform. This approach is designed for processing petabytes of unstructured data efficiently and at scale, making it ideal for PII discovery and removal before model training. AWS Glue's serverless nature eliminates infrastructure management, and Spark provides distributed processing capabilities. AWS Glue Data Wrangler in SageMaker Canvas is for data preparation and feature engineering, but it's not optimized for petabyte-scale PII detection and removal from unstructured data. Amazon SageMaker Clarify focuses on bias detection and explainability, not PII identification and masking. Amazon Comprehend's DetectEntities API can identify entities, including PII, but it's primarily for natural language processing tasks and might not be the most cost-effective or scalable solution for direct PII removal from petabytes of unstructured data within a data pipeline context compared to AWS Glue's dedicated transform.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed