A bank has 10 years of customer data in CSV files on-premises. The data science team wants to quickly perform transformations and explore the data before building a production model in SageMaker, with minimal development effort. Which approach minimizes effort and enables fast transformation and insights?
Choose an answer
Tap an option to check your answer.
Correct answer: Upload the CSV files to an Amazon S3 bucket, grant SageMaker access, then import the data into SageMaker Data Wrangler to transform and explore..
Why this is the answer
The correct approach involves uploading CSV files to an Amazon S3 bucket, granting SageMaker access, and then importing the data into SageMaker Data Wrangler for transformation and exploration. SageMaker Data Wrangler is designed for data preparation and feature engineering, offering a visual interface to quickly transform and analyze data. S3 is the standard, scalable, and secure storage service for data in AWS, and SageMaker needs access to this data. The first incorrect option is wrong because Data Wrangler does not support direct upload of files from on-premises; data must first be in an accessible AWS service like S3. The third option is overly complex and introduces QuickSight unnecessarily for initial exploration when Data Wrangler itself provides robust exploration capabilities. The fourth option is also overly complex; while using a SageMaker Studio notebook for insights is valid, Data Wrangler's built-in analysis tools can provide quick insights without needing to move to a separate notebook for initial exploration. The core requirement is minimal effort and fast transformation/exploration, which the correct answer best addresses.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed