A company uses Amazon DataZone for data governance and a business catalog. Their data resides in an Amazon S3 data lake and they use AWS Glue with the AWS Glue Data Catalog. A data engineer must make AWS Glue Data Quality scores available in the Amazon DataZone portal. Which approach satisfies this requirement?
Choose an answer
Tap an option to check your answer.
Correct answer: Create a data quality ruleset with Data Quality Definition language (DQDL) rules that apply to a specific AWS Glue table. Schedule the ruleset to run daily. Configure the Amazon DataZone project to have an AWS Glue data source. Enable the data quality configuration for the data source..
Why this is the answer
The correct approach involves creating a data quality ruleset using DQDL for the specific AWS Glue table and scheduling it to run daily. This ensures regular data quality evaluation. Then, configuring the Amazon DataZone project to use an AWS Glue data source and enabling its data quality configuration allows DataZone to ingest and display these scores. AWS Glue Data Quality is designed to work directly with the AWS Glue Data Catalog, which is where the company's data lake metadata resides. Incorrect options: Using an Amazon Redshift data source is incorrect because the data resides in an S3 data lake and uses AWS Glue for its catalog, not Redshift. Defining data quality rulesets inside AWS Glue ETL jobs is less efficient for centralized data quality monitoring within DataZone compared to using standalone DQDL rulesets associated with tables. While ETL jobs can perform quality checks, DataZone integrates directly with Glue Data Quality rulesets for reporting.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed