You receive monthly CSVs from a third party; the schema changes every third month. Requirements: scheduled execution, non-developer analysts can modify transformations, and a graphical transformation designer. Which solution fits these requirements?
Choose an answer
Tap an option to check your answer.
Correct answer: Use Dataprep by Trifacta to build and maintain transformation recipes and schedule them.
Why this is the answer
Dataprep by Trifacta is the best fit because it directly addresses all requirements. It offers a graphical interface for building transformation recipes, enabling non-developer analysts to modify them easily. It supports scheduled execution for monthly CSVs and is designed to handle schema drift, which is crucial given the schema changes every third month. Loading into BigQuery and using SQL is less ideal because while SQL can normalize schemas, it lacks the graphical transformation designer and ease of modification for non-developers that Dataprep provides. Dataflow with Python requires developer expertise for pipeline creation and modification, failing the "non-developer analysts" requirement. Apache Spark on Dataproc also requires coding knowledge (Spark SQL) and a more technical skill set than what non-developer analysts typically possess, making it unsuitable for their direct modification.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed