A company streams CSV records in real time and wants to store them on Amazon S3 in Apache Parquet format. Which option requires the least development effort to convert incoming CSV data to Parquet before storing in S3?
Choose an answer
Tap an option to check your answer.
Correct answer: Use Amazon Kinesis Data Streams and Amazon Kinesis Data Firehose to transform incoming CSV into Parquet..
Why this is the answer
The correct answer is to use Amazon Kinesis Data Streams and Amazon Kinesis Data Firehose to transform incoming CSV into Parquet. Kinesis Data Firehose natively supports converting incoming data formats (like CSV) to Apache Parquet or ORC before delivering to S3, requiring minimal development effort. You configure the input and output formats directly in Firehose. Using Apache Kafka Streams on EC2 or Apache Spark Structured Streaming on EMR would require significant development effort to set up and manage the Kafka/Spark clusters, write the transformation logic, and handle S3 integration. While AWS Glue can convert data, using it with Kinesis Data Streams would typically involve setting up a Glue job to read from Kinesis and write to S3, which is more complex than Firehose's direct integration.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed