Design a highly available, resilient data-lake pipeline for millions of IoT sensors streaming structured and unstructured data. What should you do?
Choose an answer
Tap an option to check your answer.
Correct answer: Stream data to Pub/Sub, and use Dataflow to send data to Cloud Storage..
Why this is the answer
For a highly available, resilient data lake pipeline, Pub/Sub is ideal for ingesting millions of IoT sensor streams due to its scalability and durability. Dataflow, a fully managed service, can then process this data in real-time or batch, handling transformations and ensuring exactly-once processing. Cloud Storage is the best choice for a data lake because it's highly scalable, durable, cost-effective, and supports both structured and unstructured data, making it suitable for raw IoT data. Using Storage Transfer Service to send data to BigQuery is incorrect because Storage Transfer Service is for moving data between storage locations, not for real-time streaming from Pub/Sub, and BigQuery is a data warehouse, not a data lake for raw, varied data. Streaming directly to Dataflow isn't the primary ingestion point; Pub/Sub provides the necessary buffering and decoupling. Dataprep is for data preparation, not the primary storage for a data lake.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed