You expect a large increase in Spark and Hadoop jobs from an on-prem datacenter and want cloud scaling with minimal operational work and code changes. Which product should you use?
Choose an answer
Tap an option to check your answer.
Correct answer: Google Cloud Dataproc.
Why this is the answer
Google Cloud Dataproc is the best choice because it's a fully managed service for Apache Spark and Hadoop clusters. It offers automatic scaling and integrates well with existing on-premises Spark/Hadoop workloads, minimizing code changes and operational overhead. Google Cloud Dataflow is a fully managed service for stream and batch data processing, but it uses Apache Beam, which would require significant code changes from existing Spark/Hadoop jobs. Google Compute Engine provides raw virtual machines, requiring you to manually set up, configure, and manage Spark/Hadoop, which increases operational work. Google Kubernetes Engine orchestrates containers, but you would still need to manage the Spark/Hadoop deployments and scaling yourself, adding complexity.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed