A company plans to run Apache Spark jobs on a provisioned Amazon EMR cluster for big data analysis and requires high reliability. The big data team wants cost-optimized, long-running workloads while preserving current performance. Which combination of choices is the MOST cost-effective? (Choose two.)
Choose an answer
Tap an option to check your answer.
Correct answer: Use Amazon S3 as a persistent data store., Use Graviton instances for core nodes and task nodes..
Why this is the answer
Using Amazon S3 as a persistent data store is cost-effective because S3 offers highly durable, scalable, and inexpensive storage, decoupling storage from compute. This allows EMR clusters to be scaled up or down independently, or even terminated, without losing data, which is ideal for long-running workloads. HDFS, while integrated with EMR, ties storage to compute, making it less flexible and potentially more expensive for persistent storage as you pay for EC2 instances even when not actively processing. Using Graviton instances for core and task nodes is cost-effective because AWS Graviton processors deliver up to 40% better price-performance over comparable x86-based instances for many workloads, including Apache Spark. This directly addresses the requirement for cost-optimized performance. Using Spot Instances for primary nodes is risky for high reliability, as primary nodes manage the cluster and losing them can disrupt the entire workload.
Pass your exam — without the endless answer hunt
Get every verified question and explanation for this exam in one place, and save hours of prep. 1,000+ certifications · 20+ languages · free to start.
Pass your exam faster → No card needed