Amazon DOP-C02: High Availability, Resilience and Disaster Recovery — Study Guide

Part of the AWS DevOps Engineer Professional DOP-C02 — Study Guide. Practice with verified answers in the Amazon exam hub, or take timed practice tests on ExamRoll.io.

Overview

High availability and disaster recovery on AWS focus on reducing downtime (RTO) and data loss (RPO) under component, Availability Zone (AZ), or Regional failures. Multi-AZ designs absorb AZ failure without data loss and minimal service impact; multi-Region designs address Regional disruptions and large-scale events. Selecting between active/active, active/passive (warm standby), and pilot-light strategies is driven by business RTO/RPO targets, consistency requirements, and cost. Achieving these objectives requires coherent design across DNS routing, compute elasticity, load balancing, database replication/failover, durable object storage with replication/versioning, centralized backup, and continuous resilience verification via fault injection.

Architectures for RTO/RPO and intelligent routing

Multi-AZ and multi-Region:

Route 53 routing policies and health checks:

Resilient patterns by objective:

Elastic Load Balancing and Auto Scaling

Elastic Load Balancing:

Auto Scaling groups:

Data layer resilience, replication, and backups

Relational databases:

Object storage and DR:

Centralized backups with AWS Backup:

Chaos engineering and AWS Fault Injection Simulator (FIS)

Chaos experiments validate that HA and DR mechanisms behave as designed. AWS FIS orchestrates controlled faults with guardrails:

Practical Problem Scenario

Expedia Group operates a global trip-search API that must provide sub-60-second RTO and near-zero RPO for critical booking data, while maintaining low latency for users in North America and Europe. The team experiences occasional Regional brownouts and deployment-induced instability, and auditors require cross-account immutable backups and documented DR drills.

Step-by-step approach:

  1. Establish multi-Region, active/active stacks
  1. Global, low-RPO datastore
  1. Resilient scaling and graceful transitions
  1. Durable object DR
  1. DNS controls for canary and failover
  1. Centralized, immutable backups
  1. Automated failover orchestration
  1. Chaos validation with AWS FIS

Why these services:


Containers and Serverless Operations · All domains · Event-Driven Architectures and Automation

Practice these questions → · Timed practice on ExamRoll.io →

Pass the whole exam — not just this question

You found this answer. Get every verified question and explanation in one place, and save hours of prep. Free to start.

Pass your exam →

Browse Amazon →

Related guides

All-in-one access

One subscription. Every exam.

Every plan unlocks unlimited answer search, practice tests, AI explanations, and the full resource library — in 20+ languages.

Monthly
24.87
Just €0.83/day
Everything included:
  • Unlimited answer search
  • Unlimited practice tests
  • AI-powered explanations
  • Full resource library
  • 20+ languages
  • Weekly content updates
  • Rewards & referrals
  • Priority support
Start free trial

No credit card required*

Best value
12 months
179.87
Just €0.49/daySave 40%
Everything included:
  • Unlimited answer search
  • Unlimited practice tests
  • AI-powered explanations
  • Full resource library
  • 20+ languages
  • Weekly content updates
  • Rewards & referrals
  • Priority support
Start free trial

No credit card required*

✓ Free plan included · ✓ Cancel anytime · ✓ All plans unlock the full product