Google PCA: Reliability, Disaster Recovery and Business Continuity — Study Guide

Part of the Google Professional Cloud Architect — Study Guide. Practice with verified answers in the Google exam hub, or take timed practice tests on ExamRoll.io.

Overview

Reliability, disaster recovery (DR), and business continuity ensure that services continue to meet agreed objectives despite component, zone, or regional failures. On Google Cloud, reliability is engineered by understanding failure domains (zonal, regional, and global), defining recovery objectives (RTO/RPO), selecting resilient service architectures (e.g., active-active), and rigorously testing recovery plans. Your design must map service criticality to explicit availability targets, durability guarantees, and validated recovery paths, balancing availability, consistency, cost, and operational complexity. Key themes include isolating single points of failure, using managed replication where possible, automating failover decisions, and continuously validating that assumptions hold in production-like conditions.

Failure Domains, Locations, and Multi-Regional Services

Trade-offs:

Objectives, Dependency Mapping, and DR Validation

Resilient Compute, Databases, and Storage Patterns

Traffic Management, Multi-Site Strategies, and Continuous Resilience

Practical Problem Scenario

Acme Tickets, a fast-growing online ticketing company, must ensure continuous operations during regional outages for its purchase API and event catalog while maintaining stringent RTO/RPO (RTO ≤ 5 minutes, RPO ≤ 1 minute). The stack includes stateless microservices, a relational order database, an analytics pipeline, and static media assets.

  1. Define service tiers, SLOs, and recovery objectives
  1. Choose regional placement and multi-site strategy
  1. Implement regional MIGs and global HTTP(S) load balancing
  1. Externalize state and configure self-healing
  1. Provision Cloud SQL for PostgreSQL with HA and cross-region read replica
  1. Schedule routine database failover tests
  1. Place static media in a dual-region Cloud Storage bucket with versioning and retention
  1. Protect health checks and egress with firewall and quotas
  1. Implement DNS as a coarse control with low TTL
  1. Automate DR runbooks and validate via game days
  1. Secure and separate backups
  1. Implement graceful degradation

This architecture delivers automated regional failover for stateless services, controlled and tested failover for stateful components, and verified recovery processes that meet Acme Tickets’ business continuity objectives.


Security · All domains · Migration

Practice these questions → · Timed practice on ExamRoll.io →

Pass the whole exam — not just this question

You found this answer. Get every verified question and explanation in one place, and save hours of prep. Free to start.

Pass your exam →

Browse Google →

Related guides

All-in-one access

One subscription. Every exam.

Every plan unlocks unlimited answer search, practice tests, AI explanations, and the full resource library — in 20+ languages.

Monthly
24.87
Just €0.83/day
Everything included:
  • Unlimited answer search
  • Unlimited practice tests
  • AI-powered explanations
  • Full resource library
  • 20+ languages
  • Weekly content updates
  • Rewards & referrals
  • Priority support
Start free trial

No credit card required*

Best value
12 months
179.87
Just €0.49/daySave 40%
Everything included:
  • Unlimited answer search
  • Unlimited practice tests
  • AI-powered explanations
  • Full resource library
  • 20+ languages
  • Weekly content updates
  • Rewards & referrals
  • Priority support
Start free trial

No credit card required*

✓ Free plan included · ✓ Cancel anytime · ✓ All plans unlock the full product