Disaster recovery: RPO/RTO, backup-restore to active-active
7 min read · about 1 h 25 min with practice3 quick checks≈4% of the testStretch: Stretch: harder material that separates the top grades
Reading is free. Sign in to tick off lessons, keep your place and track your mastery.
Disaster recovery (DR) questions give you a recovery point objective (RPO), a recovery time objective (RTO) and a budget signal. They ask for the cheapest strategy that still meets both objectives. They are some of the most reliable marks on the exam once you learn the strategy ladder and which replication feature matches each data store. They are also where over-engineered distractors catch most candidates.
By the end you’ll be able to
Translate RPO/RTO requirements into the cheapest adequate strategy: backup/restore, pilot light, warm standby or active-active
Choose cross-Region replication mechanisms per data store and know their typical RPO (Aurora Global Database, DynamoDB global tables, S3 CRR with RTC)
Automate recovery with AWS Backup cross-Region copies, CloudFormation and AMIs, and Route 53 failover
Pre-provision service quotas and capacity in the recovery Region
What the exam asks
Translate RPO/RTO into backup and restore, pilot light, warm standby or multi-site active-active. Then pick the cheapest one that qualifies.
Choose the cross-Region replication feature for each data store:
Aurora Global Database
DynamoDB global tables
RDS cross-Region read replicas
S3 Cross-Region Replication (CRR) with Replication Time Control (RTC)
AWS Backup copies
Protect backups from deletion and ransomware (cross-account copies, Vault Lock), and recover from logical corruption, which replicas copy faithfully.
Make the recovery Region ready: AMIs, infrastructure as code, service quotas, and traffic failover with Route 53 or Global Accelerator.
Estimate simple transfer times to test whether a restore fits inside the RTO.
Core ideas
RPO and RTO
RPO: how much data, measured in time, the business can lose. It is decided by how often data leaves the primary site (backup frequency or replication lag).
RTO: how long the business can be down. It is decided by how much of the recovery environment already runs and how automated the failover is.
The less you pre-provision, the cheaper the strategy and the longer the RTO.
The four strategies
Strategy
What runs in the DR Region
Typical RPO / RTO
Relative cost
Backup and restore
Nothing but backups (snapshots, AMIs, S3 copies) plus IaC templates
Hours / hours to a day
Lowest
Pilot light
Data layer live and replicating; application servers off (AMIs ready, IaC to launch)
Seconds–minutes / tens of minutes to hours
Low
Warm standby
A copy that can take some traffic now, then scales out
vii.Check your understanding
3 questions on disaster recovery: RPO/RTO, backup-restore to active-active. Every option is explained once you answer.
Sign in to try the quick check
Answers are checked on our side, every option is explained, and your result feeds your mastery for this topic. It’s free.
The first 3 of 10 cards for this topic. Sign in and finish the lesson to review them with spaced repetition.
PromptCard 1 of 3
RPO vs RTO: what drives each?
scaled-down but fully working
Seconds / minutes
Medium
Multi-site active-active
Full production capacity serving traffic in 2+ Regions
Near zero / near zero
Highest
Pilot light vs warm standby is the classic exam distinction.
A pilot light environment cannot serve requests until you launch compute.
A warm standby environment can serve reduced traffic immediately.
If the stem says “must handle some production traffic without provisioning”, the answer is warm standby.
Cross-Region data replication
Data store
DR feature
Key facts
Aurora
Aurora Global Database
Storage-level replication, typically under 1 second of lag; promote the secondary in about a minute; secondary serves local reads
DynamoDB
Global tables
Multi-Region, multi-active (write in every Region); asynchronous; last writer wins
RDS (non-Aurora)
Cross-Region read replica
Asynchronous; promote manually to become a standalone writer
RDS
Automated backups + snapshot copies (AWS Backup)
RPO depends on backup frequency; restore creates a new instance
S3
CRR (+ RTC for an SLA of 99.99% of objects within 15 minutes)
Versioning required on both buckets; new objects only, so use S3 Batch Replication for existing objects
EBS / EC2
Snapshot and AMI copies (AWS Backup copy rules, Data Lifecycle Manager)
Point-in-time; RPO = schedule
EFS / FSx
AWS Backup cross-Region copies; EFS replication
Managed
Backups that survive the worst day
AWS Backup gives you central backup plans (schedules, retention), copy rules to another Region or another account in your organization, and Vault Lock. In compliance mode, Vault Lock makes recovery points immutable, even for the root user.
Replication is not backup. Multi-AZ standbys, read replicas, global tables and CRR all copy deletes and corruption. To undo a bad batch job, use point-in-time recovery (PITR). PITR restores a new DB instance to a moment before the damage.
Readiness in the recovery Region
Infrastructure: keep AMIs copied and describe the stack in CloudFormation, deployed with StackSets across Regions. An untested runbook is not a DR plan.
Service quotas are per Region. Raise them in the DR Region in advance. On-Demand Capacity Reservations guarantee capacity for critical instance types.
Traffic: Route 53 failover records with health checks, or Global Accelerator when clients cache DNS or need static IPs (TCP/UDP).
Transfer-time arithmetic
Throughput depends on the link:
1 Gbps ≈ 125 MB/s ≈ 10.8 TB per day.
10 Gbps ≈ 108 TB per day.
If the restore data must cross the link after the disaster, divide the data size by the daily throughput and compare the result with the RTO. If it does not fit, the data must already be in AWS before the disaster.
Worked examples
Exam technique
Write the numbers down. Note the RPO, the RTO and the cost qualifier.
Climb the ladder from the bottom. Start at backup and restore and move up only when a number is missed. The first rung that meets both objectives is the answer to MOST cost-effective.
Ask what is running. Only data means pilot light. Data plus small compute means warm standby. Full capacity everywhere means active-active.
“Clients cache DNS” or “fixed IPs” ⇒ Global Accelerator, not a lower TTL.
Common mistakes
Quick recap
RPO is driven by how often data leaves the primary site. RTO is driven by how much of the recovery environment already runs.
The ladder: backup and restore, then pilot light (data live), then warm standby (small working copy), then active-active. Pick the lowest rung that meets both numbers.
Aurora Global Database gives about 1 second RPO and about 1 minute RTO with one writer. DynamoDB global tables are multi-writer.
S3 CRR needs versioning on both buckets; add RTC for an SLA and Batch Replication for existing objects.
AWS Backup copy rules to another account and Region, plus Vault Lock, give ransomware-resistant backups.
Replicas copy corruption: use PITR for logical errors.
Pre-raise quotas, copy AMIs, keep IaC ready, and fail traffic over with Route 53 or Global Accelerator.