High availability across AZs: removing single points of failure
7 min read · about 1 h 25 min with practice3 quick checks≈4% of the testCore: Core: tested on most papers
Reading is free. Sign in to tick off lessons, keep your place and track your mastery.
Task statement 2.2 asks you to design highly available, fault-tolerant architectures. The most common version gives you an architecture, asks what happens when one Availability Zone (AZ) fails, and wants the managed feature that removes the weak point with the least effort. Spotting single points of failure quickly earns marks across all four domains.
By the end you’ll be able to
Identify and remove single points of failure in a given architecture (single instance, single AZ, single NAT gateway, single VPN tunnel)
Compare RDS Multi-AZ instance deployment, Multi-AZ DB cluster and Aurora replicas for failover time and read capacity
Use RDS Proxy to pool connections and speed up failover for Lambda and bursty applications
Improve resilience of applications that cannot be changed (ASG with min=max=1 self-healing, ELB health checks, EFS for shared state, Elastic IPs)
Apply immutable infrastructure and blue/green practices to reduce deployment risk
What the exam asks
Find the single point of failure (SPOF): one instance, one AZ, one NAT gateway, a Single-AZ database, a file server on one instance, or a Lambda function in one subnet.
Pick the database HA feature: RDS Multi-AZ DB instance, Multi-AZ DB cluster, Aurora Replicas or read replicas, by failover time and read capacity.
Use RDS Proxy for connection storms and faster, DNS-independent failover.
Harden legacy applications without code changes: self-healing Auto Scaling groups, ELB health checks, EFS for shared state, Elastic IP re-association.
Size for static stability and deploy immutably (blue/green, fast rollback).
The qualifier usually decides: MOST highly available, automatically, LEAST operational overhead, without modifying the application.
Core ideas
Know the blast radius of every resource
Resource
Failure scope
What makes it highly available
EC2 instance
One AZ
Auto Scaling group across 2+ AZs behind a load balancer
EBS volume
One AZ
Snapshots (Regional); Multi-Attach works only within one AZ
Zonal NAT gateway
One AZ
One NAT gateway per AZ; each AZ’s private route table uses its own
ALB / NLB
Node per enabled AZ
Enable subnets in 2+ AZs
RDS Single-AZ
One AZ
Multi-AZ deployment
Aurora
Storage: 6 copies across 3 AZs
Add an Aurora Replica in another AZ for fast compute failover
ElastiCache (Valkey / Redis OSS)
Node per AZ
vii.Check your understanding
3 questions on high availability across AZs: removing single points of failure. Every option is explained once you answer.
Sign in to try the quick check
Answers are checked on our side, every option is explained, and your result feeds your mastery for this topic. It’s free.
The first 3 of 10 cards for this topic. Sign in and finish the lesson to review them with spaced repetition.
PromptCard 1 of 3
Can the standby in an RDS Multi-AZ DB instance deployment serve read traffic?
Replicas + Multi-AZ automatic failover
FSx for Windows File Server
Single-AZ or Multi-AZ
Multi-AZ deployment
S3 Standard, EFS Regional, DynamoDB, SQS, Lambda
Multi-AZ by design
Nothing extra (VPC-attached Lambda: subnets in 2+ AZs)
S3 One Zone-IA, S3 Express One Zone, EFS One Zone
One AZ
Not for AZ-failure requirements
Database high availability compared
Option
Standby serves reads?
Failover
Use when
RDS Multi-AZ DB instance
No
Automatic; typically 1–2 minutes (endpoint flips to the standby)
HA for any RDS engine
RDS Multi-AZ DB cluster
Yes: two readers, reader endpoint
Automatic; faster (AWS quotes typically under 35 seconds)
RDS for MySQL/PostgreSQL needing faster failover and reads, same engine
Aurora with replicas
Yes: up to 15, reader endpoint
Automatic; typically about 30 seconds
Fastest managed failover and read scaling
RDS read replica
Yes
Manual promotion; asynchronous
Read scaling and DR, not automatic HA
Aurora promotion rule: the replica with the lowest tier number (0 is highest priority) is promoted. On a tie, the largest instance wins. With no replicas, Aurora must create a new writer on a best-effort basis. That is slower and may fail during an AZ-wide event.
RDS Proxy
RDS Proxy is a managed, multi-AZ connection pool.
It pools and reuses connections, so Lambda bursts do not exhaust max_connections.
During failover it keeps client connections open and routes to the new writer without waiting for DNS. That also fixes applications that cache DNS.
It supports IAM authentication and Secrets Manager.
Adopting it is a configuration change: point the connection string at the proxy endpoint.
Static stability
Keep enough capacity running that losing one AZ causes no degradation before Auto Scaling reacts.
Per-AZ capacity = required ÷ (AZs − 1).
If peak needs 6 instances: across 3 AZs that is 3 per AZ, so 9 in total; across 2 AZs it is 6 per AZ, so 12 in total. More AZs mean less over-provisioning.
Self-healing and health checks
Auto Scaling health checks use only EC2 status checks by default. Turn on ELB health checks (with a grace period) so the group replaces instances whose application fails.
EC2 recover action: a CloudWatch alarm on StatusCheckFailed_System moves an EBS-backed instance to new hardware, keeping its instance ID, private IP and Elastic IP. It covers hardware failure, not AZ failure.
Single-instance legacy servers: use an Auto Scaling group with min = max = desired = 1 across 2+ AZs and put state on EFS. User data re-associates a fixed Elastic IP at boot; the instance profile must allow it.
Immutable infrastructure and blue/green
Replace servers instead of patching them in place:
Bake a new AMI for each release.
Create a new launch template version.
Roll out the new version, using either approach:
Rolling: an Auto Scaling instance refresh.
Blue/green: a green Auto Scaling group in its own target group, with traffic shifted by ALB weighted target groups or Route 53 weighted records.
To roll back, shift the weight back. It takes minutes.
Worked examples
Exam technique
Draw AZ boxes on your noteboard and place every component. Anything in only one box is a SPOF.
“Automatically” rules out read replica promotion, manual AMI restores and scripts that someone must run.
“Cannot change the application” points to configuration fixes: Auto Scaling groups, an RDS Proxy endpoint, Multi-AZ settings, EFS mounts, Elastic IP re-association.
Map keywords to features:
faster failover + reads + stay on RDS ⇒ Multi-AZ DB cluster
connection storms / DNS caching ⇒ RDS Proxy
SMB + AD + HA ⇒ FSx for Windows Multi-AZ
shared Linux files across AZs ⇒ EFS Regional
If two options both survive an AZ loss, pick the managed one over self-managed EC2.
Common mistakes
Quick recap
Put every tier in 2+ AZs: compute, NAT, database, cache, files and Lambda subnets.
The Multi-AZ DB instance standby is failover only. A Multi-AZ DB cluster adds two readable standbys and faster failover.
Aurora promotes the lowest tier number, then the largest instance.
RDS Proxy pools connections and hides failover from DNS caches, with no code changes.