How Domain 2 questions are built
| Pattern | Typical stem | What decides it |
|---|---|---|
| Traffic spikes lose data | “Orders are lost when traffic surges” | Put a queue between tiers |
| One event, many consumers | “Three systems must each process every upload” | SNS fan-out or EventBridge |
| Job too long or too big | “Processing takes 45 minutes” | Not Lambda: Fargate, ECS, Batch or Step Functions |
| Survive an AZ failure | “The app must stay available if an AZ fails” | Multi-AZ everything, including NAT and the database |
| Survive a Region failure | “RPO 1 minute, RTO 15 minutes, lowest cost” | The cheapest DR strategy that meets both numbers |
| Legacy app, no code changes | “The application cannot be modified” | Infrastructure-level fixes (ASG, Multi-AZ, EFS, RDS Proxy) |
Decision rules: loose coupling (task 2.1)
Messaging
| Requirement | Answer |
|---|---|
| Buffer work between tiers; consumers pull at their own pace | SQS standard |
| Strict order and exactly-once processing | SQS FIFO (order is kept per message group ID) |
| Push one message to many subscribers | SNS topic, fan-out to several SQS queues |
| Route events by content, from AWS services, SaaS apps or other accounts | EventBridge rules |
| Run a task on a schedule without a server | EventBridge Scheduler |
| Existing app uses JMS, AMQP, MQTT or STOMP and must not change | Amazon MQ |
| Real-time stream, several consumers, replay | Kinesis Data Streams (Domain 3) |
SQS tuning rules the exam tests:
- Messages processed twice → the visibility timeout is shorter than the processing time. Increase it.
- Messages that keep failing → a dead-letter queue with a maxReceiveCount.
- Too many empty responses and wasted API calls → long polling (wait time up to 20 seconds).
- Scale consumers → an Auto Scaling policy on backlog per instance (queue depth ÷ number of instances). Raw CPU is the wrong metric.
Serverless and containers
- Lambda runs for at most 15 minutes. Longer work goes to ECS or EKS on Fargate, AWS Batch, or is split into steps with Step Functions.
- Step Functions Standard: long-running (up to a year), exactly-once, human approval steps. Express: high-volume, short (up to 5 minutes), at-least-once.
- API Gateway protects backends with throttling and usage plans and can cache responses.
- Lambda to RDS under bursty load → RDS Proxy to pool connections.
- Containers without managing servers → Fargate. Existing Kubernetes tooling → EKS. Otherwise ECS is the simpler choice.
Load balancing and Auto Scaling
| Requirement | Answer |
|---|---|
| HTTP routing by path or host, Lambda targets, user authentication | ALB |
| TCP/UDP, static IP per AZ, extreme throughput, PrivateLink | NLB |
| Insert third-party firewalls or appliances transparently | GWLB |
| Keep a steady target such as 50% CPU or 1,000 requests per target | Target tracking scaling |
| Known daily or weekly peaks | Scheduled scaling (or predictive scaling) |
| Slow-booting instances | Warm pools, or lifecycle hooks during launch |
| Finish work before an instance terminates | Lifecycle hook on termination |
Stateless design
Move session state to ElastiCache or DynamoDB, shared files to EFS or S3, and configuration to Parameter Store. Then any instance can serve any request, and the Auto Scaling group can replace instances freely. Sticky sessions are a workaround, not an HA design.
Decision rules: high availability and DR (task 2.2)
Inside a Region
| Component | Highly available version |
|---|---|
| EC2 | Auto Scaling group across at least 2 AZs behind a load balancer |
| Legacy single instance that cannot scale out | ASG with min = max = 1 across AZs (self-healing), plus EFS for its data |
| NAT | One NAT gateway per AZ, each private route table using its own AZ’s NAT gateway |
| RDS | Multi-AZ instance deployment (the standby is not readable; failover typically 1–2 minutes) |
| Faster failover plus readable standbys | Multi-AZ DB cluster, or Aurora |
| Aurora | Replicas in other AZs act as failover targets (typically under 30 seconds); up to 15 replicas |
| Hybrid link | Two VPN tunnels, or DX with a VPN backup, or two DX connections at two locations |
Route 53
| Requirement | Routing policy |
|---|---|
| Active-passive DR | Failover with health checks |
| Send users to the lowest-latency Region | Latency |
| Canary release, 10% to the new version | Weighted |
| Content or legal rules by country | Geolocation |
| Shift traffic by distance with a bias | Geoproximity |
| Return several healthy IPs | Multivalue answer |
Use alias records for the zone apex and AWS targets. A CNAME cannot be created at the apex.
Disaster recovery strategies
Choose the cheapest strategy whose RPO and RTO both meet the requirement:
| Strategy | What runs in the DR Region | Typical RPO / RTO | Cost |
|---|---|---|---|
| Backup and restore | Only backups (AWS Backup copies, snapshots, S3 CRR) | Hours / hours | Lowest |
| Pilot light | Data replicated live; compute switched off or not yet created | Minutes / tens of minutes | Low |
| Warm standby | A scaled-down but working full stack | Seconds–minutes / minutes | Medium |
| Multi-site active-active | Full production in both Regions | Near zero / near zero | Highest |
Replication tools: Aurora Global Database (typically under 1 second of lag, promotion in about a minute), DynamoDB global tables (multi-Region, multi-active), S3 CRR with Replication Time Control (99.99% of objects within 15 minutes), cross-Region read replicas for RDS, and AWS Backup cross-Region copies. Remember to raise service quotas in the DR Region in advance.
Observability and automation
- Memory and disk metrics are not collected by default. Install the CloudWatch agent.
- To see where latency comes from across microservices, use X-Ray.
- To deploy the same stack to many accounts and Regions, use CloudFormation StackSets. To detect manual changes, use drift detection.
- For throttling errors from AWS APIs, use exponential backoff with jitter, and a queue to absorb bursts.
Common traps
Pacing in Domain 2
DR questions reward writing down the numbers. Note the RPO, RTO and qualifier on your note board (“RPO 5m / RTO 30m / cost”), then check each option against them. That takes 20 seconds and prevents most over-engineering errors. Messaging items are usually quick (60 seconds) once you have spotted the ordering, fan-out or duration keyword.