How Domain 3 questions are built
| Pattern | Typical stem | What decides it |
|---|---|---|
| Storage by access pattern | “Hundreds of Linux instances need a shared file system” | Block vs file vs object, and the protocol |
| Performance symptom | “The database CPU is high from repeated reads” | Cache or read replicas, not a bigger instance |
| Global users | “Users worldwide report slow downloads / need fixed IPs” | CloudFront vs Global Accelerator |
| Many VPCs or on-premises links | “40 VPCs need to talk to each other and to the data centre” | Transit Gateway, Direct Connect |
| Data at volume | “Millions of clickstream events per minute” | Kinesis, Firehose, MSK, then S3 and Athena |
| Moving data in | “80 TB, 100 Mbps link, one week” | Bandwidth maths → Snowball Edge or DataSync |
Storage (task 3.1)
| Trigger | Answer |
|---|---|
| General-purpose block storage; tune IOPS without growing size | EBS gp3 |
| Highest IOPS, sub-millisecond latency, critical databases | EBS io2 Block Express |
| Large sequential reads (logs, big data), low cost | EBS st1 (HDD; cannot be a boot volume) |
| Coldest, cheapest block storage | EBS sc1 |
| Temporary scratch data at extreme IOPS; loss is acceptable | Instance store |
| Shared POSIX file system for Linux across AZs | EFS |
| SMB, Windows, Active Directory, DFS | FSx for Windows File Server |
| HPC, sub-millisecond, linked to S3 data | FSx for Lustre |
| NFS + SMB + iSCSI together, NetApp features | FSx for NetApp ONTAP |
| Migrate ZFS or Linux NFS file servers | FSx for OpenZFS |
| Single-digit-millisecond object access next to compute | S3 Express One Zone |
S3 performance: each prefix supports at least 3,500 write and 5,500 read requests per second, so spread keys across prefixes for more. Use multipart upload for large objects (recommended over 100 MB, required over 5 GB) and byte-range fetches for parallel downloads. Use Transfer Acceleration for long-distance uploads to one bucket.
Hybrid storage (Storage Gateway): S3 File Gateway (NFS/SMB into S3), Volume Gateway cached (primary data in AWS, hot data cached locally) or stored (primary data on premises, asynchronous backup to AWS), and Tape Gateway (replaces physical tape libraries).
Compute (task 3.2)
| Trigger | Answer |
|---|---|
| Lowest latency between nodes, HPC, tightly coupled jobs | Cluster placement group, plus ENA or EFA for MPI |
| Critical instances must not share hardware | Spread placement group (at most 7 running instances per AZ per group) |
| HDFS, Cassandra, Kafka: isolate groups of nodes by rack | Partition placement group |
| Many batch jobs with queues and dependencies | AWS Batch |
| Spark or Hadoop on large datasets | EMR |
| Lambda is too slow on CPU-bound work | Increase Lambda memory (CPU scales with it) |
| Single-digit-millisecond latency to users in a metro area | Local Zones |
| Ultra-low latency to 5G mobile devices | Wavelength |
| AWS infrastructure on premises (data residency, local processing) | Outposts |
Instance families: C for compute-bound work, R/X for memory, I/D for storage IOPS, P/G/Inf/Trn for accelerators, and Graviton for price-performance.
Databases and caching (task 3.3)
| Trigger | Answer |
|---|---|
| Read-heavy relational load | Read replicas and the reader endpoint |
| Fastest failover, up to 15 low-lag replicas, storage grows automatically | Aurora |
| Unpredictable or intermittent relational load | Aurora Serverless v2 |
| Too many connections from Lambda or bursty clients | RDS Proxy |
| Key-value at any scale, single-digit-millisecond latency | DynamoDB |
| Microsecond reads on DynamoDB | DAX |
| Global users writing locally | DynamoDB global tables or Aurora Global Database (single writer) |
| MongoDB-compatible documents | DocumentDB |
| Relationships and graphs (fraud rings, social links) | Neptune |
| Cassandra (CQL) workloads | Keyspaces |
| Complex analytics over large structured data | Redshift |
| Full-text search, log analytics | OpenSearch Service |
| Session store, leaderboards, cache in front of RDS | ElastiCache (Valkey or Redis OSS) |
Caching strategy: use lazy loading (fill the cache on a miss, with a TTL) when stale data is acceptable and not every item is read. Use write-through when reads must always be fresh. Choose Valkey or Redis OSS for replication, Multi-AZ, persistence and sorted sets. Choose Memcached for a simple, multi-threaded cache with no persistence.
Networking (task 3.4)
| Trigger | Answer |
|---|---|
| Cache static and dynamic HTTP content near users | CloudFront |
| Static anycast IPs, TCP/UDP, fast Regional failover, IP allow-listing by clients | Global Accelerator |
| A few VPCs, simple point-to-point | VPC peering (not transitive; no overlapping CIDRs) |
| Many VPCs and on-premises networks, transitive hub | Transit Gateway |
| Consistent bandwidth and latency to on premises | Direct Connect (with a VPN backup, or two locations for resilience) |
| One DX connection reaching VPCs in several Regions | Direct Connect gateway |
| More VPN throughput | Several tunnels with ECMP on a Transit Gateway; accelerated VPN |
| Quick, encrypted, low-cost hybrid link | Site-to-Site VPN |
Plan CIDRs that do not overlap with on premises or with other VPCs you may connect in future. Add a secondary CIDR when a VPC runs out of addresses.
Data ingestion and transformation (task 3.5)
| Trigger | Answer |
|---|---|
| Real-time stream, several custom consumers, replay, ordering per shard | Kinesis Data Streams |
| Load streaming data into S3, Redshift or OpenSearch with no code; convert to Parquet | Amazon Data Firehose |
| Existing Apache Kafka | Amazon MSK |
| Video from cameras | Kinesis Video Streams |
| Catalogue and transform data in S3 (CSV to Parquet) | Glue crawlers, Data Catalog and Glue ETL |
| Serverless SQL queries directly on S3 | Athena |
| Row- and column-level permissions on a data lake | Lake Formation |
| Dashboards | Amazon Quick |
Moving data in: the bandwidth rule
1 Gbps moves about 10 TB a day at full use, and 100 Mbps moves about 1 TB a day. Real links run at perhaps 70–80% of that. So 80 TB over a 100 Mbps link takes about 100 days, which means Snowball Edge. The same 80 TB over a mostly free 10 Gbps Direct Connect takes about a day, which means DataSync.
- Online, scheduled or incremental transfer from NFS, SMB or HDFS → DataSync
- Ongoing on-premises access to cloud storage → Storage Gateway
- Partners send files over SFTP → Transfer Family
- Database migration with minimal downtime → DMS with change data capture (plus Schema Conversion for a different engine)
- Lift-and-shift servers → Application Migration Service
Common traps
Pacing in Domain 3
Most items are trigger-matching and take 45–75 seconds. Spend the time you save on the transfer-time and network topology items. For transfer questions, write out the division (data ÷ daily capacity) instead of guessing.