How Domain 4 questions are built
| Pattern | Typical stem | What decides it |
|---|---|---|
| Purchase option | “Steady 24/7 baseline plus unpredictable spikes” | Commitment for the baseline, flexible pricing for the rest |
| Storage class | “Accessed often for 30 days, rarely after, kept 7 years” | Lifecycle transitions, minimum durations, retrieval time |
| Right service | “The job runs 10 minutes an hour” | Pay per use (Lambda, Fargate) over idle EC2 |
| Data transfer | “The NAT gateway bill is high” / “cross-AZ charges” | Endpoints, locality, CloudFront |
| Visibility | “Finance needs cost per project” / “alert at 80% of budget” | Tags, Cost Explorer, Budgets |
Compute purchasing (task 4.2)
| Workload | Cheapest adequate option |
|---|---|
| Steady, predictable, 1–3 years, may change family, Region or move to Fargate/Lambda | Compute Savings Plan |
| Steady, fixed instance family in one Region | EC2 Instance Savings Plan or Standard RI (deepest commitment discount) |
| Steady, but may need to exchange instance types | Convertible RI |
| Fault-tolerant, stateless, flexible timing (batch, CI, rendering, big data) | Spot (deepest discount; can be interrupted with a 2-minute notice) |
| Short-term, unpredictable, cannot be interrupted | On-Demand |
| Guarantee capacity in an AZ for an event or DR | On-Demand Capacity Reservation (combine with a Savings Plan for a discount) |
| Per-socket or per-core BYOL licensing, host-level visibility | Dedicated Host |
| Single-tenant hardware only, no licensing need | Dedicated Instances |
The mixed pattern: cover the steady baseline with a Savings Plan or RIs, then scale above it with On-Demand, or with Spot if the work tolerates interruption. A mixed-instances Auto Scaling group with several instance types and capacity-optimized Spot allocation lowers the chance of interruptions.
Design choices that cut compute cost:
- Rightsize with Compute Optimizer before you commit to anything. Buying a Savings Plan for oversized instances locks in the waste.
- Move to Graviton instances for better price-performance.
- Use Lambda or Fargate for spiky or mostly idle work, and EC2 with commitments for steady high utilisation.
- For non-production, use scheduled scaling to zero, stop and start automation, or hibernation outside working hours.
- Choose the simplest load balancer that meets the need. Do not add a GWLB unless there are appliances to insert.
Storage (task 4.1)
S3 storage classes
| Class | Use when | Retrieval | Minimum storage duration |
|---|---|---|---|
| Standard | Frequent access | Milliseconds | None |
| Intelligent-Tiering | Unknown or changing access pattern | Milliseconds for the frequent, infrequent and archive instant tiers | None (small monitoring fee per object) |
| Standard-IA | Infrequent access, needs fast retrieval, multi-AZ | Milliseconds, retrieval fee | 30 days |
| One Zone-IA | Infrequent access, re-creatable data | Milliseconds, retrieval fee | 30 days |
| Glacier Instant Retrieval | Rare access (about once a quarter), but milliseconds needed | Milliseconds | 90 days |
| Glacier Flexible Retrieval | Archive, minutes to hours is fine | 1–5 min expedited, 3–5 h standard, 5–12 h bulk | 90 days |
| Glacier Deep Archive | Compliance archive, rarely or never read | Within 12 h standard, within 48 h bulk | 180 days |
Lifecycle rules the exam tests:
- Objects must stay 30 days in Standard before a lifecycle rule can move them to Standard-IA or One Zone-IA.
- Deleting or moving objects before their minimum duration still charges you for the full minimum.
- IA classes bill small objects as if they were 128 KB, so millions of tiny files can cost more in IA than in Standard.
- Add rules to expire noncurrent versions and abort incomplete multipart uploads.
- If downloaders should pay for data transfer, use Requester Pays.
Block, file and backup
- gp2 → gp3: cheaper per GB, with IOPS set independently of size. Change it in place with no downtime.
- Delete unattached EBS volumes and old snapshots. Automate snapshots with Data Lifecycle Manager. Move long-term snapshots to the snapshot archive tier (much cheaper storage, 90-day minimum, restore takes hours).
- EFS lifecycle management moves cold files to Infrequent Access and Archive. EFS One Zone suits data that does not need multi-AZ.
- FSx: single-AZ and HDD options are cheaper where the requirements allow.
- Backups: AWS Backup with retention rules and cold storage. Use Tape Gateway to Glacier Deep Archive to replace physical tape.
Databases (task 4.3)
| Situation | Cost-optimized choice |
|---|---|
| Steady production RDS or Aurora | Reserved DB instances |
| Intermittent or unpredictable relational load, dev/test | Aurora Serverless v2 |
| Aurora where I/O is a large share of the bill | Aurora I/O-Optimized |
| DynamoDB with unpredictable or new traffic | On-demand capacity |
| DynamoDB with steady, predictable traffic | Provisioned capacity with auto scaling (plus reserved capacity) |
| DynamoDB table with rarely accessed data | Standard-IA table class |
| Repeated reads overloading the primary | ElastiCache or read replicas, not a bigger instance |
| Old data in expensive tables | Export to S3 and query with Athena, or use DynamoDB TTL to expire items |
| Commercial engine licence costs | Migrate to Aurora or open-source engines with DMS and Schema Conversion |
Networking (task 4.4)
What is charged
| Traffic | Charged? |
|---|---|
| Data into AWS from the internet | Free |
| Within one AZ over private IPs | Free |
| Between AZs | Charged, in each direction |
| Between Regions | Charged |
| Out to the internet | Charged (cheaper through CloudFront, and S3/EC2 to CloudFront is free) |
| Through a NAT gateway | Hourly charge plus a per-GB processing charge |
| Gateway endpoint (S3, DynamoDB) | Free |
| Interface endpoint | Hourly charge plus per GB |
| VPC peering | No hourly charge; normal cross-AZ or cross-Region data rates apply |
| Transit Gateway | Per-attachment hourly charge plus per-GB processing |
Decision rules:
- Private subnets pulling large volumes from S3 or DynamoDB → a gateway endpoint removes the NAT processing charge.
- A few VPCs with heavy traffic between them → peering is cheaper than Transit Gateway. Many VPCs → Transit Gateway is simpler, and simplicity is often the stated requirement.
- One NAT gateway per AZ adds hourly cost but avoids cross-AZ charges and a single point of failure. Choose shared NAT only when the stem accepts lower availability, for example in dev/test.
- Large, steady hybrid transfer → Direct Connect (lower egress rate, consistent bandwidth). Small or occasional → VPN.
- Heavy downloads from origin → CloudFront caching cuts both egress and origin load.
- API spend growing from abuse → API Gateway throttling and usage plans.
Cost-management tools
| Need | Tool |
|---|---|
| Analyse spend, forecast, Savings Plans recommendations | Cost Explorer |
| Alert or act when spend or usage crosses a threshold | AWS Budgets (with budget actions) |
| Most detailed line-item data for analysis | Cost and Usage Report (CUR), queried with Athena |
| Cost per team, project or environment | Cost allocation tags (activate them in Billing) |
| Share volume discounts and RIs across accounts | Consolidated billing in AWS Organizations |
| Rightsizing recommendations | Compute Optimizer |
| Idle resources and best-practice checks | Trusted Advisor |
Common traps
Pacing in Domain 4
Handle every cost item in two passes. First, cross out any option that breaks a requirement (retrieval time, durability, availability, interruption). Then rank the rest by cost. Most cost items take 60–90 seconds. S3 lifecycle items with several time periods are worth 2 minutes. Sketch the timeline on your note board.