What the exam asks
- Collect the right metrics: the defaults, memory and disk through the CloudWatch agent, custom metrics, and metrics built from logs.
- Build alarms that match the business need. Composite alarms cut noise; alarm actions notify, scale or recover.
- Trace requests across distributed services with X-Ray, and centralise logs with subscription filters.
- Keep infrastructure consistent with CloudFormation (drift detection, StackSets, deletion policies) and operate fleets with Systems Manager.
- Handle service quotas and throttling: raise quotas ahead of time, alarm on usage, retry with backoff and jitter, and buffer bursts in queues.
Core ideas
CloudWatch metrics and alarms
- Default EC2 metrics: CPU, network, disk I/O for instance store, and status checks. Basic monitoring reports every 5 minutes; detailed monitoring reports every 1 minute. Neither mode includes memory or disk-space usage, which need the CloudWatch agent. You can install the agent at scale with Systems Manager.
- Custom metrics are published with PutMetricData, down to high resolution of 1 second.
- Alarm actions:
- SNS notifications
- Auto Scaling policies
- EC2 stop, terminate, reboot or recover (recover keeps the instance ID, private IP and Elastic IP on new hardware)
- Systems Manager OpsItems
- Composite alarms combine alarms with AND/OR/NOT, so you page only when, for example, latency AND errors are high.
- CloudWatch Synthetics canaries probe endpoints on a schedule. Anomaly detection learns a baseline and alarms on deviations from it.
CloudWatch Logs
- Metric filters turn matching log lines (for example
PAYMENT_FAILED) into a metric that you can alarm on. Nothing needs to change in the application. - Logs Insights runs interactive queries for investigation.
- Subscription filters stream log events in near real time to Kinesis Data Streams, Amazon Data Firehose or Lambda. A cross-account destination lets many accounts stream into one central logging account.