What the exam asks
- Pick the ingestion service from the keywords. Custom real-time processing, replay and several independent consumers point to Kinesis Data Streams. Zero-admin delivery to S3, Redshift, OpenSearch or an HTTP endpoint points to Data Firehose. Existing Kafka clients point to MSK. Camera video points to Kinesis Video Streams.
- Fix throughput problems. Expect throttling from a hot partition key, too few shards, or read contention between consumers.
- Assemble a pipeline with the least operational overhead. A typical chain is stream, then transform, then store as Parquet, then query.
- Choose between a stream and a queue. The deciding facts are ordering, replay and how many consumers need each record.
A typical stem reads: “A company collects clickstream data... The data must be available in Amazon S3 for analysis in near real time... Which solution meets these requirements with the LEAST operational overhead?”
Core ideas
Kinesis Data Streams (KDS)
| Property | What to remember |
|---|---|
| Capacity unit | Shard (provisioned mode). Each shard accepts up to 1 MB/s or 1,000 records/s of writes and serves 2 MB/s of reads, which standard consumers share. |
| On-demand mode | No shard planning. Capacity scales automatically and you pay per GB. Choose it for unpredictable or spiky traffic. |
| Partition key | A hash of the key selects the shard. Ordering is guaranteed per shard, so records with the same key stay in order. |
| Retention | 24 hours by default. You can extend it to 7 days, or up to 365 days. Any consumer can replay from any point inside the retention window. |
| Consumers | Many applications read the same records independently: Lambda, KCL apps on EC2 or ECS, Managed Service for Apache Flink, or Firehose. Enhanced fan-out gives each registered consumer its own 2 MB/s per shard with push delivery and lower latency. |
| Producers | SDK PutRecord/, the Kinesis Producer Library (batching and of small records) or the Kinesis Agent. |