What the exam asks
- Pick the engine from the access pattern. Relationships several hops deep, existing MongoDB drivers, CQL, SQL analytics over terabytes, full-text search, sub-millisecond lookups.
- Design keys and indexes. Which partition key avoids hot partitions, and whether a new query needs a GSI or an LSI.
- Choose the capacity mode. On-demand or provisioned with auto scaling, and when DAX is the real fix.
- Add the right feature. Streams to react to changes, TTL for expiry, global tables for multi-Region writes, PITR for recovery, export to S3 for analytics, transactions for all-or-nothing writes.
Core ideas
How DynamoDB stores data
A table holds items of up to 400 KB each. Every item has a primary key:
- Partition key only (simple key). DynamoDB hashes the value to choose a physical partition, so every value must be unique.
- Partition key + sort key (composite key). Items that share a partition key are stored together, sorted by the sort key. One
Querycan then return all orders for customer 42, newest first.
Query reads a single partition-key value and is efficient. Scan reads the entire table and filters afterwards, and you pay for every item it reads. For a production access pattern, “scan the table” is almost always the wrong answer.
Each partition has a fixed throughput ceiling. A key with few distinct values, such as a status flag, a contestant ID or today’s date, sends all its traffic to one partition. That partition throttles even when the table as a whole has spare capacity. There are two fixes. Use a high-cardinality key, or use write sharding: add a random or calculated suffix to the key and sum across the shards when you read.
Secondary indexes
| Global secondary index (GSI) | Local secondary index (LSI) | |
|---|---|---|
| Keys | Any partition key and sort key | Same partition key as the table, different sort key |
| When created | At any time | Only when the table is created |