What the exam asks
- Build the lake. Store raw and curated data in S3, discover schemas with Glue crawlers, keep metadata in the Glue Data Catalog, and transform CSV or JSON into partitioned, columnar Parquet.
- Query it. Use Athena for serverless ad hoc SQL. Use Redshift for frequent, complex warehouse queries, and Redshift Spectrum to join the warehouse with data in S3. Use Athena Federated Query to read live data from other databases.
- Secure it. Use Lake Formation for database, table, column, row and cell permissions, tag-based access control and cross-account sharing.
- Process it. Choose EMR, Glue or Athena from the framework and the effort involved.
- Show it. Use Amazon Quick (formerly Amazon QuickSight) for BI dashboards and OpenSearch Service for search and log analytics.
Core ideas
The lake, layer by layer
| Layer | Service | Exam trigger |
|---|---|---|
| Storage | Amazon S3 (raw, curated, analytics prefixes) | “durable, low-cost, any format” |
| Catalog | Glue crawlers and the Glue Data Catalog | “discover schema”, “make new datasets queryable” |
| Transform | Glue ETL (serverless Spark), Glue DataBrew (visual, no code), EMR | “convert to Parquet”, “clean”, “existing Spark/Hadoop” |
| Govern | AWS Lake Formation | “column-level”, “row-level”, “share with other accounts”, “central permissions” |
| Query | Athena, Redshift and Redshift Spectrum | “ad hoc SQL, pay per query” versus “complex, frequent BI queries” |
| Visualise | Amazon Quick, OpenSearch Dashboards | “interactive business dashboards” versus “log search” |
| Third-party data | AWS Data Exchange | “subscribe to external datasets delivered to S3” |
Glue
- infer schemas and partitions and create or update tables in the . Athena, EMR, Redshift Spectrum and Lake Formation all use the catalog as their shared, Hive-compatible metastore.