Serverless patterns: Lambda, API Gateway and Step Functions
8 min read · about 1 h 20 min with practice3 quick checks≈3% of the testCore: Core: tested on most papers
Reading is free. Sign in to tick off lessons, keep your place and track your mastery.
Serverless questions test one skill above all: spotting the limit or keyword that decides between Lambda, API Gateway, Step Functions and a container or EC2 alternative. The right answer is usually a serverless design wired correctly. The distractors break a hard limit or add servers the scenario did not need.
By the end you’ll be able to
Design event-driven serverless flows (S3, SQS, DynamoDB Streams, EventBridge to Lambda) and recognise the 15-minute Lambda timeout limit
Manage concurrency with reserved and provisioned concurrency; reduce cold starts; connect Lambda to VPC resources through RDS Proxy
Choose REST, HTTP or WebSocket APIs in API Gateway and apply throttling, usage plans, caching and endpoint types (edge, regional, private)
Orchestrate multi-step processes with Step Functions (Standard vs Express, retries, parallel/map states, human approval)
What the exam asks
Wire event sources to AWS Lambda and handle failures correctly for each invocation model.
Recognise when Lambda is ruled out: over 15 minutes, GPUs, persistent connections or very large payloads.
Manage concurrency and cold starts, protect relational databases with RDS Proxy, and give Lambda private network access.
Choose a REST, HTTP or WebSocket API, its endpoint type, throttling, usage plans and caching.
Orchestrate with AWS Step Functions and choose Standard or Express workflows.
Core ideas
Lambda limits that decide answers
Setting
Value
What it means in a question
Maximum timeout
15 minutes
“Runs for 40 minutes” ⇒ AWS Fargate, AWS Batch, or split the work with Step Functions
Memory
128 MB to 10,240 MB; CPU scales with memory
CPU-bound and slow ⇒ raise memory. Waiting on I/O ⇒ more memory will not help
Ephemeral storage (/tmp)
512 MB default, up to 10,240 MB
Large scratch files ⇒ raise ephemeral storage or mount Amazon EFS
Package size
250 MB unzipped (.zip) or a 10 GB container image
Heavy dependencies ⇒ container image
Synchronous payload
6 MB request and 6 MB response
Large uploads ⇒ presigned URL straight to Amazon S3
Lambda has no GPU option. GPU work goes to EC2 GPU instances, directly or as ECS/EKS capacity.
Three invocation models
Synchronous (API Gateway, ALB, function URLs, SDK calls): the caller waits and handles errors and retries.
Asynchronous (S3 event notifications, SNS, EventBridge): Lambda queues the event and retries a failed invocation twice. Send failures to an on-failure destination (for example an SQS queue, SNS topic, EventBridge bus or another function) or to a dead-letter queue.
Event source mapping (SQS, Kinesis Data Streams, DynamoDB Streams, Amazon MSK, Amazon MQ): Lambda polls the source and invokes the function with batches, and failure handling belongs to the source. For SQS, put a on the queue, enable so only failed messages return, and set the visibility timeout to at least six times the function timeout.
vii.Check your understanding
3 questions on serverless patterns: Lambda, API Gateway and Step Functions. Every option is explained once you answer.
Sign in to try the quick check
Answers are checked on our side, every option is explained, and your result feeds your mastery for this topic. It’s free.
The first 3 of 11 cards for this topic. Sign in and finish the lesson to review them with spaced repetition.
PromptCard 1 of 3
What is the maximum AWS Lambda timeout, and what replaces Lambda for longer jobs?
DLQ with a redrive policy
partial batch responses
Concurrency and cold starts
Concurrency = requests per second × average duration in seconds. For example, 400 requests per second at 0.25 seconds each needs about 100 concurrent executions. All functions in a Region share the account’s concurrency quota, which is 1,000 by default and can be raised.
Control
What it does
Cost
Keyword in the stem
Reserved concurrency
Sets capacity aside for one function and caps it at that number
No charge
“Guarantee capacity”, “one function starves another”, “protect a downstream system”
“Cold starts”, “consistent low latency”. Schedule it with Application Auto Scaling for predictable peaks
Reserved concurrency does not remove cold starts. Provisioned concurrency does (Lambda SnapStart also shortens start-up for Java, Python and .NET).
Lambda in a VPC
Attach a function to a VPC only when it must reach private resources such as an RDS database. Its network interfaces never get public IP addresses, so:
Internet access: keep the function in private subnets and route 0.0.0.0/0 to a NAT gateway. A public subnet does not help.
S3 or DynamoDB: add a free gateway endpoint. Other AWS services use interface endpoints.
Relational databases: put Amazon RDS Proxy in front of RDS or Aurora. It pools and reuses connections so thousands of executions do not exhaust max_connections, speeds up failover, and supports IAM authentication with credentials in Secrets Manager.
API Gateway: pick the API type
Requirement
REST API
HTTP API
WebSocket API
Usage plans, API keys, per-client quotas
Yes
No
No
Response caching
Yes
No
No
AWS WAF web ACL
Yes
No
No
Private endpoint (reachable only from VPCs)
Yes
No
No
Native JWT/OIDC authorizer
No (Cognito user pool authorizer instead)
Yes
No
Lowest cost and latency for a simple proxy
No
Yes
Not applicable
Server pushes messages to connected clients
No
No
Yes
REST API endpoint types: edge-optimized (through CloudFront points of presence, for global clients), Regional (same-Region clients, or your own CloudFront in front) and private (only from VPCs, through an execute-api interface endpoint plus a resource policy).
Other facts that decide answers:
Throttling works at account, stage, method and usage-plan level, and callers over the limit get HTTP 429. API keys meter callers but do not authorize them, so pair them with IAM, Cognito or Lambda authorizers.
Caching (REST only) stores responses per stage for a TTL (default 300 seconds, maximum 3,600).
Integration timeout is about 29 seconds by default for REST APIs and 30 seconds at most for HTTP APIs. Payloads are capped at 10 MB.
Direct service integrations send to SQS, start Step Functions or write to DynamoDB with no Lambda function in between.
Step Functions: Standard or Express
Standard
Express
Maximum duration
1 year
5 minutes
Execution semantics
Exactly-once
At-least-once (asynchronous) or at-most-once (synchronous)
Pricing
Per state transition
Per execution, duration and memory
History
Visual execution history for 90 days
CloudWatch Logs
Run a job (.sync) and wait for callback (task token)
Supported
Not supported
Typical use
Order and approval workflows, orchestrating ECS, Batch or Glue jobs, auditable processes
High-volume, short, idempotent event processing
Useful states: Choice (branch), Parallel (fixed branches at once), Map (same steps per item; Distributed Map fans out across millions of S3 objects), Wait (pause without paying for compute), plus Retry with exponential backoff and Catch for fallbacks. Human approval uses the callback pattern: the workflow passes a task token, for example in an email link, and pauses until SendTaskSuccess or SendTaskFailure arrives with that token.
Worked examples
Exam technique
Scan for numbers first. Durations over 15 minutes (Lambda), over 5 minutes (Express) or over about 30 seconds behind an API, payloads over 6 MB or 10 MB, and waits of hours or days eliminate options before you weigh the qualifier. A 504 from a slow API means “go asynchronous”: return a job ID, work in the background, let the client poll.
“LEAST operational overhead” + event-driven + short tasks ⇒ serverless. EC2, cron or self-hosted options are usually the “works but fails the qualifier” distractors.
Map the symptom to the control: cold starts ⇒ provisioned concurrency; noisy neighbour ⇒ reserved concurrency; “too many connections” ⇒ RDS Proxy; downstream overload ⇒ cap concurrency and buffer with SQS.
Remove glue code. An option where API Gateway or EventBridge calls SQS or Step Functions directly usually beats one that adds a forwarding Lambda function.
Common mistakes
Quick recap
Lambda: 15-minute maximum, up to 10,240 MB memory (CPU scales with it), 6 MB synchronous payload, no GPUs.
Async invocations retry twice and use on-failure destinations; SQS triggers use the queue’s DLQ and partial batch responses.
Reserved concurrency guarantees and caps capacity for free; provisioned concurrency removes cold starts for a fee.
VPC Lambda: private subnets + NAT gateway, gateway endpoints for S3 and DynamoDB, RDS Proxy for relational databases.
REST for usage plans, caching, WAF and private endpoints; HTTP for cheap JWT-protected APIs; WebSocket for server push.
Long work behind an API goes asynchronous.
Step Functions Standard: up to 1 year, exactly-once, callbacks and .sync. Express: 5 minutes, high volume, idempotent.