Memory hook: Find who invokes the function before deciding who owns retries.
Must remember
- Synchronous invokers receive a result/error and own their retry policy. Asynchronous Lambda invocation queues an event and has configurable retry/age handling and destinations. Event-source mappings poll sources such as SQS and streams; failure handling depends on the source. Never apply one retry rule to every invocation type.
- SQS messages remain hidden for the visibility timeout while processing. Set an adequate timeout for function execution and retry behaviour. A partial batch response can identify failed items so successful items need not be retried; the handler must implement the required response correctly. Poison messages need a bounded retry and DLQ strategy.
- Stream ordering follows shard/partition semantics. A failed batch can stall progress; investigate record age and use supported partial-batch, batch-bisection or discard settings deliberately. Consumers must tolerate duplicate processing and checkpoints/replays.
- Make handlers idempotent using a durable request key and conditional state transition. Reusing an execution environment permits connection reuse and cached static data, but request-specific identity or mutable state must not leak between invocations.
- Reserved concurrency caps/reserves function concurrency within account constraints. Provisioned concurrency keeps execution environments prepared for eligible versions/aliases. Memory allocation affects CPU availability; measure duration and total cost, not just configured MB.
- An execution role permits outbound AWS calls. Resource-based permissions authorise supported invokers. A VPC attachment provides network access, not internet access by itself: check routes, endpoints, DNS and security groups. Use an RDS Proxy when connection behaviour justifies it, rather than opening unbounded connections.
- SDK credential chains should obtain temporary role credentials. Handle pagination, service throttling and transient errors with bounded exponential backoff and jitter. Distinguish retriable failures from invalid requests and access denial. Reuse clients where safe; never log secrets or complete sensitive payloads.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| One failed SQS item reprocesses successful siblings | Implement supported partial batch failure reporting. |
| Warm latency is fine but cold starts breach the objective | Measure and evaluate provisioned concurrency or supported startup optimisations. |
| A successful API returns only part of a list | Follow pagination tokens. |
Traps
- A Lambda DLQ for asynchronous invocation is not the same as the source queue's redrive policy.
- A function in a public subnet does not automatically receive public internet access.
- Retries without idempotency can duplicate external actions.
Active recall
1. Who retries a synchronous invocation?
The caller or calling service according to its policy; do not assume Lambda queues it asynchronously.
2. Why can a short visibility timeout cause duplicate work?
The message can become visible before processing finishes, allowing another consumer to receive it.
3. Does reserved concurrency remove cold starts?
No. It controls concurrency allocation; provisioned concurrency addresses prepared environments.
4. A function can reach RDS but times out calling a public API. What should be checked?
Its VPC outbound path, including NAT or appropriate endpoints, DNS and security rules.
5. What should a repeated payment event do?
Return or recover the recorded result using a durable idempotency key, rather than charge again.