certslothcertsloth
DOP-C02/Topic 09

AWS / Professional

Event Automation and Resilience Engineering

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: An event is a trigger, not permission to repeat a destructive repair forever.

Must remember

  • EventBridge patterns select events; schedules run on time; target roles permit actions. Design retries, dead-letter handling, deduplication and event-age limits. Replays can repeat side effects, so action handlers need idempotency and current-state checks.
  • Step Functions Standard workflows support durable orchestration patterns; Express has different execution/history semantics. Choose retry/catch, timeout, heartbeat and callback behaviour explicitly. Distinguish a task failure from an entire workflow failure.
  • SSM Automation runbooks can coordinate remediation and approvals. Use concurrency/error thresholds, scoped document parameters and a tested rollback. A Config noncompliance event should not trigger a repair loop fighting an application deployment.
  • Reliability includes dependency timeouts, bulkheads, circuit breakers, queues and load shedding. Scale on a metric related to useful demand, such as backlog per worker, and account for warmup/cooldown and downstream limits. More workers can overload a database.
  • Recovery drills need measured RTO/RPO, restored keys/secrets/configuration, dependency ordering and traffic failback. Chaos experiments should have a hypothesis, bounded scope, stop conditions and observability. Test AZ/dependency failures rather than merely stopping a random server.
  • Track deployment frequency, lead time, change failure rate and recovery time alongside SLOs. Error budgets connect reliability evidence to release decisions. Post-incident reviews should improve controls and runbooks, not just add an alarm for the last symptom.

Choose under exam pressure

Requirement Choice and reason
Replay archived business events Idempotent consumers and explicit replay scope.
Repair thousands of resources Rate-limited runbooks with failure thresholds and approvals.
Queue backlog grows while database saturates Address downstream capacity and backpressure before adding unlimited workers.

Traps

  • Exactly-once-looking orchestration does not guarantee every external side effect occurs once.
  • An alarm threshold is not a service-level objective by itself.
  • Failover without a failback plan leaves recovery incomplete.

Active recall

1. Why should remediation check present state?

The event can be duplicated or stale, and another process may already have repaired the issue.

2. What should stop a chaos experiment?

Predefined safety thresholds or loss of required observability, not only the end of a timer.

3. Why can faster scaling reduce reliability?

New workers may overwhelm constrained dependencies.

4. What does a callback token pattern support?

A workflow waiting for an external process to report completion under controlled permissions and time limits.

5. Which evidence proves the RTO?

A timed restoration of useful application service, including dependencies and verification.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.