Memory hook: Work from the symptom to the failing layer; prove recovery with a user-visible result.
Must remember
- EC2 system status failures point toward underlying AWS infrastructure; instance status failures can indicate guest/network configuration. Inspect instance events, console output and OS logs. Reboot, recover, stop/start and replacement have different identity and storage consequences.
- EBS bottlenecks involve IOPS, throughput, queue depth and instance limits. Burstable instance credits can constrain sustained CPU. Database latency may come from locks, connections, slow queries, memory or storage; adding CPU blindly can miss the cause.
- For an unhealthy load-balancer target, verify health-check path, port, protocol, expected response, application binding and security-group chain. Check ASG health-check type, grace period and warmup before increasing capacity. Lifecycle hooks need completion or a deliberate timeout outcome.
- A failed connection needs both forward and return paths. Check address family, DNS answer, route tables, internet/NAT gateways, security groups, stateless NACL ephemeral ports and endpoint policies. Flow Logs report accepted/rejected traffic metadata; Reachability Analyzer reasons about supported network configuration rather than sending a real application request.
- For CloudFront stale or denied content, inspect cache policy, origin access, TLS, origin health and error caching. DNS TTL means an updated record may not instantly change every client. Regional/private-zone resolution may differ from a public lookup.
- Backups should be restored into a test environment and verified for data, permissions, encryption-key access and application startup. Record RPO as allowable data loss and RTO as time to restore useful service. Replication can copy deletion or corruption; it is not a substitute for history.
- Use EventBridge/alarms to start bounded, observable repair workflows. Suppress duplicate actions, preserve evidence, escalate unknown failures and test rollback. A green infrastructure dashboard is insufficient if customers still cannot complete a transaction.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Instance runs but ALB marks it unhealthy | Validate the actual health-check request and response. |
| Only large packets fail over a hybrid link | Investigate MTU/path MTU and permitted control traffic. |
| Backups exist but recovery time is unknown | Run and time a realistic restore drill. |
Traps
- NACLs need return traffic; security groups track established flows.
- A larger fleet cannot fix all requests hitting an unavailable database.
- A backup without its required KMS access may not be recoverable by the intended account.
Active recall
1. Which status check is more closely associated with host infrastructure?
EC2 system status.
2. Why might healthy new instances immediately be replaced?
Incorrect health-check settings or insufficient startup/grace timing can classify them as unhealthy.
3. What does a rejected Flow Logs record prove?
That the recorded traffic was rejected at the observed network interface path; investigate the applicable controls and context.
4. Why restore instead of merely listing backups?
Only a restore exercise tests data usability, dependencies, permissions and recovery time.
5. An alarm fires repeatedly during recovery. What should automation prevent?
Conflicting duplicate repairs; use idempotency, state checks and bounded execution.