certslothcertsloth
DVA-C02/Topic 09

AWS / Associate

Observability, Debugging and Optimisation

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Trace the slow request, distinguish its dependency, then measure the fix.

Must remember

  • Emit structured logs with request/correlation IDs and meaningful severity. Metrics describe rates, durations and saturation; traces follow work across components. Use CloudWatch Logs Insights for focused queries and tracing instrumentation for dependency latency. Redact credentials and personal data.
  • Instrument useful business outcomes as well as infrastructure health. A queue backlog may reveal a consumer bottleneck while CPU is low. Record failure categories, retries and downstream latency to separate cause from symptom.
  • For AccessDenied, identify the caller, action, resource, condition context and every applicable policy layer, including KMS or endpoint policies. For timeouts, check DNS, routes, security rules, connection limits and downstream health. For throttling, inspect quotas, concurrency, partition distribution and client retry behaviour.
  • Diagnose Lambda using errors, duration, throttles, concurrency and event-source age. Diagnose DynamoDB using throttled operations, key distribution, consumed capacity and query patterns. Diagnose APIs using integration latency versus total request latency, status codes and authorised/unauthorised request behaviour.
  • Caching changes freshness and invalidation requirements. Reuse safe connections, batch eligible operations and avoid repeated full-table scans. Queue long-running work when an interactive response need not wait. Tune memory or capacity using representative load tests and percentile latency.
  • A trace sample is not every request; missing instrumentation is not proof that a dependency was never called. Debug with a hypothesis, a narrow observation and a measurable acceptance condition. Preserve a rollback path and compare cost per successful operation after the change.

Choose under exam pressure

Requirement Choice and reason
Find which dependency consumed most request time Distributed trace with instrumented spans.
Find repeated errors for one correlation ID Structured logs and a scoped query.
Traffic spike creates retries and more throttling Bound retries with jitter and address the actual capacity/access-pattern limit.

Traps

  • Average latency can hide a bad tail.
  • Increasing timeout can conceal a failed dependency without fixing it.
  • Logging an access token creates a credential exposure.

Active recall

1. Why use a correlation ID?

To connect evidence for one operation across services without relying on timestamps alone.

2. A table throttles despite spare aggregate capacity. Why?

A hot partition/key or operation-specific limit can bottleneck before overall capacity is exhausted.

3. What distinguishes a trace from a log entry?

A trace represents related spans and timing across an operation; logs record individual events/details.

4. Why benchmark before and after a change?

To verify the suspected cause and quantify quality, latency and cost effects.

5. Should every error be retried indefinitely?

No. Classify transient versus permanent failures, bound retries and handle final failure explicitly.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.