certslothcertsloth
← SOA-C03 overview

CloudOps Engineer Associate / STUDY TOOLS

SOA-C03 quick review

Reviewed 10 October 2026 · CloudOps Engineer Associate

Memory hook: Observe, locate the failing layer, repair safely, verify the user outcome.

Monitoring and optimisation 22%; reliability 22%; provisioning and automation 22%; security 16%; networking 18%.

Use this as a final revision pass after the chapters. Each task below maps to the published exam outline; the outline itself is not an exhaustive list of possible questions. Recheck the official guide for your booked exam version, especially beta releases.

Must remember by exam objective

1.1 — Metrics, alarms and filters

  • CloudWatch metrics/alarms observe performance; Logs Insights queries logs; metric filters turn matching logs into metrics; subscriptions forward logs. The CloudWatch agent supplies guest memory/disk/custom telemetry. Composite alarms combine alarm states. Configure missing-data behavior and permissions deliberately.

1.2 — Diagnose and remediate operational faults

  • Distinguish EC2 system status from guest instance status, application health from target health and symptom from dependency failure. Check stack/events/logs and correlate errors before remediation. EventBridge/SSM automation needs safe inputs, scoped roles, duplicate handling and verified completion.

1.3 — Optimise compute, storage and databases

  • Measure CPU/credits, memory, EBS IOPS/throughput/queue depth, network, database locks/connections and query plans. Rightsize the bottleneck; a bigger application tier does not fix a locked database. Storage class, caching and data-transfer changes must preserve required performance and recovery.

2.1 — Elastic capacity

  • Use launch templates, ASG desired/min/max capacity, suitable scaling metrics, warmup and lifecycle hooks. ELB health checks and ASG replacement settings differ. Graceful draining protects active requests; scheduled/predictive capacity and reactive policies serve different load patterns.

2.2 — Resilient environments

  • Spread useful capacity and dependencies across AZs. Validate target health, database failover mode, connection retry behavior and quotas. Multi-AZ regional availability differs from regional disaster recovery. Replacement, recovery and stop/start have different identity/storage implications.

2.3 — Backup and restore

  • Define RPO/data loss and RTO/restoration time. EBS snapshots/AMIs, database backups/PITR and AWS Backup have resource-specific restore/copy behavior. Include KMS keys, roles and configuration in restore tests. Replication can copy corruption, so retain required historical recovery.

3.1 — Provision and maintain resources

  • CloudFormation parameters, mappings, conditions, outputs and dependencies shape provisioning. Change sets preview; drift detects supported differences; StackSets scale across accounts/Regions. Understand update/replacement, rollback, DeletionPolicy and UpdateReplacePolicy. Container operations need image access, task/node roles, health and networking.

3.2 — Automate resource operations

  • SSM nodes need agent, identity and endpoint connectivity. Session Manager handles supported sessions, Run Command actions, State Manager desired associations, Automation runbooks and Patch Manager patch policy. Maintenance windows schedule; patch baselines select approval. Rate-limit fleet operations and report failed nodes.

4.1 — Security policies and compliance

  • Identify actual identity, trust, grants and ceilings such as SCPs/boundaries. Control iam:PassRole and service roles. Config evaluates supported resource compliance; CloudTrail audits activity. Findings do not fix resources unless a remediation is configured; emergency access must be deliberate and auditable.

4.2 — Protect data and infrastructure

  • Separate transport encryption, storage encryption and key access. Secrets Manager rotation and Parameter Store configuration have distinct purposes. Use least privilege, private paths, scoped SGs and suitable WAF/GuardDuty/Inspector/Macie controls. Test restore access before scheduling deletion of a needed key.

5.1 — Network connectivity and efficiency

  • Routes and address families establish paths; SGs are stateful, NACLs stateless. Public IPv4 needs addressing, IGW route and permission; private outbound often uses NAT or service endpoints. Peering is nontransitive; Transit Gateway needs attachment and subnet routes. Compare cross-AZ/NAT/endpoint costs.

5.2 — DNS and content delivery

  • Route 53 TTL and negative caching delay observed DNS changes. Private-zone association and Resolver rules determine answers. CloudFront OAC secures eligible S3 origins; signed viewer access is a different boundary. Inspect cache key, origin request policy, TLS, error caching and origin health.

5.3 — Network fault diagnosis

  • Trace DNS answer, forward and return routes, SG/NACL rules, endpoint policy, app listener and TLS. Flow Logs give flow metadata; Reachability Analyzer reasons about supported configuration, not a real application response. Small packets working while large packets fail suggests MTU/path-MTU issues.

Choose under exam pressure

Deciding clue Recall the distinction
EC2 running but ALB unhealthy Check exact health path/port/protocol/status and SG chain.
SSM session unavailable Check agent, instance role, required endpoints/DNS and logs.
CloudFormation rollback Find earliest meaningful event; final rollback is often a consequence.
Current state differs from template Drift detection; no automatic repair implied.
Backups exist but recovery unknown Restore and verify data, key access, dependencies and measured time.

Traps

  • An alarm notification does not guarantee remediation.
  • A public subnet does not give Lambda a public IP.
  • Stopped instances can retain charged storage and other resources.
  • Deleting a stack can intentionally retain resources.

Verification cues

  • Read EC2 status checks and console output, target-health reason codes, ASG activity and CloudFormation events in that order for a failed launch.
  • For one network fault, draw both paths and compare them to Flow Logs; test using the same resolver context as the workload.
  • Explain the readiness signal that proves an application is useful after restore, beyond resource-running status.

Last-pass active recall

1. Which is data loss: RPO or RTO?

RPO. RTO is the target elapsed restoration time.

2. Does a security group need an explicit return-port rule?

Established response traffic is statefully tracked; NACLs need matching stateless return permissions.

3. Can a service role do what the initiating user cannot?

Yes, within its own permissions; control who may pass/use that role.

4. Why might rising EBS size not fix latency?

The bottleneck may be IOPS, throughput, instance limits or application behavior.

5. Why can monitoring appear healthy during an outage?

The telemetry path may be broken or the checks may not observe the failing business operation.

Sources and version check

The numbered chapters provide worked distinctions and further technical sources. These are original revision notes and original recall scenarios, not real exam questions.

Every topic at a glance

Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.

01 · Monitoring & Audit

Memory hook: Metrics show symptoms, logs explain events, traces follow requests, and audit records identify changes.

Must remember

Separate the evidence questions

  • CloudWatch: how is the workload behaving? Use metrics, logs, dashboards and alarms for operational evidence. CloudTrail: which identity called which API, against what resource and when? AWS Config: what resource configuration existed, and did an evaluated rule consider it compliant?
  • These sources complement each other. A slow API can require a latency alarm, application logs and a trace; identifying an administrator's change requires audit evidence. Check that the needed events, resources and retention were actually configured before promising historical answers.
  • AWS X-Ray follows instrumented requests across services and downstream calls. Its trace map helps find latency, errors and bottlenecks. Traces are not a replacement for every application log or API audit event; instrumentation and sampling affect visibility.

Metrics and logs

  • A metric is a numeric time series identified by namespace, name and dimensions. Choose meaningful statistics and evaluation periods: average latency can conceal slow tail requests, while a total error count without request volume may mislead.
  • Alarms evaluate metric conditions; actions notify or invoke supported responses. Composite alarms combine alarm states to reduce noisy paging. Treat missing data intentionally rather than assuming missing means healthy. Metric streams continuously deliver selected metric updates to downstream consumers.
  • Logs Insights queries log events. Metric filters count matching events into metrics. Subscriptions forward matching logs to supported destinations; export writes log data to S3 for a different processing workflow. These are distinct mechanisms.
  • The CloudWatch agent collects additional guest-OS and application signals. Standard EC2 metrics do not automatically reveal every filesystem or memory measurement.
  • Container Insights and Lambda Insights add workload-specific visibility; Contributor Insights identifies prominent contributors; Application Insights helps correlate application problems. Their scope and collection costs need deliberate configuration.

Events, audit and compliance

  • EventBridge rules match events on buses and deliver them to targets. Scheduler invokes targets on time-based schedules. An archive retains selected events for replay; a replay can repeat a business action, so consumers still need idempotency.
  • A successfully accepted event can fail to match a rule or fail later delivery. Check source/detail pattern, bus, target configuration, resource permissions and retry/dead-letter behavior separately.
  • CloudTrail management events describe control-plane activity; selected data events provide supported resource-level activity. CloudTrail Insights detects unusual supported API activity. A rule reacting to an API call via EventBridge needs the appropriate event path and coverage.
  • Config recording tracks selected resource configurations; rules evaluate compliance. Notifications and remediation are separately configured. A remediation role can mutate resources, so evaluation should not be confused with automatic repair.

Additional published-scope tools

  • Amazon Managed Service for Prometheus stores and queries compatible operational metrics, especially for container workloads using PromQL. Amazon Managed Grafana visualizes metrics, logs and traces from multiple sources. The dashboard layer is different from the metric storage/query layer.
  • AWS Health Dashboard reports AWS service events and account-relevant impacts. Combine it with workload telemetry: a healthy AWS status does not prove your application or configuration is healthy.
  • Keep logs and traces useful: redact sensitive data, set retention, scope collection and correlate request identifiers. Broad logging can create both sensitive-data exposure and substantial ingestion charges.

Choose under exam pressure

Clue in the requirement Choose or investigate
Alert on errors or latency CloudWatch metric alarm with an action
Determine who deleted a database CloudTrail audit events
Review a resource's historical configuration compliance Config history and rule evaluations
Find the slow downstream call in a distributed request X-Ray tracing
Query container metrics using PromQL Managed Service for Prometheus
Visualize several telemetry sources together Managed Grafana
Trigger work for matching application events EventBridge rule and target
Investigate a relevant AWS service disruption AWS Health Dashboard

Traps

  • An alarm with no action is not an email subscription; rule creation alone does not establish every permission needed for delivery.
  • A trace sample or a log metric is not a complete security audit record.
  • A Config rule can report a problem without correcting it. Replaying an event can repeat side effects rather than merely replaying a picture of history.

Practise this topic

02 · ELB & Auto Scaling

Memory hook: A load balancer chooses a destination; an Auto Scaling group maintains capacity; application state must survive either decision.

Must remember

Pick the right traffic layer

  • Application Load Balancer (ALB) understands HTTP/HTTPS at layer 7. A listener receives traffic; ordered rules choose forward, redirect, authentication or fixed-response actions. Conditions can include host and path.
  • Target groups hold instances, IPs or supported Lambda destinations, with health checks/attributes. Path rules can select different groups.
  • Lower numbered listener priorities run first. A fixed response comes from the ALB and can succeed even when the application is unavailable.
  • Network Load Balancer (NLB) handles TCP/UDP/TLS at layer 4 and supplies static addresses per enabled AZ, with optional Elastic IPs for supported internet-facing deployments. It fits non-HTTP traffic and IP allowlist requirements.
  • NLB is not an HTTP path router. Modern NLBs support security groups in supported configurations; the old claim that NLBs never have security groups is unsafe.
  • Gateway Load Balancer (GWLB) steers traffic through network appliances using GENEVE/UDP 6081, rather than routing application URLs.
  • An internet-facing load balancer can forward to private backends; public users do not require public backend IPv4.

Connection behavior and health

  • Stickiness uses supported cookies/affinity mechanisms to favor a target. It does not replicate session memory or guarantee that target will survive.
  • Cross-zone load balancing allows a node to route across enabled AZs. ALB has it enabled at the load-balancer level, with target-group-level controls; NLB/GWLB default differently. Check transfer cost rules for the specific type.
  • Deregistration delay drains in-flight requests during removal. Too little can interrupt long work; excessive values can slow scale-in and deployments.
  • TLS termination uses listener certificates, often from ACM. SNI lets a compatible client indicate its hostname so a listener chooses the correct certificate. ALB certificates are regional; CloudFront ACM certificates use us-east-1.
  • Target health checks detect application reachability on the configured port/path. EC2 status health and application health answer different questions.
  • Health routing is not authorization: all-unhealthy/fail-open behavior can send traffic to unhealthy targets.

Scaling and replacement

  • Vertical scaling changes machine size and may require interruption; horizontal scaling changes the number of workers. Externalized state makes horizontal replacement safer.
  • AWS Auto Scaling plans coordinate resources. EC2 Auto Scaling manages instance groups; Application Auto Scaling handles supported dimensions such as ECS task counts. Plans can migrate to direct policies.
  • Launch templates describe AMIs, type, user data, interfaces and related launch settings. Updating a template does not automatically update every already-running instance.
  • An ASG maintains desired capacity between minimum and maximum bounds. Enable appropriate ELB health when application failures should cause replacement; EC2-only checks may miss a broken web process.
  • Instance refresh rolls out launch changes with health, warmup and capacity constraints.
  • Target tracking maintains a chosen metric target; step scaling changes capacity according to alarm severity; scheduled scaling anticipates known times; predictive scaling forecasts recurring demand.
  • Choose a demand-related metric: CPU can fit compute-bound servers, ALB request count per target can fit web workers, and queue backlog per worker can fit asynchronous processing.
  • Warmup/cooldown reduce unstable decisions during startup; grace periods do not prove readiness.
  • Multi-AZ subnets with desired capacity one do not provide two active replicas. Production resilience needs sufficient surviving capacity and dependencies.

Choose under exam pressure

Requirement Decision and reason
Several web apps under paths or hostnames ALB listener rules and target groups
Static addresses for TCP/UDP clients NLB
Third-party network inspection fleet GWLB
Scale before a known daily opening Scheduled scaling
Maintain utilization near a target Target tracking
Replace instances with a broken web process ASG using relevant ELB health
Roll out a new AMI to existing capacity Controlled instance refresh

Traps

  • Scaling and healing differ. Replacing an unhealthy instance can preserve the same desired capacity without adding demand capacity.
  • A cookie is not a session database. Target loss still destroys target-local state.
  • A successful listener test can bypass dependencies. Fixed responses do not prove backend, cache or database health.

Practise this topic

03 · EC2 Instance Storage

Memory hook: Choose block, file or local scratch first, then check its failure boundary, performance and recovery behavior.

Must remember

EBS and instance store

  • EBS is persistent block storage in one AZ. The attached EC2 instance must be in the same AZ. It can outlive an instance, but delete-on-termination settings determine each volume's lifecycle.
  • Root/data volumes can have different deletion behavior. Stopped instances and unattached EBS volumes can retain storage charges.
  • gp3 separates capacity from configurable IOPS/throughput and suits general workloads. gp2 performance depends on size and burst behavior. More GiB is not always the best way to solve an I/O bottleneck.
  • io1/io2 Provisioned IOPS SSD suits demanding transactional random I/O and predictable performance. st1 and sc1 HDD target sequential throughput and colder sequential data; they are not boot-volume choices.
  • Volume and instance EBS limits both constrain performance; more provisioned IOPS cannot bypass instance bandwidth.
  • Multi-Attach is supported for eligible io1/io2 volumes and compatible Nitro instances in the same AZ, with engine/OS/Region restrictions. It is not supported for gp3 or as a boot-volume shortcut.
  • Concurrent block writes require a cluster-aware filesystem/application and correct coordination/fencing. Ordinary filesystems mounted read-write from unrelated servers can corrupt data.
  • Instance store is host-local and ephemeral on host loss, stop or termination, unlike reboot. Use it for reconstructable cache, scratch or replicated data, never the only durable copy.

Snapshots, images and encryption

  • Standard EBS snapshots store incremental changed blocks. AWS manages dependencies, preserving later snapshots' restorability when older snapshots are deleted.
  • Restore a snapshot into a new volume to change AZ; copy snapshots for cross-Region recovery.
  • An EBS-backed AMI adds boot/launch metadata and references image snapshots. Deregistering the AMI and deleting snapshots are distinct cleanup decisions.
  • Snapshot-created volumes can require initialization for full performance. Fast Snapshot Restore removes that initialization penalty for enabled snapshot/AZ combinations, with additional charges. It does not make the data newer.
  • Snapshot Archive converts an incremental snapshot into a full archived snapshot. It trades lower long-term storage rates for retrieval delay, retrieval costs and a minimum storage duration; it is poor short-lived lab storage.
  • Recycle Bin retains eligible deleted snapshots/AMIs for recovery, delaying final deletion and potentially retaining charges.
  • KMS encryption covers EBS data and associated snapshots. Copying an unencrypted snapshot into an encrypted one supports migration; an existing unencrypted volume does not simply gain in-place encryption.
  • Sharing an encrypted snapshot requires appropriate snapshot permissions and access to a suitable customer-managed KMS key. AWS-managed default keys cannot be shared as though they were your own cross-account keys.

EFS: files instead of blocks

  • EFS is managed NFS for concurrent Linux filesystem access. Regional EFS distributes storage across AZs; One Zone has a different failure boundary.
  • Mount targets provide network access in VPC subnets. NFS reachability, security groups, file permissions and any configured IAM access controls all matter.
  • Elastic throughput adapts to demand; Provisioned throughput sets a chosen throughput level; Bursting links available throughput/credits to the storage workload. These are distinct from storage classes.
  • Standard, IA and Archive classify data by access/storage economics. Lifecycle policies can move cold files; retrieval and minimum-duration charges can outweigh savings for short retention.
  • EFS fits shared uploads/home directories, not every block-database workload; storage integrations covers specialized Windows/HPC filesystems.

Choose under exam pressure

Requirement or clue Decision and reason
Persistent instance boot disk Supported SSD EBS volume
Predictable high random transactional I/O Provisioned IOPS SSD plus adequate instance capability
Large sequential scans at lower cost Appropriate HDD EBS type
Linux shared uploads across AZs Regional EFS
Rebuildable high-performance scratch Instance store
Repeatable bootable server image AMI
Recover a disk in another AZ Snapshot restore to a new volume
Long-retained rarely restored backup Evaluate archive retrieval and minimums

Traps

  • Multi-Attach is not managed shared NFS. Shared blocks still require write coordination and remain in one AZ.
  • A copy is only as current as its recovery point. More snapshots or faster initialization does not inherently improve the latest available RPO.
  • Deleting compute is incomplete cleanup. Volumes, snapshots, AMIs, retention policies and filesystem data each have separate lifecycles.

Practise this topic

04 · RDS, Aurora & ElastiCache

Memory hook: Replicas add read capacity, failover preserves service, backups recover history, proxies manage connections, and caches avoid repeated work.

Must remember

RDS: identify the actual bottleneck

  • RDS provides managed relational engines. AWS operates much of the platform; you choose sizing, access and retention. Ordinary RDS provides no general guest-OS administration.
  • RDS Custom allows supported OS/database customization within automation boundaries. Verify engine/platform fit; use EC2 when full control is essential.
  • CPU/memory, storage, IOPS/throughput and connections are different bottlenecks; more disk does not fix CPU-heavy queries.
  • Storage autoscaling increases storage toward a configured maximum; it does not shrink it. Monitor both consumption and the cap rather than treating it as unlimited.
  • A DB subnet group describes eligible subnets; a parameter group configures engine settings. Neither creates running capacity or HA.

Read scaling, availability and recovery differ

  • A classic Multi-AZ DB-instance deployment synchronously replicates to a standby for automatic failover. That standby does not serve application reads.
  • A Multi-AZ DB cluster has a different architecture with readable standby instances. Always identify the deployment type before applying the shortcut “Multi-AZ is not for reads.”
  • Read replicas principally offload reads, generally using asynchronous replication. Lag matters for read-after-write requirements. Promotion/cutover is a separate decision from the usual read-scaling purpose.
  • Automatic failover changes the serving database; clients still need reconnection/retries and handling for interrupted transactions.
  • Automated backups and transaction logs support point-in-time recovery (PITR). Restore produces a new database resource requiring a cutover; it does not undo selected changes in the running database.
  • Manual snapshots persist until deleted and can support copying/sharing workflows. Encryption/key access and regional transfer/storage charges remain relevant.
  • Replicas may reproduce accidental deletes or corrupt application writes. Replication is not backup history. RPO asks how much data loss is acceptable; RTO asks how long recovery may take.

Aurora: storage, endpoints and capacity

  • Aurora separates compute instances from distributed cluster storage spanning multiple AZs. A one-instance Aurora cluster still has distributed storage, but adding an appropriately placed reader improves compute failover options.
  • The writer/cluster endpoint follows the current writer. The reader endpoint distributes new connections among available readers; it does not redistribute every query inside one existing connection.
  • Instance endpoints target individual instances. Custom endpoints select instance subsets, useful when reporting should use a different capacity group from ordinary application reads.
  • Replica Auto Scaling changes the number of readers. Aurora Serverless v2 adjusts compute capacity. Neither should be confused with scaling the writer's count.
  • Serverless v2 auto-pause/zero-capacity requires eligible versions/configuration; idle cost is not universally zero. Serverless v1 is retired; see service availability.
  • Aurora Global Database uses cross-Region replication for global reads and recovery. It is a regional-disaster design, distinct from adding another reader in the same Region; asynchronous replication can imply unreplicated-write risk.
  • Aurora cloning uses copy-on-write storage sharing for fast development/test copies. Shared unchanged pages and divergent writes differ from independent regional disaster-recovery copies.
  • Babelfish for Aurora PostgreSQL supports many SQL Server interfaces to reduce application migration changes, but compatibility assessment is essential.
  • Aurora ML integrations expose supported ML capabilities from database workflows; they do not turn every database engine into a model-training service.

Secure connections and protect capacity

  • KMS protects data at rest, TLS protects data in transit, and supported IAM database authentication provides temporary authentication tokens. Database users/grants still authorize SQL operations.
  • Security groups control network reachability. Permit the application SG on the required engine port; a private subnet alone is not a complete authorization design.
  • Ports: PostgreSQL/Aurora PostgreSQL 5432, MySQL/MariaDB/Aurora MySQL 3306, SQL Server 1433, Oracle 1521, Redis OSS/Valkey 6379, Memcached 11211.
  • RDS Proxy pools/reuses database connections and can help manage failover and connection surges. It is not a query-result cache and does not remove all database capacity limits.
  • IAM authentication and Secrets Manager solve different integration needs; see encryption and secrets.

ElastiCache and cache correctness

  • Redis OSS/Valkey offer richer structures such as sorted sets, with replication, failover and persistence capabilities depending on deployment.
  • Memcached is a simpler multithreaded distributed cache. Do not assume every modern Serverless feature matches a classic node-based engine comparison; choose the actual deployment's capabilities.
  • Lazy loading/cache-aside: read cache, fetch the database on a miss, then populate. It avoids caching never-read objects but adds miss latency and can return stale data.
  • Write-through: update the cache alongside database writes. It improves hit availability at extra write work, but still requires a coherent failure/invalidation strategy.
  • TTL expires entries. Stagger expiry or use appropriate request coordination to avoid a cache stampede when many entries expire together.
  • An external session store makes web servers replaceable. Required session durability and failover must still be designed; stickiness alone does not preserve lost memory.
  • Cache authentication/user controls, supported IAM authentication, TLS, and SGs address distinct layers. A cache is not safe merely because it has no public endpoint.

Choose under exam pressure

Requirement or symptom Decision and reason
Primary overloaded by tolerant reporting reads Read replicas; account for lag
Automatic recovery from AZ failure Appropriate Multi-AZ deployment
Recover before an accidental update PITR or suitable snapshot
Burst of short-lived Lambda connections RDS Proxy
Repeated expensive reads or fast shared sessions Appropriate cache strategy
Isolate Aurora analytics readers Custom endpoint and selected capacity
Variable Aurora compute demand Serverless v2, with supported limits/cost model
Database recovery after regional loss Cross-Region copies or Global Database design

Traps

  • “Multi-AZ” is not one architecture. Distinguish classic DB-instance standbys from readable DB-cluster instances.
  • Encryption is not one switch. At-rest KMS, network TLS, identity authentication and SQL privileges solve different problems.
  • Caches and replicas copy failures too. Neither automatically supplies a clean historical recovery point.

Practise this topic

05 · Disaster Recovery & Migrations

Memory hook: RPO limits how much data may be lost; RTO limits how long useful service may be unavailable.

Must remember

Objectives before recovery patterns

  • Recovery point objective (RPO) measures acceptable data loss as a time interval. Recovery time objective (RTO) measures acceptable recovery duration. Fast restoration cannot recreate transactions that never reached the available recovery data.
  • Measure detection, decision, provisioning, data recovery, dependencies, traffic switching and validation. A database restore benchmark alone does not prove users can work within the RTO.
  • Backup and restore: retain recovery data, then rebuild as needed. Pilot light: keep the minimal core running and add missing serving capacity. Warm standby: maintain a functional reduced-capacity workload and scale it. Multi-site active: serve from multiple sites while managing routing, consistency and failure behavior.
  • Recovery patterns trade standing cost and complexity against recovery work; their names do not guarantee an RPO/RTO. Select using tested behavior and actual requirements.
  • Availability across AZs and regional disaster recovery solve different failure scopes. Multi-AZ replication does not automatically protect against a Region-wide outage, accidental deletion or corrupted application data.

Backups need more than a schedule

  • AWS Backup plans define rules; resource selections determine what is protected; vaults contain recovery points. A plan with no selections schedules no useful protection for the intended resource.
  • Cross-Region or cross-account copies require compatible resource support, IAM and encryption-key arrangements. Verify the copy, its retention and the destination's ability to restore. A copied recovery point that cannot be decrypted is not a working recovery plan.
  • Vault Lock provides retention controls; compliance-mode protection can intentionally prevent early deletion. Governance controls and compliance immutability have different bypass properties. This pack does not create retention that obstructs immediate cleanup.
  • Keep historical recovery points where required: replication can quickly copy an unwanted deletion or bad write. Backup retention, replication and application validation protect against different failure modes.
  • Validate service quotas, instance availability, subnet addresses, dependencies and key access in the standby Region before a disaster. A successful Terraform plan does not reserve all future capacity or raise every quota.

Migrate with a deliberate cutover

  • DMS supports data movement with full load and change data capture (CDC) for supported endpoints. SCT helps convert schema/code for heterogeneous migrations; data replication does not automatically convert every stored procedure or engine feature.
  • A low-downtime database migration commonly uses initial load, ongoing CDC, validation, a controlled write cutover and a rollback decision. Monitor replication lag and reconcile data; “CDC enabled” is not proof of zero data loss.
  • RDS/Aurora paths include compatible dump/restore, snapshots, replicas and DMS. Check engine/version support and downtime tolerance before selecting a path.
  • VM Import/Export moves supported VM images. Application Migration Service (MGN), now documented as AWS Transform MGN, replicates servers for rehosting and cutover, with staging resources and costs. Rehosting differs from moving only a database.
  • Application Discovery Service is a historical inventory/discovery tool closed to new customers, not a universal replacement for MGN. VMware Cloud on AWS preserves VMware-based operating assumptions but requires current commercial availability and capacity checks. It remains named in the published exam list; that does not authorize a new subscription in this lab.
  • For large datasets, estimate effective transfer time from data size and throughput, including validation and changes during transfer. Compare DataSync, suitable uploads and supported alternatives; transfer/storage notes cover the tool boundaries. Snow Family remains a published exam concept despite lifecycle changes; no physical job is created here. Consult service status.

Choose under exam pressure

Clue in the requirement Choose or investigate
Restore rarely, tolerate substantial recovery work Backup and restore
Keep critical data/core services ready but provision serving tiers later Pilot light
Need a functional secondary environment before failure Warm standby
Minimize database migration downtime Full load plus CDC, validation and controlled cutover
Rehost whole servers with limited rewriting MGN
Convert between database engine families Schema assessment/conversion plus a compatible data migration
Centralize protection across supported resources AWS Backup plans, selections and restore tests
A standby cannot scale during disaster Inspect quotas, capacity, address space and dependencies

Traps

  • Faster compute during restore improves no data that is absent from the backup.
  • A replica is not automatically a historical backup, and a backup is not automatically immediately available serving capacity.
  • Retention locks deliberately constrain deletion. Terraform lifecycle flags cannot bypass a genuine compliance retention requirement.

Practise this topic

06 · Security & Encryption

Memory hook: Protect the connection, the stored data, the permission to use it and the evidence of misuse separately.

Must remember

Encryption and key control

  • TLS protects transit; encryption at rest protects stored data. Neither prevents an already-authorized compromised application from reading plaintext. Authentication, authorization, secret handling and monitoring remain necessary.
  • KMS: AWS-owned keys are managed within services; AWS-managed keys are visible in your account but have service-controlled administration; customer-managed keys provide your own policy/lifecycle control. Symmetric encryption, asymmetric operations and HMAC keys solve different cryptographic tasks.
  • A KMS key policy is central to authorization. IAM permissions alone are not a universal substitute for a suitable key policy. Supported rotation keeps older material available for decrypting existing ciphertext; rotating a key does not automatically re-encrypt every stored object.
  • Multi-Region KMS keys share related key material, enabling supported regional cryptographic use, but policies, grants, aliases and lifecycle remain regional decisions. Creating a replica does not copy every administrative setting or automatically replicate application data.
  • Encrypted snapshot/AMI sharing needs resource permissions and appropriate customer-key access for the recipient. S3 replication of SSE-KMS objects needs explicit replication configuration, source decryption and destination encryption permissions with the correct destination key. A generic S3 copy policy is insufficient.
  • CloudHSM provides dedicated hardware security modules and more direct cryptographic control, with greater administration/capacity responsibility. Choose it for a requirement that specifically needs that control or interface; ordinary managed encryption requirements often fit KMS better.

Configuration, secrets and certificates

  • Parameter Store provides hierarchical configuration, Standard/Advanced tiers and KMS-backed SecureString. Secrets Manager supplies secret versions, supported rotation workflows and optional regional replication. Rotation requires the relevant integration and permissions; merely storing a secret does not rotate a database password.
  • Applications should retrieve secrets using a role, with caching and refresh behavior appropriate to rotation. Terraform's sensitive flag controls some display behavior; it does not encrypt local state or prevent an authorized reader from recovering supplied values.
  • ACM manages certificates. An ALB uses a certificate in its Region; CloudFront's ACM certificate must be in us-east-1. Validate domain ownership and consider the whole client-to-edge-to-origin TLS path rather than securing only one connection.
  • AWS Private CA is an adjacent distinction: it issues certificates for a private trust hierarchy, such as internal services. Private certificates are not automatically trusted by public browsers, and a private CA introduces charges. It is not separately named in the current in-scope list, unlike ACM.

Filtering, detection and investigation

  • WAF filters supported HTTP requests with web ACLs, IP sets and rules, including rate-based rules. Shield Standard supplies baseline DDoS protection; Shield Advanced adds paid capabilities. Firewall Manager centrally manages supported security policies across an organization. None replaces least privilege or secure application logic.
  • DDoS resilience combines edge absorption, caching, rate controls, suitable scaling and protected origins. Keep expensive origin work from being the first line of defense. Network Firewall handles network inspection; WAF targets supported web request paths.
  • GuardDuty detects suspicious activity. Inspector finds vulnerabilities in supported workloads. Macie discovers sensitive data in S3. Select based on the finding needed, not the generic word “security.”
  • Security Hub, including its security-posture capabilities, consolidates findings and evaluates supported security controls. Detective helps investigate relationships and activity surrounding suspicious behavior. Aggregating a finding, investigating it and automatically remediating it are different steps.
  • Artifact provides AWS compliance reports and agreements. It does not certify your application's configuration. Audit Manager can collect and organize evidence for assessments; it is useful adjacent context rather than an explicitly named service in the current list, and it does not replace the auditor's judgment.
  • Modern Inspector remains relevant; Inspector Classic is retired. Consult service status for generation-specific dates. Never infer that a current service is unavailable solely because an older namesake ended support.

Operational boundaries

  • Shared responsibility changes with the service: AWS operates underlying infrastructure, while you still control data classification, identities and workload configuration. Managing EC2 also includes guest-OS responsibilities that a fully managed service takes off your hands.
  • Choose retention deliberately. Customer-key deletion has a waiting period; secret recovery settings and replicas affect deletion; immutable compliance retention can intentionally prevent removal. Such retention is valuable when required by a real workload and incompatible with this disposable lab's default.

Choose under exam pressure

Clue in the requirement Choose or investigate
Managed encryption with controlled key permissions Customer-managed KMS key and appropriate policies
Dedicated HSM control or required cryptographic integration CloudHSM
Automatically rotate supported database credentials Secrets Manager with configured rotation
Sensitive information found in S3 objects Macie
Vulnerable supported packages or images Modern Inspector
Suspicious account/workload activity GuardDuty
Consolidated findings and posture checks Security Hub
Investigate connected security events and entities Detective
Obtain AWS's compliance documentation Artifact
Block abusive HTTP requests at CloudFront WAF rules/IP sets/rate controls

Traps

  • Encryption is not authorization, and a resource share without key access can remain unusable.
  • A managed certificate is not a domain registration; a private CA certificate is not automatically public trust.
  • A detection service is not automatically a remediation engine. Enabling broad scans or organization controls can change account behavior and spending.

Practise this topic

07 · Networking: VPC

Memory hook: A working connection needs the right address, a forward route, permission and a return path.

Must remember

Addresses and routes come first

  • A VPC is a regional network boundary; a subnet occupies one AZ. Plan nonoverlapping CIDRs before connecting VPCs and on-premises networks. AWS reserves five IPv4 addresses in an ordinary subnet, so a /24 supplies 251 usable IPv4 addresses. Default-VPC conveniences should not be assumed in a custom VPC.
  • A route table selects the most specific matching destination route. A subnet is called public when it has an internet-gateway route, but an IPv4 instance also needs usable public addressing and suitable security rules. A public IP without the route, or the route without the public IP, is insufficient.
  • An internet gateway (IGW) supports the VPC's internet path. A bastion is a deliberate administrative hop; Session Manager can avoid inbound SSH by using the managed agent's outbound service connectivity and IAM permissions.
  • Private addressing does not itself guarantee isolation from every network. Check all routes: peering, transit, VPN and service endpoints may create intentional private connectivity.

Stateful and stateless filters

  • Security groups use stateful allow rules on interfaces/resources. Return traffic for an allowed connection is tracked; there is no explicit SG deny rule. Referencing another SG is useful for tier-to-tier access without maintaining individual IP lists.
  • NACLs apply ordered stateless allow/deny rules at subnet boundaries. Both directions need appropriate rules; HTTPS responses often need outbound ephemeral ports. Lower-numbered matching rules determine the decision, so an earlier deny can defeat a later allow.
  • A permitted filter cannot compensate for a missing route, and a correct route cannot override a denying filter. Trace the actual source/destination seen at each network hop.

Egress and service endpoints

  • NAT instances require operating-system management, routing, security rules and appropriate source/destination-check changes. NAT gateways reduce appliance management; time, processing and associated address charges still matter. NAT permits outbound-initiated connectivity, not unsolicited inbound sessions.
  • The traditional zonal public NAT gateway sits in a public subnet; same-AZ routing with one per required AZ avoids a single-AZ egress dependency. Current AWS also offers regional NAT gateways, which can automatically expand across AZs and do not require a hosting public subnet. Know which mode a question describes; regional mode currently does not provide private NAT. A single regional resource is not a promise of one-AZ pricing.
  • Gateway endpoints for S3 and DynamoDB add route-table targets without an endpoint-hour fee. They are not general transit access for clients in peered VPCs or on-premises networks.
  • Interface endpoints and AWS PrivateLink provide private access to supported services through endpoint networking and DNS, typically using private ENIs and SGs for interface endpoints. They can expose a particular service instead of granting full VPC-to-VPC routing. Hourly/per-AZ and data charges require comparison against the actual traffic pattern.
  • Endpoint policies, where supported, limit use through that endpoint. They do not override missing IAM permissions or a denying bucket/resource policy. Network reachability and API authorization are separate checks.

Connect networks and resolve names

  • VPC peering connects compatible nonoverlapping networks with explicit routes; it is not transitive. A–B and B–C do not establish an A–C path through B. Transit Gateway supplies a routed hub for many VPCs and on-premises connections; attachment and route-table configuration still control permitted paths.
  • Site-to-Site VPN connects networks using encrypted tunnels over IP connectivity. VPN CloudHub supports compatible hub-and-spoke VPN site communication. Client VPN supplies remote-user access, with authentication, authorization and routes; it is not the same workload as linking two corporate networks.
  • Direct Connect provides dedicated connectivity and more predictable network characteristics, with physical provisioning considerations. It is not encrypted by default. Use appropriate application TLS, supported MACsec or VPN designs when encryption is required. Direct Connect Gateway connects eligible virtual-interface designs to multiple VPCs/Regions; it is not an automatic transitive VPC router.
  • Route 53 Resolver inbound endpoints let external networks query supported AWS DNS namespaces. Outbound endpoints and forwarding rules send matching VPC queries toward external DNS. Direction follows the query, and DNS resolution still needs underlying network reachability. See DNS notes.

IPv6, observation and cost

  • Amazon-provided public IPv6 addresses are globally routable; routing and filters control reachability. An egress-only IGW permits outbound-initiated internet flows for public IPv6 addresses. Reaching IPv4-only destinations from IPv6 is a separate translation requirement.
  • AWS also supports private IPv6 through IPAM, including ULA and private GUA ranges. IGWs and egress-only IGWs drop these private ranges; internet access requires a suitable intermediary with public addressing. “Outbound-only” and “private IPv6 address” are different properties.
  • Flow logs summarize supported IP flows, including accepted/rejected traffic, rather than packet payloads. Delivering to S3 and querying with Athena supports analysis. Traffic mirroring copies supported packet traffic for inspection; Network Firewall provides managed network filtering/inspection. WAF specializes in supported HTTP request paths.
  • Add the whole path's cost: public IPv4, NAT processing, cross-AZ/Region transfer, endpoints, load balancers and inspection. Service quotas, subnet address capacity and standby-region limits can prevent scale-out even when an architecture diagram looks sound.

Choose under exam pressure

Clue in the requirement Choose or investigate
Private VPC workloads only need same-Region S3 Gateway endpoint and suitable policies
Expose one supported private service to another account PrivateLink rather than broad routed connectivity
Many VPCs need transitive routing Transit Gateway
Remote employees need authenticated private access Client VPN
On-premises DNS must query private AWS names Resolver inbound endpoint plus connectivity
VPC clients need corporate DNS zones Resolver outbound endpoint and rules
Outbound-only internet access using public IPv6 addresses Egress-only IGW
Inspect actual packet content Suitable mirroring/inspection design

Traps

  • A subnet name, SG rule or public IP alone does not establish a complete path.
  • A gateway endpoint is not a replacement for all interface endpoints or for hybrid connectivity.
  • Older “all NAT gateways are zonal” shorthand is incomplete. Preserve the availability mode and routing assumptions in the question.

Practise this topic

08 · CloudFormation and Systems Manager Operations

Memory hook: Inspect the intended change, the actual state and the identity performing it.

Must remember

  • A CloudFormation stack tracks declared resources; dependencies control ordering. Change sets preview updates, drift detection compares supported properties with actual state, and StackSets apply stacks across selected accounts/Regions. These are different operations.
  • Know update-in-place versus replacement, rollback states and why a retained resource can survive stack deletion. DeletionPolicy controls supported deletion/retention behaviour; update-replacement retention is a separate concern. Imported resources and existing physical names need careful ownership checks.
  • Template parameters vary inputs; mappings select fixed values; conditions select resources/properties; outputs expose results. Resolve circular dependencies by reconsidering resource references and ordering. A creation signal or wait condition is not automatically satisfied by an EC2 instance entering running state.
  • Use service roles with least privilege and understand iam:PassRole. A user may initiate an operation while a service role performs it. Read stack events from the earliest meaningful failure, not only the final rollback summary.
  • Systems Manager managed nodes need a working agent, identity permissions and connectivity to required endpoints. Session Manager avoids opening SSH/RDP ports for supported access. Run Command executes commands, State Manager maintains associations, Automation coordinates runbooks, Patch Manager applies patch policies.
  • Maintenance windows define when approved work may run; patch baselines define approved patches. Inventory and compliance show state; remediation requires a configured action. Restrict runbook parameters and use approvals, rate controls and failure thresholds where the impact demands them.
  • Schedule start/stop and cleanup by explicit ownership tags. Stopped compute can leave charged storage and addresses. Cost Explorer, budgets, tags and anomaly detection provide evidence and alerts, not immediate universal spending caps.

Choose under exam pressure

Requirement Choice and reason
Same declared baseline across many accounts StackSets with appropriate delegated permissions.
Find console edits to managed resources Drift detection for supported properties.
Run an approved multi-step repair Systems Manager Automation runbook.

Traps

  • Drift detection does not automatically repair drift.
  • A successful stack update does not prove application readiness.
  • A service role can perform actions the initiating user cannot perform directly; control who may pass it.

Practise this topic

09 · Operational Troubleshooting and Recovery Drills

Memory hook: Work from the symptom to the failing layer; prove recovery with a user-visible result.

Must remember

  • EC2 system status failures point toward underlying AWS infrastructure; instance status failures can indicate guest/network configuration. Inspect instance events, console output and OS logs. Reboot, recover, stop/start and replacement have different identity and storage consequences.
  • EBS bottlenecks involve IOPS, throughput, queue depth and instance limits. Burstable instance credits can constrain sustained CPU. Database latency may come from locks, connections, slow queries, memory or storage; adding CPU blindly can miss the cause.
  • For an unhealthy load-balancer target, verify health-check path, port, protocol, expected response, application binding and security-group chain. Check ASG health-check type, grace period and warmup before increasing capacity. Lifecycle hooks need completion or a deliberate timeout outcome.
  • A failed connection needs both forward and return paths. Check address family, DNS answer, route tables, internet/NAT gateways, security groups, stateless NACL ephemeral ports and endpoint policies. Flow Logs report accepted/rejected traffic metadata; Reachability Analyzer reasons about supported network configuration rather than sending a real application request.
  • For CloudFront stale or denied content, inspect cache policy, origin access, TLS, origin health and error caching. DNS TTL means an updated record may not instantly change every client. Regional/private-zone resolution may differ from a public lookup.
  • Backups should be restored into a test environment and verified for data, permissions, encryption-key access and application startup. Record RPO as allowable data loss and RTO as time to restore useful service. Replication can copy deletion or corruption; it is not a substitute for history.
  • Use EventBridge/alarms to start bounded, observable repair workflows. Suppress duplicate actions, preserve evidence, escalate unknown failures and test rollback. A green infrastructure dashboard is insufficient if customers still cannot complete a transaction.

Choose under exam pressure

Requirement Choice and reason
Instance runs but ALB marks it unhealthy Validate the actual health-check request and response.
Only large packets fail over a hybrid link Investigate MTU/path MTU and permitted control traffic.
Backups exist but recovery time is unknown Run and time a realistic restore drill.

Traps

  • NACLs need return traffic; security groups track established flows.
  • A larger fleet cannot fix all requests hitting an unavailable database.
  • A backup without its required KMS access may not be recoverable by the intended account.

Practise this topic

10 · Containers on AWS

Memory hook: Separate the container image, application task, compute capacity and permissions before choosing an orchestrator.

Must remember

Understand what each resource represents

  • A Docker image packages application layers; a container is a running instance. Rebuilding an image and replacing running containers are separate operations.
  • ECR stores images and versions/tags. Immutable tags prevent accidental tag replacement; digests identify specific content. Lifecycle rules remove old registry artifacts, but stopping tasks does not clean the registry.
  • ECS task definitions describe containers, CPU/memory, roles, networking and logging. A task is an execution; a service maintains desired running tasks and replaces failures.
  • ECS on EC2 leaves host capacity/patching and placement choices to your design. Fargate removes server provisioning while still requiring task sizing, subnets, security groups and suitable network access.
  • An empty cluster or registered task definition is not running compute. Conversely, an idle-looking service with desired tasks can keep billing.

Learn the three ECS role boundaries

  • Task role: permissions used by application code, such as reading S3 or DynamoDB. Give each workload only its required data/API scope.
  • Execution role: actions the ECS/Fargate agent performs for the task, such as pulling a private ECR image and delivering configured logs.
  • EC2 instance role: host/agent permissions for EC2-backed capacity. Do not use the host role as a substitute for per-task application permissions.
  • A successful image pull does not prove the application can reach its database. IAM, routing, DNS, security groups and the data service's policies are separate checks. ECS IAM roles

Match scaling and storage to the workload

  • ALB routes HTTP(S) to tasks and evaluates target health. Tasks using awsvpc networking register as IP targets, not host instance targets.
  • Service auto scaling changes desired task count. Capacity scaling adds/removes EC2 hosts where needed. More desired tasks do not help if the cluster lacks capacity to place them.
  • Fargate manages underlying capacity, but tasks still need valid CPU/memory combinations, quotas and reachable image/log endpoints. A private subnet may need NAT or appropriate service endpoints.
  • EventBridge/Scheduler can start standalone tasks for a schedule or event. Choose this for intermittent work rather than an always-running service; target execution permissions and iam:PassRole matter.
  • EFS supplies shared persistent files outside disposable containers. Mount targets, access points, permissions and NFS security-group paths matter. Container-local writable storage is not a shared persistent database.
  • Separate application scaling from downstream limits: scaling containers can overload a fixed-capacity database. Service scaling

Recognize Kubernetes and hybrid variants

  • EKS provides a managed Kubernetes control plane. It fits requirements for Kubernetes APIs, controllers and ecosystem compatibility, rather than merely “we have a container image.”
  • Compute options include managed node groups, self-managed EC2, supported Fargate profiles and current managed options such as EKS Auto Mode. Responsibility and feature support differ.
  • CSI storage drivers connect Kubernetes storage to AWS services. EBS is AZ-scoped block storage; EFS supports shared file access. Check the chosen compute/storage combination rather than assuming every volume works with every node type.
  • ECS Anywhere runs registered external machines under the regional ECS control plane; you still manage those machines and connectivity.
  • EKS Anywhere is customer-managed Kubernetes for supported on-premises/edge environments, including disconnected designs. EKS Distro is the Kubernetes component distribution, not a hosted control plane. EKS Hybrid Nodes instead connect customer-managed nodes to an AWS-managed regional control plane. EKS deployment choices, ECS external instances
  • Remove controller-managed load balancers and persistent storage appropriately before removing Kubernetes controllers/cluster infrastructure, or external resources may be orphaned.
  • App Runner historically simplified managed web-container deployment; App2Container analyzed/containerized existing applications. They are restricted for new customers as documented in the availability record; recognize their purposes without treating them as new sandbox defaults.

Choose under exam pressure

Requirement in the question Best direction
Containers with minimal host management Fargate
Kubernetes compatibility EKS
Specialized EC2 hosts or detailed capacity control EC2-backed orchestration
Container code needs S3 access Task role
ECS must pull a private image Execution role
Occasional scheduled container job Event-driven standalone task
Existing external machines under ECS control ECS Anywhere
Customer-managed disconnected Kubernetes EKS Anywhere

Traps

  • Task count, host count and Kubernetes control-plane availability are separate decisions.
  • Giving a role permissions does not create a route or open a security-group path.
  • Container-local state can disappear during replacement; scaling makes that weakness more visible.
  • A distribution of Kubernetes software is not the same product as an AWS-managed Kubernetes cluster.

Practise this topic

11 · Route 53

Memory hook: DNS chooses an answer that clients may cache; it does not inspect or balance every application request.

Must remember

Resolution, records and ownership

  • A recursive resolver finds an answer on a client's behalf, following cached information and authoritative DNS as needed. An authoritative hosted zone contains the records for its namespace.
  • A maps a name to IPv4; AAAA to IPv6; CNAME to another hostname; NS identifies authoritative name servers; MX identifies mail servers; TXT carries text/verification data; SOA carries zone metadata.
  • TTL controls how long a DNS answer may be cached. Lower TTL can improve the responsiveness of future changes but increases queries and cannot instantly invalidate previously cached answers.
  • Domain registration and DNS hosting are separate services. A third-party registrar can delegate a domain to Route 53 by using the correct Route 53 name servers. Moving DNS does not necessarily require transferring registration.
  • Creating a same-named public hosted zone does not configure delegation; resolvers need the correct authoritative chain.
  • A CNAME cannot occupy a zone apex alongside its required SOA/NS records. Route 53 Alias A/AAAA can map an apex to supported targets such as an ALB or CloudFront distribution.
  • Alias is an AWS DNS feature with supported target rules, not permission to point any apex record at an arbitrary hostname. Its TTL/health behavior depends on the target.

Every routing policy answers a different question

Requirement or clue Routing policy and meaning
Ordinary answer for one service Simple: basic response without specialized selection
Gradual rollout or relative distribution Weighted: choose among records according to relative weights
Best response latency among configured Regions Latency: use measured latency information, not just map distance
Primary/standby DNS recovery Failover: prefer healthy primary, otherwise secondary
Country/continent or location-specific content Geolocation: select by user location; define a default
Shift a geographic catchment boundary Geoproximity: resource/user location with adjustable bias
Known client-source CIDR requirements IP-based: use CIDR collections to choose destinations
Several healthy IP answers Multivalue: return a small set of healthy answers; not a full load balancer
  • Weights are relative, not required to total 100. Resolver caching and client reuse mean a 90/10 policy does not guarantee nine of each ten HTTP requests use one endpoint.
  • Latency and geography differ: the geographically nearest endpoint need not have the lowest measured network latency.
  • Geolocation is not authorization or residency enforcement; storage locations, cache distribution and access controls need separate policies.
  • Geoproximity bias changes the area attracted to a resource; it is not the same as changing an exact percentage weight.
  • Traffic Flow can compose visual traffic policies, with separate pricing/management considerations. It is a configuration facility rather than another universal per-request proxy.

Health and failover

  • An endpoint health check probes a supported publicly reachable endpoint using its configured protocol and conditions.
  • A calculated health check combines child checks using a threshold; a CloudWatch alarm-based health check derives health from supported alarm/metric behavior instead of a direct public probe.
  • Route 53 public health checkers cannot directly reach an ordinary private-only IP. A suitable metric/alarm integration is one way to express private resource health.
  • Health checks are separate resources and must be correctly associated with the DNS design. Supported Alias targets may provide Evaluate Target Health behavior instead of needing a duplicate direct endpoint probe.
  • Failover depends on detection time, DNS answers, resolver caches and application reconnection. A small TTL is not a promise that all clients switch instantly.
  • Secondary endpoints still need usable capacity and sufficiently current data; health checks preserve neither transactions nor state.

Private zones and hybrid DNS

  • Private hosted zones answer within associated VPCs through the appropriate resolver context. Required VPC DNS attributes, associations and application resolver configuration must be correct.
  • Split-view DNS uses the same namespace with different internal and external answers. A private zone can intentionally shadow public names.
  • If an associated matching private zone lacks the requested name/type, resolution can return NXDOMAIN, rather than automatically falling back to the public zone.
  • Route 53 VPC Resolver is the current name for the VPC service historically called Route 53 Resolver. Its inbound endpoints receive queries from on-premises/other connected networks.
  • Outbound endpoints and forwarding rules send matching VPC queries to external DNS servers. Think “inbound to the VPC” and “outbound from the VPC.”
  • Endpoints do not create the underlying VPN/Direct Connect route. DNS ports, security groups, routes, forwarding rules and resilient endpoint placement are separate requirements.
  • Private zones do not support every public-zone routing policy.

Choose under exam pressure

Situation Decision and reason
Apex domain must reach an eligible AWS load balancer Alias A/AAAA, not apex CNAME
Keep registrar but use Route 53 DNS Update authoritative delegation
Hybrid clients must resolve AWS private names Inbound Resolver endpoint and private connectivity
VPC clients must resolve corporate names Outbound endpoint with matching forwarding rules
Internal name exists publicly but fails inside VPC Inspect private-zone shadowing and record/type
Exact request-level canary split required DNS weighting alone cannot guarantee it
Private backend needs failover health Appropriate alarm/metric-based health integration

Traps

  • Resolution is not connectivity. A correct address does not open a firewall or establish a route.
  • Caching limits immediate control. DNS updates do not terminate established connections or flush every resolver.
  • Private and public evidence differ. A laptop using public DNS cannot prove a VPC-only record is absent.

Practise this topic

12 · CloudFront and Global Accelerator

Memory hook: CloudFront caches HTTP content, Global Accelerator routes network connections, and replication creates durable regional copies.

Must remember

Match the origin and cache behavior to the application

  • CloudFront accepts viewer HTTP(S) requests at edge locations. A distribution can use different origins and cache behaviors for static assets and dynamic API paths.
  • S3 REST origin plus OAC: CloudFront signs origin requests, and a bucket policy permits the intended distribution. The bucket can remain private. For SSE-KMS objects, the key policy also needs appropriate access.
  • An S3 website endpoint is a custom origin, cannot use OAC, and does not provide origin HTTPS. Do not choose it when the requirement is private S3 access through signed origin requests. S3 origin restrictions
  • ALB/EC2 origins serve dynamic HTTP applications. CloudFront can pass requests through without caching user-specific responses; “CDN” does not mean static files only.
  • Current VPC origins support eligible private-subnet ALBs, NLBs and EC2 instances, subject to service/Region constraints. Do not memorize the outdated rule that every application origin must be public. Public custom origins still need protection against bypassing the distribution. VPC origins

Keep caching separate from authorization

  • The cache key decides which requests can reuse a response. Include only relevant headers, cookies and query parameters: unnecessary variation lowers hit rate; missing user-specific variation can expose private content.
  • The origin request policy decides what reaches the backend. Forwarding a value and including it in the cache key are different decisions. Disable caching where that is the safest fit for personalized data.
  • TTL bounds freshness. Versioned object names make immutable assets easy to update without replacing the contents behind a cached URL. Invalidations remove cached paths earlier and may add cost.
  • Signed URLs suit individual restricted resources or clients that do not support cookies. Signed cookies suit access to multiple restricted files without rewriting each URL, such as a media presentation with many segments.
  • OAC authorizes CloudFront to S3; signed URLs/cookies authorize viewers to CloudFront. They solve different boundaries and may be used together. Signed access choices
  • Geo restriction controls viewer countries based on location signals. Price classes constrain the eligible edge footprint to trade price against reach. Neither determines the S3 bucket's storage Region or guarantees data residency.
  • TLS applies on both viewer and origin connections. For a CloudFront custom hostname using ACM, the viewer certificate is obtained in us-east-1; an ALB's certificate is regional to that ALB.

Distinguish routing, resilience and edge code

  • Global Accelerator provides static anycast IPs and routes TCP/UDP through AWS networking to supported healthy endpoints. Endpoint groups are regional; traffic dials and endpoint weights influence traffic distribution.
  • It does not cache S3 objects or replace application authorization. Its control API uses us-west-2, although application endpoints can be elsewhere. Global Accelerator overview
  • S3 CRR stores durable copies in another Region. CloudFront caches can expire or evict objects, so caching alone is not regional disaster recovery.
  • CloudFront Functions is for lightweight viewer request/response logic, such as URL or header changes. Lambda@Edge supports richer processing and origin events, with different limits and replicated-resource lifecycle behavior.
  • CloudFront changes/deletion need propagation. Lambda@Edge replicas can delay cleanup; avoid treating edge code as an ordinary instantly deleted regional function. See serverless services and S3 replication.

Choose under exam pressure

Requirement in the question Best direction
Millions of identical global HTTP downloads CloudFront caching
Private S3 origin with controlled viewer access OAC plus signed viewer access
Many restricted media segments Signed cookies
Static global IPs for TCP/UDP applications Global Accelerator
Durable regional recovery copy S3 CRR
Personalized API behind an ALB Appropriate cache policy or caching disabled
Simple viewer URL/header rewrite CloudFront Functions

Traps

  • A CDN cache is not a backup, and a replica is not an authorization mechanism.
  • Country filtering is not the same as choosing where original data is stored.
  • Allowing all CloudFront traffic to a public origin does not necessarily restrict access to your specific distribution.
  • A long TTL on a personalized response can be a security mistake, not just a freshness mistake.

Practise this topic

Search across every published topic.