certslothcertsloth
SOA-C03/Topic 05

AWS / Associate

Disaster Recovery & Migrations

6 min read5 recall promptsReviewed 2026-10-10

Memory hook: RPO limits how much data may be lost; RTO limits how long useful service may be unavailable.

Must remember

Objectives before recovery patterns

  • Recovery point objective (RPO) measures acceptable data loss as a time interval. Recovery time objective (RTO) measures acceptable recovery duration. Fast restoration cannot recreate transactions that never reached the available recovery data.
  • Measure detection, decision, provisioning, data recovery, dependencies, traffic switching and validation. A database restore benchmark alone does not prove users can work within the RTO.
  • Backup and restore: retain recovery data, then rebuild as needed. Pilot light: keep the minimal core running and add missing serving capacity. Warm standby: maintain a functional reduced-capacity workload and scale it. Multi-site active: serve from multiple sites while managing routing, consistency and failure behavior.
  • Recovery patterns trade standing cost and complexity against recovery work; their names do not guarantee an RPO/RTO. Select using tested behavior and actual requirements.
  • Availability across AZs and regional disaster recovery solve different failure scopes. Multi-AZ replication does not automatically protect against a Region-wide outage, accidental deletion or corrupted application data.

Backups need more than a schedule

  • AWS Backup plans define rules; resource selections determine what is protected; vaults contain recovery points. A plan with no selections schedules no useful protection for the intended resource.
  • Cross-Region or cross-account copies require compatible resource support, IAM and encryption-key arrangements. Verify the copy, its retention and the destination's ability to restore. A copied recovery point that cannot be decrypted is not a working recovery plan.
  • Vault Lock provides retention controls; compliance-mode protection can intentionally prevent early deletion. Governance controls and compliance immutability have different bypass properties. This pack does not create retention that obstructs immediate cleanup.
  • Keep historical recovery points where required: replication can quickly copy an unwanted deletion or bad write. Backup retention, replication and application validation protect against different failure modes.
  • Validate service quotas, instance availability, subnet addresses, dependencies and key access in the standby Region before a disaster. A successful Terraform plan does not reserve all future capacity or raise every quota.

Migrate with a deliberate cutover

  • DMS supports data movement with full load and change data capture (CDC) for supported endpoints. SCT helps convert schema/code for heterogeneous migrations; data replication does not automatically convert every stored procedure or engine feature.
  • A low-downtime database migration commonly uses initial load, ongoing CDC, validation, a controlled write cutover and a rollback decision. Monitor replication lag and reconcile data; “CDC enabled” is not proof of zero data loss.
  • RDS/Aurora paths include compatible dump/restore, snapshots, replicas and DMS. Check engine/version support and downtime tolerance before selecting a path.
  • VM Import/Export moves supported VM images. Application Migration Service (MGN), now documented as AWS Transform MGN, replicates servers for rehosting and cutover, with staging resources and costs. Rehosting differs from moving only a database.
  • Application Discovery Service is a historical inventory/discovery tool closed to new customers, not a universal replacement for MGN. VMware Cloud on AWS preserves VMware-based operating assumptions but requires current commercial availability and capacity checks. It remains named in the published exam list; that does not authorize a new subscription in this lab.
  • For large datasets, estimate effective transfer time from data size and throughput, including validation and changes during transfer. Compare DataSync, suitable uploads and supported alternatives; transfer/storage notes cover the tool boundaries. Snow Family remains a published exam concept despite lifecycle changes; no physical job is created here. Consult service status.

Choose under exam pressure

Clue in the requirement Choose or investigate
Restore rarely, tolerate substantial recovery work Backup and restore
Keep critical data/core services ready but provision serving tiers later Pilot light
Need a functional secondary environment before failure Warm standby
Minimize database migration downtime Full load plus CDC, validation and controlled cutover
Rehost whole servers with limited rewriting MGN
Convert between database engine families Schema assessment/conversion plus a compatible data migration
Centralize protection across supported resources AWS Backup plans, selections and restore tests
A standby cannot scale during disaster Inspect quotas, capacity, address space and dependencies

Traps

  • Faster compute during restore improves no data that is absent from the backup.
  • A replica is not automatically a historical backup, and a backup is not automatically immediately available serving capacity.
  • Retention locks deliberately constrain deletion. Terraform lifecycle flags cannot bypass a genuine compliance retention requirement.

Active recall

1. Nightly snapshots restore in five minutes, but the business accepts only one minute of data loss. Which target fails?

RPO fails: the snapshot may omit many hours of transactions. A fast restore helps recovery duration, but meeting the tighter data-loss target requires sufficiently current recoverable data.

2. A secondary Region has replicated data but no application tier. Is that necessarily warm standby?

No. A minimal core with serving components still to provision is closer to pilot light. Warm standby maintains a functional reduced-capacity workload that can be scaled for service.

3. DMS has finished a full load, but the source database still accepts writes. Is it safe to switch clients immediately?

Not merely because the full load finished. Assess CDC lag, validation, source write control, dependent clients and rollback before the planned cutover; otherwise recent changes or split writes can be lost.

4. A restore exercise cannot launch enough instances in the standby Region. What was missing beyond backups?

Regional quota and capacity planning, along with relevant subnet addresses and dependencies. Recovery readiness includes the ability to provision and operate the target workload, not just possession of its data.

5. A replicated database faithfully copied an accidental deletion. Why might a separate recovery point still be valuable?

Historical backups or point-in-time recovery can provide a state from before the mistake. Replication keeps another copy current, which can also propagate the unwanted change.

Terraform anchor: Preconditions can reject inconsistent declared RPO/RTO assumptions; actual recovery tests, validated data and regional readiness provide the operational proof.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.