Memory hook: Replicas add read capacity, failover preserves service, backups recover history, proxies manage connections, and caches avoid repeated work.
Must remember
RDS: identify the actual bottleneck
- RDS provides managed relational engines. AWS operates much of the platform; you choose sizing, access and retention. Ordinary RDS provides no general guest-OS administration.
- RDS Custom allows supported OS/database customization within automation boundaries. Verify engine/platform fit; use EC2 when full control is essential.
- CPU/memory, storage, IOPS/throughput and connections are different bottlenecks; more disk does not fix CPU-heavy queries.
- Storage autoscaling increases storage toward a configured maximum; it does not shrink it. Monitor both consumption and the cap rather than treating it as unlimited.
- A DB subnet group describes eligible subnets; a parameter group configures engine settings. Neither creates running capacity or HA.
Read scaling, availability and recovery differ
- A classic Multi-AZ DB-instance deployment synchronously replicates to a standby for automatic failover. That standby does not serve application reads.
- A Multi-AZ DB cluster has a different architecture with readable standby instances. Always identify the deployment type before applying the shortcut “Multi-AZ is not for reads.”
- Read replicas principally offload reads, generally using asynchronous replication. Lag matters for read-after-write requirements. Promotion/cutover is a separate decision from the usual read-scaling purpose.
- Automatic failover changes the serving database; clients still need reconnection/retries and handling for interrupted transactions.
- Automated backups and transaction logs support point-in-time recovery (PITR). Restore produces a new database resource requiring a cutover; it does not undo selected changes in the running database.
- Manual snapshots persist until deleted and can support copying/sharing workflows. Encryption/key access and regional transfer/storage charges remain relevant.
- Replicas may reproduce accidental deletes or corrupt application writes. Replication is not backup history. RPO asks how much data loss is acceptable; RTO asks how long recovery may take.
Aurora: storage, endpoints and capacity
- Aurora separates compute instances from distributed cluster storage spanning multiple AZs. A one-instance Aurora cluster still has distributed storage, but adding an appropriately placed reader improves compute failover options.
- The writer/cluster endpoint follows the current writer. The reader endpoint distributes new connections among available readers; it does not redistribute every query inside one existing connection.
- Instance endpoints target individual instances. Custom endpoints select instance subsets, useful when reporting should use a different capacity group from ordinary application reads.
- Replica Auto Scaling changes the number of readers. Aurora Serverless v2 adjusts compute capacity. Neither should be confused with scaling the writer's count.
- Serverless v2 auto-pause/zero-capacity requires eligible versions/configuration; idle cost is not universally zero. Serverless v1 is retired; see service availability.
- Aurora Global Database uses cross-Region replication for global reads and recovery. It is a regional-disaster design, distinct from adding another reader in the same Region; asynchronous replication can imply unreplicated-write risk.
- Aurora cloning uses copy-on-write storage sharing for fast development/test copies. Shared unchanged pages and divergent writes differ from independent regional disaster-recovery copies.
- Babelfish for Aurora PostgreSQL supports many SQL Server interfaces to reduce application migration changes, but compatibility assessment is essential.
- Aurora ML integrations expose supported ML capabilities from database workflows; they do not turn every database engine into a model-training service.
Secure connections and protect capacity
- KMS protects data at rest, TLS protects data in transit, and supported IAM database authentication provides temporary authentication tokens. Database users/grants still authorize SQL operations.
- Security groups control network reachability. Permit the application SG on the required engine port; a private subnet alone is not a complete authorization design.
- Ports: PostgreSQL/Aurora PostgreSQL 5432, MySQL/MariaDB/Aurora MySQL 3306, SQL Server 1433, Oracle 1521, Redis OSS/Valkey 6379, Memcached 11211.
- RDS Proxy pools/reuses database connections and can help manage failover and connection surges. It is not a query-result cache and does not remove all database capacity limits.
- IAM authentication and Secrets Manager solve different integration needs; see encryption and secrets.
ElastiCache and cache correctness
- Redis OSS/Valkey offer richer structures such as sorted sets, with replication, failover and persistence capabilities depending on deployment.
- Memcached is a simpler multithreaded distributed cache. Do not assume every modern Serverless feature matches a classic node-based engine comparison; choose the actual deployment's capabilities.
- Lazy loading/cache-aside: read cache, fetch the database on a miss, then populate. It avoids caching never-read objects but adds miss latency and can return stale data.
- Write-through: update the cache alongside database writes. It improves hit availability at extra write work, but still requires a coherent failure/invalidation strategy.
- TTL expires entries. Stagger expiry or use appropriate request coordination to avoid a cache stampede when many entries expire together.
- An external session store makes web servers replaceable. Required session durability and failover must still be designed; stickiness alone does not preserve lost memory.
- Cache authentication/user controls, supported IAM authentication, TLS, and SGs address distinct layers. A cache is not safe merely because it has no public endpoint.
Choose under exam pressure
| Requirement or symptom | Decision and reason |
|---|---|
| Primary overloaded by tolerant reporting reads | Read replicas; account for lag |
| Automatic recovery from AZ failure | Appropriate Multi-AZ deployment |
| Recover before an accidental update | PITR or suitable snapshot |
| Burst of short-lived Lambda connections | RDS Proxy |
| Repeated expensive reads or fast shared sessions | Appropriate cache strategy |
| Isolate Aurora analytics readers | Custom endpoint and selected capacity |
| Variable Aurora compute demand | Serverless v2, with supported limits/cost model |
| Database recovery after regional loss | Cross-Region copies or Global Database design |
Traps
- “Multi-AZ” is not one architecture. Distinguish classic DB-instance standbys from readable DB-cluster instances.
- Encryption is not one switch. At-rest KMS, network TLS, identity authentication and SQL privileges solve different problems.
- Caches and replicas copy failures too. Neither automatically supplies a clean historical recovery point.
Active recall
1. A classic Multi-AZ PostgreSQL instance needs faster reporting reads and automatic AZ failover. Does its standby satisfy both?
No. Its standby supports failover but is not a reporting endpoint. Add suitable read capacity or choose a deployment designed with readable instances. Confirm consistency requirements before routing reports to an asynchronous replica.
2. Thousands of Lambda invocations open brief database connections. Why might a cache be the wrong first fix?
RDS Proxy pools connections. A cache only reduces suitable database work; it does not pool remaining connections or necessarily address connection exhaustion.
3. A DELETE reaches the writer and every replica. Which recovery feature addresses the mistake?
PITR or an appropriate prior snapshot, followed by validation/cutover. Replicas can reproduce the deletion; failover to an affected replica does not rewind history.
4. One Aurora client keeps a long-lived connection to the reader endpoint. Will each query use a different reader?
No. The endpoint balances new connections; the existing connection remains attached to its selected instance. Connection-pool behavior and custom endpoint design matter when distributing or isolating workloads.
5. An encrypted private database accepts valid IAM tokens but rejects a SQL operation. What layer remains?
Database-level authorization, such as the user's grants, still applies. KMS secures storage, private networking/SGs govern reachability, and IAM authentication proves an identity; none automatically grants permission to every table or operation.
Terraform anchor: Guard optional database references, protect local state containing secrets, and treat deletion/backup settings as explicit lifecycle decisions.
Sources
- RDS Multi-AZ deployment types — instance and cluster distinctions.
- Aurora cluster architecture — compute and distributed storage.
- Aurora endpoints — connection routing and failover.
- RDS Proxy — connection management.
- ElastiCache engine comparison — deployment-specific engine capabilities.