Memory hook: Start with the required queries and consistency, then choose the data model, resilience and operating model.
Must remember
Ask the questions in a useful order
- Access pattern: known-key lookup, relational join, relationship traversal, full-text search, time-window aggregation or large-object retrieval?
- Correctness: what must be transactional, how current must a read be, and can users tolerate eventual consistency or stale cached data?
- Scale and latency: reads versus writes, predictable versus bursty demand, hot keys, record sizes, geographic access and tail latency.
- Resilience: AZ failure, Region failure and accidental data changes are different failure cases. Specify RPO/RTO and test recovery instead of selecting a service by a generic “highly available” label.
- Operations and migration: engine/API compatibility, managed versus self-managed responsibilities, required extensions, licensing and allowed application changes.
- Cost: include idle capacity, replicas, indexes, cache clusters, I/O, backups and transfer—not only the smallest instance price. “Serverless” and “NoSQL” are properties, not complete selection criteria. AWS database selection guide
Relational, keyed and cached data
- RDS fits a supported relational engine when SQL compatibility and limited application change matter. Aurora offers MySQL/PostgreSQL-compatible relational choices with its cluster-storage and endpoint architecture.
- Separate read scaling from failover availability. A read replica, a standby and a backup have different jobs; details depend on deployment type. Relational does not imply that every write can scale horizontally without redesign. See relational resilience.
- DynamoDB fits modeled key-value/document access patterns with managed scaling. Partition/sort keys and indexes must support the required queries; scanning a huge table for routine lookups is a design warning.
- On-demand capacity can fit unpredictable requests; provisioned capacity can fit predictable utilization. Strong reads, transactions, indexes and global replication have specific costs and behavior. “NoSQL” does not mean transactions are impossible. See DynamoDB keys and consistency.
- ElastiCache reduces repeated database work and latency. Valkey/Redis OSS and Memcached have different data structures, persistence/replication and operational capabilities; choose the engine/configuration rather than assuming every cache is interchangeable.
- Lazy loading/cache-aside: on a miss, read the authoritative store and populate cache. Write-through: update cache along with application writes, trading additional work for more immediately populated entries. Both need an invalidation/expiry strategy.
- A session store can move user state out of disposable web instances. Decide what happens when a cache entry expires or a node fails; some sessions can be reconstructed, while others require stronger persistence.
- S3 stores durable objects and data-lake files. Use it for large media or immutable data, often with keyed metadata elsewhere. It is not a drop-in transactional row database.
Purpose-built models
- DocumentDB provides MongoDB-compatible document APIs. It can fit JSON-document workloads, but compatibility is version/feature-specific: check operators, indexing behavior and drivers before assuming an unchanged migration.
- “Document” describes the data/query model; simply storing JSON does not require DocumentDB if a keyed DynamoDB item or an S3 object meets the actual access pattern. DocumentDB compatibility
- Neptune serves graph workloads built around connected entities and traversals, such as fraud relationships, identity links and recommendation paths. Prefer a graph model when relationship navigation is central, not merely because the dataset contains foreign keys. Neptune concepts
- Keyspaces provides managed Apache Cassandra-compatible wide-column access. The cue is an existing Cassandra/CQL ecosystem or a suitable distributed wide-column model, with service-specific compatibility constraints. Keyspaces concepts
- Time-series storage matches timestamped measurements, retention and time-window queries. Timestream for LiveAnalytics is closed to new customers; Timestream for InfluxDB is a distinct offering with a different operating model. Preserve the time-series concept without treating the names as interchangeable. LiveAnalytics availability
- QLDB historically supplied an immutable, cryptographically verifiable journal under a central owner. Support ended July 31, 2025. Recognize the historical integrity/ledger pattern, but choose a supported current design for new systems. See the availability record.
Transactional stores are not every analytical tool
An application's primary database handles its operational reads/writes. Large historical aggregation may belong in S3/Athena or Redshift; relevance search may need OpenSearch. Copying suitable data into a specialized read/analytics system can protect the transactional workload, but introduces freshness and pipeline responsibilities. See analytics choices.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Existing relational application with joins | Compatible RDS/Aurora engine |
| Known-key low-latency application records | DynamoDB |
| Repeated hot reads or external session state | ElastiCache with explicit miss/recovery behavior |
| Large media objects | S3, often with database metadata |
| MongoDB-compatible document queries | Evaluate DocumentDB compatibility |
| Multi-hop fraud/relationship traversal | Neptune |
| Existing Cassandra/CQL workload | Keyspaces |
| Timestamped measurements and time windows | Supported time-series design |
| Large analytical aggregation | See Athena/Redshift choices |
Traps
- A replica can copy an accidental deletion; it is not a historical recovery point.
- A cache hit can be fast and stale at the same time.
- API compatibility is not a promise that every engine feature or performance characteristic matches.
- Choosing one database for every workload can increase complexity rather than reduce it.
Active recall
1. A relational application requires joins and minimal code changes. Should unpredictable traffic force a DynamoDB rewrite?
No. Preserve the required data/query semantics and evaluate a compatible managed relational design. Capacity behavior is one criterion, not permission to ignore compatibility.
2. An application repeatedly asks for the same product details. Why might caching help, and what decision remains?
Caching removes repeated authoritative reads and lowers latency. You must still decide expiry/invalidation, acceptable staleness and the behavior on cache loss or a miss.
3. Fraud analysis follows accounts through devices, addresses and counterparties. Why consider Neptune?
The query's central operation is traversing relationships. A graph model can express that directly, while simple known-key lookup alone would not justify the same choice.
4. An existing MongoDB client connects successfully to DocumentDB. Does that prove the migration is compatible?
No. Test the actual query operators, indexes, driver behavior and workload requirements against supported features. A connection test covers only one small part of compatibility.
5. A global replica contains the same accidental update as the primary. What design requirement was overlooked?
Historical recovery. Replication improves geographic availability/access but can reproduce mistakes; backups or point-in-time recovery provide a separate recovery path.
Terraform anchor: Manage database infrastructure and disposable seed fixtures deliberately; avoid making Terraform continually overwrite ordinary application-owned records.