Memory hook: Measure the bottleneck; price the whole path; commit only to the baseline.
Must remember
Translate performance objectives into measurable latency percentiles, throughput, concurrency and freshness. Average latency can hide severe tail delays. Correlate traces, service metrics, database query plans and resource saturation to locate the limiting component. Scaling a web tier will not fix a serialized database lock or a single hot partition.
Caching moves reads away from slower dependencies but changes freshness, invalidation and failure behavior. Select TTL and eviction based on data semantics; protect against cache stampedes and cold-cache load. Read replicas distribute eligible reads, while connection pooling/proxying addresses connection pressure. Neither automatically increases the writer’s transaction capacity. Sharding/partition-key changes can improve distribution but introduce query and operational trade-offs.
Match storage performance to access pattern, not just capacity. IOPS, throughput, latency and request size interact. A migration from an expensive disk class to a cheaper one is acceptable only if required sustained/burst behavior remains supported. Small-file/request-heavy object access may be dominated by request/transition costs rather than bytes stored.
Cost allocation starts with accounts, activated cost-allocation tags and consistent ownership. Cost and Usage Report/Data Exports support detailed analysis; Cost Explorer supports exploration; budgets and anomaly detection notify. They do not universally prevent spend. Allocate shared networking, security and platform costs using an agreed showback/chargeback method rather than treating untagged costs as free.
Rightsize, schedule idle environments and remove unnecessary retained resources before making commitments. Savings Plans exchange eligible spend commitment for rates; Reserved Instances have offering-specific discount/capacity attributes; Spot trades interruption/capacity risk for lower price. Do not confuse a billing discount with a guarantee that a required AZ/type is available. Evaluate utilization and coverage separately, and preserve flexibility for uncertain migration demand.
Compare full data paths: NAT processing, endpoint hourly/data charges, cross-AZ/Region transfer, replication, log ingestion/retention, requests, licenses and operator effort. Centralizing egress can reduce duplicated infrastructure while increasing transit/transfer or failure dependencies. Choose the cheapest design that satisfies the full requirement, then validate assumptions with actual usage.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Database high latency despite idle app CPU | Inspect query plans, locks, connections and storage before scaling web servers. |
| Stable measured compute baseline | Evaluate an appropriate commitment after rightsizing. |
| Large NAT bill from supported AWS-service traffic | Evaluate endpoint routing and its complete cost/availability trade-offs. |
Traps
- Lowest hourly instance price is not lowest solution cost.
- A cache can worsen recovery if every instance misses simultaneously.
- Commitment coverage and commitment utilization answer different questions.
Active recall
1. Why measure percentiles?
They expose slow-tail behavior hidden by an average.
2. What does high coverage indicate?
A large share of eligible usage benefits from commitments, not necessarily that every purchased commitment is fully used.
3. Why can a hot partition resist horizontal scaling?
Load is concentrated on one key/range or serialized access path.
4. Why rightsize before committing?
To avoid locking in unnecessary baseline spend.
5. What belongs in a network cost comparison?
Hourly resources, processing, transfer, routing topology, availability and operational effort.