Memory hook: Start with the workload's requirements, then explain what the design protects, what it costs and how you know it works.
Must remember
Six pillars, six different questions
- Operational excellence: can the team operate and improve the workload? Prefer repeatable deployments, observability, runbooks, incident learning and clear response ownership.
- Security: are identities, data and workload boundaries protected? Use least privilege, appropriate encryption, traceability and a response plan. Security includes application behavior and the handling of credentials, not only network rules.
- Reliability: can the workload withstand and recover from failures while meeting its objectives? Remove unjustified single points of failure, test recovery, plan quotas/capacity and automate suitable repair. Durable data and continuously available service are different outcomes.
- Performance efficiency: do resources suit changing demand? Measure bottlenecks and choose appropriate data models, compute, storage, networking and scaling; larger instances do not fix every bottleneck.
- Cost optimization: does spending create useful value? Evaluate utilization, commitments, transfer, managed-service tradeoffs and operating effort; the lowest resource price can hide wider costs.
- Sustainability: can the workload deliver useful work with a smaller resource footprint? Improve utilization, reduce unnecessary processing/data movement and match capacity to demand. Deleting useful redundancy without considering the reliability objective is not a complete optimization argument.
- Operate, protect, recover, perform, justify cost, reduce waste: explain when a choice helps one pillar but constrains another.
Shared responsibility is service-dependent
- AWS is responsible for security of the cloud; customers remain responsible for their configuration and use of it. The boundary shifts with the service model, not with whether a console calls it “managed.”
- With EC2, the customer manages the guest OS, patching and application configuration alongside data and identities. A managed database offloads substantial engine/infrastructure work, but the customer still chooses access, data protection settings and application behavior within the service's capabilities.
- With S3, AWS operates storage infrastructure; customers control data classification, authorization and protection/retention settings. Encryption does not fix overly broad authorized access.
- Map compliance requirements to controls and evidence. AWS reports do not certify your application; see security notes.
Review tools support judgment
- Well-Architected Tool records reviews, answers, risks and milestones. Assign owners and prioritize evidence-backed remediation; completing a questionnaire does not automatically make the workload resilient or compliant.
- Trusted Advisor supplies checks and recommendations, with availability varying by support/features. A recommendation is an investigation lead, not permission to change production. Review workload purpose, peak demand and recovery constraints before accepting a cost or capacity suggestion.
- Reference architectures and AWS Architecture Center examples illustrate patterns, service boundaries and failure domains. Adapt them to access patterns, compliance, team skills, latency, RPO/RTO and budget. A three-Region diagram is not a requirement for every application.
- Specify traffic, consistency, tolerable data loss, recovery time and failover capacity. Replace “multi-AZ therefore highly available” with concrete failure behavior and tested recovery.
Read a design question systematically
- Extract hard constraints first: residency, compatibility, loss/outage tolerance, access pattern and operational requirements. Eliminate options that violate them before comparing cost or convenience.
- Match the stated optimization: operational effort, resilience, cost or latency can favor different solutions. Avoid unneeded capabilities.
- Distinguish the four exam domains from the six pillars. The published SAA-C03 outline organizes assessment around secure, resilient, high-performing and cost-optimized architecture; the framework is the broader review lens. The official guide explicitly says its content list is non-exhaustive, so no notes pack can guarantee every possible question.
- Validate data and failure paths: nodes, AZs, dependencies, credentials and operator mistakes. Include teardown and ownership for disposable environments.
Choose under exam pressure
| Clue in the requirement | Decision habit |
|---|---|
| Fragile single component threatens the availability target | Review reliability and failure domains |
| Repeated manual recovery steps create mistakes | Improve operational automation and runbooks |
| Idle resources consume budget and power | Assess utilization, cost and sustainability together |
| Required data residence conflicts with a cheaper Region | Satisfy the hard constraint before optimizing price |
| A reference design has more Regions than needed | Reassess against actual RPO/RTO and operational complexity |
| A recommendation conflicts with failover capacity | Validate workload purpose before applying it |
| A managed service stores sensitive customer data | Identify the customer's remaining access/configuration duties |
Traps
- Two subnet AZs do not prove there are two simultaneously healthy application targets.
- A passing configuration check does not measure real recovery, application correctness or compliance.
- “AWS manages the service” does not remove the customer's data, identity and application responsibilities.
Active recall
1. An ALB spans two AZs but its ASG has one live target. What reliability claim is unsafe?
Claiming uninterrupted service through target or AZ failure is unsafe. The target may be replaced elsewhere, but detection, launch and initialization create a gap without simultaneous healthy capacity.
2. A low-traffic internal tool tolerates a business day of recovery. Must it copy a three-Region active-active reference design?
No. A simpler design may satisfy the real recovery, security and performance constraints with less cost and complexity. Demonstrate that through tested recovery rather than assuming the reference's resource count is mandatory.
3. A team moves from EC2-hosted software to a managed service. Which responsibilities should it reassess instead of assuming they disappeared?
Reassess the service-specific boundary: identities, data classification, authorization, application configuration and relevant backup/retention choices generally remain customer decisions, while underlying operational tasks may shift to AWS.
4. Trusted Advisor flags lightly used capacity that serves as a recovery standby. Why might immediate deletion be a mistake?
Observed low use may be intentional. Evaluate failover capacity and recovery time first; removing it could satisfy a narrow cost suggestion while violating the business's resilience requirement.
5. A Terraform precondition accepts a declared five-minute RTO. What evidence is still needed?
A realistic recovery exercise measuring detection, decisions, dependencies, data recovery, routing and successful user operations. The precondition compares declared values; it cannot prove that the live process completes in time.
Terraform anchor: Typed review data and preconditions make assumptions explicit and reviewable; human judgment and operational tests determine whether those assumptions are sound.
Sources
- Well-Architected Framework defines the six review pillars.
- AWS shared responsibility explains how customer duties vary with service choice.
- Well-Architected Tool describes workload reviews and milestones.
- Trusted Advisor explains recommendations and access considerations.
- Published SAA-C03 blueprint distinguishes exam domains from framework pillars and identifies the guide's limits.