Memory hook: A key names an object, versioning preserves history, and replication places eligible copies elsewhere.
Must remember
Objects, access and current limits
- S3 stores objects: data plus metadata identified by bucket and full key. It is not an EBS block device or NFS filesystem.
- General-purpose buckets use flat keys; slash-separated prefixes resemble folders but also support policy/lifecycle selection.
- Names may use the shared global namespace within a partition or current account regional namespaces. The bucket's chosen Region determines data location; global naming does not imply global replication.
- Strong read-after-write consistency covers object writes/deletes and corresponding reads/listing. Asynchronous replication and CDN caches still have independent delay.
- Current limits, checked 2026-10-09: AWS's upload guide states a 50 TB multipart maximum; its specification expresses the bound as 48.8 TiB. Older 5 TB course limits are outdated. A single PUT remains 5 GB.
- Multipart supports 10,000 parts, each 5 MiB–5 GiB, except no minimum for the final part. Choose adequate part sizes; unfinished parts bill until completed/aborted. See S3 operations.
Policies and website delivery
- IAM identity policies authorize callers; bucket policies attach rules to resources. Bucket actions need bucket ARNs; object actions need appropriate object ARNs.
- Explicit deny wins. Block Public Access restricts public exposure but does not grant private application access.
- Bucket owner enforced Object Ownership disables ACLs and simplifies ownership; policies handle access.
- Static websites serve HTML/JavaScript/assets, not server-side PHP or similar code. S3 website endpoints do not directly provide HTTPS.
- Private-origin HTTPS delivery commonly uses CloudFront with S3 REST-origin/OAC. Website endpoints are custom origins with different access behavior; see edge delivery and object security.
Versions and replication
- Versioning preserves prior data when a key is overwritten. Version ID, object key and bucket name identify different things.
- A normal delete usually creates a delete marker; earlier versions remain billable. Deleting a specific version permanently removes that version.
- Suspending versioning does not erase earlier history. Manage noncurrent-version retention separately from current objects.
- CRR copies across Regions; SRR stays in one Region. Both need versioning, rules and appropriate permissions.
- Live replication is asynchronous and normally covers eligible writes after configuration. Existing objects need explicit backfill, such as Batch Replication.
- Delete-marker and encrypted-object replication require suitable rule/key configuration. Never assume every deletion or protection setting propagates identically.
- A replica is not historical backup by itself; a CDN cache is not durable regional replication.
Storage classes: access, failure boundary and cost
| Requirement | Class and tradeoff |
|---|---|
| Frequent access and regional resilience | Standard |
| Unknown/changing access | Intelligent-Tiering, with monitoring/tiering economics |
| Infrequent, immediate access | Standard-IA, with retrieval/minimum charges |
| Re-creatable infrequent data; one AZ acceptable | One Zone-IA |
| Archive requiring immediate access | Glacier Instant Retrieval |
| Archive tolerating restore delay | Glacier Flexible Retrieval |
| Long archive tolerating longer recovery | Glacier Deep Archive |
| Low-latency object access colocated in an AZ | Express One Zone, using directory buckets |
- Durability means avoiding data loss; availability means successful access now. A durability figure is not an uptime percentage.
- One Zone classes accept a different failure boundary from regional classes. Avoid storing the only irreplaceable copy there when AZ-loss survival is required.
- Intelligent-Tiering adapts eligible access tiers; optional archive tiers change retrieval behavior. It is not guaranteed cheaper for every object.
- Retrieval, requests, minimum duration/size and transfer charges can outweigh lower storage rates. Short-lived tiny objects are poor archival candidates.
- Directory buckets/Express One Zone have different feature support from general-purpose buckets; do not assume identical versioning or replication behavior.
Choose under exam pressure
| Requirement | Decision and reason |
|---|---|
| Recover overwritten data | Versioning with suitable retention |
| Durable copy of new writes in another Region | CRR with permissions/monitoring |
| Replicate older objects too | Explicit backfill |
| Private static content globally over HTTPS | CloudFront and private S3 REST origin |
| Only copy must survive AZ loss | Suitable regional storage |
| Archived objects need immediate reads | Glacier Instant Retrieval |
Traps
- Invisible is not deleted. Versions, markers and multipart parts have independent retention.
- Strong consistency does not control caches or replication.
- Storage price is not total workload price. Access and retention charges can reverse apparent savings.
Active recall
1. A deleted key disappears from ordinary listing, but storage charges remain. What should be checked?
Old versions, delete markers and incomplete multipart uploads. Hiding a key with a delete marker does not permanently remove historical data.
2. A new CRR rule does not copy last year's objects. Is enabling versioning enough?
No. Versioning is a prerequisite; an explicit backfill such as Batch Replication handles existing objects. Validate eligibility, permissions and encryption configuration.
3. S3 returns current data while CloudFront serves an old response. Is S3 consistency broken?
No. CloudFront has its own cached representation and freshness policy. Diagnose TTL/invalidation and cache keys rather than assuming the S3 write failed.
4. Irreplaceable data must survive AZ loss and remain quickly accessible. Why reject One Zone-IA?
Its single-AZ boundary conflicts with the resilience requirement for the only copy. Choose a suitable regional design before optimizing its storage cost.
5. A private static site needs HTTPS. Is an exposed S3 website endpoint sufficient?
No. A suitable CloudFront distribution with private S3 REST origin/OAC fits; website endpoints have different access behavior and no direct HTTPS support.
Terraform anchor: Terraform addresses, S3 keys and version IDs differ; preserve ownership during refactors and account for retained versions during cleanup.