Reviewed 10 October 2026 · Data Engineer Associate
Memory hook: Ingest with identity, transform reproducibly, publish trustworthy data.
Ingestion/transformation 34%; data stores 26%; operations 22%; security/governance 18%.
Use this as a final revision pass after the chapters. Each task below maps to the published exam outline; the outline itself is not an exhaustive list of possible questions. Recheck the official guide for your booked exam version, especially beta releases.
Must remember by exam objective
1.1 — Perform data ingestion
- Select batch, micro-batch, stream or CDC by volume, freshness and replay needs. DMS full load/CDC moves supported database changes; Kinesis/MSK retain streams; Firehose buffers delivery. Preserve offsets, event IDs and timestamps. Hot partition keys can bottleneck one shard despite spare aggregate capacity.
1.2 — Transform and process data
- Glue/EMR/Spark transform data; Lambda fits bounded events; Flink handles stateful streaming. Event time, watermarks and late-data policy affect results. Wide Spark operations shuffle; skew/tiny files/driver collection can dominate. Broadcast only genuinely small data that fits executor memory.
1.3 — Orchestrate data pipelines
- Step Functions, Glue workflows and MWAA/Airflow coordinate dependencies, schedules, retries and backfills. Persist checkpoints and make sinks idempotent. A Glue bookmark tracks supported inputs; it does not guarantee exactly-once business output after a partial write.
1.4 — Apply programming concepts
- Use modular, versioned code, packaged dependencies, parameters and appropriate SDK paginators. SQL WHERE filters rows, HAVING filters groups, windows preserve row granularity and joins alter cardinality. Use deterministic deduplication keys/tie-breakers; test nulls, duplicates and boundary timestamps.
2.1 — Choose a data store
- Relational stores fit transactions/relationships; DynamoDB fits modeled key access; S3 lakes retain flexible files; Redshift serves analytics; OpenSearch serves search. Select engines by queries, consistency, scale and cost. File/object storage, catalog metadata and analytical tables are separate layers.
2.2 — Understand data cataloging systems
- Glue Data Catalog stores metadata; crawlers infer schemas; Schema Registry manages supported streaming schema compatibility. Catalog discovery does not prove data quality or permissions. Keep business meaning, ownership and lineage alongside physical schema.
2.3 — Manage the lifecycle of data
- Choose retention, storage class, partition/file compaction, archive and deletion for raw/curated/serving layers. Track versions/snapshots and incomplete uploads. Deleting catalog metadata can leave files; deleting current files can break retained table snapshots. Verify supported cleanup and recovery semantics.
2.4 — Design data models and schema evolution
- Star schemas join facts and dimensions; normalization reduces update anomalies while denormalization favors selected reads. SCD type 1 overwrites; type 2 keeps history. Test schema additions, removals and type changes against consumers. Iceberg-style table evolution depends on engine support.
3.1 — Automate data processing by using AWS services
- Automate code/infrastructure deployment, validation and orchestration with scoped identities and environment separation. Version scripts, dependencies, schemas and job parameters. A green scheduler means only its configured steps passed; verify sink commit and expected fresh output.
3.2 — Analyze data by using AWS services
- Athena benefits from pruning, Parquet/ORC, compression and column selection; workgroups constrain query use. Redshift sort/distribution choices affect scans and joins; Spectrum queries external data. Query plans and bytes scanned explain cost better than returned-row count alone.
3.3 — Maintain and monitor data pipelines
- Monitor freshness, lag/backlog, event age, job duration/errors, retries, rejected rows, resource saturation and query cost. Correlate a bad output through lineage and logs. Backfills need deterministic write ownership and reconciliation so they do not race live ingestion.
3.4 — Ensure data quality
- Validate completeness, uniqueness, validity, consistency, freshness and business totals. Quarantine bad records with reasons; define stop/publish thresholds. Reconcile source and sink while accounting for legitimate late arrivals. Glue Data Quality can execute supported rules; weak rules provide weak assurance.
4.1 — Apply authentication mechanisms
- Use temporary IAM/service roles and appropriate database authentication instead of embedded keys. Separate human, pipeline and consumer identities. Secrets belong in managed stores, not job arguments or logs. Establish identity before troubleshooting resource authorization.
4.2 — Apply authorization mechanisms
- Align IAM, S3, KMS, Lake Formation and engine permissions. Lake Formation can apply supported table/column/row and tag-based controls; cross-account sharing needs recipient configuration. A catalog grant alone may not permit every data access path.
4.3 — Ensure data encryption and masking
- Use TLS and at-rest encryption with an accessible key. Mask/tokenize selected fields where appropriate; encryption is not anonymization. Scope raw/curated access and avoid leaking sensitive fields through rejected records, temporary files and logs.
4.4 — Prepare logs for audit
- CloudTrail and relevant service/query logs support evidence; scope necessary data events and retention deliberately. Protect destinations and key access, centralize where required and preserve caller correlation. A job log is not a complete record of every source-object read.
4.5 — Understand data privacy and governance
- Classify data, record lawful use/provenance, ownership, residency and retention. Carry deletion/consent requirements through derived tables, indexes, caches and backups as applicable. Govern sharing and test recovery using versioned code, reproducible transforms and reconciled outputs.
Choose under exam pressure
| Deciding clue | Recall the distinction |
|---|---|
| Late event belongs to yesterday’s window | Use event-time and defined watermark/lateness behavior. |
| One Spark task dominates completion | Inspect skew and shuffle; more workers may not fix one partition. |
| Latest row per key | Window ranking with a deterministic tie-breaker. |
| Daily job succeeds but no data arrives | Freshness/reconciliation alert, not only job exit status. |
| Millions of tiny lake files | Plan compaction and suitable partition/file layout. |
Traps
- Bookmarks/checkpoints are not universal end-to-end exactly-once guarantees.
- A WHERE predicate on the right side of a LEFT JOIN can remove unmatched rows.
- A crawler discovering a column does not make old consumers compatible.
- A query returning few rows may still scan a large dataset.
Verification cues
- Read a query plan and explain pruning, column reads, joins and shuffle before recommending capacity.
- Trace an event from source offset through checkpoint to committed business key in the sink.
- For a failed access, identify the job role, table grant, object policy and encryption-key permission.
Last-pass active recall
1. How does SCD type 2 help historical reporting?
It preserves versioned dimension history so facts can match the state valid at the time.
2. What does a Glue crawler store?
Discovered metadata/schema in the catalog; it does not relocate all underlying data.
3. Why quarantine instead of silently discard?
Preserve traceability, reconciliation and a controlled repair/replay path.
4. Why can a retry duplicate rows?
The prior attempt may have committed part of the output before failing.
5. Why is encryption insufficient for privacy?
Authorized users and derived copies can still expose personal data; classification, minimization and access controls remain necessary.
Sources and version check
The numbered chapters provide worked distinctions and further technical sources. These are original revision notes and original recall scenarios, not real exam questions.
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Data and analytics
Memory hook: Separate ingestion, storage, metadata, permissions, computation and visualization so each requirement has a clear owner.
Must remember
Catalog and govern the lake before choosing the dashboard
- S3 holds data-lake objects. Glue Data Catalog holds metadata describing datasets, locations, schemas and partitions; a catalog entry does not contain or transform all of the underlying data.
- Glue crawlers discover supported schemas/partitions; Glue ETL jobs clean, transform and move data. Schema discovery and transformation are separate compute activities with separate costs.
- A known small schema can be declared directly instead of running a crawler. A newly declared column does not populate missing values in existing JSON objects.
- Lake Formation centrally governs lake access, including supported fine-grained permissions over cataloged data. IAM, storage access and key permissions remain relevant; enabling governance does not mean every existing access path is automatically secured. Glue concepts, Lake Formation
Athena and Redshift are different query choices
- Athena is serverless SQL for supported data sources, especially S3. For occasional queries, it avoids keeping a warehouse cluster running purely to wait for work.
- Partitioning allows queries with appropriate predicates to skip irrelevant data. Columnar formats such as Parquet/ORC, compression and selecting only needed columns reduce scanned work.
- Many tiny files add overhead; poor partition choices can also hurt performance. A
LIMITor a post-read filter is not a universal guarantee of low scan cost. - Workgroups isolate settings and controls such as query output location and scan limits. Keep result objects outside the table's source-data prefix so later scans do not mix source JSON with generated result formats.
- Federated queries access supported non-S3 sources through connectors. Some connectors introduce Lambda execution, spill storage and source-system load; “serverless SQL” does not make all dependencies free. Athena data optimization
- Redshift is an analytical warehouse suited to substantial joins, aggregations and recurring BI workloads. Columnar/parallel execution and table distribution/sort design matter; simply adding nodes is not every query optimization.
- Redshift Spectrum queries external S3 data without first loading it into Redshift tables. This can combine warehouse data with a larger lake rather than copying every cold dataset into the warehouse.
- Snapshots and cross-region copies provide recovery options. Storage, encryption/key access and restore time still matter; an available snapshot does not prove the required recovery time.
- Provisioned versus serverless deployment changes operations and billing, not the need to size, govern and monitor the analytical workload. Spectrum
Processing, search and streaming
- EMR runs frameworks such as Spark and Hadoop with supported cluster/serverless execution choices. Choose it when the workload needs that processing ecosystem or detailed framework control; idle cluster capacity can remain costly.
- OpenSearch indexes data for relevance search, aggregations and log analytics. An index is not automatically the authoritative transactional record; ingestion, retention and recovery need design.
- MSK supplies managed Kafka-compatible infrastructure. It fits existing Kafka clients, tooling and stream semantics; it does not by itself execute all application analytics.
- Managed Service for Apache Flink runs stateful streaming computations with windows, event-time handling and checkpoints. It can consume streams and maintain continuous results rather than rerunning isolated batch queries.
- Kinesis Data Streams provides a retained stream for independent consumers; Firehose provides managed buffered delivery. Choose the ingestion/replay and delivery functions separately. See messaging and stream choices.
- The old Kinesis Data Analytics for SQL service is retired. That is not the same product as current Managed Service for Apache Flink. Flink concepts, availability record
Current BI naming and external datasets
- The current exam service list names Amazon Quick. AWS describes Amazon Quick Sight as the continuing BI/visualization capability within Quick; older material calls it Amazon QuickSight. Preserve that name mapping rather than assuming the dashboard capability disappeared.
- For the familiar exam pattern, choose Quick Sight/QuickSight for interactive dashboards, business-user visualization and embedded analytics. SPICE stores imported analytical datasets in memory; refresh behavior determines when new source data appears. Direct query and imported data have different performance/freshness tradeoffs.
- Quick also includes broader research, automation, connected knowledge and application features. Its existence does not turn a dashboard into the database or ETL engine. No subscription is created for these labs. Current Quick terminology
- AWS Data Exchange manages access/entitlements to externally shared datasets and marketplace data products. It fits acquiring or sharing third-party data, not converting file formats or operating a query engine.
- Data grants/subscriptions, dataset revisions and allowed use must be understood before incorporating external data into analytics. Receiving a dataset does not automatically catalog, cleanse or visualize it. Data Exchange
Trace one complete pipeline
A plausible design is Kinesis/MSK → Flink or Firehose → S3 → Glue/Lake Formation → Athena/Redshift → Quick Sight. These are choices, not mandatory boxes: a simple batch dataset may need only S3, a catalog, Athena and a dashboard. Specify latency, duplicate handling, failure recovery, schema evolution and access controls at every handoff.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Occasional SQL over S3 files | Athena |
| Repeated warehouse joins and analytical reporting | Evaluate Redshift |
| Query cold S3 data from Redshift | Spectrum |
| Discover schemas or transform data | Glue crawler or ETL, respectively |
| Govern lake permissions centrally | Lake Formation |
| Preserve Kafka ecosystem | MSK |
| Continuous stateful window aggregation | Managed Service for Apache Flink |
| Search relevance over indexed documents | OpenSearch |
| Dashboards and embedded BI | Amazon Quick Sight, formerly QuickSight |
| Acquire/manage external dataset access | Data Exchange |
Traps
- Metadata, data files, query execution and visualization are separate layers.
- “Real time” requires a latency definition; buffered delivery and stateful event processing are not identical.
- A cache/imported BI dataset can be fast but stale until refreshed.
- Adding a schema field is not an ETL transformation, and an S3 bucket policy alone is not a complete lake-governance design.
- Serverless query/processing choices can still incur scans, capacity minimums, storage, request and transfer charges.
02 · SQS, SNS, Kinesis and Amazon MQ
Memory hook: Queue work for competing workers, fan out events to independent consumers, and retain streams when replay matters.
Must remember
SQS separates arrival rate from processing rate
- Standard queues provide at-least-once delivery and best-effort ordering. More workers can process different messages concurrently; producer and consumer availability no longer have to match exactly.
- Retention determines how long an unprocessed message can remain. Visibility timeout temporarily hides a received message while a worker processes it. Receiving does not delete it; delete only after successful processing.
- Set visibility around realistic processing/retry behavior. A worker can extend visibility for long work. If it crashes or visibility expires before deletion, another attempt can occur.
- Long polling waits for available messages, reducing empty receives and their cost. It does not extend retention or the worker's processing deadline.
- A DLQ receives messages after the configured receive-count failure threshold. Investigate and correct the cause before redriving; a DLQ is not successful completion.
- Queue resource policies authorize cross-service or cross-account senders, often with source ARN/account conditions. Encryption and permission to send are separate controls. Visibility behavior
FIFO ordering is scoped, and business effects still need protection
- FIFO queues preserve order within a message group. Different groups can progress independently; one global group limits parallelism.
- Deduplication IDs prevent duplicate sends within the five-minute deduplication interval. Content-based deduplication hashes the message body, not its attributes; repeated identical bodies need deliberate identity semantics.
- Deduplicated enqueueing is not a transaction with a payment provider or database. A worker may commit a side effect, crash before acknowledging, then receive the message again. Use an idempotency key and atomic/conditional business-state handling.
- Moving a message out to a DLQ can interrupt an application's intended sequence. If later messages must never overtake a failed operation, design the failure workflow accordingly. FIFO deduplication
SNS distributes copies; it does not create a worker backlog by itself
- SNS pushes messages to subscribers. Filter policies select messages by configured attributes or payload fields; they are not authorization policies.
- SNS plus SQS fan-out gives each subscriber its own copy, buffer, scaling and retry boundary. A single shared queue instead distributes work among competing consumers.
- SNS FIFO topics with compatible FIFO queue subscriptions support ordered fan-out; do not assume all subscriber types preserve the same ordering guarantees.
- S3 events can publish through SNS for independent consumers. Direct S3-to-SQS-FIFO notification is unsupported; EventBridge is a supported routing alternative. See object events. Messaging choices
Streams and broker compatibility answer different requirements
- Kinesis Data Streams retains records so independent consumers can read and replay them. A partition key maps records to a shard; ordering is scoped to the relevant shard/key, not the entire distributed workload.
- Provisioned mode means planning shard capacity; on-demand mode reduces that capacity-management work. Neither excuses a design that sends all traffic through one hot partition key.
- Consumers track processing position/checkpoints. Reading a record does not delete it for other applications. Retention and consumer lag determine whether replay remains possible. Kinesis Data Streams
- Amazon Data Firehose provides managed delivery into supported destinations, with buffering and optional transformation/format conversion. Choose it when delivery is the goal; choose a stream plus consumers when custom processing/replay is the goal. Firehose's buffering is not a general-purpose long-term replay contract. Firehose
- Amazon MQ manages supported ActiveMQ/RabbitMQ brokers. Existing JMS/AMQP-style applications that must preserve broker semantics may fit MQ better than rewriting around SQS/SNS.
- ASG worker scaling: approximate acceptable backlog per instance as acceptable queueing delay divided by average processing time. Scale on backlog per worker, not raw queue length alone; keep retries and downstream limits in the design.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Buffer bursts before independent jobs | SQS |
| Per-order sequencing with parallel unrelated orders | FIFO with a group per order/workflow |
| Every application must receive the event | SNS fan-out with separate queues |
| Replay telemetry through several consumers | Kinesis Data Streams |
| Managed streaming delivery into S3 | Firehose |
| Preserve an existing broker protocol | Amazon MQ |
| Scale workers while meeting queueing latency | Backlog per worker |
Traps
- Visibility, retention and delay are different clocks.
- FIFO does not make an external payment automatically exactly once.
- A filter decides delivery selection; a resource policy decides whether delivery is authorized.
- A partition key's ordering benefit can become a throughput bottleneck when too much work shares that key.
03 · Advanced S3
Memory hook: Measure access, choose an economical storage class, automate aging, and make event consumers tolerate retries.
Must remember
Storage cost is more than the monthly rate
- Standard suits frequent access and has no minimum storage duration. Intelligent-Tiering responds to changing access patterns; monitoring charges and small-object eligibility matter. Its optional archive tiers require restore workflows.
- Standard-IA and One Zone-IA: 30-day minimum storage duration; retrieval charges; 128 KB minimum billable object size. One Zone-IA fits recreatable data because it does not survive loss of its AZ.
- Glacier Instant Retrieval: immediate access, but a 90-day minimum. Glacier Flexible Retrieval: restore before reading, with minutes-to-hours retrieval choices and a 90-day minimum. Deep Archive: hours-scale restore and a 180-day minimum.
- Those minimums are billing commitments, not deletion locks. Deleting or transitioning early can leave a remaining-duration charge. This is why a short lab uses Standard even when an archive class has a lower advertised storage price. Storage-class comparison
Lifecycle, observation and bulk actions
- Lifecycle filters by prefix, tags or size, then transitions or expires matching objects. Treat current versions, noncurrent versions, delete markers and incomplete multipart uploads as separate cleanup concerns.
- In a versioned bucket, ordinary expiration can create a delete marker while old versions remain billable. Noncurrent-version expiration removes those older versions.
- Current lifecycle defaults do not transition objects smaller than 128 KB. Explicit size filters can change eligibility, but many tiny transitions may cost more than they save. Do not confuse the storage-class minimum billing duration with the object's age before a lifecycle transition is eligible. Lifecycle constraints
- Storage Class Analysis observes access patterns to inform Standard-to-IA choices; it does not perform the transition. Storage Lens aggregates usage/activity across buckets and accounts; advanced metrics can cost extra.
- Batch Operations uses a manifest and an execution role to apply supported actions across large existing object sets. Think “managed bulk job,” not “event notification for future writes.”
- Requester Pays makes authenticated requesters pay qualifying requests and downloads; the owner still pays storage. It is not anonymous public access and does not shift every possible charge.
Notifications and throughput
- S3 events can target SNS, SQS standard queues or Lambda; EventBridge adds filtering and more destinations. Direct S3 notifications cannot target SQS FIFO.
- Expect duplicate and potentially out-of-order notifications. Make processing idempotent; use separate subscriber queues when each consumer needs its own retry history. Avoid a function repeatedly triggering itself by writing back to its input prefix. Notification destinations
- Multipart upload sends large objects as independent parts, allowing parallelism and retrying only failed parts. Unfinished parts keep consuming storage until the upload is completed or aborted.
- Byte-range GET fetches selected bytes and can parallelize downloads. Transfer Acceleration improves long-distance client-to-S3 transfer paths through edge networking; evaluate its additional cost.
- S3 scales per partitioned prefix. Parallel requests and sensible prefix distribution improve throughput; KMS quotas can also constrain SSE-KMS workloads. Randomizing every object-name prefix is not a universal requirement. Performance guidance
See storage foundations for versioning, replication and durability versus availability, and S3 access controls for encryption and policy evaluation.
Choose under exam pressure
| Requirement in the question | Best direction |
|---|---|
| Predictable aging and expiration | Lifecycle |
| Unknown or changing object access | Evaluate Intelligent-Tiering |
| Rare data must be readable immediately | IA or Glacier Instant, depending on retention/access pattern |
| Millions of existing objects need one supported action | Batch Operations |
| Independent processing and audit consumers | Fan-out with separate queues |
| Distant clients upload large objects | Test acceleration and multipart upload |
| Investigate storage growth across accounts | Storage Lens |
Traps
- Cheap archival storage can be expensive for objects deleted tomorrow or retrieved frequently.
- Lifecycle is asynchronous; it is neither an exact timer nor a complete replacement for explicit teardown.
- A prefix is part of an object key, not a filesystem directory with independent throughput hardware.
- Delivering an event once and applying a business side effect once are different guarantees.
04 · Security & Encryption
Memory hook: Protect the connection, the stored data, the permission to use it and the evidence of misuse separately.
Must remember
Encryption and key control
- TLS protects transit; encryption at rest protects stored data. Neither prevents an already-authorized compromised application from reading plaintext. Authentication, authorization, secret handling and monitoring remain necessary.
- KMS: AWS-owned keys are managed within services; AWS-managed keys are visible in your account but have service-controlled administration; customer-managed keys provide your own policy/lifecycle control. Symmetric encryption, asymmetric operations and HMAC keys solve different cryptographic tasks.
- A KMS key policy is central to authorization. IAM permissions alone are not a universal substitute for a suitable key policy. Supported rotation keeps older material available for decrypting existing ciphertext; rotating a key does not automatically re-encrypt every stored object.
- Multi-Region KMS keys share related key material, enabling supported regional cryptographic use, but policies, grants, aliases and lifecycle remain regional decisions. Creating a replica does not copy every administrative setting or automatically replicate application data.
- Encrypted snapshot/AMI sharing needs resource permissions and appropriate customer-key access for the recipient. S3 replication of SSE-KMS objects needs explicit replication configuration, source decryption and destination encryption permissions with the correct destination key. A generic S3 copy policy is insufficient.
- CloudHSM provides dedicated hardware security modules and more direct cryptographic control, with greater administration/capacity responsibility. Choose it for a requirement that specifically needs that control or interface; ordinary managed encryption requirements often fit KMS better.
Configuration, secrets and certificates
- Parameter Store provides hierarchical configuration, Standard/Advanced tiers and KMS-backed SecureString. Secrets Manager supplies secret versions, supported rotation workflows and optional regional replication. Rotation requires the relevant integration and permissions; merely storing a secret does not rotate a database password.
- Applications should retrieve secrets using a role, with caching and refresh behavior appropriate to rotation. Terraform's sensitive flag controls some display behavior; it does not encrypt local state or prevent an authorized reader from recovering supplied values.
- ACM manages certificates. An ALB uses a certificate in its Region; CloudFront's ACM certificate must be in us-east-1. Validate domain ownership and consider the whole client-to-edge-to-origin TLS path rather than securing only one connection.
- AWS Private CA is an adjacent distinction: it issues certificates for a private trust hierarchy, such as internal services. Private certificates are not automatically trusted by public browsers, and a private CA introduces charges. It is not separately named in the current in-scope list, unlike ACM.
Filtering, detection and investigation
- WAF filters supported HTTP requests with web ACLs, IP sets and rules, including rate-based rules. Shield Standard supplies baseline DDoS protection; Shield Advanced adds paid capabilities. Firewall Manager centrally manages supported security policies across an organization. None replaces least privilege or secure application logic.
- DDoS resilience combines edge absorption, caching, rate controls, suitable scaling and protected origins. Keep expensive origin work from being the first line of defense. Network Firewall handles network inspection; WAF targets supported web request paths.
- GuardDuty detects suspicious activity. Inspector finds vulnerabilities in supported workloads. Macie discovers sensitive data in S3. Select based on the finding needed, not the generic word “security.”
- Security Hub, including its security-posture capabilities, consolidates findings and evaluates supported security controls. Detective helps investigate relationships and activity surrounding suspicious behavior. Aggregating a finding, investigating it and automatically remediating it are different steps.
- Artifact provides AWS compliance reports and agreements. It does not certify your application's configuration. Audit Manager can collect and organize evidence for assessments; it is useful adjacent context rather than an explicitly named service in the current list, and it does not replace the auditor's judgment.
- Modern Inspector remains relevant; Inspector Classic is retired. Consult service status for generation-specific dates. Never infer that a current service is unavailable solely because an older namesake ended support.
Operational boundaries
- Shared responsibility changes with the service: AWS operates underlying infrastructure, while you still control data classification, identities and workload configuration. Managing EC2 also includes guest-OS responsibilities that a fully managed service takes off your hands.
- Choose retention deliberately. Customer-key deletion has a waiting period; secret recovery settings and replicas affect deletion; immutable compliance retention can intentionally prevent removal. Such retention is valuable when required by a real workload and incompatible with this disposable lab's default.
Choose under exam pressure
| Clue in the requirement | Choose or investigate |
|---|---|
| Managed encryption with controlled key permissions | Customer-managed KMS key and appropriate policies |
| Dedicated HSM control or required cryptographic integration | CloudHSM |
| Automatically rotate supported database credentials | Secrets Manager with configured rotation |
| Sensitive information found in S3 objects | Macie |
| Vulnerable supported packages or images | Modern Inspector |
| Suspicious account/workload activity | GuardDuty |
| Consolidated findings and posture checks | Security Hub |
| Investigate connected security events and entities | Detective |
| Obtain AWS's compliance documentation | Artifact |
| Block abusive HTTP requests at CloudFront | WAF rules/IP sets/rate controls |
Traps
- Encryption is not authorization, and a resource share without key access can remain unusable.
- A managed certificate is not a domain registration; a private CA certificate is not automatically public trust.
- A detection service is not automatically a remediation engine. Enabling broad scans or organization controls can change account behavior and spending.
05 · Ingestion, Streaming and Transformation
Memory hook: A durable checkpoint and an idempotent sink matter more than a promise of no retries.
Must remember
- Choose batch, micro-batch or streaming using freshness, volume, latency and recovery needs. Full loads copy a dataset; CDC captures changes. Preserve source offsets, timestamps and identifiers so replay, deduplication and audit are possible.
- Kinesis partition keys determine shard placement and per-key ordering; skew can overload a shard. Firehose handles supported delivery/buffering, not arbitrary consumer replay. MSK provides managed Kafka infrastructure; consumer groups and offsets have their own processing semantics.
- Glue jobs and EMR/Spark transform data; Lambda fits bounded event processing; Managed Service for Apache Flink handles stateful streaming. Flink checkpoints preserve recoverable state; event time, processing time, watermarks and late-event policies affect window correctness.
- In Spark, narrow transformations avoid a shuffle; wide operations such as many joins/aggregations redistribute data. Partition skew, tiny files, excessive shuffles and driver collection can dominate runtime. Broadcast a genuinely small dimension when appropriate; do not broadcast a dataset that exhausts executor memory.
- Glue bookmarks track supported processed inputs, not universally exactly-once business output. A failed job can have written partial results. Use transactional tables, staging/commit patterns or idempotent upserts as required.
- Step Functions, Glue workflows or managed Airflow coordinate dependencies and retries according to operational needs. Keep configuration separate from code, package dependencies reproducibly and unit-test transformations plus integration contracts. An SDK paginator is necessary when an API returns continuation tokens.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Correct results for late-arriving event timestamps | Event-time windows with a defined lateness policy. |
| A few keys dominate stream traffic | Reconsider partitioning and downstream ordering requirements. |
| A join shuffles a huge fact table against a tiny dimension | Evaluate a safe broadcast join. |
Traps
- Successful source reads do not prove a committed sink write.
- A bookmark is not a substitute for business-key deduplication.
- Increasing workers may not fix one skewed partition.
06 · SQL, Data Models and Schema Evolution
Memory hook: Store for the queries you need, and treat schema changes as a compatibility contract.
Must remember
- Relational OLTP favours transactional access; warehouses favour analytical scans and aggregations; data lakes retain multiple formats; lakehouse table formats add supported transactional and metadata features on object storage. Choose by query, consistency and operational requirements.
- A star schema joins fact measurements to descriptive dimensions. Normalisation reduces update anomalies; denormalisation can improve selected reads at the cost of duplication. Slowly changing dimensions may overwrite values (type 1) or preserve versioned history (type 2).
WHEREfilters rows before grouping;HAVINGfilters groups. An inner join keeps matches; a left join preserves left rows, but a right-table predicate inWHEREcan accidentally remove unmatched rows. Window functions calculate over related rows without collapsing them likeGROUP BY.- Example:
ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY updated_at DESC)can select a latest row, but equal timestamps need a deterministic tie-breaker. UseIS NULL, not= NULL; null and an empty string are different values. - Partition pruning reduces scanned partitions; Parquet/ORC and compression reduce bytes read; column selection avoids unnecessary columns. Athena workgroups control query settings and limits. Redshift distribution, sort design and query plans influence joins/scans; Spectrum queries external data.
- Glue Data Catalog stores metadata; crawlers infer supported schemas. Inference can be wrong for mixed formats or evolving columns. Glue Schema Registry supports compatibility control for supported streaming integrations. Test additions, removals and type changes against old and new consumers.
- Open table formats such as Iceberg support schema/partition evolution and transactional operations where the engine supports them. Catalog metadata, table snapshots and physical files have separate cleanup/retention lifecycles.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Preserve a customer's historical address per sale | A history-aware dimension/model, such as SCD type 2. |
| Only one date is needed from a large lake | Partition pruning plus columnar scan. |
| Keep every customer even without orders | Left join, preserving null-extended rows in later filters. |
Traps
- A crawler discovering a column does not make every consumer compatible.
- A query that returns fewer rows may still scan the same bytes.
- Deleting table metadata may leave the underlying S3 data.
07 · Data Quality, Governance and Pipeline Operations
Memory hook: Trust the dataset only when freshness, correctness, access and recovery have evidence.
Must remember
- Define completeness, uniqueness, validity, consistency, freshness and reconciliation rules. Compare source/sink counts and totals, account for legitimate late data and quarantine bad records with a reason. Glue Data Quality can evaluate supported rules; passing a weak rule set is weak assurance.
- Monitor job failures, duration, backlog, event age, lag, throughput, rejected rows and query cost. Alert on missing scheduled data as well as explicit failures. Use structured logs and lineage to trace a bad dashboard value back through transformations.
- Lake Formation governs supported lake access with table/column/row controls and tag-based permissions. IAM, S3, KMS and Lake Formation controls must align; a catalog permission alone may not establish every required data path. Cross-account sharing needs recipient-side configuration and grants.
- Encrypt data in transit/at rest, isolate network access, scope roles and avoid secrets in job arguments/logs. Mask or tokenise sensitive fields where appropriate; encryption is not anonymisation. Keep raw, curated and serving layers deliberately separated.
- Classify data, record consent/licensing and provenance, define residency and retention, and propagate deletion requirements to derived datasets where required. CloudTrail and relevant service logs support audit; scope data events and retention to the evidence requirement.
- Recovery uses versioned code/configuration, source replay or snapshots, reproducible transformations and tested output reconciliation. Define RPO/RTO, reprocess a bounded range and prevent a backfill from racing live ingestion into inconsistent output.
- Optimise after measuring: appropriate file size, partitioning, compression, worker type/count, query patterns and storage lifecycle. A cheaper worker that repeatedly spills or retries can raise total cost.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A successful job delivers no new rows | Alert on freshness/reconciliation, not just exit status. |
| Malformed records block a daily batch | Quarantine with traceability and an explicit quality threshold. |
| Backfill several months while live ingestion continues | Isolate/reconcile writes using deterministic keys and commit rules. |
Traps
- Encryption does not remove personal-data classification.
- A green job can publish logically wrong results.
- Broad catalog grants can expose more data than intended.