Reviewed 10 October 2026. Use the linked official exam guide for your exam version. These are condensed revision notes; the topic pages provide worked distinctions and more recall practice. Google’s 2026 guides use newer Gemini Enterprise Agent Platform names while some APIs and documentation still use Vertex AI.
Memory hook: Define semantics before scaling bytes.
1. Design secure, reliable data systems — 1.1–1.4
Begin with source/sink contracts, business grain, freshness, quality, access, residency, RPO/RTO and cost. Separate development and production; grant dataset/table roles independently from job execution. Plan encryption/key availability, permitted raw-data staging and deletion. A data catalog describes and discovers assets; actual access enforcement is separate.
Choose flexible formats and interfaces where portability matters, without assuming service behavior is identical across clouds. Validate generated SQL and model-enriched data against source evidence. Migration plan: assess → baseline → CDC if needed → reconcile → controlled cutover → observe → retire after fallback. Database Migration Service supports selected migration paths; Datastream captures selected source changes; transfer products move bytes, not application semantics.
2. Ingest and process correctly — 2.1–2.3
- Pub/Sub is managed messaging; Kafka organizes retained logs into partitions with offsets and consumer groups. Ordering is scoped to the relevant key/partition, not an assumed global sequence. Acknowledgment and transport deduplication do not automatically make external side effects exactly once.
- Dataflow executes Beam. Event time reflects the event, processing time the worker; fixed windows separate intervals, sliding windows overlap, sessions group activity separated by inactivity. Watermarks estimate progress; triggers decide emissions; late-data and accumulation policies determine revised results.
- A hot key serializes otherwise parallel work. Fan out/pre-aggregate only when semantics permit. Bound state and stream joins; retain a dead-letter/reconciliation path for invalid records. Preserve schema and event IDs across replay.
- Spark/Hadoop code fits Dataproc/Managed Spark; Data Fusion fits visual integration; Dataform manages SQL dependencies/assertions. ETL transforms before loading; ELT after. Choose persistent clusters for justified steady workloads and job-based/serverless execution for intermittent demand.
- CI tests transforms/schema/retry behavior; an orchestrator coordinates execution rather than carrying huge datasets in task metadata. AI enrichment needs versioned outputs, confidence/validation, bounded cost and a failure path.
3. Store by access pattern — 3.1–3.4
| Pattern | Choice and caution |
|---|---|
| Object lake | Cloud Storage with location, lifecycle, IAM and retention design |
| Large analytical SQL | BigQuery; warehouse storage and execution are separate |
| Familiar relational / demanding PostgreSQL | Cloud SQL / AlloyDB; check compatibility and HA model |
| Scalable strongly consistent relational transactions | Spanner; design keys and transaction locality |
| Wide-column/key ranges | Bigtable; avoid hot row-key prefixes |
| Application documents | Firestore; model queries/indexes and transaction boundaries |
| Low-latency cache | Memorystore; define durability and source of truth |
BigLake supports governed access to supported lake data; Dataplex/catalog capabilities supply discovery, metadata, lineage and quality. Federated governance lets domain owners manage data within common standards. An open lake without ownership or permissions is not a governed platform.
4. Prepare analysis, AI and sharing — 4.1–4.3
Partition prunes relevant partitions; clustering narrows blocks; nested/repeated structures can reduce large joins. Inspect bytes, shuffle, skew, stages and slot use. LIMIT does not reliably cap scan cost; dry runs and maximum-bytes controls address different needs. Materialized views accelerate eligible repeated computations; BI Engine accelerates suitable interactive queries; result caching has eligibility rules.
Row access policies filter records; column controls/masking protect fields; authorized views expose selected projections without broad base-table access. Looker centralizes metric definitions. BigQuery sharing/Analytics Hub distributes governed listings; exported copies have different revocation/lifecycle behavior.
For ML, split before fitting transforms, preserve point-in-time features and monitor training-serving skew. For RAG: authorized documents → extraction/chunking → versioned embeddings/index → filtered retrieval/reranking → grounded response evaluation. Retrieval similarity is neither truth nor permission.
5. Automate and recover — 5.1–5.5
Composer/Airflow manages DAG dependencies/backfills; Workflows coordinates APIs; simple SQL schedules may need neither full platform. Use deterministic partition targets, run IDs, checkpoints and idempotent/transactional publishing. Bound retries; distinguish intentional backfills from duplicate execution.
On-demand BigQuery charges by processed data; editions/reservations allocate capacity for workload requirements. Separate interactive and batch needs, and tune scans before buying capacity. Monitor completeness, freshness, rejected records, lag, billing and quota as well as exit codes. Retain replayable inputs and test corrupt/missing data recovery. Replication is availability; historical backups/PITR protect earlier valid state.
Traps to catch
- A successful pipeline can produce empty, late or duplicated results.
- More workers do not fix hot keys; more slots do not fix every bad query.
- A data replica can reproduce deletion. An authorized metadata viewer may still lack data access.
Last-pass self-check
1. An event arrives after its hourly watermark. Is it automatically invalid?
No. Watermarks estimate completeness; the configured allowed lateness and triggers determine whether/how the result is updated.
2. Why is LIMIT 10 not a reliable BigQuery cost control?
The underlying scan can still read many bytes before producing ten rows; prune partitions/columns and inspect bytes.
3. Kafka order across all partitions?
No general global ordering guarantee; the relevant ordering boundary is a partition.
4. What does a green DAG fail to prove?
Freshness, completeness, correct totals, privacy compliance and end-to-end business correctness.
5. What must a low-downtime migration verify before cutover?
Schema/application compatibility, reconciled baseline, acceptable CDC lag, target performance and a safe fallback for new writes.
Sources
- Official exam guide
- Dataflow programming model
- BigQuery performance
- Apache Kafka design
- BigQuery sharing
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Cloud Storage, Transfers and Retention
Memory hook: Object key is not a disk path; retention is not a backup.
Must remember
Cloud Storage stores objects in buckets. Choose location for latency, resilience and data requirements; bucket names and object names are separate identifiers. Object storage is not a normal mutable block filesystem. Version/generation identifiers and preconditions support safe concurrent updates.
| Class | Remember |
|---|---|
| Standard | Frequent access; no minimum storage duration. |
| Nearline | Infrequent access; 30-day minimum storage duration. |
| Coldline | Rare access; 90-day minimum. |
| Archive | Very rare access; 365-day minimum. |
Lower storage price can be offset by retrieval, operations, transfer and early-deletion charges. Archive remains online object storage; do not import another vendor's restore-wait assumptions. Autoclass and lifecycle rules automate supported class/deletion behavior under different models; review compatibility and costs.
Uniform bucket-level access uses IAM rather than per-object ACLs. Public access prevention restricts public grants. Signed URLs provide time-bound access to a specific operation/object under the signing authority; treat the URL as a credential. Encryption is default, while customer-managed keys add key-control and availability responsibilities.
Versioning retains older generations under its model; soft delete supplies a recovery window; retention policies/holds prevent deletion under configured rules. A locked retention policy can be irreversible: understand it conceptually instead of experimenting on a disposable-cost assumption. Lifecycle deletion may be blocked by retention/holds.
Use gcloud storage for object operations, Storage Transfer Service for managed supported transfers and Transfer Appliance for suitable very large offline transfers. Check checksums, permissions, location compatibility and transfer progress. A successful upload is not a tested application restore.
Review details
Cloud Storage provides strong global consistency for object writes, overwrites, deletions and listing. Do not describe object listing as eventually consistent. Access-policy changes take time to propagate, and publicly cached objects can remain stale until their cache lifetime expires: those are different consistency boundaries.
A generation precondition can make a write conditional on the expected object version and prevent lost updates. Single-object operations are atomic, but a batch of several independent object requests is not a multi-object transaction. Pin a generation when a sequence of range reads must use the same version.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Frequently accessed application objects | Standard storage in an appropriate location. |
| Restrict access consistently | Uniform bucket-level access and least-privilege IAM. |
| Bulk recurring transfer | Storage Transfer Service where the source/destination are supported. |
Traps
- Cheap storage classes can incur minimum-duration and retrieval charges.
- A signed URL can be used by whoever possesses it until its conditions expire.
02 · Databases, Analytics and Data Pipelines
Memory hook: Choose by access pattern, not by the largest feature list.
Must remember
| Requirement | Candidate |
|---|---|
| Managed familiar relational engine | Cloud SQL; assess engine, HA, replicas and backups. |
| PostgreSQL-compatible demanding enterprise workload | AlloyDB where its architecture fits. |
| Horizontally scalable relational transactions | Spanner, with deliberate schema/key and location design. |
| Document-oriented application data | Firestore, with document/query/index design. |
| High-throughput wide-column/key access | Bigtable; design row keys to avoid hotspots. |
| Analytical SQL over large datasets | BigQuery; separate from OLTP assumptions. |
| Durable object data | Cloud Storage. |
Availability replicas, read scaling and backup/PITR solve different problems. Verify automatic failover behavior, replication scope/lag and tested restoration for the chosen product. A read replica does not necessarily protect against a destructive write replicated from the primary.
Pub/Sub transports asynchronous messages; Dataflow runs Apache Beam batch/streaming processing; Dataproc runs managed Spark/Hadoop-style workloads. Select by existing code, operational burden, state/window requirements and integration. Inspect job status and logs, not merely the existence of a job resource.
BigQuery organizes projects, datasets, tables and jobs. Partition pruning and clustering can reduce scanned data; selecting needed columns and applying suitable filters controls cost/performance. Load jobs, streaming and external data access differ in latency and economics. bq and SQL query tools operate under the actual caller's IAM permissions and chosen processing location.
Migrate data with supported transfer/replication tools, consistent snapshots or exports, schema conversion where needed and a controlled cutover. Match storage/data-processing locations to avoid unsupported operations, latency or transfer cost. Protect secrets, database network access and service identities; a private IP is not database authorization.
Review details
Cloud SQL regional HA maintains a failover standby with synchronous cross-zone protection; it is different from an asynchronous read replica serving read traffic. Read-after-write through a lagging replica can return old state. Spanner supplies strongly consistent relational transactions; Firestore provides strongly consistent document reads/queries and supported transactions. Bigtable supports strongly consistent reads in a single cluster, while multi-cluster replication/routing can introduce eventual consistency. Choose the documented mode rather than equating “NoSQL” with eventual consistency.
Cloud SQL/AlloyDB auth proxies or connectors simplify supported authenticated TLS connections, but do not create a missing private route. Database users/object privileges are separate from cloud-resource administration. Bound application connection pools across maximum instances.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Global relational consistency and scale | Evaluate Spanner rather than treating all SQL services as interchangeable. |
| Petabyte analytical scans | BigQuery with partition/clustering design. |
| Existing Spark pipeline | Dataproc may reduce migration work. |
| Managed event-time stream processing | Dataflow/Beam with appropriate windows and state. |
Traps
- Bigtable row-key design can create severe hotspots.
- BigQuery is not a direct substitute for a high-frequency transactional database.
03 · IAM, Service Accounts and Data Security
Memory hook: Grant the action to the actual caller at the narrowest scope.
Must remember
An IAM binding associates a principal with a role on a resource, optionally under conditions. Basic roles are broad; predefined roles are service-oriented; custom roles bundle supported permissions for specific needs. Inherited allow grants can broaden access; deny policies and organization constraints have distinct effects. Removing one narrow grant does not remove an inherited grant elsewhere.
A service account is both an identity used by workloads and a resource whose use can be controlled. Attaching/acting as a service account and creating short-lived impersonated credentials require different permissions. Grant API permissions to the runtime identity, not merely to the human who deployed it.
Prefer attached workload identity, service-account impersonation or Workload Identity Federation over downloaded long-lived keys where supported. External federation exchanges trusted external identity for controlled Google access. Protect audience, attribute mappings/conditions and role bindings; a permissive trust mapping can expose many unintended callers.
Use Secret Manager for secrets and Cloud KMS for encryption-key control. Default encryption does not mean everyone should read the data. Customer-managed keys introduce key IAM, location, rotation, availability and destruction considerations. Separation between key administrators and data users reduces excessive privilege.
VPC Service Controls creates supported service perimeters to reduce data exfiltration; it complements IAM rather than replacing it. Identity-Aware Proxy controls supported application/tunnel access. Audit logs identify activity under their service-specific categories/settings; enable required data-access visibility deliberately and protect the sink destination.
For permission denied, establish the real principal, resource project, required permission, inherited/conditional/deny policies and any perimeter/org-policy restriction. Granting Owner to “make it work” conceals the diagnosis and creates risk.
Review details
Service Account User (roles/iam.serviceAccountUser) includes acting as the service account for supported resource attachment. Service Account Token Creator (roles/iam.serviceAccountTokenCreator) supports generating impersonated credentials and supported signing operations. Granting attachment rights is not the same as giving the human direct access to every resource the account can read, but launching code as that account can be a privilege-escalation path.
Application Default Credentials searches supported credential locations; it is not an IAM role. Workforce federation serves external people, workload federation software. To diagnose a denied call, distinguish authentication failure, missing permission, inherited deny/boundary, organization policy and VPC Service Controls. Principal access boundaries constrain eligible resources for supported access; they do not grant permissions.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Application accesses a bucket | Grant the workload identity the necessary bucket role. |
| CI outside Google Cloud | Federated short-lived credentials with narrowly defined trust. |
| Reduce exfiltration through supported managed APIs | VPC Service Controls plus IAM and data controls. |
Traps
- Service-account use permission is not identical to the service account’s resource permissions.
- A VPC Service Controls perimeter does not replace IAM.
04 · Batch, Streaming and Processing Semantics
Memory hook: Event time says when; watermark estimates completeness.
Must remember
Pub/Sub buffers messages between independent producers and consumers. A partition/key strategy, acknowledgments, retries and dead-letter handling affect processing correctness. Separate transport delivery guarantees from the final business effect: a retried database write still needs deduplication or a transactional/idempotent design.
Dataflow runs Apache Beam pipelines for batch and streaming. Event time is the timestamp of the real event; processing time is when the worker handles it. Windows group records, watermarks estimate event-time progress, and triggers decide when to emit results. Allowed lateness and accumulation choices determine whether delayed events revise earlier answers. A watermark is an estimate, not proof that no older message will arrive.
Choose Managed Service for Apache Spark/Dataproc when preserving Spark/Hadoop libraries and existing jobs matters. Ephemeral clusters or serverless execution reduce idle cost; persistent clusters may suit repeated tightly scheduled workloads. Data Fusion offers visual integration; Dataform organizes SQL transformations and assertions in BigQuery. ELT loads before transformation; ETL transforms before loading. Neither excuses storing prohibited raw sensitive data.
Design a schema contract, malformed-record path, idempotency key and reconciliation counts before increasing throughput. Handle evolving fields compatibly, monitor freshness and backlogs, and avoid logging raw secrets. For stream joins, reason about state size, time bounds and late data. Hot keys can bottleneck an otherwise well-provisioned pipeline; redistribute or pre-aggregate where semantics permit.
AI enrichment is another processing dependency: bound timeouts/cost, retain model/version provenance, validate structured output and isolate failed records. A model response is not automatically correct or safe to publish.
Review details
Kafka uses partitioned retained logs, offsets and consumer groups; order is within a partition, not universal across a topic. Pub/Sub hides more broker operations and uses subscriptions/acknowledgments; do not import every Kafka partition-management assumption. Select by ecosystem compatibility, delivery/access patterns and operating effort.
Beam fixed windows divide time into distinct intervals; sliding windows overlap; session windows group activity separated by a gap. Triggers control output timing, and accumulation modes control whether subsequent panes include earlier results. “Exactly once” must be scoped to the documented processing/storage boundary; an external email or payment still needs idempotency.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Portable batch and streaming transformations | Beam on Dataflow. |
| Preserve existing Spark jobs | Managed Spark/Dataproc. |
| Late transactions change hourly totals | Event-time windows with a deliberate late-data/trigger policy. |
Traps
- Exactly-once transport does not guarantee exactly-once external side effects.
- Adding workers does not fix a single hot key.
05 · BigQuery Models, SQL and Capacity
Memory hook: Partition prunes; cluster narrows; slots execute.
Must remember
BigQuery separates managed storage from query execution. Choose on-demand bytes-processed economics or capacity-based reservations/editions according to workload predictability and isolation needs. Reservations assign capacity; they do not remove the need to tune wasteful SQL. Separate interactive reporting from batch work when latency and priority differ.
Partition large tables on a useful date/time or supported range key. Filters must permit partition pruning. Clustering organizes blocks around selected columns and can reduce scanning within partitions. Avoid unnecessary SELECT *, repeated large joins and unbounded date ranges. Inspect the execution plan, scanned bytes, skew, shuffle and slot utilization before selecting a remedy. A LIMIT does not necessarily reduce bytes read.
Denormalized nested/repeated structures can reduce repeated joins for suitable analytical relationships; normalize where reuse and update correctness demand it. OLTP transaction design and analytical models optimize different access patterns. Materialized views precompute supported results with managed refresh; BI Engine accelerates suitable interactive analysis. Result caching depends on eligibility and unchanged inputs.
External/federated queries access data without fully loading it into native tables, but latency, supported operations, source load and network locality still matter. BigLake supports governed access across supported lake data. Use compatible formats, partition layout and regional placement.
For dashboards, Looker supplies governed business definitions and controlled access. Row policies restrict rows; column policy tags and masking address sensitive fields. An authorized view can expose a controlled projection without granting broad access to underlying datasets. Test as the consumer identity, not as an administrator.
Review details
A dry run estimates supported query processing without executing the query; a maximum-bytes-billed setting can reject an excessive on-demand query. Capacity/reservation management solves a different problem. Batch priority suits jobs that can tolerate scheduling delay; isolate business-critical interactive workloads appropriately.
SQL correctness before performance: check join multiplicity, null handling, aggregate grain and partition filters. WHERE acts before grouping; HAVING filters aggregates. A window function retains row detail while computing over its partition. A faster query with duplicated totals is still wrong.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Repeated daily time-range reporting | Partition on the relevant date and filter it. |
| Predictable isolated query capacity | Appropriate reservations and workload assignments. |
| Restricted consumer dataset | Authorized views and suitable row/column controls. |
Traps
- LIMIT is not a reliable query-cost cap.
- Clustering and partitioning help only when query patterns can exploit them.
06 · Governance, Sharing and AI-Ready Data
Memory hook: Catalog describes; policy controls; lineage explains.
Must remember
An enterprise data platform needs ownership, discoverability, classification, lineage and quality rules, not just a storage bucket. Dataplex Universal Catalog helps organize metadata and governance across supported assets. Metadata access does not automatically grant data access. Assign accountable stewards and measure quality dimensions such as completeness, freshness, validity and uniqueness.
Separate development, test and production identities and datasets. Apply least privilege at organization/project/dataset/table scope as appropriate. Sensitive Data Protection discovers/classifies and can de-identify supported content; masking is not equivalent to irreversible anonymization. Keep residency constraints, retention, deletion and access evidence together.
BigQuery sharing, associated with the Analytics Hub name, distributes governed shared datasets without ordinary full-copy workflows. Publisher and subscriber permissions, refresh behavior and commercialization rules remain separate design decisions. An exported file may escape later centralized revocation, so sharing mechanisms affect control.
Prepare AI features with point-in-time correctness: a model must not learn information unavailable at prediction time. Fit preprocessing on training data and apply it consistently. Data leakage can produce impressive offline scores and poor production results. BigQuery ML enables supported model workflows using SQL; embeddings encode similarity for retrieval, not guaranteed factual truth.
For retrieval-augmented generation, preserve document provenance and permissions through chunking, embedding, indexing and retrieval. A vector search result must still be authorized for the requesting user. Evaluate retrieval quality and final answer quality separately; stale indexes, missing context and prompt injection require explicit defenses.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Discover and trace enterprise data | Catalog, owners, classification and lineage. |
| SQL-oriented supported ML | BigQuery ML. |
| Share governed analytics datasets | BigQuery sharing with publisher/subscriber access design. |
Traps
- A catalog entry is not a grant to read the underlying data.
- Removing names alone does not prove a dataset is anonymous.
07 · Orchestration, Deployment and Data Operations
Memory hook: DAG orders work; monitoring proves it completed correctly.
Must remember
Cloud Composer runs managed Apache Airflow for dependency-driven data workflows. A DAG describes dependencies, scheduling and retries; the actual heavy processing belongs in services such as Dataflow, BigQuery or Spark. Workflows coordinates API calls with managed execution and branching when a full Airflow environment is unnecessary.
Parameterize environments and keep code/configuration in version control. Test transformations against representative fixtures, validate schema contracts, and separate unit, integration and reconciliation tests. Use CI/CD service identities with limited access and approval appropriate to production risk. Secrets belong in managed secret storage, not DAG source or task logs.
Retries must be safe. Write to staging then publish atomically where supported, use deterministic partition targets or transaction keys, and distinguish a retry from intentional backfill. Backfills can overwhelm quotas or overwrite historical data if date boundaries are wrong. Track run IDs, source offsets and output partitions.
Monitor job success, freshness, completeness, lag, rejected records and resource use. A green task can still publish an empty or stale table. Alert on user-facing data SLOs and actionable causes. Correlate logs with workflow and pipeline IDs; restrict sensitive payloads and retention.
For cost, terminate idle clusters, select suitable worker sizes/autoscaling, tune query scans, and isolate reservations or priorities where necessary. Compare steady capacity to per-job processing, including operator effort and recovery time. Quotas, regional capacity and dependencies are part of scheduling, not afterthoughts.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Complex recurring Airflow DAGs | Composer. |
| Coordinate a modest sequence of service APIs | Workflows. |
| Successful job but missing sales rows | Reconciliation/quality checks, not only process exit status. |
Traps
- The orchestrator should not carry a large dataset through task metadata.
- Blind retries can duplicate output or overwrite valid partitions.
08 · Migration, Consistency and Recovery
Memory hook: Copy a baseline, capture change, reconcile, then cut over.
Must remember
Start with data volume, transfer window, source load, supported engines and acceptable downtime. Storage Transfer Service moves supported object/file data; Transfer Appliance addresses appropriate offline bulk transfer constraints. Database Migration Service supports specific database migrations, while Datastream captures supported database changes for downstream pipelines. Product support must match the exact source and target versions.
A common migration sequence is full load, continuous change capture, reconciliation, controlled write cutover and observation. Check row counts, checksums or business totals, schema conversion, character encoding, time zones and sequence behavior. CDC transports changes; it does not automatically make a non-compatible application/schema portable.
Define RPO as tolerable data loss and RTO as restoration time. Replication supports availability but may reproduce bad writes or deletion. Backups and point-in-time recovery provide historical recovery subject to retention and service limits. Test restoration into an isolated environment with realistic permissions, keys and dependencies.
ACID addresses transactional properties; eventual consistency describes convergence behavior. Select consistency and transaction boundaries for the actual business invariant. A distributed pipeline may need deduplication and compensating actions even when each individual database operation is transactional.
Plan regional outages and missing/corrupt data separately. Managed database failover may change endpoints or connections; applications need reconnection and retry logic. Redis/cache recovery should not become the sole recovery strategy for authoritative data. Preserve replayable source records where lawful and economical, and document who can approve cutover or rollback.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Low-downtime supported relational migration | Baseline plus CDC, validation and controlled cutover. |
| Accidental corruption replicated everywhere | Recover from verified history rather than another damaged replica. |
| Recover stream outputs | Replay retained input with deterministic/idempotent processing. |
Traps
- A replica is not a complete backup strategy.
- A green transfer job does not prove business reconciliation passed.