Reviewed 10 October 2026. Use the linked official exam guide for your exam version. These are condensed revision notes; the topic pages provide worked distinctions and more recall practice. Google’s 2026 guides use newer Gemini Enterprise Agent Platform names while some APIs and documentation still use Vertex AI.
Memory hook: Workload → engine → connections → recovery → measured operations → cutover.
1. Design and size — 1.1–1.4
Measure peak transactions, reads/writes, joins, latency, connections, working-set RAM, IOPS/throughput, storage growth and maintenance. Size from representative workload and failure headroom, not database file size alone. Include compute, license, storage, backup, network and operator cost. Managed services reduce administration but impose engine/version/extension boundaries; self-managed, partner or bare-metal offerings can fit specialized constraints.
| Workload | Candidate and decisive distinction |
|---|---|
| Familiar MySQL/PostgreSQL/SQL Server | Cloud SQL; validate extensions, HA, replication and engine limitations |
| Demanding PostgreSQL-compatible system | AlloyDB; primary/HA and read pools have different roles |
| Horizontally scaled relational transactions | Spanner; distributed schema, key distribution and transaction locality matter |
| High-throughput wide-column/key-range access | Bigtable; row keys and app-profile routing determine performance/consistency |
| Document-oriented applications | Firestore; model document transactions and indexed queries |
| Low-latency cache/session data | Memorystore; define durability, eviction and source of truth |
| Analytical warehouse | BigQuery; not a replacement for every OLTP workload |
SQL versus NoSQL is about access and transaction patterns, not absence of schema. Vector search compares embeddings; choose dimensions/model version, metric, index and metadata filtering. It does not replace exact key lookup, transactional integrity or authorization. A multi-store design earns its extra operational burden only when distinct access patterns justify it.
2. Access, operations and recovery — 2.1–2.5
Connection diagnosis has three gates: network → identity → engine privileges. Private IP/PSC/private services access provide supported paths; they do not grant SELECT. An auth proxy/connector can handle supported authentication/TLS but is not a VPN. Cloud IAM administers resources and supported IAM database authentication; database roles govern engine objects. Secret rotation needs coordinated clients, not only a new stored secret.
Connection budget = per-instance pool × maximum application instances, plus jobs/admin/headroom. Set timeouts and bounded retries; connection storms during autoscaling or failover can overwhelm a healthy database. Scope audit and query logs without exposing secrets unnecessarily. Key IAM, location and availability matter for CMEK-backed production and restore.
Cloud SQL regional HA synchronously protects committed writes across its primary/standby zones and supports failover; the standby is not a read-serving replica. Read replicas generally serve asynchronous read scaling; regional disaster recovery must account for lag and promotion. Spanner quorum/location, AlloyDB HA/read pools, Bigtable routing/replication and Firestore location semantics differ—avoid treating every secondary as equivalent.
Backups preserve recovery points; transaction logs support eligible PITR; exports support portability and selected recovery uses. RPO bounds data loss; RTO includes provisioning, restore/replay, permissions/keys, connection changes and application validation. A replicated bad write remains bad on replicas. Test restore into a safe target before relying on the retention setting.
For slowness: query plan → rows scanned/selectivity → indexes/joins → locks/transaction duration → I/O/memory/CPU → connections → replication lag. Indexes trade faster selected reads against write cost and storage. Shorten transactions and standardize lock ordering before assuming bigger hardware solves contention. Read replicas do not scale all writes; hot keys can serialize an otherwise large system. Automate supported exports, upgrades and maintenance with alerts, ownership and rollback planning.
3. Migrate with a data-safe fallback — 3.1
Inventory engine/version, stored code, extensions, schema, character sets/collation, constraints, sequences and dependencies. Homogeneous migration still needs validation; heterogeneous migration also needs schema/query/application conversion. Offline export/import is simple when downtime is allowed; supported full-copy plus CDC minimizes final outage.
Sequence: validate target/connectivity/permissions → initial copy → catch up changes → compare counts/checksums/business totals → stop/control source writes → confirm lag threshold → switch applications/jobs/caches → validate transactions and performance. Define go/no-go criteria and source-retention period. Reverse replication/fallback is product/path-specific; changing DNS back after target writes can lose data without reconciliation.
4. Implement and prove HA — 4.1
Provision through reviewed IaC with monitoring, backup policy, deletion safeguards and least-privileged identities. Exercise zone loss, region loss, primary promotion, read-replica scaling, maintenance and unavailable keys/credentials. Test whether clients reconnect and whether a stale read violates a business invariant. Measure the service outcome and retained data, not just a successful control-plane failover operation.
Traps to catch
- Standby ≠ read replica; replica ≠ backup; export ≠ automatically complete PITR.
- Cloud administrator ≠ every SQL privilege; proxy ≠ network route.
- Schema conversion success ≠ application compatibility; fallback needs new-write handling.
Last-pass self-check
1. Why can a Cloud SQL HA standby not solve a read-heavy reporting bottleneck?
It is maintained for failover rather than query-serving read scale; assess a compatible read replica and acceptable lag.
2. A query is blocked behind a long transaction. First remedy?
Identify and reduce the locking/transaction problem; more CPU may not remove serialization.
3. Why cap pools when Cloud Run scales?
Every new application instance may create more database connections; total pools can exceed database capacity.
4. After migration, why is DNS-only fallback dangerous?
New writes on the target may not exist at the source; reconcile or use supported reverse replication before returning traffic.
5. Which recovery test catches a forgotten CMEK permission?
A realistic restore and application validation using the intended recovery identity and key dependency.
Sources
- Official exam guide
- Published objective groups (PDF)
- Cloud SQL HA
- Cloud SQL replicas
- Database migrations
- Bigtable schema design
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Choose a database from the workload
Memory hook: Access pattern before product name.
Must remember
- Cloud SQL offers managed MySQL, PostgreSQL and SQL Server; AlloyDB targets demanding PostgreSQL-compatible workloads with managed architecture and read pools. Check extension and feature compatibility rather than assuming all PostgreSQL systems are interchangeable.
- Spanner provides horizontally scalable relational transactions with strong consistency. Bigtable suits high-throughput key/range access; Firestore suits document-centric applications; Memorystore supplies low-latency caching. BigQuery is an analytical warehouse.
- Assess transaction shape, joins, consistency, data size, peak concurrency, working-set memory, IOPS, throughput and growth. Benchmark representative traffic rather than sizing solely from database file size.
- Managed services reduce operations but impose supported engines, versions, extensions and configuration boundaries. Self-managed or partner/bare-metal offerings can meet requirements outside those boundaries at higher operational responsibility.
- Include region availability, data residency, organization policy, licensing, network egress and backup retention in the design. Multiple databases may be appropriate when each serves a distinct access pattern.
- Vector retrieval can support AI grounding, but choose index type, filtering, distance metric and embedding lifecycle deliberately. A vector index does not replace transactional correctness or permission checks.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Global relational transactions with horizontal growth | Evaluate Spanner with schema and latency testing. |
| Existing compatible PostgreSQL application | Compare Cloud SQL and AlloyDB before rewriting for a different data model. |
Traps
- NoSQL is not a synonym for no schema design.
- A cache is not automatically a durable system of record.
02 · Connectivity, identities and protection
Memory hook: Reachability, identity, privileges: three gates.
Must remember
- Private connectivity removes a public path but does not grant database permission. Plan VPC routing, DNS, private services access or Private Service Connect according to the chosen database.
- Cloud SQL/AlloyDB connectors and auth proxies simplify supported authenticated TLS connections. They do not magically create a missing private network route or remove engine-level authorization.
- IAM controls cloud resource operations and supported IAM database authentication; database users/roles control SQL privileges. Separate administration, application access and migration identities.
- Pool connections to avoid exhausting database limits during application autoscaling. Set bounded pool sizes, connection lifetimes, timeouts and retry behavior; multiply per-instance pools by maximum instance count.
- Use TLS in transit and appropriate managed/CMEK encryption at rest. Plan key permissions, rotation and availability before restricting keys; secret rotation must coordinate clients and database credentials.
- Audit sensitive access and changes at the relevant cloud and engine layers. Log enough to investigate without exposing query parameters or credentials unnecessarily.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Many short-lived application instances | Bounded connection pooling and a database connection budget. |
| Application can reach the host but SELECT fails | Inspect database authentication and object privileges. |
Traps
- Cloud IAM administrator access does not imply every SQL data privilege.
- An auth proxy is not a general-purpose VPN.
03 · Availability, backups and deployment
Memory hook: Replica for continuity; backup for history.
Must remember
- RPO defines acceptable data loss; RTO defines acceptable recovery time. Map both to actual failover, restoration, DNS, application reconnection and validation steps.
- Cloud SQL regional HA uses a standby for failover; a read replica serves reads and has different promotion/replication behavior. Cross-region replicas support regional recovery but can lag.
- Spanner placement and quorum, AlloyDB HA/read pools, Bigtable replication and Firestore location choices have different consistency and failover semantics. Test the chosen service rather than assuming every replica behaves alike.
- Backups, transaction logs/PITR and exports solve different needs. Replication can copy accidental deletion; recoverability requires retained history and restore permissions.
- Automate provisioning through reviewed infrastructure code, include backup policy and monitoring, and protect production deletion paths. Schedule maintenance windows and rehearse compatible upgrades.
- Test zone loss, regional loss, credential/key availability and restoration to a clean environment. Measure data correctness and application health after recovery, not merely resource creation.
Review details
In Cloud SQL regional HA, committed writes are synchronously protected across the primary/standby zones, and the standby serves failover rather than ordinary read scaling. Read replicas use different replication/promotion semantics and can return stale data. Test client reconnection after failover rather than assuming existing TCP sessions survive.
For Bigtable, single-cluster routing versus multi-cluster routing changes availability/consistency behavior; app profiles matter. For Spanner, replica placement and quorum drive availability/latency. A Firestore location is a deployment choice with service-specific replication behavior. Record the chosen service's actual guarantees rather than assuming all “replicas” are identical.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Accidental DELETE propagated to replicas | Restore using retained backups/PITR to a safe target. |
| Read load is exhausting a primary | A compatible read replica/read pool, accounting for lag and consistency needs. |
Traps
- A standby is not necessarily a query-serving read endpoint.
- A stated SLA is not a tested end-to-end application RTO.
04 · Diagnose and optimize operations
Memory hook: Find the bottleneck before adding capacity.
Must remember
- Correlate CPU, memory, storage latency, IOPS, connection count, locks, replication lag and query latency. A busy CPU and blocked queries require different remedies.
- Use query plans and query insights to find scans, poor selectivity, missing indexes and expensive joins. Indexes improve selected reads but consume space and slow writes; validate changes against the workload.
- Resolve lock contention by shortening transactions, using consistent access order and choosing appropriate isolation. Increasing machine size does not necessarily remove a logical deadlock.
- Scale vertically for per-node resource pressure and horizontally where the engine supports it. Read replicas do not distribute all writes; hot keys can defeat otherwise large capacity.
- Budget compute, storage, backups, network and licenses. Right-size after measuring peaks and growth; check quotas before scaling and maintenance before committing to a savings strategy.
- Automate exports, supported maintenance and upgrades with least privilege, retries, alerting and an owner. Monitor user-facing SLOs as well as database vitals; alert on symptoms with a useful response.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A query scans millions of rows for one record | Inspect its filter and index plan before upgrading the instance. |
| Writes bottleneck on one key | Redesign the key/access pattern or transaction flow; more replicas may not help. |
Traps
- Every slow query does not need another index.
- An automated task without failure alerts can silently stop protecting data.
05 · Migration and cutover
Memory hook: Assess, copy, catch up, switch, verify.
Must remember
- Inventory engines, versions, schema, extensions, collation, stored code, dependencies and operational requirements. Homogeneous moves differ from heterogeneous moves requiring DDL/DML or application conversion.
- Offline export/import accepts downtime; continuous replication/CDC reduces the final outage for supported source-target combinations. Database Migration Service supports specific paths, not every possible engine conversion.
- Establish secure connectivity and permissions, take the initial copy, then catch up changes. Validate row counts, checksums where appropriate, constraints, sequences, queries and performance.
- Cutover usually requires stopping or controlling source writes, confirming acceptable lag, redirecting applications and checking business transactions. Coordinate caches, connection pools, DNS and scheduled jobs.
- Define a go/no-go threshold and fallback before migration. Returning to the source after new target writes requires reconciliation or supported reverse replication; merely changing DNS can lose data.
- Keep source retention, backups, monitoring and rollback responsibility explicit until acceptance. Remove temporary access only after the agreed recovery window.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A small database can tolerate a maintenance outage | Tested export/import may be simpler than continuous replication. |
| Minimal downtime is essential | A supported initial-copy-plus-CDC migration with rehearsed cutover. |
Traps
- Schema conversion does not prove application behavior is compatible.
- Near-zero downtime is not zero planning or zero risk.