Reviewed 10 October 2026. Use the linked official exam guide for your exam version. These are condensed revision notes; the topic pages provide worked distinctions and more recall practice. Google’s 2026 guides use newer Gemini Enterprise Agent Platform names while some APIs and documentation still use Vertex AI.
Memory hook: Bootstrap → build → release → measure → improve.
1. Organization and platform foundations — 1.1–1.5
Separate environments/projects and define owners, inherited IAM, organization policy, location constraints and billing. Shared VPC centralizes network control; PSC exposes supported private services; peering is not automatic transit. Give central logging/metrics appropriate cross-project IAM without granting workload administrators the ability to erase all evidence.
Terraform/Infrastructure Manager provision infrastructure; Config Connector and GitOps reconcile declared state; Helm packages Kubernetes objects. Assign one clear owner per resource. Protect state, review plans, parameterize environments and remove temporary environments through an owned lifecycle. Golden images/bootstrap scripts standardize Cloud Workstations/VM tooling; staged fleet patching and upgrades reduce blast radius. AI coding/operations assistance still needs review and evidence.
2. CI/CD and supply-chain trust — 2.1–2.4
| Function | Remember |
|---|---|
| Build/test source | Cloud Build; triggers and its build identity need controlled scope |
| Store artifacts | Artifact Registry; immutable digests, access and cleanup policy |
| Promote releases and manage rollouts | Cloud Deploy with supported targets, approvals and verification |
| Render/deploy Kubernetes configuration | Skaffold/Kustomize and the chosen delivery pipeline |
| Reconcile Git intent | A GitOps controller such as Argo CD; approved changes remain essential |
| Find known weaknesses / prove origin / enforce deployment trust | Artifact Analysis / provenance and attestations / Binary Authorization |
Build once, test, scan, record provenance, then promote the same digest. A clean scan does not prove trustworthy source; a signature does not prove no vulnerability. SLSA describes assurance of the build/source process. Separate build, deploy and runtime identities; prefer federation over exported keys.
Secret Manager stores credentials; Parameter Manager stores configuration; KMS manages cryptographic operations; Certificate Manager serves certificate-management needs. Runtime injection avoids baking secrets into image layers, but access and logs still require controls. Rotation must coordinate consumers. Deploy canary/blue-green/rolling according to capacity, risk and reversibility; feature flags separate activation from deployment. ML promotion also needs quality/data gates. Schema changes must remain compatible with the rollback version.
3. Reliability engineering — 3.1–3.3
Define user-relevant SLIs, target SLOs and any contractual SLA. For a request-based 99.9% SLO over 1,000,000 valid requests, the error budget is 1,000 bad requests. For time-based 99.9% over 30 days, it is 43.2 minutes. Do not mix denominators. Burn rate = observed bad-event fraction ÷ allowed bad-event fraction; 1% errors under a 99.9% objective burns at 10×.
Use an error-budget policy to balance releases and reliability work. Capacity planning includes quotas, reservations, startup delays, dependency capacity and failure headroom. MIG/HPA/Cloud Run autoscaling adjust different layers; reservations and Dynamic Workload Scheduler address specific capacity/scheduling needs. During an incident, mitigate through rollback, traffic drain/redirection or capacity as supported by evidence; then investigate and reduce repeat toil.
4. Observability and troubleshooting — 4.1–4.5
- Metrics quantify trends, logs explain events, traces connect latency across spans, profiles locate CPU/memory use. OpenTelemetry supplies portable instrumentation; Ops Agent serves VMs; Managed Service for Prometheus supports Prometheus metrics.
- Propagate trace IDs and record safe structured fields: service, version, severity, resource and outcome. Synthetic monitors exercise a real user path. Multi-project metrics scopes and log sinks require explicit access.
- Logs Explorer uses Logging queries; Log Analytics/BigQuery support analytical use cases. Sinks route new matching entries and need destination permission. Redact sensitive payloads; sampling/exclusions reduce cost but also evidence. High-cardinality labels can inflate metric cost.
- Alert on actionable symptoms and multi-window budget burn with ownership/runbooks. For a regression, compare deployment versions, request traces, dependency saturation, quota failures and recent configuration before scaling blindly. If telemetry itself fails, inspect agents/exporter credentials, collectors, sinks and ingestion delay.
5. Performance and FinOps — 5.1–5.2
Measure useful transactions/jobs per cost, not just cheap machine-hours. Right-size requests and instance types, remove idle infrastructure, tune concurrency, reduce wasteful scans/log retention, and assess network tiers/data transfer. Spot suits interruption-tolerant work; commitments suit stable demand after optimization. Validate Active Assist recommendations against rare peak or recovery workloads before acceptance.
Traps to catch
- CI is not CD. Artifact storage is not rollout control. A passed build is not user-visible health.
- Infrastructure rollback does not restore lost data. Extra replicas do not fix a single serialized bottleneck.
- Observability cost reduction must preserve required security and incident evidence.
Last-pass self-check
1. Why promote a digest rather than rebuild the same Git commit in each environment?
Dependencies or build conditions can change; a digest identifies the already tested immutable artifact.
2. A 99.9% request SLO currently has 1% errors. Burn rate?
10×, because 1% observed bad requests divided by the allowed 0.1% equals 10.
3. Which service stores artifacts and which manages release promotion?
Artifact Registry stores; Cloud Deploy manages supported delivery/rollout workflows.
4. CPU rises in one process while distributed latency is normal elsewhere. Which tool helps?
A profiler plus process/application context to find the costly code path.
5. Why can blindly following a rightsizing recommendation hurt reliability?
Observed history may omit failover headroom, rare peaks, scheduled jobs or resource requirements not captured by the recommendation.
Sources
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Projects, Organizations, Billing and CLI
Memory hook: Project contains; billing pays; IAM permits.
Must remember
The resource hierarchy places projects under folders and an organization where available. Projects contain resources, enabled APIs and quotas; billing accounts fund linked projects but are not simply their parent in the IAM hierarchy. Project names, unique IDs and numeric project numbers serve different purposes.
IAM allow policies can be inherited from ancestors; organization policies constrain allowed configurations rather than granting permissions. Cloud Identity/Google Workspace manages organizational identities and groups. Prefer group-based grants to many individual bindings, then review inherited permissions and applicable deny constraints.
Enable required service APIs in the intended project. A quota limits a metric such as resource count or API rate; requesting an increase does not guarantee physical capacity in a specific zone. Select regions for latency, availability, data-location requirements and product support, considering regional versus zonal resource scope.
Budgets alert; they are not automatically a hard spending cap. Export billing data for analysis, use labels and project organization for allocation, and review idle resources, retained disks, public addresses and network transfer. Linking billing and granting access to billing reports require appropriate billing permissions, distinct from workload administration.
gcloud config list and gcloud auth list inspect the active CLI configuration/identity. Named configurations help switch contexts; explicit --project, region and zone flags reduce ambiguity. Cloud Shell supplies a managed command environment but actions still use identity and authorization. Application Default Credentials used by libraries can differ from gcloud's active login: diagnose the actual caller.
Recall drill: explain why a user can administer a VM yet cannot view billing, or can view a project but cannot invoke an API that has not been enabled.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Central identity administration | Cloud Identity/Workspace groups with appropriate IAM grants. |
| Warn at a spend threshold | A billing budget and notifications, plus a separate control process if needed. |
| Prevent prohibited resource configurations | Organization policy constraints. |
Traps
- A budget is not a guaranteed automatic shutdown.
- Changing the gcloud project does not rewrite every script’s explicit project flag.
02 · IAM, Service Accounts and Data Security
Memory hook: Grant the action to the actual caller at the narrowest scope.
Must remember
An IAM binding associates a principal with a role on a resource, optionally under conditions. Basic roles are broad; predefined roles are service-oriented; custom roles bundle supported permissions for specific needs. Inherited allow grants can broaden access; deny policies and organization constraints have distinct effects. Removing one narrow grant does not remove an inherited grant elsewhere.
A service account is both an identity used by workloads and a resource whose use can be controlled. Attaching/acting as a service account and creating short-lived impersonated credentials require different permissions. Grant API permissions to the runtime identity, not merely to the human who deployed it.
Prefer attached workload identity, service-account impersonation or Workload Identity Federation over downloaded long-lived keys where supported. External federation exchanges trusted external identity for controlled Google access. Protect audience, attribute mappings/conditions and role bindings; a permissive trust mapping can expose many unintended callers.
Use Secret Manager for secrets and Cloud KMS for encryption-key control. Default encryption does not mean everyone should read the data. Customer-managed keys introduce key IAM, location, rotation, availability and destruction considerations. Separation between key administrators and data users reduces excessive privilege.
VPC Service Controls creates supported service perimeters to reduce data exfiltration; it complements IAM rather than replacing it. Identity-Aware Proxy controls supported application/tunnel access. Audit logs identify activity under their service-specific categories/settings; enable required data-access visibility deliberately and protect the sink destination.
For permission denied, establish the real principal, resource project, required permission, inherited/conditional/deny policies and any perimeter/org-policy restriction. Granting Owner to “make it work” conceals the diagnosis and creates risk.
Review details
Service Account User (roles/iam.serviceAccountUser) includes acting as the service account for supported resource attachment. Service Account Token Creator (roles/iam.serviceAccountTokenCreator) supports generating impersonated credentials and supported signing operations. Granting attachment rights is not the same as giving the human direct access to every resource the account can read, but launching code as that account can be a privilege-escalation path.
Application Default Credentials searches supported credential locations; it is not an IAM role. Workforce federation serves external people, workload federation software. To diagnose a denied call, distinguish authentication failure, missing permission, inherited deny/boundary, organization policy and VPC Service Controls. Principal access boundaries constrain eligible resources for supported access; they do not grant permissions.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Application accesses a bucket | Grant the workload identity the necessary bucket role. |
| CI outside Google Cloud | Federated short-lived credentials with narrowly defined trust. |
| Reduce exfiltration through supported managed APIs | VPC Service Controls plus IAM and data controls. |
Traps
- Service-account use permission is not identical to the service account’s resource permissions.
- A VPC Service Controls perimeter does not replace IAM.
03 · State, Backends and Drift
Memory hook: One binding, one writer, protected history.
Must remember
State maps configured instances to remote object IDs and stores attributes needed for planning. The default local backend uses a local state file. Remote backends place state in shared storage; locking, encryption and access-control behavior depend on the selected backend.
Locking prevents concurrent writers corrupting a state snapshot. Do not disable locks just to get past a busy run. Force-unlock is a recovery operation for your own stale lock after confirming no writer remains. It does not roll back infrastructure.
Configure storage in a backend block; backend configuration cannot refer to normal variables or resource outputs. Supply authentication through the recommended external credential mechanism. Hard-coded or command-line backend secrets may be cached or captured in plans. init -migrate-state transfers state after a reviewed backend change; -reconfigure treats the configuration as new rather than migrating existing state.
Drift is found when providers read real objects during planning. Decide whether configuration should restore the intended value or be updated to accept the external change. A refresh-only plan proposes state/output updates without proposing remote infrastructure changes; applying it still writes state.
CLI workspaces provide separate state instances for one configuration, but shared credentials/backend access can remain. Use separate roots and permissions when environments need strong isolation. Consumers of terraform_remote_state need access to the underlying snapshot, even though the data source exposes only root outputs.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Team collaboration | Shared protected state with supported locking and a controlled apply process. |
| Deliberate external change | Review drift and update configuration or deliberately reconcile state. |
Traps
- Remote state storage does not always mean remote execution.
- Copying state into Git exposes historical secrets and provides poor locking.
04 · GKE and Container Operations
Memory hook: Google runs the control plane; choose who manages the nodes.
Must remember
GKE Standard gives greater node-pool control; Autopilot manages more infrastructure and applies its supported workload/configuration model. Regional control planes improve control-plane availability; workload replicas and node placement still determine application resilience. Private networking choices govern node addresses and API access separately.
Authenticate to the intended cluster and verify the kubectl context/namespace before changes. Deployments manage stateless replicas and rollouts; StatefulSets manage stable identities and storage patterns; Services discover/expose workloads. Inspect Pods, events, logs, Services and EndpointSlices before changing cluster capacity blindly.
Store images in Artifact Registry and grant the actual pulling identity appropriate access. Image-pull failures can result from a wrong image path, permissions, network restrictions or missing artifacts. Kubernetes ServiceAccounts and Google IAM service accounts are different identities; Workload Identity Federation for GKE connects supported workload identity to Google API access without embedding static keys.
HPA changes Pod replicas; VPA recommends/adjusts resource requests under its configured mode; cluster/node autoscaling changes node capacity. Requests affect scheduling, limits constrain use, and disruption budgets influence voluntary maintenance. Do not expect a Pod autoscaler to manufacture node capacity instantly.
Node pools group node configuration. Plan upgrades, surge/disruption settings, maintenance windows and compatibility. A node count can be healthy while a workload is Pending because of affinity, taints, quota or a PVC. Persistent-volume topology and access modes must match placement.
GKE Enterprise features can help manage fleets, policy and multi-cluster environments. Choose them for actual governance/operational needs rather than assuming every small application requires a fleet platform.
Review details
Probe distinction: readiness removes an unhealthy Pod from eligible service endpoints; liveness restarts a failed container; startup permits slow initialization before normal probes take over. Do not use an aggressive liveness probe to respond to every downstream database timeout.
Read-only diagnosis sequence: kubectl config current-context → kubectl get pods,svc → kubectl describe pod POD → kubectl logs POD. Inspect events for image authorization, scheduling, volume attachment and probe failures. Image pulling often uses a different identity from the application making a Google API call; granting the latter access does not necessarily fix the former.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Minimal node administration | Autopilot when the workload fits its model. |
| Application needs Google API access | Workload Identity Federation with narrowly scoped IAM. |
| Pods Pending after replica increase | Check requests, placement, storage and node capacity. |
Traps
- A regional control plane does not automatically replicate your database.
- Kubernetes RBAC and Google IAM govern different parts of the access path.
05 · Cloud Run, Functions and Event Delivery
Memory hook: Revision receives; identity authorizes; retries repeat.
Must remember
Cloud Run runs containerized services without managing a VM fleet. Services receive requests; Jobs run finite tasks. A revision is an immutable deployment configuration. Split traffic between revisions for controlled release and rollback, and distinguish deploying a revision from sending it production traffic.
Scale and concurrency settings affect latency, cost and downstream connection pressure. Minimum instances reduce cold-start exposure while keeping capacity allocated; maximum instances can protect a backend but do not replace admission control or guarantee unlimited availability. Keep request handlers stateless and externalize durable state.
Cloud Run functions is the current function-oriented experience associated with the older Cloud Functions name. Select generation/runtime/event support deliberately. Eventarc routes supported events to targets; Pub/Sub decouples message producers and subscribers. Cloud Storage object events can trigger processing. Match region, trigger identity, target invocation permissions and event schema.
Design handlers for retries and duplicate delivery. Persist an idempotency key/result or make the operation naturally repeatable. A successful HTTP response acknowledges work; returning success before durable completion can lose business processing. Timeouts and retry policies must match the work and poison-event handling.
Invocation identity and runtime identity are distinct: one calls the service, the other determines what code may access. Restrict ingress and use authenticated invocation where appropriate. VPC egress configuration permits access to private dependencies; it does not automatically make all inbound requests private.
For a rollout issue, inspect the revision receiving traffic, request logs, container startup/listening port, service account, secrets/configuration, concurrency and backend limits. A healthy previous revision offers a rollback option only if data/schema compatibility remains intact.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Request-driven stateless container | Cloud Run service. |
| Finite batch container | Cloud Run Job. |
| Process object-created events | Eventarc/function or Cloud Run handler with idempotent processing. |
Traps
- Deploying a new revision and shifting traffic are separate operations.
- Event-driven does not mean duplicate-free business execution.
06 · SRE, Release Safety and Operational Excellence
Memory hook: Measure reliability from the user’s side.
Must remember
An SLI measures service behavior, an SLO sets a target and an SLA defines an external commitment/remedy. An error budget is the tolerated unreliability implied by an SLO over its window. Use it to make release/reliability tradeoffs; do not invent a universal acceptable percentage for every business.
Alert on symptoms that require action, with ownership and runbooks. Burn-rate-style alerts compare how quickly the budget is being consumed over suitable windows. Metrics, logs, traces and profiles answer complementary questions. Avoid paging for every transient utilization spike when users are unaffected.
CI validates changes; delivery/deployment moves approved artifacts through environments. Canary releases limit initial exposure; blue/green provides separate environments; rolling updates replace incrementally. Automated rollback needs trustworthy health signals and compatible data/schema. Feature flags separate activation from deployment but require lifecycle cleanup and secure access.
Use load testing for capacity, penetration/security testing for authorized attack resistance and chaos experiments for controlled failure hypotheses. Define blast radius, stop conditions and recovery before experiments. A test in staging may miss production-scale bottlenecks; justify what conclusions it supports.
Incident response should restore service, communicate clearly and preserve enough evidence for a blameless root-cause review. Fix system conditions, not only the person who made the last change. Reduce toil with bounded automation and improve runbooks through actual exercises.
Review capacity, quotas, dependencies, support plans, cost allocation and sustainability continuously. Gemini Cloud Assist and other tools can help investigate or propose changes, but verify recommendations against evidence and the intended environment. Operational excellence is sustained ownership, not a one-time dashboard installation.
Review details
For a request-based SLO, budget is allowed bad-event fraction × valid requests. A 99.9% target across 1,000,000 requests permits 1,000 bad requests. For a time-based 99.9% target over 30 days, the equivalent allowed bad time is 43.2 minutes. Do not mix request and time denominators.
Burn rate is the observed bad-event fraction divided by the budget's allowed bad-event fraction. With a 99.9% objective, a 1% error fraction burns at 10×. Combine suitable short and long windows to distinguish urgent sustained budget consumption from transient noise; the exact alert policy follows the service and response needs.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Rapid releases consume reliability tolerance | Use an explicit error-budget policy and stabilize the service. |
| Risky new version | Canary/staged deployment with meaningful user-health signals. |
| Repeated manual incidents | Automate understood recovery and remove the root cause. |
Traps
- An SLA and an internal SLO can have different targets and consequences.
- Rollback cannot automatically undo an incompatible data migration.
07 · Secure Build, Release and Platform Automation
Memory hook: Build once; attest it; promote the same artifact.
Must remember
Cloud Build executes builds/tests; Artifact Registry stores versioned artifacts; Cloud Deploy manages delivery to supported targets with releases, rollouts, approvals and promotion. Skaffold/Kustomize help render and deploy Kubernetes workloads. GitOps controllers reconcile declared state from version control; they still need trustworthy commits, restricted credentials and a recovery process.
Build once and promote an immutable digest through environments. Scanning identifies known vulnerabilities; provenance records origin and build process; signatures/attestations provide evidence; Binary Authorization can enforce configured deployment policy. SLSA describes supply-chain assurance levels. A clean vulnerability scan alone does not establish trusted provenance or absence of malicious logic.
Use short-lived workload federation for external automation where supported. Separate build, deploy and runtime identities; give each only needed permissions. Secret Manager stores secrets, Parameter Manager serves configuration use cases, and KMS manages cryptographic keys. Runtime secret injection avoids embedding credentials in images; build-time secrets can leak through layers, caches and logs.
Choose rolling, blue/green or canary release from capacity, reversibility and exposure needs. Define success metrics before shifting traffic, including latency/errors and business behavior. ML releases also need model/data quality and drift checks. Feature flags decouple code deployment from activation but create configuration/cleanup work. Database changes require backward-compatible migration if rollback to old code must work.
Bootstrap projects with reviewed Terraform/Infrastructure Manager or blueprints, remote state protection and policy checks. Config Connector manages supported resources through Kubernetes configuration; Helm packages Kubernetes resources. Separate production and temporary environments, enforce expiry/cleanup, and control fleet upgrades. Cloud Workstations offers managed developer environments; AI coding assistants still require review, tests and secret-handling discipline.
Review details
A useful release sequence is source review → repeatable build → unit/integration/security tests → digest and provenance → staging verification → approval where needed → progressive rollout → observed success or rollback. Configure trigger permissions and environment-specific identities as carefully as runtime IAM. Cloud Build logs and audit/deployment records establish which actor promoted which artifact.
Certificate Manager handles supported certificate-management needs, distinct from Secret Manager values and KMS key operations. Parameter Manager holds configuration. A GitOps controller continually reconciles desired state; manually patching production without changing its declared source can be reverted by the controller.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Reproducible promotion | Immutable artifact digest plus provenance. |
| Enforce trusted artifacts at deployment | Binary Authorization policy and attestations. |
| External CI without long-lived keys | Workload Identity Federation with scoped impersonation. |
Traps
- Cloud Build and Cloud Deploy have different responsibilities.
- Rolling back code cannot automatically undo an incompatible database change.
08 · Telemetry, Incident Diagnosis and FinOps
Memory hook: Metrics point; logs explain; traces connect; profiles locate cost.
Must remember
Collect platform and application telemetry with the Ops Agent, OpenTelemetry and supported integrations. Managed Service for Prometheus handles Prometheus-style metrics; Cloud Monitoring supports dashboards, alerts and SLO observation. Synthetic checks exercise a user journey externally. A server reporting healthy does not prove login and checkout work.
Use structured logs with severity, service, deployment version and trace correlation. Logs Explorer queries events; routing sinks send selected entries to destinations such as BigQuery, Pub/Sub or Cloud Storage. Exclusions/sampling reduce cost but can remove evidence. Protect audit/security logs and redact sensitive fields before export. Multi-project logging and metrics scopes need explicit IAM and ownership.
Distributed traces connect spans across services. Follow the critical path: downstream timeouts, repeated retries, lock contention or network hops may dominate a request. Profiles identify CPU/memory hot spots; high average CPU does not reveal which code path wastes it. Correlate a regression with a deployment before assuming infrastructure is undersized.
Alert on actionable user impact and rapid error-budget consumption. Route incidents to an owner with a runbook; avoid paging on every harmless transient. Stabilize by rollback, traffic drain or capacity increase as evidence supports, then investigate root causes. AI-assisted analysis is a hypothesis source, not authority to change production without verification.
FinOps joins engineering, finance and product decisions. Attribute costs by project/labels, remove idle capacity, rightsize requests, tune logs and data transfer, and select commitments for stable demand. Spot capacity suits interruption-tolerant work. Dynamic Workload Scheduler and reservations address specific scheduling/capacity needs. Compare cost per useful transaction/job, not only the cheapest VM hour.
Review details
Use the four golden signals—latency, traffic, errors and saturation—plus service-specific indicators. High-cardinality metric labels such as unbounded user IDs can increase cost and make analysis harder. Synthetic checks test an external path; they complement, rather than replace, real-user telemetry.
If observability is missing, trace the telemetry path itself: agent/exporter → credentials/network → ingestion → filter/exclusion → sink/destination → query/permissions. A logs-based metric does not automatically become a useful alert. FinOps recommendations are evidence to investigate; rare peaks and failure headroom may not be visible in a short observation window.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Requests slow across many services | Correlated distributed traces. |
| High CPU inside one process | A profiler plus application context. |
| Growing observability bill | Review retention, volume, cardinality, exclusions and sampling without losing required evidence. |
Traps
- An alert without an owner/runbook often creates noise.
- A recommendation is not evidence that a workload can tolerate its proposed reduction.