Memory hook: Database runs the transaction; warehouse explains the pattern.
Must remember
Structured data follows a defined schema, semi-structured data carries flexible structure such as JSON, and unstructured data includes images/audio. First-party data comes from an organization’s own relationships; second-party data is another party’s directly shared data; third-party data is obtained through external aggregation/providers. Permission and quality matter regardless of source.
Cloud SQL manages familiar relational engines; AlloyDB targets PostgreSQL-compatible workloads; Spanner addresses globally scalable relational transactions; Firestore stores application documents; Bigtable serves large low-latency wide-column/key-based workloads. BigQuery supports analytical queries. Cloud Storage stores objects and can support a data lake; a lake is not automatically a governed warehouse.
Cloud Storage Standard suits frequent access; Nearline, Coldline and Archive trade cheaper storage for minimum-duration/retrieval considerations. Autoclass can manage supported transitions automatically. Select total cost and access needs rather than the cheapest storage line item.
Data pipelines collect, process, store, analyze and activate information. Pub/Sub decouples event delivery; Dataflow transforms batch/stream data; managed Spark supports Spark processing. Looker makes governed metrics and dashboards available to decision makers. Catalogs, lineage, ownership, retention and quality controls prevent an accessible lake becoming an untrusted data dump.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Operational relational application | Cloud SQL/AlloyDB according to engine and scale needs. |
| Large-scale SQL analytics | BigQuery. |
| Low-latency stream transformation | Pub/Sub with an appropriate Dataflow pipeline. |
Traps
- Object storage is not a relational database.
- A dashboard is only as trustworthy as its definitions and input data.
Active recall
1. Warehouse versus database?
Analytical patterns across data versus operational application transactions, broadly speaking.
2. When choose Spanner?
When supported globally scalable relational consistency/availability requirements justify it.
3. What does Looker contribute?
Governed business definitions, exploration and visualization.
4. Why govern data?
To keep it discoverable, trustworthy, authorized and retained appropriately.
5. Why might Archive cost more than expected?
Retrieval and minimum-storage-duration charges can outweigh lower storage rates.