certslothcertsloth
ADP/Topic 03

Google Cloud / Associate

Basic ML with SQL and managed tools

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Train on history; test on unseen data.

Must remember

  • Choose classification for categories, regression for numeric targets, forecasting for future time values and clustering for unlabeled groups. A deterministic SQL rule may be sufficient before introducing ML.
  • BigQuery ML lets analysts create and use models through SQL. CREATE MODEL trains supported model types; ML.EVALUATE measures quality; ML.PREDICT runs inference for applicable models.
  • Separate training, validation and test data. Split time-series data chronologically and avoid features that reveal the target or would not exist at prediction time.
  • Accuracy can mislead on rare events: compare precision, recall and the cost of false positives versus false negatives. Regression needs magnitude-sensitive error metrics; forecasts need time-aware evaluation.
  • AutoML automates parts of model training; pretrained models avoid training a general capability from scratch. BigQuery remote models use connections and authorized access to supported external model services.
  • Register models with versions and lineage so a deployment can be traced to its training data and evaluation. Monitor performance after deployment and plan controlled retraining.

Review details

Precision = TP/(TP+FP) asks how many predicted positives were correct; recall = TP/(TP+FN) asks how many actual positives were found. Missing rare fraud may demand recall; wasting analyst time may demand precision. A 99% accurate model that predicts “not fraud” for every row can still be useless on a 1% fraud dataset.

Model lifecycle in SQL: train with CREATE MODEL, evaluate on separate data, then call the appropriate inference function. ML.PREDICT serves supported prediction models; ML.FORECAST serves supported forecasting models. A remote generative model uses its supported function/API and authorized connection, not an assumption that every model supports the same SQL operation.

Choose under exam pressure

Requirement Choice and reason
Predict whether an invoice will be late Classification with a time-correct split and no future-payment leakage.
Generate summaries from warehouse text A supported remote generative model, with access, cost and quality checks.

Traps

  • Excellent training accuracy is not evidence of generalization.
  • AutoML still needs suitable data, target definitions and evaluation.

Active recall

1. Which command evaluates a supported BigQuery ML model?

ML.EVALUATE.

2. What is leakage?

Information unavailable at prediction time entering training or evaluation and inflating results.

3. Which metric matters when missed fraud is expensive?

Recall is important, but assess precision and investigation cost too.

4. Why store model versions?

Reproducible deployment, comparison, audit and rollback.

5. Must every prediction use an LLM?

No; select the simplest model or rule that meets the requirement.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.