Memory hook: Batch for deadlines; online for requests.
Must remember
- Batch inference processes a collection without per-request interactive latency; online inference serves requests within a response target. Choose the simplest mode that meets the business deadline.
- Register model versions and package compatible prebuilt or custom serving containers. Include preprocessing/postprocessing and feature definitions so online behavior matches evaluation.
- Managed endpoints, Cloud Run and GKE offer different control and operations tradeoffs. Public/private access, authentication, networking and accelerator support must fit the model and consumers.
- Capacity planning includes model load time, memory, concurrency, tokens, request distribution and downstream feature latency. Autoscaling cannot instantly remove cold-start or quota constraints.
- Canary traffic limits exposure to a new version; A/B tests compare variants; shadow traffic can evaluate without returning new-model answers. Keep a tested rollback target and compatible features.
- Feature serving freshness and availability are part of the prediction SLO. Monitor latency percentiles and errors per version, not only average service latency.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Score all accounts before tomorrow morning | Batch prediction may be cheaper and simpler than permanent online capacity. |
| Safely replace a production model | A versioned canary with quality/latency thresholds and rollback. |
Traps
- A model that fits in storage may not fit in serving memory.
- Autoscaling the model does not automatically scale the feature database.
Active recall
1. Why use Model Registry?
To manage versioned model artifacts and controlled deployment lineage.
2. What does a canary limit?
The share of production traffic exposed to a new version.
3. What can cause prediction drift without changing weights?
Changed input data, preprocessing, features or serving configuration.
4. Why inspect tail latency?
Averages can hide slow requests that violate user-facing requirements.
5. When is a private endpoint useful?
When network access must be restricted to approved private paths in addition to identity controls.