Memory hook: A model artifact is only one component of a reliable prediction service.
Must remember
- Match real-time, asynchronous, batch or supported serverless inference to traffic, payload, deadline and cold-start tolerance. Multi-model/multi-container deployment options trade consolidation against isolation and operational complexity. Check model/framework compatibility before choosing.
- Package preprocessing, model and postprocessing consistently. Register model versions with evaluation evidence and approval status. SageMaker Pipelines coordinates supported ML steps; CI/CD can promote approved versions through environments using immutable artifacts and scoped roles.
- Use shadow tests to observe a candidate without exposing its answers to users, or canary/blue-green strategies to control live exposure. Define rollback alarms on latency, errors and meaningful model outcomes. A successful endpoint deployment does not prove prediction quality.
- Autoscaling should reflect the actual bottleneck and startup time. Monitor invocation latency/errors, queue depth, utilisation and saturation. Optimisations include batching, model compilation/quantisation where supported, smaller models and suitable inference hardware, each validated for quality loss.
- Model Monitor and related tools help observe data/model quality and drift with appropriate baselines and captured data. Clarify supports bias/explainability capabilities. Delayed ground-truth labels require a separate outcome-collection process; input drift alone is not an accuracy measurement.
- Protect training and inference with separate least-privilege roles, encrypted storage, private connectivity where needed, restricted container/network access and redacted logs. Track prompts/tool calls for AI workloads with privacy-aware retention. Retraining needs governed triggers and approval, not blind reaction to every distribution change.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Evaluate a new model without affecting user responses | Shadow deployment with comparative evaluation. |
| Predictions are slow only after scale-out | Investigate loading/warmup and scaling policy. |
| Inputs shifted but labels are unavailable | Report drift evidence and collect outcomes before asserting accuracy loss. |
Traps
- A registered model is not necessarily approved for production.
- Automatic retraining can amplify bad incoming data.
- Autoscaling capacity is not unlimited or instantaneous.
Active recall
1. What makes a rollback viable?
A retained compatible artifact, configuration and tested traffic-switch process.
2. Why monitor business outcomes?
Infrastructure can be healthy while predictions fail the business objective.
3. What must accompany model versioning?
Feature/preprocessing versions, environment, data lineage and evaluation evidence.
4. Does input drift always require retraining?
No. Investigate whether it affects relevant performance and whether the new data is trustworthy.
5. Why separate training and inference roles?
They need different data and service permissions, reducing unnecessary access.