Memory hook: Version, evaluate, promote, observe, recover.
Reviewed 10 October 2026. Read this once, then answer the last-pass checks without looking.
Must remember by domain
| Domain | Rapid revision |
|---|---|
| MLOps infrastructure | Workspace organizes work; datastore describes a connection; data asset versions a reference; environment defines dependencies; component packages a step; registry shares supported assets. Pin source data as well as the reference. Compute instances suit development; scalable targets run jobs. |
| IaC and identity | Deploy workspaces/targets/networking with Bicep or CLI and reviewed GitHub Actions. Federated workload identity avoids stored deployment secrets. Workspace RBAC and storage/registry access remain separate. Private endpoints require working DNS and permitted dependency egress. |
| Model lifecycle | Experiment with notebooks/AutoML; make training repeatable with scripts, components, pipelines and MLflow. Record parameters, metrics, artifacts and data/code versions. Hyperparameter sweeps tune configured choices; distributed training also needs communication/data partitioning. Package preprocessing/feature-retrieval specification with the model. |
| Deployment and monitoring | Managed online endpoints serve low-latency requests; batch endpoints process asynchronous collections. Validate schema/authentication/networking, then allocate traffic progressively and retain rollback. Data drift changes inputs; concept drift changes input–target relationships; training-serving skew changes feature preparation. Healthy HTTP responses do not establish model quality. |
| GenAIOps infrastructure | Foundry project access, model deployment, connected resources and networking must be reproducible. Compare serverless APIs, managed compute and provisioned throughput by supported model and demand. Version prompts, templates, tools and retrieval settings alongside model versions. |
| Quality and observability | Fixed held-out datasets distinguish relevance, groundedness, coherence, fluency, safety and task success. Calibrate model judges with human review. Track inference quality and operational latency/throughput/token cost separately; trace retrieval and tool dependencies. |
| Optimization | Improve chunking, overlap, filters, embeddings, hybrid retrieval and reranking against relevance measurements. Fine-tuning changes learned behavior; RAG supplies fresh authorized knowledge. Validate synthetic training examples and retain independent evaluation cases. |
Release order and traps
Reproduce training → compare candidate with baseline → responsible-AI/security checks → register immutable version and runtime → staged deployment → observe useful outcomes → promote or roll back. A drift alert triggers investigation; retraining does not automatically authorize production promotion. A versioned data pointer is insufficient if its underlying file is overwritten.
Last-pass self-check
1. Fast endpoint, declining accuracy: which monitoring is missing?
Prediction/business-quality monitoring, including delayed labels and subgroup performance.
2. What must travel with a model artifact?
Compatible runtime, preprocessing/features, input contract and lineage.
3. Why evaluate prompts on a fixed set?
To distinguish genuine behavior changes from different test populations.
4. Fresh policy documents: tune or retrieve?
Retrieve current authorized evidence; tuning is not a freshness mechanism.
5. When should automated retraining replace the champion?
Only after defined evaluation, safety and release gates pass.
Sources
- Official exam scope and version
- Product documentation
- Product documentation
- Product documentation
- Product documentation
- Product documentation
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Reproducible ML infrastructure
Memory hook: Code, data, environment: version all three.
Must remember
- An Azure Machine Learning workspace organizes jobs, models and assets. Datastores describe supported storage connections; data assets version references to data; neither automatically guarantees that underlying mutable data stays unchanged.
- Compute instances suit development; compute clusters and other supported targets run scalable jobs. Configure identity, networking, quota and auto-shutdown/scale-down to match workload and cost requirements.
- Environments define runtime dependencies; components package reusable steps; registries share supported assets across workspaces. Pin versions/digests so a pipeline can be reproduced.
- Use Bicep/Azure CLI and reviewed configuration to provision infrastructure. Separate development/test/production and pass environment-specific settings without hard-coded credentials.
- GitHub Actions can authenticate through federated identity rather than stored long-lived secrets. Grant the workflow identity only required deployment and data permissions.
- Private endpoints, managed networking and storage access must work together. A private workspace alone does not prove that compute can reach every required package, registry or datastore.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Reuse a tested transformation across teams | A versioned component/environment in a supported registry. |
| Provision identical workspaces repeatedly | Reviewed IaC with separate identities and approved environment parameters. |
Traps
- A versioned asset pointing to overwritten files is not truly reproducible data.
- Network isolation can break dependency downloads if required paths are not designed.
02 · Train, register and deploy models
Memory hook: Experiment freely; promote with evidence.
Must remember
- Track runs with MLflow, including parameters, metrics, artifacts and source/data versions. Notebooks support exploration; training scripts and components support repeatable jobs.
- AutoML explores supported model choices; hyperparameter sweeps optimize configured training choices. Distributed training requires compatible frameworks, data partitioning and sufficient communication bandwidth.
- Build pipelines for preparation, training and evaluation. Avoid leakage, use appropriate validation splits and compare models against an agreed business baseline.
- Register the model with its runtime and feature retrieval/preprocessing specification. A model artifact without compatible features and dependencies is not a deployable solution.
- Managed online endpoints suit low-latency requests; batch endpoints suit asynchronous collections. Test authentication, network paths, input schema, resource limits and failure behavior.
- Progressive rollout, traffic allocation and known-good versions support safe rollback. Responsible AI evaluation and release gates should precede promotion; archive superseded artifacts according to policy rather than losing lineage.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Nightly scoring of a large dataset | A batch endpoint/workflow with completion and correctness checks. |
| Interactive prediction with a strict latency target | A sized online endpoint and measured tail latency. |
Traps
- The highest validation score is not sufficient if the model violates latency, fairness or cost requirements.
- Registering a model does not automatically deploy it.
03 · Monitor traditional ML in production
Memory hook: Healthy endpoint is not healthy prediction.
Must remember
- Monitor service latency, throughput, errors and resource saturation separately from prediction quality. An endpoint can return fast, valid JSON with poor business outcomes.
- Data drift changes input distribution; concept drift changes the relationship between features and target; training-serving skew comes from inconsistent feature preparation or availability.
- Collect appropriate inference data with privacy, sampling and retention controls. Delayed labels mean quality measurements can lag operational symptoms.
- Use thresholds and comparisons that account for seasonality and sample size. Investigate missing features, upstream schema changes and model version differences before retraining automatically.
- Configure alerts and retraining triggers with named owners and evaluation gates. A retrained model can fail acceptance and should not automatically replace the current champion.
- Measure subgroup quality and business error cost. Document intended use, limitations and the rollback or human-review path when confidence degrades.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Endpoint latency is normal but error outcomes rise | Investigate model quality, data changes and labels rather than only compute. |
| Input drift alert fires during a known seasonal event | Compare expected patterns and outcome evidence before promoting a new model. |
Traps
- Drift is not automatic proof that retraining will improve performance.
- More frequent retraining can amplify bad or mislabeled data.
04 · Foundry infrastructure and prompt lifecycle
Memory hook: Version the prompt like application code.
Must remember
- Foundry resources/projects organize supported AI development and deployment. Configure managed identities, RBAC, private networking and connected resources rather than relying on a developer’s broad login.
- Choose supported serverless/model API or managed-compute deployment by model, region, data requirements, control and cost. Availability, quotas and supported deployment types vary across models.
- Provisioned throughput can suit predictable sustained demand; consumption pricing suits other patterns. Measure token rates, concurrency, latency and utilization before choosing capacity.
- Pin or deliberately manage model versions and deployment names. A provider model update can change output behavior even when application code stays the same.
- Keep prompts, templates, tool definitions and retrieval settings in Git. Compare variants on the same evaluation set; capture the full configuration required to reproduce a response.
- Use Bicep/CLI and CI workflows for repeatable infrastructure and staged releases. Keep secrets in supported secure stores, avoid logging private prompt content indiscriminately and retain rollback configurations.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| High-volume steady token demand | Evaluate provisioned throughput with utilization and capacity evidence. |
| Prompt change improves one demo but harms other users | Run a representative regression evaluation before release. |
Traps
- A model deployment name does not guarantee immutable behavior forever.
- Private networking does not replace model/project authorization.
05 · Evaluate and optimize generative systems
Memory hook: Retrieve well, judge carefully, tune last.
Must remember
- Create representative evaluation datasets with expected evidence, outcomes and risk cases. Map fields correctly and evaluate groundedness, relevance, coherence, fluency, safety and task completion as distinct properties.
- Built-in and custom evaluators need calibration; human review catches ambiguous judgments and domain errors. Add automated evaluation to CI and sample production behavior under privacy controls.
- Trace retrieval, model and tool steps with correlation IDs. Monitor latency, throughput, token/cost usage, errors and detailed safe diagnostics so a slow dependency is not mistaken for a model issue.
- Tune RAG chunk size/overlap, embedding model, similarity threshold, filters, hybrid retrieval and reranking using measured relevance. A/B tests must compare equivalent user/task populations.
- Fine-tuning changes behavior through additional training; use suitable curated or validated synthetic examples, supported methods and held-out evaluation. Do not use tuning as a substitute for fresh authorized knowledge retrieval.
- Manage customized models through versioned deployment, monitoring and rollback. Optimize end-to-end useful outcomes rather than minimizing tokens at the expense of accuracy or safety.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Answers are fluent but unsupported | Inspect retrieval and groundedness, not only prompt tone. |
| Fine-tuning improves training examples but hurts unseen tasks | Investigate overfitting/data quality and reject promotion until evaluation passes. |
Traps
- A low similarity threshold may add irrelevant context rather than improve recall usefully.
- Synthetic examples can replicate the same model’s errors and bias.