certslothcertsloth
AI-300/Topic 05

Azure / Associate

Evaluate and optimize generative systems

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Retrieve well, judge carefully, tune last.

Must remember

  • Create representative evaluation datasets with expected evidence, outcomes and risk cases. Map fields correctly and evaluate groundedness, relevance, coherence, fluency, safety and task completion as distinct properties.
  • Built-in and custom evaluators need calibration; human review catches ambiguous judgments and domain errors. Add automated evaluation to CI and sample production behavior under privacy controls.
  • Trace retrieval, model and tool steps with correlation IDs. Monitor latency, throughput, token/cost usage, errors and detailed safe diagnostics so a slow dependency is not mistaken for a model issue.
  • Tune RAG chunk size/overlap, embedding model, similarity threshold, filters, hybrid retrieval and reranking using measured relevance. A/B tests must compare equivalent user/task populations.
  • Fine-tuning changes behavior through additional training; use suitable curated or validated synthetic examples, supported methods and held-out evaluation. Do not use tuning as a substitute for fresh authorized knowledge retrieval.
  • Manage customized models through versioned deployment, monitoring and rollback. Optimize end-to-end useful outcomes rather than minimizing tokens at the expense of accuracy or safety.

Choose under exam pressure

Requirement Choice and reason
Answers are fluent but unsupported Inspect retrieval and groundedness, not only prompt tone.
Fine-tuning improves training examples but hurts unseen tasks Investigate overfitting/data quality and reject promotion until evaluation passes.

Traps

  • A low similarity threshold may add irrelevant context rather than improve recall usefully.
  • Synthetic examples can replicate the same model’s errors and bias.

Active recall

1. What does groundedness measure?

Whether claims are supported by the supplied evidence, within the evaluator’s limitations.

2. Why use hybrid retrieval?

Lexical and semantic methods can recover complementary relevant results.

3. What should a trace reveal?

The sequence, timing and outcomes of retrieval, model and tool operations.

4. When consider fine-tuning?

When measured behavior/task needs remain after simpler approaches and suitable training data exists.

5. How optimize cost safely?

Measure quality alongside model size, token use, caching, retrieval and capacity changes.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.