Memory hook: Measure the answer, the experience and the harm; averages can hide who fails.
Must remember
- Use representative held-out tasks, stable baselines and subgroup analysis. Compare quality, robustness, safety, latency, cost and business outcomes. A benchmark unrelated to the real workflow is weak evidence of suitability.
- BLEU emphasises n-gram precision against references and is associated with translation. ROUGE uses overlap/recall-oriented measures often applied to summaries. BERTScore uses contextual representations for semantic similarity. None alone proves factual correctness or useful business outcomes.
- LLM-as-a-judge can scale evaluation but introduces judge bias, model/version sensitivity and prompt dependence. Calibrate against human judgement and inspect disagreements. Bedrock evaluation capabilities support supported automated and human evaluation approaches.
- Evaluate RAG retrieval separately from answer groundedness and relevance. Evaluate an agent's tool use, action validity and completed task, not just its final wording. Include adversarial inputs and safe refusal/escalation tests.
- Responsible AI includes fairness, robustness, safety, privacy, transparency, explainability, accountability and veracity. Representative data, label review, human audits and subgroup metrics help find harmful differences hidden by aggregate accuracy.
- A transparent system exposes relevant workings and limitations; an explanation helps a person understand a particular result or behaviour. Model cards document intended use, evidence, limitations and risk. An open-source model is not automatically interpretable, safe or appropriately licensed.
- Consider intellectual-property rights, deceptive or biased outputs, environmental impact and user trust. Human-centred design needs clear AI disclosure, feedback and contestability where appropriate. Guardrails reduce specified risks but cannot certify that every response is harmless or true.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Summary wording differs but meaning is similar | Use semantic and human evaluation alongside overlap metrics. |
| High overall accuracy but poor outcomes for one group | Subgroup analysis and fairness investigation. |
| High-consequence decision | Human oversight, documented limits and a suitable error/appeal process. |
Traps
- Fairness has multiple definitions and trade-offs.
- A hallucination score is evidence, not a guarantee.
- Removing all sensitive columns does not necessarily remove proxy bias.
Active recall
1. Why can accuracy conceal unfairness?
Large groups dominate averages; inspect meaningful subgroups and error costs.
2. Can BLEU verify a factual claim?
No. Reference overlap does not independently establish truth.
3. What is the danger of using only one model as a judge?
Its own biases and blind spots can systematically favour poor answers.
4. What belongs in a model card?
Purpose, training/evaluation context, limitations, risk considerations and operating assumptions.
5. Why evaluate retrieval before generation?
A grounded answer requires relevant authorised evidence; missing evidence is a different failure from misuse of good evidence.