certslothcertsloth
AIP-C01/Topic 08

AWS / Professional

GenAI Testing and Failure Diagnosis

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Freeze the test conditions, separate failure stages and verify the final business action.

Must remember

  • Maintain representative golden tasks, adversarial cases, no-answer cases, multi-turn interactions and tenant-boundary tests. Keep development and final evaluation sets separate. Record model, prompt, retrieval/index and tool versions to reproduce regressions.
  • Use deterministic schema/contract checks where possible, semantic metrics where useful and calibrated human or model judgement for subjective outcomes. An LLM judge needs a rubric, bias checks and validation against human decisions.
  • Decompose RAG failures into ingestion, retrieval, reranking, context assembly and answer generation. Decompose agent failures into planning, tool selection, argument formation, authorisation, execution and result interpretation. Fixing the wrong stage can increase cost without improving correctness.
  • Test prompt injection through user input, retrieved documents and tool output. Check output encoding, command/query injection, sensitive-data leakage, refusal boundaries and side effects. Guardrails require configuration and regression tests; they are not a universal policy engine.
  • Load-test realistic context sizes, concurrency, streaming and downstream dependencies. Monitor quota use and degradation paths. Canary a new prompt/model/index and retain an explicit rollback; a provider model update can change behaviour even if application code is unchanged.
  • Evaluate against business acceptance criteria: task completion, correct external state, user satisfaction, safety, latency and total cost. Document residual limitations and keep a human escalation option for cases that cannot be reliably automated.

Choose under exam pressure

Requirement Choice and reason
A new prompt improves averages but fails rare critical tasks Gate on risk-weighted cases and subgroup/task categories.
Generated JSON is malformed Enforce/validate supported structured output and handle validation failures.
Agent says it booked a meeting but calendar is empty Verify tool outcome and external state, not prose.

Traps

  • A unit test with mocked model output cannot measure model quality.
  • A benchmark score can hide unsafe tail cases.
  • Evaluation datasets can become contaminated by repeated tuning.

Active recall

1. Why freeze the retrieval index for a model comparison?

Otherwise changed evidence confounds the effect of the model change.

2. What should a judge rubric specify?

The criteria, scoring meaning, required evidence and handling of uncertain or conflicting cases.

3. What is a meaningful end-to-end agent assertion?

The intended authorised business action occurred correctly, with no unintended side effects.

4. Why include no-answer cases?

To test whether the system abstains instead of inventing unsupported information.

5. Which artifacts belong in rollback planning?

Compatible model/prompt versions, retrieval/index configuration, tool contracts and application code.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.