Reviewed 10 October 2026 · AI Practitioner
Memory hook: Predict, generate or act; then prove value, safety and access.
AI/ML 20%; generative AI 24%; foundation-model applications 28%; responsible AI 14%; security and governance 14%.
Use this as a final revision pass after the chapters. Each task below maps to the published exam outline; the outline itself is not an exhaustive list of possible questions. Recheck the official guide for your booked exam version, especially beta releases.
Must remember by exam objective
1.1 — AI terminology and learning approaches
- AI is the broad category; ML learns patterns; deep learning uses multilayer neural networks; generative AI produces content. Supervised learning uses labels, unsupervised learning finds structure, and reinforcement learning optimizes rewards. Classification predicts categories; regression predicts numeric values.
1.2 — Use cases and service selection
- Choose managed task APIs when they fit: Textract for document extraction, Rekognition for image/video analysis, Comprehend for language insights, Transcribe for speech-to-text, Polly for text-to-speech, Translate for translation and Lex for conversational interfaces. Stable rules/arithmetic may need no ML.
1.3 — The ML lifecycle and its metrics
- Collect and label data, split training/validation/test sets, transform consistently, train, evaluate, deploy and monitor. Prevent leakage and training/serving skew. Accuracy can hide rare-class failure; precision measures positive-prediction reliability, recall measures detection of actual positives, and regression uses error measures such as MAE/RMSE.
2.1 — Generative and agentic AI concepts
- Foundation models are broadly pretrained; tokens consume context and generation budgets; embeddings represent similarity. Transformers commonly serve language tasks and diffusion serves many image-generation tasks. Agents choose actions/tools; fixed workflows follow explicit control paths. Neither agent memory nor a retrieved page is automatically trustworthy.
2.2 — Business value and model limitations
- Measure verified business outcomes, adoption, latency, safety and cost per successful task. Hallucinations, bias, stale training knowledge, limited context and non-determinism constrain use. A larger model or lower temperature does not guarantee truth; high-impact decisions need appropriate review.
2.3 — AWS AI platforms and cost trade-offs
- Bedrock offers managed foundation-model application capabilities; SageMaker AI offers custom ML development/training/hosting control. Compare on-demand inference, eligible provisioned throughput, batch processing and custom endpoint cost using utilization and requirements. Include storage, retrieval, tokens and failed retries.
3.1 — Foundation-model application design
- A RAG application ingests, cleans, chunks, embeds, indexes, retrieves and generates with citations. Knowledge Bases can manage supported parts; vector stores retain searchable representations. Match model modality, context, language, latency, privacy, availability and license to the task.
3.2 — Prompting and prompt management
- Prompt with task, context, examples, output schema and constraints. Zero-shot has no worked example; few-shot includes demonstrations. Version templates and test on held-out tasks. Temperature/top-p affect sampling; asking the model to ignore attacks is weaker than enforceable tool/data boundaries.
3.3 — Model training and adaptation
- Prompting changes context; RAG adds retrieved facts; fine-tuning changes weights for learned behavior; continued pretraining adapts to domain text; distillation transfers capabilities to a smaller model. Full pretraining is a much larger investment. None replaces lawful, representative, well-labeled data.
3.4 — Evaluation methods
- Evaluate retrieval relevance separately from answer groundedness and task completion. BLEU/ROUGE measure reference overlap; semantic metrics and human review answer different questions. Calibrate LLM judges against people and inspect disagreements. Agents need correct tool selection, arguments and verified external outcomes.
4.1 — Responsible AI practices
- Assess fairness, robustness, safety, privacy, security, accountability, veracity and human oversight across the lifecycle. Inspect subgroup errors and proxy bias; data balancing alone cannot prove fairness. Provide escalation, feedback and a way to stop unsafe actions.
4.2 — Transparency and explanations
- Transparency communicates purpose, limits and evidence; explainability helps understand behavior/results. Model cards record intended use and evaluation context; SageMaker Clarify supports bias/explainability analysis. A fluent explanation is not proof of the model’s internal reasoning or factual correctness.
5.1 — Securing AI applications
- Use least-privilege workload roles, TLS, KMS, scoped private connectivity and managed secrets. Prompt injection comes through user text, retrieved content or tool output; poisoning corrupts data. Guardrails apply configured filters, but authorization must constrain tools before actions occur.
5.2 — Governance and compliance
- Track data rights, consent, provenance, lineage, residency and retention across sources, embeddings, prompts, logs and memory. CloudTrail audits supported APIs; Config evaluates configuration; Artifact supplies AWS evidence. Service certification does not certify a customer’s entire AI use case.
Choose under exam pressure
| Deciding clue | Recall the distinction |
|---|---|
| Frequently changing internal policy answers | RAG over authorized current sources. |
| Stable specialized response style | Evaluate prompting, then supported fine-tuning. |
| Custom training/hosting control | SageMaker AI; managed FM access points to Bedrock. |
| A tool changes financial records | Policy checks, validated parameters and appropriate approval before execution. |
| Model sounds confident but fails the task | Measure verified task outcomes and groundedness, not confidence alone. |
Traps
- Embeddings are not encryption.
- RAG does not change model weights; tuning is not a live document lookup.
- Guardrails reduce configured risks without guaranteeing safety or correctness.
Verification cues
- Draw one RAG request from caller identity through document authorization to cited answer.
- For each metric, explain which failure it can reveal and which it cannot.
Last-pass active recall
1. Why can high accuracy miss fraud?
Rare fraud cases can be overwhelmed by correct negatives; examine recall, precision and error cost.
2. When should you retrieve rather than tune?
When answers need current external facts or per-user authorized evidence.
3. Can a document instruct an agent to bypass policy?
No. Treat source text as data and enforce tool permissions separately.
4. Does open-source mean explainable?
No. Availability of code/weights does not establish interpretability, safety or appropriate licensing.
5. What must an agent evaluation verify beyond its final sentence?
The chosen tools, valid arguments, authorization, actual action result, task completion, latency and total cost.
Sources and version check
The numbered chapters provide worked distinctions and further technical sources. These are original revision notes and original recall scenarios, not real exam questions.
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · AI, Machine Learning and Service Selection
Memory hook: Predict a label, predict a number, find a group, or generate content: these are different jobs.
Must remember
- AI is the broad field; machine learning learns patterns from data; deep learning uses multilayer neural networks. Generative AI creates content; agentic AI combines models with tools and control logic to pursue tasks. These categories overlap.
- Supervised learning uses labelled examples: classification predicts a category, regression a numeric value. Unsupervised learning finds structure without target labels, such as clustering. Reinforcement learning improves a policy through rewards from interaction; it is not simply a model with more labelled rows.
- Tabular and time-series data are different from unstructured text, image and audio. Preserve sequence in time-series validation; future information must not leak into training features. Labels must represent the real business outcome.
- Training changes model parameters; inference uses the resulting model. A deterministic tax rule is often better implemented as ordinary code than probabilistic prediction. Choose traditional ML when structured prediction or explainability fits; choose a foundation model when its general language, visual or other learned capabilities help.
- Speech to text: Transcribe. Text to speech: Polly. Language translation: Translate. Text entities/sentiment: Comprehend. Conversational interfaces: Lex. Document extraction: Textract. Image/video analysis: Rekognition. Custom model lifecycle: SageMaker AI. Managed foundation-model applications: Bedrock.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Predict a house price | Regression. |
| Group similar customers without labels | Clustering. |
| An exact legal formula must always give the same answer | Deterministic application logic, with validation. |
Traps
- Classification confidence is not proof of truth.
- A large language model is not automatically better for a small tabular prediction.
- Selecting a managed service does not remove data-quality responsibility.
02 · The ML Lifecycle and MLOps
Memory hook: Split before learning transformations; evaluate before deployment; monitor after deployment.
Must remember
- Start with a measurable business problem, an acceptable error cost and a baseline. Collect authorised, representative data; clean missing/invalid values; transform features; split training, validation and test sets; train; evaluate; deploy; monitor and retrain when justified.
- Fit scalers, imputers and feature selection on training data, then apply the learned transformation to validation/test data. Duplicate entities or future events across splits can create data leakage and unrealistic scores.
- Overfitting means learning training-specific patterns that generalise poorly; regularisation, representative data and simpler models can help. Underfitting means the model cannot capture useful patterns; improve features or capacity. Hyperparameters control training choices; model parameters are learned.
- Precision = TP/(TP+FP): of predicted positives, how many are correct? Recall = TP/(TP+FN): of actual positives, how many were found? F1 combines precision and recall. Accuracy can hide failure on a rare class. For regression, error metrics must match the business penalty.
- Batch inference handles offline collections; real-time endpoints serve interactive calls; asynchronous inference accepts queued work with later results; serverless inference reduces endpoint-management effort for supported patterns. Match latency, payload, traffic and cost.
- MLOps versions data, code and model artifacts; tracks experiments; automates repeatable pipelines; approves releases; monitors drift and operating health. SageMaker AI supports these lifecycle stages. A technically better score is insufficient if cost per user rises beyond the value of the improvement.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Missing a disease is especially costly | Prioritise recall while measuring the resulting false positives. |
| Many alerts are wrong | Investigate precision and the decision threshold. |
| Millions of records need results tomorrow | Evaluate batch inference. |
Traps
- The test set is not a tuning dataset.
- Data drift does not prove accuracy declined; collect outcome evidence.
- A model needs operational and business metrics as well as statistical scores.
03 · Foundation Models and Generative AI
Memory hook: A model predicts plausible output; your application must establish whether it is useful and supported.
Must remember
- A foundation model is pretrained on broad data and can support multiple downstream tasks. Transformers underpin many language models; diffusion models are common for image generation. Multimodal models process or generate more than one modality.
- A token is a model-specific text or data unit, not necessarily a word. A context window limits what a request can include; input and output consume capacity and may have different prices. Longer prompts and answers can increase cost and latency.
- Embeddings represent content as vectors for similarity operations. Chunking divides source content into retrievable pieces. Embeddings do not encrypt text and do not themselves generate a final answer.
- Bedrock offers managed access to supported models and application capabilities. SageMaker AI/JumpStart support a more custom model-development and hosting path. Model choice depends on modality, languages, quality, context/output length, region, licensing, latency, privacy and cost.
- Context engineering selects and organises instructions, retrieved evidence, tool results, conversation state and memory within the available context. Dumping an entire document collection into every request is costly and can obscure relevant facts.
- Generation can hallucinate, vary between calls and be hard to explain. Lower temperature generally reduces sampling variability; it does not guarantee correct or identical answers. Evaluate with real tasks rather than choosing the largest model by default.
- Compare token-based on-demand use, provisioned throughput where supported, caching, batch options and custom-model infrastructure. Business value includes task completion, time saved, user satisfaction and cost per successful interaction, not just tokens per second.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Quickly call several supported foundation models | Bedrock managed model APIs. |
| Need custom training and hosting control | SageMaker AI, with additional operating responsibilities. |
| Relevant facts are missing from the prompt | Improve context/retrieval before buying a larger model. |
Traps
- An embedding is not a lossless copy or a privacy boundary.
- A large context window does not guarantee reliable use of every supplied fact.
- Fluent wording is not evidence of factual correctness.
04 · Prompt Engineering, RAG and Fine-Tuning
Memory hook: Prompt for instructions, retrieve for current evidence, tune for learned behaviour.
Must remember
- A good prompt defines the task, relevant context, output format and constraints. Zero-shot uses no demonstration; one-shot/few-shot include examples. Templates standardise inputs. Version prompts and evaluation sets so improvements can be compared and rolled back; Bedrock Prompt Management supports managed prompt versions.
- Structured reasoning prompts can help decompose a task, but an explanation is not proof of a model's internal process or correctness. Prefer checkable intermediate results and concise justifications. Negative instructions alone are weak enforcement.
- RAG retrieves relevant content, adds it to a prompt and generates a grounded response. A typical path is ingest → clean → chunk → embed → index → retrieve → generate → cite. Bedrock Knowledge Bases manages supported parts of this workflow.
- Vector storage can use supported OpenSearch, Aurora/PostgreSQL or other integrations; the exact service feature matters. Hybrid keyword/vector retrieval, metadata filtering and reranking can improve results. Enforce user access before evidence enters the prompt.
- Fine-tuning changes model weights using task/domain examples; it can improve style and behaviour but is not a live fact lookup. Continued pretraining adapts using additional domain text. Instruction tuning uses instruction/response examples. RLHF uses human preference feedback. Distillation trains a smaller model from a larger model's outputs or behaviour.
- Curate representative, licensed, deduplicated data and hold out evaluation examples. Full pretraining has far higher data/compute requirements than prompt changes or retrieval. Prompt caching can reuse eligible repeated context; it does not repair stale source facts.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Answer questions about policies updated daily | RAG over authorised current documents. |
| Consistent specialised output style across many tasks | Evaluate prompting, then fine-tuning if needed. |
| Reduce a successful model's serving cost | Evaluate a smaller model or distillation against quality targets. |
Traps
- RAG does not alter model weights.
- Fine-tuning does not automatically grant access to new documents.
- A retrieved document can contain hostile instructions; treat it as untrusted data.
05 · Agents, Tools and AWS AI Platforms
Memory hook: The model proposes; tools act; policy decides whether an action is allowed.
Must remember
- An agent uses a model, instructions, state and tools to work toward a goal. A fixed workflow follows predefined steps; an agent can choose among actions dynamically. Use the simpler workflow when its decision structure is sufficient.
- A tool has an interface/schema, permissions and observable results. Validate arguments and outputs. Require human approval for consequential actions where appropriate. Retries need idempotency so a timeout does not create duplicate purchases or updates.
- MCP standardises how supported clients connect to tools and resources; it does not grant trust or make every exposed tool safe. Multi-agent designs may use a supervisor, delegation or peer interaction. More agents introduce coordination, latency and failure modes, not guaranteed quality.
- Short-term state tracks the current task; longer-term memory persists selected facts. Apply data minimisation, isolation, retention and deletion policies. Never mix one customer's retrieved data or memory with another's context.
- Bedrock Agents provides managed agent capabilities; AgentCore provides services for operating agents, including runtime and identity-related capabilities. Strands Agents is an agent-development framework. Choose components based on control and operating requirements, not similar names.
- Current AWS objectives also mention Amazon Quick for business AI experiences and Kiro for AI-assisted development. These are not substitutes for the underlying identity, model evaluation and application security controls. Product branding evolves; use the exam's current terminology.
- Evaluate whole-task completion, valid tool choice, argument accuracy, loop rate, latency and total cost. Add execution limits, safe failure paths and traces so an agent that repeats a tool indefinitely can be diagnosed and stopped.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Known approval steps and deterministic branches | Workflow orchestration. |
| A model must choose among approved business tools | An agent with scoped permissions and validation. |
| Several specialised agents share work | Explicit orchestration, state boundaries and end-to-end evaluation. |
Traps
- Tool discovery is not tool authorisation.
- Memory can preserve incorrect or sensitive information.
- Successful text generation is not the same as successful completion of an external action.
06 · Evaluation and Responsible AI
Memory hook: Measure the answer, the experience and the harm; averages can hide who fails.
Must remember
- Use representative held-out tasks, stable baselines and subgroup analysis. Compare quality, robustness, safety, latency, cost and business outcomes. A benchmark unrelated to the real workflow is weak evidence of suitability.
- BLEU emphasises n-gram precision against references and is associated with translation. ROUGE uses overlap/recall-oriented measures often applied to summaries. BERTScore uses contextual representations for semantic similarity. None alone proves factual correctness or useful business outcomes.
- LLM-as-a-judge can scale evaluation but introduces judge bias, model/version sensitivity and prompt dependence. Calibrate against human judgement and inspect disagreements. Bedrock evaluation capabilities support supported automated and human evaluation approaches.
- Evaluate RAG retrieval separately from answer groundedness and relevance. Evaluate an agent's tool use, action validity and completed task, not just its final wording. Include adversarial inputs and safe refusal/escalation tests.
- Responsible AI includes fairness, robustness, safety, privacy, transparency, explainability, accountability and veracity. Representative data, label review, human audits and subgroup metrics help find harmful differences hidden by aggregate accuracy.
- A transparent system exposes relevant workings and limitations; an explanation helps a person understand a particular result or behaviour. Model cards document intended use, evidence, limitations and risk. An open-source model is not automatically interpretable, safe or appropriately licensed.
- Consider intellectual-property rights, deceptive or biased outputs, environmental impact and user trust. Human-centred design needs clear AI disclosure, feedback and contestability where appropriate. Guardrails reduce specified risks but cannot certify that every response is harmless or true.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Summary wording differs but meaning is similar | Use semantic and human evaluation alongside overlap metrics. |
| High overall accuracy but poor outcomes for one group | Subgroup analysis and fairness investigation. |
| High-consequence decision | Human oversight, documented limits and a suitable error/appeal process. |
Traps
- Fairness has multiple definitions and trade-offs.
- A hallucination score is evidence, not a guarantee.
- Removing all sensitive columns does not necessarily remove proxy bias.
07 · AI Security, Privacy and Governance
Memory hook: Protect the data path and the action path, then keep evidence of both.
Must remember
- Apply least-privilege IAM roles to models, tools, data stores and logs. Separate end-user identity from workload identity. Agent identity and policy features support controls, but the application still needs tenant isolation and scoped tool permissions.
- Use TLS in transit, suitable encryption at rest and controlled KMS key access. PrivateLink provides supported private connectivity; it does not replace identity authorisation. Macie helps discover sensitive S3 data. Secrets should not be included in prompts or source code.
- Prompt injection tries to turn untrusted text into instructions. It may arrive in a user message, retrieved page or tool result. Poisoning corrupts training or indexed data. Jailbreaking seeks to defeat safety behaviour. Validate inputs/outputs, restrict actions and treat retrieved content as data.
- Bedrock Guardrails can apply configured content, topic, sensitive-information and other supported controls. Grounding and output validation can help detect unsupported statements. Do not use model self-reported confidence as the sole authority for high-risk decisions.
- Track provenance, licences and lineage from source data through transformations, model versions and outputs. Define residency, retention and deletion requirements for prompts, logs, embeddings and memory as well as primary datasets.
- CloudTrail supplies supported API audit events; Config evaluates resource configuration; Inspector assesses supported workload vulnerabilities; Artifact supplies AWS compliance evidence; Trusted Advisor highlights supported recommendations. Choose evidence according to the question.
- Governance needs accountable owners, approval gates, review cadence, staff training and documented exceptions. Use a risk framework appropriate to the application and shared-responsibility model. Service compliance does not certify that your own data collection or use is lawful.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| A retrieved document tells an agent to reveal secrets | Treat as injection; enforce permissions and tool constraints outside the model. |
| Need evidence of configuration compliance | Config plus the relevant audit process. |
| Need to prove where training material came from | Data lineage, provenance and licensing records. |
Traps
- A private network does not prevent an authorised application from leaking data.
- Encrypting vectors does not resolve the right to retain their source data.
- Logging every prompt without redaction can create a second sensitive-data store.