Memory hook: Authorise evidence before retrieval, and authorise actions before tool execution.
Must remember
- Build ingestion with source permissions, stable IDs, text/layout extraction, chunking, embeddings and index updates/deletes. Azure AI Search supports lexical, vector and hybrid retrieval with supported semantic ranking/enrichment capabilities. Vector similarity, ranking and grounded answer quality are different measurements.
- Enforce tenant/document permissions using trusted identity context before evidence enters the model. Preserve source references for citations and freshness. Tune chunk size, overlap, filters and reranking on held-out questions; more retrieved text is not always better.
- Agents need explicit roles/goals, conversation state, memory, typed tool schemas and limits. Integrate APIs, search, knowledge stores, custom functions and content-analysis tools through supported interfaces. A tool result is untrusted input and may itself contain injection.
- Multi-agent systems require ownership of tasks/shared state, clear handoffs, failure propagation and end-to-end evaluation. Supervisor/delegation patterns add cost and latency. Prefer a deterministic workflow where known branches are sufficient; use rules for exact constraints rather than relying on model obedience.
- Record project/deployment configuration, prompts, index versions and tool contracts in CI/CD. Test with managed identities/private networking and the actual application roles. Treat model, prompt and retrieval changes as releases with canary/rollback and compatibility checks.
- Monitor tokens, quotas, time to first useful output, completion latency, tool failures and cost per successful task. Bound retry/reflection loops, validate structured output, cache with identity/freshness context and evaluate smaller-model routing where quality permits.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Current internal knowledge with citations | RAG with ACL-aware retrieval and support checks. |
| Several tools need controlled sequencing | A workflow or bounded agent orchestration. |
| A model requests an irreversible action | Validate policy and required approval before execution. |
Traps
- Reflection can repeat an error rather than correct it.
- A cache without tenant context can leak data.
- Changing an embedding model can require index migration.
Active recall
1. Why separate retrieval and generation evaluation?
A missing document and a misused correct document have different causes.
2. What should a tool schema define?
Allowed operation, typed arguments, required fields and output/error contract.
3. Why track prompt/index versions?
They affect behaviour even when application code is unchanged.
4. When should a deterministic rule replace model judgement?
When an exact enforceable constraint or calculation is required.
5. What should multi-agent tests measure?
The final authorised outcome, coordination errors, latency and total cost.