Memory hook: Retrieve with permissions; remember with boundaries.
Must remember
- Full-text search matches lexical terms; vector search matches embedding similarity; hybrid search combines signals. Choose dimensions, data type, distance metric and index from the embedding model and workload.
- DiskANN and supported sharded/vector-partitioned designs trade search performance, recall and scale. Evaluate filtered searches and index behavior using realistic tenant distributions.
- RAG retrieves authorized, fresh evidence for generation. Chunking, embedding version, relevance and citations matter; a similar vector is not proof of a correct answer.
- Store conversation state separately from durable semantic memory where appropriate. Use tenant/user keys, concurrency control, TTL/retention and deletion rules; do not let one user’s conversation become another user’s retrieval context.
- Cosmos DB in Fabric and Azure Cosmos DB serve different integration/operational choices. Mirroring exposes supported operational data for Fabric analytics; verify lag, schema handling and security.
- Fabric T-SQL/Spark can analyze mirrored JSON data, while the Cosmos DB Spark connector supports appropriate reads/writes. Analytics access must preserve data governance and avoid overwhelming the operational workload.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Need exact product codes and semantic descriptions | Evaluate hybrid lexical/vector retrieval with filters. |
| Analyze operations without repeated application queries | Supported Fabric mirroring with governance and freshness checks. |
Traps
- An embedding model change may require re-embedding and index compatibility review.
- Vector-only memory can lose exact identifiers, chronology and authorization context.
Active recall
1. What does hybrid retrieval combine?
Lexical matching and semantic vector similarity, with an appropriate ranking strategy.
2. Why partition agent memory by tenant?
To enforce isolation and control query scope/cost.
3. What does RAG change compared with fine-tuning?
It supplies external evidence at request time rather than updating model weights.
4. Why monitor mirroring lag?
Analytical results may not represent the latest operational state.
5. How validate retrieval quality?
Use representative questions, relevance judgments, permission tests and grounded-answer evaluation.