Memory hook: Choose keys and indexes for the query, then make cache freshness explicit.
Must remember
- Cosmos DB for NoSQL SDK operations need endpoint/authentication, database/container and partition-key context. Point reads by ID and partition key differ from cross-partition queries. Request Units reflect work; examine query metrics, indexing policy and chosen consistency before adding throughput.
- Store embeddings with compatible dimensions and supported vector indexing/query configuration. Filter by trusted tenant/metadata constraints. Vector similarity is not access control or proof of answer correctness. A model/dimension change can require index/data migration.
- A change feed processor tracks supported changes through leases/checkpoints and distributed workers. Design idempotent handlers and verify the feed mode's treatment of updates/deletes; do not assume every mode captures every historical operation.
- PostgreSQL needs sensible tables/types, keys and indexes plus connection pooling. pgvector supports vector similarity operations and index choices with recall/performance trade-offs. Metadata predicates, candidate count, index parameters, memory and compute affect latency and quality.
- Bound connection counts; reuse pools safely and set timeouts. A larger database can still be bottlenecked by poor queries or too many short-lived connections. Inspect query plans and resource metrics before scaling.
- Azure Managed Redis supplies cache operations and supported vector search capabilities. Set TTL, eviction and invalidation deliberately. Cache-aside loads on a miss; write-through updates a cache during writes. Protect against stampedes and include identity/context in sensitive result cache keys.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Known Cosmos item ID and partition key | Prefer a point read where it meets the requirement. |
| Embedding search is slow | Inspect vector index, filters, candidate settings and resource limits. |
| Frequently requested data changes occasionally | Cache with a defined invalidation/TTL strategy. |
Traps
- A cache is not automatically the source of truth.
- Cosmos change-feed modes have different semantics.
- Vector dimensions must match the chosen embedding/index configuration.
Active recall
1. Why can a Cosmos query consume many RUs?
Cross-partition access, inefficient predicates/indexing, consistency and result/work volume can increase cost.
2. What does a change-feed checkpoint provide?
Progress tracking for recovery/rebalancing, not universal exactly-once side effects.
3. Why use a PostgreSQL connection pool?
To reuse bounded connections and reduce connection overhead/saturation.
4. What is a cache stampede?
Many concurrent misses trigger duplicate expensive backend work.
5. Why include tenant context in a cache key/control?
To prevent one user's authorised result being served to another.