Memory hook: Target the partition; measure the RU; protect the retrieval.
Reviewed 10 October 2026. Read this once, then answer the last-pass checks without looking.
Scope/version: This guide follows the refreshed Cosmos DB AI Developer scope, including vector/full-text retrieval, agent memory and Fabric integration. Older DP-420 outlines omit some of these areas.
Must remember by domain
| Domain | Rapid revision |
|---|---|
| Model/throughput | Account → database → container → item. Logical partition is determined by key value; physical partitions manage underlying distribution. Pick high-cardinality, balanced keys aligned with queries. Hierarchical keys support suitable tenant hierarchies; synthetic sharding spreads writes but can complicate reads. Manual/autoscale provisioned throughput and serverless suit different demand. |
| Consistency | Strong prioritizes latest-read consistency; bounded staleness caps lag; session preserves session guarantees using tokens; consistent prefix prevents out-of-order observations; eventual converges. Stronger read guarantees can increase work/latency. Consistency is independent of the supported transaction boundary. |
| SDK operations | Reuse a client, set preferred regions, target id+partition key for point reads and use continuation tokens for paging. ETag/If-Match rejects stale updates. Transactional batch is atomic within one logical partition; bulk optimizes throughput without global atomicity. TTL controls expiry, not exact job scheduling. |
| Change processing | Processor distributes work with a lease container; pull mode gives direct control. Latest-version feed differs from all-versions-and-deletes history and prerequisites. Design at-least-once handling with idempotent side effects and post-success checkpoints. Copy-container jobs move supported data; AI-generated code still needs key/security review. |
| Security/recovery | Entra data-plane RBAC differs from account management RBAC. Private endpoint/firewall permission is not data permission. Configure backups, restore rights, regional read/write behavior and conflict/failover policy. Periodic backup and continuous PITR differ; a replica cannot restore arbitrary past state. |
| Optimization/fleets | Observe request charge, 429/retry-after, status/substatus, diagnostics and per-partition utilization. Point reads, partition filters, smaller projections and appropriate indexes reduce work. Composite indexes support selected query shapes. A hot partition may throttle despite account headroom. Fleet governance does not remove per-account/partition limits. |
| AI/analytics | Full text = lexical matching; vectors = semantic similarity; hybrid combines them. Match dimension/type/metric/index to the embedding model. DiskANN/sharding trade recall, latency and scale. Agent state needs tenant/user boundaries, TTL and concurrency. Fabric mirroring and connectors support distinct analytics paths with their own freshness/security. |
Diagnosis and traps
Locate the slow request → inspect charge/diagnostics → isolate client/network versus service time → inspect partition key/index/query → bound retries/concurrency → change capacity if evidence supports it. A high-cardinality random key can make required queries expensive. Vector similarity is not authorization; access filtering must precede evidence entering the model.
Last-pass self-check
1. Need atomic updates across arbitrary partition keys: transactional batch?
No. Its atomic boundary is one logical partition.
2. What should a stale ETag update do?
Fail its conditional check so the client can reload/reconcile instead of overwriting silently.
3. All item deletes needed: default latest-version feed enough?
No. Choose a supported mode/retention design that captures required deletes.
4. Why does RU scaling sometimes fail to solve throttling?
A hot partition, costly query or excessive retry concurrency can remain the bottleneck.
5. Why record embedding model/version?
Dimension or semantic-space changes can require regenerating vectors and indexes.
Sources
- Official exam scope and version
- Product documentation
- Product documentation
- Product documentation
- Product documentation
- Product documentation
Every topic at a glance
Open any topic to revisit its essential facts, decisions and exam traps. Use the full topic for active recall and supporting references.
01 · Resource model, consistency and partitioning
Memory hook: Good keys spread writes and target reads.
Must remember
- An account contains databases and containers; items live within logical partitions determined by their partition-key values. Choose keys from access patterns, cardinality and write distribution, not only a field that looks unique.
- Hierarchical partition keys use multiple key levels to support suitable tenant and scale patterns. Synthetic keys combine or distribute values deliberately; random sharding can reduce hotspots but complicate reads.
- Embed related data when it is read/updated together and bounded; reference when independent growth or update patterns make duplication expensive. Model for the actual queries rather than copying a relational schema unchanged.
- The five consistency choices are strong, bounded staleness, session, consistent prefix and eventual. Session consistency preserves supported session guarantees through tokens; consistency is not the same as transactional scope.
- Request Units normalize operation cost. Provisioned manual/autoscale and serverless models suit different traffic patterns; evaluate account, database and container throughput settings and service constraints.
- Use a long-lived SDK client where recommended, appropriate connection mode and region preferences. Capacity planning includes hot partitions, item size, indexing, query shape and concurrency.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Most requests target one tenant’s records | A tenant-aware key design, with growth/hot-tenant analysis and HPKs where suitable. |
| Occasional unpredictable traffic | Evaluate serverless against provisioned/autoscale pricing and supported limits. |
Traps
- High total RU capacity cannot always rescue one hot logical partition.
- A unique random partition key can make common multi-item queries expensive.
02 · SDK operations, change feed and concurrency
Memory hook: Point read cheaply; retry safely; process changes twice safely.
Must remember
- A point read uses item id and partition key and is usually cheaper than an equivalent query. Parameterize queries and use paging/continuation tokens for bounded result processing.
- Create, replace, patch and delete have different update semantics. ETags with conditional requests implement optimistic concurrency so a stale writer does not silently overwrite a newer version.
- Transactional batches operate within a supported logical partition boundary. Bulk execution improves throughput across operations but is not one global atomic transaction.
- TTL expires eligible items according to container/item configuration; expiration is not a precise scheduling mechanism. Account for expired data in retention, change processing and recovery design.
- Change feed supports reactive processing. The processor coordinates work through a lease container; pull mode gives the application more direct control. Latest-version versus all-versions-and-deletes modes have different prerequisites and event semantics.
- Make consumers idempotent and checkpoint only after successful work. Copy-container jobs support suitable data movement; Agent Kit/AI-assisted scaffolding can accelerate development but generated partitioning, security and retry choices still require review.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Fetch one known item | Use id plus partition key for a point read. |
| Prevent lost updates | Conditional writes using the current ETag and conflict handling. |
Traps
- Bulk writes do not make operations across arbitrary partitions atomic.
- Change-feed processing must not assume a side effect will only ever be attempted once.
03 · Secure, distribute and recover data
Memory hook: Identity plus network; replicas plus history.
Must remember
- Prefer supported Entra identity and data-plane RBAC for applications, with narrowly scoped permissions. Control-plane roles and data access roles solve different problems; avoid distributing broad account keys.
- Restrict network access using supported firewall/private endpoint controls and correct private DNS. Authentication still applies on a private path; test public-network settings independently.
- Use encryption, auditing and supported data masking according to the API/feature. Dynamic masking changes what selected users see; it does not replace authorization or encryption against privileged access.
- Multi-region reads improve locality; multi-region writes require conflict and consistency design. Configure preferred regions, failover priorities and supported partition/failover options, then test client behavior.
- Periodic and continuous backup/PITR options have different capabilities and retention. Replication copies current state; recovery history is needed for accidental corruption or deletion.
- Test restored data, identities, endpoints and application reconnection. Keep recovery permissions and retained backups usable while protecting production deletion paths.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| An application needs only one container | Scoped data-plane permissions through a supported identity. |
| Recover a valid state before an accidental update | A supported backup/PITR restore workflow. |
Traps
- A private endpoint does not grant database access.
- Multi-region writes do not remove the need to understand conflicts and consistency.
04 · RU efficiency, diagnostics and fleets
Memory hook: Measure per request and per partition.
Must remember
- Inspect request charge, latency, status/substatus, query metrics and SDK diagnostics. Separate client/network delay from service processing and throttling.
- Target partition keys, use point reads, reduce unnecessary fields and avoid unbounded scans. Index policies trade read efficiency against write/storage cost; composite indexes support particular query shapes.
- Analyze index utilization and query plans before raising throughput. Large items, fan-out queries, excessive indexing and hot tenants can each drive RU consumption differently.
- Azure Monitor metrics/alerts and diagnostic logs in Log Analytics connect service behavior to application SLOs. Monitor end-to-end latency and errors rather than only provisioned RU values.
- Fleet capabilities manage supported groups of accounts, throughput policies and analytics across subscriptions. Fleet-level governance does not eliminate per-account/partition limits or the need to check feature support.
- Load-test representative peak traffic, retries and regional behavior. Reuse SDK clients, respect retry-after guidance and cap concurrency to avoid amplifying a throttling incident.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| One tenant is throttled while the account has spare capacity | Inspect partition distribution and tenant-specific load. |
| A query is expensive despite few returned rows | Inspect scanned data, partition fan-out and index support. |
Traps
- Few returned records do not imply a cheap query.
- More aggressive retries can worsen throttling and latency.
05 · Vector retrieval, agent memory and Fabric
Memory hook: Retrieve with permissions; remember with boundaries.
Must remember
- Full-text search matches lexical terms; vector search matches embedding similarity; hybrid search combines signals. Choose dimensions, data type, distance metric and index from the embedding model and workload.
- DiskANN and supported sharded/vector-partitioned designs trade search performance, recall and scale. Evaluate filtered searches and index behavior using realistic tenant distributions.
- RAG retrieves authorized, fresh evidence for generation. Chunking, embedding version, relevance and citations matter; a similar vector is not proof of a correct answer.
- Store conversation state separately from durable semantic memory where appropriate. Use tenant/user keys, concurrency control, TTL/retention and deletion rules; do not let one user’s conversation become another user’s retrieval context.
- Cosmos DB in Fabric and Azure Cosmos DB serve different integration/operational choices. Mirroring exposes supported operational data for Fabric analytics; verify lag, schema handling and security.
- Fabric T-SQL/Spark can analyze mirrored JSON data, while the Cosmos DB Spark connector supports appropriate reads/writes. Analytics access must preserve data governance and avoid overwhelming the operational workload.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Need exact product codes and semantic descriptions | Evaluate hybrid lexical/vector retrieval with filters. |
| Analyze operations without repeated application queries | Supported Fabric mirroring with governance and freshness checks. |
Traps
- An embedding model change may require re-embedding and index compatibility review.
- Vector-only memory can lose exact identifiers, chronology and authorization context.