Memory hook: Authenticate, select the deployment, send a bounded request, inspect the result and handle failure.
Must remember
- A Foundry project organises supported AI application resources/connections. Choose a model and deployment option using capability, Region, quota, throughput, latency and cost. A catalogue model name and your deployment identifier are not always interchangeable.
- In the portal, deploy an available model, test representative prompts and inspect output/usage. A successful playground example does not validate application authentication, networking or error handling. Record the model/deployment and configuration used.
- System instructions define application behaviour; user input provides the request; retrieved text/tool output is untrusted data. Use clear tasks, context, constraints and output schemas. Zero-shot uses no examples; few-shot includes demonstrations. Keep prompts versioned with evaluation cases.
- A lightweight client needs the supported SDK, endpoint/project configuration, a credential and deployment/model selection. Prefer Entra/default-credential patterns for supported keyless access, with the required roles. Create a client, submit messages/input, inspect text/structured results and usage, and handle authentication, rate-limit and timeout errors.
- Streaming returns incremental output; it requires correct accumulation, cancellation and partial-failure handling. Never put keys in source code. SDK shapes evolve, so use the current language quickstart for exact imports and methods rather than memorising an old preview signature.
- A single agent adds a goal/instructions, tools and conversation state. Test it in the portal, then integrate the supported agent client lifecycle: create/reference the agent, submit a user turn, process required tool actions under policy and collect the final result. Bound loops and verify real tool outcomes.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Prototype model suitability | Portal deployment/playground with representative tests. |
| Production app needs credentials without a stored key | Supported Entra/managed-identity authentication. |
| Model requests a tool action | Validate and authorise the action before execution. |
Traps
- Deployment ID and base model name may differ.
- A successful portal call does not prove the app identity has permission.
- Agent text saying “done” is not proof an external action succeeded.
Active recall
1. What configuration must a client know?
The intended endpoint/project, deployment, authentication method and request contract.
2. Why keep a request bounded?
To control context/output limits, latency, cost and failure handling.
3. What should happen on throttling?
Apply bounded retry/backoff and inspect quota/throughput rather than retrying indefinitely.
4. Why version prompts?
To reproduce evaluations and roll back a behaviour regression.
5. What should an agent test verify beyond its prose?
Correct authorised tool use, resulting state and safe handling of failure.