Memory hook: Version the prompt like application code.
Must remember
- Foundry resources/projects organize supported AI development and deployment. Configure managed identities, RBAC, private networking and connected resources rather than relying on a developer’s broad login.
- Choose supported serverless/model API or managed-compute deployment by model, region, data requirements, control and cost. Availability, quotas and supported deployment types vary across models.
- Provisioned throughput can suit predictable sustained demand; consumption pricing suits other patterns. Measure token rates, concurrency, latency and utilization before choosing capacity.
- Pin or deliberately manage model versions and deployment names. A provider model update can change output behavior even when application code stays the same.
- Keep prompts, templates, tool definitions and retrieval settings in Git. Compare variants on the same evaluation set; capture the full configuration required to reproduce a response.
- Use Bicep/CLI and CI workflows for repeatable infrastructure and staged releases. Keep secrets in supported secure stores, avoid logging private prompt content indiscriminately and retain rollback configurations.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| High-volume steady token demand | Evaluate provisioned throughput with utilization and capacity evidence. |
| Prompt change improves one demo but harms other users | Run a representative regression evaluation before release. |
Traps
- A model deployment name does not guarantee immutable behavior forever.
- Private networking does not replace model/project authorization.
Active recall
1. Why version tool definitions with prompts?
Tool schemas and available actions influence behavior and reproducibility.
2. What should a prompt experiment hold constant?
Evaluation inputs, scoring criteria and other relevant configuration.
3. Why measure tokens rather than only requests?
Request sizes and generated outputs vary substantially in cost and capacity use.
4. What is a safe release unit?
A tested combination of model, prompt, retrieval, tools, configuration and application code.
5. Why use managed identity?
To obtain scoped service access without embedding credentials in code.