certslothcertsloth
AI-103/Topic 03

Azure / Associate

Text, Speech, Vision and Content Understanding

2 min read5 recall promptsReviewed 2026-10-10

Memory hook: Interpret existing evidence, generate new content and extract structured fields as different tasks.

Must remember

  • Text analysis includes entity/key-phrase extraction, sentiment, summarisation and sensitive-content detection. Use supported Foundry Tools or model prompting according to accuracy, format and control requirements. Validate returned JSON/types before using results downstream.
  • Azure Speech supports speech recognition and synthesis; translation and multimodal audio models address related but different tasks. Match language, audio format, streaming/batch mode and latency. A spoken prompt can be transcribed before a text model or sent to a compatible multimodal model.
  • Vision-capable models interpret image inputs for captions, questions or visual evidence. Image-generation models create new visual outputs from prompts/references. Supported models may offer editing/masks; accepting an image does not mean a model can generate one.
  • Content Understanding uses configured analyzers to extract information from supported documents, images, audio and video. Define desired fields/schema, submit source content and consume the resulting structured output/evidence. Layout, timestamps and provenance can matter as much as the extracted value.
  • Long-running analysis may return an operation identifier; poll/await completion according to the SDK rather than treating acceptance as a final result. Validate confidence/evidence and route uncertain high-impact fields for review. An extracted invoice total still needs a business validation rule.
  • Protect input data, apply content/safety controls and preserve consent/licensing. Embedded text in images or documents can contain prompt injection. Accessibility captions should describe useful visible information without inventing details. Measure performance on the actual languages, document layouts and recording conditions.

Choose under exam pressure

Requirement Choice and reason
Extract invoice fields into a schema A suitable Content Understanding analyzer.
Generate spoken output Speech synthesis.
Answer a question about an existing image A vision-capable multimodal model with grounded evaluation.

Traps

  • Recognition and synthesis run in opposite directions.
  • A job accepted response is not completed analysis.
  • OCR text can contain hostile instructions.

Active recall

1. What is the difference between vision analysis and image generation?

Interpreting an existing image versus creating a new one.

2. Why validate extracted JSON?

Models/services can return missing, unexpected or invalid values relative to the application contract.

3. Why retain timestamps from video/audio extraction?

To locate and verify the source evidence for a result.

4. What should low-confidence critical fields trigger?

A defined validation or human-review path.

5. Does support for one language/layout imply equal quality on all inputs?

No. Evaluate representative content and operating conditions.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.