Building an AI solution goes beyond calling a model. The focus is on creating production-ready AI applications and agents with Microsoft Foundry that can use enterprise data, interact with tools, process different content types, and collaborate to complete real tasks.
Core capabilities to know
| Area | What you should understand |
|---|---|
| Generative AI apps | Build conversational applications using models, APIs, and SDKs |
| Grounding | Connect models to your own data for relevant, fact-based responses |
| Agents + tools | Let agents retrieve information and take actions |
| Multi-agent systems | Orchestrate specialized agents to collaborate on workflows |
| Multimodal AI | Process text, documents, vision, and speech |
| Production | Deploy, publish, monitor, secure, and apply responsible AI safeguards |
Exam focus: Understand not just what these capabilities do, but when and why you would use them together in an Azure AI solution.
Think in solution flows
A useful mental model for AI-103 is:
User → AI App/Agent → Model → Data + Tools → Action/Response
For more complex solutions:
User → Orchestrator → Agent A + Agent B + Agent C → Tools/Data → Result
An agent therefore isn't simply a chatbot. It combines a model's reasoning capabilities with instructions, knowledge, and tools so it can perform useful work.
Preparing effectively
The course assumes working knowledge of Python, REST APIs/SDKs, Azure fundamentals, and generative AI concepts. Hands-on practice is important: build applications in Microsoft Foundry, connect models to data, add tools to agents, experiment with multimodal inputs, and create multi-agent workflows.
Key takeaway: Think beyond prompts and models. AI-103 is about assembling the components required for an end-to-end AI solution:
Models → Grounding → Tools → Agents → Orchestration → Production
Related Snipps
- What Microsoft Foundry Provides for Building AI Agents — Core Foundry capabilities for tools, conversations, models, security, storage, and monitoring.
- Microsoft Foundry SDK vs. Foundry Tools SDKs — Understand platform-level and specialized SDKs when developing Foundry solutions.
- Building Production-Ready AI Agents with Microsoft Agent Framework Course — Explores agent architecture, tools, workflows, memory, monitoring, and deployment.
Before building an AI application or agent, understand how Microsoft Foundry organizes models, tools, knowledge, and development resources. These relationships form the foundation for everything that follows.
Foundry architecture at a glance
Think of the structure as:
Foundry Resource → Project → Models + Agents + Tools + Knowledge
The Foundry resource is the underlying Azure resource and infrastructure boundary. A project lives within that resource and organizes the models, agents, tools, and knowledge used by an AI solution. The resource is the foundation; the project is the development workspace.
Microsoft Foundry provides access to generative models alongside Foundry Tools for specialized AI capabilities such as Language, Speech, Translation, and Document Intelligence. These services complement models when an application needs capabilities such as speech recognition or structured information extraction.
Know which SDK fits
| Need | Typical choice |
|---|---|
| Direct model/chat interaction | OpenAI SDK |
| Agents, tools and grounding | Microsoft Foundry SDK |
| Specialized AI capability | Service-specific SDK |
| Universal HTTP integration | REST API |
Key distinction: use the OpenAI SDK when targeting a model directly; move toward the Foundry SDK when working with the broader agentic platform, including tools and grounding.
For development, Visual Studio Code with the Microsoft AI Toolkit is the recommended combination presented in the course. In the Foundry portal, Discover is primarily for finding models, tools and templates, while Build is where deployed resources are configured and tested.
Remember the six Responsible AI principles
Fairness • Reliability & Safety • Privacy & Security • Inclusiveness • Transparency • Accountability
These are not an afterthought: they influence grounding, prompts, guardrails, UX, evaluation, and ongoing operation of the solution.
Memory model: Resource hosts → Project organizes → Model reasons → Knowledge grounds → Tools act → Responsible AI governs.
Related Snipps
- What Microsoft Foundry Provides for Building AI Agents — Explains models, tools, knowledge, security, monitoring, and managed agent capabilities.
- Building Production-Ready AI Agents with Microsoft — Covers agents, tools, orchestration, monitoring, and production-oriented architecture.
- Choosing the Right Platform for Building AI Agents — Places Foundry and Azure within the broader range of agent-development approaches.
Choosing a model is more than picking the most capable option. In Microsoft Foundry, the practical lifecycle is Select → Deploy → Evaluate: balance model capability, performance, cost, data-location requirements, and measured output quality.
1. Select the model
Use the Model Catalog and Leaderboard to narrow candidates by capabilities, supported languages, context window, fine-tuning support, and benchmarks.
| Benchmark | What it tells you |
|---|---|
| Quality | Usefulness and overall response quality |
| Safety | Susceptibility to harmful or adversarial inputs |
| Throughput | How quickly the model processes and returns output |
| Cost | Price based on input/output token usage |
Key trade-off: a larger model may deliver better results, while a smaller model can offer higher throughput and lower cost.
2. Choose the deployment
Know these deployment choices:
-
Global → broadest capacity and potentially highest throughput
-
Data Zone → processing stays within a geographic zone such as the EU or US
-
Regional → maximum control over the processing region, but capacity is constrained to it
-
Standard → usage/token-based; suitable for general workloads
-
Provisioned → predictable, guaranteed throughput
-
Batch → high-volume, non-interactive processing where latency is less important
-
Developer → lightweight testing of fine-tuned models
Remember: Global Standard is the general-purpose choice highlighted for obtaining the largest available quota.
3. Evaluate before production
Model performance isn't just speed. Evaluate quality, relevance, fluency, and groundedness.
Manual evaluation: run representative prompts and edge cases against models side-by-side.
Automated evaluation: use a larger prompt dataset, expected behavior, and evaluators to systematically score responses and safety.
High-value distinctions to remember:
Throughput = speed/capacity of responses · Fluency = natural, linguistically correct output · Groundedness = whether responses align with known information
Mental model:Requirements → Catalog/Benchmarks → Deployment → Test Dataset → Evaluators → Compare → Improve
A chat application becomes much easier to design once you understand three decisions: which endpoint to use, which API to call, and where conversation state is maintained.
Endpoint & API choices
| Choice | Remember this |
|---|---|
| Azure OpenAI endpoint | Direct model access; typically use the OpenAI SDK |
| Foundry project endpoint | Higher-level access to models, tools, and agents |
| Chat Completions API | Client resends the conversation history |
| Responses API | Can link turns using a previous response ID |
Key concept: LLMs are inherently stateless. Chat history must therefore be supplied or referenced. With Chat Completions, your application maintains and resends the message array. With Responses, server-side context can be continued by passing the previous response identifier.
Core pattern
response = client.responses.create(
model="my-deployment",
instructions="You are a helpful assistant.",
input=user_input,
previous_response_id=last_response_id
)
print(response.output_text)
last_response_id = response.id
For authentication, prefer Microsoft Entra ID over embedded API keys. DefaultAzureCredential is useful because the same application code can obtain credentials across local development and Azure-hosted environments.
Also know the configuration controls: temperature influences response variability, while max tokens constrains output size.
For applications performing other I/O while waiting for model responses, use the asynchronous client + await to avoid blocking execution.
Remember: responses.create() generates a response; response.output_text retrieves its text; response.id can connect the next turn.
Generative AI models are powerful, but their trained knowledge is limited. Tools extend models beyond text generation, allowing them to access real-time information, take actions, ground responses in facts, extend functionality, and build intelligent workflows.
Know the Tools
| Tool | Purpose |
|---|---|
code_interpreter |
Generate and run code for calculations and data analysis |
web_search |
Find current information on the internet |
file_search |
Search files and ground responses in specific knowledge |
function |
Call custom functions implemented by your application |
Remember: current information → web_search · uploaded/private documents → file_search · calculations/code → code_interpreter · application-specific actions → function
Responses API
Tools are provided through the tools collection. The model can determine which available tool is appropriate for a request.
response = client.responses.create(
model=model_name,
input="Answer the user's request using the available tools.",
tools=[
{"type": "code_interpreter", "container": {"type": "auto"}},
{"type": "web_search"},
{"type": "file_search", "vector_store_ids": [vector_store.id]}
]
)
print(response.output_text)
Core flow: User → Responses API → Model → Tool → Result → Model → Response
For file_search, documents are stored in a vector store and prepared for semantic retrieval:
Files → Chunking → Embeddings → Vector Store → Retrieval → Model
This lets the model answer using relevant document content rather than relying only on its trained knowledge. Uploaded company policies or private documents → File Search + Vector Store.
Function Calling
Functions are different because the application executes the function, not the model. The model identifies the required function and returns a function-call request:
User → Model → Function Call → Application → Function → Result → Model → Response
The application executes the requested code and returns its result. This process can run in a loop when multiple tool calls are needed.
Key distinction: built-in tools extend the model with predefined capabilities; function calling connects the model to your own application logic and actions.
Related Snipps on Snippset
- What Microsoft Foundry Provides for Building AI Agents — Overview of Foundry capabilities for models, agents, tools, data, security, and monitoring.
- Microsoft Foundry SDK vs. Foundry Tools SDKs — Explains the difference between platform-level Foundry development and specialized AI service SDKs.
- Building Production-Ready AI Agents with Microsoft Agent Framework Course — Covers tools, memory, workflows, orchestration, monitoring, and production-ready AI agents.
Getting better results from a generative AI model does not automatically mean fine-tuning it. The key is identifying what is wrong with the output and choosing the least complex technique that solves it.
The decision you should remember
| Problem | Technique | Why |
|---|---|---|
| Instructions, tone or output format | Prompt engineering | Fastest and simplest |
| Missing or external knowledge | RAG | Retrieves relevant information at runtime |
| Consistently wrong behavior/style | Fine-tuning | Changes how the model responds |
| Knowledge + behavior problems | RAG + fine-tuning | Both may be required |
Start with prompt engineering. System instructions can define the model's role, constraints, tone and expected output. Examples and few-shot prompting can further improve consistency.
RAG = give the model knowledge
Retrieval-Augmented Generation is appropriate when the required information is too large, specialized or dynamic to include directly in the prompt:
Question → Vectorize → Retrieve relevant chunks → Question + context → LLM → Answer
The important distinction is that RAG does not retrain the model. Relevant external information is retrieved and added to the model's context at runtime.
Fine-tuning = change behavior
Fine-tuning trains a base model using many examples of desired input/output behavior. A supervised dataset commonly contains conversations such as:
{"messages":[
{"role":"user","content":"Suggest a destination."},
{"role":"assistant","content":"Absolutely! What kind of trip interests you?"}
]}
Microsoft Foundry supports different customization methods depending on the selected model, including supervised fine-tuning and Direct Preference Optimization (DPO). Model and region must support the chosen method.
Key distinction: Fine-tuning is comparatively expensive and time-consuming, produces a new fixed model that must be deployed, and must be repeated when training requirements change.
Remember
Prompt → instructions. RAG → knowledge. Fine-tuning → behavior.
When uncertain, try prompting first, use RAG for missing knowledge, and reserve fine-tuning for behavior that prompting cannot reliably achieve.
Generative AI is non-deterministic: the same system can produce unexpected or harmful outputs. Responsible AI therefore needs to be engineered into the complete application lifecycle—not added as a final check.
The core lifecycle
MAP → MEASURE → MITIGATE → DEPLOY & MONITOR ↻
| Step | Engineering focus |
|---|---|
| Map | Identify harms, attack vectors, risky inputs and misuse scenarios |
| Measure | Evaluate actual model outputs against those risks |
| Mitigate | Add multiple, layered controls |
| Monitor | Observe production behavior and feed new risks back into Map |
Defense in depth
Think of safety as an AI request/response pipeline:
Input → UX Controls → System Prompt + Grounding → Guardrails → Model → Guardrails → Output
UX controls limit the attack surface, for example through input or conversation limits. System prompts and grounding constrain model behavior, but prompts alone are not a security boundary.
Microsoft Foundry guardrails add enforcement around the model and can intervene before input reaches the model and after output is generated. Controls can target:
Jailbreaks · Hate · Violence · Sexual content · Self-harm · Protected material · Groundedness · PII
Guardrails have configurable blocking thresholds, allowing stricter policies for higher-risk applications.
Model refusal ≠ Guardrail blocking
This distinction matters:
Model refusal: Request → Model → "I can't help with that."
The model received and processed the request.
Guardrail blocking: Request → Guardrail → BLOCKED ⛔ → Model
The unsafe request never reaches the model.
What to remember
There is no single safety control. Combine UX restrictions, system instructions, grounding, guardrails, appropriate model selection, evaluations, and production monitoring.
Treat responsible AI like security engineering: identify → test → defend → monitor → repeat.
AI agents add an orchestration layer above an LLM: model + instructions + tools + conversation state work together in an agentic loop to complete multi-step tasks rather than simply return a response.
Core Architecture
Application → Foundry Project → Agent → Model + Instructions + Tools
Microsoft Foundry Agent Service provides a managed runtime for conversation state, tool calling, and agent lifecycle. Agents can invoke built-in tools such as File Search for grounded retrieval and Code Interpreter for Python-based analysis, plus APIs and custom functions.
Build & Integrate
Typical flow:
Create project → Deploy model → Define agent → Configure instructions/tools → Test → Consume from application
The Foundry portal is useful for prototyping; SDK/code-based configuration improves repeatability, version control, and CI/CD.
from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential
project = AIProjectClient(
endpoint=PROJECT_ENDPOINT,
credential=DefaultAzureCredential()
)
openai = project.get_openai_client()
Key Distinctions to Remember
Project endpoint ≠ model endpoint. When working through Foundry, AIProjectClient connects to the Foundry project endpoint. The project client can provide an OpenAI-compatible client for Responses, Conversations, and related operations.
Conversation state can be server-side. Conversations are durable objects containing messages, tool calls, and tool outputs, allowing the same conversation to continue across requests.
Tools determine agent capabilities. File Search provides grounding/RAG; Code Interpreter executes analysis; custom functions/APIs allow agents to take actions.
Managed Agent Service vs. direct model API: the managed runtime handles infrastructure such as conversation management, tool execution/orchestration, and agent lifecycle, while application code operates at a higher abstraction.
Exam mental model: know what belongs to the project, agent, model, conversation, and tool layers—and which SDK/client or endpoint your application is actually targeting.
AI agents become much more powerful when they can act on external systems, not just generate answers. Custom tools let a Foundry agent use application logic, databases, APIs, calculations, and workflows.
The Core Tool-Calling Pattern
Prompt → Agent → function_call → App executes tool → Result → Agent → Answer
A custom function tool has a name, description, and parameters. The agent uses these definitions to determine when a function is needed and what arguments to provide.
# 1. Create the agent
agent = project_client.agents.create_version(...)
# 2. Ask the agent
response = openai_client.responses.create(
conversation=conversation.id,
input="What's the weather in Zurich?",
extra_body={"agent": agent}
)
# 3. Check whether the agent wants to use a function
for item in response.output:
if item.type == "function_call":
# YOUR application executes the function
result = call_function(item.name, item.arguments)
# Return the result to the agent
send_function_result(item.call_id, result)
Key concept: the LLM does not execute your local function. It returns a function_call containing the requested function and arguments. Your application dispatches and executes it, then returns the result so the agent can continue reasoning.
Choose the Right Tool
| Need | Use |
|---|---|
| Local application code | Custom function |
| REST API described with OpenAPI | OpenAPI tool |
| Remote/serverless compute | Azure Functions |
| Low-code workflow | Logic Apps |
Remember
Agent = decides what to call → Application = executes it → Agent = uses the result
One prompt can trigger multiple function calls, allowing an agent to combine several operations before generating its final response.
Model Context Protocol (MCP) standardizes how AI agents discover and invoke external tools. An MCP server publishes its available tools; an MCP client discovers them and makes them usable by an agent.
Remote MCP — direct integration
If the MCP server is remotely accessible, it can be registered directly with the Foundry agent using MCPTool.
# Define the remote MCP server
mcp_tool = MCPTool(
server_label="docs",
server_url="https://.../mcp",
require_approval="always"
)
# Give it to the agent
agent = project_client.agents.create_version(
...,
tools=[mcp_tool]
)
The agent can discover and use tools from that server. If approval is required:
for item in response.output:
if item.type == "mcp_approval_request":
approval = McpApprovalResponse(
approval_request_id=item.id,
approve=True
)
💡 Key point: Remote MCP can be registered directly with the agent. Approval can provide a control point before a requested MCP tool is executed.
Local MCP — your application is the bridge
A cloud-hosted agent cannot directly reach an MCP server running on your machine. Your application therefore connects to the server as the MCP client.
# MCP SERVER — expose tools
@mcp.tool()
def get_inventory(product):
return ...
# MCP CLIENT — discover & call tools
session = ClientSession(...)
tools = await session.list_tools()
result = await session.call_tool(...)
# AGENT — expose tools as functions
agent_tool = FunctionTool(...)
💡 Key point: Local MCP requires your application to perform the MCP communication. The discovered tools are exposed to the agent as FunctionTools.
How to remember it
Ask one question: Can the agent reach the MCP server directly?
Remote MCP: Yes → register it with MCPTool.Agent → MCPTool → Remote MCP Server
Local MCP: No → your application bridges the connection.Agent → FunctionTool → Your App → Local MCP Server
⭐ In short: Remote = agent talks to MCP directly. Local = your application acts as the bridge.
An LLM can identify entities or PII itself—but an agent can instead delegate these tasks to a specialized Azure Language tool through MCP. This separates agent reasoning from deterministic NLP processing.
Architecture: Prompt → Agent → discover/select MCP tool → Azure Language → tool result → final response
What to know
Azure Language MCP Server exposes Azure Language capabilities as tools that an agent can dynamically discover and invoke. Core capabilities include PII detection, language detection, and Named Entity Recognition (NER); additional Language capabilities are also exposed through MCP.
The important distinction is:
-
Agent/LLM: reasons about the request and chooses an appropriate tool.
-
MCP: standardizes tool discovery and invocation.
-
Azure Language: performs the specialized NLP operation.
Tool selection is not hard-coded. The MCP server advertises available tools and their descriptions; the agent matches the user's intent to those descriptions. Good agent instructions further guide when Azure Language should be used.
Minimal mental model
# Agent is already configured with Azure Language MCP
response = openai_client.responses.create(
input="Find and redact PII in this text...",
extra_body={"agent": {"name": "text-agent"}}
)
print(response.output_text)
Behind this simple call:
Agent
└─ discovers MCP tools
└─ selects PII tool
└─ Azure Language analyzes text
└─ result returns to Agent
Watch for approval: MCP tool calls can require user/application approval. Either handle the approval request in code or configure appropriate tools for automatic approval.
Remember: MCP exposes tools; the agent selects them; Azure Language executes the NLP task.
Voice-enabled generative AI is essentially a two-way conversion pipeline: speech → text lets an application understand spoken input, while text → speech turns generated responses back into audio. Azure AI Foundry provides specialized models for both inference tasks.
Know which model solves which problem:
| Task | Model type | Data flow |
|---|---|---|
| Transcription | Speech-to-text | Audio → Text |
| Speech synthesis | Text-to-speech (TTS) | Text → Audio |
A transcription model such as GPT-4o-mini-transcribe accepts audio and returns text. A TTS model such as GPT-4o-mini-tts performs the reverse and can also follow instructions affecting characteristics such as tone.
The implementation pattern is straightforward: deploy the appropriate model in Foundry → create an authenticated Azure OpenAI client → call the corresponding audio API → handle text or binary audio output. Streaming is useful for TTS because audio bytes can be consumed as they arrive rather than waiting for the complete response.
# Speech → Text
with open("speech.wav", "rb") as audio:
text = client.audio.transcriptions.create(
model="gpt-4o-mini-transcribe",
file=audio
)
# Text → Speech
with client.audio.speech.with_streaming_response.create(
model="gpt-4o-mini-tts",
voice="alloy",
input="Hello from Azure AI"
) as audio:
audio.stream_to_file("speech.mp3")
Remember the direction: Transcribe = audio in, text out. TTS = text in, audio out. The audio side is binary data, so applications must correctly read input files or stream/write generated audio.
Azure Content Understanding converts unstructured documents, images, audio, and video into structured, application-ready data. The central concept is the analyzer: a reusable configuration defining how content is processed and what information is returned.
Core architecture
Input → Analyzer → AI processing → Structured output
An analyzer combines:
-
Base analyzer → modality-specific foundation, e.g. document, image, audio, or video.
-
Field schema → defines the information and data types to return.
-
Models/configuration → controls AI-powered processing.
-
Output → content, structured fields, grounding and confidence information.
Key distinction: the schema defines WHAT to extract; the analyzer defines HOW that schema is applied repeatedly.
Prebuilt vs. custom
Use a prebuilt analyzer when the scenario matches an existing type, such as invoice or receipt. Build a custom analyzer when application-specific fields are required.
Invoice
├─ VendorName: string
├─ Total: number
└─ LineItems: array<object>
├─ Description: string
└─ Quantity: number
Fields can use different generation methods:
extract → retrieve information from the sourceclassify → select from predefined categoriesgenerate → derive new information from the content
For example, quantities can be extracted from invoice rows while TotalQuantity can be generated from those values.
Multimodal processing
| Input | Example output |
|---|---|
| Document | Text, tables, fields, totals |
| Image/slide | Text, summary, chart data |
| Audio | Transcript, speakers, actions |
| Video | Transcript, visuals, participants, tasks |
The important pattern stays the same across modalities: define schema → build analyzer → analyze content → consume structured results.
Studio, Foundry & applications
Microsoft Foundry is the broader AI development platform; Content Understanding provides multimodal analysis capabilities within that ecosystem. Content Understanding Studio is the specialized experience for designing, testing, and evaluating analyzers.
Production applications normally use the API/SDK directly:
client = ContentUnderstandingClient(endpoint, DefaultAzureCredential())
# Build reusable analyzer
client.begin_create_analyzer(
analyzer_id="invoice-analyzer",
analyzer_definition=schema
).result()
# Analyze new content
result = client.begin_analyze_binary(
analyzer_id="invoice-analyzer",
binary_input=document
).result()
Content can be supplied as binary data or a downloadable URL, and analyzer operations typically follow an asynchronous begin_* → poll → result pattern.
Remember
Prebuilt analyzer → ready-made scenario
Base analyzer → foundation for customization
Schema → fields, types and extraction behavior
Analyzer → reusable processing configuration
Grounding → where extracted information came from
Confidence (0–1) → application decision signal
Typical decision flow:
Choose modality/analyzer → customize schema if needed → build → analyze → inspect fields → validate using grounding/confidence.
Comments