Develop a generative AI chat app with Microsoft Foundry | AI-103 | Episode 3
A chat application becomes much easier to design once you understand three decisions: which endpoint to use, which API to call, and where conversation state is maintained.
Endpoint & API choices
| Choice | Remember this |
|---|---|
| Azure OpenAI endpoint | Direct model access; typically use the OpenAI SDK |
| Foundry project endpoint | Higher-level access to models, tools, and agents |
| Chat Completions API | Client resends the conversation history |
| Responses API | Can link turns using a previous response ID |
Key concept: LLMs are inherently stateless. Chat history must therefore be supplied or referenced. With Chat Completions, your application maintains and resends the message array. With Responses, server-side context can be continued by passing the previous response identifier.
Core pattern
response = client.responses.create(
model="my-deployment",
instructions="You are a helpful assistant.",
input=user_input,
previous_response_id=last_response_id
)
print(response.output_text)
last_response_id = response.id
For authentication, prefer Microsoft Entra ID over embedded API keys. DefaultAzureCredential is useful because the same application code can obtain credentials across local development and Azure-hosted environments.
Also know the configuration controls: temperature influences response variability, while max tokens constrains output size.
For applications performing other I/O while waiting for model responses, use the asynchronous client + await to avoid blocking execution.
Remember: responses.create() generates a response; response.output_text retrieves its text; response.id can connect the next turn.
Comments