Learning by Patrik

Develop a generative AI chat app with Microsoft Foundry | AI-103 | Episode 3

A chat application becomes much easier to design once you understand three decisions: which endpoint to use, which API to call, and where conversation state is maintained.

Endpoint & API choices

Choice Remember this
Azure OpenAI endpoint Direct model access; typically use the OpenAI SDK
Foundry project endpoint Higher-level access to models, tools, and agents
Chat Completions API Client resends the conversation history
Responses API Can link turns using a previous response ID

Key concept: LLMs are inherently stateless. Chat history must therefore be supplied or referenced. With Chat Completions, your application maintains and resends the message array. With Responses, server-side context can be continued by passing the previous response identifier.

Core pattern

response = client.responses.create(
    model="my-deployment",
    instructions="You are a helpful assistant.",
    input=user_input,
    previous_response_id=last_response_id
)

print(response.output_text)
last_response_id = response.id
 

For authentication, prefer Microsoft Entra ID over embedded API keys. DefaultAzureCredential is useful because the same application code can obtain credentials across local development and Azure-hosted environments.

Also know the configuration controls: temperature influences response variability, while max tokens constrains output size.

For applications performing other I/O while waiting for model responses, use the asynchronous client + await to avoid blocking execution.

Remember: responses.create() generates a response; response.output_text retrieves its text; response.id can connect the next turn.

Comments