Learning by Patrik

Select, deploy, and evaluate Microsoft Foundry models | AI-103 | Episode 2

Choosing a model is more than picking the most capable option. In Microsoft Foundry, the practical lifecycle is Select → Deploy → Evaluate: balance model capability, performance, cost, data-location requirements, and measured output quality.

1. Select the model

Use the Model Catalog and Leaderboard to narrow candidates by capabilities, supported languages, context window, fine-tuning support, and benchmarks.

Benchmark What it tells you
Quality Usefulness and overall response quality
Safety Susceptibility to harmful or adversarial inputs
Throughput How quickly the model processes and returns output
Cost Price based on input/output token usage

Key trade-off: a larger model may deliver better results, while a smaller model can offer higher throughput and lower cost.

2. Choose the deployment

Know these deployment choices:

  • Global → broadest capacity and potentially highest throughput

  • Data Zone → processing stays within a geographic zone such as the EU or US

  • Regional → maximum control over the processing region, but capacity is constrained to it

  • Standard → usage/token-based; suitable for general workloads

  • Provisioned → predictable, guaranteed throughput

  • Batch → high-volume, non-interactive processing where latency is less important

  • Developer → lightweight testing of fine-tuned models

Remember: Global Standard is the general-purpose choice highlighted for obtaining the largest available quota.

3. Evaluate before production

Model performance isn't just speed. Evaluate quality, relevance, fluency, and groundedness.

Manual evaluation: run representative prompts and edge cases against models side-by-side.

Automated evaluation: use a larger prompt dataset, expected behavior, and evaluators to systematically score responses and safety.

High-value distinctions to remember:
Throughput = speed/capacity of responses · Fluency = natural, linguistically correct output · Groundedness = whether responses align with known information

Mental model:
Requirements → Catalog/Benchmarks → Deployment → Test Dataset → Evaluators → Compare → Improve

Foundry
Models
Deployment
Evaluation
Azure

Comments