Select, deploy, and evaluate Microsoft Foundry models | AI-103 | Episode 2
Choosing a model is more than picking the most capable option. In Microsoft Foundry, the practical lifecycle is Select → Deploy → Evaluate: balance model capability, performance, cost, data-location requirements, and measured output quality.
1. Select the model
Use the Model Catalog and Leaderboard to narrow candidates by capabilities, supported languages, context window, fine-tuning support, and benchmarks.
| Benchmark | What it tells you |
|---|---|
| Quality | Usefulness and overall response quality |
| Safety | Susceptibility to harmful or adversarial inputs |
| Throughput | How quickly the model processes and returns output |
| Cost | Price based on input/output token usage |
Key trade-off: a larger model may deliver better results, while a smaller model can offer higher throughput and lower cost.
2. Choose the deployment
Know these deployment choices:
-
Global → broadest capacity and potentially highest throughput
-
Data Zone → processing stays within a geographic zone such as the EU or US
-
Regional → maximum control over the processing region, but capacity is constrained to it
-
Standard → usage/token-based; suitable for general workloads
-
Provisioned → predictable, guaranteed throughput
-
Batch → high-volume, non-interactive processing where latency is less important
-
Developer → lightweight testing of fine-tuned models
Remember: Global Standard is the general-purpose choice highlighted for obtaining the largest available quota.
3. Evaluate before production
Model performance isn't just speed. Evaluate quality, relevance, fluency, and groundedness.
Manual evaluation: run representative prompts and edge cases against models side-by-side.
Automated evaluation: use a larger prompt dataset, expected behavior, and evaluators to systematically score responses and safety.
High-value distinctions to remember:
Throughput = speed/capacity of responses · Fluency = natural, linguistically correct output · Groundedness = whether responses align with known information
Mental model:Requirements → Catalog/Benchmarks → Deployment → Test Dataset → Evaluators → Compare → Improve
Comments