Summary

Completed

Evaluating generative AI applications requires more than a single score. You need representative data, the right evaluator mix, and a disciplined way to interpret the results. Microsoft Foundry supports portal and SDK-based evaluation workflows, while the built-in evaluator catalog helps you assess writing quality, similarity to ground truth, RAG behavior, safety, and agent behavior.

The most effective evaluation practice combines automated runs with targeted human review. Use real-world data when possible, supplement it with synthetic data when coverage is limited, use AI red teaming or other adversarial testing when you need to probe safety and security risks, and compare runs against a stable baseline before you decide a change improved the system.

Before you operationalize the workflow, confirm each evaluator's required inputs, target support, preview status, and region support in current Microsoft Learn guidance. That check matters most for cloud evaluation, safety and red-team workflows, custom evaluators, graders, and some agent-focused evaluators.

Once you have results, translate them into action. Improve retrieval when groundedness or relevance is weak, strengthen safety instructions and filtering when harm evaluators surface risk, and add custom evaluators when your business criteria go beyond the built-in catalog.

Learn more