This browser is no longer supported.
Upgrade to Microsoft Edge to take advantage of the latest features, security updates, and technical support.
Choose the best response for each of the questions below.
What is the primary purpose of evaluating generative AI applications?
To increase the speed of AI model training.
To measure the quality, safety, and reliability of AI systems.
To reduce the cost of AI development.
To replace human evaluators with automated systems.
Which of the following isn't a characteristic of good evaluation data?
Diversity.
Representativeness.
High quality.
Homogeneity.
Which statement best reflects how evaluation data should be prepared in Microsoft Foundry?
Every built-in evaluator uses the same input fields, so one fixed schema always works.
Different evaluators can require different inputs, so confirm the mapping for fields such as query, response, context, ground truth, or tool calls.
Only custom evaluators ever use context or ground truth.
You should remove edge cases so they don't distort the average score.
When comparing two evaluation runs after a prompt change, what gives you the clearest evidence that the change helped?
Compare the new run against a baseline while keeping the dataset and evaluator set stable.
Run the new prompt on a different dataset with different evaluators.
Look only at the overall pass/fail rate and ignore row-level results.
Change the thresholds at the same time so more results pass.
You must answer all questions before checking your work.
Was this page helpful?
Need help with this topic?
Want to try using Ask Learn to clarify or guide you through this topic?