Plan evaluations for grounded answers

Completed

A round of Preview tests shows how the agent answers at the time you test it, and those answers can change when a content owner edits a document, you rewrite a source description, or you add a source. In this unit, you learn how evaluations let you rerun saved test questions, which knowledge questions to keep, what an evaluation result shows, and when to run it again.

How evaluations work

The Evaluate tab is a preview feature that provides structured, repeatable testing for an agent. An evaluation is a named test set of conversations. Each conversation is a test case with one or more user messages and, optionally, an expected response. You can write conversations yourself, generate them with AI, or upload them from a CSV file.

When you run an evaluation, a test method scores each response. The General quality test method is an AI-based check of whether a response meets quality standards such as relevance and completeness. General quality doesn't compare a response with the expected response. After several runs, you can compare scores to see whether a change improved or reduced the quality of answers.

Learn more in Evaluate an agent (preview).

Choose which questions to keep

Start from the questions you tested in Preview. You already know the expected source and supporting passage for each one, so they make good evaluation cases. Keep a question when:

  • It covers a source employees rely on. Keep at least one in-scope question for each knowledge source, so a problem with any source shows up.
  • It uses employees' own words. A paraphrase shows whether the agent still finds the answer when a question doesn't repeat a document title.
  • It sounds like it belongs to one source, but another source answers it. These questions show whether adding or editing a source changes which source the agent picks.
  • No source should answer it. An unsupported question shows whether the agent starts inventing answers.
  • It failed before and you fixed it. A fixed question shows whether a later change undoes the fix.

For example, for your help agent, you might keep the remote work policy question and its paraphrase, the laptop model question, the docking-station sign-in question, and the pet insurance question.

If you generate conversations with AI, review them before you keep them, and add the questions you planned yourself. You know which source should answer each planned question. Because an evaluation runs the same questions every time, use questions that don't contain real personal or sensitive information.

Know what a result shows

A General quality Pass means the response met quality standards such as relevance and completeness. It doesn't show that the answer came from the intended source or that it matches your expected response. For example, the laptop model question could pass with a clear answer drawn from the purchasing policy instead of the purchasing records table.

Keep the expected source and supporting passage with each question, either in the expected response or in your own notes. When a result looks wrong, open the test case to read the agent's response. Then ask the same question in Preview and use the activity trace to see which source the agent searched.

Evaluations run under a user profile that you set in the Configure test set panel, and the results show which profile ran each evaluation. When you compare runs, check that the same profile ran them. An evaluation doesn't replace tests by employees whose access differs from yours.

Decide when to run it again

Run the evaluation again after changes that can affect knowledge answers:

  • You add, remove, or edit a knowledge source, including its name or description.
  • You change the agent's instructions.
  • A content owner makes a major update to content that your questions depend on.
  • You're about to publish the agent.

Compare each run's score with the previous one. When the score drops, open the test cases that fail, and check those questions in Preview to find the cause.

Reflect: Which change to your agent's knowledge is most likely to happen without you knowing about it? How would a regular evaluation help you notice it?

Learn more in View evaluation results for an agent (preview).