Validate grounded answers and refine source selection
Adding a source makes its content available, but you still need evidence that the agent uses the right source for each question before people rely on its answers. In this unit, you learn how to plan test questions, use the activity trace in Preview to see what the agent searched, and match a weak answer to its cause and fix.
Build a test question set
A test is useful only when you know the expected result before you read the answer. Before you open Preview, write down each question, the source that should answer it, and the passage or record that supports the answer. A useful set includes four kinds of questions:
- An in-scope question for each source: A question whose answer the source owner can point to.
- A paraphrase in users' own words: The agent uses document titles to help select SharePoint content, so a question that repeats a title word for word is an easy case. A paraphrase tests whether employees still get the answer when they ask in their own words.
- A question that belongs to a different source: A question that sounds like it belongs to one source but is answered by another. It tests whether the agent picks the right one.
- An unsupported question: A question that no configured source answers. For this question, a good outcome is an agent that doesn't invent an answer.
For example, a first question set for the help agent might look like this:
| Question | Expected source | What it tests |
|---|---|---|
| What does the remote work policy say about working from home full time? | Policy library in SharePoint | In-scope retrieval |
| Can I skip the office and do my job from my house every day? | Policy library in SharePoint | Paraphrase findability |
| Which laptop models am I allowed to order? | Purchasing records table in Dataverse | Selection between the policy library and the table |
| How was the recurring sign-in error on docking stations resolved? | Known issue resolution history (Copilot connector) | In-scope retrieval |
| What's the reimbursement limit for pet insurance? | None, because no policy covers it | Unsupported question |
Plan who runs the tests, too. In Microsoft 365, each item that a synced Copilot connector indexes carries an access control list that enforces the source system's permissions. Microsoft 365 Search and Copilot show items only to users who have access in the source system. Your own tests can't show what an employee with different access sees. Ask your admin how the connection is set up for your agent, and plan tests with employees whose permissions differ from yours.
Record which account and environment you use for each test. If results change later, you can then tell whether the account, the environment, or a content or configuration change caused it. Use test questions and records that don't contain real sensitive data.
Inspect the activity trace
The activity trace shows how the agent produces each answer, including the knowledge sources it searches. To inspect a knowledge answer:
- On the Preview tab, confirm that the End user preview toggle is off. The maker view is the default, and the activity trace appears alongside the chat pane.
- Ask a question from your test set.
- In the trace, find the Knowledge node. It appears as a Searched knowledge collapsible card that lists each knowledge source accessed and shows citation links when they're available.
- Select a source to see the specific content retrieved.
- Compare the consulted source and the retrieved content with your expected source, your supporting passage, and the agent's answer.
An answer might include citations. When it does, open each citation and check the cited content against the answer. Not every answer includes citations, and a citation doesn't prove that the answer is complete.
The trace also shows other node types, such as Tool and Ran action, when the agent invokes a tool. Recall that knowledge grounds answers in content, while tools suit live data and actions. If a question you expected to answer from knowledge invokes a tool instead, treat that path as a source-selection signal.
The End user preview toggle switches between the maker view and the end-user view. It doesn't show what another user can access.
Learn more in Use the activity trace to debug your agent.
Match the observation to a cause
Each pattern in the trace points to a likely cause. The cause determines what you change and who usually makes the change.
| Observation | Likely cause | What to change | Who usually owns the fix |
|---|---|---|---|
| The intended source doesn't appear in the Searched knowledge card. | Source selection | Confirm that the source appears in the Knowledge section of the Build tab. Check that its name and description identify these questions, and look for another source that covers similar terms. For SharePoint, check the quality of document titles. | Maker adjusts the description or scope. Content owner fixes titles. |
| The intended source appears, but the retrieved content is wrong or stale. | Content | Update the authoritative document or record. SharePoint content changes reach the agent automatically. Update the source configuration only when its location or connection settings change. | Content owner |
| Relevant content is retrieved, but the answer is incomplete or goes beyond it. | Instructions or presentation | Revise the agent's instructions, which influence how the agent interprets and presents knowledge. | Maker |
| The unsupported question gets an answer anyway. | Grounding scope | Check whether the answer draws on configured content. When Search all websites is on, the agent can supplement answers with web content. Turn it off, or narrow the sources or the instructions. | Maker |
| An Error node appears, such as a knowledge source access error. | Access or connection | Review the error details, record the identity and environment, and address the root cause on the Build tab or in Settings. | Maker, working with the admin or source owner |
| The answer works for you but not for another employee. | Audience access | Check that employee's access to the source content. | Source owner or admin |
Fix each problem where it starts. Rewriting instructions doesn't update a stale policy document, and a better description doesn't give a user access to a record.
For example, suppose the laptop question returns a paragraph from the purchasing policy, and the Searched knowledge card lists only the policy library. The agent doesn't search the purchasing records table, so the cause is source selection. If the table's description reads Purchasing data, rewrite it to name the questions it answers, such as which laptop models employees can order.
You usually see the last row's problem only when someone else tests the agent. Before employees rely on the agent, ask a few of them to test it in the channel where you publish it, in addition to your own Preview tests.
Retest after each change
Change one factor at a time. If you revise a description while a content owner updates a document, an improved answer doesn't show which change worked.
To make a change, select the source or component in the trace, and Copilot Studio opens it on the Build tab. A turn in progress finishes with the previous configuration, and the next turn picks up your edits. If a change doesn't take effect after a couple of turns, select New chat in the preview header to start a conversation with no prior context.
Next, rerun the full question set, including the unsupported question. Compare the same fields as in the first round: expected source, consulted source, retrieved content, and answer. The full rerun catches a change that fixes one answer and breaks another, such as removing a source that a different question still depends on.
As you test, select the thumbs up or thumbs down icon on a response and enter a comment. Use that feedback to track quality patterns across test rounds.
Reflect: Pick the source in your agent whose owner you work with least often. If a test shows stale content from that source, how does the correction reach the owner, and how do you know when the fix is ready to retest?
The agent answers as intended when each in-scope question comes from the intended source, the unsupported question gets no invented answer, and the employees who test it see the same results you do. That tested question set is also the starting point for a repeatable evaluation.
Learn more in Test an agent.