Test tool use and plan evaluations

Completed

Your help agent now has the Equipment request status and submission tool added and described, but a response that looks right doesn't prove that the intended tool ran with the right inputs. In this unit, you learn how to confirm tool use in Preview with the activity trace, and how to decide which tool-use cases to keep checking as your agent changes.

Plan positive and negative requests before you test

A test request is only useful when you know what should happen. Decide the expected outcome for each request before you send it in the Preview tab, so that you have something specific to compare the agent's behavior against.

Plan two kinds of requests. A positive request should lead the agent to call the tool. A negative request shouldn't, either because another source answers it or because the agent needs more information first.

For each request, write down the tool and action you expect the agent to call, if any, the input values you expect it to send, and the result you expect back. The following table shows a plan for the help agent.

Request Expected tool Expected inputs Expected result
"What's the status of request 1042?" Equipment request status and submission, Get item if 1042 is the SharePoint item ID Item ID 1042 The status currently stored in item 1042, if the caller has access
"I need a new laptop dock." Equipment request status and submission, Create item, after the employee confirms The details the employee confirmed A new item in the list and its request number
"Who qualifies for a second monitor?" None, because knowledge answers it None An answer grounded in the policy documents
"What's the status of my request?" None until the employee provides a request number None The agent asks for the request number

The last two rows are negative requests. The policy question tests the boundary between the tool and knowledge, and the request without a number tests the instruction to ask for a missing request number.

For this example, 1042 is the SharePoint item ID of a test record the caller is allowed to read. Before you send the status request, open that item and note its current status, so that the lookup has a known value to compare against. If your request numbers use a separate column, configure and test Get items with a filter instead.

Learn more in Test an agent.

Confirm what ran in the activity trace

The response text tells you what the agent said. It doesn't show which tool, if any, produced the answer, so a plausible status isn't evidence that the lookup ran. The activity trace provides that evidence.

The trace appears alongside the chat in the maker view, which is the default when End user preview is turned off. It updates in real time as the conversation runs, and each step appears as a node:

  • Ran action: A collapsible card for a connector or tool invocation that shows the connector, the filled input parameters, and the duration.
  • Tool and Connector: Nodes that represent tool execution.
  • Knowledge: A Searched knowledge card that appears when the agent searches knowledge.
  • Complete: A green checkmark that appears when all tool calls finish.

Select a node to see the input the agent sent, the output it received, processing time and metrics, and any error details. With those details, compare the trace with the expected outcome for each planned request:

  • Tool: For a positive request, confirm that the tool execution node names the tool and action you expected. For a negative request, confirm that no tool node appears. For the policy question, expect a Knowledge node in its place.
  • Inputs: Check that the input parameters hold the values you expected, such as request number 1042.
  • Result: Compare the output with the item in the list. A status that doesn't match the item points to a problem even when the response reads well.

For the submit request, check the order of events too. The Ran action card for Create item should appear only after the employee confirms the details, and the new item should then appear in the list.

Learn more in Use the activity trace to debug your agent.

Diagnose tool problems by layer

When a check fails, the trace evidence points to the layer that needs a fix. The following table organizes common problems by what you see in the trace.

Layer Evidence in the trace Where to fix
Routing No tool node for a positive request, or a node for a different tool or action The tool's name and description, your agent's instructions, or whether the tool is added at all
Inputs The expected tool node, with missing or wrong input values The Inputs panel in the Tool details dialog, and the part of the description that states what the tool needs
Results The expected tool and inputs, with wrong or incomplete output The tool's configuration, including parameters and selected actions, and the data in the list
Connection An Error node, such as an authentication failure or a timeout on a long-running tool call The credentials in the Details panel, service availability, or your administrator for access questions

Check the layers in order, because a later layer only produces evidence when the earlier ones work. If the agent never calls the tool, the trace has no inputs or output to inspect.

To go from evidence to a fix, select the tool or source in the trace to open that component on the Build tab. After you make the fix, send the same request again and compare the new trace with your expected outcome.

Learn more in Manage and delete tools in an agent.

Choose cases to keep for repeatable evaluation

Preview tells you whether the tool works in the conversation you're testing now. As your agent changes, evaluation lets you check the same behavior again.

The Evaluate tab, currently in preview, provides structured, repeatable testing across test cases. Evaluations help answer questions such as "Does the agent call tools when it should?" and "Did a configuration change improve or degrade response quality?" The two tabs complement each other: you use Preview to explore and refine, and you use Evaluate to validate at scale.

After you add a tool, consider keeping at least these cases:

  • A positive case: A request that should call the tool, such as "What's the status of request 1042?"
  • A negative case: A request that shouldn't call the tool, such as the policy question.
  • An input case: A request that checks whether the agent extracts the request number from other wording, such as "Can you check on 1042?"
  • A known-result lookup: A read-only lookup against an authorized test item whose status you keep stable. A lookup doesn't change the item, but you still need to check the evaluation account's access and the sensitivity of the data before rerunning it.

Treat write actions differently. Each run of a case that creates an item adds another item to the list. Don't keep a write case for repeated runs unless it runs against test data and you account for the items each run creates.

Keep the expected outcome with each case so you can compare results after changes to the tool, its description, instructions, or knowledge. Check the tool-call evidence in Preview when you need to confirm which action ran with which inputs; a good-looking answer or evaluation result alone doesn't establish that. Before you reuse a case, confirm that the account running it has a working connection and access to the test data. Building, testing, and evaluating agents might consume Copilot Credits, so keep the retained set focused on the cases that matter.

Learn more in Preview and test an agent and Evaluate an agent.

Reflect: For a tool in an agent you're planning or building, which four cases would you keep for repeatable evaluation? Which of them can you rerun against live data without changing anything?