Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Generate test scenarios from one or more agent, prompt, or reference-file sources, simulate multi-turn conversations against the agent, and evaluate the conversations in one evaluation run. Use this workflow when you don't have hand-authored scenarios or representative conversation data.
To provide your own scenarios instead, see Simulate conversations with the Microsoft Foundry SDK.
Prerequisites
- Complete the cloud evaluation prerequisites and client setup.
- Install the packages for your language:
- Python:
pip install "azure-ai-projects>=2.5.0" python-dotenv - C#:
dotnet add package Azure.AI.Projects --prereleaseanddotnet add package Azure.Identity - JavaScript:
npm install @azure/ai-projects @azure/identity dotenv
- Python:
- Deploy a model to generate scenarios, simulate users, and run AI-assisted evaluators.
Set these environment variables:
FOUNDRY_PROJECT_ENDPOINT: Your Foundry project endpoint.FOUNDRY_MODEL_NAME: The model deployment used for scenario generation, user simulation, and AI-assisted evaluators.FOUNDRY_AGENT_NAME: Optional. The agent name. The example usesMyAgentwhen this variable isn't set.
How synthetic conversation evaluation works
The azure_ai_synthetic_data_generation_with_simulation data source runs the complete workflow:
- Generate test-case scenarios from one or more agent, prompt, or reference-file sources.
- Save the generated scenarios as a dataset.
- Simulate multi-turn conversations against the agent.
- Save the generated conversations as a dataset.
- Evaluate the conversations with conversation-level evaluators.
Generation and simulation happen in the same evaluation run. You don't need to create a separate synthetic-data generation job.
The generation_sources array accepts multiple sources in one job. Combining sources can produce broader scenario coverage. For example, use an agent definition to reflect the agent's instructions and persona, a prompt to steer scenario difficulty or domain, and a reference file to ground scenarios in longer source material. All generated scenarios feed the same multi-turn conversation simulation step.
You can combine these source types:
- Agent definition (
"type": "agent"): Uses a deployed agent's name, version, and instructions. - Prompt (
"type": "prompt"): Uses inline text to describe the domain or steer the generated scenarios. - Reference file (
"type": "file"): Uses an uploaded file ID to ground scenarios in source material.
Configure the client and agent
Create a project client and an agent to evaluate:
import os
import time
from pprint import pprint
from dotenv import load_dotenv
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import (
AzureAIDataSourceConfig,
PromptAgentDefinition,
TestingCriterionAzureAIEvaluator,
)
SEED_COUNT = 1
CONVERSATIONS_PER_SEED = 1
MAX_TURNS = 2
DESIRED_TURNS = 1
load_dotenv()
endpoint = os.environ["FOUNDRY_PROJECT_ENDPOINT"]
model_deployment_name = os.environ["FOUNDRY_MODEL_NAME"]
agent_name = os.environ.get("FOUNDRY_AGENT_NAME", "MyAgent")
credential = DefaultAzureCredential()
project_client = AIProjectClient(endpoint=endpoint, credential=credential)
client = project_client.get_openai_client()
agent = project_client.agents.create_version(
agent_name=agent_name,
definition=PromptAgentDefinition(
model=model_deployment_name,
instructions="You are a helpful customer service agent. Be empathetic and solution-oriented.",
),
)
Configure the evaluation
Use the synthetic_data_gen scenario for the evaluation group. Conversation-level evaluators receive the generated conversation through {{item.messages}}. Evaluators that assess tool use also receive {{item.tool_definitions}}.
data_source_config = AzureAIDataSourceConfig(
type="azure_ai_source",
scenario="synthetic_data_gen",
)
testing_criteria = [
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="tool_use_quality",
evaluator_name="builtin.tool_use_quality",
initialization_parameters={"model": model_deployment_name},
data_mapping={
"messages": "{{item.messages}}",
"tool_definitions": "{{item.tool_definitions}}",
},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="output_quality",
evaluator_name="builtin.output_quality",
initialization_parameters={"model": model_deployment_name},
data_mapping={
"messages": "{{item.messages}}",
"tool_definitions": "{{item.tool_definitions}}",
},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="deflection_rate",
evaluator_name="builtin.deflection_rate",
initialization_parameters={"model": model_deployment_name},
data_mapping={
"messages": "{{item.messages}}",
"tool_definitions": "{{item.tool_definitions}}",
},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="customer_satisfaction",
evaluator_name="builtin.customer_satisfaction",
initialization_parameters={"model": model_deployment_name},
data_mapping={"messages": "{{item.messages}}"},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="task_completion",
evaluator_name="builtin.task_completion",
initialization_parameters={"model": model_deployment_name},
data_mapping={"messages": "{{item.messages}}"},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="coherence",
evaluator_name="builtin.coherence",
initialization_parameters={"model": model_deployment_name},
data_mapping={"messages": "{{item.messages}}"},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="groundedness",
evaluator_name="builtin.groundedness",
initialization_parameters={"model": model_deployment_name},
data_mapping={"messages": "{{item.messages}}"},
),
]
eval_object = client.evals.create(
name="Synthetic Multi-turn Evaluation",
data_source_config=data_source_config,
testing_criteria=testing_criteria,
)
Generate, simulate, and evaluate
Create one run with the azure_ai_synthetic_data_generation_with_simulation data source:
eval_run = client.evals.runs.create(
eval_id=eval_object.id,
name="synthetic-multiturn-run",
data_source={
"type": "azure_ai_synthetic_data_generation_with_simulation",
"synthetic_data_generation_configuration": {
"test_case_count": SEED_COUNT,
"output_test_case_dataset_name": f"{agent_name}-synthetic-scenarios",
"generation_sources": [
{
"type": "prompt",
"prompt": (
"Generate customer-support scenarios that test difficult "
"billing and refund conversations."
),
},
{
"type": "agent",
"agent_name": agent.name,
"agent_version": agent.version,
},
],
},
"model_configuration": {
"model": model_deployment_name,
},
"default_simulation_configuration": {
"max_num_turns": MAX_TURNS,
"conversation_repetitions": CONVERSATIONS_PER_SEED,
"desired_num_turns": DESIRED_TURNS,
"enable_conversation_dataset_generation": True,
"output_conversation_dataset_name": f"{agent_name}-synthetic-conversations",
},
"target": {
"type": "azure_ai_agent",
"name": agent.name,
"version": agent.version,
},
},
extra_body={"evaluation_level": "conversation"},
)
Note
output_test_case_dataset_name and output_conversation_dataset_name are optional. Specify them when you want recognizable dataset names; otherwise, omit them and the service generates the names.
The OpenAI clients' typed methods don't currently expose all Azure-specific evaluation properties. Python uses extra_body to add evaluation_level to the top level of the REST request body. The TypeScript sample uses as any for the Azure-specific evaluation configuration, criteria, data source, and evaluation_level; the client forwards these properties to the service. C# sets the properties directly in its protocol request body.
The configuration uses these parameters:
| Parameter | Description |
|---|---|
test_case_count |
Number of synthetic scenarios to generate. |
generation_sources |
One or more agent, prompt, or reference-file sources used together to generate scenarios. This example combines a steering prompt with the agent's name, version, and instructions. |
output_test_case_dataset_name |
Optional name for the dataset that stores the generated scenarios. If omitted, the service generates the name. |
model_configuration.model |
Model deployment used to generate scenarios and simulate the user. |
max_num_turns |
Maximum number of turns in each conversation. |
conversation_repetitions |
Conversations to simulate for each generated scenario. |
desired_num_turns |
Preferred number of turns in each conversation. |
enable_conversation_dataset_generation |
Whether to save the simulated conversations as a dataset. |
output_conversation_dataset_name |
Optional name for the dataset that stores the generated conversations. If omitted, the service generates the name. |
Get the results
Synthetic conversation runs can take several minutes. Poll until the run reaches a terminal state, then inspect its output items and report URL:
while True:
run = client.evals.runs.retrieve(
run_id=eval_run.id,
eval_id=eval_object.id,
)
if run.status in ("completed", "failed", "canceled"):
break
print(f"Waiting for simulation to complete... current status: {run.status}")
time.sleep(10)
if run.status != "completed":
raise RuntimeError(f"Simulation run failed: {run.error}")
print(f"Result Counts: {run.result_counts}")
if run.result_counts.errored:
raise RuntimeError(
f"{run.result_counts.errored} evaluation item(s) errored"
)
expected_conversations = SEED_COUNT * CONVERSATIONS_PER_SEED
print(f"Expected up to: {expected_conversations} conversations")
output_items = list(
client.evals.runs.output_items.list(
run_id=run.id,
eval_id=eval_object.id,
)
)
if output_items:
pprint(output_items[0])
print(f"Eval Run Report URL: {run.report_url}")
Delete the evaluation when you no longer need it:
client.evals.delete(eval_id=eval_object.id)
client.close()
project_client.close()
credential.close()
Next steps
- For the complete runnable and recorded example, see sample_synthetic_multiturn_evaluation.py.
- For the complete JavaScript/TypeScript example, see syntheticMultiturnEvaluation.ts.
- To provide your own scenario dataset, see Simulate conversations with the Microsoft Foundry SDK.
- To interpret evaluation results, see Get cloud evaluation results.