Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Prepare preference data and create a direct preference optimization (DPO) job in Microsoft Foundry. For method comparisons and training concepts, see the fine-tuning overview.
Prerequisites
- A Foundry resource with a DPO-supported model and training region.
- The Foundry User role for training, or the Foundry Owner role if you also deploy the fine-tuned model.
- The client packages and credentials described in the Python procedure.
- Your resource endpoint and credentials, and a Bash-compatible shell for the REST examples.
- Azure Developer CLI version 1.22.1 or later.
Important
The Foundry RBAC roles were recently renamed. Foundry User, Foundry Owner, Foundry Account Owner, and Foundry Project Manager were previously named Azure AI User, Azure AI Owner, Azure AI Account Owner, and Azure AI Project Manager. You might still see the previous names in some places while the rename rolls out. The role IDs and core permissions are unchanged by the rename.
Prepare preference data
Start with the Orca preference-pair sample dataset on GitHub, or prepare your own data.
Save separate training.jsonl and validation.jsonl files, with one example per line. Each example contains an input and exactly two outputs:
| Field | Content |
|---|---|
input |
A messages array with the prompt. Optional tools and parallel_tool_calls also belong here. |
preferred_output |
The preferred response, including at least one assistant message. |
non_preferred_output |
The rejected response to the same input, including at least one assistant message. |
Output messages use the assistant or tool role. Both responses must answer the same prompt, and the preference must match your intended task.
{"input": {"messages": [{"role": "system", "content": "You are a chatbot assistant. Given a user question with multiple choice answers, provide the correct answer."}, {"role": "user", "content": "Question: Janette conducts an investigation to see which foods make her feel more fatigued. She eats one of four different foods each day at the same time for four days and then records how she feels. She asks her friend Carmen to do the same investigation to see if she gets similar results. Which would make the investigation most difficult to replicate? Answer choices: A: measuring the amount of fatigue, B: making sure the same foods are eaten, C: recording observations in the same chart, D: making sure the foods are at the same temperature"}]}, "preferred_output": [{"role": "assistant", "content": "A: measuring the amount of fatigue"}], "non_preferred_output": [{"role": "assistant", "content": "D: making sure the foods are at the same temperature"}]}
Reference: DPO dataset examples.
Keep validation and final test examples out of the training set. Check supported models before using a base model or an SFT-fine-tuned model for DPO.
Configure hyperparameters
Start with your selected model's defaults. Compare validation results before changing the balance between learning your preferences and staying close to the starting model.
Use the settings available in the job configuration. The available controls and accepted values depend on your selected model.
The v1 examples below omit optional hyperparameters. To override defaults, use supported settings under method.dpo.hyperparameters in the v1 fine-tuning API.
Keep the default settings in your model's YAML configuration for your first job. Use the configuration sample for your selected model rather than copying settings from another model.
Create a DPO job in the portal
Use your prepared datasets to create the job:
- Open Build > Fine-tune in the Foundry portal, and select Fine-tune.
- Select a supported model and Direct Preference Optimization.
- Upload the training and validation datasets, and resolve any validation errors.
- Select a supported training type, and leave hyperparameters at their defaults for your first job.
- Review the configuration and create the job. Retain its ID.
Create a DPO job with Python
Initialize the client with the Microsoft Foundry SDK for your project, or use the OpenAI SDK with a resource endpoint. Both clients use the same upload and job commands below.
Install azure-ai-projects, azure-identity, and openai. Sign in with a credential supported by DefaultAzureCredential, and set FOUNDRY_PROJECT_ENDPOINT to your project endpoint:
import os
from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential
project = AIProjectClient(
endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"],
credential=DefaultAzureCredential(),
)
client = project.get_openai_client()
Reference: AIProjectClient and DefaultAzureCredential.
Set FINE_TUNING_MODEL to an identifier from supported model IDs and versions that supports DPO with this API.
Set FINE_TUNING_TRAINING_TYPE to the API value for a supported training type. For example, Global training uses GlobalStandard.
For DPO after SFT, use the supported fine-tuned model's identifier, not its deployment name.
Upload the files and submit the job with default hyperparameters:
import os
with open("training.jsonl", "rb") as training:
training_file = client.files.create(file=training, purpose="fine-tune")
with open("validation.jsonl", "rb") as validation:
validation_file = client.files.create(
file=validation, purpose="fine-tune"
)
client.files.wait_for_processing(training_file.id)
client.files.wait_for_processing(validation_file.id)
job = client.fine_tuning.jobs.create(
model=os.environ["FINE_TUNING_MODEL"],
training_file=training_file.id,
validation_file=validation_file.id,
method={"type": "dpo"},
extra_body={
"trainingType": os.environ["FINE_TUNING_TRAINING_TYPE"],
},
)
print(job.id, job.status)
Reference: Fine-tuning API.
Retain the returned job ID and continue to monitor the job.
Create a DPO job with REST
Set AZURE_OPENAI_ENDPOINT and AZURE_OPENAI_API_KEY for your resource. Run these commands in a Bash-compatible shell.
Upload each dataset with purpose=fine-tune, and retain the distinct file IDs:
curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/files" \
-H "api-key: $AZURE_OPENAI_API_KEY" \
-F "purpose=fine-tune" -F "file=@training.jsonl"
curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/files" \
-H "api-key: $AZURE_OPENAI_API_KEY" \
-F "purpose=fine-tune" -F "file=@validation.jsonl"
Reference: Files API.
Replace <SUPPORTED_MODEL_ID> with an identifier from supported model IDs and versions that supports DPO with this API. For DPO after SFT, use the supported fine-tuned model's identifier, not its deployment name.
Replace <SUPPORTED_TRAINING_TYPE> with the API value for a supported training type. For example, Global training uses GlobalStandard.
Replace the file IDs and submit the job with default hyperparameters:
curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/fine_tuning/jobs" \
-H "api-key: $AZURE_OPENAI_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "<SUPPORTED_MODEL_ID>",
"training_file": "<TRAINING_FILE_ID>",
"validation_file": "<VALIDATION_FILE_ID>",
"trainingType": "<SUPPORTED_TRAINING_TYPE>",
"method": {"type": "dpo"}
}'
Reference: v1 fine-tuning API.
Retain the returned job ID and continue to monitor the job.
Use the Azure Developer CLI
Install Azure Developer CLI version 1.22.1 or later, then install the fine-tuning extension and sign in:
azd ext install azure.ai.finetune
azd auth login
Reference: Azure Developer CLI.
Download the YAML configuration and its data or grader files for your method:
| Method | Configuration samples |
|---|---|
| SFT | Supervised fine-tuning. |
| DPO | Direct preference optimization. |
| RFT | Reinforcement fine-tuning and graders. |
Review the model, method, data paths, and training type. From the configuration directory, initialize the project, submit the job, and inspect its status:
azd ai finetuning init -e <project-endpoint>
azd ai finetuning jobs submit -f <path-to-job-yaml>
azd ai finetuning jobs show -i <job-id>
Use a project endpoint in the form https://<account>.services.ai.azure.com/api/projects/<project>. Replace <job-id> with the ID returned by submission.
Reference: Fine-tuning CLI samples and commands.
Pause, resume, or cancel
Run lifecycle commands only when your model, method, and job state support them. Pause and resume aren't available for every model or method.
azd ai finetuning jobs pause -i <job-id>
azd ai finetuning jobs resume -i <job-id>
azd ai finetuning jobs cancel -i <job-id>
Reference: Fine-tuning CLI commands.
Review training results
Review the metrics reported for your DPO job using the monitoring workflow below. DPO learns from preference pairs; token-prediction accuracy isn't a measure of whether responses match your preferences.
Compare responses from candidate checkpoints on held-out preference examples. Use the same rubric as your dataset labels, and record how often each candidate meets your preferred behavior. This is a task evaluation, not a token-accuracy metric.
Don't choose a model based only on training loss. For evaluation risks, see challenges and limitations.
Monitor the job and select a checkpoint
Inspect training progress before choosing a model or checkpoint to deploy.
In the portal, open the job details:
- Check the Status and event logs. Jobs can queue before training starts; inspect error details if a job fails.
- Open Monitor to compare training and validation metrics using the guidance above.
- Open Checkpoints to inspect available model versions and their metrics. Compare candidates on held-out tasks before choosing one to deploy.
Inspect logs and checkpoints with Python
Use the Python client from your submission example. Replace <JOB_ID> with your job ID. Repeat the status check until the job finishes; don't resubmit a queued job.
Retrieve the status, event logs, available checkpoints, and result-file IDs:
job = client.fine_tuning.jobs.retrieve("<JOB_ID>")
print("Status:", job.status)
print("Error:", job.error)
print("Model:", job.fine_tuned_model)
print("Result files:", job.result_files)
for event in client.fine_tuning.jobs.list_events(job.id).data:
print(event.created_at, event.message)
checkpoints = client.fine_tuning.jobs.checkpoints.list(job.id)
print(checkpoints.model_dump_json(indent=2))
Reference: Fine-tuning API.
After completion, if the job returns a CSV metrics file, replace <RESULT_FILE_ID> with its ID and download it:
with open("results.csv", "wb") as result_file:
result_file.write(client.files.content("<RESULT_FILE_ID>").read())
Reference: Files API.
Inspect logs and checkpoints with REST
Use the resource endpoint and key for REST. Replace <JOB_ID> with your job ID. Repeat the status check until the job finishes; don't resubmit a queued job.
Retrieve the job, event logs, and available checkpoints:
JOB_URL="$AZURE_OPENAI_ENDPOINT/openai/v1/fine_tuning/jobs/<JOB_ID>"
curl "$JOB_URL" -H "api-key: $AZURE_OPENAI_API_KEY"
curl "$JOB_URL/events" -H "api-key: $AZURE_OPENAI_API_KEY"
curl "$JOB_URL/checkpoints" -H "api-key: $AZURE_OPENAI_API_KEY"
Reference: Fine-tuning API.
After completion, inspect result_files in the job response. If it includes a CSV metrics file, replace <RESULT_FILE_ID> with its ID and download it:
curl "$AZURE_OPENAI_ENDPOINT/openai/v1/files/<RESULT_FILE_ID>/content" \
-H "api-key: $AZURE_OPENAI_API_KEY" --output results.csv
Reference: Files API.
Use the Azure Developer CLI job workflow to inspect the job status.
Checkpoints appear as training progresses; a queued job might have none. Inspect the returned checkpoint identifiers and metrics rather than assuming the latest version performs best.
When training succeeds, retain the trained model or your chosen checkpoint. Compare candidates on held-out tasks before selecting one for deployment.
Deploy the model
Use the Foundry portal to deploy the model, regardless of how you submit the training job. Follow Deploy fine-tuned models for your model's supported serving option and inference procedure.
After deployment, use Run evaluations from the Foundry portal to compare candidates on held-out tasks.
For SDK evaluation, see Cloud evaluation with the Microsoft Foundry SDK.
Stop training and clean up
Cancel an unneeded job from its portal view. Delete uploaded files separately through the portal when you no longer need them.
Cancel an unneeded job with client.fine_tuning.jobs.cancel("<JOB_ID>"). Delete unused uploaded files separately with client.files.delete("<FILE_ID>").
Reference: Job cancellation and file deletion.
Cancel an unneeded job with the fine-tuning cancellation API. Delete unused uploaded files separately with the Files API.
Cancel an unneeded job with the CLI commands above. For available cleanup operations, see the fine-tuning CLI reference.
Delete unused deployments with the deployment cleanup procedure. Cancelling training doesn't delete deployments or stop their charges.
Retain any data and job records you need to reproduce the experiment.