Edit

Fine-tune a model with direct preference optimization in Microsoft Foundry

Prepare preference data and create a direct preference optimization (DPO) job in Microsoft Foundry. For method comparisons and training concepts, see the fine-tuning overview.

Prerequisites

  • Your resource endpoint and credentials, and a Bash-compatible shell for the REST examples.

Important

The Foundry RBAC roles were recently renamed. Foundry User, Foundry Owner, Foundry Account Owner, and Foundry Project Manager were previously named Azure AI User, Azure AI Owner, Azure AI Account Owner, and Azure AI Project Manager. You might still see the previous names in some places while the rename rolls out. The role IDs and core permissions are unchanged by the rename.

Prepare preference data

Start with the Orca preference-pair sample dataset on GitHub, or prepare your own data.

Save separate training.jsonl and validation.jsonl files, with one example per line. Each example contains an input and exactly two outputs:

Field Content
input A messages array with the prompt. Optional tools and parallel_tool_calls also belong here.
preferred_output The preferred response, including at least one assistant message.
non_preferred_output The rejected response to the same input, including at least one assistant message.

Output messages use the assistant or tool role. Both responses must answer the same prompt, and the preference must match your intended task.

{"input": {"messages": [{"role": "system", "content": "You are a chatbot assistant. Given a user question with multiple choice answers, provide the correct answer."}, {"role": "user", "content": "Question: Janette conducts an investigation to see which foods make her feel more fatigued. She eats one of four different foods each day at the same time for four days and then records how she feels. She asks her friend Carmen to do the same investigation to see if she gets similar results. Which would make the investigation most difficult to replicate? Answer choices: A: measuring the amount of fatigue, B: making sure the same foods are eaten, C: recording observations in the same chart, D: making sure the foods are at the same temperature"}]}, "preferred_output": [{"role": "assistant", "content": "A: measuring the amount of fatigue"}], "non_preferred_output": [{"role": "assistant", "content": "D: making sure the foods are at the same temperature"}]}

Reference: DPO dataset examples.

Keep validation and final test examples out of the training set. Check supported models before using a base model or an SFT-fine-tuned model for DPO.

Configure hyperparameters

Start with your selected model's defaults. Compare validation results before changing the balance between learning your preferences and staying close to the starting model.

Use the settings available in the job configuration. The available controls and accepted values depend on your selected model.

The v1 examples below omit optional hyperparameters. To override defaults, use supported settings under method.dpo.hyperparameters in the v1 fine-tuning API.

Keep the default settings in your model's YAML configuration for your first job. Use the configuration sample for your selected model rather than copying settings from another model.

Create a DPO job in the portal

Use your prepared datasets to create the job:

  1. Open Build > Fine-tune in the Foundry portal, and select Fine-tune.
  2. Select a supported model and Direct Preference Optimization.
  3. Upload the training and validation datasets, and resolve any validation errors.
  4. Select a supported training type, and leave hyperparameters at their defaults for your first job.
  5. Review the configuration and create the job. Retain its ID.

Create a DPO job with Python

Initialize the client with the Microsoft Foundry SDK for your project, or use the OpenAI SDK with a resource endpoint. Both clients use the same upload and job commands below.

Install azure-ai-projects, azure-identity, and openai. Sign in with a credential supported by DefaultAzureCredential, and set FOUNDRY_PROJECT_ENDPOINT to your project endpoint:

import os
from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential

project = AIProjectClient(
    endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"],
    credential=DefaultAzureCredential(),
)
client = project.get_openai_client()

Reference: AIProjectClient and DefaultAzureCredential.

Set FINE_TUNING_MODEL to an identifier from supported model IDs and versions that supports DPO with this API.

Set FINE_TUNING_TRAINING_TYPE to the API value for a supported training type. For example, Global training uses GlobalStandard.

For DPO after SFT, use the supported fine-tuned model's identifier, not its deployment name.

Upload the files and submit the job with default hyperparameters:

import os

with open("training.jsonl", "rb") as training:
    training_file = client.files.create(file=training, purpose="fine-tune")
with open("validation.jsonl", "rb") as validation:
    validation_file = client.files.create(
        file=validation, purpose="fine-tune"
    )
client.files.wait_for_processing(training_file.id)
client.files.wait_for_processing(validation_file.id)

job = client.fine_tuning.jobs.create(
    model=os.environ["FINE_TUNING_MODEL"],
    training_file=training_file.id,
    validation_file=validation_file.id,
    method={"type": "dpo"},
    extra_body={
        "trainingType": os.environ["FINE_TUNING_TRAINING_TYPE"],
    },
)
print(job.id, job.status)

Reference: Fine-tuning API.

Retain the returned job ID and continue to monitor the job.

Create a DPO job with REST

Set AZURE_OPENAI_ENDPOINT and AZURE_OPENAI_API_KEY for your resource. Run these commands in a Bash-compatible shell.

Upload each dataset with purpose=fine-tune, and retain the distinct file IDs:

curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/files" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -F "purpose=fine-tune" -F "file=@training.jsonl"

curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/files" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -F "purpose=fine-tune" -F "file=@validation.jsonl"

Reference: Files API.

Replace <SUPPORTED_MODEL_ID> with an identifier from supported model IDs and versions that supports DPO with this API. For DPO after SFT, use the supported fine-tuned model's identifier, not its deployment name.

Replace <SUPPORTED_TRAINING_TYPE> with the API value for a supported training type. For example, Global training uses GlobalStandard.

Replace the file IDs and submit the job with default hyperparameters:

curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/fine_tuning/jobs" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "<SUPPORTED_MODEL_ID>",
    "training_file": "<TRAINING_FILE_ID>",
    "validation_file": "<VALIDATION_FILE_ID>",
    "trainingType": "<SUPPORTED_TRAINING_TYPE>",
    "method": {"type": "dpo"}
  }'

Reference: v1 fine-tuning API.

Retain the returned job ID and continue to monitor the job.

Use the Azure Developer CLI

Install Azure Developer CLI version 1.22.1 or later, then install the fine-tuning extension and sign in:

azd ext install azure.ai.finetune
azd auth login

Reference: Azure Developer CLI.

Download the YAML configuration and its data or grader files for your method:

Method Configuration samples
SFT Supervised fine-tuning.
DPO Direct preference optimization.
RFT Reinforcement fine-tuning and graders.

Review the model, method, data paths, and training type. From the configuration directory, initialize the project, submit the job, and inspect its status:

azd ai finetuning init -e <project-endpoint>
azd ai finetuning jobs submit -f <path-to-job-yaml>
azd ai finetuning jobs show -i <job-id>

Use a project endpoint in the form https://<account>.services.ai.azure.com/api/projects/<project>. Replace <job-id> with the ID returned by submission.

Reference: Fine-tuning CLI samples and commands.

Pause, resume, or cancel

Run lifecycle commands only when your model, method, and job state support them. Pause and resume aren't available for every model or method.

azd ai finetuning jobs pause -i <job-id>
azd ai finetuning jobs resume -i <job-id>
azd ai finetuning jobs cancel -i <job-id>

Reference: Fine-tuning CLI commands.

Review training results

Review the metrics reported for your DPO job using the monitoring workflow below. DPO learns from preference pairs; token-prediction accuracy isn't a measure of whether responses match your preferences.

Compare responses from candidate checkpoints on held-out preference examples. Use the same rubric as your dataset labels, and record how often each candidate meets your preferred behavior. This is a task evaluation, not a token-accuracy metric.

Don't choose a model based only on training loss. For evaluation risks, see challenges and limitations.

Monitor the job and select a checkpoint

Inspect training progress before choosing a model or checkpoint to deploy.

In the portal, open the job details:

  1. Check the Status and event logs. Jobs can queue before training starts; inspect error details if a job fails.
  2. Open Monitor to compare training and validation metrics using the guidance above.
  3. Open Checkpoints to inspect available model versions and their metrics. Compare candidates on held-out tasks before choosing one to deploy.

Inspect logs and checkpoints with Python

Use the Python client from your submission example. Replace <JOB_ID> with your job ID. Repeat the status check until the job finishes; don't resubmit a queued job.

Retrieve the status, event logs, available checkpoints, and result-file IDs:

job = client.fine_tuning.jobs.retrieve("<JOB_ID>")
print("Status:", job.status)
print("Error:", job.error)
print("Model:", job.fine_tuned_model)
print("Result files:", job.result_files)

for event in client.fine_tuning.jobs.list_events(job.id).data:
    print(event.created_at, event.message)

checkpoints = client.fine_tuning.jobs.checkpoints.list(job.id)
print(checkpoints.model_dump_json(indent=2))

Reference: Fine-tuning API.

After completion, if the job returns a CSV metrics file, replace <RESULT_FILE_ID> with its ID and download it:

with open("results.csv", "wb") as result_file:
    result_file.write(client.files.content("<RESULT_FILE_ID>").read())

Reference: Files API.

Inspect logs and checkpoints with REST

Use the resource endpoint and key for REST. Replace <JOB_ID> with your job ID. Repeat the status check until the job finishes; don't resubmit a queued job.

Retrieve the job, event logs, and available checkpoints:

JOB_URL="$AZURE_OPENAI_ENDPOINT/openai/v1/fine_tuning/jobs/<JOB_ID>"
curl "$JOB_URL" -H "api-key: $AZURE_OPENAI_API_KEY"
curl "$JOB_URL/events" -H "api-key: $AZURE_OPENAI_API_KEY"
curl "$JOB_URL/checkpoints" -H "api-key: $AZURE_OPENAI_API_KEY"

Reference: Fine-tuning API.

After completion, inspect result_files in the job response. If it includes a CSV metrics file, replace <RESULT_FILE_ID> with its ID and download it:

curl "$AZURE_OPENAI_ENDPOINT/openai/v1/files/<RESULT_FILE_ID>/content" \
  -H "api-key: $AZURE_OPENAI_API_KEY" --output results.csv

Reference: Files API.

Use the Azure Developer CLI job workflow to inspect the job status.

Checkpoints appear as training progresses; a queued job might have none. Inspect the returned checkpoint identifiers and metrics rather than assuming the latest version performs best.

When training succeeds, retain the trained model or your chosen checkpoint. Compare candidates on held-out tasks before selecting one for deployment.

Deploy the model

Use the Foundry portal to deploy the model, regardless of how you submit the training job. Follow Deploy fine-tuned models for your model's supported serving option and inference procedure.

After deployment, use Run evaluations from the Foundry portal to compare candidates on held-out tasks.

Stop training and clean up

Cancel an unneeded job from its portal view. Delete uploaded files separately through the portal when you no longer need them.

Cancel an unneeded job with client.fine_tuning.jobs.cancel("<JOB_ID>"). Delete unused uploaded files separately with client.files.delete("<FILE_ID>").

Reference: Job cancellation and file deletion.

Cancel an unneeded job with the fine-tuning cancellation API. Delete unused uploaded files separately with the Files API.

Cancel an unneeded job with the CLI commands above. For available cleanup operations, see the fine-tuning CLI reference.

Delete unused deployments with the deployment cleanup procedure. Cancelling training doesn't delete deployments or stop their charges.

Retain any data and job records you need to reproduce the experiment.