Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Currently viewing:
Foundry (classic) portal version - Switch to version for the new Foundry portal
Prepare preference data
Start with the Orca preference-pair sample dataset on GitHub, or prepare your own data.
Save separate training.jsonl and validation.jsonl files, with one example per line. Each example contains an input and exactly two outputs:
| Field | Content |
|---|---|
input |
A messages array with the prompt. Optional tools and parallel_tool_calls also belong here. |
preferred_output |
The preferred response, including at least one assistant message. |
non_preferred_output |
The rejected response to the same input, including at least one assistant message. |
Output messages use the assistant or tool role. Both responses must answer the same prompt, and the preference must match your intended task.
{"input": {"messages": [{"role": "system", "content": "You are a chatbot assistant. Given a user question with multiple choice answers, provide the correct answer."}, {"role": "user", "content": "Question: Janette conducts an investigation to see which foods make her feel more fatigued. She eats one of four different foods each day at the same time for four days and then records how she feels. She asks her friend Carmen to do the same investigation to see if she gets similar results. Which would make the investigation most difficult to replicate? Answer choices: A: measuring the amount of fatigue, B: making sure the same foods are eaten, C: recording observations in the same chart, D: making sure the foods are at the same temperature"}]}, "preferred_output": [{"role": "assistant", "content": "A: measuring the amount of fatigue"}], "non_preferred_output": [{"role": "assistant", "content": "D: making sure the foods are at the same temperature"}]}
Reference: DPO dataset examples.
Keep validation and final test examples out of the training set. Check supported models before using a base model or an SFT-fine-tuned model for DPO.
How to use direct preference optimization fine-tuning
- Prepare
jsonldatasets in the preference format. - Select the model and then select the method of customization Direct Preference Optimization.
- Upload datasets – training and validation. Preview as needed.
- Select hyperparameters, defaults are recommended for initial experimentation.
- Review the selections and create a fine-tuning job.
Create a DPO job with code
Set AZURE_OPENAI_ENDPOINT and AZURE_OPENAI_API_KEY for your resource. These examples use gpt-4.1-mini-2025-04-14 and Global training; check model support and training types before running them.
Install openai. Upload the files and submit the job with default hyperparameters:
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["AZURE_OPENAI_ENDPOINT"].rstrip("/") + "/openai/v1/",
api_key=os.environ["AZURE_OPENAI_API_KEY"],
)
with open("training.jsonl", "rb") as training:
training_file = client.files.create(file=training, purpose="fine-tune")
with open("validation.jsonl", "rb") as validation:
validation_file = client.files.create(
file=validation, purpose="fine-tune"
)
client.files.wait_for_processing(training_file.id)
client.files.wait_for_processing(validation_file.id)
job = client.fine_tuning.jobs.create(
model="gpt-4.1-mini-2025-04-14",
training_file=training_file.id,
validation_file=validation_file.id,
method={
"type": "dpo",
"dpo": {
"hyperparameters": {
"n_epochs": "auto",
"batch_size": "auto",
"learning_rate_multiplier": "auto",
"beta": "auto",
}
},
},
extra_body={"trainingType": "GlobalStandard"},
)
print(job.id, job.status)
Reference: OpenAI Python client.
For DPO after SFT, set model to the supported fine-tuned model's identifier, not its deployment name.