Edit

Fine-tune a model with supervised fine-tuning in Microsoft Foundry

Prepare your data and create a supervised fine-tuning (SFT) job in Microsoft Foundry. For method comparisons and training concepts, see the fine-tuning overview.

Prerequisites

  • Your resource endpoint and credentials, and a Bash-compatible shell for the REST examples.

Important

The Foundry RBAC roles were recently renamed. Foundry User, Foundry Owner, Foundry Account Owner, and Foundry Project Manager were previously named Azure AI User, Azure AI Owner, Azure AI Account Owner, and Azure AI Project Manager. You might still see the previous names in some places while the rename rolls out. The role IDs and core permissions are unchanged by the rename.

Prepare your data

Start with the text-only GSM8K sample dataset on GitHub, or prepare your own data.

Save separate training.jsonl and validation.jsonl files with one conversation per line. Include the prompt and the desired assistant response in each example:

{"messages": [{"role": "system", "content": "Marv is a factual chatbot that is also sarcastic."}, {"role": "user", "content": "What's the biggest city in France?"}, {"role": "assistant", "content": "Paris, as if everyone doesn't know that already."}]}

Reference: Chat Completions message format.

Check the data-format, file-size, and minimum-example requirements for your selected model. Keep validation and final test examples out of the training set.

For specialized data formats, see vision fine-tuning or tool calling.

Create an SFT job in the portal

Use your prepared datasets to create the job:

  1. Sign in to the Foundry portal and select your project.
  2. Go to Build > Fine-tune, and select Fine-tune.
  3. Select a supported base model and Supervised fine-tuning.
  4. Select a training type supported by your model and resource.
  5. Select Existing dataset or Upload new dataset for the training and validation files. Inspect the preview and resolve validation errors.
  6. Configure the available hyperparameters, such as Batch size, Learning rate multiplier, and Number of epochs. Keep your model's defaults for your first job. See the hyperparameter descriptions below.

Configure hyperparameters

Hyperparameters control how the job updates the model during training. Available settings and accepted values depend on your selected model. Start with its defaults, and compare validation results before changing a setting.

Review these common settings in the job configuration:

Setting Description
Batch size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
Learning rate multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
Number of epochs The number of complete passes through the training dataset. More passes give the model more opportunities to learn, but can increase overfitting.

Set supported values under method.supervised.hyperparameters in the job body:

Hyperparameter Value Description
batch_size Integer or auto The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Number or auto Scales the training learning rate. A smaller multiplier can help reduce overfitting.
n_epochs Integer or auto The number of complete passes through the training dataset. More passes can improve learning, but can increase overfitting.

The v1 examples use auto for service-selected values. If your model doesn't accept auto, use values supported by that model.

For the request schema, see the v1 fine-tuning API.

Configure supported settings in your SFT YAML file. The CLI SFT sample uses these names:

Hyperparameter Description
batch_size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
epochs The number of complete passes through the training dataset. This CLI configuration uses epochs, rather than the v1 API's n_epochs.

Use the configuration sample for your selected model rather than copying settings from another model.

Optionally, set a suffix to identify the resulting model. Review the configuration and select Submit. Retain the job ID.

Create an SFT job with Python

Initialize the client with the Microsoft Foundry SDK for your project, or use the OpenAI SDK with a resource endpoint. Both clients use the same upload and job commands below.

Install azure-ai-projects, azure-identity, and openai. Sign in with a credential supported by DefaultAzureCredential, and set FOUNDRY_PROJECT_ENDPOINT to your project endpoint:

import os
from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential

project = AIProjectClient(
    endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"],
    credential=DefaultAzureCredential(),
)
client = project.get_openai_client()

Reference: AIProjectClient and DefaultAzureCredential.

Set FINE_TUNING_MODEL to an identifier from supported model IDs and versions that supports SFT with this API.

Set FINE_TUNING_TRAINING_TYPE to the API value for a supported training type. For example, Global training uses GlobalStandard.

Configure hyperparameters

Hyperparameters control how the job updates the model during training. Available settings and accepted values depend on your selected model. Start with its defaults, and compare validation results before changing a setting.

Review these common settings in the job configuration:

Setting Description
Batch size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
Learning rate multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
Number of epochs The number of complete passes through the training dataset. More passes give the model more opportunities to learn, but can increase overfitting.

Set supported values under method.supervised.hyperparameters in the job body:

Hyperparameter Value Description
batch_size Integer or auto The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Number or auto Scales the training learning rate. A smaller multiplier can help reduce overfitting.
n_epochs Integer or auto The number of complete passes through the training dataset. More passes can improve learning, but can increase overfitting.

The v1 examples use auto for service-selected values. If your model doesn't accept auto, use values supported by that model.

For the request schema, see the v1 fine-tuning API.

Configure supported settings in your SFT YAML file. The CLI SFT sample uses these names:

Hyperparameter Description
batch_size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
epochs The number of complete passes through the training dataset. This CLI configuration uses epochs, rather than the v1 API's n_epochs.

Use the configuration sample for your selected model rather than copying settings from another model.

Upload training.jsonl and validation.jsonl, then submit the job. The example includes hyperparameters in the job body and uses auto to request service-selected values:

import os

with open("training.jsonl", "rb") as training:
    training_file = client.files.create(file=training, purpose="fine-tune")
with open("validation.jsonl", "rb") as validation:
    validation_file = client.files.create(
        file=validation, purpose="fine-tune"
    )
client.files.wait_for_processing(training_file.id)
client.files.wait_for_processing(validation_file.id)

job = client.fine_tuning.jobs.create(
    model=os.environ["FINE_TUNING_MODEL"],
    training_file=training_file.id,
    validation_file=validation_file.id,
    suffix="my-model",
    method={
        "type": "supervised",
        "supervised": {
            "hyperparameters": {
                "batch_size": "auto",
                "learning_rate_multiplier": "auto",
                "n_epochs": "auto",
            }
        },
    },
    extra_body={
        "trainingType": os.environ["FINE_TUNING_TRAINING_TYPE"],
    },
)
print(job.id, job.status)

Reference: Files API and v1 fine-tuning API.

Retain the returned job ID and continue to monitor the job.

Create an SFT job with REST

Set AZURE_OPENAI_ENDPOINT and AZURE_OPENAI_API_KEY for your resource. Run these commands in a Bash-compatible shell.

Upload each dataset with purpose=fine-tune, and retain the distinct file IDs:

curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/files" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -F "purpose=fine-tune" -F "file=@training.jsonl"

curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/files" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -F "purpose=fine-tune" -F "file=@validation.jsonl"

Reference: Files API.

Replace <SUPPORTED_MODEL_ID> with an identifier from supported model IDs and versions that supports SFT with this API.

Replace <SUPPORTED_TRAINING_TYPE> with the API value for a supported training type. For example, Global training uses GlobalStandard.

Configure hyperparameters

Hyperparameters control how the job updates the model during training. Available settings and accepted values depend on your selected model. Start with its defaults, and compare validation results before changing a setting.

Review these common settings in the job configuration:

Setting Description
Batch size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
Learning rate multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
Number of epochs The number of complete passes through the training dataset. More passes give the model more opportunities to learn, but can increase overfitting.

Set supported values under method.supervised.hyperparameters in the job body:

Hyperparameter Value Description
batch_size Integer or auto The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Number or auto Scales the training learning rate. A smaller multiplier can help reduce overfitting.
n_epochs Integer or auto The number of complete passes through the training dataset. More passes can improve learning, but can increase overfitting.

The v1 examples use auto for service-selected values. If your model doesn't accept auto, use values supported by that model.

For the request schema, see the v1 fine-tuning API.

Configure supported settings in your SFT YAML file. The CLI SFT sample uses these names:

Hyperparameter Description
batch_size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
epochs The number of complete passes through the training dataset. This CLI configuration uses epochs, rather than the v1 API's n_epochs.

Use the configuration sample for your selected model rather than copying settings from another model.

Replace the file IDs and submit the job. The request includes hyperparameters in the job body and uses auto to request service-selected values:

curl -X POST "$AZURE_OPENAI_ENDPOINT/openai/v1/fine_tuning/jobs" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "<SUPPORTED_MODEL_ID>",
    "training_file": "<TRAINING_FILE_ID>",
    "validation_file": "<VALIDATION_FILE_ID>",
    "trainingType": "<SUPPORTED_TRAINING_TYPE>",
    "method": {
      "type": "supervised",
      "supervised": {
        "hyperparameters": {
          "batch_size": "auto",
          "learning_rate_multiplier": "auto",
          "n_epochs": "auto"
        }
      }
    },
    "suffix": "my-model"
  }'

Reference: v1 fine-tuning API.

Retain the returned job ID and continue to monitor the job.

Create an SFT job with the Azure Developer CLI

Install Azure Developer CLI version 1.22.1 or later, then install the fine-tuning extension and sign in:

azd ext install azure.ai.finetune
azd auth login

Reference: Azure Developer CLI.

Download an SFT configuration sample and its data files for your model. Set the model from supported model IDs and versions, and review the data paths and supported training type.

Configure hyperparameters

Hyperparameters control how the job updates the model during training. Available settings and accepted values depend on your selected model. Start with its defaults, and compare validation results before changing a setting.

Review these common settings in the job configuration:

Setting Description
Batch size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
Learning rate multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
Number of epochs The number of complete passes through the training dataset. More passes give the model more opportunities to learn, but can increase overfitting.

Set supported values under method.supervised.hyperparameters in the job body:

Hyperparameter Value Description
batch_size Integer or auto The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Number or auto Scales the training learning rate. A smaller multiplier can help reduce overfitting.
n_epochs Integer or auto The number of complete passes through the training dataset. More passes can improve learning, but can increase overfitting.

The v1 examples use auto for service-selected values. If your model doesn't accept auto, use values supported by that model.

For the request schema, see the v1 fine-tuning API.

Configure supported settings in your SFT YAML file. The CLI SFT sample uses these names:

Hyperparameter Description
batch_size The number of training examples processed together in one forward and backward pass. Larger batches produce less frequent updates with lower variance.
learning_rate_multiplier Scales the training learning rate. A smaller multiplier can help reduce overfitting.
epochs The number of complete passes through the training dataset. This CLI configuration uses epochs, rather than the v1 API's n_epochs.

Use the configuration sample for your selected model rather than copying settings from another model.

The following excerpt shows how to include hyperparameters in the job configuration. It uses values from the official CLI sample, not defaults for every model. Keep the rest of your model's configuration and adjust these values to its supported settings.

model: <SUPPORTED_MODEL_ID>
method:
  type: supervised
  supervised:
    hyperparameters:
      epochs: 4
      batch_size: 8
      learning_rate_multiplier: 0.1

Reference: SFT CLI configuration.

Replace <SUPPORTED_MODEL_ID> with your selected model ID. From the configuration directory, initialize the project, submit the complete job configuration, and inspect its status:

azd ai finetuning init -e <project-endpoint>
azd ai finetuning jobs submit -f <path-to-job-yaml>
azd ai finetuning jobs show -i <job-id>

Use a project endpoint in the form https://<account>.services.ai.azure.com/api/projects/<project>. Replace <job-id> with the ID returned by submission.

Reference: Fine-tuning CLI samples and commands.

Pause, resume, or cancel

Run lifecycle commands only when your model, method, and job state support them. Pause and resume aren't available for every model or method.

azd ai finetuning jobs pause -i <job-id>
azd ai finetuning jobs resume -i <job-id>
azd ai finetuning jobs cancel -i <job-id>

Reference: Fine-tuning CLI commands.

Review training metrics

SFT measures how well the model predicts the target responses in your examples. Compare training and validation results rather than judging the model on training performance alone.

Metric availability depends on your model. Use the metrics reported for your job:

Metrics What to look for
Training loss. Loss on the current training batch. Look for a decreasing trend.
Full validation loss. Loss across the validation set at the end of an epoch. If training loss falls but validation loss rises, inspect failing validation examples.
Training mean token accuracy. The fraction of target tokens correctly predicted in the training batch. Look for an increasing trend.
Full validation mean token accuracy. Token accuracy across the validation set at the end of an epoch. Confirm that improvements also help your task.

The corresponding metric names are train_loss, full_valid_loss, train_mean_token_accuracy, and full_valid_mean_token_accuracy. Batch validation metrics, when reported, aren't the same as full-validation metrics.

Compare the available checkpoints using validation metrics and held-out task results. Checkpoint creation and retention depend on the model and training workflow.

For explanations of overfitting and evaluation risks, see challenges and limitations.

Monitor the job and select a checkpoint

Inspect training progress before choosing a model or checkpoint to deploy.

In the portal, open the job details:

  1. Check the Status and event logs. Jobs can queue before training starts; inspect error details if a job fails.
  2. Open Monitor to compare training and validation metrics using the guidance above.
  3. Open Checkpoints to inspect available model versions and their metrics. Compare candidates on held-out tasks before choosing one to deploy.

Inspect logs and checkpoints with Python

Use the Python client from your submission example. Replace <JOB_ID> with your job ID. Repeat the status check until the job finishes; don't resubmit a queued job.

Retrieve the status, event logs, available checkpoints, and result-file IDs:

job = client.fine_tuning.jobs.retrieve("<JOB_ID>")
print("Status:", job.status)
print("Error:", job.error)
print("Model:", job.fine_tuned_model)
print("Result files:", job.result_files)

for event in client.fine_tuning.jobs.list_events(job.id).data:
    print(event.created_at, event.message)

checkpoints = client.fine_tuning.jobs.checkpoints.list(job.id)
print(checkpoints.model_dump_json(indent=2))

Reference: Fine-tuning API.

After completion, if the job returns a CSV metrics file, replace <RESULT_FILE_ID> with its ID and download it:

with open("results.csv", "wb") as result_file:
    result_file.write(client.files.content("<RESULT_FILE_ID>").read())

Reference: Files API.

Inspect logs and checkpoints with REST

Use the resource endpoint and key for REST. Replace <JOB_ID> with your job ID. Repeat the status check until the job finishes; don't resubmit a queued job.

Retrieve the job, event logs, and available checkpoints:

JOB_URL="$AZURE_OPENAI_ENDPOINT/openai/v1/fine_tuning/jobs/<JOB_ID>"
curl "$JOB_URL" -H "api-key: $AZURE_OPENAI_API_KEY"
curl "$JOB_URL/events" -H "api-key: $AZURE_OPENAI_API_KEY"
curl "$JOB_URL/checkpoints" -H "api-key: $AZURE_OPENAI_API_KEY"

Reference: Fine-tuning API.

After completion, inspect result_files in the job response. If it includes a CSV metrics file, replace <RESULT_FILE_ID> with its ID and download it:

curl "$AZURE_OPENAI_ENDPOINT/openai/v1/files/<RESULT_FILE_ID>/content" \
  -H "api-key: $AZURE_OPENAI_API_KEY" --output results.csv

Reference: Files API.

Use the Azure Developer CLI job workflow to inspect the job status.

Checkpoints appear as training progresses; a queued job might have none. Inspect the returned checkpoint identifiers and metrics rather than assuming the latest version performs best.

When training succeeds, retain the trained model or your chosen checkpoint. Compare candidates on held-out tasks before selecting one for deployment.

Deploy the model

Use the Foundry portal to deploy the model, regardless of how you submit the training job. Follow Deploy fine-tuned models for your model's supported serving option and inference procedure.

After deployment, use Run evaluations from the Foundry portal to compare candidates on held-out tasks.

Stop training and clean up

Cancel an unneeded job from its portal view. Delete uploaded files separately through the portal when you no longer need them.

Cancel an unneeded job with client.fine_tuning.jobs.cancel("<JOB_ID>"). Delete unused uploaded files separately with client.files.delete("<FILE_ID>").

Reference: Job cancellation and file deletion.

Cancel an unneeded job with the fine-tuning cancellation API. Delete unused uploaded files separately with the Files API.

Cancel an unneeded job with the CLI commands above. For available cleanup operations, see the fine-tuning CLI reference.

Delete unused deployments with the deployment cleanup procedure. Cancelling training doesn't delete deployments or stop their charges.

Retain any data and job records you need to reproduce the experiment.