Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Fine-tuning adapts a pretrained model to your task through additional training on task-specific data. It adjusts the model's weights rather than adding examples to each prompt. You build on the model's existing capabilities without training a model from scratch.
Fine-tuning can help when you have high-quality, task-specific training data and want to:
- Improve accuracy and relevance. Training on more examples than fit in a prompt can teach task-specific patterns.
- Reduce prompt overhead. Fewer prompt examples can lower token costs and latency.
- Align style and structure. Responses can follow a consistent tone, format, or schema.
- Improve tool use. Training examples can improve tool selection and argument accuracy.
- Use retrieved context effectively. Models can learn to prioritize relevant context and ignore irrelevant information.
- Specialize a smaller model. Teacher-generated examples can help a smaller model meet task requirements at lower cost and latency.
Benefits aren't guaranteed; held-out evaluation helps you compare the fine-tuned model with your baseline. Fine-tuning adds training and hosting costs and doesn't replace retrieval for current information or application-level safety controls.
Choose your training approach
Microsoft Foundry provides two training approaches. Managed fine-tuning runs a training job from your data and settings. Interactive training (preview) lets you control training steps in Python.
Both use Foundry-managed training infrastructure; you don't provision training GPUs.
Note
Interactive training is in preview and requires explicit access approval. Request access through the preview sign-up form and wait for approval before creating training sessions.
| Managed fine-tuning | Interactive training (preview) | |
|---|---|---|
| How it works | Submit data and settings; Foundry runs the training job. | Write a Python loop; Foundry executes training and sampling operations. |
| Best for | Beginners and AI engineers who prefer predefined workflows over lower-level training operations. | Experienced ML practitioners building custom training workflows. |
| Your control | Data, supported hyperparameters, and RFT graders. | Losses, rewards, rollouts, gradient accumulation, updates, and checkpoints. |
| Training methods | Model-specific SFT, DPO, and RFT; distillation through teacher-generated SFT data. | Recipes for SFT, reinforcement learning, preference learning, distillation, and custom losses. |
| Example use cases | Distill a larger model, learn from support conversations, or improve responses with a grader. | Collect tool-use rollouts, apply custom rewards, or change training-loop update and evaluation behavior. |
Tip
Start with managed fine-tuning if you're new to model customization. It provides predefined workflows without requiring you to implement individual training operations.
Model requirements can constrain this choice. A model supported for one approach, method, or modality isn't automatically supported for another. Cookbook recipe support also doesn't establish service availability. Check the supported models table below for available combinations.
Customization methods
Customization methods determine how a model learns from training data or feedback, such as desired responses, preferences, or rewards.
| Method | Training signal | When it fits |
|---|---|---|
| Supervised fine-tuning (SFT) | Prompts and desired responses. | You can demonstrate the behavior the model should learn. |
| Preference fine-tuning, using direct preference optimization (DPO) | Preferred and rejected responses to the same input. | Quality depends on preferences such as tone or completeness, rather than one correct answer. |
| Reinforcement fine-tuning (RFT) | Prompts and a grader that scores generated responses. | Multiple solutions are possible, and you can reliably score their quality. |
The predefined methods in this table apply only to managed fine-tuning. Interactive training exposes training primitives for building custom training workflows.
Training types
Training types define where training runs and how capacity, pricing, and service guarantees apply. They're separate from deployment types, which determine how you serve a trained model.
| Training type | Description |
|---|---|
| Standard | Training occurs in the current Foundry resource's region and provides guarantees for data residency. Ideal for workloads where data must remain in a specific region. |
| Data zone | Training occurs within a supported data zone rather than a single region. Ideal for workloads where data must remain within that zone. Only the US data zone is currently supported. |
| Global | Provides more affordable pricing compared to Standard by using capacity beyond your current region. Data and weights are copied to the region where training occurs. Ideal if data residency is not a restriction and you want faster queue times. |
| Developer | Provides significant cost savings by using idle capacity for training. There are no latency or SLA guarantees, so jobs in this tier might be automatically preempted and resumed later. There are no guarantees for data residency either. Ideal for experimentation and price-sensitive workloads. |
Deployment options
Foundry offers three deployment options with different serving infrastructure and billing models.
- Serverless: Foundry hosts the fine-tuned model and manages the serving infrastructure.
- Managed compute (preview): Dedicated GPU capacity runs a matching base model with one or more compatible LoRA adapters.
- Fireworks on Foundry (preview): The Fireworks runtime serves one adapter merged into a full-weight copy of the base model.
Serverless and Fireworks offer different deployment types, which determine the processing geography, capacity, and billing model. Availability depends on the model, deployment option, and artifact.
| Deployment type | Description |
|---|---|
| Standard | Runs inference in the deployment's region with per-token billing and applicable fine-tuned-model hosting charges. |
| Data Zone Standard | Runs inference within a supported data zone with per-token billing. Only the US data zone is supported. |
| Global Standard | Uses global capacity with per-token billing. Custom model weights might temporarily be stored outside the resource's geography. |
| Developer | Runs inference globally with per-token billing and no hourly hosting fee or data-residency guarantee. For evaluation only; lasts 24 hours and has no availability SLA. |
| Provisioned Throughput | Provides regional capacity billed in provisioned throughput units (PTUs) for predictable throughput. |
| Global Provisioned | Uses global capacity with PTU-based billing for imported custom models. |
Supported models
Your requirements might call for a specific model or training approach. Use the table to check which combinations support your customization method, training type, and deployment option.
For supported serverless deployments, Global Standard, Data Zone Standard, and Developer are available unless a legend states otherwise. Check Standard deployment regions and Provisioned Throughput deployment regions for those deployment types.
| Model ID | Approach | Training types | Deployment options |
|---|---|---|---|
MAI-Code-1.1-Flash(preview) |
Managed: ✅ (RFT) Interactive: ❌ |
Standard: ❌ Data zone: ❌ Global: ✅ Developer: ✅ |
Serverless: ✅2 Managed compute: ❌ Fireworks: ❌ |
Muse-Glimmer-30B(preview) |
Managed: ✅ (SFT, RFT) Interactive: ✅ |
Standard: ❌ Data zone: ❌ Global: ✅ Developer: ✅ |
Serverless: ❌ Managed compute: ✅ Fireworks: ✅1 |
Llama-3.3-70B-Instruct |
Managed: ✅ (SFT) Interactive: ❌ |
Standard: ❌ Data zone: ✅ Global: ✅ Developer: ❌ |
Serverless: ✅2 Managed compute: ❌ Fireworks: ❌ |
Qwen3-32B |
Managed: ✅ (SFT) Interactive: ❌ |
Standard: ❌ Data zone: ✅ Global: ✅ Developer: ❌ |
Serverless: ✅2 Managed compute: ❌ Fireworks: ❌ |
Qwen3.6-35B-A3B(preview) |
Managed: ✅ (SFT, RFT) Interactive: ✅ |
Standard: ❌ Data zone: ❌ Global: ✅ Developer: ✅ |
Serverless: ❌ Managed compute: ✅ Fireworks: ✅1 |
Qwen3.8-27B(preview) |
Managed: ✅ (SFT, RFT) Interactive: ✅ |
Standard: ❌ Data zone: ❌ Global: ✅ Developer: ✅ |
Serverless: ❌ Managed compute: ✅ Fireworks: ✅1 |
gpt-oss-20b |
Managed: ✅ (SFT) Interactive: ✅ |
Standard: ❌ Data zone: ✅ Global: ✅ Developer: ❌ |
Serverless: ✅2 Managed compute: ❌ Fireworks: ❌ |
gpt-oss-120b(preview) |
Managed: ✅ (SFT, RFT) Interactive: ✅ |
Standard: ❌ Data zone: ❌ Global: ✅ Developer: ✅ |
Serverless: ❌ Managed compute: ✅ Fireworks: ✅1 |
Ministral-3B(2411) |
Managed: ✅ (SFT) Interactive: ❌ |
Standard: ❌ Data zone: ✅ Global: ✅ Developer: ❌ |
Serverless: ✅2 Managed compute: ❌ Fireworks: ❌ |
gpt-4o-mini(2024-07-18) |
Managed: ✅ (SFT) Interactive: ❌ |
Standard: ✅ Data zone: ✅ Global: ✅ Developer: ✅ |
Serverless: ✅5 Managed compute: ❌ Fireworks: ❌ |
gpt-4o(2024-08-06) |
Managed: ✅ (SFT, DPO) Interactive: ❌ |
Standard: ✅ Data zone: ✅ Global: ✅ Developer: ✅ |
Serverless: ✅5 Managed compute: ❌ Fireworks: ❌ |
gpt-4.1(2025-04-14) |
Managed: ✅ (SFT, DPO) Interactive: ❌ |
Standard: ✅ Data zone: ✅ Global: ✅ Developer: ✅ |
Serverless: ✅ Managed compute: ❌ Fireworks: ❌ |
gpt-4.1-mini(2025-04-14) |
Managed: ✅ (SFT, DPO) Interactive: ❌ |
Standard: ✅ Data zone: ✅ Global: ✅ Developer: ✅ |
Serverless: ✅ Managed compute: ❌ Fireworks: ❌ |
gpt-4.1-nano(2025-04-14) |
Managed: ✅ (SFT, DPO) Interactive: ❌ |
Standard: ✅ Data zone: ✅ Global: ✅ Developer: ✅ |
Serverless: ✅ Managed compute: ❌ Fireworks: ❌ |
o4-mini(2025-04-16) |
Managed: ✅ (RFT) Interactive: ❌ |
Standard: ✅ Data zone: ✅ Global: ✅ Developer: ✅ |
Serverless: ✅ Managed compute: ❌ Fireworks: ❌ |
gpt-5(2025-08-07)3 |
Managed: ✅ (RFT) Interactive: ❌ |
Standard: ✅ Data zone: ✅ Global: ✅ Developer: ❌ |
Serverless: ✅4 Managed compute: ❌ Fireworks: ❌ |
MedImageInsight-Premium(preview) |
Managed: ✅ (SFT) Interactive: ❌ |
Global | Serverless: ✅2 Managed compute: ❌ Fireworks: ❌ |
CXRReportgen-Premium(preview) |
Managed: ✅ (SFT) Interactive: ❌ |
Global | Serverless: ✅2 Managed compute: ❌ Fireworks: ❌ |
Legend
- 1 Only Provisioned Throughput deployment type is supported.
- 2 Only Global Standard deployment type is supported.
- 3 GPT-5 reinforcement fine-tuning requires an invitation. Contact your Microsoft account team for enrollment.
- 4 Only Standard deployment type is supported. See Standard deployment regions.
- 5 Developer deployment type is not supported.
Supported global training regions
Note
Global training provides more affordable training per token, but doesn't offer data residency.
The table lists the Foundry project regions required to use Global managed fine-tuning and interactive training (preview). Choose a region supported for your model and training approach, including the restrictions below. Standard and Data zone training have separate regional support.
| Foundry project region | Managed fine-tuning | Interactive training (preview) |
|---|---|---|
| Australia East | ✅ | ✅ |
| Brazil South | ✅1 | ❌ |
| Canada Central | ✅1 | ❌ |
| Canada East | ✅1 | ✅ |
| East US | ✅ | ✅ |
| East US 2 | ✅ | ✅ |
| France Central | ✅1 | ❌ |
| Germany West Central | ✅1 | ❌ |
| Italy North | ✅1 | ❌ |
| Japan East | ✅1,2 | ❌ |
| Korea Central | ✅ | ✅ |
| North Central US | ✅ | ✅ |
| Norway East | ✅1 | ❌ |
| Poland Central | ✅1,3 | ❌ |
| South Africa North | ✅1 | ❌ |
| South Central US | ✅1 | ❌ |
| South India | ✅1 | ❌ |
| Southeast Asia | ✅1 | ❌ |
| Spain Central | ✅1 | ❌ |
| Sweden Central | ✅ | ✅ |
| Switzerland North | ✅ | ✅ |
| Switzerland West | ✅1 | ❌ |
| UAE North | ✅4 | ✅ |
| UK South | ✅ | ✅ |
| West Europe | ✅1 | ❌ |
| West US | ✅1 | ❌ |
| West US 3 | ✅ | ✅ |
- 1 Managed fine-tuning for
Qwen3.6-35B-A3B,Qwen3.8-27B,Muse-Glimmer-30B, andgpt-oss-120bisn't supported in this region. - 2 Vision fine-tuning isn't supported in Japan East.
- 3 Managed fine-tuning for
gpt-4.1-nanoisn't supported in Poland Central. - 4 UAE North supports Global managed fine-tuning only for
Qwen3.6-35B-A3B,Qwen3.8-27B,Muse-Glimmer-30B, andgpt-oss-120b.
Supported Standard deployment regions
Only the following models support serverless Standard deployments. Check regional availability in the table.
| Model ID | East US 2 | North Central US | Sweden Central |
|---|---|---|---|
gpt-4o-mini |
❌ | ✅ | ✅ |
gpt-4o |
✅ | ❌ | ✅ |
gpt-4.1 |
❌ | ✅ | ✅ |
gpt-4.1-mini |
❌ | ✅ | ✅ |
gpt-4.1-nano |
❌ | ✅ | ✅ |
o4-mini |
✅ | ❌ | ✅ |
gpt-5 |
❌ | ✅ | ✅ |
Supported Provisioned Throughput deployment regions
Only the following models support serverless Provisioned Throughput deployments, in the listed regions.
| Model ID | North Central US | Sweden Central |
|---|---|---|
gpt-4o-mini |
✅ | ✅ |
gpt-4o |
✅ | ✅ |
gpt-4.1 |
❌ | ✅ |
Challenges and limitations
Fine-tuning is an iterative engineering process, not a guaranteed improvement.
| Challenge | How to address it |
|---|---|
| Data quality and coverage | Use accurate, consistent, representative examples. Inspect missing scenarios, duplicates, bias, and sensitive information before uploading data. |
| Overfitting and evaluation leakage | Separate training, validation, and final test data. Compare checkpoints on held-out tasks, not just training loss. |
| Misleading reward signals | Test RFT graders against correct, incorrect, and adversarial responses. Rising reward doesn't establish better task performance if the model exploits the grader. |
| Experimentation and cost | Budget for repeated runs, grading, evaluation, hosting, and inference. Stop experiments that don't improve the target metric. |
| Model and serving compatibility | Confirm method, tier, version, and deployment eligibility before training. An adapter isn't interchangeable across similarly named base models. |
| Maintenance | Reassess quality as task data changes or a base model approaches retirement. Retraining and redeployment can be necessary. |
| Safety and governance | Review data handling and managed safety evaluation. Keep application-level safety controls and evaluate the final serving runtime. |
| Interactive runtime responsibilities | Maintain your driver, environments, tokenization, checkpoints, and recovery logic. In-session sampling isn't proof of production serving behavior. |
Next steps
Start with the guide for your training approach, then deploy and evaluate your fine-tuned model.
- Customize a model with supervised fine-tuning.
- Customize a model with direct preference optimization.
- Customize a model with reinforcement fine-tuning.
- Get started with interactive training.
- Customize a premium healthcare AI model.
- Deploy fine-tuned models.
- Run evaluations from the Foundry portal.
- Explore fine-tuning samples on GitHub.