Edit

Azure OpenAI reasoning models (classic)

Currently viewing: Foundry (classic) portal version - Switch to version for the new Foundry portal

Azure OpenAI reasoning models are designed to tackle reasoning and problem-solving tasks with increased focus and capability. These models spend more time processing and understanding the user's request, making them exceptionally strong in areas like science, coding, and math compared to previous iterations.

Key capabilities of reasoning models:

  • Complex Code Generation: Capable of generating algorithms and handling advanced coding tasks to support developers.
  • Advanced Problem Solving: Ideal for comprehensive brainstorming sessions and addressing multifaceted challenges.
  • Complex Document Comparison: Perfect for analyzing contracts, case files, or legal documents to identify subtle differences.
  • Instruction Following and Workflow Management: Particularly effective for managing workflows requiring shorter contexts.

Prerequisites

  • An Azure OpenAI reasoning model deployed.

  • If you use the REST examples:

    • Install the Azure CLI. For more information, see Install the Azure CLI.

    • Sign in with az login, then generate a bearer token and store it in the AZURE_OPENAI_AUTH_TOKEN environment variable.

      az account get-access-token --resource https://cognitiveservices.azure.com --query accessToken -o tsv
      

Usage

These models don't currently support the same set of parameters as other models that use the chat completions API.

Chat completions API

using Azure.Identity;
using OpenAI;
using OpenAI.Chat;
using System.ClientModel.Primitives;

#pragma warning disable OPENAI001 //currently required for token based authentication

BearerTokenPolicy tokenPolicy = new(
    new DefaultAzureCredential(),
    "https://ai.azure.com/.default");

ChatClient client = new(
    model: "o4-mini",
    authenticationPolicy: tokenPolicy,
    options: new OpenAIClientOptions()
    {

        Endpoint = new Uri("https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1")
    }
);

ChatCompletionOptions options = new ChatCompletionOptions
{
    MaxOutputTokenCount = 100000
};

ChatCompletion completion = client.CompleteChat(
         new DeveloperChatMessage("You are a helpful assistant"),
         new UserChatMessage("Tell me about the bitter lesson")
    );

Console.WriteLine($"[ASSISTANT]: {completion.Content[0].Text}");

Reasoning effort

Note

Reasoning models have reasoning_tokens as part of completion_tokens_details in the model response. These are hidden tokens that aren't returned as part of the message response content but are used by the model to help generate a final answer to your request. reasoning_effort can be set to low, medium, or high for all reasoning models except o1-mini. The higher the effort setting, the longer the model will spend processing the request, which will generally result in a larger number of reasoning_tokens.

Developer messages

Developer messages ("role": "developer") are functionally the same as system messages.

Adding a developer message to the previous code example would look as follows:


using Azure.Identity;
using OpenAI;
using OpenAI.Chat;
using System.ClientModel.Primitives;

#pragma warning disable OPENAI001 //currently required for token based authentication

BearerTokenPolicy tokenPolicy = new(
    new DefaultAzureCredential(),
    "https://ai.azure.com/.default");

ChatClient client = new(
    model: "o4-mini",
    authenticationPolicy: tokenPolicy,
    options: new OpenAIClientOptions()
    {

        Endpoint = new Uri("https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1")
    }
);

ChatCompletionOptions options = new ChatCompletionOptions
{
    ReasoningEffortLevel = ChatReasoningEffortLevel.Low,
    MaxOutputTokenCount = 100000
};

ChatCompletion completion = client.CompleteChat(
         new DeveloperChatMessage("You are a helpful assistant"),
         new UserChatMessage("Tell me about the bitter lesson")
    );

Console.WriteLine($"[ASSISTANT]: {completion.Content[0].Text}");

Reasoning summary

When using the latest reasoning models with the Responses API you can use the reasoning summary parameter to receive summaries of the model's chain of thought reasoning.

Important

Attempting to extract raw reasoning through methods other than the reasoning summary parameter are not supported, may violate the Acceptable Use Policy, and may result in throttling or suspension when detected.

using OpenAI;
using OpenAI.Responses;
using System.ClientModel.Primitives;
using Azure.Identity;

#pragma warning disable OPENAI001 //currently required for token based authentication

BearerTokenPolicy tokenPolicy = new(
    new DefaultAzureCredential(),
    "https://ai.azure.com/.default");

OpenAIResponseClient client = new(
    model: "o4-mini",
    authenticationPolicy: tokenPolicy,
    options: new OpenAIClientOptions()
    {
        Endpoint = new Uri("https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1")
    }
);

OpenAIResponse response = await client.CreateResponseAsync(
    userInputText: "What's the optimal strategy to win at poker?",
    new ResponseCreationOptions()
    {
        ReasoningOptions = new ResponseReasoningOptions()
        {
            ReasoningEffortLevel = ResponseReasoningEffortLevel.High,
            ReasoningSummaryVerbosity = ResponseReasoningSummaryVerbosity.Auto,
        },
    });

// Get the reasoning summary from the first OutputItem (ReasoningResponseItem)
Console.WriteLine("=== Reasoning Summary ===");
foreach (var item in response.OutputItems)
{
    if (item is ReasoningResponseItem reasoningItem)
    {
        foreach (var summaryPart in reasoningItem.SummaryParts)
        {
            if (summaryPart is ReasoningSummaryTextPart textPart)
            {
                Console.WriteLine(textPart.Text);
            }
        }
    }
}

Console.WriteLine("\n=== Assistant Response ===");
// Get the assistant's output
Console.WriteLine(response.GetOutputText());

Note

Even when enabled, reasoning summaries are not guaranteed to be generated for every step/request. This is expected behavior.

Python lark

GPT-5 series reasoning models have the ability to call a new custom_tool called lark_tool. This tool is based on Python lark and can be used for more flexible constraining of model output.

Responses API

{
  "model": "gpt-5-2025-08-07",
  "input": "please calculate the area of a circle with radius equal to the number of 'r's in strawberry",
  "tools": [
    {
      "type": "custom",
      "name": "lark_tool",
      "format": {
        "type": "grammar",
        "syntax": "lark",
        "definition": "start: QUESTION NEWLINE ANSWER\nQUESTION: /[^\\n?]{1,200}\\?/\nNEWLINE: /\\n/\nANSWER: /[^\\n!]{1,200}!/"
      }
    }
  ],
  "tool_choice": "required"
}

Microsoft Entra ID:

from openai import OpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider

token_provider = get_bearer_token_provider(
    DefaultAzureCredential(), "https://ai.azure.com/.default"
)

client = OpenAI(  
  base_url = "https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/",  
  api_key=token_provider,
)

response = client.responses.create(  
    model="gpt-5",  # replace with your model deployment name  
    tools=[  
        {  
            "type": "custom",
            "name": "lark_tool",
            "format": {
                "type": "grammar",
                "syntax": "lark",
                "definition": "start: QUESTION NEWLINE ANSWER\nQUESTION: /[^\\n?]{1,200}\\?/\nNEWLINE: /\\n/\nANSWER: /[^\\n!]{1,200}!/"
            }
        }  
    ],  
    input=[{"role": "user", "content": "Please calculate the area of a circle with radius equal to the number of 'r's in strawberry"}],  
)  

print(response.model_dump_json(indent=2))  

API Key:

import os
from openai import OpenAI

client = OpenAI(  
  base_url = "https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/",
  api_key=os.getenv("AZURE_OPENAI_API_KEY")  
)

response = client.responses.create(  
    model="gpt-5",  # replace with your model deployment name  
    tools=[  
        {  
            "type": "custom",
            "name": "lark_tool",
            "format": {
                "type": "grammar",
                "syntax": "lark",
                "definition": "start: QUESTION NEWLINE ANSWER\nQUESTION: /[^\\n?]{1,200}\\?/\nNEWLINE: /\\n/\nANSWER: /[^\\n!]{1,200}!/"
            }
        }  
    ],  
    input=[{"role": "user", "content": "Please calculate the area of a circle with radius equal to the number of 'r's in strawberry"}],  
)  

print(response.model_dump_json(indent=2))  
  

Output:

{
  "id": "resp_689a0cf927408190b8875915747667ad01c936c6ffb9d0d3",
  "created_at": 1754926332.0,
  "error": null,
  "incomplete_details": null,
  "instructions": null,
  "metadata": {},
  "model": "gpt-5",
  "object": "response",
  "output": [
    {
      "id": "rs_689a0cfd1c888190a2a67057f471b5cc01c936c6ffb9d0d3",
      "summary": [],
      "type": "reasoning",
      "encrypted_content": null,
      "status": null
    },
    {
      "id": "msg_689a0d00e60c81908964e5e9b2d6eeb501c936c6ffb9d0d3",
      "content": [
        {
          "annotations": [],
          "text": ""strawberry" has 3 r's, so the radius is 3.\nArea = πr<sup>2</sup> = π × 3<sup>2</sup> = 9π ≈ 28.27 square units.",
          "type": "output_text",
          "logprobs": null
        }
      ],
      "role": "assistant",
      "status": "completed",
      "type": "message"
    }
  ],
  "parallel_tool_calls": true,
  "temperature": 1.0,
  "tool_choice": "auto",
  "tools": [
    {
      "name": "lark_tool",
      "parameters": null,
      "strict": null,
      "type": "custom",
      "description": null,
      "format": {
        "type": "grammar",
        "definition": "start: QUESTION NEWLINE ANSWER\nQUESTION: /[^\\n?]{1,200}\\?/\nNEWLINE: /\\n/\nANSWER: /[^\\n!]{1,200}!/",
        "syntax": "lark"
      }
    }
  ],
  "top_p": 1.0,
  "background": false,
  "max_output_tokens": null,
  "max_tool_calls": null,
  "previous_response_id": null,
  "prompt": null,
  "prompt_cache_key": null,
  "reasoning": {
    "effort": "medium",
    "generate_summary": null,
    "summary": null
  },
  "safety_identifier": null,
  "service_tier": "default",
  "status": "completed",
  "text": {
    "format": {
      "type": "text"
    }
  },
  "top_logprobs": null,
  "truncation": "disabled",
  "usage": {
    "input_tokens": 139,
    "input_tokens_details": {
      "cached_tokens": 0
    },
    "output_tokens": 240,
    "output_tokens_details": {
      "reasoning_tokens": 192
    },
    "total_tokens": 379
  },
  "user": null,
  "content_filters": null,
  "store": true
}

Chat Completions

{
  "messages": [
    {
      "role": "user",
      "content": "Which one is larger, 42 or 0?"
    }
  ],
  "tools": [
    {
      "type": "custom",
      "name": "custom_tool",
      "custom": {
        "name": "lark_tool",
        "format": {
          "type": "grammar",
          "grammar": {
            "syntax": "lark",
            "definition": "start: QUESTION NEWLINE ANSWER\nQUESTION: /[^\\n?]{1,200}\\?/\nNEWLINE: /\\n/\nANSWER: /[^\\n!]{1,200}!/"
          }
        }
      }
    }
  ],
  "tool_choice": "required",
  "model": "gpt-5-2025-08-07"
}

Availability

Region availability

Model Region Limited access
gpt-6.1-sol Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-6-astra Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-6-sol Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-6-luna Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-5.6-sol Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-5.6-terra Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-5.6-luna Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-chat-latest Model availability No access request needed.
gpt-5.5 Model availability No access request needed. Quota request required depending on quota tier. Tier 5 and Tier 6 subscriptions have quota by default.
gpt-5.4-mini Global Standard:
East US2
Sweden Central
South Central US
Poland Central
No access request needed.
gpt-5.4-nano Global Standard:
East US2
Sweden Central
South Central US
Poland Central

Datazone Standard:
East US2
South Central US
No access request needed.
gpt-5.4-pro Model availability Request access: Limited access model application. If you already have access to a limited access model no request is required.
gpt-5.4 Model availability Request access: Limited access model application. If you already have access to a limited access model no request is required.
gpt-5.3-codex Model availability Request access: Limited access model application. If you already have access to a limited access model no request is required.
gpt-5.2-codex Model availability Request access: Limited access model application. If you already have access to a limited access model no request is required.
gpt-5.2 Model availability Request access: Limited access model application. If you already have access to a limited access model no request is required.
gpt-5.1-codex-max Model availability Access is no longer restricted for this model.
gpt-5.1 Model availability Access is no longer restricted for this model.
gpt-5.1-chat Model availability No access request needed.
gpt-5.1-codex Model availability Access is no longer restricted for this model.
gpt-5.1-codex-mini Model availability No access request needed.
gpt-5-pro Model availability Access is no longer restricted for this model.
gpt-5-codex Model availability Access is no longer restricted for this model.
gpt-5 Model availability Access is no longer restricted for this model.
gpt-5-mini Model availability No access request needed.
gpt-5-nano Model availability No access request needed.
o3-pro Model availability Request access: Limited access model application. If you already have access to a limited access model no request is required.
codex-mini Model availability No access request needed.
o4-mini Model availability No access request needed to use the core capabilities of this model.

Request access: o4-mini reasoning summary feature
o3 Model availability Request access: Limited access model application
o3-mini Model availability. Access is no longer restricted for this model.
o1 Model availability. Access is no longer restricted for this model.

API and feature support

Feature gpt-6-astra, 2026-09-03 gpt-6-sol, 2026-09-22 gpt-6-luna, 2026-09-22 gpt-6.1-sol, 2026-09-29
Structured outputs ✅ ✅ ✅ ✅
Context window 1,050,000 tokens 1,050,000 tokens 1,050,000 tokens 1,050,000 tokens
Maximum input tokens 922,000 tokens 922,000 tokens 922,000 tokens 922,000 tokens
Maximum output tokens 128,000 tokens 128,000 tokens 128,000 tokens 128,000 tokens
Input modalities Text and images Text and images Text and images Text and images
Output modalities Text Text Text Text
Chat Completions API ✅ ✅ ✅ ✅
Responses API ✅ ✅ ✅ ✅
Streaming ✅ ✅ ✅ ✅
Functions/tools ✅ ✅ ✅ ✅
Reasoning effort (including none) ✅ ✅ ✅ ✅
Verbosity ✅ ✅ ✅ ✅
logprobs ✅ ✅ ✅ ✅
temperature ✅ ✅ ✅ ✅
top_p ✅ ✅ ✅ ✅

Note

  • To avoid timeouts background mode is recommended for o3-pro.
  • o3-pro does not currently support image generation.

Unsupported parameters

Reasoning models other than GPT-6 Astra don't support the following parameters:

  • temperature, top_p, presence_penalty, frequency_penalty, logprobs, top_logprobs, logit_bias, max_tokens

Markdown output

By default the o3-mini and o1 models will not attempt to produce output that includes markdown formatting. A common use case where this behavior is undesirable is when you want the model to output code contained within a markdown code block. When the model generates output without markdown formatting you lose features like syntax highlighting, and copyable code blocks in interactive playground experiences. To override this new default behavior and encourage markdown inclusion in model responses, add the string Formatting re-enabled to the beginning of your developer message.

Adding Formatting re-enabled to the beginning of your developer message does not guarantee that the model will include markdown formatting in its response, it only increases the likelihood. We have found from internal testing that Formatting re-enabled is less effective by itself with the o1 model than with o3-mini.

To improve the performance of Formatting re-enabled you can further augment the beginning of the developer message which will often result in the desired output. Rather than just adding Formatting re-enabled to the beginning of your developer message, you can experiment with adding a more descriptive initial instruction like one of the examples below:

  • Formatting re-enabled - please enclose code blocks with appropriate markdown tags.
  • Formatting re-enabled - code output should be wrapped in markdown.

Depending on your expected output you may need to customize your initial developer message further to target your specific use case.