Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Microsoft-Decision-1 is available in Microsoft Foundry for classification,
routing, ranking, grading, and binary decisions. Use it to triage support
tickets, prioritize incidents, filter content, evaluate model output, or
select a model, tool, or agent for a request.
Unlike a generative large language model (LLM), Microsoft-Decision-1 doesn't
generate a free-form response or written rationale. It makes a focused
judgment about text or JSON and returns a typed, numerical decision. This
approach is useful when your application needs a predictable response shape
and probabilities that it can threshold or rank.
In this article, you deploy Microsoft-Decision-1 in Microsoft Foundry, call
the decision API, and build a support-ticket classifier.
Prerequisites
An Azure subscription with a valid payment method. If you don't have an Azure subscription, create a paid Azure account.
Access to Microsoft Foundry.
Permission to create and manage model deployments. The Cognitive Services Contributor role lets you deploy models. For more information, see Azure RBAC roles.
An authentication method: Microsoft Entra ID (recommended) or an API key.
Python 3.10 or later to run the quickstart.
The Azure Identity library:
pip install azure-identity
Choose a question type
Microsoft-Decision-1 accepts a state value and one or more typed questions.
The state can be text or JSON. Choose the question type based on the decision
your application needs. The request pattern is similar to a system-one API:
your application supplies the state and bounded questions, and the model
returns decisions instead of generated prose.
| Question | Type | Result |
|---|---|---|
| Is this true? | noul |
A probability from 0 through 1. |
| Which option is it? | choice |
One selected option and the probability for each option. |
| How much? | score |
A value on an ordered scale and the probability for each level. |
The response answers object uses the question names from your request. The
response also includes the deployed model name and token usage.
Use Noul for gates and filters
Use a noul question for a yes-or-no condition, such as whether a tool call
touches production data, a document contains a prompt injection, or a support
reply promises a refund.
{
"state": "I refunded your $200. You don't need to contact billing.",
"questions": {
"promises_refund": {
"type": "noul",
"instructions": "Does this reply promise the customer a refund?",
"criteria": {
"true": "Commits to returning money",
"false": "Makes no commitment about money"
}
}
}
}
The criteria field is optional. Define the true and false conditions to make
borderline questions more precise. Choose a probability threshold based on
your workload, and validate it against labeled examples before you automate a
gate.
Use Choice for classification and routing
Use a choice question to select exactly one label from a fixed set. Common
uses include routing tickets, classifying intent or document type, selecting a
model or tool, and sorting alerts by cause.
{
"state": "The API returns 500 on every call.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Charges, invoices, and refunds",
"engineering": "Bugs, errors, and outages",
"support": "How-to and account questions"
}
}
}
}
Describe each option clearly. The response contains the selected option and a probability for every option. Use a low top probability as a signal to escalate the decision to a person or another model.
Use Score for ranking and grading
Use a score question to place an item on an ordered scale. Common uses
include prioritizing a queue, grading model output, and measuring qualities
such as severity, helpfulness, or frustration.
{
"state": "Checkout is down for all EU customers since 09:00.",
"questions": {
"severity": {
"type": "score",
"instructions": "How severe is this incident?",
"criteria": [
"Cosmetic",
"Minor",
"Major",
"Critical"
]
}
}
}
List the levels from lowest to highest. The returned score is the
probability-weighted average of the level indexes. For example, a four-level
scale runs from 0 through 3, and the result can fall between levels.
Validate scores on your own labeled examples. Prefer scores for relative
ordering and thresholds rather than as absolute, calibrated ratings.
Microsoft-Decision-1 at a glance
| Model name | Model version | Deployment type | API type |
|---|---|---|---|
Microsoft-Decision-1 |
1 | DataZoneStandard in selected regions; GlobalStandard |
Decision |
To minimize network latency between your application and Foundry, create the
Foundry resource in a region close to your application and users, and call
that resource's endpoint. For locality-sensitive workloads, use
DataZoneStandard in the data zone closest to your application when it's
available. This deployment type keeps inference processing within that data
zone and can reduce long-distance routing. Benchmark your workload because a
deployment type doesn't guarantee lower latency.
GlobalStandard is also supported and provides access to global capacity.
Because inference can be processed in any supported Azure region, latency
might be higher or more variable. For more information, see
Deployment types for Microsoft Foundry Models
and region availability for Foundry Models sold by Azure.
Deploy Microsoft-Decision-1
Deploy the model from the Foundry model catalog.
- Open the Foundry portal, and go to your project.
- Select Model catalog.
- Search for and select Microsoft-Decision-1.
- Select Deploy, configure the deployment, and then select Deploy.
- On the deployment details page, copy the resource endpoint, deployment name, and API key.
For more information about the portal workflow, see Deploy Microsoft Foundry Models.
Alternatively, deploy the model by using the Azure CLI. Replace
<ACCOUNT_NAME>, <RESOURCE_GROUP>, and <DEPLOYMENT_NAME> with your values.
az cognitiveservices account deployment create \
--name <ACCOUNT_NAME> \
--resource-group <RESOURCE_GROUP> \
--deployment-name <DEPLOYMENT_NAME> \
--model-name "Microsoft-Decision-1" \
--model-format Microsoft \
--model-version "1" \
--sku-name GlobalStandard \
--sku-capacity 1
Reference: az cognitiveservices account deployment create
The example uses GlobalStandard. To keep inference processing within the
data zone closest to your application, change --sku-name to
DataZoneStandard when that deployment type is available in your region.
To list all available deployments on your resource:
az cognitiveservices account deployment list \
--resource-group <RESOURCE_GROUP> \
--name <ACCOUNT_NAME> \
--output table
Reference: az cognitiveservices account deployment list
Call the decision API
Send requests to the Microsoft provider's decision endpoint on your Foundry resource:
<your-foundry-resource-endpoint>/providers/microsoft/v1/systemone
The request's model value is your deployment name, not the underlying model
name. For example, a deployment named pi-decision-1 uses
"model": "pi-decision-1". The response's model field identifies the
underlying model, so it can contain a different value, such as
microsoft-decision-1.
Set the values that you copied from the deployment details page:
export AZURE_ENDPOINT="<your-foundry-resource-endpoint>"
export DEPLOYMENT_NAME="<your-deployment-name>"
For Microsoft Entra ID authentication, sign in with the Azure CLI, and get an access token:
az login
export AZURE_ENTRA_TOKEN=$(az account get-access-token \
--resource https://cognitiveservices.azure.com \
--query accessToken \
--output tsv)
Reference: az account get-access-token
Call the endpoint with Microsoft Entra ID:
curl "$AZURE_ENDPOINT/providers/microsoft/v1/systemone" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AZURE_ENTRA_TOKEN" \
-d '{
"model": "'"$DEPLOYMENT_NAME"'",
"state": "The API returns 500 on every call.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Charges, invoices, and refunds",
"engineering": "Bugs, errors, and outages",
"support": "How-to and account questions"
}
}
}
}'
To use an API key instead, replace the Authorization header with:
export AZURE_API_KEY="<your-api-key>"
-H "api-key: $AZURE_API_KEY"
For more information about authentication, see Configure Microsoft Entra ID authentication.
Classify support tickets
This quickstart uses a choice question to route customer support requests to
the billing, technical, or account team. It also measures accuracy and median
request latency across a small set of labeled examples.
Create a file named
support_ticket_classifier.py.Add the following code:
import json import os import urllib.error import urllib.request from collections.abc import Mapping from statistics import median from time import perf_counter from typing import Any from azure.identity import DefaultAzureCredential OPTIONS = { "billing": "Charges, invoices, refunds, or subscription payments", "technical": "Software errors, bugs, or integration failures", "account": "Sign-in, password, or account-access problems", } EXAMPLES = [ ("I was charged twice.", "billing"), ("The integration crashes during checkout.", "technical"), ("I cannot sign in after resetting my password.", "account"), ] AZURE_ENDPOINT = os.environ["AZURE_ENDPOINT"].rstrip("/") DEPLOYMENT_NAME = os.environ["DEPLOYMENT_NAME"] TIMEOUT_SECONDS = 60 TOKEN_SCOPE = "https://cognitiveservices.azure.com/.default" CREDENTIAL = DefaultAzureCredential() class DecisionAPIError(RuntimeError): """Raised when the API can't return a valid classification.""" def read_choice( payload: Any, options: Mapping[str, str], ) -> str: if not isinstance(payload, dict): raise DecisionAPIError("The API returned an invalid response.") answers = payload.get("answers") if not isinstance(answers, dict): raise DecisionAPIError("The API returned no answers object.") answer = answers.get("team") if not isinstance(answer, dict) or answer.get("type") != "choice": raise DecisionAPIError("The API returned an invalid answer.") choice = answer.get("choice") if not isinstance(choice, str) or choice not in options: raise DecisionAPIError( f"The API selected an unsupported team: {choice!r}." ) return choice def predict( text: str, options: Mapping[str, str], ) -> str: if not text.strip(): raise ValueError("text must not be empty.") if len(options) < 2: raise ValueError("options must contain at least two choices.") body = json.dumps( { "model": DEPLOYMENT_NAME, "state": text, "questions": { "team": { "type": "choice", "instructions": ( "Which team should handle this customer support " "request? Select exactly one team based on the " "primary problem." ), "criteria": dict(options), } }, } ).encode("utf-8") access_token = CREDENTIAL.get_token(TOKEN_SCOPE).token request = urllib.request.Request( f"{AZURE_ENDPOINT}/providers/microsoft/v1/systemone", data=body, headers={ "Authorization": f"Bearer {access_token}", "Content-Type": "application/json", "Accept": "application/json", }, method="POST", ) try: with urllib.request.urlopen( request, timeout=TIMEOUT_SECONDS, ) as response: response_body = response.read() except urllib.error.HTTPError as exc: raise DecisionAPIError( f"The API returned HTTP {exc.code}." ) from exc except urllib.error.URLError as exc: raise DecisionAPIError( f"Couldn't reach the API: {exc.reason}." ) from exc try: payload = json.loads(response_body) except (json.JSONDecodeError, UnicodeDecodeError) as exc: raise DecisionAPIError( "The API returned invalid JSON." ) from exc return read_choice(payload, options) def main() -> None: correct = 0 latencies_ms = [] for text, expected in EXAMPLES: start = perf_counter() predicted = predict(text, OPTIONS) elapsed_ms = (perf_counter() - start) * 1000 correct += int(predicted == expected) latencies_ms.append(elapsed_ms) print( f"Expected={expected}, predicted={predicted}, " f"latency={elapsed_ms:.1f} ms" ) accuracy = correct / len(EXAMPLES) print(f"Accuracy: {accuracy:.1%}") print( "Median request latency: " f"{median(latencies_ms):.1f} ms" ) if __name__ == "__main__": main()Set the environment variables from Call the decision API.
Run the sample:
python support_ticket_classifier.py
The script prints the expected and predicted team for each request. It then
prints accuracy and median latency for the sample set. Replace EXAMPLES with
representative, labeled examples from your workload before you use the
measurements to make deployment decisions.
Combine decisions in one request
Questions that use the same state can share one request. For example, route a
ticket with choice, rate its severity with score, and check whether it's a
repeat contact with noul.
{
"model": "<deployment-name>",
"state": {
"message": "Checkout is down again. This is my third request.",
"customer_contact_count": 3
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Charges, invoices, and refunds",
"engineering": "Bugs, errors, and outages",
"support": "How-to and account questions"
}
},
"severity": {
"type": "score",
"instructions": "How severe is this issue?",
"criteria": ["Cosmetic", "Minor", "Major", "Critical"]
},
"repeat_contact": {
"type": "noul",
"instructions": "Has the customer contacted support before?"
}
}
}
Combining related questions reduces the number of requests and keeps the decision context consistent.
Interpret decisions
Microsoft-Decision-1 returns numerical decisions without a written rationale.
Use the probabilities to decide whether your application acts automatically,
requests another evaluation, or sends the item to a person.
When to use Microsoft-Decision-1
Use Microsoft-Decision-1 when your application needs one or more bounded,
typed decisions about the same input. It works well for high-volume workflows
where your application needs to act on a probability, selected option, or
ordered score instead of displaying generated text.
Common use cases include:
- Gates and filters: Check whether content meets a condition before a workflow continues.
- Classification and routing: Send tickets, alerts, documents, or requests to one option from a known set.
- Ranking and prioritization: Order incidents, leads, search results, or review queues by an application-defined scale.
- Evaluation and grading: Score model output against an ordered rubric.
- Model, tool, or agent selection: Choose the component that should handle a request.
Use a generative LLM instead when the application needs to create, summarize,
rewrite, or explain content. You can also combine the two approaches. Use
Microsoft-Decision-1 to route or gate a request, and then send approved
requests to a generative model.
Known limitations and risks
- Fairness: The model might reflect biases from its base model and training data. Don't use its scores as the sole basis for decisions about individuals.
- Wording sensitivity: Scores can change based on how you phrase or order questions and options. Poorly framed questions still return scores.
- Reliability: Calibration is strongest on familiar task types. The model doesn't provide explanations and might rely on outdated knowledge.
- Harmful content: When you use the model as a safety filter, it might miss subtle harmful content or flag benign content.
Best practices
Before you use Microsoft-Decision-1 in an application, evaluate the model
for your intended scenario.
- Validate the model on data that's representative of your use case.
- Tune the
state,instructions, and options for each primitive. You can improve task definition with structured or code-enriched state and parameterized options. Use coding agents to generate candidate configurations, evaluate them against labeled examples, and compare results. Review agent-generated changes before you use them in production. - Set confidence thresholds based on the cost of false positives and false
negatives. Include an abstention option, such as
cannot tell, when the model shouldn't make a forced choice. - Use clear, neutral wording for questions and options. Consider randomizing option order, and test whether changing the order affects results.
- For consequential decisions about people, such as decisions involving credit, employment, housing, healthcare, or legal matters, use the model for decision support with meaningful human review. Don't use it as the sole decision-maker.
- Avoid sensitive attributes that aren't necessary for the use case. Monitor outcomes for errors and unfair disparities.
- Tell affected users when AI contributes to a decision.
Troubleshoot
| Error | Possible cause | Resolution |
|---|---|---|
| 400 Bad Request | The request uses an invalid question type or shape. | Check state, questions, type, and criteria. |
| 401 Unauthorized | The credential is missing, invalid, or expired. | Refresh the token, or verify the API key and authentication header. |
| 403 Forbidden | The identity doesn't have access to the deployment. | Verify the role assignment and deployment access. |
| 404 Not Found | The resource endpoint or API path is incorrect. | Verify the endpoint on the deployment details page. |
| 422 Unprocessable Entity | The deployment name or question definition isn't supported. | Verify the model value and each question definition. |
| 429 Too Many Requests | The deployment rate limit was exceeded. | Retry with exponential backoff, or request more quota. |