Query foundation and embedding models

Databricks provides several ways to query foundation models. Choose an interface based on whether you need a provider-independent API, provider-specific features, or batch inference.

Ways to query a model

Approach When to use it
Unified APIs Use interfaces compatible with OpenResponses and OpenAI Chat Completions. Databricks translates each request to the downstream model's native format, so you can switch models across providers without changing client code.
Provider-native APIs Use provider-specific features or existing OpenAI, Anthropic, or Google SDK code.
ai_query Run batch inference from SQL or Python.

Quickstart

Query a model service in two steps:

Step 1: Pick a model

Azure Databricks provides models out of the box, such as claude-sonnet-4-5 or gpt-5-6-sol. Azure Databricks-provided models are registered in Unity Catalog under system.ai.

Step 2: Send a request using the unified OpenAI-compatible API

Use the MLflow Chat Completions API with the OpenAI Python SDK:

from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
  api_key=DATABRICKS_TOKEN,  # your personal access token
  base_url="https://<workspace-url>/ai-gateway/mlflow/v1"  # your Databricks workspace instance
)

chat_completion = client.chat.completions.create(
  messages=[
    {"role": "user", "content": "What is Databricks?"},
  ],
  model="system.ai.claude-sonnet-4-5",
  max_tokens=256
)

print(chat_completion.choices[0].message.content)

Requirements

Query model services with unified APIs

Unified APIs offer an OpenAI-compatible interface to query models on Azure Databricks. Use unified APIs to seamlessly switch between models from different providers without changing your code.

MLflow Chat Completions API

MLflow Chat Completions API

Python

from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
  api_key=DATABRICKS_TOKEN,
  base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)

chat_completion = client.chat.completions.create(
  messages=[
    {"role": "user", "content": "Hello!"},
    {"role": "assistant", "content": "Hello! How can I assist you today?"},
    {"role": "user", "content": "What is Databricks?"},
  ],
  model="system.ai.gpt-5-6-sol",
  max_tokens=256
)

print(chat_completion.choices[0].message.content)

REST API

curl \
  -u token:$DATABRICKS_TOKEN \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{
    "model": "system.ai.gpt-5-6-sol",
    "max_tokens": 256,
    "messages": [
      {"role": "user", "content": "Hello!"},
      {"role": "assistant", "content": "Hello! How can I assist you today?"},
      {"role": "user", "content": "What is Databricks?"}
    ]
  }' \
  https://<workspace-url>/ai-gateway/mlflow/v1/chat/completions

Replace <workspace-url> with your Azure Databricks workspace URL.

MLflow Embeddings API

MLflow Embeddings API

Python

from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
  api_key=DATABRICKS_TOKEN,
  base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)

embeddings = client.embeddings.create(
  input="What is Databricks?",
  model="<model-service>"
)

print(embeddings.data[0].embedding)

REST API

curl \
  -u token:$DATABRICKS_TOKEN \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-service>",
    "input": "What is Databricks?"
  }' \
  https://<workspace-url>/ai-gateway/mlflow/v1/embeddings

Replace <workspace-url> with your Azure Databricks workspace URL and <model-service> with the fully qualified name of your model service.

Supervisor API

Supervisor API

The Supervisor API (/mlflow/v1/responses) is an OpenResponses-compatible, provider-agnostic API for building agents in Beta. Workspace admins can enable it from the Previews page. See Manage Azure Databricks previews. Pick the best model for your agent use case across providers, without changing your code.

Python

from openai import OpenAI
import os

DATABRICKS_TOKEN = os.environ.get('DATABRICKS_TOKEN')

client = OpenAI(
  api_key=DATABRICKS_TOKEN,
  base_url="https://<workspace-url>/ai-gateway/mlflow/v1"
)

response = client.responses.create(
  model="<model-service>",
  input=[{"role": "user", "content": "What is Databricks?"}]
)

print(response.output_text)

REST API

curl \
  -u token:$DATABRICKS_TOKEN \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-service>",
    "input": [
      {"role": "user", "content": "What is Databricks?"}
    ]
  }' \
  https://<workspace-url>/ai-gateway/mlflow/v1/responses

Replace <workspace-url> with your Azure Databricks workspace URL and <model-service> with the fully qualified name of your model service.

Query model services with ai_query

You can use the ai_query function to query model services directly from SQL or Python. This allows you to capture usage tracking information for your batch inference workloads.

To query a model service with ai_query, run ai_query against a model service:

SELECT ai_query(
  'system.ai.claude-sonnet-4-5',
  'Summarize the following text: ' || text_column
) AS summary
FROM my_table
LIMIT 10

The usage tracking system table (system.ai_gateway.usage) captures requests made through ai_query to model services. These requests also appear in the built-in usage dashboard.

For full ai_query syntax and parameter reference, see ai_query function. For best practices and supported models, see Use ai_query.

Limitations

  • ai_query support for Unity Gateway is only available for Azure Databricks-provided models. Pass the system.ai model service name (for example, system.ai.claude-sonnet-4-5 or system.ai.gpt-5-6-sol). Model services that you create in Unity Gateway are not yet supported.
  • Only usage tracking applies to ai_query batch inference workloads. Other Unity Gateway features such as rate limits, guardrails implemented with service policies, inference tables, and fallbacks do not apply.

Query model services with native APIs

Native APIs offer provider-specific interfaces to query models on Azure Databricks. Use native APIs to access the latest provider-specific features.

Each native API works only with model services whose underlying model uses the matching API format:

To query a model service regardless of its underlying model, use the unified APIs instead.

Next steps