Available models in Unity Gateway

Databricks hosts the latest partner and open-weight models. Use Unity Gateway to access these models for applications, coding agents, and AI workloads on your data.

Start querying models.

Models available through Databricks

The following models are hosted by Databricks. Each model name links to its specifications.

Model assets in system.ai are listed globally. Seeing a model there does not mean it is available in your workspace. Availability depends on your workspace region, cross-geo settings, and model availability. Check the Unity Gateway UI for the corresponding model service and confirm that you have permission to query it.

Listing a model does not send customer data to it for inference. See Model availability by region for regional details.

Note

On Azure Databricks, OpenAI and Google Gemini models, along with Zhipu AI GLM 5.3, GLM 5.3 Flash, and GLM 5.2, Moonshot AI Kimi K3, Thinking Machine Labs Inkling, and DeepSeek V4.1 Flash and V4 Flash, are available through ADI Services, provided by Databricks. For these model descriptions, see Detailed list of models supported by Databricks Foundation Model APIs. For setup and access, see ADI Services.

Provider Model family Models
Anthropic Claude Haiku Claude Haiku 4.5
Anthropic Claude Sonnet Claude Sonnet 5.5, Claude Sonnet 5, Claude Sonnet 4.6, Claude Sonnet 4.5
Anthropic Claude Fable Claude Fable 5.1, Claude Fable 5
Anthropic Claude Opus Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Opus 4.1
OpenAI GPT GPT-6 Astra, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5 Pro, GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5.4 nano, GPT-5.2, GPT-5.1, GPT-5, GPT-5 mini, GPT-5 nano
OpenAI GPT Codex GPT-5.3 Codex
OpenAI GPT OSS GPT OSS 120B, GPT OSS 20B
Google Gemini Flash Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Flash Lite, Gemini 3 Flash
Google Gemini Pro Gemini 3.1 Pro Preview
Google Gemini image generation Gemini 3.1 Flash Image, Gemini 3 Pro Image
Google Gemma Gemma 3 12B
Meta Llama Llama 4 Maverick, Llama 3.3 70B Instruct, Llama 3.1 8B Instruct
Alibaba Cloud Qwen OpenJev (Qwen3.5 4B), Qwen3.5 122B A10B, Qwen3-Embedding-0.6B, Qwen3-Next 80B A3B Instruct
Moonshot AI Kimi Kimi K3
Zhipu AI GLM GLM 5.3, GLM 5.3 Flash, GLM 5.2
DeepSeek DeepSeek V4.1 Flash, V4 Flash (0731)
Thinking Machine Labs Inkling Inkling
Alibaba GTE GTE Large (En)

For deprecated models and recommended replacements, see Deprecated and retired models.

Use a model

Your task How to use the models
Build an application or agent Call a model through the Unity Gateway API. System-provided model services in system.ai let you start without creating a custom model service. See Query models.
Apply AI to your data Use ai_query to send prompts to a supported model from SQL, notebooks, or batch pipelines. See Use ai_query.
Use a coding agent Connect your coding agent to models through Unity Gateway. See Supported coding agents.

Supported models and model identifiers vary by interface. See the linked guides for requirements and examples.

System-provided model services

Azure Databricks provides ready-to-use, pay-per-token model services in system.ai, such as system.ai.claude-opus-5. New services are added as foundation models become available.

  • A system user owns them, and you cannot delete them.
  • By default, only metastore administrators can modify them. A metastore administrator can delegate management by granting the MANAGE privilege.

To restrict access, see Discover and govern access to model services.

Choose serving capacity

Option Recommended for How it works
Pay-per-token Workloads with variable or intermittent usage. Pay for the tokens you use without provisioning dedicated serving capacity.
On-demand provisioned throughput Workloads that need dedicated capacity with flexibility to adjust it as requirements change. Pay for allocated capacity with no term commitment.

Unity Gateway model services can route to provisioned-throughput destinations. See each option's documentation for eligible models and setup requirements.

For batch inference with ai_query, use the AI Functions guidance to select supported models and serving options. Model and regional support vary by serving option.

Connect other model providers

You can also use Unity Gateway with your own model provider account. Connect the model provider and govern access using your organization's credentials.

See Connect an external model provider.

Related information