Muistiinpano
Tämän sivun käyttö edellyttää valtuutusta. Voit yrittää kirjautua sisään tai vaihtaa hakemistoa.
Tämän sivun käyttö edellyttää valtuutusta. Voit yrittää vaihtaa hakemistoa.
Important
The Databricks CLI commands for AI Runtime are in Public Preview. A workspace admin must enable the AI Runtime preview before you can submit workloads. See Requirements.
The databricks air CLI runs single-node and multi-node GPU workloads on AI Runtime, the on-demand serverless GPU compute platform. Define a workload's compute, environment, and command in YAML, then submit it with databricks air run.
AI Runtime handles GPU provisioning and environment setup and provides run status, logs, system metrics, and MLflow tracking. The YAML and CLI workflow works naturally with coding agents, which can define, submit, and monitor workloads.
Use the CLI to:
- Develop and iterate on Python applications from your local environment.
- Define repeatable single-node or multi-node workloads in YAML files that you can source control.
- Submit and manage workloads from a terminal or automated development workflow.
Typical workloads include LLM fine-tuning and post-training, distributed training, batch inference, Ray applications, and hyperparameter search.
To submit a workload, follow the quickstart.
The Databricks CLI commands for AI Runtime meet the same compliance standards as AI Runtime workloads. For the supported standards and exceptions, see compliance security profile.
Note
The databricks air commands and workload configuration differ from the legacy Python-based air CLI in the databricks-air package. See Python CLI migration.
How a workload runs
You provide a workload YAML file, application code when needed, and the command that starts your application. When you submit a workload:
- The workload definition is validated against the supported schema and configuration constraints.
- If
code_sourceis configured, the selected code is packaged as a snapshot and uploaded. - The requested GPU compute and environment are prepared on each node.
- The command provided in the workload YAML runs on each node.
The submission returns a Job Run ID that identifies the workload. Use it to check status, inspect logs and configuration, or cancel the workload. You can also view the Jobs run page and MLflow experiment for execution details, system metrics, logs, and artifacts.
This workflow provides managed GPU provisioning, reproducible environment setup, and built-in observability without requiring you to configure or maintain GPU infrastructure.
Workload configuration
A workload YAML file describes the following components:
- Experiment: The MLflow experiment and run that track the workload.
- Environment: The managed environment, dependencies, or custom Docker image to use.
- Compute: The accelerator type and number of accelerators to provision.
- Code source: An optional snapshot of application code to make available to the workload.
- Command: The shell command that starts your application.
- Runtime settings: Options such as retries, timeouts, environment variables, secrets, and permissions.
For configuration fields and examples, see Workload YAML reference.
Upload application code
Use code_source to upload application files with a workload. AI Runtime makes the uploaded directory available on every node and sets $CODE_SOURCE_PATH to its location. For example, the following command runs train.py from the uploaded directory:
command: python "$CODE_SOURCE_PATH/train.py"
You can upload the current working tree or select a Git branch or commit. For a complete example, see the quickstart.
Develop and manage workloads
Development typically follows this process:
- Develop the application code and workload YAML locally.
- Validate the YAML locally with
databricks air run --file <file> --dry-run. - Submit the workload with
databricks air run --file <file>. Add--watchto stream logs until the run finishes. - Use
databricks air get,databricks air list, anddatabricks air logsto inspect workloads and logs. - Optionally stop a running workload with
databricks air cancel.
For all commands and flags, see air command group.
For scheduled runs, multi-task workflows, or promotion across environments, you can use databricks air convert-to-dabs as a starting point for a bundle. See Schedule GPU workloads and compose tasks.
Track workloads with MLflow
Each submitted workload is represented by a Databricks job run and tracked through an MLflow experiment:
- Databricks job run: Shows the workload status, execution attempts, and output on the Jobs run page.
- MLflow experiment: Contains the MLflow runs with automatically captured system metrics, workload configuration, and logs. Your application can also use the MLflow APIs to log custom parameters, metrics, and artifacts.
Use the Job Run ID with commands such as databricks air get, databricks air logs, and databricks air cancel. For more details about tracking and observability, see Track runs with MLflow and the Jobs run page.
Run single-node and distributed workloads
AI Runtime runs the configured command once on each provisioned node. For a distributed workload, it provides coordination variables such as NUM_NODES, NODE_RANK, LOCAL_WORLD_SIZE, MASTER_ADDR, and MASTER_PORT.
This execution model is framework-independent. Your command must start and configure the distributed framework, for example by invoking torchrun, a Ray bootstrap script, or another framework launcher. For examples, see Databricks CLI examples for AI Runtime.
Use an AI agent
The databricks-ai-runtime agent skill provides instructions that help AI agents, such as Claude Code, Cursor, and Codex, work with AI Runtime through the Databricks CLI. Use it to create workload YAML, submit GPU training jobs, check run status, stream logs, and cancel runs. It also includes guidance for custom Docker images.
After installing the Databricks CLI and configuring authentication, run the following command:
databricks aitools install --skills-only --skills databricks-ai-runtime --experimental
Follow the prompts to choose your AI agent and installation scope. For installation options, see aitools command group.