Muistiinpano
Tämän sivun käyttö edellyttää valtuutusta. Voit yrittää kirjautua sisään tai vaihtaa hakemistoa.
Tämän sivun käyttö edellyttää valtuutusta. Voit yrittää vaihtaa hakemistoa.
Important
This feature is in Public Preview. To use it, a workspace admin must enable the AI Runtime preview from the Previews page. See Manage Azure Databricks previews.
Use Declarative Automation Bundles to schedule an AI Runtime GPU workload and combine it with other work, such as upstream data preparation. Define the workload and its schedule in YAML, then deploy them as a job. For a training pipeline, run preprocessing on CPU compute and start GPU training after the data is ready.
This guide starts with a workload that prints "Hello world", then builds a scheduled preprocessing and training pipeline. If you already have an AI Runtime workload YAML, see Convert an AI Runtime workload to a bundle.
How it works
- A bundle contains your code and a
databricks.ymlconfiguration. - A job groups tasks and defines when they run.
- An
ai_runtime_taskruns your command on serverless GPU compute.
The depends_on field connects tasks into a directed acyclic graph (DAG), and the schedule field schedules your job to run on a cadence.
databricks bundle deploy uploads the code and creates or updates the job. databricks bundle run starts a job run immediately. On each run, AI Runtime provisions the GPU compute, runs your command, and records the run in the named MLflow experiment.
Requirements
- A workspace in a supported region. See Requirements.
- The latest Databricks CLI, authenticated to your workspace. Update an existing installation before using these examples.
- For the preprocessing example, an existing Unity Catalog volume. The job's run identity needs
USE CATALOG,USE SCHEMA,READ VOLUME, andWRITE VOLUMEprivileges for that volume. See Privileges for Unity Catalog volumes.
Hello world example with ai_runtime_task
This example runs a shell command on one A10 GPU.
Create a directory for the bundle. In that directory, create
command.shwith the following contents:#!/usr/bin/env bash set -euo pipefail echo "Hello world"Create
databricks.ymlin the same directory:bundle: name: hello-ai-runtime resources: jobs: hello: name: hello-ai-runtime tasks: - task_key: hello environment_key: default ai_runtime_task: experiment: hello-ai-runtime deployments: - command_path: ${workspace.file_path}/command.sh compute: accelerator_type: GPU_1xA10 accelerator_count: 1 environments: - environment_key: default spec: environment_version: '6' targets: dev: mode: development default: trueFrom the bundle directory, validate, deploy, and run the job:
databricks bundle validate --target dev databricks bundle deploy --target dev databricks bundle run hello --target devOpen the run URL printed by the CLI and view the
hellotask's output. It containsHello world. The run also appears in thehello-ai-runtimeMLflow experiment.
More complex example: Schedule a preprocessing and training pipeline
This example prepares a small dataset on CPU compute, then trains a linear model on one A10 GPU. Both tasks use the same volume file to pass data between them. The job is scheduled for 09:00 UTC daily, with the schedule paused until you test it.
Create the project
Create a separate directory with the following layout:
scheduled-training/ ├── databricks.yml ├── prep.py ├── command.sh └── src/ └── train.pyCreate
prep.pywith the following contents. Replace<catalog>,<schema>, and<volume>with your volume's names. This source notebook normalizes the input values and writes the prepared data to the volume:# Databricks notebook source import json from pathlib import Path data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json") inputs = [value / 10 for value in range(10)] data = {"x": inputs, "y": [2 * value + 1 for value in inputs]} data_path.parent.mkdir(parents=True, exist_ok=True) data_path.write_text(json.dumps(data)) print(f"Prepared {len(inputs)} rows at {data_path}")Create
src/train.pywith the following contents. Use the same volume path as inprep.py:import json from pathlib import Path import torch data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json") data = json.loads(data_path.read_text()) x = torch.tensor(data["x"], device="cuda").reshape(-1, 1) y = torch.tensor(data["y"], device="cuda").reshape(-1, 1) model = torch.nn.Linear(1, 1).to("cuda") optimizer = torch.optim.SGD(model.parameters(), lr=0.1) for _ in range(200): optimizer.zero_grad() loss = torch.nn.functional.mse_loss(model(x), y) loss.backward() optimizer.step() print(f"Trained on {x.device}: loss={loss.item():.4f}")Create
command.shto run the training code. The bundle below packagessrc/, soCODE_SOURCE_PATHpoints to the extractedsrcdirectory:#!/usr/bin/env bash set -euo pipefail cd "$CODE_SOURCE_PATH" python train.pyCreate
databricks.ymlwith the following contents. The bundle packagessrc/for the GPU task. Theprepnotebook runs on serverless CPU compute, anddepends_onstartstrainonly afterprepsucceeds.max_concurrent_runs: 1prevents runs of this job from overwriting each other's input file:bundle: name: scheduled-training artifacts: code: type: tgz path: . include: [src] files: - source: ./dist/code.tgz resources: jobs: train_pipeline: name: scheduled-training max_concurrent_runs: 1 schedule: quartz_cron_expression: '0 0 9 * * ?' timezone_id: UTC pause_status: PAUSED tasks: - task_key: prep notebook_task: notebook_path: ./prep.py - task_key: train depends_on: - task_key: prep environment_key: training ai_runtime_task: experiment: scheduled-training code_source_path: ./dist/code.tgz deployments: - command_path: ${workspace.file_path}/command.sh compute: accelerator_type: GPU_1xA10 accelerator_count: 1 environments: - environment_key: training spec: environment_version: '6' dependencies: - torch targets: dev: mode: development default: true
Test and enable the schedule
From
scheduled-training/, deploy the bundle and start a manual run:databricks bundle validate --target dev databricks bundle deploy --target dev databricks bundle run train_pipeline --target devOpen the run URL printed by the CLI. Confirm that
prepsucceeds beforetrainstarts. Theprepoutput reports 10 prepared rows, and thetrainoutput reportsTrained on cuda:0and the loss.In
databricks.yml, change the schedule'spause_statusfromPAUSEDtoUNPAUSED. The schedule block is now:schedule: quartz_cron_expression: '0 0 9 * * ?' timezone_id: UTC pause_status: UNPAUSEDDeploy the change to activate the daily schedule:
databricks bundle deploy --target devSetting
pause_status: UNPAUSEDexplicitly enables this schedule even for a development target. In Jobs & Pipelines, open the deployed job and confirm that its schedule is active for 09:00 UTC. See Run jobs on a schedule.
To pause scheduled runs, set pause_status: PAUSED and deploy again. To remove the example job and the files uploaded by the bundle, run the following from its bundle directory:
databricks bundle destroy --target dev
The prepared data in your volume remains. Delete the scheduled-training data directory when you no longer need it.
Additional resources
- Convert an existing AI Runtime workload YAML into a bundle. See Convert an AI Runtime workload to a bundle.
- Configure code packaging, compute, environments, and task parameters. See Configure AI Runtime bundle tasks.
- Schedule an existing notebook on GPU. See Schedule with the Jobs API and Declarative Automation Bundles.
- Track training runs. See Experiment tracking and observability.
- Checkpoint model, optimizer, and data pipeline state. See Improve training performance and resiliency on AI Runtime.