Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Important
This feature is in Public Preview.
Note
If you're using Ray on Databricks classic compute with Databricks Runtime ML, see Ray on Databricks.
Ray is an open source framework for scaling Python workloads. You can run Ray on AI Runtime without provisioning or managing the underlying GPU infrastructure.
Why use Ray on AI Runtime
Ray is commonly used to distribute Python workloads such as model training, batch inference, data processing, and hyperparameter tuning. On AI Runtime, Ray can access data in Unity Catalog volumes, track experiments with MLflow, and run as part of production workflows.
You can use common Ray libraries such as Ray Core, Ray Data, Ray Train, and Ray Tune on AI Runtime. Existing Ray workloads can continue to use the libraries' standard APIs.
How Ray works with AI Runtime
Every Ray application runs on a Ray cluster, which can contain one or more nodes. AI Runtime provisions the nodes, and Ray schedules tasks and actors across their available CPU and GPU resources. How you start Ray depends on the execution interface:
| Interface | Nodes | Start Ray |
|---|---|---|
| Notebook | The single node attached to the notebook | Call ray_init() from the serverless_gpu package |
| Databricks CLI | A fixed set of one or more nodes | Start the head on node 0 and join the remaining nodes as workers using a reusable bootstrap script |
Use Ray in notebooks
When your notebook is connected to AI Runtime GPU compute, use ray_init() from the serverless_gpu package to start Ray on the attached node:
from serverless_gpu import ray_init
ray_init()
ray_init() calls the standard ray.init() API, configures the Ray dashboard for access through the Databricks driver proxy, and prints the dashboard URL in the notebook output. Use the dashboard to inspect Ray jobs, tasks, actors, logs, and resource usage while your code runs.
Note
ray_init() requires environment version 5 and above. Databricks AI v6 includes Ray. If you use Standard v6, install Ray before calling ray_init(). See Set up your environment.
Notebook examples
| Example | Description |
|---|---|
| Ray Core hello world | Submit asynchronous GPU tasks on attached compute, inspect scheduling in the Ray dashboard, and retrieve the results. |
| CIFAR-10 hyperparameter tuning with Ray Tune | Run concurrent fractional-GPU trials for a PyTorch image classifier and use the asynchronous successive halving algorithm (ASHA) to stop underperforming configurations early. |
| Qwen2.5-32B batch inference with Ray Data and vLLM | Run multilingual batch inference with eight persistent vLLM replicas on 8 H100 GPUs and save the results as Parquet in a Unity Catalog volume. |
Use Ray with the Databricks CLI
The Databricks CLI commands for AI Runtime support single-node and multi-node Ray workloads. In the workload configuration, specify a fixed accelerator_type and the total num_accelerators. The Ray hello world examples include a reusable bootstrap script. In the workload's command, set RAY_ENTRYPOINT to the path of your Python file and run the bootstrap script. The bootstrap script does the following:
- On node 0, it starts the Ray head and runs the configured entry point.
- On every other node, it starts a Ray worker and keeps it connected while the entry point runs.
- After the entry point finishes, it stops the Ray processes.
In the application, connect to the cluster with ray.init(address="auto").
With the Standard environment, add ray or the required extra, such as ray[data], ray[train], or ray[tune], to the workload dependencies. Databricks AI v6 includes Ray.
CLI examples
| Example | Description |
|---|---|
| Ray hello world examples | Minimal single-node and multi-node examples for Ray Core, Ray Train, Ray Data, and Ray Tune, including the cluster bootstrap pattern. |
| Distributed training with Ray Train | Fine-tune a large language model across 8 H100 GPUs on a single node using Ray Train and PyTorch. |
| Batch inference with Ray Data and vLLM | Run large-scale batch inference using Ray Data for distributed data loading and vLLM for efficient model serving. |
| Hyperparameter search with Ray Tune | Search LoRA fine-tuning hyperparameters for Qwen2.5 with Ray Tune, running one trial per GPU and stopping underperforming trials early with the ASHA scheduler. |
Work with data and Databricks features
Ray workloads can access data in Unity Catalog volumes, record runs with MLflow, and run as Lakeflow Jobs tasks defined with Declarative Automation Bundles.
Access data in Unity Catalog
For tabular data in a Unity Catalog table, use ray.data.read_databricks_tables() to read a table or run a SQL query through a Databricks SQL warehouse. Specify the warehouse ID, catalog, and schema because AI Runtime does not provide a local Spark session. You can use SQL to perform joins, aggregations, and filters before Ray processes the results.
For file-based and unstructured data, use a Unity Catalog volume. Ray workers can access files using their /Volumes/<catalog>/<schema>/<volume>/... paths. For example, Ray Data can read Parquet files with ray.data.read_parquet() and write results with Dataset.write_parquet().
Track and orchestrate workloads
Use MLflow to record parameters, metrics, and artifacts from a Ray workload. See Experiment tracking and observability.
To schedule the workload or compose it with CPU data-preparation tasks, use Lakeflow Jobs and Declarative Automation Bundles. See Productionize training workloads.
Migrate an existing Ray workload
Migrating a self-managed Ray workload involves replacing its infrastructure configuration and service integrations with the AI Runtime execution model.
Assess your workload
Before migrating, review the limitations. Confirm that the workload can run on a fixed-size cluster with one accelerator type.
Migrate the workload
To migrate a self-managed Ray workload:
- Choose a notebook for interactive, single-node development or the Databricks CLI for single-node or multi-node jobs.
- Map the existing cluster resources to an AI Runtime accelerator type and accelerator count. See Hardware options.
- Replace the existing cluster startup with
ray_init()in a notebook or the documented head and worker bootstrap in a Databricks CLI workload. - Declare the workload's Python dependencies and update its data, checkpoint, model, and output paths.
- Validate the workload on the smallest applicable configuration. Confirm its data access, resource requests, and outputs before testing the intended fixed topology.
Limitations
Ray on AI Runtime has the following limitations:
- The Ray cluster has a fixed node count for the duration of a workload. Ray cluster autoscaling is not supported.
- All nodes in a workload use the same
accelerator_type. A Ray cluster cannot contain separate CPU-only and GPU worker groups or scale CPU and GPU capacity independently. - A notebook starts Ray on its attached single node. Use the Databricks CLI for a multi-node Ray cluster.
- Ray dashboard access through the Databricks driver proxy is currently available only for notebook workloads.
Tip
If CPU preparation and GPU processing can run as separate stages, use a multi-task Lakeflow Job and pass data between the tasks through a Unity Catalog volume.