Bemærk
Adgang til denne side kræver godkendelse. Du kan prøve at logge på eller ændre mapper.
Adgang til denne side kræver godkendelse. Du kan prøve at ændre mapper.
In this article, you deploy the infrastructure required to run Ray AI workloads with Kueue on Azure Kubernetes Service (AKS). The Terraform module provisions an AKS cluster with GPU node pools, installs the KubeRay operator and Kueue controller via Helm, creates Azure Blob Storage with pre-staged datasets, and configures workload identity for secure access.
Important
Open-source software is mentioned throughout AKS documentation and samples. Software that you deploy is excluded from AKS service-level agreements, limited warranty, and Azure support. As you use open-source technology alongside AKS, consult the support options available from the respective communities and project maintainers to develop a plan.
Microsoft takes responsibility for building the open-source packages that we deploy on AKS. That responsibility includes having complete ownership of the build, scan, sign, validate, and hotfix process, along with control over the binaries in container images. For more information, see Vulnerability management for AKS and AKS support coverage.
Prerequisites
- An Azure subscription. If you don't have one, create a free account.
- Azure CLI version 2.70 or later, authenticated with
az login. If you have multiple subscriptions, set the correct one withaz account set --subscription <subscription-id>. - Terraform version 1.6 or later.
- kubectl version 1.28 or later. Install with
az aks install-cli. - Python version 3.10 or later (required for Aurora data generation).
- GPU quota in your target Azure region. The default configuration provisions a
Standard_ND96amsr_A100_v4node (8×A100 80 GB GPUs).
Clone the repository
Clone the Azure/AKS repository and navigate to the infrastructure module:
git clone https://github.com/Azure/AKS
cd AKS/examples/kueue-and-ray-on-aks/1-infrastructure
Install Python dependencies
The Aurora data upload step requires Python packages. Install them before deploying:
pip install -r terraform/requirements-generator.txt
Check GPU quota
Verify that your subscription has quota for the GPU SKU in your target region. Replace eastus2 with your preferred region:
LOCATION=<your-azure-region>
az vm list-usage --location "$LOCATION" --output table | grep -i "NDAMSv4_A100"
Note
If you don't have GPU quota or want to validate the infrastructure without GPUs, deploy with gpu_enabled=false:
terraform apply -var="subscription_id=<your-subscription-id>" -var="gpu_enabled=false"
The CPU-only path is useful for validating infrastructure and queue configuration. Workloads that request GPUs stay Pending.
Deploy with Terraform
Initialize and apply the Terraform module:
cd terraform
terraform init
terraform apply -var="subscription_id=<your-subscription-id>"
The Terraform module creates:
- An AKS cluster with OIDC issuer and workload identity enabled
- A system node pool and a GPU node pool (when
gpu_enabled=true) - KubeRay operator and Kueue controller (Helm releases from MCR)
- GPU monitoring DaemonSet (when
gpu_enabled=true) - Azure Blob Storage account with
auroraandllm-pipelinecontainers - User-assigned managed identity with Storage Blob Data Contributor role
- Federated identity credential for the
ray-workloadServiceAccount - Pre-staged datasets (Aurora WeatherBench2 data and viggo NLG dataset)
For the full Terraform module configuration, see the 1-infrastructure directory in the repository.
Connect to the cluster
Retrieve the credentials for your new AKS cluster:
eval "$(terraform output -raw get_credentials_command)"
Verify the deployment
Confirm that the operators are running and nodes are available:
Verify the KubeRay operator:
kubectl -n kuberay-system get podsExpected output:
NAME READY STATUS RESTARTS AGE kuberay-operator-xxxxxxxxxx-xxxxx 1/1 Running 0 5mVerify the Kueue controller:
kubectl -n kueue-system get podsExpected output:
NAME READY STATUS RESTARTS AGE kueue-controller-manager-xxxxxxxxxx-xxxxx 1/1 Running 0 5mVerify the nodes:
kubectl get nodesExpected output (with
gpu_enabled=true):NAME STATUS ROLES AGE VERSION aks-gpupool-xxxxxxxx-vmssxxxxxx Ready <none> 5m v1.35.x aks-system-xxxxxxxx-vmssxxxxxx Ready <none> 10m v1.35.x aks-system-xxxxxxxx-vmssxxxxxx Ready <none> 10m v1.35.xVerify GPU monitoring (only when
gpu_enabled=true):kubectl -n gpu-monitoring get daemonsetsExpected output:
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE gpu-monitoring-a100 1 1 1 1 1 <none> 5m
Clean up resources
To delete all Azure resources that you created by using this module:
terraform destroy -var="subscription_id=<your-subscription-id>"