Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Confidential GPUs provide hardware-based security and isolation for workloads running on Azure Kubernetes Service (AKS). They help protect sensitive data and computations from unauthorized access, even from privileged users. To create a node pool with Confidential GPUs, you need to skip the automatic GPU driver installation and install the NVIDIA GPU Operator to provide the signed drivers.
This article shows you how to create an AKS node pool with a NCCads_H100_v5 sizes series VM size.
By enabling confidential computing on GPUs, you have more options and flexibility to run your workloads securely and efficiently on the cloud. These virtual machines (VMs) are ideal for inferencing, fine-tuning, or training small-to-medium sized models. You can also use them for applied AI workloads, such as:
- GPU-accelerated analytics and databases
- Batch inferencing with heavy pre-processing and post-processing
- Machine learning (ML) development
- Video processing
- AI/ML web services
Limitations
- Confidential GPUs aren't supported for Windows nodes.
- You can't use Confidential GPUs with node auto-provisioning (NAP).
Prerequisites
- This article assumes you have an existing AKS cluster. If you don't have a cluster, create one using the Azure CLI, Azure PowerShell, or the Azure portal.
- You need the Azure CLI version 2.72.2 or later installed to set the
--gpu-driverfield. Runaz --versionto find the version. If you need to install or upgrade, see Install Azure CLI.
Note
GPU-enabled VMs contain specialized hardware subject to higher pricing and region availability. For more information, see the pricing tool and region availability.
Get the credentials for your cluster
Get the credentials for your AKS cluster by using the az aks get-credentials command. The following example gets the credentials for the cluster myAKSCluster in the myResourceGroup resource group:
az aks get-credentials --resource-group myResourceGroup --name myAKSCluster
Note
The NVIDIA GPU Operator isn't compatible with multiple OS versions on the same AKS cluster.
Create a node pool with Confidential GPUs
The NVIDIA GPU Operator automates the management and deployment of all NVIDIA software components needed to provision a GPU, including driver installation, the NVIDIA device plugin for Kubernetes, the NVIDIA container runtime, and more. Because the NVIDIA GPU Operator handles these components, you don't need to separately install the NVIDIA device plugin on your AKS cluster. This setup also means you should skip automatic GPU driver installation to use the NVIDIA GPU Operator on AKS.
Important
Open-source software is mentioned throughout AKS documentation and samples. Software that you deploy is excluded from AKS service-level agreements, limited warranty, and Azure support. As you use open-source technology alongside AKS, consult the support options available from the respective communities and project maintainers to develop a plan.
Microsoft takes responsibility for building the open-source packages that we deploy on AKS. That responsibility includes having complete ownership of the build, scan, sign, validate, and hotfix process, along with control over the binaries in container images. For more information, see Vulnerability management for AKS and AKS support coverage.
The following example shows how to create an AKS node pool with Confidential GPUs by using the Azure NCCads_H100_v5 VM size.
Skip automatic GPU driver installation by creating an NVIDIA GPU-enabled node pool by using the
az aks nodepool addcommand and setting the API field--gpu-driverto the valuenone. Setting this API field tononeduring node pool creation skips the default GPU driver installation. For an example, see Skip GPU driver installation. This setting doesn't change any existing nodes. You can scale the node pool to zero and then back up to make the change take effect.az aks nodepool add \ --resource-group myResourceGroup \ --cluster-name myAKSCluster \ --name gpunp \ --node-count 1 \ --node-vm-size Standard_NCC40ads_H100_v5 \ --node-taints sku=gpu:NoSchedule \ --gpu-driver noneInstall the NVIDIA GPU Operator. The GPU Operator might rely on autodetection of the confidential GPU SKU, but you can enforce signed NVIDIA GPU driver installation by explicitly pinning the
-signedtag. For example:Note
NVIDIA publishes a compatible signed driver image with your Ubuntu version, so you should verify this image when installing the GPU Operator.
helm install gpu-operator nvidia/gpu-operator \ -n gpu-operator --create-namespace \ --set driver.enabled=true \ --set driver.repository=nvcr.io/nvidia \ --set driver.version=550.90.07-signed-ubuntu22.04 \ --set toolkit.enabled=true \ --waitThe output should show a
STATUS: deployedmessage.Verify that the NVIDIA GPU Operator is running by checking the status of the pods in the
gpu-operatornamespace:kubectl get pods -n gpu-operatorExample output:
NAME READY STATUS RESTARTS AGE gpu-operator-xxxx 1/1 Running 0 5m nvidia-driver-daemonset-xxxxx 2/2 Running 0 4m nvidia-container-toolkit-daemonset-xxxxx 1/1 Running 0 3m nvidia-device-plugin-daemonset-xxxxx 1/1 Running 0 2m nvidia-dcgm-exporter-xxxxx 1/1 Running 0 2m nvidia-operator-validator-xxxxx 1/1 Running 0 2m gpu-feature-discovery-xxxxx 1/1 Running 0 2mVerify that the NVIDIA GPU driver is loaded and the module is signed:
DRIVER_POD=$(kubectl get pod -n gpu-operator -l app=nvidia-driver-daemonset -o jsonpath='{.items[0].metadata.name}') kubectl exec -n gpu-operator -it $DRIVER_POD -c nvidia-driver-ctr -- nvidia-smi kubectl exec -n gpu-operator -it $DRIVER_POD -c nvidia-driver-ctr -- modinfo nvidia | grep -i sigIn the output,
nvidia-smishould show the H100 GPU and a driver version of 550.90.07.modinfoshould show signer fields. For example:signer: NVIDIA CERTIFICATE sig_key: <key-id> sig_hashalgo: sha512Missing signer fields indicate that the module isn't signed despite the image tag.
Confirm that the GPU is schedulable:
kubectl get node <node-name> -o jsonpath='{.status.capacity.nvidia\.com/gpu}'The output should show the number of GPUs available on the node. For this example, it should show
1for a single GPU.
Clean up resources
To clean up the resources you created for the NVIDIA GPU Operator, delete the namespace:
kubectl delete namespace gpu-operator
You can also delete the node pool that you created for the NVIDIA GPU Operator by using the az aks nodepool delete command. For example:
az aks nodepool delete \
--resource-group myResourceGroup \
--cluster-name myAKSCluster \
--name gpupool
Next steps
To learn more about confidential features on AKS, see the following resources: