Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
As an AI engineer or developer, you might need to prototype and deploy AI workloads with a range of different model weights. AKS provides the option to deploy inferencing workloads by using open-source presets that the KAITO model catalog supports and manages out of the box, or to dynamically download models from the HuggingFace registry at runtime onto your AKS cluster.
In this article, you learn how to onboard a sample HuggingFace model for inferencing with the AI toolchain operator add-on without having to manage custom images on Azure Kubernetes Service (AKS).
Prerequisites
An Azure account with an active subscription. If you don't have an account, you can create one for free.
An AKS cluster with the AI toolchain operator add-on enabled. For more information, see Enable KAITO on an AKS cluster.
Quota for the
Standard_NVadsA10_v5virtual machine (VM) family in your Azure subscription. If you don't have quota for this VM family, request a quota increase.Note
Currently, only the HuggingFace runtime supports inference with the KAITO custom model deployment template.
Choose an open-source language model from HuggingFace
In this example, we use the HuggingFaceTB SmolLM2-1.7B-Instruct small language model. Alternatively, you can choose from thousands of text-generation models supported on HuggingFace.
Connect to your AKS cluster using the
az aks get-credentialscommand.az aks get-credentials --resource-group <resource-group-name> --name <aks-cluster-name>Clone the KAITO project GitHub repository using the
git clonecommand.git clone https://github.com/kaito-project/kaito.git
Deploy your model inferencing workload using the KAITO workspace template
Navigate to the
kaitodirectory and copy the sample deployment YAML manifest. Replace the default values in the following fields with your model's requirements:instanceType: The VM size for your inference service deployment. For larger model sizes you can choose a VM in theStandard_NCads_A100_v4family with higher memory capacity.MODEL_ID: Your model's specific HuggingFace identifier, which can be found afterhttps://huggingface.co/in the model card URL."--torch_dtype": Set to"bfloat16", which is supported on all GPU SKUs that KAITO supports.- For this example, use
Standard_NV36ads_A10_v5as the instance type and the HuggingFaceTB SmolLM2-1.7B-Instruct model.
apiVersion: kaito.sh/v1beta1 kind: Workspace metadata: name: workspace-custom-llm resource: instanceType: "Standard_NV36ads_A10_v5" # Required VM SKU based on model requirements labelSelector: matchLabels: apps: custom-llm inference: template: spec: containers: - name: custom-llm-container image: mcr.microsoft.com/aks/kaito/kaito-base:0.2.0 # KAITO base image which includes hf runtime livenessProbe: failureThreshold: 3 httpGet: path: /health port: 5000 scheme: HTTP initialDelaySeconds: 600 periodSeconds: 10 successThreshold: 1 timeoutSeconds: 1 readinessProbe: failureThreshold: 3 httpGet: path: /health port: 5000 scheme: HTTP initialDelaySeconds: 30 periodSeconds: 10 successThreshold: 1 timeoutSeconds: 1 resources: requests: nvidia.com/gpu: 1 # Request 1 GPU; adjust as needed limits: nvidia.com/gpu: 1 # Optional: Limit to 1 GPU command: - "accelerate" args: - "launch" - "--num_processes" - "1" - "--num_machines" - "1" - "--gpu_ids" - "all" - "tfs/inference_api.py" - "--pipeline" - "text-generation" - "--trust_remote_code" - "--allow_remote_files" - "--pretrained_model_name_or_path" - "HuggingFaceTB/SmolLM2-1.7B-Instruct" # The model's HuggingFace identifier - "--torch_dtype" - "bfloat16" volumeMounts: - name: dshm mountPath: /dev/shm volumes: - name: dshm emptyDir: medium: MemorySave these changes to your
custom-model-deployment.yamlfile.Run the deployment in your AKS cluster using the
kubectl applycommand.kubectl apply -f custom-model-deployment.yaml
Test your custom model inferencing service
Track the live resource changes in your KAITO workspace using the
kubectl get workspacecommand.kubectl get workspace workspace-custom-llm -wNote
Note that machine readiness can take up to 10 minutes, and workspace readiness up to 20 minutes. Proceed to the next step only once the workspace status shows
Ready.Once the workspace is ready, port forward the inference service to your local machine in a separate terminal.
kubectl port-forward svc/workspace-custom-llm 5000:80Install the OpenAI Python client.
pip install openaiSave the following script as
test_inference.pyand run it usingpython test_inference.py:from openai import OpenAI client = OpenAI( base_url="http://127.0.0.1:5000/v1", api_key="unused", ) response = client.chat.completions.create( model="workspace-custom-llm", messages=[{"role": "user", "content": "What sport should I play in rainy weather?"}], max_tokens=400, stream=True, ) print("".join(chunk.choices[0].delta.content or "" for chunk in response))
Clean up resources
If you no longer need these resources, you can delete them to avoid incurring extra Azure compute charges.
Delete the KAITO inference workspace using the kubectl delete workspace command.
kubectl delete workspace workspace-custom-llm
Next steps
In this article, you learned how to onboard a Hugging Face model for inferencing with the AI toolchain operator add-on directly to your AKS cluster. To learn more about AI and machine learning on AKS, see the following articles: