Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
This article shows you how to package a bring-your-own (BYO) model, push it to an OCI-compatible registry, and deploy it on Foundry Local on Azure Local.
Use this article when your model isn't available in the Foundry catalog and you want to deploy a custom, fine-tuned, or third-party model from your own registry.
Important
- Foundry Local is available in preview. Preview releases provide early access to features that are in active deployment.
- Features, approaches, and processes can change or have limited capabilities before general availability (GA).
Prerequisites
Before you begin, make sure that you have:
- A running Foundry Local on Azure Local environment.
kubectlinstalled and configured for your cluster.- Access to an OCI-compatible registry, such as Azure Container Registry.
- ORAS installed. For installation steps, see ORAS.
- A model package that matches the runtime and workload requirements. For packaging requirements and limitations, see Bring your own models in Foundry Local on Azure Local.
Choose a deployment pattern
Foundry Local on Azure Local supports two BYO deployment patterns:
- Inline custom model: Define the model source directly in the
ModelDeploymentresource. Use this option for a single deployment. - Named model resource: Create a
Modelresource and reference it from one or moreModelDeploymentresources. Use this option when you want to reuse the same model definition across deployments.
Prepare your model files
Organize your model files in a directory that matches the runtime you want to use.
- For ONNX Runtime generative workloads, include the ONNX model file and required tokenizer and generation configuration files.
- For ONNX Runtime predictive workloads, include the ONNX model file.
- For vLLM workloads, include the Hugging Face model configuration, weights, and tokenizer files.
For detailed file requirements, see Bring your own models in Foundry Local on Azure Local.
Package and push the model
Organize your model files in a working directory.
Example directory layout:
my-model/ config.json model.safetensors tokenizer.json tokenizer_config.jsonPackage the model files as a
.tar.gzarchive.tar -czf my-model.tar.gz -C my-model .Sign in to your registry and push the archive by using ORAS.
oras login myregistry.azurecr.io -u <username> -p <password> oras push myregistry.azurecr.io/models/my-model:v1 my-model.tar.gz
The model cache job unpacks the .tar.gz archive after it pulls the artifact from your external registry.
Create registry credentials in Kubernetes
Create a Kubernetes secret in the same namespace where you plan to deploy the model.
kubectl create secret generic my-registry-secret \
-n foundry-local-operator \
--from-literal=username=<registry-username> \
--from-literal=password=<registry-password-or-token>
Deploy the model by using inline custom model configuration
Use this option when you want to define the model source directly in the ModelDeployment manifest.
Create a file named
modeldeployment-byo-inline.yaml.apiVersion: foundrylocal.azure.com/v1 kind: ModelDeployment metadata: name: my-custom-model namespace: foundry-local-operator spec: workloadType: generative compute: gpu runtime: vllm model: custom: registry: myregistry.azurecr.io repository: models/my-model tag: v1 credentials: secretRef: name: my-registry-secret usernameKey: username passwordKey: password vllm: preferences: gpu_memory_utilization: 0.92 max_model_len: 8192 resources: requests: cpu: "2" memory: "8Gi" limits: cpu: "4" memory: "16Gi" gpu: 1Apply the manifest.
kubectl apply -f modeldeployment-byo-inline.yaml
If you scale this deployment beyond a single replica, the operator automatically enables inference-aware routing across the replicas. The operator deploys an Endpoint Picker (EPP) that scores replicas on queue depth, KV-cache utilization, and prefix-cache locality before each request is forwarded. The HTTPRoute targets the InferencePool instead of the Service. On a single-replica deployment, EPP is off by default (nothing to balance). You only need to set spec.vllm.epp.enabled if you want to override this default - for example, set false on a multi-replica deployment to keep plain Gateway round-robin, or set true on a single-replica deployment that you intend to scale up. See Inference-aware routing with the Endpoint Picker (EPP) (concept-inference-runtimes#inference-aware-routing-with-the-endpoint-picker-epp) for prerequisites and trade-offs.
Deploy the model by using a named model resource
Use this option when you want to reuse the same model definition across multiple deployments.
Create a file named
model-byo.yaml.apiVersion: foundrylocal.azure.com/v1 kind: Model metadata: name: my-custom-model namespace: foundry-local-operator spec: displayName: "My Custom Model" source: custom: registry: myregistry.azurecr.io repository: models/my-model tag: v1 credentials: secretRef: name: my-registry-secret usernameKey: username passwordKey: password variants: - id: my-model-gpu compute: gpu priority: 10 requirements: minMemory: "16Gi" minVRAM: "8Gi" supportedCompute: [gpu] capabilities: task: chat-completion streaming: trueApply the model resource.
kubectl apply -f model-byo.yamlCreate a file named
modeldeployment-byo-ref.yaml.apiVersion: foundrylocal.azure.com/v1 kind: ModelDeployment metadata: name: my-custom-deployment namespace: foundry-local-operator spec: workloadType: generative compute: gpu runtime: vllm model: ref: my-custom-model resources: requests: cpu: "2" memory: "8Gi" limits: cpu: "4" memory: "16Gi" gpu: 1Apply the deployment manifest.
kubectl apply -f modeldeployment-byo-ref.yaml
If you scale this deployment beyond a single replica, inference-aware routing across the replicas is enabled automatically. The operator deploys an Endpoint Picker (EPP) that scores replicas on queue depth, KV-cache utilization, and prefix-cache locality before each request is forwarded, and the HTTPRoute targets the InferencePool instead of the Service. On a single-replica deployment EPP is off by default (nothing to balance). You only need to set spec.vllm.epp.enabled if you want to override this - for example, set false on a multi-replica deployment to keep plain Gateway round-robin, or set true on a single-replica deployment that you intend to scale up. See Inference-aware routing with the Endpoint Picker (EPP) (concept-inference-runtimes#inference-aware-routing-with-the-endpoint-picker-epp) for prerequisites and trade-offs.
Verify the deployment
Check the
ModelDeploymentstatus.kubectl get modeldeployment -n foundry-local-operator kubectl describe modeldeployment my-custom-model -n foundry-local-operatorIf you used a named model resource, check the
Modelstatus.kubectl get model -n foundry-local-operator kubectl describe model my-custom-model -n foundry-local-operatorReview pod status and model cache job progress.
kubectl get pods -n foundry-local-operator kubectl get jobs -n foundry-local-operator
The deployment is ready when the ModelDeployment reaches the Running state.
Troubleshoot common issues
- If the model cache job fails, verify the registry hostname, credentials, and artifact path.
- If validation fails, confirm that the package contains the required files for the selected runtime and workload type.
- If the deployment doesn't start, review the
ModelDeploymentstatus message and pod logs.
For model lifecycle details, see Inference operator and model lifecycle in Foundry Local on Azure Local. For field-level details, see ModelDeployment and operator configuration reference for Foundry Local on Azure Local.