An Azure service that provides access to OpenAI’s GPT-3 models with enterprise capabilities.
Hello Mohamed Hussein,
I understand that your deployment of the Nvidia LLM model (deepseek-r1-distill-llama-8b--1) on Azure AI Foundry has failed with an unknown error.
Could you please confirm whether you are using shared compute for deployment or if the deployment is consuming quota from your subscription?
I addition to that here are the recommended troubleshooting steps to consider,
Validate Model Name - Ensure that the model's name and deployment parameters are correctly specified in the configuration.
Check Resource Availability - Verify that the selected SKU (Standard_NC96ads_A100_v4) has sufficient available resources in your Azure region. Check your Azure subscription quotas to ensure you haven’t reached usage limits.
Network Connectivity - Confirm there are no network issues that might be preventing the deployment process. If using private endpoints, verify that outbound traffic rules allow access to required Azure services.
Review Deployment Logs - Check the deployment logs in Azure AI Foundry for any detailed error messages.
Review Azure Machine Learning logs and Azure Monitor for further insights into potential failures.
Make sure that all necessary libraries and dependencies for running the model are installed.
If using a custom environment, validate the conda.yaml or requirements.txt file.
Model Compatibility - Ensure that the Nvidia LLM model you are deploying is compatible with the selected compute instance and environment.
Deployment Configuration - Verify that all configuration settings for your deployment are correct.
Check the model path, environment settings, and dependency configurations to ensure proper setup.
I hope this information helps.
Kindly consider upvoting the comment if the information provided is helpful. This can assist other community members in resolving similar issues.