LLM deployment failed

Mohamed Hussein 715 Reputation points
2025-03-20T01:59:35.7066667+00:00

Good Day,

Looks child tag is not related to my questions, but i could not find a tag for Azure Ai Foundry

Anyway, I've tried to deploy the new Nvidia LLM model, and got unknown error

https://husazaifoundryproject-oxjiz.eastus.inference.ml.azure.com/v1/chat/completions

Deployment info

Namedeepseek-r1-distill-llama-8b--1

Provisioning state__Failed

Last updated onMar 20, 2025 3:47 AM

Created byMohamed Hussein

Created onMar 20, 2025 3:47 AM

Traffic allocation0%

Instance count1

Compute typeTemporary - 6d 23h 52m left

SKUStandard_NC96ads_A100_v4

Azure OpenAI in Foundry Models

1 answer

Sort by: Most helpful
  1. SRILAKSHMI C 19,825 Reputation points Microsoft External Staff Moderator
    2025-03-20T08:29:11.9733333+00:00

    Hello Mohamed Hussein,

    I understand that your deployment of the Nvidia LLM model (deepseek-r1-distill-llama-8b--1) on Azure AI Foundry has failed with an unknown error.

    Could you please confirm whether you are using shared compute for deployment or if the deployment is consuming quota from your subscription?

    I addition to that here are the recommended troubleshooting steps to consider,

    Validate Model Name - Ensure that the model's name and deployment parameters are correctly specified in the configuration.

    Check Resource Availability - Verify that the selected SKU (Standard_NC96ads_A100_v4) has sufficient available resources in your Azure region. Check your Azure subscription quotas to ensure you haven’t reached usage limits.

    Network Connectivity - Confirm there are no network issues that might be preventing the deployment process. If using private endpoints, verify that outbound traffic rules allow access to required Azure services.

    Review Deployment Logs - Check the deployment logs in Azure AI Foundry for any detailed error messages.

    Review Azure Machine Learning logs and Azure Monitor for further insights into potential failures.

    Make sure that all necessary libraries and dependencies for running the model are installed.

    If using a custom environment, validate the conda.yaml or requirements.txt file.

    Model Compatibility - Ensure that the Nvidia LLM model you are deploying is compatible with the selected compute instance and environment.

    Deployment Configuration - Verify that all configuration settings for your deployment are correct.

    Check the model path, environment settings, and dependency configurations to ensure proper setup.

    I hope this information helps.

    Kindly consider upvoting the comment if the information provided is helpful. This can assist other community members in resolving similar issues. 

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.