Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
This article helps you diagnose and resolve problems with Microsoft Discovery when you don't have a specific error code to look up. It covers deployment, supercomputers, workspaces, networking, bookshelves and knowledge bases, and tools and agents.
Start with the general troubleshooting checklist. Then go to the section that matches your symptom for likely causes and solutions. If you already have a specific error code or message, see Microsoft Discovery error codes. Before you open a support request, also review Known issues for Microsoft Discovery.
Prerequisites
- An Azure account with an active subscription. Create an account for free.
- A Microsoft Discovery environment that's deployed or being deployed.
- The required Discovery role assignments. See Role assignments for Microsoft Discovery.
- Access to the Activity Log and to Log Analytics to query platform logs. See Query supercomputer logs and Query workspace logs.
Troubleshooting checklist
Work through these steps first. They identify the cause of most problems and provide the evidence you need for the sections that follow.
Capture the correlation ID and read the operation history
Get the correlation ID for the failed operation and read the control-plane operations. The summary error is often generic, but the operation history shows the first operation that actually failed. See Get operation correlation ID from Activity Log.
Check the resource provisioning state
Confirm the resource's provisioningState. Succeeded means the operation completed, Accepted usually means the operation is still in progress, and Failed means you must inspect the operation history before you retry. Some resource types must be deleted and recreated after they enter a terminal Failed state.
Verify the deploying identity's roles
Confirm the required roles are assigned to the identity that runs the operation. For a pipeline, check the service principal or managed identity rather than the interactive user. Azure ownership permissions don't replace the Discovery data-plane and supporting service roles described in Role assignments in Microsoft Discovery.
Confirm the region and API version
Confirm the target region is supported and the request uses the API version that ships with the current Toolbox or infrastructure quickstart.
Query the relevant logs
Query supercomputer, workspace, and bookshelf logs to locate the failing phase. See Query supercomputer logs, Query workspace logs, and Query bookshelf logs.
Potential quick workarounds
These workarounds can restore service quickly. For a permanent fix, use the cause and solution sections that follow.
Retry a transient failure
- If an operation fails with an
InternalServerError, a gateway timeout, or other kind of timeout, wait a few minutes. - Check the operation history for a terminal failure before you retry. Don't delete partially created resources unless the resource-specific guidance requires recreation.
Pin a job to a known-good node pool
- In the project preference, set the correct node pool explicitly (
nodePoolId). - Resubmit the job or tool run.
Deployment fails because the identity is missing required roles
Deployment reaches an operation that needs a data-plane, network-perimeter, storage, or Foundry permission that the deploying identity doesn't have.
Solution: Assign the full role set to the deploying identity
- Assign the platform administrator persona roles to the identity that runs the deployment. Include the network security perimeter, storage data-plane, and Foundry roles required for the resources in your deployment.
- Preview the assignments with the
Set-DiscoveryRoleAssignments.ps1script and its-WhatIfparameter, and then run the script without-WhatIf. See Assign persona roles with a PowerShell script. - Wait several minutes for role and network-perimeter propagation.
- Rerun the deployment. See Role assignments for Microsoft Discovery.
Deployment fails or is blocked in a specific region
A region can be restricted or withdrawn for Discovery managed resources even when it appears selectable. Resource-group location policies can also block where you create virtual networks or resource groups. Capacity or quota might also contribute to deployment failure.
Solution: Deploy to a supported region and add policy exemptions
- Deploy to a supported region.
- If a required role is eligible through Privileged Identity Management (PIM), activate it before deploying.
- Create the resource group in the correct region. If an organizational policy blocks the deployment, ask the policy administrator whether an exemption is permitted at the required scope.
For a region that appears selectable but isn't supported, see Known issues for Microsoft Discovery.
Deployment appears stuck and never finishes
The deployment fails internally but keeps retrying components, so it looks like it's still running.
Solution: Find and fix the first internal failure
- Open the deployment operation history and find the first operation that failed.
- Resolve that root cause, which is commonly a missing role, a network feature, or a subnet or quota limit.
- Rerun the deployment.
A Discovery resource is stuck in a terminal failed state
Some Discovery resources can't be updated after they enter a terminal Failed state, so further operations return a conflict.
Solution: Delete and recreate the failed resource
- Confirm from the operation history that the resource is in a terminal
Failedstate and can't be retried. - Delete only the failed resource. For a workspace, leave the resource group, storage, and supercomputer in place.
- Recreate the resource with a new name and the correct prerequisites.
- To free a supercomputer that's bound to a stuck workspace, remove the supercomputer association from within the workspace instead of deleting the supercomputer. See Manage Microsoft Discovery workspaces and Delete Microsoft Discovery resources.
Supercomputer creation fails with a webhook timeout
During the internal Helm install of the supercomputer, the kueue-webhook-service in the kueue-system namespace returns a 504 Gateway Timeout to the network-operator pre-upgrade hook, and the install fails. This condition is transient.
Solution: Retry creation and confirm the latest release
- Retry supercomputer creation.
- Confirm you're on the latest platform release.
- If it persists, inspect the pod behind
kueue-webhook-servicein thekueue-systemnamespace. See Query supercomputer logs.
The managed cluster fails during node-pool provisioning
The Activity Log shows failures during managed cluster provisioning because the supercomputer subnet is too small for node-pool scaling.
Solution: Increase the subnet size and redeploy
- Increase the AKS and supercomputer subnet to at least
/24. - Redeploy the supercomputer.
- Validate that node-pool creation succeeds.
Supercomputer jobs terminate unexpectedly
AKS node OS upgrades can evict or kill running jobs and pods, so a subset of long-running jobs terminate, sometimes without surfaced logs. Node or GPU driver health issues can also cause a "node became not ready" failure.
Solution: Use a resilient release and pin the node pool
- Confirm you're on a release with job and node resiliency improvements.
- Set the correct
nodePoolIdin the project preference. - If a node or GPU node pool is unhealthy, ensure the latest node image and recreate the node pool.
- Review the affected run in the supercomputer logs. See Debug task execution.
Tool runs fail because resources aren't allowlisted
A storage account or supercomputer exists but isn't associated with the workspace, so tool runs that reference it are rejected with a "Resources not allowed" error.
Solution: Allowlist the storage account and supercomputer in the workspace
- Associate (allowlist) the storage account and the supercomputer in the workspace.
- Rerun the tool. See Manage Microsoft Discovery workspaces.
Workspace creation fails because of insufficient quota
The workspace's Azure Container Apps environment requires enough quota to create. For the Enterprise SKU, the environment needs at least 80 cores.
Solution: Increase the container environment quota
- Request or confirm at least 80 cores of Azure Container Apps quota in the target region.
- Retry workspace creation. See Quota and reservations.
The workspace host doesn't resolve over a private endpoint
Private DNS isn't resolving the workspace data-plane host from inside the virtual network after you enable network isolation or disable public access, so investigations are unreachable over Private Link.
Solution: Fix private DNS resolution
Verify the private DNS zone and A-records for the data-plane host exist.
Confirm the private endpoint is approved and linked to the correct virtual network and subnet.
Validate resolution from a host inside the virtual network:
nslookup <workspace-host>.workspace.discovery.azure.comApply the same checks to a bring-your-own Cosmos DB private endpoint when project creation depends on it. See Network security for Microsoft Discovery.
Network security group rules block deployment and usage
DenyAllInbound and DenyAllOutbound rules applied to every subnet block the traffic Discovery needs, so deployment fails and, even after deployment, chat or agent operations fail.
Solution: Scope network security group rules to the correct subnets
- Attach the required allow rules to the correct network security group for the supercomputer subnets. See Plan network security groups for a Microsoft Discovery supercomputer.
- Remove blanket deny rules that apply to every subnet.
- If agent operations still appear blocked by rules, create an Azure support request with the affected subnet and network security group details.
Supercomputer deployment stalls because a required public IP feature isn't registered
For a network configuration that uses a public IP address, the supercomputer deployment can get stuck when the bring-your-own public IP network feature isn't registered on the subscription. This condition causes an internal failure that the deployment keeps retrying. If Azure Policy blocks public IP creation, see Known issues for Microsoft Discovery instead.
Solution: Register the network feature and redeploy
Run the following commands to register the feature and re-register the provider:
az feature register --namespace Microsoft.Network --name AllowBringYourOwnPublicIpAddress az provider register -n Microsoft.NetworkRun the following command to confirm that the feature state is
Registered, and then redeploy:az feature show --namespace Microsoft.Network --name AllowBringYourOwnPublicIpAddress -o table
Knowledge base creation fails validation
The data plane rejects an invalid version string or bookshelf name. This rejection surfaces as a chunked-encoding or 400 error on the version request. A bookshelf that's still deploying shows provisioningState as Accepted rather than Succeeded.
Solution: Use a valid name and version, and wait for provisioning
- Use a simple integer version number, such as
1, instead of a longer string such as1.0. - Ensure the bookshelf name uses lowercase letters and dashes only.
- Wait for
provisioningStateto reachSucceededbefore you continue. See Index a bookshelf knowledge base.
Bookshelf indexing fails or silently drops documents
Indexing can fail on a very large single document or from a transient database timeout. In some cases, a large document is dropped from the index even though indexing reports success.
Solution: Check indexing logs and split large documents
- Check the bookshelf indexing logs for the affected document. See Query bookshelf indexing logs.
- Split very large PDFs, and confirm document sizes against the current ingest limits.
- Confirm sufficient embedding and retrieval model quota in the region.
- Ensure you're on a platform release with ingest improvements, and then retry.
Project or agent creation fails with a 500 error
Agent creation in the backing Foundry resource returns a 500 error, often because the project managed identity is missing a Foundry role or because a bring-your-own Cosmos DB private endpoint can't be reached.
Solution: Verify Foundry roles and Cosmos DB connectivity
- Verify the Foundry User role is assigned to the project managed identity.
- Validate the bring-your-own Cosmos DB private endpoint connectivity and DNS.
- If project deletion is blocked for a project already in a failed state, use platform cleanup actions, and then redeploy. See Agent creation.
Tool execution fails with an invalid message id
An invalid or mismatched message identifier passed from the agent to the tool prevents tool invocation, so prompts that use tools can't complete.
Solution: Update to a fixed release and collect definitions
- Ensure you're on a release that includes the message ID fix and diagnostic logging.
- Apply any recommended prompt-level improvements.
- Collect the agent, workflow, and tool definitions to reproduce the issue. See Debug task execution.
Advanced troubleshooting and data collection
If you create an Azure support request, collect the following information so support can investigate quickly:
- The correlation ID and the deployment operation list for the failed operation.
- The Activity Log entries and relevant supercomputer, workspace, or bookshelf logs.
- The role assignments on the deploying identity or project managed identity.
- The region, API version, and Toolbox or template version.
- For networking issues, the
nslookupoutput from inside the virtual network, the private endpoint and DNS configuration, and the network security group rules.