Edit

Known issues with H-series and N-series virtual machines

Applies to: ✔️ Linux VMs ✔️ Windows VMs ✔️ Flexible scale sets ✔️ Uniform scale sets

This article attempts to list recent common issues and their solutions when using the HB-series and N-series HPC and GPU VMs.

AMD NVv5 V710 and NGv1 V620 Linux GPU drivers NULL Pointer Dereference (CVE-2026-43603) might cause denial-of-service

A vulnerability in the AMD Linux GPU driver (CVE-2026-43603) could allow a local user to trigger a kernel crash and denial of service through a NULL pointer dereference. Azure plans to release updated Linux guest drivers for NVv5 V710 and NGv1 V620 by September 28, 2026. For more information, see the AMD Security Bulletin SB-6034 Linux GPU Driver NULL Pointer Dereference.

NVIDIA-SMI Not Showing Full Telemetry on NCv6 (RTX Pro 6000) virtual machines

Running the nvidia-smi will not show full telemetry of the RTX Pro 6000 Blackwell GPU(s) on a NCv6-series virtual machine. Specifically, power and utilization statistics will not be exposed. This is due to the use of an SRIOV-based exposure of the GPU(s) to the virtual machine, as opposed to passthrough mode. This is a known limitation of NVIDIA's SRIOV driver supporting "vGPU" functionality.

InfiniBand RDMA and NUMA Node Affinity on HBv5 virtual machines

On certain HBv5 virtual machines (VMs), the InfiniBand RDMA device names (such as mlx5_[0-3]) may not align correctly with their respective NUMA node affinities. Ideally, each RDMA device should be mapped as follows:

  • mlx5_0 is on NUMA node: 0
  • mlx5_1 is on NUMA node: 4
  • mlx5_2 is on NUMA node: 8
  • mlx5_3 is on NUMA node: 12

However, an incorrect mapping example could be:

  • mlx5_0 is on NUMA node: 4
  • mlx5_1 is on NUMA node: 8
  • mlx5_2 is on NUMA node: 12
  • mlx5_3 is on NUMA node: 0

This misalignment can lead to performance degradation, particularly when running multinode MPI workloads.

To confirm whether your RDMA devices are correctly mapped to NUMA nodes, execute the following script:

for d in /sys/class/infiniband/*;
do
dev=$(basename "$d")
node=$(cat "$d/device/numa_node")
echo "$dev is on NUMA node: $node"
done

Compare the output with the ideal mapping listed above.

Solution: Persistent device naming with Udev rules

To remediate the misalignment issue, follow these steps:

  1. Create a new file in /etc/udev/rules.d/, for example: 99-rdma-persistent-naming.rules
  2. Add the following lines to the file:
    ACTION=="add", SUBSYSTEMS=="pci", KERNELS=="0101:00:00.0", PROGRAM="rdma_rename %k NAME_FIXED mlx5_ib0"
    ACTION=="add", SUBSYSTEMS=="pci", KERNELS=="0102:00:00.0", PROGRAM="rdma_rename %k NAME_FIXED mlx5_ib1"
    ACTION=="add", SUBSYSTEMS=="pci", KERNELS=="0103:00:00.0", PROGRAM="rdma_rename %k NAME_FIXED mlx5_ib2"
    ACTION=="add", SUBSYSTEMS=="pci", KERNELS=="0104:00:00.0", PROGRAM="rdma_rename %k NAME_FIXED mlx5_ib3"
    
  3. Reload udev rules and trigger device events:
    # udevadm control --reload
    # udevadm trigger --type=devices --action=add
    

This solution ensures that RDMA device naming persists across VM reboots.

Cache topology on Standard_HB120rs_v3

lstopo displays incorrect cache topology on the Standard_HB120rs_v3 VM size. It may display that there’s only 32 MB L3 per nonuniform memory access (NUMA) node. However, in practice, there's indeed 120 MB L3 per NUMA as expected since the same 480 MB of L3 to the entire VM is available as with the other constrained-core HBv3 VM sizes. This incorrect display is a cosmetic error and shouldn't affect workloads.

Access Restriction on queue pair 0

To prevent low-level hardware access that can result in security vulnerabilities, Queue Pair 0 isn't accessible to guest VMs. This restriction should only affect actions typically associated with administration of the ConnectX InfiniBand network interface card (NIC) and running some InfiniBand diagnostics like ibdiagnet, but not end-user applications.

Accelerated Networking on InfiniBand-equipped virtual machines

Azure Accelerated Networking allows enhanced throughput and latencies over the Azure Ethernet network. Though Accelerated Networking is separate from the InfiniBand network, its use may affect behavior of certain MPI implementations when running jobs over InfiniBand. Specifically, the InfiniBand interface on some VMs may have a slightly different name (mlx5_1 as opposed to earlier mlx5_0). This issue may require tweaking of the MPI command lines, especially when using the UCX interface (commonly with OpenMPI and HPC-X).

For more information on this issue, see the TechCommunity article with instructions on how to address any observed issues.

Cache Cleaning

On HPC systems, it's often useful to clean up memory after a job finishes before the next user is assigned the same node. After running applications in Linux, you may find that your available memory reduces while your buffer memory increases, despite not running any applications.

Screenshot of command prompt before cleaning

Using numactl -H shows which NUMAnodes the memory is buffered with (possibly all). In Linux, users can clean the caches in three ways to return buffered or cached memory to ‘free’. You need to be root or have sudo permissions.

sudo echo 1 > /proc/sys/vm/drop_caches [frees page-cache]
sudo echo 2 > /proc/sys/vm/drop_caches [frees slab objects e.g. dentries, inodes]
sudo echo 3 > /proc/sys/vm/drop_caches [cleans page-cache and slab objects]

Screenshot of command prompt after cleaning