Core component of SQL Server for storing, processing, and securing data
SQL Server: Unexpected CPU Spike on Primary Replica — Troubleshooting Assistance Needed
Problem description
I am experiencing an issue with my SQL Server virtual machine where the primary replica of an Always On availability group in the UK South region is experiencing unexpected CPU utilization spikes up to 100%. This occurs under a workload that runs at approximately 30% CPU on the secondary replica. The spike was first observed on September 3, 2026, at 16:00 UTC, detected via external monitoring tools. The issue is not currently active, and no error codes have been generated. I am seeking assistance to diagnose and resolve the root cause of these CPU spikes.
Environment
A Windows Server–based virtual machine running SQL Server in the UK South region, deployed in an availability zone, configured as part of an Always On availability group with at least one primary and one secondary replica. Azure VM agent and platform metrics are enabled.
What I've already tried
I have reviewed monitoring data confirming that the sqlservr process reached 100% CPU on the primary replica while the secondary remained near 30% under the same workload. I declined to collect additional diagnostics when prompted. An initial support engineer attempted to retrieve platform data but reported 'No data found.' Automated diagnostics flagged high-CPU instances over the past 24 hours, with average max CPU above 85% in ten-minute windows and peak 30-minute average CPU about 99.8%. I have also reviewed platform diagnostic insights regarding VM resizing constraints and performed basic troubleshooting steps such as monitoring VM-level metrics and reviewing configuration settings.
Current status
I am seeking guidance from an engineer to diagnose the cause of the CPU spikes on the primary replica and to identify potential solutions or configuration adjustments to prevent recurrence. I am prepared to provide additional diagnostic data, including SQL Server audit logs, Extended Events traces, process information during the incident, and performance metrics as needed.