Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
You can configure the cluster autoscaler on your Azure Red Hat OpenShift with hosted control planes cluster to automatically add and remove nodes across all autoscaling-enabled node pools. Cluster-wide autoscaling helps keep your applications schedulable during demand spikes and reduces costs by removing underutilized nodes during low-demand periods.
Prerequisites
- An existing Azure Red Hat OpenShift with hosted control planes cluster. For more information, see Create an Azure Red Hat OpenShift with hosted control planes cluster.
- Azure CLI version 2.67.0 or later is installed and authenticated.
Run
az --versionto check your installed version. For more information, see Install Azure CLI.
- The ARO HCP CLI extension is installed. If you need to install it, download the wheel file for the ARO HCP CLI extension.
Understand cluster autoscaling
The cluster autoscaler monitors your cluster for pods that can't be scheduled because of insufficient resources and adjusts the number of nodes accordingly.
How the cluster autoscaler scales nodes
The cluster autoscaler increases the number of nodes in your cluster when:
- Pods are pending because they can't be scheduled on any current worker nodes due to insufficient resources.
- Another node is necessary to meet deployment needs.
The cluster autoscaler doesn't increase the cluster beyond the maximum limits that you set.
For example, if you set the maximum number of nodes on a cluster to be 50, with two autoscaling node pools and one manually scaled node pool, the cluster autoscaler can provision up to 50 nodes across the two autoscaling node pools combined. Nodes in the manually scaled node pool don't count toward that limit.
The cluster autoscaler also removes nodes that are underutilized. Every 10 seconds, it checks whether a node's resource use is below the defined threshold (default is 50%) and whether all pods on the node can be moved to other nodes. For the full list of conditions that prevent node removal, see the Kubernetes Cluster Autoscaler FAQ.
How pod priority affects cluster autoscaling
The podPriorityThreshold property controls which pods are important enough to justify adding new nodes.
- Pods below the threshold are treated as best-effort. The cluster autoscaler doesn't add new nodes to run these pods. They stay pending until spare resources become available on existing nodes.
- Pods at or above the threshold are treated as standard pods. If these pods are pending because of insufficient resources, the cluster autoscaler adds new nodes.
The default podPriorityThreshold is -10, which maps to the Kubernetes expendable-pods-priority-cutoff concept.
Cluster-wide autoscaling compared to per-node-pool autoscaling
Azure Red Hat OpenShift with hosted control planes provides two levels of autoscaling configuration:
- Cluster-wide autoscaling governs the overall behavior of the autoscaler, such as the total number of nodes across all node pools and whether low-priority pods trigger scale-up.
- Per-node-pool autoscaling sets the boundaries for how many nodes an individual node pool can scale to.
Both levels work together: the cluster autoscaler respects both the per-pool min and max bounds and the cluster-wide total number of nodes.
For more information about configuring per-node-pool autoscaling, see Scale and configure node pools in Azure Red Hat OpenShift with hosted control planes.
Configure cluster autoscaling
You can configure the cluster to automatically add and remove nodes across all autoscaling-enabled node pools.
Set the following environment variables if you haven't already:
CUSTOMER_RG_NAME="<resource-group-name>" CLUSTER_NAME="<cluster-name>"Run the following command to retrieve the current autoscaling configuration for your cluster:
az aro hcp cluster show \ --name "${CLUSTER_NAME}" \ --resource-group "${CUSTOMER_RG_NAME}" \ --query "properties.autoscaling"Example output:
{ "maxNodesTotal": 0, "maxPodGracePeriodSeconds": 600, "maxNodeProvisionTimeSeconds": 900, "podPriorityThreshold": -10 }Update the autoscaling configuration by running the command.
Set--max-nodes-totalto a value that reflects your expected peak capacity plus extra room.
For example, if your autoscaling node pools total 30 nodes at peak, a value of 40-50 provides room for burst scaling while capping costs. For a description of each argument including defaults and valid ranges, see Cluster autoscaling arguments reference.Important
The default
--max-nodes-totalvalue of0means there's no limit on the total number of nodes the autoscaler can provision across all autoscaling-enabled node pools. To control costs, set--max-nodes-totalto an explicit value that reflects your expected maximum capacity.az aro hcp cluster update \ --name "${CLUSTER_NAME}" \ --resource-group "${CUSTOMER_RG_NAME}" \ --max-nodes-total 50You can set multiple autoscaling arguments in a single command:
az aro hcp cluster update \ --name "${CLUSTER_NAME}" \ --resource-group "${CUSTOMER_RG_NAME}" \ --max-nodes-total 50 \ --max-pod-grace-period 600 \ --max-node-provision-time 900 \ --pod-priority-threshold -10Verify that the autoscaling configuration was applied correctly. Run the same command that you used to view the current configuration:
az aro hcp cluster show \ --name "${CLUSTER_NAME}" \ --resource-group "${CUSTOMER_RG_NAME}" \ --query "properties.autoscaling"Example output:
{ "maxNodesTotal": 50, "maxPodGracePeriodSeconds": 600, "maxNodeProvisionTimeSeconds": 900, "podPriorityThreshold": -10 }
Cluster autoscaling arguments reference
The following table describes the cluster-wide autoscaling arguments for the az aro hcp cluster update command.
All arguments are optional. Set them when you create the cluster or update them after creation.
| Argument | Default | Valid range | Description |
|---|---|---|---|
--max-nodes-total |
0 (no limit) |
Minimum: 0 |
Maximum number of nodes the cluster autoscaler can provision across all autoscaling-enabled node pools. A value of 0 means no maximum limit. Manually scaled node pools don't count toward this total. |
--max-pod-grace-period |
600 |
Minimum: 0 |
Grace period in seconds for pod termination before a node is scaled down. The default is 600 seconds (10 minutes). |
--max-node-provision-time |
900 |
Minimum: 0, Maximum: 3600 |
Maximum time in seconds the cluster autoscaler waits for a node to be provisioned before considering the provisioning unsuccessful. The default is 900 seconds (15 minutes). The maximum configurable value is 3600 seconds (60 minutes). |
--pod-priority-threshold |
-10 |
No constraint | Pods with a priority value below this threshold are treated as best-effort and don't trigger node scale-up. Pods at or above this value are treated as standard pods and trigger scale-up when pending. The default is -10. This value maps to the Kubernetes PriorityClass system. |
Note
These arguments are distinct from per-node-pool autoscaling arguments (--min-replicas and --max-replicas), which you configure on individual node pool resources by using az aro hcp cluster nodepool update.
For more information, see Cluster-wide autoscaling compared to per-node-pool autoscaling.