SAS Workload Management can be configured to automatically scale the number of nodes in the Kubernetes cluster to intelligently meet the demands of the current or anticipated workload. Depending on how you set your configuration parameters, you can configure at least three different ways to scale.
Figure 1. Conceptual view of how SAS Workload Management can trigger Kubernetes node autoscaling through pending jobs, time-based minimum capacity, and dynamic resource thresholds.
Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.
Let’s review all three, starting with the initial integration with the Kubernetes Cluster Autoscaler available since SAS Viya 2023.05.
At a high level, on every job scheduling cycle, SAS Workload Orchestrator scans each job which is in a pending state and uses the job's queue definition to check for nodes with sufficient resources to run the job.
First, it looks for regular, non-scalable host types. If none are available, SAS Workload Orchestrator checks for scalable host types. If no hosts are found to run the job across all host types associated with its queue, then SAS Workload Orchestrator starts a scaling request to force the Kubernetes Cluster Autoscaler to provision a new node.
Administrators can configure a set of parameters to tune how aggressive this behavior is. The ones I find the most useful are:
This autoscaling behavior is reactive because it cannot anticipate workload demands and reacts to resource exhaustion.
Starting with the SAS Viya 2026.01 release, for any scalable host type defined in SAS Workload Orchestrator, you can set a minimum number of nodes that must be available for that host type. This can be useful to support interactive users. In cloud environments, nodes can take a few minutes to start and become ready to host compute jobs. You do not want your SAS Data and AI Studio programmers to face timeouts when they log on in the morning, waiting while the first node of a scalable host type comes online. Defining at least one or two minimum nodes makes them immediately available to your users, while still leveraging autoscaling to expand the node count as demand grows. At the same time, we all know that cloud resources come with a price tag, and it would be a waste of money to leave that minimum number of nodes up and running 24x7. You can couple this newer setting with the existing capability to specify start and end times for time-based configuration settings. For example, you can set a host type to require a minimum number of nodes only during peak hours, then go down to zero during nights and weekends. The result is more predictable performance when you need it most, while preserving cost efficiency during off‑peak hours.
This new capability introduces proactive fixed configuration; you can predefine when to scale up your nodes based on predicted workloads, anticipating user and job demands.
Starting with the SAS Viya 2026.07 release, you can define an even more advanced autoscaling behavior. Dynamic resource threshold-based scaling uses the average dynamic resource usage across host types to scale nodes. Unlike scaling based on pending jobs or a minimum node count, it responds proactively to real-time pressure from CPU, memory, disk, I/O, and other dynamic resources, helping you effectively handle diverse and unpredictable workloads. You can also define custom resources to monitor, which can be used for scaling thresholds. See Define User-Defined Resources in SAS Environment Manager: User’s Guide.
This capability moves proactive configuration a step forward. You do not need a pending job to trigger scaling. SAS Workload Management monitors resource utilization in near real time and dynamically prepares the right number of nodes so that they are already there when additional jobs need a host to run. Dynamic scaling settings anticipate computing workload demands even for unpredictable use cases.
To avoid scaling for short-lived spikes, you can define how long a measured resource must remain above the threshold before scaling occurs. By default, this wait time is 60 seconds, so scaling is driven by sustained pressure rather than transient activity.
This new, dynamic capability works alongside the time‑based minimum node count so a baseline capacity can always be enforced when you need it.
For more information, see Scale Thresholds in SAS Viya Platform: Workload Management.
So far, we have covered different ways to configure when to start an autoscaling request. SAS Viya 2026.05 introduced an additional feature to influence how to drive scaling. Queue-based scaling can now trigger a multiple-node scaling event, which starts as many nodes as required based on the number of pending jobs. Before this, each scaling event could only start one new node, followed by the default 300-second wait time before a second scaling request for another node went in. When you have a large burst of jobs all coming in at the same time, this behavior cannot keep pace with the workload.
Now, to accommodate pending jobs, the number of nodes to be scaled up is calculated based on the job limits of the queue and host type, and the resources requested in the pending jobs.
Let’s say, for example, that 20 batch jobs are submitted at once to a queue configured to run, at most, 5 jobs per host on a scalable host type. If the existing nodes remain full and unavailable for more than the configured waiting time, then SAS Workload Management will trigger a scaling request for enough new nodes, so that all 20 jobs can be started as soon as possible.
Figure 2. Multi-node scaling reduces wait time for a job burst by requesting multiple nodes in a single event instead of waiting for one node per scaling request.
The dynamic resource threshold-based scaling capability has been designed to support multi-node scaling, too; an administrator can define the number of nodes to scale per request (the default value is 1). After the average of the monitored resource exceeds the defined threshold continuously for the defined wait time, SAS Workload Orchestrator starts the configured number of nodes in a single event.
SAS Workload Orchestrator sends scaling requests to the Kubernetes Cluster Autoscaler without any limit on the maximum number of nodes that can be active. Capping the maximum number of nodes, to control cloud spending, can be enforced with proper node pool configuration by the Kubernetes administrator within the cloud provider interface.
Similarly, after jobs complete and resource utilization drops, the Kubernetes Cluster Autoscaler removes unused nodes according to its own configuration and heuristics. To support this process, SAS Workload Orchestrator tries to place new jobs on a small number of nodes instead of spreading them across multiple empty nodes, allowing the Cluster Autoscaler to shut down as many nodes as possible.
Finally, as already discussed in previous articles, defining the proper Kubernetes ClusterRole and ClusterRoleBinding is mandatory to give SAS Workload Orchestrator the proper permissions to collect the deep telemetry required for dynamic threshold triggers.
Together, these autoscaling enhancements help SAS Workload Management provide capacity more predictably, respond faster to changing demand, and use cloud resources more efficiently.
Rather than relying only on queued jobs to trigger growth, administrators can combine baseline capacity, real-time resource monitoring, and multi-node scaling events to match compute availability more closely to actual and predicted workload needs.
For more information about this topic and other aspects of SAS Workload Orchestrator in action, visit learn.sas.com to view the Architecture and Administration for SAS® Workload Management on SAS® Viya® workshop.
Find more articles from SAS Global Enablement and Learning here.
Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.
The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.