BookmarkSubscribeRSS Feed

SAS Workload Management advanced Kubernetes node autoscaling

Started Wednesday by
Modified Wednesday by
Views 70

SAS Workload Management can be configured to automatically scale the number of nodes in the Kubernetes cluster to intelligently meet the demands of the current or anticipated workload. Depending on how you set your configuration parameters, you can configure at least three different ways to scale.

 

  1. Pending jobs (no more nodes available): reactive.
  2. Minimum node count (time-based): proactive, fixed.
  3. Resource threshold monitoring: proactive, dynamic.

 

01_ER_20260922_01_SWO_Advanced_Scaling.png

Figure 1. Conceptual view of how SAS Workload Management can trigger Kubernetes node autoscaling through pending jobs, time-based minimum capacity, and dynamic resource thresholds.

 

 

Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.

 

 

Let’s review all three, starting with the initial integration with the Kubernetes Cluster Autoscaler available since SAS Viya 2023.05.

 

 

Pending jobs

 

At a high level, on every job scheduling cycle, SAS Workload Orchestrator scans each job which is in a pending state and uses the job's queue definition to check for nodes with sufficient resources to run the job.

 

First, it looks for regular, non-scalable host types. If none are available, SAS Workload Orchestrator checks for scalable host types. If no hosts are found to run the job across all host types associated with its queue, then SAS Workload Orchestrator starts a scaling request to force the Kubernetes Cluster Autoscaler to provision a new node.

 

Administrators can configure a set of parameters to tune how aggressive this behavior is. The ones I find the most useful are:

 

  • The number of seconds that the host manager waits before honoring a new scaling request for the same host type. How long do you want to wait between submitting consecutive scaling requests? Cloud environments might take a few minutes to provision a new node, and then it might be a few more before the node can accept a job. If you set the value too low, you risk triggering a runaway number of new Kubernetes node requests before the first node even becomes ready. (default = 300 sec)
  • The minimum amount of time that each job is in a pending state before it is considered to trigger the scaling request. Do you want to request a scaling event as soon as a job is pending, or maybe wait 15 seconds to see if any running jobs terminate and make node resources available? (default = 0, i.e. no waiting)
  • The minimum number of jobs pending in a queue before triggering a scaling event. Do you want to provision a new node as soon as a single job is unable to run, or do you want to wait for a few jobs to accumulate in the queue before asking for a new node? (default = 0, i.e. no additional jobs have to be pending)

 

This autoscaling behavior is reactive because it cannot anticipate workload demands and reacts to resource exhaustion.

 

 

Time-based minimum node count

 

Starting with the SAS Viya 2026.01 release, for any scalable host type defined in SAS Workload Orchestrator, you can set a minimum number of nodes that must be available for that host type. This can be useful to support interactive users. In cloud environments, nodes can take a few minutes to start and become ready to host compute jobs. You do not want your SAS Data and AI Studio programmers to face timeouts when they log on in the morning, waiting while the first node of a scalable host type comes online. Defining at least one or two minimum nodes makes them immediately available to your users, while still leveraging autoscaling to expand the node count as demand grows. At the same time, we all know that cloud resources come with a price tag, and it would be a waste of money to leave that minimum number of nodes up and running 24x7. You can couple this newer setting with the existing capability to specify start and end times for time-based configuration settings. For example, you can set a host type to require a minimum number of nodes only during peak hours, then go down to zero during nights and weekends. The result is more predictable performance when you need it most, while preserving cost efficiency during off‑peak hours.

 

This new capability introduces proactive fixed configuration; you can predefine when to scale up your nodes based on predicted workloads, anticipating user and job demands.

 

 

Resource threshold monitoring

 

Starting with the SAS Viya 2026.07 release, you can define an even more advanced autoscaling behavior. Dynamic resource threshold-based scaling uses the average dynamic resource usage across host types to scale nodes. Unlike scaling based on pending jobs or a minimum node count, it responds proactively to real-time pressure from CPU, memory, disk, I/O, and other dynamic resources, helping you effectively handle diverse and unpredictable workloads. You can also define custom resources to monitor, which can be used for scaling thresholds. See Define User-Defined Resources in SAS Environment Manager: User’s Guide.

 

This capability moves proactive configuration a step forward. You do not need a pending job to trigger scaling. SAS Workload Management monitors resource utilization in near real time and dynamically prepares the right number of nodes so that they are already there when additional jobs need a host to run. Dynamic scaling settings anticipate computing workload demands even for unpredictable use cases.

 

To avoid scaling for short-lived spikes, you can define how long a measured resource must remain above the threshold before scaling occurs. By default, this wait time is 60 seconds, so scaling is driven by sustained pressure rather than transient activity.

 

This new, dynamic capability works alongside the time‑based minimum node count so a baseline capacity can always be enforced when you need it.

 

For more information, see Scale Thresholds in SAS Viya Platform: Workload Management.

 

 

Multiple-Node scaling

 

So far, we have covered different ways to configure when to start an autoscaling request. SAS Viya 2026.05 introduced an additional feature to influence how to drive scaling. Queue-based scaling can now trigger a multiple-node scaling event, which starts as many nodes as required based on the number of pending jobs. Before this, each scaling event could only start one new node, followed by the default 300-second wait time before a second scaling request for another node went in. When you have a large burst of jobs all coming in at the same time, this behavior cannot keep pace with the workload.

 

Now, to accommodate pending jobs, the number of nodes to be scaled up is calculated based on the job limits of the queue and host type, and the resources requested in the pending jobs.

 

Let’s say, for example, that 20 batch jobs are submitted at once to a queue configured to run, at most, 5 jobs per host on a scalable host type. If the existing nodes remain full and unavailable for more than the configured waiting time, then SAS Workload Management will trigger a scaling request for enough new nodes, so that all 20 jobs can be started as soon as possible.

 

02_ER_20260922_02_SWO_MultiNode_Scaling.png

Figure 2. Multi-node scaling reduces wait time for a job burst by requesting multiple nodes in a single event instead of waiting for one node per scaling request.

 

The dynamic resource threshold-based scaling capability has been designed to support multi-node scaling, too; an administrator can define the number of nodes to scale per request (the default value is 1). After the average of the monitored resource exceeds the defined threshold continuously for the defined wait time, SAS Workload Orchestrator starts the configured number of nodes in a single event.

 

 

Additional considerations

 

SAS Workload Orchestrator sends scaling requests to the Kubernetes Cluster Autoscaler without any limit on the maximum number of nodes that can be active. Capping the maximum number of nodes, to control cloud spending, can be enforced with proper node pool configuration by the Kubernetes administrator within the cloud provider interface.

 

Similarly, after jobs complete and resource utilization drops, the Kubernetes Cluster Autoscaler removes unused nodes according to its own configuration and heuristics. To support this process, SAS Workload Orchestrator tries to place new jobs on a small number of nodes instead of spreading them across multiple empty nodes, allowing the Cluster Autoscaler to shut down as many nodes as possible.

 

Finally, as already discussed in previous articles, defining the proper Kubernetes ClusterRole and ClusterRoleBinding is mandatory to give SAS Workload Orchestrator the proper permissions to collect the deep telemetry required for dynamic threshold triggers.

 

 

Closing

 

Together, these autoscaling enhancements help SAS Workload Management provide capacity more predictably, respond faster to changing demand, and use cloud resources more efficiently.

 

Rather than relying only on queued jobs to trigger growth, administrators can combine baseline capacity, real-time resource monitoring, and multi-node scaling events to match compute availability more closely to actual and predicted workload needs.

 

For more information about this topic and other aspects of SAS Workload Orchestrator in action, visit learn.sas.com to view the Architecture and Administration for SAS® Workload Management on SAS® Viya® workshop.

 

 

Find more articles from SAS Global Enablement and Learning here.

Contributors
Version history
Last update:
Wednesday
Updated by:

Viya Copilot Motion Graphic.gifViya Copilot Motion Graphic

Ready to see what SAS Viya Copilot can do?

Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.

Get Started →

SAS AI and Machine Learning Courses

The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.

Get started