BookmarkSubscribeRSS Feed

Building a Layered Node Scaling Strategy with SAS Workload Orchestrator for Predictable Performance

Started 3 hours ago by
Modified Wednesday by
Views 27

SAS Workload Orchestrator (SWO) plays a unique role within the SAS Viya platform by bridging business workload priorities with the elastic infrastructure capabilities of Kubernetes. While Kubernetes provides the foundation for container orchestration and scaling, SWO adds workload awareness, helping ensure that analytics jobs receive the resources they need while maintaining service levels across diverse users and workloads. Recent enhancements to SWO’s scaling capabilities further strengthen this value proposition by giving administrators multiple, complementary ways to manage capacity: Time-Based Minimums for predictable demand patterns, Dynamic Resource Threshold Scaling for sustained resource pressure, and Reactive Pending-Job Scaling for sudden bursts of workload demand. 

A common question we hear from customers is: “Which scaling features should we use and when?” 

 

Analytics workloads don't behave like a single traffic pattern. The same cluster might need to absorb a predictable 8 a.m. usage surge, an unplanned afternoon burst of ad-hoc data science jobs, and a nightly batch pipeline that submits significant number of jobs at once. No single scaling feature can cover all three well for optimum performance. That's why SWO's scaling features are designed to be layered, a “defense in depth” approach to capacity management. This blog provides a practical decision framework, how the features interact when more than one is active, and a worked example of running all three together on the same SWO host type.

The three features at a glance:

 

Screenshot 2026-09-30 at 2.12.07 PM.png

  

Must Read SAS Workload Management advanced Kubernetes node autoscaling GEL blog that is written by Edoardo Riva, to learn more about these three SWO Scaling capabilities.

 
Dynamic Resource Threshold Scaling is the latest scaling feature and it deserves a special mention. Its support for Disk I/O utilization and SASWORK utilization (refer dynamic resource for the full list) is a capability the open-source Kubernetes autoscaling stack does not provide out of the box. Stock Kubernetes autoscaling i.e. the Horizontal Pod Autoscaler, the Cluster Autoscaler, or opensource options such as Karpenter considers CPU and memory as the signals that matter. None of them natively watch disk I/O throughput or a SAS-specific concept like SASWORK utilization. Getting there in an open-source stack means building your own metrics pipeline: exporting node/pod-level I/O metrics, wiring them through a Prometheus adapter, and authoring custom HPA or provisioner rules around them. That's real, ongoing engineering investment most platform teams don't take on.


SAS workloads are frequently I/O and SASWORK-intensive long before they're CPU-intensive. Example: data-heavy transformations, large temporary work files, concurrent SAS sessions competing for local scratch space. A scaling strategy that only watches CPU and memory can miss exactly the pressure that matters most for these jobs. SWO exposes Disk I/O rate and SASWORK utilization as first-class, built-in dynamic resources, configurable in SAS Environment Manager right alongside CPU and memory, which means no external telemetry pipeline required.

If you're evaluating orchestration or autoscaling options for I/O- or storage-intensive SAS workloads, put this on your checklist explicitly: ask whether the approach you're considering can scale on disk I/O and SASWORK pressure natively, not just CPU and memory. For SAS analytics environments, this is often the difference between infrastructure that reacts to the right signal and infrastructure that's still waiting for a CPU spike that may never come.

 

Decision framework at a glance

 

Use this as a starting checklist when deciding what to turn on for a given host type:

  • Do you know exactly when demand will spike? (Morning usage rush, scheduled batch window, end-of-month close) → Configure a Time-Based Minimum that pre-warms nodes ahead of the window. As a rule of thumb, 20–30 minutes before the surge is enough for most cloud providers' VM boot times.
  • Is demand genuinely unpredictable ex: ad-hoc analysis, exploratory model training, variable data volumes? → Configure Dynamic Resource Threshold Scaling on the resource that actually constrains those workloads (often disk I/O or SASWORK-adjacent storage pressure before CPU becomes the bottleneck). Set the threshold with a buffer i.e. 70–80% utilization is a reasonable starting point so there's still headroom to absorb the spike while the new node boots.
  • Do jobs sometimes arrive in large, synchronized bursts rather than trickling in one at a time? (Nightly risk aggregation, a scheduled batch pipeline that submits dozens of jobs simultaneously) → Ensure Autoscaling is enabled on the host type, this will ensure multi-node scaling will be triggered when a burst of pending jobs is in a SWO Queue. It provisions the whole set of needed nodes in one event instead of trickling out one node at a time.
  • For all the other cases, the Reactive Scaling based on the pending jobs is the failsafe that guarantees a pending job eventually gets a node, even if neither of the above proactive features are enabled.

 

prasadpz_1-1790791059796.png

How the three features interact:

Running more than one feature at the same time is normal and expected. Three rules matter:

  • Time-Based Minimum is always enforced. If a minimum node count is configured for a window, that baseline capacity is enforced regardless of what current utilization looks like. Threshold scaling and reactive scaling can still add nodes on top of that floor, but they can never scale below it during the window.
  • Multiple Scale Thresholds on the same host type are evaluated as OR, not AND. If you configure separate thresholds for, say, memory and disk I/O, either one breaching independently is enough to trigger a scale-up. There's no single blended, weighted metric i.e. just the first threshold to breach wins.
  • Multi-node capability applies within Reactive and both Proactive scaling features. While the Reactive scaling automatically calculates the one or more nodes to scale up at a time, both Proactive scaling options works based on the configured value to add more than one node per event. 

 

Worked example: one host type, all three features on a day

 

Picture a single scalable host type serving a mixed interactive/batch analytics environment:

  • 7:30 a.m. Time-Based Minimum pre-warms the host type to its configured floor ahead of the 8 a.m. login surge. Analysts log in at 8:00am and get immediate compute; no cold start.
  • 11:15 a.m. A data scientist kicks off an ad-hoc, I/O-heavy feature-engineering job. Average disk I/O across the host type climbs past the configured 75% threshold and stays there for over a minute. Dynamic Resource Threshold Scaling fires, a new node is already online by the time the workload actually needs it.
  • 11:45 a.m. Utilization drops back under the threshold once the job finishes; the extra node becomes eligible for the Cluster Autoscaler to reclaim during its normal idle-node cleanup.
  • 8:00 p.m. A nightly batch pipeline submits 20 jobs at once against a queue limited to 5 jobs per host. Reactive Pending-Job Scaling's multi-node capability calculates that four (4) additional nodes are needed and requests them together in one event, instead of scaling one node at a time and falling further behind.
  • Anything else, any time if a job somehow still can't find a host across all of the above (an edge case not covered by a configured threshold or window), Reactive Pending-Job Scaling is still running underneath as the fallback that guarantees the job doesn't wait forever.

No single trigger above would have handled that whole day well on its own. Together, they cover it without any manual intervention and ensure capacity is made available to provide optimum performance for your workloads.

 

Tuning pitfalls to avoid

 

  • Don't set thresholds too close to 100%. A threshold configured at 95% still means the job that trips it waits for the new node. The buffer may protect the next job, but not the one that caused the spike. That's why 70–80% is the more practical starting point.
  • Tune 'Autoscaler delay after new host' value to your cloud provider's real VM boot time, not a guess. Too short a delay risks requesting a second, third, and fourth node before the first one has even finished booting, a scaling “thrashing” loop that costs money without adding useful capacity.
  • Remember the precedence rule. If you're troubleshooting “why didn't my node count go below what I expected,” check whether a Time-Based Minimum window is active before assuming threshold scaling is misbehaving.
  • Confirm your ClusterRole/ClusterRoleBinding is correctly scoped. Dynamic Resource Threshold Scaling depends on SWO collecting cluster-wide telemetry. Without the right permissions, thresholds simply won't have the data they need to evaluate.
  • Cap total node count at the cloud-provider node-pool level, not in SWO. None of the three features enforce a maximum on their own.

 

Conclusion

 

None of SWO's three scaling features is trying to be the one answer for optimum performance. Reactive scaling is the safety net every environment can utilize, and since 2026.05 it can absorb bursts in bulk instead of trickling out one node at a time. Time-Based Minimums buy you predictability for the demand you already know is coming. Dynamic Resource Threshold scaling buys you resilience against the unplanned demand. Layered together, they turn Kubernetes' native “wait until something breaks, then react” model into something closer to what analytics workloads actually need: capacity that shows up before your users notice it was missing, without paying to keep it running around the clock. Together, these features enable a more proactive, resilient, and cost-conscious scaling strategy that helps SAS Viya environments stay responsive under a wide range of operational conditions. 

 

Further Reading:

 

 

 

Contributors
Version history
Last update:
Wednesday
Updated by:

Viya Copilot Motion Graphic.gifViya Copilot Motion Graphic

Ready to see what SAS Viya Copilot can do?

Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.

Get Started →

SAS AI and Machine Learning Courses

The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.

Get started

Article Tags