SAS Workload Orchestrator (SWO) plays a unique role within the SAS Viya platform by bridging business workload priorities with the elastic infrastructure capabilities of Kubernetes. While Kubernetes provides the foundation for container orchestration and scaling, SWO adds workload awareness, helping ensure that analytics jobs receive the resources they need while maintaining service levels across diverse users and workloads. Recent enhancements to SWO’s scaling capabilities further strengthen this value proposition by giving administrators multiple, complementary ways to manage capacity: Time-Based Minimums for predictable demand patterns, Dynamic Resource Threshold Scaling for sustained resource pressure, and Reactive Pending-Job Scaling for sudden bursts of workload demand.
A common question we hear from customers is: “Which scaling features should we use and when?”
Analytics workloads don't behave like a single traffic pattern. The same cluster might need to absorb a predictable 8 a.m. usage surge, an unplanned afternoon burst of ad-hoc data science jobs, and a nightly batch pipeline that submits significant number of jobs at once. No single scaling feature can cover all three well for optimum performance. That's why SWO's scaling features are designed to be layered, a “defense in depth” approach to capacity management. This blog provides a practical decision framework, how the features interact when more than one is active, and a worked example of running all three together on the same SWO host type.
Must Read SAS Workload Management advanced Kubernetes node autoscaling GEL blog that is written by Edoardo Riva, to learn more about these three SWO Scaling capabilities.
Dynamic Resource Threshold Scaling is the latest scaling feature and it deserves a special mention. Its support for Disk I/O utilization and SASWORK utilization (refer dynamic resource for the full list) is a capability the open-source Kubernetes autoscaling stack does not provide out of the box. Stock Kubernetes autoscaling i.e. the Horizontal Pod Autoscaler, the Cluster Autoscaler, or opensource options such as Karpenter considers CPU and memory as the signals that matter. None of them natively watch disk I/O throughput or a SAS-specific concept like SASWORK utilization. Getting there in an open-source stack means building your own metrics pipeline: exporting node/pod-level I/O metrics, wiring them through a Prometheus adapter, and authoring custom HPA or provisioner rules around them. That's real, ongoing engineering investment most platform teams don't take on.
SAS workloads are frequently I/O and SASWORK-intensive long before they're CPU-intensive. Example: data-heavy transformations, large temporary work files, concurrent SAS sessions competing for local scratch space. A scaling strategy that only watches CPU and memory can miss exactly the pressure that matters most for these jobs. SWO exposes Disk I/O rate and SASWORK utilization as first-class, built-in dynamic resources, configurable in SAS Environment Manager right alongside CPU and memory, which means no external telemetry pipeline required.
If you're evaluating orchestration or autoscaling options for I/O- or storage-intensive SAS workloads, put this on your checklist explicitly: ask whether the approach you're considering can scale on disk I/O and SASWORK pressure natively, not just CPU and memory. For SAS analytics environments, this is often the difference between infrastructure that reacts to the right signal and infrastructure that's still waiting for a CPU spike that may never come.
Use this as a starting checklist when deciding what to turn on for a given host type:
Running more than one feature at the same time is normal and expected. Three rules matter:
Picture a single scalable host type serving a mixed interactive/batch analytics environment:
No single trigger above would have handled that whole day well on its own. Together, they cover it without any manual intervention and ensure capacity is made available to provide optimum performance for your workloads.
None of SWO's three scaling features is trying to be the one answer for optimum performance. Reactive scaling is the safety net every environment can utilize, and since 2026.05 it can absorb bursts in bulk instead of trickling out one node at a time. Time-Based Minimums buy you predictability for the demand you already know is coming. Dynamic Resource Threshold scaling buys you resilience against the unplanned demand. Layered together, they turn Kubernetes' native “wait until something breaks, then react” model into something closer to what analytics workloads actually need: capacity that shows up before your users notice it was missing, without paying to keep it running around the clock. Together, these features enable a more proactive, resilient, and cost-conscious scaling strategy that helps SAS Viya environments stay responsive under a wide range of operational conditions.
Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.
The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.