BookmarkSubscribeRSS Feed

Enhanced Observability for Compute Workloads

Started ‎09-01-2026 by
Modified ‎09-01-2026 by
Views 92

After reading the post by Bruno Mueller entitled “How to customize your Programming Run-Time Servers pod names” I was inspired to look at what this might mean for observability of the SAS Viya platform.

 

So, how can we take advantage of this approach with the SAS observability tools, and do you need to make any configuration changes?

 

I’m talking about SAS® Enterprise Session Monitor and the SAS® Viya® Monitoring for Kubernetes framework.

 

In his post Bruno explains how you can update the SAS Viya platform configuration so that you can identify different workloads based on the different launcher contexts and by changing the default pod names for the different programming runtime pods.

 

Let’s start by looking at SAS Enterprise Session Monitor.

 

 

SAS Enterprise Session Monitor

 

The good news is that “out-of-the-box” SAS Enterprise Session Monitor (ESM) has good observability of the different compute workloads. For example, it differentiates between SAS Studio compute sessions and SAS Batch Server jobs. If we look at the ESM main dashboard, there is a gauge showing the workload by type (the 'Load by Type' gauge).

 

Below is an image taken while generating some simple workload.

 

MG_202608-01_esm-dashboard.png

Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.

 

But I wanted to understand the impact of changing the pod names, does this have any impact on ESM.

 

To investigate this, I first started two SAS Studio sessions (for users sastest1 and sastest2). This was using the default configuration. Therefore, the pods have the default names. For example.

 

NAME                                                         READY   STATUS    RESTARTS   AGE   USERNAME      JOB-TYPE         REQUESTED-BY-CLIENT
sas-compute-server-b8728159-18ea-43df-84c5-8001d76ec49e-25   2/2     Running   0          34m   sastest1      compute-server   sas.studio
sas-compute-server-65efde45-8a93-4ecf-814a-4746151fde87-26   2/2     Running   0          33m   sastest2      compute-server   sas.studio
 

 

I then updated the pod prefix to be “studio-cs-“ as described in Bruno’s blog. I then started two more SAS Studio sessions, this time for the gatedemo001 and gatedemo002 users. Again, looking at the pods, this gave the following result.

 

NAME                                                         READY   STATUS    RESTARTS   AGE   USERNAME      JOB-TYPE         REQUESTED-BY-CLIENT
sas-compute-server-b8728159-18ea-43df-84c5-8001d76ec49e-25   2/2     Running   0          34m   sastest1      compute-server   sas.studio
sas-compute-server-65efde45-8a93-4ecf-814a-4746151fde87-26   2/2     Running   0          33m   sastest2      compute-server   sas.studio
studio-cs-4b26e5e5-e50e-4462-b8d9-a0a7cb1612c5-31            2/2     Running   0          24m   gatedemo001   compute-server   sas.studio
studio-cs-b6cfbc96-9201-4ec1-891a-0733ec054913-32            2/2     Running   0          22m   gatedemo002   compute-server   sas.studio

 

You can see that the new sessions have the custom pod names assigned. As Bruno described, this allows you to distinguish between different compute contexts. For example, if using workload orchestration, you could set different names for the contexts associated with the different queues. Thus, improving visibility.

 

I then started several batch sessions using the SAS Viya CLI to submit the jobs. This gave the following.

 

NAME                                                         READY   STATUS    RESTARTS   AGE   USERNAME      JOB-TYPE         REQUESTED-BY-CLIENT
sas-compute-server-b8728159-18ea-43df-84c5-8001d76ec49e-25   2/2     Running   0          34m   sastest1      compute-server   sas.studio
sas-compute-server-65efde45-8a93-4ecf-814a-4746151fde87-26   2/2     Running   0          33m   sastest2      compute-server   sas.studio
studio-cs-4b26e5e5-e50e-4462-b8d9-a0a7cb1612c5-31            2/2     Running   0          24m   gatedemo001   compute-server   sas.studio
studio-cs-b6cfbc96-9201-4ec1-891a-0733ec054913-32            2/2     Running   0          22m   gatedemo002   compute-server   sas.studio
sas-batch-server-2aa84357-4a40-41ec-b10b-fe3a27d6f0a0-52     2/2     Running   0          14s   sastest1      sas-batch-job    sas.cli
sas-batch-server-30ad6ec0-4452-400e-8603-a9f6eb5bb25b-49     2/2     Running   0          14s   sastest1      sas-batch-job    sas.cli
sas-batch-server-67f0a41c-75fb-44f0-ab66-2393771bdc50-50     2/2     Running   0          14s   sastest1      sas-batch-job    sas.cli
sas-batch-server-f58106a4-b24f-497e-a5ea-de056d1e0459-51     2/2     Running   0          14s   sastest1      sas-batch-job    sas.cli

 

As can be seen, all the sessions were running under the sastest1 user. Above you can see the default pod names for the ‘sas-batch-server’ sessions.

 

So, what does this look like in SAS Enterprise Session Monitor?

 

In ESM if we look at the ‘Live View’ and filter on the compute node(s) and the process categories of CMP and Batch we can see the following.

 

MG_202608-02_esm-live-view.png

Looking at the image, you can see 4 SAS Studio sessions and 4 batch sessions. The SAS Studio sessions were all identified as compute sessions, this is indicated by the type of CMP. The first two processes are for sastest1 and sastest2, these were using the default pod names. You can also see that the process name is ‘sas.studio’.

 

The SAS Studio sessions for the gatedemo users had the custom pod names, this led to the process name ‘studio-cs’ being assigned.

 

For the batch sessions, you can see these were all associated with the ‘Batch’ category type. The process name was the name of the SAS program being run.

 

As can be seen, there is no special configuration in ESM required to support the approach of use custom pod names.

 

See the Enterprise Session Monitor documentation: here

 

 

SAS Viya Monitoring for Kubernetes

 

When it comes to the Grafana dashboards in SAS Viya Monitoring for Kubernetes, having unique pod names allows you to write regex (regular expression) statements that can target specific compute pods. Perhaps you want to report on a specific type of workload.

 

As an example, you could use a Time Series gauge to show a history of the running compute pods. For this you could use the following two queries:

 

sum(kube_pod_status_phase{namespace="$namespace", phase="Running", pod=~"sas-compute-server-.*"})

 

And

 

sum(kube_pod_status_phase{namespace="$namespace", phase="Running", pod=~"sas-batch-server-.*"})

 

Using the example above of changing the pod names for the SAS Studio sessions the query would look like the following.

 

sum(kube_pod_status_phase{namespace="$namespace", phase="Running",pod=~"studio-cs-.*"})

 

You might do this to get more fine grain reporting. For example, this could be used to report on usage by business unit or department. Assuming these users are running under different contexts.

 

However, the use of custom pod names may not be needed in all cases and/or it can be better to target the pod labels rather than the actual pod names. This can provide a more flexible and maintainable approach when it comes to developing the Grafana dashboard.

 

As a side note, while I have focused on the queries for Grafana dashboards, another use case for using the custom pod names is for custom alerts (PrometheusRules) on specific types of workloads. In this case you'd also have to create or change any alert rules that refer to custom pod names using the same approach (wildcards in your promql expression).

 

So, let’s look at the standard compute pod labels, then I will show you how I made use of them.

 

We will start by looking at the labels on the compute server pods, the SAS Studio pods.

 

Pod Labels
app=sas-workload-orchestrator
launcher.sas.com/job-type=compute-server
launcher.sas.com/requested-by-client=sas.studio
launcher.sas.com/username=username
sas.com/created-by=sas-launcher
sas.com/deployment=sas-viya
swo.sas.com/containerName=sas-programming-environment
swo.sas.com/jobID=18
swo.sas.com/jobName=sas-compute-server-xxx-xxx-xxx-xxx…
swo.sas.com/queueName=default
swo.sas.com/realm=viya

 

I have highlighted two of the labels that could be used in a query. Being the job-type and client. It should be noted that the context isn’t part of the standard labels, this is where using custom pod names could be useful.

 

Now looking at the labels on a batch job we can see the following.

 

Pod Labels
app=sas-workload-orchestrator
launcher.sas.com/job-type=sas-batch-job
launcher.sas.com/requested-by-client=sas.cli
launcher.sas.com/username=username
sas.com/config-init=true
sas.com/created-by=sas-launcher
sas.com/deployment=sas-viya
swo.sas.com/containerName=sas-programming-environment
swo.sas.com/jobID=21
swo.sas.com/jobName=high-cpu-compute-load
swo.sas.com/queueName=default
swo.sas.com/realm=viya

 

Here we can see the job-type label is 'sas-batch-job' and that I was using the SAS Viya CLI to submit the batch job (the requested-by-client label is 'sas.cli').

 

All this made me think about what an enhanced view of the compute workloads and workload orchestration queues might look like. The SAS Viya Monitoring for Kubernetes project has a couple of dashboards that may provide all you need (SAS Launched Jobs- Node Activity and User Activity).

 

However, it should be noted that these dashboards only report on SAS Jobs launched via Workload Orchestrator. They use metrics data from the Workload Orchestrator. Therefore, any SAS jobs launched outside of Workload Orchestrator will not be shown/available on those two dashboards.

 

Here is my summary version, my ‘Sessions and Jobs’ panel. This is using two time series reporting objects (gauges).

 

MG_202608-03_grafana-sessions-and-queues-1024x330.png

Focusing on the 'Compute Session – By Type'

 

MG_202608-04_grafana-sessions-by-type.png

As stated, this is a ‘Time series’ gauge, it has 3 queries defined.

 

The first is to calculate the total number of all compute sessions (Total Compute – all types).

 

count(
  kube_pod_labels{
    namespace="$namespace",
    label_launcher_sas_com_job_type=~".+"
  }
  and on(namespace,pod)
  kube_pod_status_phase{
    namespace="$namespace",
    phase="Running"
  } == 1
)

 

This is counting all pods with the label ‘launcher.sas.com/job-type’, where the label is non-blank. In the Prometheus data store the dots, dashes and forward slash characters are all stored as an underscore. Hence, label ‘launcher.sas.com/job-type’ looks like: label_launcher_sas_com_job_type

 

In the image above you can see that there were 20 sessions at the peak load.

 

The second query counts the batch server pods, targeting the pod name.

 

sum(kube_pod_status_phase{namespace="$namespace", phase="Running", pod=~"sas-batch-server-.*"})

 

The third query counts the SAS Studio sessions.

 

count(
  kube_pod_labels{
    namespace="$namespace",
    label_launcher_sas_com_requested_by_client="sas.studio"
  }
  and on(namespace,pod)
  kube_pod_status_phase{
    namespace="$namespace",
    phase="Running"
  } == 1
)

 

This is using the launcher.sas.com/requested-by-client=sas.studio label.

 

My second gauge was reporting on the Workload Orchestration queues.

 

MG_202608-05_grafana-sessions-by-queue.png

In this image we can see that 3 queues were detected (default, high-priority and low-priority).

 

This has a single query, targeting the swo.sas.com/queueName label. In the query this translates to: label_swo_sas_com_queue_name

 

This query provides a count for each distinct value of the queue name label.

 

sum by (label_swo_sas_com_queue_name) (
  kube_pod_labels{
    namespace="$namespace",
    label_swo_sas_com_queue_name=~".+"
  }
  * on(namespace,pod) group_left()
  kube_pod_status_phase{
    namespace="$namespace",
    phase="Running"
  }
)

 

Looking at the image, you can see that I created two queues, a low-priority queue and a high-priority queue. The SAS Studio sessions were using the default queue. However, I could have created a “sas-studio” queue for better visibility / management reasons.

 

For example.

 

MG_202608-06_grafana-sessions-by-queue2.png

In this example the 'SAS Studio launcher context' was updated to use a queue called 'sas-studio'. After creating the 'sas-studio' queue and updating the context, I started several SAS Studio sessions and reran my batch load test program.

 

In this image, you can now see the SAS Studio sessions and the batch workload separately. The batch submissions were using the default, low-priority and high-priority queues. At the peak activity there were 5 SAS Studio session, 1 job running in the default queue, 3 high-priority jobs and 7 low-priority jobs running.

 

There was no need to adjust the query, the gauge automatically reported on the new queue.

 

 

Conclusion

 

Both tools provide great observability by default, but having unique pod names for the different programming runtime contexts gives the opportunity to enhance this.

 

SAS Enterprise Session Monitor has excellent detection and knowledge of the different internal SAS platform process types. As shown above, no additional customization was required when using the custom pod names.

 

The same is true for SAS Viya Monitoring for Kubernetes, but there is always the opportunity to create a custom dashboard to meet your requirements. I hope the Grafana dashboard example, the sample queries, has given you a taste of what is possible.

 

I think that having dedicated queues for workload types, such as SAS Studio, provides better observability. Especially, as shown here, if you wanted to provide a Grafana panel to monitor and track the running workloads.

 

MG_202608-07_grafana-sessions-and-queues-3-hours-1024x332.png

If you would like to learn more about building a Grafana dashboard to monitor SAS Viya see this post: Building a Grafana dashboard to monitor SAS Viya

 

Also see the SAS Viya Observability workshop.

 

I hope this is useful and thanks for reading.

 

 

Find more articles from SAS Global Enablement and Learning here.

Contributors
Version history
Last update:
‎09-01-2026 05:05 PM
Updated by:

Viya Copilot Motion Graphic.gifViya Copilot Motion Graphic

Ready to see what SAS Viya Copilot can do?

Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.

Get Started →

SAS AI and Machine Learning Courses

The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.

Get started

Article Tags