Is this your first time using statistical procedures within SAS software? Are you new to statistics in general? Has it been a while since your last statistics course? Need a review of the multitude of statistical procedures found in SAS? If you answer yes to any of these questions, then this series is for you. In part 1, we discussed aspects of exploring and describing continuous variables. We investigated PROC SGPLOT, MEANS, UNIVARIATE, and CORR. In part 2, our discussion turned to the modeling aspects of continuous variables. Our focus was on PROC REG, GLM, GLMSELECT, and PLM. In part 3, we took our analysis to categorical variables. Specifically, we discussed procedures that allow us to investigate and explore any categorical variables in our data. In part 4, we discussed modeling categorical variables using the popular procedure PROC LOGISTIC. In part 5, we moved away from the classical programming aspects of statistical procedures and explored utilizing graphical interfaces to access descriptive and graphical procedures. In part 6, we used SAS Visual Analytics to perform automated explanations and then transform these explanations into linear and logistic regressions.
In this post, let’s move to a different SAS Viya interface that also allows point and click capabilities for analyses like linear and logistic regression, SAS Model Studio. Model Studio is a browser-based, low-code and no-code machine learning platform that lets you build, compare and deploy predictive models in a single environment. Model Studio has the means to go much further than just linear and logistic regressions. It supports both business analysts and data scientists by automating data preparation, model training and deployment while allowing customization with Python, R, and SAS code. You can also bring in machine learning algorithms like forests and neural networks for predictive modeling. To keep with the focus of our series, let’s just look at the creation of linear and logistic regressions.
From the SAS Landing Page in SAS Viya, go to the Applications Menu in the upper left corner and select Build Models. This opens the browser to SAS Model Studio. Model Studio is project based. Meaning that you will create a project, load data into this project, and perform all analysis and exploration within this project.
Start by clicking on the New Project button in the upper right of the interface. The message box that appears will request specific information for the new project. Start by giving the project a name.
Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.
The Type area is very important. There are three different types of projects that SAS Model Studio can create: Data Mining and Machine Learning, Forecasting, Text Analytics. Your choice here will limit your project to only this type of analysis. You will not be able to change this decision within the project. Please make sure that for linear and logistic regression you choose Data Mining and Machine Learning.
Projects within SAS Model Studio use pipelines. For those that are past or current users of SAS Enterprise Guide, pipelines are like process flows that allow you to follow the progress and process performed within an analysis. SAS R&D have created several pipeline templates structured for very specific styles of analyses. These templates can be edited by you if you would like or so need. For our example, let’s leave this as blank template. This will just place a Data node in the pipeline and then we will have the freedom to start from scratch with our pipeline.
Next is your data. Clicking browse next to the data line allows you to select a data set that is already in CAS memory or allows you to import a data set that you want to analyze. For our example, we will be using the vs_bank data set. Within this data set, we have both a target variable that is continuous (for our linear regression) and one that is binary (for our logistic regression). It should be noted that we will not be able to perform both analyses on the same project. You will see why when we complete the data section of the project.
Clicking Save both creates our new project based on our selections and opens it to the data area of the project.
When first viewing the Data tab of the project, you may notice that roles have already been assigned by SAS to some variables. I would like to take control of that. I first selected all variables and set the role to Rejected. Now we may proceed with our choices.
I see the variable b_tgt with the label tgt Binary New Product. Let’s select this variable and give it the role of target. With this being a binary target, we are proceeding first with a logistic regression. I next make the variables cat_input1, cat_input2, and all variables starting with the prefix logi_rfm as input role. If you try to make a second variable as a target, you will see a warning about this, and no pipelines can be run until the warning is cleared. This is why we will have to separate the linear and logistic regressions within our projects.
Now, let’s move to the Pipelines tab. With just the Data node present, we can take full control of what our pipeline will look like. You have multiple ways to bring more nodes to the pipeline. You can either right click on the node and select add a child node and then choose from the provided list.
Or you can click on the node icon on the left side of the screen and choose from there.
Under Data Mining Preprocessing, you can select items like clustering, imputation, and transformations. These can be used if needed within your analysis. Supervised Learning gives you access to various modeling structures like decision trees, forests, neural networks, and of course linear and logistic regressions. Note that if you use the node area, you will have to drag and drop the selected node into the pipeline. If you use the right click method, SAS Model Studio will already know where you would like to place the node within the pipeline.
Let’s start with Logistic Regression.
On the right side of the screen will appear properties and options for logistic regression. From the target link function to effect options, to selection options, you can customize elements of the logistic regression. Running the pipeline will execute the analysis based on these options.
Right clicking on the Logistic Regression node when running is complete allows you to inspect the results. From here you are given various items of output from the analysis including an area dedicated to assessments.
Now let’s switch to linear regression. As mentioned before, we cannot have two target variables selected at the same time. You are likely thinking that we can just go back to the data tab and switch the roles. If you try this, you will notice that the role area is blocked from changes now that a pipeline has been executed. We will create a new project.
In this project, we will make int_tgt the target role. All other variables will be set to be rejected. Next, we’ll modify the variable roles so that the same input variables described above will be used in this regression. Moving to Pipelines, we will add a linear regression node to the pipeline using the methods explained earlier.
Like with logistic regression, we have effect options and selection options we can implement on the analysis. Running the pipeline with this node added we can retrieve the results just like with the logistic regression node. (Right click the node and choose results.)
If you are interested in trying any of the pipeline templates that have been previously made by SAS R&D, I would suggest looking for the templates that mention class or interval targets. There are both basic and advanced templates of both varieties.
Interested in using SAS code to perform any of these statistical procedures? Please go back to the previous parts of this series. If you are not a coder, SAS has ways for you to access our statistical procedures using several graphical user interfaces. Try out SAS Visual Analytics and SAS Information Catalog and see what aspects you enjoy. See you in the next installment of this series.
Find more articles from SAS Global Enablement and Learning here.
Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.
The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.