Is this your first time using statistical procedures within SAS software? Are you new to statistics in general? Has it been a while since your last statistics course? Need a review of the multitude of statistical procedures found in SAS? If you answer yes to any of these questions, then this series is for you. In part 1, we discussed aspects of exploring and describing continuous variables. We investigated PROC SGPLOT, MEANS, UNIVARIATE, and CORR. In part 2, our discussion turned to the modeling aspects of continuous variables. Our focus was on PROC REG, GLM, GLMSELECT, and PLM. In part 3, we took our analysis to categorical variables. Specifically, we discussed procedures that allow us to investigate and explore any categorical variables in our data. In part 4, we discussed modeling categorical variables using the popular procedure PROC LOGISTIC. In part 5, we moved away from the classical programming aspects of statistical procedures and explored utilizing graphical interfaces to access descriptive and graphical procedures. In part 6, we used SAS Visual Analytics to perform automated explanations and then transform these explanations into linear and logistic regressions. In part 7, we left SAS Visual Analytics and created both a linear and logistic regressions within SAS Model Studio.
In this post, let’s go back and look at some of the results from our graphical explorations and model analyses. We’ll point out some of the pitfalls that can quickly make your analysis or explanations questioned and untrustworthy.
Let’s start with correlation. First and foremost, you need to chant this statement and etch it into your brain, “Correlation does not mean causation!” Too many researchers see a significant and strong correlation between two variables and immediately interpret it as one of these variables causing the other to happen. A causative relationship cannot be assessed from a correlation. To determine causation, one could perform an experiment where you isolate the variable in question, and all other aspects of the experiment are held constant. Then can you truly determine the causation relationship between the variables of interest. Until that point, you merely have correlation established which means that one of your variables of interest may be useful within a model for the other.
Correlation has another issue of which you must be mindful. Interpretations of the pearson correlation coefficient do not make sense and are not valid unless the relationship between the variables are linear. Consider a relationship between our predictor variable and response that is truly quadratic. It is possible for you to collect specific subsamples of the data where the calculated correlation would be positive (if more of the positively sloped side is sampled), negative (if more of the negatively sloped side is sampled), and zero (if the sample is balanced from both the positive and negatively sloped sides).
Within SAS Visual Analytics, it is very easy to get SAS to create a correlation matrix where a heat map indicates the level of correlation among the variables chosen. Is it ok to just accept these correlations and begin your description? No! You will need to explore the scatterplots of these variables to determine if the underlying relationship is linear. Ignoring this check step will cause you to make statements about correlations that cannot be allowed.
Look at this example using the BASEBALL data set from SASHELP.
Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.
A quick glance suggests that the strongest correlations appear to be with the variable, Career RBIs. Specifically, the correlation between career RBIs and log salary is presented as 0.6294. But what type of relationship exists between these two variables? Within SAS Visual Analytics, we can quickly change this plot to a scatterplot matrix to assess the underlying relationship.
Looking at the scatterplot for log salary and career RBIs, I am not convinced that we have a linear relationship. I see what appears to be a line of asymptote which appears to be creating a curvilinear relationship. To be honest, several of these scatterplots show curvilinear relationships. Career runs and career hits is the one scatterplot that appears the most linear of the ones provided and the only one I would report its correlation.
Let’s move to Linear Regression. Regression analysis is widely taught in statistical courses and is one that is performed most prominently. Despite this common use, there are a couple of pitfalls that analysts fall into during their analysis.
First is the p-value associated with the intercept. Ever since your introductory statistics courses, you have been focusing your attention on the p-values associated with the parameter estimates table. You compare each of these values to your alpha of choice and determine to keep or drop a variable from the model. While this can be a useful approach for the explanatory variables in the model, it is not a prudent way to assess the intercept.
You should not be looking at p-value value significance to determine whether to include an intercept in a model. Before a non-intercept model can be considered, you must determine if the origin is a valid point of the data set. In other words, if all predictor variables are at zero, will that yield a response of zero? There are other times when an intercept can be removed from the model. One other example would be in Bayesian analysis when you are trying to assist in the parameters converging during the MCMC procedure. Simply using the p-value of the intercept is not enough evidence to move to a non-intercept model.
Secondly, please do not forget to assess your residuals from the linear regression. There are assumptions in linear regression that must be confirmed before we can accept the output analysis. Look at this example from the BASEBALL data set with log salary as the response variable.
From this residual plot, I am concerned about the curvilinear relationship between the residual and the predicted value. This indicates that one of my predictor variables is missing a potential polynomial term. The assumption that has been violated is that the relationship between the average response and the predictor has a nonlinear aspect that we have missed. We have misidentified the model. If we proceed with this model, we will not be able to trust the output and would be missing an important piece of the relationship among the predictors and the response.
Finally, let’s look at a pitfall within logistic regression. We just saw an example within linear regression where polynomials would be helpful. Sadly, when analysts perform logistic regression, they typically do not check the functional form of the predictor variables. It is assumed that when you apply the logit transformation the sigmoidal relationship between the predictor and the probability of the event changes to a linear relationship between the predictor and the logit. Do you take the time to check this assumption? How would you do that? Empirical logit plots to the rescue!
Using a binning technique and a quick calculation, you can create these plots that will assess the relationship between the predictor and the logit.
Ideally, we want the image to mimic the upper left, indicating a linear relationship. If we see the other two images, we would be helping our model if we included polynomial forms of the predictor variable in question.
The code that creates such an image looks like this. Suppose that we are looking at predicting a mother having a low-birth-weight child. We are interested in the relationship her weight has to this event. Recall that logit is the log of the odds of the event. Since weight is continuous, we will bin the variable so that we can get estimates of both the probabilities of the event and the non-event. From this we calculate the logit using a small correction to avoid any computational issues. Finally, we plot the calculated empirical logit for each bin against the average weight from that bin.
proc rank data=cdal.birth groups=13 out=work.ranks;
var LWT;
ranks bin;
run;
proc means data=work.ranks noprint nway;
class bin;
var LOW LWT;
output out=work.bins sum(LOW)=Events mean(LWT)=avgLWT n(LOW)=cases;
run;
data work.plot;
set work.bins;
elogit=log((Events+(sqrt(cases)/2))/(cases-Events+(sqrt(cases)/2)));
run;
proc sgplot data=work.plot;
loess y=elogit x=avgLWT;
yaxis label="Estimated Logit";
title "Estimated Logit Plot of Mother's Weight";
run;
From the resulting image, we can determine that we might be assisted within the model to include a polynomial for age.
These are not the only pitfalls that you might run across while you perform your explorations and analyses. It is best to make sure that you are diligent in your work so that you can utilize the output that is provided to you from SAS. If you would like more information about these analyses, check out the following classes:
See you in the next installment of this series.
Find more articles from SAS Global Enablement and Learning here.
Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.
The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.