Some SAS analytic procedures don’t accept CLASS variables. If you want to use a categorical predictor with these procedures, you’ll need to convert it into numeric dummy variables. Most users create these manually in a DATA step, but for variables with many levels this quickly becomes tedious. A faster approach is to use PROC GLMSELECT as a dummy‑variable generator. With the OUTDESIGN= option, GLMSELECT can produce a full design matrix for any categorical variable using GLM coding, reference coding, or effect coding. In this post, I’ll walk through how dummy variables work, why they’re needed, and how to generate them efficiently using GLMSELECT.
Dummy variables, also called design variables or indicator variables, are numeric variables that are used to represent the levels of a categorical variable. They typically take only the values 0 and 1 (and sometimes negative 1) and are used to indicate the presence or absence of characteristics or membership in a group.
For example, a categorical variable Gender with three levels (Female, Male, and Other) can be dummy-coded (sometimes called “one-hot encoding”) as three numeric variables:
GLM coding
| Gender | Female dummy variable | Male dummy variable | Other dummy variable |
| Female | 1 | 0 | 0 |
| Male | 0 | 1 | 0 |
| Other | 0 | 0 | 1 |
Why turn one variable into several that contain the same information? Because statistical procedures can’t use text values like “Female,” “Male,” or “Other” directly in a mathematical model. When you include a categorical variable in a regression using a CLASS statement, SAS automatically creates the necessary dummy variables behind the scenes. For example, PROC GLM generates one indicator variable for each level of the categorical predictor, as in the Gender example above. This approach is known as GLM coding.
Because GLM coding creates one dummy variable per level, the model is overparameterized when an intercept is included. SAS handles this by setting the last parameter estimate to zero, making the intercept the estimate for that reference group.
GLM coding is just one approach to creating dummy variables. SAS procedures such as GLMSELECT, GENMOD, LOGISTIC, and MIXED also support other coding schemes that change the number of dummy variables and the interpretation of their coefficients. A widely used option is reference coding, which produces one fewer dummy variable than GLM coding:
Reference coding
| Gender | Female dummy variable | Male dummy variable |
| Female | 1 | 0 |
| Male | 0 | 1 |
| Other | 0 | 0 |
In regression models, the coefficients for reference‑coded dummy variables represent differences between each group (Female or Male) and the reference level (Other), which is the level coded with all zeros.
Another common coding scheme is effect coding. Like reference coding, it uses one fewer dummy variable than GLM coding, but the reference level is coded with –1s instead of 0s:
Effect coding
| Gender | Female dummy variable | Male dummy variable |
| Female | 1 | 0 |
| Male | 0 | 1 |
| Other | –1 | –1 |
With effect coding, each coefficient shows how that group’s mean differs from the overall mean across all levels of the categorical variable. The reference category is assigned –1s so that the group effects sum to zero, which is why this approach is sometimes referred to as “deviation‑from‑the‑mean” coding.
Why dummy coding is needed
Because these procedures only work with numeric variables, any categorical predictors must be dummy‑coded first. The table below shows several SAS procedures that require numeric inputs and what they’re typically used for.
| Procedure/option requiring numeric variables | What it is commonly used for |
| PROC REG | Regression analysis. PROC REG produces a greater variety of model diagnostic plots than PROC GLM or GLMSELECT, so models fit in those procedures can be refit in PROC REG to graphically check model assumptions. REG also can calculate collinearity diagnostics like the Condition Index and Variance Inflation Factors. For explanations of these collinearity tools see: What is collinearity and why does it matter? | Global Enablement & Learning Blog. |
| PROC PRINCOMP | Principal component analysis. This technique is often used in reducing the number of variables to be used in other models. For a review of PCA, see: How many principal components should I keep? Part 1: common approaches | Global Enablement & Learnin.... |
| PROC VARCLUS | Variable clustering. Can be used for variable-reduction by clustering variables and keeping one variable per cluster. This approach reduces multicollinearity. |
| PROC FASTCLUS | Clustering observations. Cluster membership could be used for segmentation. Predictive models could be separately fit to each cluster/segment, producing predictions tailored to important subgroups. |
| SELECTION=SCORE (“best subsets selection” in PROC PHREG and PROC LOGISTIC) | A variable selection approach that looks at all possible models and ranks them by the score chi-squared statistic. This assesses many more potential models than the stepwise selection methods Forward, Backward, and Stepwise selection. |
There are many ways to create dummy variables in SAS. Here are two common DATA‑step approaches before we switch to the GLMSELECT method. I’ll dummy code the variable “Origin” from the sashelp.cars data. Origin has three levels: Asia, Europe, and USA.
Dummy coding in a DATA step
The first approach uses IF, THEN, and ELSE statements. The second approach uses logical tests in which SAS will assign a 1 when the statement in parentheses is true and assign a 0 when the parentheses term is false.
data dummy1;
set sashelp.cars;
if origin="Asia" then Origin_A=1; else Origin_A=0;
if origin="Europe" then Origin_E=1; else Origin_E=0;
if origin="USA" then Origin_U=1; else Origin_U=0;
run;
data dummy2;
set sashelp.cars;
Origin_A=(origin="Asia");
Origin_E=(origin="Europe");
Origin_U=(origin="USA");
run;
The above code is not difficult to write, but with 20+ levels it becomes tedious and error‑prone. GLMSELECT makes this much easier.
Dummy coding with GLMSELECT
To generate dummy variables, create a regression model with PROC GLMSELECT, putting categorical variables in the CLASS and MODEL statements. Use the OUTDESIGN option in the PROC statement to save the dummy variables to a SAS data set. If you want to save all the inputs, use the ADDINPUTVARS option for OUTDESIGN to include all the variables in the output data set. The variable names used in the model will be automatically saved to a macrovariable called _GLSMOD. %PUT can be used to write the value of this macrovariable to the SAS log.
GLMSELECT uses GLM-coding by default, but you can change the coding scheme with the PARAM= option in the CLASS statement. The SHOWCODING option prints the coding scheme used for each CLASS variable, which is helpful for verifying GLM, reference, or effect-coding. GLMSELECT performs stepwise variable selection by default, so this should be turned off using the SELECTION=NONE option in the MODEL statement.
ods select ClassLevelInfo ClassLevelCoding;
proc glmselect data=sashelp.cars outdesign (addinputvars)=x;
class origin/showcoding;
model MSRP= origin /selection=none;
run;
proc contents data=x;
run;
proc print data=x (obs=15);
var "Origin Asia"n "Origin Europe"n "Origin USA"n;
run;
%put &_GLSMOD; /* writes the variable names from GLMSELECT model to the log */
PROC GLMSELECT output:
Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.
PROC CONTENTS output (without ADDINPUTVARS):
PROC CONTENTS output (with ADDINPUTVARS):
PROC PRINT output:
SAS Log:
To switch to reference or effect coding, add param=ref or param=effect to the CLASS statement in PROC GLMSELECT:
class origin/param=ref showcoding;
class origin/param=effect showcoding;
Dummy coding doesn’t need to be a separate preprocessing chore. When a SAS procedure won’t accept CLASS variables, you don’t need to hand code anything. PROC GLMSELECT can generate the design matrix for you, and you can pass it directly into your downstream analysis. It’s a small trick that saves time across many analyses.
Links
Find more articles from SAS Global Enablement and Learning here.
In SAS, there are other PROCs could do the same thing:
PROC GLMMOD
PROC LOGISTIC
PROC TRANSREG
PROC GLIMMIX
Check Rick's blogs:
https://blogs.sas.com/content/iml/2016/02/24/create-a-design-matrix-in-sas.html
https://blogs.sas.com/content/iml/2016/02/22/create-dummy-variables-in-sas.html
Thanks for the info, especially PROC GLMMOD which I was not familiar with!
GLMMOD is how we have done this for decades, but the truth is, we should be able to get the design matrix out of all the analytic procs. Why not?
Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.
The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.