BookmarkSubscribeRSS Feed

A Hidden Gem in SAS: Using PROC GLMSELECT to Dummy Code Categorical Variables

Started ‎07-23-2026 by
Modified ‎07-23-2026 by
Views 367

Some SAS analytic procedures don’t accept CLASS variables. If you want to use a categorical predictor with these procedures, you’ll need to convert it into numeric dummy variables. Most users create these manually in a DATA step, but for variables with many levels this quickly becomes tedious. A faster approach is to use PROC GLMSELECT as a dummy‑variable generator. With the OUTDESIGN= option, GLMSELECT can produce a full design matrix for any categorical variable using GLM coding, reference coding, or effect coding. In this post, I’ll walk through how dummy variables work, why they’re needed, and how to generate them efficiently using GLMSELECT.

 

Dummy variables, also called design variables or indicator variables, are numeric variables that are used to represent the levels of a categorical variable.  They typically take only the values 0 and 1 (and sometimes negative 1) and are used to indicate the presence or absence of characteristics or membership in a group.

 

For example, a categorical variable Gender with three levels (Female, Male, and Other) can be dummy-coded (sometimes called “one-hot encoding”) as three numeric variables:

 

GLM coding

 

Gender Female dummy variable Male dummy variable Other dummy variable
Female 1 0 0
Male 0 1 0
Other 0 0 1

 

Why turn one variable into several that contain the same information? Because statistical procedures can’t use text values like “Female,” “Male,” or “Other” directly in a mathematical model. When you include a categorical variable in a regression using a CLASS statement, SAS automatically creates the necessary dummy variables behind the scenes. For example, PROC GLM generates one indicator variable for each level of the categorical predictor, as in the Gender example above. This approach is known as GLM coding.

 

Because GLM coding creates one dummy variable per level, the model is overparameterized when an intercept is included. SAS handles this by setting the last parameter estimate to zero, making the intercept the estimate for that reference group.

 

GLM coding is just one approach to creating dummy variables. SAS procedures such as GLMSELECT, GENMOD, LOGISTIC, and MIXED also support other coding schemes that change the number of dummy variables and the interpretation of their coefficients. A widely used option is reference coding, which produces one fewer dummy variable than GLM coding:

 

Reference coding

 

Gender Female dummy variable Male dummy variable
Female 1 0
Male 0 1
Other 0 0

 

In regression models, the coefficients for reference‑coded dummy variables represent differences between each group (Female or Male) and the reference level (Other), which is the level coded with all zeros.

 

Another common coding scheme is effect coding. Like reference coding, it uses one fewer dummy variable than GLM coding, but the reference level is coded with –1s instead of 0s:

 

Effect coding

 

Gender Female dummy variable Male dummy variable
Female 1 0
Male 0 1
Other –1 –1

 

With effect coding, each coefficient shows how that group’s mean differs from the overall mean across all levels of the categorical variable. The reference category is assigned –1s so that the group effects sum to zero, which is why this approach is sometimes referred to as “deviation‑from‑the‑mean” coding.

 

 

Why dummy coding is needed

 

Because these procedures only work with numeric variables, any categorical predictors must be dummy‑coded first. The table below shows several SAS procedures that require numeric inputs and what they’re typically used for.

 

Procedure/option requiring numeric variables What it is commonly used for
PROC REG Regression analysis. PROC REG produces a greater variety of model diagnostic plots than PROC GLM or GLMSELECT, so models fit in those procedures can be refit in PROC REG to graphically check model assumptions.  REG also can calculate collinearity diagnostics like the Condition Index and Variance Inflation Factors.  For explanations of these collinearity tools see:  What is collinearity and why does it matter? | Global Enablement & Learning Blog.
PROC PRINCOMP Principal component analysis.  This technique is often used in reducing the number of variables to be used in other models. For a review of PCA, see: How many principal components should I keep? Part 1: common approaches | Global Enablement & Learnin....
PROC VARCLUS Variable clustering.  Can be used for variable-reduction by clustering variables and keeping one variable per cluster. This approach reduces multicollinearity.
PROC FASTCLUS Clustering observations. Cluster membership could be used for segmentation.  Predictive models could be separately fit to each cluster/segment, producing predictions tailored to important subgroups.
SELECTION=SCORE (“best subsets selection” in PROC PHREG and PROC LOGISTIC) A variable selection approach that looks at all possible models and ranks them by the score chi-squared statistic. This assesses many more potential models than the stepwise selection methods Forward, Backward, and Stepwise selection.

 

There are many ways to create dummy variables in SAS. Here are two common DATA‑step approaches before we switch to the GLMSELECT method. I’ll dummy code the variable “Origin” from the sashelp.cars data. Origin has three levels: Asia, Europe, and USA.

 

 

Dummy coding in a DATA step

 

The first approach uses IF, THEN, and ELSE statements.  The second approach uses logical tests in which SAS will assign a 1 when the statement in parentheses is true and assign a 0 when the parentheses term is false.

 

data dummy1;
   set sashelp.cars;
   if origin="Asia" then Origin_A=1; else Origin_A=0;
   if origin="Europe" then Origin_E=1; else Origin_E=0;
   if origin="USA" then Origin_U=1; else Origin_U=0;
run;

 

data dummy2;
   set sashelp.cars;
   Origin_A=(origin="Asia");
   Origin_E=(origin="Europe");
   Origin_U=(origin="USA");
run;

 

The above code is not difficult to write, but with 20+ levels it becomes tedious and error‑prone.  GLMSELECT makes this much easier.

 

 

Dummy coding with GLMSELECT

 

To generate dummy variables, create a regression model with PROC GLMSELECT, putting categorical variables in the CLASS and MODEL statements.  Use the OUTDESIGN option in the PROC statement to save the dummy variables to a SAS data set.  If you want to save all the inputs, use the ADDINPUTVARS option for OUTDESIGN to include all the variables in the output data set.  The variable names used in the model will be automatically saved to a macrovariable called _GLSMOD.  %PUT can be used to write the value of this macrovariable to the SAS log.

 

GLMSELECT uses GLM-coding by default, but you can change the coding scheme with the PARAM= option in the CLASS statement.  The SHOWCODING option prints the coding scheme used for each CLASS variable, which is helpful for verifying GLM, reference, or effect-coding. GLMSELECT performs stepwise variable selection by default, so this should be turned off using the SELECTION=NONE option in the MODEL statement.

 

ods select ClassLevelInfo ClassLevelCoding;
proc glmselect data=sashelp.cars outdesign (addinputvars)=x;
   class origin/showcoding;
   model MSRP= origin /selection=none;
run;

proc contents data=x;
run;

proc print data=x (obs=15);
   var "Origin Asia"n "Origin Europe"n "Origin USA"n;
run;

%put &_GLSMOD;       /* writes the variable names from GLMSELECT model to the log */

 

PROC GLMSELECT output:

 

01_TE_blog20-taelna-origin-lvls-and-dummy-vars.png

Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.

 

 

PROC CONTENTS output (without ADDINPUTVARS):

 

02_TE_blog20-taelna-proc-contents-1.png

 

 

PROC CONTENTS output (with ADDINPUTVARS):

 

03_TE_blog20-taelna-proc-contents-long.png

 

 

PROC PRINT output:

 

04_TE_blog20-taelna-origin-dummy-proc-print-obs15.png

 

 

SAS Log:

 

05_TE_blog20-taelna-sas-log2-put.png

 

 

To switch to reference or effect coding, add param=ref or param=effect to the CLASS statement in PROC GLMSELECT:

 

class origin/param=ref showcoding;

 

06_TE_blog20-taelna-ref-coding-dummys.png

 

class origin/param=effect showcoding;

 

07_TE_blog20-taelna-effect-coding-dummys.png

 

Dummy coding doesn’t need to be a separate preprocessing chore. When a SAS procedure won’t accept CLASS variables, you don’t need to hand code anything. PROC GLMSELECT can generate the design matrix for you, and you can pass it directly into your downstream analysis. It’s a small trick that saves time across many analyses.

 

 

Links

 

 

 

Find more articles from SAS Global Enablement and Learning here.

Comments

In SAS, there are other PROCs could do the same thing:

PROC GLMMOD

PROC LOGISTIC

PROC TRANSREG

PROC GLIMMIX

Check Rick's blogs:

https://blogs.sas.com/content/iml/2016/02/24/create-a-design-matrix-in-sas.html

https://blogs.sas.com/content/iml/2016/02/22/create-dummy-variables-in-sas.html

Thanks for the info, especially PROC GLMMOD which I was not familiar with!

GLMMOD is how we have done this for decades, but the truth is, we should be able to get the design matrix out of all the analytic procs. Why not?

Contributors
Version history
Last update:
‎07-23-2026 03:30 PM
Updated by:

Viya Copilot Motion Graphic.gifViya Copilot Motion Graphic

Ready to see what SAS Viya Copilot can do?

Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.

Get Started →

SAS AI and Machine Learning Courses

The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.

Get started