I am analyzing the effects of air pollution on daily respiratory symptoms in a sample of 47 families. All families included a participating mother and target child. 39 of the 47 families also included a participating father. Each family member completed an evening survey across 56 consecutive days.
RespSx_Sum is my count outcome, representing the sum count of 5 respiratory symptoms endorsed (0=absent, 1=present) by an individual on a given day. I am estimating conditional GLMM’s via PROC GLIMMIX in SAS with random intercepts at the family and individual levels. I control for significant linear and quadratic trends in RespSx_Sum across study days (TIME, TIME2), weekday vs. weekend effects, as well as other meteorological covariates (RHMEAN, TEMPMEAN). My air pollution predictors are continuous and have the same values for all family members within a family on a given day; they are disaggregated to test within- (LAG0_COAQI_PMC) and between-family effects (COAQI_GMC).
Let’s say I want to test whether increases in air pollution on a given day predict higher counts of respiratory symptoms that day and also whether sensitivity to air pollution differs across family roles (mother, father, child).
Using a Poisson distribution in PROC GLIMMIX with LAPLACE approximation indicated mild overdispersion – Pearson X2/df was between 1 – 1.5 in the models. So, I used the following negative binomial model:
PROC GLIMMIX data = DATA.ENVIRO_LONG NOCLPRINT NOITPRINT METHOD=LAPLACE;
CLASS familyid ID ROLE (ref = "1") ; *Reference is mothers;
MODEL RespSx_Sum = ROLE time time2 weekend RHMEAN_GMC TEMPMEAN_GMC LAG0_RHMEAN_PMC LAG0_TEMPMEAN_PMC
COAQI_GMC LAG0_COAQI_PMC LAG0_COAQI_PMC*ROLE / SOLUTION LINK=LOG DIST=NEGBINOMIAL CHISQ cl;
random intercept / sub=familyid ;
random intercept / sub=ID(familyid);
covtest / wald;
RUN;
My question is: how do I account for autocorrelation AR(1) between responses measured closer together in time at the within-person level? And how does this affect the scale parameter? I know I cannot use a repeated statement in this case, and I cannot use RESIDUAL on a random statement using LAPLACE approximation.
Your advice is MUCH appreciated!!!
The short answer is, you might not model AR(1) with method=LAPLACE in PROC GLIMMIX.
Alternatives:
1. fit a random coefficients model including slopes, which indirectly models the correlation involving the time values. For example,
random intercept time / sub=ID(familyid) type=un;
for over-dispersion, you might add random _residual_; statement.
2. Do not use the METHOD=LAPLACE option, use the default RSPL instead.
then you might use something like
random _residual_ / subject=ID(family) type=AR(1);
I do not know the values for TIME, but be aware that AR(1) does not use the actual time values, rather the order of the time values in the computation of the parameters.
For over dispersions, not sure if it is necessary in this case, because the variance in AR(1) might have taken care of that. Or, you might consider using ARH(1).
Hope this helps,
Jill
For over
1)Could you try Generalized Estimating Equations model ? Like:
PROC GENMOD
PROC GEE
2)It looks like you want to do panel data analysis. Check:
PROC PANEL
PROC COUNTREG
can be used for repeated count responses (panel data).
But they are all under SAS/ETS module , suggest you to post it at Forecasting Forum:
https://communities.sas.com/t5/SAS-Forecasting-and-Econometrics/bd-p/forecasting_econometrics
The short answer is, you might not model AR(1) with method=LAPLACE in PROC GLIMMIX.
Alternatives:
1. fit a random coefficients model including slopes, which indirectly models the correlation involving the time values. For example,
random intercept time / sub=ID(familyid) type=un;
for over-dispersion, you might add random _residual_; statement.
2. Do not use the METHOD=LAPLACE option, use the default RSPL instead.
then you might use something like
random _residual_ / subject=ID(family) type=AR(1);
I do not know the values for TIME, but be aware that AR(1) does not use the actual time values, rather the order of the time values in the computation of the parameters.
For over dispersions, not sure if it is necessary in this case, because the variance in AR(1) might have taken care of that. Or, you might consider using ARH(1).
Hope this helps,
Jill
For over
Thank you for your detailed response. Adding a random _residual_ / subject=ID(family) type=AR(1); to my GLMM model seemed to work. However, I'm wondering if this changes the inference space of my model from a conditional GLMM to a marginal quasi-GEE?
It's your turn to help shape SAS Innovate 2027. Share your expertise and inspire the SAS community.
ANOVA, or Analysis Of Variance, is used to compare the averages or means of two or more populations to better understand how they differ. Watch this tutorial for more.
Find more tutorials on the SAS Users YouTube channel.