Hello,
I am looking at attrition differences across market segments (attrition is a binary variable, 0 they left, 1 they didn't). All my my results are highly significant - even when I look at differences between two market segments that I know, from past analysis, have minimal differences in attrition. I am wondering if my huge sample size (8,000,000) is overwhelming the test. When I limit the data to a sample of 10,000, the results are more in line with what I have seen in the past, but I only want to limit the sample size if it is more accurate, not because it is giving me the results I want.
My code:
PROC NPAR1WAY WILCOXON DATA = WORK.ATTR_TEMP_SS_T2 ;
CLASS segment1;
VAR att_60day;
RUN;