I have a data set with both continuous and categorical variables. I need to find extreme values and replace them as missing values for the continuous variables. I've gotten this far:
/* Calculate Median and IQR */
PROC UNIVARIATE DATA = kddcup98 NOPRINT;
VAR DemAge
DemMedHomeValue
DemMedIncome
DemPctVeterans
GiftAvg36
GiftAvgAll
GiftAvgCard36
GiftAvgLast
GiftCnt36
GiftCntAll
GiftCntCard36
GiftCntCardAll
GiftTimeFirst
GiftTimeLast
PromCnt12
PromCnt36
PromCntAll
PromCntCard12
PromCntCard36
PromCntCardAll
TARGET_D;
OUTPUT OUT = boxStats p25 = p25 p75 = p75 QRANGE = iqr;
RUN;
DATA _null_;
SET boxStats;
CALL symput ('p25',p25);
CALL symput ('p75',p75);
CALL symput ('iqr', iqr);
RUN;
%PUT &p25;
%PUT &p75;
%PUT &iqr;
DATA trimmed;
SET kddcup98;
ARRAY change _numeric_;
DO OVER change;
IF (change > &p75 + 1.5 * &iqr) OR (change < &p25 - 1.5 * &iqr) THEN change = .;
END;
RUN;
/* List Variables with Missing Values */
PROC MEANS DATA=trimmed NMISS N;
TITLE 'trimmed Variables with Number of Missing Values (NMISS) and Number of Numeric Values (N)';
RUN;
The only problem is that is miscalculates the number of extreme values. In some cases, it considers most of the values as extreme.