a simple way to generate a validation data set and a training data set with 20/80 split
data rand;
set have;
random=ranuni(0);
run;
data train_plus_validate;
set rand;
original_y=y;
if rand>0.8 then y=.;
run;
proc reg data=train_plus_validate;
model y=x1 x2 x3;
output out=predicteds p=predicted;
run;
By setting y to missing in the validation data set, these observations are not used in creating the regression model, but you do get predicted values for each of them (that's what PROC REG does if the y is missing); you can then see how close the predicted values are to variable ORIGINAL_Y.
--
Paige Miller