I'm baffled by a "data set X is not sorted" ERROR message. Below is the relevant excerpt of the .LOG file showing (in my view) that the data involved *was* PROC SORTed in the exact way the subsequent PROC RANK step required. The program involved has never generated this error before in processing similar data, and it only did in this particular iteration of the do-loop involved (loop 15 of 37). The only "explanation" I can think of is that I was running 2 SAS programs concurrently on rather large data sets and maybe that affected program performance. But that isn't very convincing. I'll re-run the program and see if it happens again when I'm not having the PC do anything else -- but whether or not the error re-occurs, I'll wonder why it ever occurred. Can anyone suggest an explanation or a line of inquiry about this? Thanks in advance.
MPRINT(MYSWAP): PROC SORT DATA=tmp;
SYMBOLGEN: Macro variable CONSTRAINTVARS resolves to mh1_r start_status_org servseta agegroup gender race hispanic smised
MPRINT(MYSWAP): BY mh1_r start_status_org servseta agegroup gender race hispanic smised swap5 rnd;
NOTE: There were 6636340 observations read from the data set WORK.TMP.
NOTE: The data set WORK.TMP has 6636340 observations and 110 variables.
NOTE: PROCEDURE SORT used (Total process time):
real time 51.99 seconds
cpu time 24.14 seconds
MPRINT(MYSWAP): PROC RANK DATA=tmp OUT=tmpr;
SYMBOLGEN: Macro variable CONSTRAINTVARS resolves to mh1_r start_status_org servseta agegroup gender race hispanic smised
MPRINT(MYSWAP): BY mh1_r start_status_org servseta agegroup gender race hispanic smised swap5 ;
MPRINT(MYSWAP): VAR rnd;
ERROR: Data set WORK.TMP is not sorted in ascending sequence. The current BY group has smised = 3 and the next BY group has smised = 2.
NOTE: The SAS System stopped processing this step because of errors.
NOTE: There were 997665 observations read from the data set WORK.TMP.
WARNING: The data set WORK.TMPR may be incomplete. When this step was stopped there were 0 observations and 0 variables.
WARNING: Data set WORK.TMPR was not replaced because this step was stopped.
NOTE: PROCEDURE RANK used (Total process time):
real time 6.63 seconds
cpu time 0.35 seconds
Take a look at these two notes:
PROC SORT:
NOTE: The data set WORK.TMP has 6636340 observations and 110 variables.
PROC RANK:
NOTE: There were 997665 observations read from the data set WORK.TMP.
I can only assume that the TMP data SORT processes is NOT the same as the TMP dataset RANK processes as the observation count doesn't match. Why would that be? Are we missing some processing between the SORT and the RANK?
No, there's no hidden/unlogged code between the PROC SORT and PROC RANK:
%do v=1 %to 37; [...]
%if &ntoswap NE %then %do;
DATA tmp;
CALL streaminit(98768976);
SET &outsw5;
rnd = RAND("UNIFORM");
PROC SORT DATA=tmp; BY &constraintvars. swap5 rnd; /*constraintvars are f(v)*/
PROC RANK DATA=tmp OUT=tmpr;
BY &constraintvars. swap5 ;
VAR rnd;
PROC SQL;
[...]
%end;
%end;
But the fact that TMP has 2 different record counts at the 2 steps is distinctive, thank you. That didn't happen for any other other iteration of the do-loop (which runs through dwindling lists of variables to sort and rank the records by). It seems to confirm that "the machine got confused" by multiple versions of the TMP data set.
Since then I've re-run the code for the same input data and it ran without the error. So I suspect that it had to do with using the same WORK space for concurrent SAS programs (even the other one didn't write a "competing' TMP data set).
If you can't reproduce the problem then I suggest you "park" it until it reappears 🙂
I would first confirm from the SAS log that there was not another step between those two that would have changed the dataset named WORK.TMP. Especially since the code seems to have been generated by a macro, so perhaps the macro had turned off the writing of notes to the SAS log.
Then I would check what type of disk is being used for the WORK directory. Perhaps there was some type of timing issue with write buffering on the disk.
You mentioned that you were running multiple programs, so make sure that each program is using a separate location for the WORK directory. That should be the default, but it is possible to override the default when the SAS session is started.
There was no intervening code. I'm not sure how to check where a second SAS execution puts its WORK directory; I'd assumed it would be in the same place as the usual one (AppData/Local/Temp/SAS Temporary Files), maybe in a different "_TDblahblahSITELICENSE" subfolder. I guess sometime I'll run 2 long programs at the same time and see what I see.
At any rate, when I re-ran the same code on the same input file, but without any other SAS program running, the program ran without error. This time, TMP had the same number of records after PROC SORT as it did on intake by PROC RANK (see LOG file excerpt below, at the same point in program execution that failed earlier). So I conclude: don't run 2 programs at once, at at least not 2 programs that work on relatively large data sets, at least not on my PC. The one I just ran writes 17MB to the WORK file, the other one is probably similar.
MPRINT(MYSWAP): PROC SORT DATA=tmp;
SYMBOLGEN: Macro variable CONSTRAINTVARS resolves to mh1_r start_status_org servseta agegroup gender race hispanic smised
MPRINT(MYSWAP): BY mh1_r start_status_org servseta agegroup gender race hispanic smised swap5 rnd;
NOTE: There were 6636340 observations read from the data set WORK.TMP.
NOTE: The data set WORK.TMP has 6636340 observations and 110 variables.
NOTE: PROCEDURE SORT used (Total process time):
real time 45.70 seconds
cpu time 26.12 seconds
MPRINT(MYSWAP): PROC RANK DATA=tmp OUT=tmpr;
SYMBOLGEN: Macro variable CONSTRAINTVARS resolves to mh1_r start_status_org servseta agegroup gender race hispanic smised
MPRINT(MYSWAP): BY mh1_r start_status_org servseta agegroup gender race hispanic smised swap5 ;
MPRINT(MYSWAP): VAR rnd;
NOTE: There were 6636340 observations read from the data set WORK.TMP.
NOTE: The data set WORK.TMPR has 6636340 observations and 110 variables.
NOTE: PROCEDURE RANK used (Total process time):
real time 32.38 seconds
cpu time 26.00 seconds
That's a very weird error. If you have time, I would actually try to reproduce it.
Simply having two separate SAS sessions running at the same time should not cause the error you showed, even if SAS runs out of resources.
From your log it looks like the first PROC SORT step silently failed, which I've never seen happen in SAS.
Using multiple concurrent SAS sessions should be fine. Each session gets its own work library. If you want to see the location of the work library, you can run:
%put %sysfunc(pathname(work)) ;
Returns:
1 %put %sysfunc(pathname(work)) ; C:\Users\Quentin\AppData\Local\Temp\SAS Temporary Files\_TD20708_XXXX_
I just realized, the XXXX is the name of my computer. So I guess on windows, the folder name is something like process ID and computer name.
Disk overload is probably the reason then. When SAS makes a new version of an existing dataset (like in your PROC SORT step) what it actually does is make a NEW file and then when the step finishes without error it deletes the old file and renames the file it just wrote.
Perhaps in your case the second step started before the delete/rename actually modified the disk, so it opened the original unsorted dataset.
Note that the reduced number of observations read the the second step was just because it stopped reading when it noticed the sorting error.
@Tom wrote:
Disk overload is probably the reason then. When SAS makes a new version of an existing dataset (like in your PROC SORT step) what it actually does is make a NEW file and then when the step finishes without error it deletes the old file and renames the file it just wrote.
Perhaps in your case the second step started before the delete/rename actually modified the disk, so it opened the original unsorted dataset.
Note that the reduced number of observations read the the second step was just because it stopped reading when it noticed the sorting error.
Put this is a PROC SORT step followed by a PROC RANK step in the same session.
Have you actually seen a case where a following step started executing before the preceding step had completed executing? I've never seen it happen.
I noticed the difference between PROC SORT and PROC RANK:
BY mh1_r start_status_org servseta agegroup gender race hispanic smised swap5 rnd; BY mh1_r start_status_org servseta agegroup gender race hispanic smised swap5 ;
try to remove 'rnd' in your proc sort.
If that still does not work ,try option NOTSORTED:
BY mh1_r start_status_org servseta agegroup gender race hispanic smised swap5 notsorted;
@thomasn528 wrote:
Thanks but no, that's not it. The error message claimed the *prior* nested sort hadn't worked. and had even yielded a differently sized data set!! (as SASKiwi noted.) The "rnd" is listed last, it's needed there, it's the item to be ranked in PROC RANK -- and this step had never malfunctioned before, in dozens and dozens of program runs-times-loop-iterations.
Having re-run the same program on the same data without error this time, this now seems to me to have (indeed) been some kind of resource shortage/malfunction caused by running two demanding programs at the same time on my little pony of a PC.
The highlighted is not what the log shows. It says that the set that Proc Rank attempted to use did not match the required sort. Not that the sort hadn't worked. Pedantic perhaps but there is a difference.
Macros and reusing of same data set names is always something to be a bit cautious of as a failure of one element can leave a data set from a previous run. I often remove "temporary" data sets at the end of a loop so this doesn't occur (and to keep disk space under control sometimes).
@ballardw wrote in part:
The highlighted is not what the log shows. It says that the set that Proc Rank attempted to use did not match the required sort. Not that the sort hadn't worked. Pedantic perhaps but there is a difference.
Macros and reusing of same data set names is always something to be a bit cautious of as a failure of one element can leave a data set from a previous run. I often remove "temporary" data sets at the end of a loop so this doesn't occur (and to keep disk space under control sometimes).
But in this case reusing the same data set name could not be the cause of the issue. The macro generates two consecutive proc statements:
PROC SORT DATA=tmp; BY &constraintvars. swap5 rnd; /*constraintvars are f(v)*/
PROC RANK DATA=tmp OUT=tmpr;
The log reports that the PROC SORT succeeded, and wrote the dataset named tmp. The immediately subsequent PROC RANK step fails, and states that the dataset tmp is not sorted. This is a mystery to me, unless PROC SORT somehow decided to sort the data using a different sortseq than PROC RANK expected.
It's your turn to help shape SAS Innovate 2027. Share your expertise and inspire the SAS community.
Learn how use the CAT functions in SAS to join values from multiple variables into a single value.
Find more tutorials on the SAS Users YouTube channel.
Ready to level-up your skills? Choose your own adventure.