@Stephen_Shell
Best practice for this forum is that you start your own post and possibly reference this one as related. Readers wont have to wade through the original posts to find the details of your question and responses to questions about your problem. Also when a solution is reached you can indicate which answer fits your problem.
Also, you want to post code in a code box opened with the </> or "running man" icon on this forum. Pasting code in the main message will get reformatted by the forum software. As posted, your data step reading the data throws errors because column 22, which is supposed to part of the Company from your input statement actually includes the start of the Debt variable and Town starts in column 31 not 39. The forum likely removed repeated blanks. An example from the log running your posted code:
1 Data Company; input Company $ 1-22 Debt 25-30 Number 33-36 Town $
1 ! 39-51;
2 datalines;
NOTE: Invalid data for Number in line 3 33-36.
RULE: ----+----1----+----2----+----3----+----4----+----5----+----6----
3 Ice Cream Delight 299.98 2310 Holly Springs
Company=Ice Cream Delight 299. Debt=2310 Number=. Town=rings _ERROR_=1
_N_=1
NOTE: Invalid data for Number in line 4 33-36.
4 Ice Cream Delight 299.98 2210 Holly Springs
Company=Ice Cream Delight 299. Debt=2210 Number=. Town=rings _ERROR_=1
_N_=2
NOTE: Invalid data for Number in line 5 33-36.
5 Ice Cream Delight 299.98 2310 Holly Springs
Company=Ice Cream Delight 299. Debt=2310 Number=. Town=rings _ERROR_=1
_N_=3
NOTE: Invalid data for Number in line 6 33-36.
6 Ice Cream Delight 300.98 2310 Holly Springs
Company=Ice Cream Delight 300. Debt=2310 Number=. Town=rings _ERROR_=1
_N_=4
NOTE: Invalid data for Number in line 7 33-36.
7 Ice Cream Delight 300.98 2510 Holly Springs
Company=Ice Cream Delight 300. Debt=2510 Number=. Town=rings _ERROR_=1
_N_=5
NOTE: Invalid data for Number in line 8 33-36.
8 Ice Cream Delight 300.98 2510 Nolly Springs
Company=Ice Cream Delight 300. Debt=2510 Number=. Town=rings _ERROR_=1
_N_=6
NOTE: Invalid data for Number in line 9 33-36.
9 Ice Cream Delight 300.98 2510 Nolly Springs
Company=Ice Cream Delight 300. Debt=2510 Number=. Town=rings _ERROR_=1
_N_=7
NOTE: Invalid data for Number in line 10 33-36.
10 Ice Cream Delight 300.98 2510 Nolly Springs
Company=Ice Cream Delight 300. Debt=2510 Number=. Town=rings _ERROR_=1
_N_=8
NOTE: The data set WORK.COMPANY has 8 observations and 4 variables.
NOTE: DATA statement used (Total process time):
real time 0.06 seconds
cpu time 0.00 seconds
Hint: at least for examples on this forum it may be better to use simple input statement with numeric values instead of specifying columns.
Then provide a few more details as to what exactly you need for output.
You say your problem is "similar" to the topic in this thread but are not particularly clear about the differences other than mentioning "51 variables".
There would be no way for us to determine what an "error" might be without rules defining what an error is.
Differences might be easier using Proc Compare which is designed for comparing two data sets (or one set with itself).
I would submit that a variable such as "debt" (or price, balance, number in inventory etc) is very likely to be time dependent and comparisons likely should be compared based on some date measure.
If you are looking for potential differences in Number (an address component? store indentification? descriptions help) and City where you expect only one value of Number per City for a given country a diagnostic tool that may be helpful is proc freq with the list option:
Consider:
Proc freq data=company;
table company *town *number/list;
run;
The company data set would not require sorting though if very large sorting might improve the run time for the proc freq.
The LIST option gives a count of the combinations of the variables on a single line and you can tell how many values of NUMBER you get for each combination of Company and Town easily.
You could also inlcude
table company * number * town / list;
Which would clearly show if number was associated with multiple towns.
I had a data source that had, and required reporting on by, city, county and Zipcode. However the data entry folks were more than a bit sloppy and would enter the county in the city field, or the State in the County field. The state name was also the name of one of the counties and several county names were the same as cities, a not uncommon occurence. So I used this technique to find counties not associated with the city name and vice versa. And the Zipcode because those would often have reversed digits or other inconsistencies. This step using proc freq was after use the SASHELP.ZIPCODE data set to identify likely errors involving Zipcodes. Caution: Zipcodes may cross county lines and frequently do in rural areas.
... View more