Anybody have some good resources on using select distinct vs group by? I feel hesitant when using them and find myself confused as to which one to use or if there are times you need to use both.
Thanks!
Paul
Run below and examine the results. May be that explains things to you already a bit more.
In a nutshell: You use DISTINCT to de-duplicate rows, you use GROUP BY to aggregate values by the variables in the group by statement.
data have;
input groupvar nvar cvar $;
datalines;
1 10 A
1 10 A
1 20 B
2 50 B
2 40 B
;
run;
proc sql;
select distinct groupvar, nvar, cvar
from have
;
select groupvar, sum(nvar) as sum_nvar, cvar
from have
group by groupvar
;
select distinct groupvar, sum(nvar) as sum_nvar, cvar
from have
group by groupvar
;
select distinct groupvar, sum(nvar) as sum_nvar, cvar
from have
group by groupvar
having sum(nvar)>80
;
quit;
Run below and examine the results. May be that explains things to you already a bit more.
In a nutshell: You use DISTINCT to de-duplicate rows, you use GROUP BY to aggregate values by the variables in the group by statement.
data have;
input groupvar nvar cvar $;
datalines;
1 10 A
1 10 A
1 20 B
2 50 B
2 40 B
;
run;
proc sql;
select distinct groupvar, nvar, cvar
from have
;
select groupvar, sum(nvar) as sum_nvar, cvar
from have
group by groupvar
;
select distinct groupvar, sum(nvar) as sum_nvar, cvar
from have
group by groupvar
;
select distinct groupvar, sum(nvar) as sum_nvar, cvar
from have
group by groupvar
having sum(nvar)>80
;
quit;
April 27 – 30 | Gaylord Texan | Grapevine, Texas
Walk in ready to learn. Walk out ready to deliver. This is the data and AI conference you can't afford to miss.
Register now and lock in 2025 pricing—just $495!
Still thinking about your presentation idea? The submission deadline has been extended to Friday, Nov. 14, at 11:59 p.m. ET.
Learn how use the CAT functions in SAS to join values from multiple variables into a single value.
Find more tutorials on the SAS Users YouTube channel.
Ready to level-up your skills? Choose your own adventure.