BookmarkSubscribeRSS Feed

From Data to Clusters (Part 2): Customer Segmentation with PROC KCLUS

Started Wednesday by
Modified Wednesday by
Views 55

In my previous post, From Data to Clusters (Part 1), I used PROC KCLUS to explore k‑means clustering on a small, interpretable dataset derived from the 2000 U.S. Census. That example showed how descriptive statistics can guide standardization choices and how k‑means assigns observations to clusters. In this second part, I apply k‑means clustering to a more complex business scenario using CommsData, a telecommunications dataset with many predictors and a clear analytic goal: segmenting customers to support marketing and predictive modeling. The focus here is on selecting meaningful variables for clustering and interpreting the results when the two‑dimensional plots are messy or only partially separable.

 

 

A Quick Review of K-means Clustering

 

K‑means clustering groups observations based on similarity across a set of input variables, so choosing appropriate inputs is essential. Useful predictors are typically continuous or, if categorical, have several levels, and ideally are not strongly correlated with one another. Because k‑means relies on distance calculations, the variables must be placed on comparable scales; otherwise, predictors with large numeric ranges will dominate the clustering. Standardization is commonly used for this purpose, and when variables differ dramatically in scale, normalizing them to a [0, 1] range can be more appropriate. These considerations help ensure that the clusters reflect meaningful patterns in the data rather than artifacts of variable scaling.

 

 

Challenges in Business Clustering

 

Businesses often use clustering techniques such as k‑means to identify segments within their customer base. Segmentation refers to grouping customers into meaningful categories that support decisions such as marketing, retention, and predictive modeling. These segments may come from an algorithm, or they may be defined manually, for example by combining age ranges with income tiers or grouping customers by total annual purchases. When clustering is used to create segments, the challenge is choosing variables that are not only appropriate for k‑means—relatively independent, continuous when possible, and placed on comparable scales—but also aligned with the business goals the segments are meant to support. Business datasets typically contain many potential predictors, and selecting the ones that reflect how the company understands customer behavior is crucial for producing clusters that are both interpretable and useful.

 

 

About CommsData

 

CommsData is a telecommunications dataset that contains detailed information about customers, including demographics, product usage, billing behavior, and interactions with customer service. With 128 predictors describing multiple aspects of customer behavior, CommsData provides an example of the complexity that often arises in business datasets. The dataset is currently available in SAS Viya for Learners at ~/Courses/CPML/commsdata.sas7bdat.

 

 

Selecting Variables for Clustering

 

When clustering customer data, the choice of input variables has a major influence on the usefulness of the resulting segments. In this example, the goal is to group customers into segments that can support marketing and predictive modeling, either as features in a model or as stratification variables. For a telecommunications provider, variables related to customer lifecycle and value, usage patterns, dissatisfaction signals, payment behavior, and basic demographics often reveal meaningful differences in how customers interact with the company. These ideas guided the selection of the twelve inputs used here. The selected variables represent several important aspects of customer behavior, including tenure, usage, value, customer service interactions, delinquency, income, and urbanicity. Exploratory analyses (not shown) confirmed that these variables have very different scales, are not strongly correlated, and include some missing values.

 

The twelve variables used in this example are summarized in the table below.

 

Variable Description Why use for segmentation?
acct_age number of months the account has been active captures customer lifecycle and value
lifetime_value customer's value captures customer lifecycle and value
avg_arpu_3m average revenue for the past 3 months captures customer lifecycle and value
voice_total_bill_mou_curr current minutes of voice billed describes usage (voice and data)
tot_mb_data_curr current MB data used describes usage (voice and data)
nbr_contracts number of contracts the customer made to the company describes customer dissatisfaction
calls_TS_acct number of tech support calls describes customer dissatisfaction
unsolv_tsupcomplnts number of unsolved tech support complaints describes customer dissatisfaction
delinq_indicator Delinquency indicator (scale: -2 to +4) describes financial stress and payment behavior
times_delinq consecutive months in default describes financial stress and payment behavior
Est_HH_Income household income demographic data
cs_ttl_urban urban population in customer's area demographic data

 

 

Using PROC KCLUS for K-Means Clustering

 

PROC KCLUS performs k‑means clustering in SAS Viya and is designed for wide, complex datasets like CommsData. It standardizes variables, handles missing values through imputation, selects initial cluster centers, and can automatically choose the number of clusters using the aligned box criterion (ABC). After selecting meaningful predictors, PROC KCLUS is run to create the clusters, which can then be examined using both numerical output and visualizations.

 

The following PROC KCLUS step runs the clustering using the twelve selected variables and lets the aligned box criterion (ABC) determine the number of clusters.

 

/* Cluster using 12 inputs, use ABC to find number of clusters */

proc kclus data=blog.commsdata impute=mean standardize=range noc=ABC
    (Align=none) init=forgy;
   input acct_age lifetime_value avg_arpu_3m voice_tot_bill_mou_curr
    tot_mb_data_curr nbr_contacts calls_TS_acct unsolv_tsupcomplnt
    delinq_indicator times_delinq Est_HH_Income cs_ttl_urban
    /level=interval;
   score out=blog.commsdata_clusters copyvars=(_all_);
run;

 

Results

 

PROC KCLUS produces several intermediate tables, including descriptive statistics, standardization details, and within‑cluster statistics. For a 12‑variable clustering problem, these add little practical value, so the focus here is on the key model results: the chosen number of clusters, the cluster assignments, and the visualizations used to explore them.

 

taelna-blog22-1.png

 

Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.

 

The Model Information table shows that all 56,557 observations were included. K‑means clustering used range standardization, mean imputation, and Forgy initialization. The aligned box criterion (ABC) selected four clusters, and the algorithm converged within the maximum of ten iterations.

 

taelna-blog22-2.png

The aligned box criterion (ABC) evaluates different numbers of clusters and chooses the value that provides the best separation between groups. The ABC Parameters table shows that the procedure compared solutions with 2 through 6 clusters. For each value of k, PROC KCLUS computes the within‑cluster sum of squared errors (SSE) for the actual data (“Input”) and for a reference distribution (“Reference”) that represents data with no cluster structure.

 

The key column is the “Gap,” which measures how much better the clustering is compared to the reference distribution. Larger gap values indicate clearer separation between clusters. In this example, the gap statistic increases from 2 to 4 clusters and then declines, meaning that cluster separation is strongest at k = 4.

 

The Estimated Number of Clusters table confirms this result: the GlobalPeak criterion selects four clusters for CommsData. This is the value PROC KCLUS uses for the final cluster assignments.

 

taelna-blog22-3.png

The descriptive statistics highlight why standardization is essential for k‑means clustering. The twelve variables have very different scales from household income in the tens of thousands to contact counts that average less than one, and their standard deviations vary just as widely. Without standardization, variables with large numeric ranges (such as voice minutes or lifetime value) would dominate the distance calculations and overwhelm the smaller‑scale variables. These summaries also show that several variables have substantial variability, meaning customers differ widely on these measures. When variables spread out this much, the clusters can appear to overlap in two‑dimensional plots used for interpretation, even though k‑means still assigns each customer to exactly one cluster.

 

PROC KCLUS also produces within‑cluster statistics for each variable, but these values are range‑standardized and not directly interpretable for a 12‑variable clustering problem, so they are omitted here. The procedure also outputs the location and scale used for range standardization, but these details are mainly procedural and do not affect cluster interpretation, so they are not shown.

 

taelna-blog22-5.png

The Cluster Summary table shows that the four clusters differ mainly in size, as indicated by the Frequency column: Cluster 1 is the largest group with 30,795 customers, and Cluster 2 is the smallest with 6,520. The minimum, maximum, average, and standard deviation values are aggregated across all twelve standardized variables, so they do not reveal variable-specific patterns. The distances to the nearest cluster centroids are all of similar magnitude, which simply indicates that no cluster is an extreme outlier relative to the others. To understand how the clusters differ on individual dimensions, we turn to selected two‑variable plots next.

 

 

Creating a Subsample for Cluster Visualization

 

Scatterplots with tens of thousands of observations are difficult to interpret because dense regions of points obscure the underlying structure. To make the cluster visualizations readable, we draw a stratified random sample of 10% of the data. Stratified sampling ensures that each cluster contributes a proportional number of observations to the sample, so the plots reflect the true cluster distribution.

 

PROC PARTITION is well‑suited for sampling CAS tables in SAS Viya, including stratified random sampling. The variable _CLUSTER_ID_ is automatically created by PROC KCLUS and contains the cluster assignment for each observation. By sampling within each value of _CLUSTER_ID_, we preserve the relative cluster sizes while reducing the data to a manageable number of points for plotting.

 

/* Sample 10% of the data for plotting using stratified sampling */  

proc partition data=blog.commsdata_clusters samppct=10 seed=12345;
   by _CLUSTER_ID_;
   output out=blog.comms_sample copyvars=(acct_age lifetime_value 
    avg_arpu_3m voice_tot_bill_mou_curr tot_mb_data_curr nbr_contacts 
    calls_TS_acct unsolv_tsupcomplnt delinq_indicator times_delinq 
    Est_HH_Income cs_ttl_urban product_plan_desc handset handset_age_grp 
    _CLUSTER_ID_);
run;

 

Results

 

taelna-blog22-6.png

The stratified sampling output lists each cluster, the number of observations in the full data, and the number selected for the 10% sample. Because sampling is done within each value of _CLUSTER_ID_, every cluster contributes a proportional number of customers to the sample. The resulting CAS table, comms_sample, contains 5,656 observations and will be used for the scatterplots that follow.

 

With the stratified sample prepared, we can now to visualize how the clusters differ across selected business dimensions. These plots compare pairs of variables that represent tenure, usage, value, customer service friction, delinquency, and income. Each scatterplot uses color to indicate cluster membership, allowing us to see how the clusters appear when projected into two‑dimensional space.

 

/* Tenure vs usage */

proc sgplot data=blog.coms_sample;
   scatter x=acct_age y=tot_mb_data_curr /group=_CLUSTER_ID_ transparency=0.7 markerattrs=
    (symbol=CircleFilled size=4);
   ellipse x=acct_age y=tot_mb_data_curr /group=_CLUSTER_ID_;
run;

/* Value vs usage */

proc sgplot data=blog.coms_sample;
   scatter x=lifetime_value y=tot_mb_data_curr/group=_CLUSTER_ID_ transparency=0.7 markerattrs=
    (symbol=CircleFilled size=4);
   ellipse x=lifetime_value y=tot_mb_data_curr/group=_CLUSTER_ID_;
run;

/* Customer service contacts vs unresolved complaints */

proc sgplot data=blog.coms_sample;
   scatter x=nbr_contacts y=unsolv_tsupcomplnt /group=_CLUSTER_ID_ transparency=0.7 markerattrs=
    (symbol=CircleFilled size=4);
   ellipse x=nbr_contacts y=unsolv_tsupcomplnt /group=_CLUSTER_ID_;
run;

/* Delinquency vs. value */

proc sgplot data=blog.coms_sample;
   scatter x=delinq_indicator y=lifetime_value /group=_CLUSTER_ID_ transparency=0.7 markerattrs=
    (symbol=CircleFilled size=4);
   ellipse x=delinq_indicator y=lifetime_value /group=_CLUSTER_ID_;
run;

 
/* Income vs. urbanicity */

proc sgplot data=blog.coms_sample;
   scatter x=Est_HH_Income y=cs_ttl_urban/group=_CLUSTER_ID_ transparency=0.7 markerattrs=
    (symbol=CircleFilled size=4);
   ellipse x=Est_HH_Income y=cs_ttl_urban/group=_CLUSTER_ID_;
run;

 

Results

 

taelna-clus2-1.png

taelna-clus2-2.png

taelna-clus2-3.png

taelna-clus2-4.png

taelna-clus2-5.png

Interpreting the Scatterplots

 

Across the five scatterplots, three variable pairs—tenure vs. usage, value vs. usage, and delinquency vs. value—show substantial overlap, with no visible cluster separation along these individual dimensions. The customer service friction plot reveals a modest pattern: clusters 1 and 4 contact tech support at similar rates, but cluster 4 has more unresolved complaints, a difference that may be relevant for churn modeling. The income vs. urbanicity plot is the only one that shows clear separation. Incomes are broadly similar across clusters, but cluster 2 occupies more rural areas, clusters 1 and 4 occupy the most urban areas, and cluster 3 falls in between.

 

 

Why the Scatterplots Look Messy

 

Even though each customer is assigned to one cluster, the clusters overlap when projected into two‑dimensional scatterplots. This happens because the clustering is based on twelve variables, and compressing that multidimensional structure into a single pair of axes hides much of the separation. Clusters often differ on combinations of variables rather than any single measure, so a group may be defined by a pattern across several variables even if no individual variable isolates it visually. Range standardization also equalizes scales for distance calculations, but it does not guarantee that clusters will appear distinct on any particular axis.

 

In this example, most variable pairs show substantial overlap, and only the income versus urbanicity plot reveals clear separation. This does not mean the clusters are uninformative; it simply reflects the fact that customer behavior varies along many dimensions at once, and two‑dimensional projections capture only a small part of that structure.

 

 

How the Clusters Can Still Be Useful

 

Even when the scatterplots appear messy, the clusters still summarize patterns in the full feature space that are not visible in any single plot. The clustering is based on twelve variables, and the algorithm groups customers according to similarities across all of them, not just the two shown in each scatterplot. Cluster membership can therefore be used as a feature in downstream models, where the model can learn how the multidimensional patterns associated with each cluster relate to outcomes such as churn or retention.

 

Cluster membership can also be used as a stratification variable. In this approach, separate predictive models are fit within each cluster, allowing the modeling process to adapt to the behavioral patterns specific to each subgroup. When customer behavior varies along many dimensions at once, stratified models can sometimes provide more accurate predictions than a single model fit to the entire population.

 

And when a variable pair does show partial separation — as with income and urbanicity — that information can provide additional context for understanding how the clusters differ.

 

 

Conclusion

 

Clustering customer data is ultimately about understanding behavior across many dimensions. In this post, the focus was on selecting business‑relevant predictors and interpreting clusters when the two‑dimensional plots are messy or only partially separable. Even when the visuals are imperfect, the clusters still capture meaningful structure in the full feature space. That structure can support modeling, stratified analysis, and other downstream work where multidimensional patterns matter. Clustering is not about finding visually obvious groups; it is about discovering patterns that help you understand and act on customer behavior.

 

 

Links

 

From Data to Clusters (Part 1): A Simple Example Using PROC KCLUS

 

 

Find more articles from SAS Global Enablement and Learning here.

Contributors
Version history
Last update:
Wednesday
Updated by:

Viya Copilot Motion Graphic.gifViya Copilot Motion Graphic

Ready to see what SAS Viya Copilot can do?

Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.

Get Started →

SAS AI and Machine Learning Courses

The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.

Get started

Article Tags