BookmarkSubscribeRSS Feed

A Hands-On Introduction to Local Model Interpretability with PROC LIME

Started ‎08-03-2026 by
Modified ‎08-03-2026 by
Views 94

Modern machine learning models such as gradient boosting, forests, neural networks, and ensemble models often deliver highly accurate predictions. However, these models are frequently criticized for being "black boxes" because it can be difficult to understand why a particular prediction was made. To address this challenge, SAS Viya provides the LIME procedure (PROC LIME), which implements the Local Interpretable Model-Agnostic Explanations (LIME) methodology introduced by Ribeiro, Singh, and Guestrin (2016). PROC LIME helps analysts explain individual predictions by approximating the behavior of a complex machine learning model in the neighborhood of a specific observation.

This post introduces the fundamentals of LIME, explains how PROC LIME works in SAS Viya, and demonstrates its use with a practical example.

 

 

Why Do We Need Local Explanations?

 

Suppose a machine learning model predicts that a customer has a 92% likelihood of responding to a marketing campaign. While the prediction itself is useful, business users often want answers to questions such as:

 

  • Which variables most influenced this prediction?
  • Did income increase or decrease the predicted response?
  • How important was customer age?
  • Why did this customer receive a different prediction than another customer?

 

While overall model performance provides a useful summary of how well the model works, it does not reveal why the model made a particular prediction for an individual observation. LIME fills this gap by generating local explanations that describe the model's behavior near a specific observation.

 

 

What is LIME?

 

Local Interpretable Model-Agnostic Explanations (LIME) is a technique that explains the predictions of any machine learning model by fitting a simple, interpretable model around the observation of interest. Instead of attempting to explain the entire machine learning model, LIME focuses on a small neighborhood around the query observation and answers the question: "What factors drove the prediction for this specific observation?"

The method is model-agnostic, meaning it can be used with:

 

  • Gradient Boosting models
  • Forest models
  • Neural Networks
  • Support Vector Machines
  • Ensemble models
  • Any model that can generate predictions

 

LIME functionality in SAS isn't limited to a single interface. Data scientists who prefer writing code can use PROC LIME to generate local explanations, while users working in SAS Model Studio can access similar functionality directly within their machine learning projects. Regardless of how it is accessed, the goal remains the same: to help uncover the factors driving a specific prediction and make complex models easier to understand.

 

 

How PROC LIME Works

 

PROC LIME explains predictions by constructing a local surrogate model around the observation being investigated.

The procedure follows four key steps.

 

Step 1: Generate Data Around the Query Observation

 

The process begins with a query observation whose prediction needs to be explained. PROC LIME generates a large number of synthetic observations near the query point.

For Interval Variables: Values are sampled from a normal distribution-

 

  • Centered at the query observation value
  • Using the variance observed in the reference data

 

For example, if a customer's age is 45, PROC LIME generates nearby ages such as: 42, 44, 47, 50 while respecting the variability present in the training data.

For Nominal Variables: Values are sampled from the empirical distribution of the reference data. For example, if a variable has categories: Gold, Silver, and Bronze then the new observations are generated by randomly sampling these categories according to their observed frequencies.

The number of generated observations is controlled by the SAMPLESIZE= option.

 

Step 2: Calculate Distances and Weights

 

Not all generated observations contribute equally to the explanation. Observations that are closer to the query point receive larger weights, while observations farther away receive smaller weights. PROC LIME calculates a mixed distance measure using:

 

Interval Variables: A normalized Euclidean distance is computed using the standard deviation of each variable. This standardization ensures variables with larger scales do not dominate the distance calculation.

Nominal Variables: A Hamming distance is used to measure mismatches between categorical values.

Combined Distance: The interval and nominal distances are combined into a mixed-distance metric. The observation weights are then calculated using an exponential kernel:

 

  • Nearby observations receive high weights.
  • Distant observations receive low weights.

 

This weighting mechanism ensures that the explanation focuses on the local region around the query observation.

Two options can be used to customize distance calculations:

 

  • EXPONENTIALKERNEL=
  • MIXEDDISTANCEWEIGHT=

 

In most applications, the default settings work effectively.

 

Step 3: Score the Generated Data

 

The generated observations are then scored using the machine learning model being explained. For each synthetic observation:

 

  • Predictor values are supplied to the model.
  • Predicted values are obtained.
  • These predictions become the target for the surrogate model.

 

As a result, the surrogate model learns to mimic the behavior of the original machine learning model in the local neighborhood.

 

Step 4: Fit a Local Surrogate Model

 

PROC LIME fits a weighted linear regression model using the generated observations.

The surrogate model:

 

  • Uses observation weights from the distance calculation.
  • Uses the machine learning model predictions as the response variable.
  • Applies a LASSO penalty for variable selection.

 

The resulting regression coefficients become the LIME values.

These coefficients quantify how sensitive the prediction is to changes in each input variable near the query observation.

 

 

Example: Explaining a Neural Network Model in SAS Viya

 

In this example, the HMEQ dataset, which contains information about mortgage applicants, is used to build a predictive model. The target variable, BAD, indicates whether an applicant eventually paid off the loan or defaulted after receiving loan approval. A neural network model is trained to estimate the probability of default and is subsequently saved as an analytic store (ASTORE) for scoring and deployment. The code used to create and store the model is reproduced below for reference and convenience.

 

proc nnet data=casuser.hmeq standardize=midrange missing=mean;

   architecture mlp;

   input reason / level=nominal;

   input debtinc delinq loan mortdue value yoj derog clage clno;

   hidden 50;

   target bad / level=nominal;

  optimization algorithm=lbfgs maxiter=500;

   train outmodel=casuser.nnetmodel1 seed=12345;

   partition fraction(validate=0.3 seed=54321);

   savestate rstore=casuser.nnet_astore;

run;

 

Once the model has been successfully trained and stored, the next step is to identify the customer record whose prediction you want to explain. This record, often referred to as the query observation, is loaded into a query table and provided as input to PROC LIME.

 

Step 1: Load the Query Observation

 

data casuser.query;

   set casuser.hmeq;

   keep clno debtinc delinq derog loan mortdue ninq value yoj reason clage bad;

   if _N_ = 10;

run;

 

PROC LIME then generates a set of perturbed observations around the query observation and fits a simple, interpretable surrogate model in the local neighborhood of that observation. The resulting local model approximates the behavior of the original neural network for that specific customer and helps explain which variables contributed to increasing or decreasing the predicted probability of default.

 

Step 2: Run PROC LIME

 

ods output ParameterEstimates=casuser.Lime_parms;

proc lime

   data=casuser.query

   referenceData=casuser.hmeq samplesize=5000  seed=12345;

   input clno debtinc delinq derog loan mortdue ninq value yoj clage / level=interval;

   input reason /level=nominal;

   predictedTarget P_Bad1;

   astoreModel rstore=casuser.nnet_astore;

run;

 

The PROC LIME step uses the records in CASUSER.HMEQ as a reference population to create 5,000 perturbed observations around the customer record being explained. These synthetic observations are then scored using the previously trained neural network model stored in the analytic store nnet_astore. Based on the resulting predictions, LIME fits a simple and interpretable local surrogate model that mimics the behavior of the neural network in the vicinity of the selected customer. This local model helps reveal which input variables are driving the prediction and whether they are increasing or decreasing the customer's predicted probability of loan default.

The PREDICTEDTARGET statement identifies P_Bad1 as the variable containing the predictions generated by the neural network model. The ASTOREMODEL statement specifies the analytic store that contains the model to be explained. Finally, the ODS OUTPUT statement saves the parameter estimates from the local surrogate model to the Lime_parms table, enabling further examination of feature importance and the creation of visual explanations such as contribution plots and LIME importance charts. The successful execution of the code produces several output tables that provide insights into the LIME explanation.

The Explainer Information table summarizes the settings that PROC LIME used to generate the explanation.

 

01_MS_Explainer-Info.png

 

LIME created a local explanation by generating 5,000 synthetic observations around the query observation (Query Centered). Similarity was measured using a Normalized Euclidean distance with an Exponential Kernel ensuring that observations closest to the query observation had the greatest influence on the explanation. The predictions from the neural network were then approximated using a LASSO regression model, which serves as the local surrogate model. The random seed was used to ensure that the results are reproducible.

 

The Parameter Estimates table contains the coefficients from the local surrogate model that LIME built to explain the neural network's prediction for the selected customer. Each coefficient indicates how a variable influences the prediction in the neighborhood of that specific observation. A positive coefficient suggests that higher values of the variable push the prediction upward, while a negative coefficient suggests that the variable reduces the predicted risk.

 

02_MS_Parameter-Estimates.png

 

For this customer, the intercept of 0.197 represents the baseline prediction from the local surrogate model. The parameter estimates show how changes in each predictor affect the prediction in the vicinity of this observation. Positive coefficients are associated with an increase in the predicted probability of default, while negative coefficients are associated with a decrease. Because the predictors are measured on different scales, the coefficients should not be interpreted as a ranking of variable importance. To understand which factors actually influenced this customer's prediction, it is more useful to examine each variable's contribution, which reflects both the coefficient and the customer's observed value. These contributions provide a clearer picture of how the prediction was formed.

 

Looking at this customer's values, several factors contribute positively to the prediction. The customer's DEBTINC (debt-to-income ratio), YOJ (years on the job), MORTDUE (mortgage balance), VALUE (property value), and CLAGE (credit age) all have positive parameter estimates and nonzero values, meaning they each add to the predicted probability of default in the local LIME model. On the other hand, although DELINQ and DEROG are associated with positive parameter estimates, this customer has a value of 0 for both variables. As a result, they do not contribute to this specific prediction.

 

Not all variables push the prediction upward. For example, CLNO (Number of Credit Lines) has a negative parameter estimate, suggesting that a higher number of credit lines is associated with a slight reduction in the predicted risk within the local explanation. Among the categorical predictors, REASON=DebtCon also has a negative parameter estimate, indicating that debt-consolidation loans are associated with a lower predicted risk relative to the reference category. However, this effect does not apply to the current customer because the loan purpose is HomeImp (Home Improvement) rather than DebtCon.

 

Overall, the table shows that the customers predicted probability of default is primarily influenced by factors such as DEBTINC, YOJ, MORTDUE, VALUE, and other aspects of the customer's credit profile. While DELINQ and DEROG are important drivers in the local model, they do not affect this prediction because the customer has no delinquencies or derogatory reports. By combining the parameter estimates with the customer's actual values, LIME provides a clear and intuitive explanation of how the neural network arrived at its prediction for this individual customer.

 

The local LIME model can be expressed approximately as:

 

03_MS_Formula.png

 

Select any image to see a larger version.
Mobile users: To view the images, select the "Full" version at the bottom of the page.

 

 

Approximate Feature Contributions

 

Multiplying estimate with query value gives an indication of impact:

 

Variable Approx. Contribution
DEBTINC 0.00020085 × 38.2636 ≈ 0.0077
YOJ 0.001104 × 12 ≈ 0.0132
CLAGE 0.0000115 × 70.49 ≈ 0.0008
CLNO -0.000151 × 21 ≈ -0.0032
LOAN 2.38E-7 × 2400 ≈ 0.0006
MORTDUE 5.96E-8 × 34863 ≈ 0.0021
VALUE 3.43E-8 × 47471 ≈ 0.0016

 

Finally, the Explainer Fidelity table shows how closely the LIME surrogate model matches the prediction from the original neural network. In this example, the original model predicted 0.21965, while the LIME model predicted 0.21989. The two values are nearly identical, and the very small RMSE of 0.00037 indicates an excellent local fit. This gives us confidence that the feature effects identified by LIME provide a reliable explanation of the prediction for this customer.

 

04_MS_Fidelity.png

 

After discussing the numerical results, we can create a LIME contribution plot to visualize how individual variables influenced the prediction for the selected customer.

 

The code first creates a new variable called Contribution by multiplying each parameter estimate from the local surrogate model by the corresponding query value. This calculation converts the model coefficients into actual feature contributions for the observation being explained. A positive contribution indicates that the variable increases the predicted probability, whereas a negative contribution indicates that it decreases the prediction. To make the output easier to interpret, the code also classifies each contribution as Positive, Negative, or Neutral.

 

data casuser.lime_contrib;

    set casuser.lime_parms;

    if not missing(REASON) then

        VariableName = catx('=', Variable, REASON);

    else

        VariableName = Variable;

    Contribution = Estimate * QueryValue;

        if Contribution > 0 then Effect='Positive';

    else if Contribution < 0 then Effect='Negative';

    else Effect='Neutral';

run;

 

Finally, PROC SGPLOT is used to create a horizontal bar chart in which bars extending to the right of zero represent variables that increase the prediction, while bars extending to the left represent variables that decrease it. The reference line at zero clearly separates positive and negative influences, making it easy to identify the key drivers behind the model's prediction.

 

title "LIME Contributions for Query Observation";

proc sgplot data=casuser.lime_contrib;

    hbarparm category=VariableName

             response=Contribution /

             group=Effect

             datalabel;

    refline 0 / axis=x lineattrs=(thickness=2);

    styleattrs datacolors=(yellow blue green);

    xaxis label="Contribution to Prediction";

    yaxis discreteorder=data display=(nolabel);

run;

title;

 

This visual representation complements the parameter estimates table by providing an intuitive view of the factors that contributed most strongly to the prediction for the selected customer.

 

05_MS_LIME-Plot.png

 

For this observation, the largest active contributors appear to be:

 

  1. YOJ (Years on Job) → positive
  2. DEBTINC (Debt-to-Income Ratio) → positive
  3. CLNO (Number of Credit Lines) → negative
  4. MORTDUE → positive
  5. VALUE → positive

 

When Should You Use PROC LIME?

LIME is particularly valuable when you need to understand the reasoning behind an individual prediction rather than the overall behavior of a model. It can help explain why a specific customer was classified as high risk, reveal the factors driving an unexpected prediction, and provide transparency into complex machine learning models.

 

 

Real-World Applications of LIME

Because LIME can be used with virtually any machine learning model, it has applications across many industries. In banking and financial services, it can explain credit risk assessments and loan approval decisions. In fraud detection, it helps analysts understand why a transaction was flagged as suspicious. Organizations also use LIME to explain customer churn predictions, marketing response models, insurance claim outcomes, and healthcare risk assessments.

 

Regardless of the application, the goal remains the same: to make individual predictions more transparent, interpretable, and actionable for decision-makers.

 

For more on model interpretability: SAS Tutorial | Interpreting Machine Learning Models in SAS

 

Model Interpretability for Models with Uninterpretable Features

 

References

 

 

Find more articles from SAS Global Enablement and Learning here.

Contributors
Version history
Last update:
‎08-03-2026 01:21 PM
Updated by:

Viya Copilot Motion Graphic.gifViya Copilot Motion Graphic

Ready to see what SAS Viya Copilot can do?

Visit the Tips & Tricks page for setup guidance, demos, and practical examples that show how Copilot supports your workflows.

Get Started →

SAS AI and Machine Learning Courses

The rapid growth of AI technologies is driving an AI skills gap and demand for AI talent. Ready to grow your AI literacy? SAS offers free ways to get started for beginners, business leaders, and analytics professionals of all skill levels. Your future self will thank you.

Get started

Article Tags