BookmarkSubscribeRSS Feed
Edie_Moyers
SAS Employee
SAS Data Maker  •  Product Update

New in SAS Data Maker: the Private Evolution method

Posted in: SAS Data Maker | Product Updates
 

If you've been generating synthetic data in SAS Data Maker and wishing for an option that's built with privacy in mind from the ground up, this update is for you. SAS Data Maker now includes the Private Evolution method: a new generator type designed for small to medium single-table data sets, particularly useful when your synthetic data will be used to train downstream classifiers and privacy is a real concern.

What is Private Evolution?

Private Evolution takes an evolutionary approach to building synthetic data. Rather than training a traditional generative model, it works by:

  • Using APIs to generate variations of your data
  • Privately evaluating those variations
  • Keeping the highest-quality samples

This cycle repeats, gradually refining the synthetic data set. One thing worth noting: the output row count is fixed to match your source data, so if you need to generate more or fewer rows than you started with, this may not be the method for you.

The method comes out of recent research by Toan Tran and colleagues (Tran et al. 2026).

Setting it up

Once you've pointed SAS Data Maker at your source data, you'll find the Private Evolution settings on the Training tab. Here's a rundown of what you can configure:

Model type
Specifies the type of generator to use. Select "Private Evolution" to use this method.
Random state
Controls the random number generator. Standard random seed control, shared across model types.
Use differential privacy
Specifies whether to use differential privacy with the Private Evolution method. On by default. This limits how much any single individual's data can influence the trained generator. You'll set an epsilon value here:
~0.01 Extremely high privacy
1–20 Standard privacy needs
20.0 (max) Notably, Private Evolution tends to hold onto high utility even at this upper bound
Maximum number of categories
Caps how much categorical detail is preserved per column. Default is 100; range is 1–100.
Initial mutation rate
Controls how large a "jump" a numerical feature can make, and doubles as the probability that a categorical feature gets swapped out. This decays over the course of training. Valid range: 0.1–0.5. Tune this if you're chasing better downstream classification accuracy.
Mutation rate decay
Specifies how the mutation rate decays. Choose Polynomial (default, and a safe bet for most cases) or Linear.
Variation method
How new candidate rows are generated each round:
  • Random walk — Nudges numerical values by a uniform-random step scaled by the mutation rate; categorical features get resampled at the mutation rate probability.
  • Normal random walk — Same idea, but step sizes follow a Gaussian distribution.
  • Reflected random walk — Like random walk, but instead of clipping numerical values at the boundary, it reflects them back in.
  • Interpolation — Samples pairs of numerical values and interpolates between them.
  • Interpolation and weighted — Same numerical handling as Interpolation, but mutated categorical features are resampled using the weighted distribution from the previous round.

Once your settings are in place, hit Start training. You'll see a Task Performance chart (give it a few seconds to appear) tracking running tasks and memory usage. You can filter it to only show tasks running longer than a threshold you set. Click into any bar for more detail. When training wraps up, you'll land on the Evaluation tab automatically.

If a generator already exists for your project, you'll be prompted about whether to create a new model with your updated settings — choose Make a copy and edit if so.

Picking your settings: what we know so far

Here's the honest answer: there's no reliable heuristic yet for guaranteeing the best settings for your specific use case. A few things do seem to hold up, though:

  • Set epsilon based on the privacy level your use case actually requires — don't just default it
  • Polynomial mutation rate decay is a reasonable starting point for most situations
  • Variation method and initial mutation rate are your main levers for squeezing out better downstream classification performance

To ground this, an ablation study tested Private Evolution on binary classification tasks (using XGBoost with default settings as the downstream model) across a few benchmark data sets:

Data set Best settings Accuracy achieved Real-data upper bound
adult Initial mutation rate 0.5, Normal random walk 0.84 0.87
artificial characters Initial mutation rate 0.3, Random walk 0.68 0.89
person activity Initial mutation rate 0.4, Interpolation and weighted 0.74 0.82

Take these as a starting point for experimentation rather than a formula — results will vary by data set, and some exploration of the mutation rate and variation method space is likely worth your time.

Where to learn more

Interested in trying SAS Data Maker and the Private Evolution method?
Explore SAS Data Maker
Have an idea for how SAS Data Maker could work better for you?
Submit a Product Suggestion