I'm using EM 7.1 and the variable clustering node gives an error if the dataset has over 100,000 observations. I can sample the dataset, and then cluster my variables, but is there an easy way to pass along only the new clustered variables along with the full dataset (train, validate, test) to the models?
On the options for the node, look under "Stopping Criteria". There is an option called "Suppress Sampling Warning". Set it to "Yes" and your problem should be solved.
April 27 – 30 | Gaylord Texan | Grapevine, Texas
Registration is open
Walk in ready to learn. Walk out ready to deliver. This is the data and AI conference you can't afford to miss. Register now and lock in 2025 pricing—just $495!