03-27-2016 10:48 PM
I have a data set with 100k records of customer information.
I'm planning to build some simple rules for matching (ex. firstname + lastname);
So the code must go like matching record 1 to the other records then if it matches then i put the ids of these records in an output table.
Greatly appreciate your advise what are the possible faster approaches to attain these matching of records across a single table.
03-27-2016 11:45 PM
It sounds like you're looking to do data linkages based on identifiers. Here's a tool that has been referenced - though I've never used it.
Statistics Canada offers a tool called G-Link as well, free but they recommend support, you can find it via google.
Additionally, here's a solution that I kind of like that uses a few of the fuzzy matching options.