flâneur — a map of the web's best reading

Choosing comparators and thresholds - Splink

moj-analytical-services.github.io · 2,060 words · saved by 1 readers

When building a Splink model, one of the most important aspects is defining the Comparisons and Comparison Levels that the model will train on. Each Comparison Level within a Comparison should contain a different amount of evidence that two records are a match, which the model can assign a Match Weight to. When considering different amounts of evidence for the model, it is helpful to explore fuzzy matching as a way of distinguishing strings that are similar, but not the same, as one another. This guide is intended to show how Splink's string comparators perform in different situations in order to help choosing the most appropriate comparator for a given column as well as the most appropriate threshold (or thresholds). For descriptions and examples of each string comparators available in Splink, see the dedicated topic guide. There are three main classes of string comparator that are considered within Splink: where String Similarity Scores are scores between 0 and 1 indicating how simil

Choosing String Comparators ¶ When building a Splink model, one of the most important aspects is defining the Comparisons and Comparison Levels that the model will train on. Each Comparison Level within a Comparison should contain a different amount of evidence that two records are a match, to which the model can assign a match weight. When considering different amounts of evidence for the model, it is helpful to explore fuzzy matching as a way of distinguishing strings that are similar, but not the same, as one another. This guide is intended to show how Splink's string comparators perform in

Explore this link on the map →

related reading