flâneur — a map of the web's best reading

Defining and customising comparisons - Splink

moj-analytical-services.github.io · 1,830 words · saved by 1 readers

A key feature of Splink is the ability to customise how record comparisons are made - that is, how similarity is defined for different data types. For example, the definition of similarity that is appropriate for a date of birth field is different than for a first name field. By tailoring the definitions of similarity, linking models are more effectively able to distinguish beteween different gradations of similarity, leading to more accurate data linking models. Note that for performance reasons, Splink requires the user to define n discrete levels (gradations) of similarity. Comparisons are defined on pairwise record comparisons. Suppose for instance your data contains first_name and surname and dob: To compare these records, at the blocking stage, Splink will set these records against each other in a table of pairwise record comparisons: When defining comparisons, we are defining rules that operate on each row of this latter table of pairwise comparisons A Splink model contains a co

Defining and customising how record comparisons are made ¶ A key feature of Splink is the ability to customise how record comparisons are made - that is, how similarity is defined for different data types. For example, the definition of similarity that is appropriate for a date of birth field is different than for a first name field. By tailoring the definitions of similarity, linking models are more effectively able to distinguish between different gradations of similarity, leading to more accurate data linking models. Comparisons and ComparisonLevels ¶ Recall that a Splink model contains a c

Explore this link on the map →

related reading