Defining and customising comparisons - Splink
A key feature of Splink is the ability to customise how record comparisons are made - that is, how similarity is defined for different data types. For example, the definition of similarity that is appropriate for a date of birth field is different than for a first name field. By tailoring the definitions of similarity, linking models are more effectively able to distinguish beteween different gradations of similarity, leading to more accurate data linking models. Note that for performance reasons, Splink requires the user to define n discrete levels (gradations) of similarity. Comparisons are defined on pairwise record comparisons. Suppose for instance your data contains first_name and surname and dob: To compare these records, at the blocking stage, Splink will set these records against each other in a table of pairwise record comparisons: When defining comparisons, we are defining rules that operate on each row of this latter table of pairwise comparisons A Splink model contains a co
Defining and customising how record comparisons are made ¶ A key feature of Splink is the ability to customise how record comparisons are made - that is, how similarity is defined for different data types. For example, the definition of similarity that is appropriate for a date of birth field is different than for a first name field. By tailoring the definitions of similarity, linking models are more effectively able to distinguish between different gradations of similarity, leading to more accurate data linking models. Comparisons and ComparisonLevels ¶ Recall that a Splink model contains a c
Explore this link on the map →related reading
- Choosing string comparators - Splinkmoj-analytical-services.github.io
- 4. Estimating model parameters - Splinkmoj-analytical-services.github.io
- 3. Blocking - Splinkmoj-analytical-services.github.io
- 3. Blocking - Splinkmoj-analytical-services.github.io
- What are Blocking Rules? - Splinkmoj-analytical-services.github.io
- ML vs. Score matchingarxiv.org
- score matching (hyvarinen)jmlr.org
- GitHub - josephg/diamond-types: The world's fastest CRDT. WIP. · GitHubgithub.com
- 聚合、联接或合并数据 - Tableauhelp.tableau.com
- GitHub - ansonyuu/matchmaking: Embedding space of names clustered based on their interests using the sentence-transformers all-MiniLM-L6-v2 model · GitHubgithub.com
- GitHub - ByteByteGoHq/system-design-101: Explain complex systems using visuals and simple terms. Help you prepare for system design interviews. · GitHubgithub.com
- 1. Data prep prerequisites - Splinkmoj-analytical-services.github.io